<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[The Inference Layer]]></title><description><![CDATA[Exploring AI inference, GPU optimization, machine learning systems, MLOps, and the engineering behind building fast and scalable intelligent systems.]]></description><link>https://the-inference-layer.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/682f46bd4329361faee55ca9/799385c0-1d63-4056-9bef-fc7de3e01ed1.png</url><title>The Inference Layer</title><link>https://the-inference-layer.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sat, 10 Oct 2026 19:23:26 GMT</lastBuildDate><atom:link href="https://the-inference-layer.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[GPU Parallelization Isn't Just Moving Code to the GPU]]></title><description><![CDATA[1. The Problem
Before getting into GPU parallelization, let's start with the problem that led us there. We had an ML inference workload with a strict end-to-end time requirement. As the amount of data]]></description><link>https://the-inference-layer.hashnode.dev/gpu-parallelization-isn-t-just-moving-code-to-the-gpu</link><guid isPermaLink="true">https://the-inference-layer.hashnode.dev/gpu-parallelization-isn-t-just-moving-code-to-the-gpu</guid><category><![CDATA[GPU]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[cuda]]></category><category><![CDATA[mlops]]></category><category><![CDATA[#parallel computing]]></category><dc:creator><![CDATA[Apoorv Gupta]]></dc:creator><pubDate>Fri, 02 Oct 2026 21:17:57 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/682f46bd4329361faee55ca9/2fb6c0d9-49c9-4b44-ba7e-d0cd7a6b9eee.jpg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>1. The Problem</h2>
<p>Before getting into GPU parallelization, let's start with the problem that led us there. We had an ML inference workload with a strict end-to-end time requirement. As the amount of data and number of processing steps grew, simply making the model faster wasn't enough. A substantial part of the execution time was being spent in the stages leading up to inference. When we profiled the workload, the preprocessing stage stood out as an obvious place to investigate. It was largely CPU-bound and relied heavily on Pandas and NumPy. That raised a fairly straightforward question:</p>
<p><em>If we already have a GPU available, can we use it to accelerate the preprocessing as well?</em></p>
<p>The first approach was exactly what you might expect: replace CPU-oriented operations with CUDA-enabled alternatives, moving from Pandas → cuDF and NumPy → CuPy wherever possible.</p>
<p>It sounds simple.<br />It wasn't.</p>
<blockquote>
<p>The interesting part wasn't getting the code to run on the GPU. It was figuring out when moving computation to the GPU actually made the pipeline faster.</p>
</blockquote>
<h2>2. The Pipeline</h2>
<p>At a high level, the pipeline looked like this:</p>
<img src="https://cdn.hashnode.com/uploads/covers/682f46bd4329361faee55ca9/e1510666-2748-40c4-b64f-39dfe830ac3f.png" alt="" style="display:block;margin:0 auto" />

<p>The data was time-series data associated with many unique IDs.<br />That made preprocessing more complicated than a simple collection of independent rows.</p>
<p>We had to deal with things such as</p>
<p>missing timestamps<br />irregular observations<br />missing values<br />feature transformations<br />model-specific preprocessing</p>
<p>The interesting part was figuring out which of these operations were actually good candidates for parallel execution.</p>
<hr />
<h1>3. The First Attempt: Pandas + NumPy → GPU</h1>
<p>A lot of the existing preprocessing code was written using Pandas and NumPy.</p>
<p>The obvious first step was to look at CUDA-enabled alternatives.</p>
<p>For example, imagine we have 10 million rows:</p>
<pre><code class="language-python">import pandas as pd
import numpy as np

df = pd.DataFrame({
    "value": np.random.rand(10_000_000),
    "temperature": np.random.rand(10_000_000) * 100,
    "humidity": np.random.rand(10_000_000) * 100
})

# Feature engineering
df["normalized_value"] = (
    df["value"] - df["value"].mean()
) / df["value"].std()

df["heat_index"] = (
    df["temperature"] * 0.7 +
    df["humidity"] * 0.3
)

df["anomaly_score"] = np.abs(
    df["normalized_value"]
)

df["is_anomaly"] = df["anomaly_score"] &gt; 2
</code></pre>
<p>Conceptually, the computation is happening on the CPU:</p>
<pre><code class="language-text">RAM --&gt; NumPy / Pandas --&gt; CPU --&gt; Result
</code></pre>
<p>Now consider moving the dataframe and numerical operations to the GPU using RAPIDS cuDF and CuPy.</p>
<pre><code class="language-python">import cudf
import cupy as cp

df = cudf.DataFrame({
    "value": cp.random.rand(10_000_000),
    "temperature": cp.random.rand(10_000_000) * 100,
    "humidity": cp.random.rand(10_000_000) * 100
})

df["normalized_value"] = (
    df["value"] - df["value"].mean()
) / df["value"].std()

df["heat_index"] = (
    df["temperature"] * 0.7 +
    df["humidity"] * 0.3
)

df["anomaly_score"] = cp.abs(
    df["normalized_value"].values
)

df["is_anomaly"] = df["anomaly_score"] &gt; 2
</code></pre>
<p>Now the data and computation can remain primarily in GPU memory:</p>
<pre><code class="language-text">                  GPU VRAM
              ┌───────────────┐
              │ value         │
              │ temperature   │
              │ humidity      │
              └───────┬───────┘
                      │
                      ▼
                     GPU
                      │
                      ▼
                  Results
</code></pre>
<p>For sufficiently large operations, this can be significantly faster.</p>
<p>But this is where I learned an important distinction.</p>
<hr />
<h1>4. GPU-Friendly Does Not Mean "Everything Is Faster"</h1>
<p>For a large enough workload, operations such as <em><strong>element-wise arithmetic, reductions, aggregations, transformations, matrix operations</strong></em> can benefit substantially from GPU parallelism.</p>
<p>But the size and structure of the operation matter.</p>
<p>If an operation is tiny, the GPU may have very little useful work to do.</p>
<p>You might see something like:</p>
<pre><code class="language-text">CPU operation --&gt; small runtime
GPU operation --&gt; slightly smaller runtime
</code></pre>
<p>When measured individually, the GPU version may be faster.</p>
<p>But when the operation is only a tiny fraction of the complete pipeline, that difference may barely move the end-to-end runtime.</p>
<p>This was one of the first lessons I learned:</p>
<blockquote>
<p><strong>A faster individual operation does not automatically mean a meaningfully faster pipeline.</strong></p>
</blockquote>
<p>For our workload, the interesting gains came from finding <strong>large, repetitive computations</strong> where the GPU had enough work to justify its parallel execution model.</p>
<h3><em><strong>Moving computation to the GPU isn't the same as parallelizing it</strong></em></h3>
<p>There is another trap here.</p>
<p>Just because we move our data and computation to the GPU doesn't mean that the computation itself has suddenly become parallel.</p>
<p>Consider a simple cumulative calculation:</p>
<pre><code class="language-python">result = np.zeros(len(data))

for i in range(1, len(data)):
    result[i] = result[i - 1] + data[i]
</code></pre>
<p>The dependency looks like:</p>
<p>result[0] -&gt; result[1] -&gt; result[2] -&gt; result[3] -&gt; ...</p>
<p>We could move the data to the GPU:</p>
<pre><code class="language-python">data = cp.asarray(data) result = cp.zeros(len(data))
for i in range(1, len(data)):
    result[i] = result[i - 1] + data[i]
</code></pre>
<p>Now the data lives on the GPU, but the fundamental dependency hasn't changed.</p>
<pre><code class="language-plaintext">i = 1 → calculate 
       ↓ 
i = 2 → calculate 
       ↓ 
i = 3 → calculate 
       ↓ 
...
</code></pre>
<p>The GPU is being used, but we haven't given it the kind of independent workload where its parallel execution model can really shine. This distinction became important for our time-series preprocessing.</p>
<h1>5. Where the GPU Really Shines: Matrix Operations</h1>
<p>A good way to see why GPUs are powerful is matrix multiplication.</p>
<p>Consider:</p>
<p><code>C = A @ B</code> which is basically how we multiply matrices in python. Its the same as: <code>cp.matmul(A, B)</code></p>
<p>Suppose:</p>
<pre><code class="language-text">A = 10,000 × 10,000
B = 10,000 × 10,000
</code></pre>
<p>Conceptually, every element of <code>C</code> is computed from a row of <code>A</code> and a column of <code>B</code>:</p>
<pre><code class="language-text">C[i][j] = Σ A[i][k] × B[k][j]
</code></pre>
<p>There are a huge number of these calculations. And importantly, many of them can be computed independently.</p>
<p>Conceptually:</p>
<pre><code class="language-text">                    Matrix C

              C[0,0] C[0,1] C[0,2] ...
                 │      │      │
                 ▼      ▼      ▼
               Work   Work   Work
                 │      │      │
                 └──────┴──────┘
                       GPU
</code></pre>
<p>This is exactly the kind of workload GPUs are designed to attack.</p>
<p>Instead of one CPU core calculating each result sequentially, the GPU can distribute enormous amounts of similar arithmetic across its execution resources.</p>
<p>For operations like this, the GPU is not merely doing the same work "a little faster." It is exploiting a fundamentally different execution model:</p>
<blockquote>
<p><strong>Many threads performing similar operations over large amounts of data.</strong></p>
</blockquote>
<p>That is where GPU parallelization starts becoming really powerful.</p>
<hr />
<h1>6. The Hidden Problem: Data Movement</h1>
<p>There is another factor that has to be considered when moving CPU-based preprocessing to the GPU: <strong>data movement</strong>. If the data repeatedly moves between CPU memory and GPU memory, we can end up with:</p>
<pre><code class="language-text">CPU
 │
 │ transfer
 ▼
GPU
 │
 │ computation
 ▼
CPU
 │
 │ transfer
 ▼
GPU
 │
 │ computation
 ▼
CPU
</code></pre>
<p>That overhead is real.</p>
<p>However, in our workload, this was not the primary bottleneck once we started working with sufficiently large operations.</p>
<p>If moving a large dataset takes an additional second or two, but we save tens of seconds or minutes in the computation itself, that transfer cost is relatively small.</p>
<p>The important distinction is therefore not:</p>
<blockquote>
<p>"Never move data between CPU and GPU."</p>
</blockquote>
<p>It is:</p>
<blockquote>
<p><strong>Avoid unnecessary transfers when the computation is too small to justify them, and make the GPU work large enough that the computational savings dominate the transfer overhead.</strong></p>
</blockquote>
<p>In our case, the bigger problem was not simply that data had to move once.</p>
<p>The problem was repeatedly switching between CPU-oriented and GPU-oriented representations across small pieces of work.</p>
<hr />
<h1>7. The Next Question: Can We Batch More Work?</h1>
<p>Once we understood that the GPU needs enough useful work, another question became important:</p>
<blockquote>
<p><strong>Are we giving the GPU enough data at once?</strong></p>
</blockquote>
<p>Part of the workload was being processed in chunks:</p>
<pre><code class="language-text">Chunk 1 → GPU
Chunk 2 → GPU
Chunk 3 → GPU
Chunk 4 → GPU
...
</code></pre>
<p>Instead, where the computation allowed it, we wanted to move toward:</p>
<pre><code class="language-text">Large batch
    │
    ▼
   GPU
    │
    ├── operation
    ├── operation
    ├── operation
    └── operation
</code></pre>
<p>The intuition was simple:</p>
<blockquote>
<p><strong>More independent work gives the GPU more opportunity to stay busy.</strong></p>
</blockquote>
<p><strong>But this introduced another problem.</strong> We had workloads where a large amount of data needed to be processed at once. In trying to make these operations more suitable for the GPU, we were sometimes converting them into larger matrix-based operations. While this exposed more parallelism, it also increased the memory footprint considerably. At that point, <strong>VRAM became the bottleneck</strong>, and the GPU was spending more effort dealing with memory pressure than doing useful computation.</p>
<p>This became particularly relevant when multiple processes were sharing the same GPU, so it is less important for the single-instance case we're focusing on here. Still, it highlights an important consideration when moving larger workloads to the GPU: <strong>GPU memory is a finite resource.</strong></p>
<p>A practical way to handle this is to understand the memory requirements of your workload, determine the amount of data the GPU can comfortably process, and introduce batching when necessary.</p>
<hr />
<h1>8. The Problem With Time-Series Dependencies</h1>
<p>Our workload involved time-series data, and time-series operations can contain dependencies between observations.</p>
<p>Imagine:</p>
<pre><code class="language-text">A → B → C → D → E
</code></pre>
<p>where:</p>
<pre><code class="language-text">B = f(A)
C = f(B)
D = f(C)
E = f(D)
</code></pre>
<p>We cannot simply treat this as:</p>
<pre><code class="language-text">A   B   C   D   E
↓   ↓   ↓   ↓   ↓
f   f   f   f   f
</code></pre>
<p>because <code>C</code> depends on the result of <code>B</code>, <code>D</code> depends on <code>C</code>, and so on.</p>
<p>This led to an important distinction:</p>
<blockquote>
<p><strong>The amount of data is not the same thing as the amount of parallelism.</strong></p>
</blockquote>
<p>A dataset can contain millions of rows while parts of its computation remain sequential or dependent.</p>
<p>So instead of asking:</p>
<blockquote>
<p>"Can this entire operation run on the GPU?"</p>
</blockquote>
<p>we started asking:</p>
<blockquote>
<p><strong>"Which parts of this operation are actually independent?"</strong></p>
</blockquote>
<p>That was a much better question.</p>
<hr />
<h1>9. Making the Time-Series Data More Regular</h1>
<p>One of the preprocessing steps was creating a consistent time grid for each unique ID. Suppose the original observations looked like:</p>
<pre><code class="language-text">10:00 → 12
10:03 → 15
10:07 →  9
</code></pre>
<p>We could create a complete timeline between the minimum and maximum observed timestamp:</p>
<pre><code class="language-text">10:00 → 12
10:01 → ?
10:02 → ?
10:03 → 15
10:04 → ?
10:05 → ?
10:06 → ?
10:07 →  9
</code></pre>
<p>So, it would look somewhat like this:</p>
<img src="https://cdn.hashnode.com/uploads/covers/682f46bd4329361faee55ca9/60c5cd9c-0dcd-4763-81bf-5c169dddfe39.jpg" alt="" style="display:block;margin:0 auto" />

<p>For the newly introduced or unavailable values, we can use an indicator or another column to represent that this is an imputed value, or synthetic value.</p>
<p>Once the time grid was regularized, the imputation logic for the newly created timestamps could be applied consistently for each unique ID.</p>
<p>At first glance, imputation can look like <strong>extra processing</strong>. If a value is missing, why not just leave it missing?</p>
<p>The problem is that missing data can change how our model interprets the underlying signal. In a time-series workload, gaps aren't just empty cells. They can change the shape, continuity, and patterns that our ML models learn from.</p>
<p>You can see how dramatically this can change our interpretation of the data in the chart below.</p>
<img src="https://cdn.hashnode.com/uploads/covers/682f46bd4329361faee55ca9/01236017-4e9b-499c-b288-e3b5521b44d4.jpg" alt="" style="display:block;margin:0 auto" />

<p>So now, instead of continuously dealing with irregular gaps, we had a much more predictable structure:</p>
<pre><code class="language-plaintext">ID 1 → 300 points
ID 2 → 300 points
ID 3 → 300 points
...
ID N → 300 points
</code></pre>
<p>Here, 300 is just an example based on the chosen time period and sampling interval. Previously, different IDs could have had different numbers of observations, such as <strong>200, 248, 289, or 232</strong>, depending on the timestamps that were actually available. This made the data irregular and meant that the amount of work required for each ID could vary.</p>
<p>This regularity made the workload much more predictable and, importantly, gave us a structure that was much easier to process in parallel on the GPU.</p>
<hr />
<h1>10. What Actually Worked</h1>
<p>The final improvement did not come from one magical CUDA call.</p>
<p>It came from combining several ideas:</p>
<h3>1. Identify operations with enough computational work</h3>
<p>Small operations were not automatically worth moving to the GPU. We focused on workloads where parallel execution could make a meaningful difference.</p>
<h3>2. Reduce unnecessary chunking</h3>
<p>Where the operation allowed it, we provided larger batches of data rather than repeatedly launching small pieces of work.</p>
<h3>3. Make the data more regular</h3>
<p>Creating a consistent time grid reduced irregularity and gave downstream processing a predictable structure.</p>
<h3>4. Separate independent computation from dependent computation</h3>
<p>Not every preprocessing operation was a good GPU candidate.</p>
<p>We had to understand which operations could actually execute independently.</p>
<h3>5. Keep the computation coherent</h3>
<p>Rather than constantly bouncing between CPU and GPU representations, we tried to keep larger portions of the computation together where possible.</p>
<pre><code class="language-text">                    Workload
                       │
              ┌────────┴────────┐
              ▼                 ▼
        Independent          Dependent
        computation          computation
              │                 │
              ▼                 ▼
             GPU          CPU / suitable
                           execution path
</code></pre>
<hr />
<h1>11. The Results</h1>
<p>The workload contained approximately:</p>
<ul>
<li><p><strong>~1 million rows</strong></p>
</li>
<li><p><strong>~3,000 unique IDs</strong></p>
</li>
<li><p><strong>~300 observations per ID</strong></p>
</li>
</ul>
<p>Initially, the preprocessing stage took approximately:</p>
<p><strong>120 seconds</strong></p>
<p>After the optimization for the preprocessing module:</p>
<p><strong>~10 seconds</strong></p>
<p>That represents:</p>
<blockquote>
<p><strong>92% reduction in preprocessing time.</strong></p>
</blockquote>
<p>However, preprocessing was not the only component of the pipeline.</p>
<p>Data fetching and other parts of the pipeline also contributed to the total runtime.</p>
<p>A simplified view was approximately:</p>
<pre><code class="language-text">BEFORE

┌─────────────────────┐
│ Data fetching ~60s  │
├─────────────────────┤
│ Preprocessing ~120s │
└─────────────────────┘
        ~180s
</code></pre>
<p>After the optimization, including improvements to the query/fetching path:</p>
<pre><code class="language-text">AFTER

┌─────────────────────┐
│ Data fetching ~40s  │
├─────────────────────┤
│ GPU processing ~10s │
└─────────────────────┘
         ~50s
</code></pre>
<p>So the preprocessing stage itself improved by roughly <strong>92%</strong>, while the overall pipeline saw an improvement of roughly <strong>60%+</strong>.</p>
<hr />
<h1>12. What I Learned</h1>
<h3>Profile before optimizing</h3>
<p>Don't assume the model is the bottleneck.</p>
<p>Measure where the time actually goes.</p>
<hr />
<h3>GPUs are not automatically faster</h3>
<p>The workload matters.</p>
<p>Large, repetitive, parallel computations are where GPUs can shine.</p>
<p>A tiny operation may technically run faster on the GPU and still make almost no difference to the application.</p>
<hr />
<h3>Parallelism requires independence</h3>
<p>A large dataset does not automatically mean a highly parallel workload.</p>
<p>Dependencies can limit what can be executed simultaneously.</p>
<hr />
<h3>Batch size matters</h3>
<p>Giving a GPU too little work can leave its parallel execution resources underutilized.</p>
<p>But simply making a batch larger isn't enough either. The computation still needs to be parallelizable.</p>
<hr />
<h3>Data movement matters, but context matters more</h3>
<p>CPU-GPU transfers have a cost.</p>
<p>But that cost needs to be considered relative to the amount of computation being accelerated.</p>
<p>For a large workload, spending a second moving data may be insignificant if the computation saves tens of seconds.</p>
<hr />
<h3>Regular data is easier to process</h3>
<p>Turning irregular time-series data into a consistent representation made the downstream computation easier to reason about and optimize.</p>
<hr />
<h1>13. The Question That Came Next</h1>
<p>At this point, I had a much better understanding of <strong>when</strong> GPU parallelization could help.</p>
<p>But I still had a fundamental question.</p>
<p>I was thinking about the GPU as:</p>
<blockquote>
<p><strong>"A device with thousands of cores that can perform lots of operations simultaneously."</strong></p>
</blockquote>
<p>That explanation is useful, but incomplete.</p>
<p>What actually happens when we launch work on the GPU?</p>
<p>How does the GPU take something like:</p>
<pre><code class="language-python">C = A @ B
</code></pre>
<p>and distribute that enormous amount of work?</p>
<p>What exactly is a <strong>thread</strong>? Why are threads grouped into <strong>blocks</strong>? And why doesn't one thread simply map to one GPU core?</p>
<p>And many more terminologies and questions...</p>
<p>Understanding those concepts changed the way I thought about GPU optimization.</p>
<p>That became the next part of the journey:</p>
<blockquote>
<p><strong>From Threads to Warps: Understanding Kernels and how a GPU Actually Executes Work</strong></p>
</blockquote>
]]></content:encoded></item></channel></rss>