The Data Starvation Crisis: Why Bandwidth, Not Compute, Dictates the Speed of Large Language Models
When most people go shopping for a graphics processing unit (GPU), they immediately flip to the performance specifications that look the most impressive on a marketing sheet. They hunt for the clock speeds measured in gigahertz, the massive counts of CUDA or Tensor cores, and the headline computing figures quantified in teraflops or petaflops. The instinct is understandable: we have been trained for decades to believe that raw mathematical computing power is the definitive metric of computational speed.
However, if you are building, training, or running Large Language Models (LLMs), this instinct is fundamentally wrong.
When executing modern AI workloads, your GPU is almost never short on math; it is perpetually starved for data.
Consider a striking real-world asymmetry from the peak of modern AI hardware: a top-tier NVIDIA H100 Tensor Core GPU can execute roughly 989 trillion floating-point operations per second (989 TFLOPS) of FP16 mathematical compute. Yet, during that exact same second, its state-of-the-art High Bandwidth Memory (HBM3) subsystem can only pull 3.35 terabytes (TB) of data out of storage and feed it into the processing cores.
At first glance, 3.35 TB/s sounds incredibly fast—fast enough to transfer dozens of ultra-high-definition movies in the blink of an eye. But when stacked against nearly a quadrillion mathematical operations per second, that memory speed is a crippling bottleneck. This structural disparity is known as the “Memory Wall.” It means that the expensive, blazing-fast processing cores spend the vast majority of their operational life sitting completely idle, twiddling their digital thumbs while waiting for data to arrive from memory.
To understand why buying a GPU for AI requires an obsessive focus on memory bandwidth rather than computing flops, we must walk directly inside the silicon architecture, dissect how an LLM processes information, and uncover why data movement—not mathematical calculation—rules the modern AI universe.
1. How a GPU Works: Massively Parallel Execution
To understand why data delivery falls so far behind, we must first look at the unique internal geography of a GPU and contrast it with a traditional Central Processing Unit (CPU).
┌──────────────────────────────────────┐ ┌──────────────────────────────────────┐
│ TRADITIONAL CPU ARCHITECTURE │ │ GPU ARCHITECTURE │
├──────────────────────────────────────┤ ├──────────────────────────────────────┤
│ ┌──────────┐ ┌──────────┐ ┌────────┐ │ │ ┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐┌───┐ │
│ │ Core 1 │ │ Core 2 │ │ Huge │ │ │ │ C ││ C ││ C ││ C ││ C ││ C ││ C ││ C │ │
│ │ (Complex)│ │ (Complex)│ │ Cache │ │ │ ├───┤├───┤├───┤├───┤├───┤├───┤├───┤├───┤ │
│ └──────────┘ └──────────┘ └────────┘ │ │ │ C ││ C ││ C ││ C ││ C ││ C ││ C ││ C │ │
│ ┌──────────────────────────────────┐ │ │ ├───┤├───┤├───┤├───┤├───┤├───┤├───┤├───┤ │
│ │ Control Logic & ALUs │ │ │ │ C ││ C ││ C ││ C ││ C ││ C ││ C ││ C │ │
│ └──────────────────────────────────┘ │ │ └───┴───┴───┴───┴───┴───┴───┴───┴───┘ │
└──────────────────────────────────────┘ │ Thousands of Tiny, Simple Cores (C) │
│ Connected to a Massive Data Highway │
└──────────────────────────────────────┘
A CPU is designed like a small team of elite, hyper-versatile engineers. It contains a relatively small number of processing cores (typically 8 to 64) optimized for sequential processing. These cores feature immensely complex control logic, massive hierarchies of local cache memory, and the ability to switch tasks instantly. A CPU excels at executing complex linear code where item A must be calculated before item B can begin.
A GPU, by contrast, is designed like an army of millions of highly disciplined factory workers. Instead of a few complex cores, a GPU contains thousands of small, mathematically simple processing units running in parallel.
The Components of a GPU Core
Inside an enterprise-grade GPU, these thousands of cores are grouped into structures NVIDIA calls Streaming Multiprocessors (SMs). Inside each SM, you will find:
- ALUs (Arithmetic Logic Units): Dedicated lanes for standard integer and floating-point math (FP32, FP16).
- Tensor Cores: specialized, hardwired matrix-multiplication engines designed specifically to multiply rows and columns of numbers simultaneously—the core mathematical operation of deep learning.
- Register Files and SRAM: Tiny, ultra-fast, local memory pockets immediately adjacent to the execution units.
Because a GPU features thousands of these lanes executing instructions simultaneously via a paradigm known as SIMT (Single Instruction, Multiple Threads), it can process vast oceans of numbers at the same time. If you need to add a billion pairs of numbers together, a CPU must cycle through them sequentially or in tiny batches. A GPU simply assigns each pair to one of its thousands of threads and processes massive chunks of the problem in a single, unified clock cycle.
This architecture works beautifully for computer graphics, where every pixel on your screen can be calculated independently of its neighbor. It also happens to be the exact mathematical foundation required for neural networks.
2. The Mechanics of LLMs: Matrix Multiplication and Token Generation
Large Language Models like GPT-4, Llama 3, or Mistral may feel like sentient entities, but underneath the natural language interface, they are simply mammoth probability engines driven by matrix calculus.
An LLM is structured as an architectural stack of layers built on the Transformer architecture. Each layer contains billions of parameters (weights). These weights are numerical values that represent the model’s learned knowledge. When you type a prompt into an LLM, your words are broken down into numerical fragments called tokens.
The fundamental cycle of an LLM involves taking the numerical vector of your token, multiplying it against a giant matrix of weight parameters at layer one, passing that result to the next layer’s matrix, and repeating this sequence through dozens of layers until the model predicts the most statistically probable next token.
Crucially, LLM operations are split into two distinct phases, each interacting with the GPU’s internal hardware in entirely different ways:
┌─────────────────────────────────────────────────────────────────────────────────┐
│ THE TWO PHASES OF LLM EXECUTION │
├────────────────────────────────────────┬────────────────────────────────────────┤
│ 1. PREFILL PHASE │ 2. DECODE (AUTOREGRESSIVE) │
├────────────────────────────────────────┼────────────────────────────────────────┤
│ • Processes the entire input prompt │ • Generates text one token at a time │
│ • Large batch of data processed once │ • Sequential: Token N requires Token │
│ • Highly Compute-Bound │ N-1 to finish computing first │
│ • Satures Tensor Cores efficiently │ • Severely Memory-Bandwidth Bound │
└────────────────────────────────────────┴────────────────────────────────────────┘
Phase 1: The Prefill Phase
When you submit a 2,000-word prompt, the GPU ingests the entire block of text at once. It parallelizes the math across all input tokens simultaneously. Because the GPU is computing a massive batch of data simultaneously, it can load a weight parameter from its memory and reuse that same parameter across thousands of tokens sitting in its local registers. This makes the prefill phase highly compute-bound. It easily saturates the Tensor Cores and utilizes the GPU’s advertised TFLOPS to their maximum potential.
Phase 2: The Decode Phase (The Autoregressive Bottleneck)
Once the prompt is understood, the model begins generating its response. This is where the disaster happens. LLMs are autoregressive, meaning they generate text one single token at a time. To generate token number 50, the model must look at tokens 1 through 49. To generate token 51, it must recalculate everything including the newly minted token 50.
During this decode phase, the batch size for a single user drops down to one single token.
To calculate that one next token, the GPU must sweep through every single weight parameter across its entire network. If you are running a 70-billion parameter model (70B), the GPU must load all 70 billion numbers from its main memory chips, move them across the internal buses, feed them into the Tensor Cores, perform a tiny mathematical calculation against the single current token, generate the new token, and then throw those parameters away.
A millisecond later, to generate the very next token, the GPU must pull those exact same 70 billion weights out of memory all over again.
3. Arithmetic Intensity and the Roofline Model
To quantify precisely why this sequential parameter sweep breaks GPU performance, computer scientists rely on a concept called Arithmetic Intensity.
$$\text{Arithmetic Intensity} = \frac{\text{Total Operational Floating-Point Math (Flops)}}{\text{Total Memory Access Required (Bytes)}}$$
Arithmetic intensity measures how much mathematical work you perform on a piece of data every time you pull it from memory.
If a program has high arithmetic intensity, it fetches a number once from memory and performs thousands of calculations on it. The processing cores are kept incredibly busy, meaning the application is Compute-Bound.
If a program has low arithmetic intensity, it fetches a massive block of data from memory, performs a single trivial addition or multiplication on it, and drops it. The processing cores spend most of their time completely starved, waiting on the memory pipeline. The application is Memory-Bound.
This dynamic is visualized using the Roofline Model, a chart that graphs an application’s attainable performance against its arithmetic intensity.
Attainable Performance (TFLOPS)
^
│ COMPUTE-BOUND REGION (Tensor Cores saturated)
│ ┌───────────────────────────────────────────────
│ / ◄── The "Roofline" (Max Hardware TFLOPS)
│ /
│ / MEMORY-BOUND REGION (Cores waiting for data)
│/
│ ▲
│ └─ Inflection Point (Where memory speed locks performance)
└─────────────────────────────────────────────────────────> Arithmetic Intensity (Flops/Byte)
The diagonal slope represents the limits imposed by memory bandwidth. The flat horizontal top represents the absolute mathematical limits of the processor’s silicon.
Let’s plug the numbers from an NVIDIA H100 (SXM5) into this equation to see exactly where the inflection point sits:
- Peak FP16 Compute: ~989 TFLOPS ($989 \times 10^{12}$ operations/second)
- Memory Bandwidth: ~3.35 TB/s ($3.35 \times 10^{12}$ bytes/second)
To find the minimum arithmetic intensity required to fully saturate the H100’s computing cores without being limited by memory, we divide the peak compute by the memory bandwidth:
$$\frac{989 \times 10^{12} \text{ Flops}}{3.35 \times 10^{12} \text{ Bytes}} \approx 295 \text{ Flops per Byte}$$
This means that for every single byte of data the H100 pulls from its VRAM, it must perform at least 295 mathematical operations on that byte to keep its processing cores fully utilized.
Now, let us look at the arithmetic intensity of the LLM decode phase. When generating text at a batch size of 1, the arithmetic intensity of matrix-vector multiplication is roughly 2 operations per parameter byte.
Compare those two numbers: The hardware requires 295 operations per byte to run at full speed, but the LLM decode phase only provides 2 operations per byte.
The LLM workload misses the compute-bound threshold by a massive factor of nearly 150. Consequently, during LLM text generation, an H100 GPU is operating at roughly less than 1% to 2% of its theoretical computing capacity. The remaining 98% of its raw computing power is completely wasted, idling in a state of deep data starvation. The speed at which tokens can flow out of the model is restricted entirely by the memory bandwidth clock.
4. The Math of Memory Bottlenecks: A Step-by-Step Breakdown
Let us do the explicit math for running an open-source model locally or on a server to see how this bandwidth bottleneck dictates your operational reality.
Imagine you want to run a quantized version of the Llama 3 70B model. You use a 16-bit precision format (FP16), meaning every single weight parameter requires 2 bytes of memory.
$$\text{Model Size in Bytes} = 70,000,000,000 \text{ parameters} \times 2 \text{ bytes} = 140 \text{ Gigabytes (GB)}$$
To generate a single token, your GPU must read 140 GB of data from its memory banks and pass it through its SMs.
Scenario A: Running on a Consumer GPU (NVIDIA RTX 4090)
Let’s look at a high-end consumer graphics card, the NVIDIA GeForce RTX 4090.
- Peak Compute: 83 TFLOPS (FP16)
- Memory Bandwidth: 1.01 TB/s (1,018 GB/s)
- VRAM Capacity: 24 GB GDDR6X
(Note: In reality, a 140 GB model won’t fit on a single 24 GB card without extreme quantization down to 2-bits, but for the sake of isolation, let’s assume we have an ideal bandwidth pipeline of 1.01 TB/s).
If your memory system can transfer data at a maximum speed of 1,018 GB per second, how long does it take to stream the required 140 GB of model weights through the chip to generate one token?
$$\text{Time per Token} = \frac{140 \text{ GB}}{1,018 \text{ GB/s}} \approx 0.1375 \text{ seconds}$$
$$\text{Max Tokens per Second} = \frac{1}{0.1375 \text{ seconds}} \approx 7.27 \text{ tokens/second}$$
Even if the RTX 4090’s compute cores were upgraded to be ten times faster, the speed would remain completely locked at ~7.27 tokens per second because the data highway simply cannot move the weights any faster.
Scenario B: Running on an Enterprise GPU (NVIDIA H100)
Now let’s look at the dedicated data-center chip, the NVIDIA H100 (SXM5), which utilizes ultra-dense High Bandwidth Memory (HBM3) stacks stacked directly on top of the silicon interposer.
- Peak Compute: 989 TFLOPS
- Memory Bandwidth: 3.35 TB/s (3,350 GB/s)
- VRAM Capacity: 80 GB HBM3
Assuming we scale the model across a cluster of cards to accommodate the 140 GB memory footprint, let’s look at what the 3.35 TB/s bandwidth speed unlocks:
$$\text{Time per Token} = \frac{140 \text{ GB}}{3,350 \text{ GB/s}} \approx 0.0418 \text{ seconds}$$
$$\text{Max Tokens per Second} = \frac{1}{0.0418 \text{ seconds}} \approx 23.92 \text{ tokens/second}$$
By moving from standard consumer GDDR6X memory to enterprise HBM3 memory, the text generation speed jumps from 7 to nearly 24 tokens per second. This performance gain has absolutely nothing to do with the H100 having higher TFLOPS; it is entirely driven by the fact that the H100 can pull data out of memory 3.3 times faster.
5. Memory Types Compared: GDDR vs. HBM
Because bandwidth is the primary factor limiting performance, the type of memory architecture integrated into a GPU determines its class and capabilities for AI workloads. Modern hardware is broadly divided into two primary camps:
┌─────────────────────────────────┐ ┌─────────────────────────────────┐
│ GDDR6X ARCHITECTURE │ │ HBM3 ARCHITECTURE │
├─────────────────────────────────┤ ├─────────────────────────────────┤
│ │ │ ┌─────────┐ ┌─────────┐ │
│ ┌───────────┐ ┌───────────┐ │ │ │ HBM3 │ │ HBM3 │ │
│ │ GDDR6X │ │ GDDR6X │ │ │ │ Stack │ │ Stack │ │
│ └─────┬─────┘ └─────┬─────┘ │ │ └────┬────┘ └────┬────┘ │
│ │ │ │ │ │ │ │
│ ┌─────┴───────────────┴─────┐ │ │ ┌──────┴─────────────┴──────┐ │
│ │ GPU Silicon │ │ │ │ GPU Silicon │ │
│ └───────────────────────────┘ │ │ ├───────────────────────────┤ │
│ │ │ │ Silicon Interposer │ │
│ Long, narrow motherboard traces │ │ └───────────────────────────┘ │
│ Wide, thin layout (384-bit bus) │ │ Ultra-wide 3D-stacked (4096-bit)│
└─────────────────────────────────┘ └─────────────────────────────────┘
| Specification Metric | GDDR6X (e.g., RTX 4090) | HBM3 / HBM3e (e.g., H100 / B200) |
|---|---|---|
| Physical Architecture | Discrete chips soldered around the main GPU socket. | 3D-stacked vertical silicon integrated onto the GPU package. |
| Bus Width | Narrow (Typically 256-bit to 384-bit) | Ultra-Wide (Typically 4,096-bit to 8,192-bit) |
| Clock Speeds | Exceptionally high to compensate for narrow bus widths. | Moderate, focused on massive parallel lane distribution. |
| Typical Bandwidth | 500 GB/s to 1.1 TB/s | 2.0 TB/s to over 8.0 TB/s |
| Manufacturing Cost | Low to Moderate (Standard PCB assembly) | Extremely High (Requires advanced silicon packaging) |
GDDR (Graphics Double Data Rate)
GDDR memory is the standard architecture found in consumer gaming GPUs and mid-tier workstations. The memory units sit as separate chips surrounding the main GPU processor, communicating over traces cut into the graphics card’s circuit board.
Because space on a circuit board is physically limited, the bus width—the number of data lanes connecting the memory to the processor—is highly restricted, typically maxing out at 384 bits. To push data through this narrow pipe fast enough, manufacturers must run the memory at extremely high, power-hungry clock frequencies.
HBM (High Bandwidth Memory)
HBM completely reimagines the physical architecture. Instead of placing chips around the circuit board, manufacturers stack memory dies vertically on top of one another like floors in a skyscraper. This entire stack is then placed immediately next to the main GPU computing die on a microscopic layer of silicon called an interposer.
This proximity allows for an immensely wide memory highway. Instead of a 384-bit bus, an HBM3 system uses an ultra-wide 4096-bit bus or wider. Because the highway has thousands of extra lanes, it can move massive volumes of data at lower clock speeds, consuming far less power while achieving bandwidth metrics that leave GDDR memory far behind.
6. How Software Engineering Fights the Memory Wall
Because hardware engineers cannot simply manufacture infinite memory bandwidth due to physical and thermodynamic constraints, the AI industry has turned to clever software optimization strategies to ease the burden on the memory pipeline.
Strategy 1: Model Quantization
If memory bandwidth limits your speed because weights are too large, the most direct solution is to make the weights smaller. Quantization shrinks the numerical precision of an LLM’s parameter weights.
[FP16 Precision] --> 16 bits per weight --> [ 1.1403, -0.4589, 0.8912 ] (100% Size)
│
▼ Quantization Process
│
[INT4 Precision] --> 4 bits per weight --> [ 1, 0, 1 ] ( 25% Size)
By compressing a model from standard FP16 (16 bits per parameter) down to INT4 (4 bits per parameter), you cut the physical storage and bandwidth footprint by 75%.
Suddenly, a 70B model shrinks from a 140 GB bandwidth requirement down to just 35 GB. This compression allows large models to fit onto much cheaper hardware configurations and slashes the amount of data that must traverse the memory bus for every token generated, boosting text generation speeds significantly.
Strategy 2: KV Caching (Key-Value Cache)
During the autoregressive decode phase, the model must evaluate all previous tokens to generate the next one. Computing attention values for every single historical token at every step requires repetitive, resource-intensive memory calls.
To bypass this overhead, modern LLM engines deploy a KV Cache. This technique stores the calculated mathematical keys and values of past tokens directly in a dedicated block of VRAM. Instead of recalculating the entire conversational history from scratch for every single word, the system simply appends the new token’s data to the existing cache, preserving precious memory bandwidth cycles.
Strategy 3: PagedAttention (vLLM)
While the KV Cache saves computing cycles, it introduces a severe memory fragmentation problem. The history of a conversation grows dynamically, making its exact size unpredictable. Traditional systems were forced to pre-allocate large, continuous blocks of VRAM to safely store these keys and values. This approach regularly resulted in up to 60% of the allocated memory sitting empty and completely wasted.
Developed by academic researchers and commercialized via systems like vLLM, PagedAttention mimics the virtual memory paging concepts used in traditional computer operating systems. It slices the KV cache up into small, non-contiguous blocks scattered throughout memory. By linking these fragments via a dynamic lookup table, it eliminates VRAM waste entirely. This allows operators to scale up their batch sizes—processing multiple users simultaneously on the same hardware—without running out of memory.
7. Buying Guide: How to Evaluate a GPU for LLM Workloads
If you are evaluating hardware to run or fine-tune LLMs, you must learn to look past basic marketing specifications. Use this step-by-step framework to identify the right hardware for your needs:
Step 1: Prioritize Total VRAM Capacity Above All Else
Before looking at speed, you must ensure the model will physically fit onto your hardware. If your GPU runs out of VRAM capacity, the system will either crash or spill the model over into your computer’s system RAM via the PCIe bus. This drop-off destroys performance, cutting generation speeds down to an unusable crawl.
$$\text{Required VRAM} \approx (\text{Parameter Count in Billions} \times \frac{\text{Quantization Bits}}{8}) \times 1.25$$
(The extra 25% provides a necessary buffer for context window data, KV caching, and basic operating system overhead).
Step 2: Calculate the Memory Bandwidth
Once you have identified cards that can comfortably hold your target model, the Memory Bandwidth spec is your primary indicator of performance.
HIGHER BANDWIDTH = HIGHER TOKENS PER SECOND (Text Generation Speed)
If you are choosing between two cards with similar computing configurations, always pick the card with the wider bus width and higher memory bandwidth.
Step 3: Understand Corporate Market Segmentation
NVIDIA manages its product stack carefully to protect its lucrative data center business. Consumer cards like the RTX 4090 or RTX 5090 feature incredibly fast core processing speeds, but their memory architectures are intentionally restricted to narrow buses with standard GDDR memory.
If your goal is scaling up heavy enterprise workloads with multiple concurrent users, consumer components will quickly fall flat. For professional, highly parallel production scale, investing in specialized architecture that utilizes HBM memory—such as NVIDIA H100, H200, B200, or AMD’s MI300X series—is essential to handle the data demands of enterprise AI.
Final Thoughts: The Data Movement Paradigm
The AI revolution has exposed an underlying truth about modern computer engineering: calculating a number is effectively free; moving a number across a slice of silicon is incredibly expensive.
The legendary processing speeds promised by modern hardware platforms are fundamentally an illusion when applied to sequential text generation. The thousands of Tensor Cores built into an enterprise accelerator are marvels of modern engineering, but without data to process, they are simply high-performance engines trapped in bumper-to-bumper traffic.
As you look to purchase, configure, or optimize infrastructure for the era of large language models, change how you view computing architecture. Look past the theoretical compute metrics, step away from the allure of teraflops, and look closely at the data highway. In the world of AI, memory bandwidth is the definitive speed limit.
Frequently asked questions (FAQs) about how GPU architecture impacts Large Language Models, the role of memory bandwidth, and the realities of the “Memory Wall.”
Core Architecture & How GPUs Work
Why are GPUs used for LLMs instead of CPUs?
Large Language Models are powered by massive layers of matrix multiplication. A CPU is built with a few highly complex cores designed to handle sequential tasks one after another. A GPU contains thousands of simpler, parallel processing cores (including specialized Tensor Cores). This architecture allows a GPU to perform thousands of mathematical calculations simultaneously, making it perfectly suited for the massive parallel math required by neural networks.
What is the difference between the “Prefill” phase and the “Decode” phase in an LLM?
- Prefill Phase: The GPU ingest your entire prompt at once. Because it processes a massive batch of tokens simultaneously, it can load model weights once and reuse them across many tokens. This phase is highly compute-bound and runs near the GPU’s maximum advertised speed (TFLOPS).
- Decode Phase: The GPU generates the output text one token at a time. To generate a single new word, it must sweep through every single weight parameter in the entire model. Because it loads billions of parameters just to perform a trivial amount of math for one token, this phase is strictly memory-bandwidth bound.
The Memory Bandwidth Bottleneck
What is “Memory Bandwidth” and why does it limit LLM speed?
Memory bandwidth is the speed at which data can be transferred from the GPU’s storage banks (VRAM) into its processing cores. During text generation, the GPU cores are so fast that they constantly finish their math instantly and have to sit idle, waiting for the next batch of model weights to travel across the memory highway. Your text generation speed (tokens per second) is almost entirely dictated by how fast your memory can feed data to the cores, not how fast the cores can calculate.
What is the “Memory Wall”?
The Memory Wall is a term used in computer science to describe the massive engineering gap between processing speed and memory speed. Over the last few decades, raw computing power (TFLOPS) has grown exponentially faster than the speed at which data can be pulled out of memory. This leaves high-performance processors severely starved for data during data-heavy workloads like AI text generation.
Why does an NVIDIA H100 run at less than 2% capacity during LLM text generation?
An H100 requires roughly 295 mathematical operations per byte of data to stay fully saturated. However, generating text one token at a time only requires about 2 operations per byte. Because the arithmetic intensity of the decode phase is so incredibly low, the processing cores spend 98% of their time waiting for data to arrive from the HBM3 memory stacks.
Hardware & Buying Advice
What is the difference between GDDR and HBM memory?
- GDDR (e.g., RTX 4090): Standard consumer graphics memory. The chips sit around the processor on a traditional circuit board, communicating over a narrow data highway (typically a 384-bit bus). It is affordable but offers lower bandwidth (around 1 TB/s).
- HBM (e.g., NVIDIA H100/H200): Enterprise-grade High Bandwidth Memory. Memory dies are stacked vertically directly next to the processor on a microscopic silicon interposer. This allows for an ultra-wide data highway (4,096-bit bus or wider), hitting massive bandwidth speeds (3.35 TB/s to over 8 TB/s).
When buying a GPU for AI, what specs should I look at first?
- VRAM Capacity: If a model cannot physically fit entirely inside your GPU’s VRAM, performance will drop to near-zero as data spills over into your computer’s system RAM.
- Memory Bandwidth: Once you know a model fits, memory bandwidth is the single most important spec determining text generation speed. Higher bandwidth equals more tokens per second.
- Compute (TFLOPS): Treat this as a secondary metric for LLM generation. It only becomes a primary factor if you are doing heavy training workloads or processing massive multi-thousand-token prompt batches.
Software & Optimizations
What is Model Quantization and how does it help?
Quantization is a compression method that shrinks the numerical precision of an LLM’s weights (for example, converting 16-bit numbers down to 4-bit numbers). This reduces the model’s total file footprint by up to 75%. Because the model takes up far fewer bytes, much less data needs to travel across the memory bus for every token generated, significantly speeding up performance on lower-bandwidth hardware.
What do KV Caching and PagedAttention do?
- KV Caching: Saves the mathematical history of past tokens in a conversation inside the VRAM so the GPU doesn’t have to recalculate the entire chat history from scratch every time it generates a single new word.
- PagedAttention: Solves the problem of VRAM fragmentation caused by KV Caching. It breaks up the saved token history into small, non-contiguous blocks (like pages in an operating system), eliminating wasted VRAM space and allowing a single GPU to handle multiple users simultaneously without running out of memory.
#GPUArchitecture, #MemoryBandwidth, #LLMPerformance, #DeepLearningHardware, #AIEngineering, #NVIDIAH100, #HardwareBottleneck, #DataStarvation, #MachineLearning, #TechDeepDive
