Skip to content
BotServBotServ
Memory BandwidthGB/sAI HardwareVRAMRAMApple SiliconPerformance

Memory Bandwidth: The AI Inference Bottleneck

What is memory bandwidth and why it matters most for AI inference? GB/s explained with RAM, VRAM, and Apple Silicon comparison.

S

schutzgeist

12 min read
Memory Bandwidth: The AI Inference Bottleneck

Memory Bandwidth: The Bottleneck in Local AI

What this article covers

  • What memory bandwidth actually is and why it’s measured in GB/s
  • Why processor speed doesn’t determine AI inference performance, memory bandwidth does
  • How much bandwidth different memory types like DDR5, GDDR6, GDDR6X, and HBM deliver
  • Why Apple Silicon performs surprisingly well for local AI
  • Practical tips to get the most out of your hardware’s bandwidth

Introduction: Understanding memory bandwidth

When you start working with local AI, you quickly encounter terms like VRAM, parameters, and quantization. One concept often gets overlooked despite controlling your actual AI speed: memory bandwidth. This article explains what memory bandwidth is, why it’s the critical factor in AI inference, and how to make the right hardware choices.

Most guides focus on how much VRAM you need to load a model at all. That matters, but it’s not enough. A model that barely fits in VRAM can still run painfully slow if bandwidth can’t ferry weights to the processor fast enough. This is where performant local AI really lives.

Why does memory bandwidth matter?

Imagine two graphics cards sitting in front of you. Both have 12 GB of VRAM. Both can load a 7B model. On paper, they look identical. In practice, one card produces 40 tokens per second while the other manages only 13. The difference? Memory bandwidth.

The first card delivers 360 GB/s. The second offers 120 GB/s. Three times the bandwidth means three times the tokens per second. No other factor, not shader unit count or clock speed, has as direct an impact on AI inference performance. Buying a GPU based only on VRAM size is flying blind.

This knowledge helps in several situations: when buying a new graphics card, comparing Mac versus PC, deciding between CPU and GPU inference, and understanding why an apparently weak chip processes large models remarkably quickly.

Memory bandwidth explained

Memory bandwidth describes how much data can move between storage and processor each second. It’s measured in gigabytes per second, abbreviated GB/s. Higher bandwidth means more data travels from storage to processor in the same time window.

The simplest analogy is a highway. More lanes mean more cars can drive simultaneously. The car speed corresponds to memory clock rate, the number of lanes corresponds to the memory interface. Together they determine bandwidth. A wide highway with many lanes moves more cars per second than a narrow country road, even if both have identical speed limits.

In AI inference, every single token the model generates requires reading all weights from memory. A 7B model in 16-bit format contains roughly 14 GB. For each token, 14 GB must move through the bus. At 100 GB/s bandwidth, memory theoretically handles about 7 tokens per second. At 400 GB/s, that’s 28. The math is straightforward, the impact enormous.

Who should read this?

This article is for beginners experimenting with local AI who want to understand why their hardware is fast or slow. If you’re deciding whether to buy a graphics card or use a Mac with Apple Silicon, this knowledge will guide your decision. If you’re already running models and wonder why performance feels sluggish, you’ll find the answer here.

No prior knowledge is required. A basic understanding of AI hardware fundamentals and rough familiarity with the difference between CPU and GPU helps but isn’t mandatory.

Key terms around memory bandwidth

TermMeaning
Memory bandwidthAmount of data per second between memory and processor, measured in GB/s
GB/sGigabytes per second, the standard unit for memory bandwidth
TB/sTerabytes per second, 1000 GB/s, found in high-end GPUs and HBM
Memory interfaceData path width in bits, for example 128-bit, 256-bit, or 384-bit
GDDR6Graphics memory standard common on consumer GPUs, up to roughly 768 GB/s
GDDR6XGDDR6 evolution used by Nvidia, up to roughly 1008 GB/s
HBMHigh Bandwidth Memory, stacked memory in professional GPUs, over 3 TB/s
LPDDRLow-Power DDR memory found in Apple Silicon and laptops
DDR5Current generation of RAM for PCs and servers
Bandwidth per wattEfficiency metric, important for mobile devices and laptops

Why is memory bandwidth so critical for AI?

Most people assume compute power is the bottleneck in AI inference. That’s true for training models, but not for inference, where you run a finished model. The constraint isn’t computation but data movement.

The reason lies in model architecture. A language model consists of billions of parameters, the weights. For each token the model predicts next, all weights must be read from memory. For a 7B model in 16-bit precision, that’s roughly 14 GB per token. The actual calculations are relatively fast, but memory must deliver the weights before any math happens.

This means inference speed depends almost entirely on how fast memory transfers weights to the processor. A GPU with massive compute but low bandwidth spends its time waiting for data. Shader units sit idle while memory catches up. Memory bandwidth is the bottleneck, not TFLOPS.

For deeper insight into VRAM’s role, the article on RAM versus VRAM provides additional background.

Bandwidth comparison

Looking at typical bandwidth numbers for common memory types gives you a feel for scale.

Memory typeHardware exampleBandwidth (approx.)
DDR4 RAMStandard desktop PC25 to 50 GB/s
DDR5 RAMModern PC, server60 to 90 GB/s
GDDR6RTX 3060, RTX 4060360 to 448 GB/s
GDDR6XRTX 40901008 GB/s
HBM3Nvidia H100over 3000 GB/s
LPDDR5Apple M2 (unified)100 GB/s
LPDDR5Apple M2 Pro200 GB/s
LPDDR5Apple M2 Ultra800 GB/s
LPDDR5XApple M3 Max400 GB/s
LPDDR5XApple M4 Max546 GB/s

The table reveals why a graphics card with modest VRAM often processes a model faster than a PC with abundant standard RAM. GDDR6 delivers four to ten times the bandwidth of DDR5. HBM3 shatters all expectations with over 3 TB/s, but comes at a proportional cost.

Apple Silicon occupies an interesting position. The M2 Ultra achieves 800 GB/s, the M4 Max over 500 GB/s. That’s substantially more than consumer cards like the RTX 3060 offer and enough to run large models smoothly. More on Apple Silicon in the section below.

What does this mean in practice?

Let’s walk through a concrete example. You want to run a 7B model in 16-bit precision. The model occupies roughly 14 GB in memory. For each token, those 14 GB must be read. The theoretical token rate is bandwidth divided by model size.

HardwareBandwidthTheoretical Tokens/s
DDR4 PC40 GB/s~2.8
DDR5 PC80 GB/s~5.7
RTX 3060360 GB/s~25.7
RTX 40901008 GB/s~72
Apple M2 Ultra800 GB/s~57
Nvidia H1003000 GB/s~214

In practice, you won’t reach these theoretical values completely because overhead, KV cache, and other factors consume performance. But the order of magnitude is correct. On a DDR4 PC, you’ll wait several seconds for a sentence, while an RTX 4090 generates text smoothly.

With quantization, you can reduce model size. A 7B model in 4-bit takes up only about 3.5 GB. At 360 GB/s bandwidth, the theoretical rate jumps above 100 tokens per second. Quantization isn’t just a tool for saving VRAM, then, but also a way to use bandwidth more efficiently.

Apple Silicon and bandwidth

Apple Silicon occupies a special position. Macs with M, Pro, Max, and Ultra chips use a unified memory approach where CPU and GPU share the same memory. The memory is LPDDR5 or LPDDR5X, and the bandwidths are remarkably high for consumer hardware.

An M2 Ultra delivers 800 GB/s, an M3 Max 400 GB/s, an M4 Max 546 GB/s. This comes from the wide memory bus that Apple has integrated directly into the chip. While an RTX 4060 with a 128-bit interface reaches 272 GB/s, Apple uses significantly wider interfaces to connect to LPDDR chips.

For local AI, this means a Mac Studio with an M2 Ultra and 192 GB of unified memory can load a 70B model and run it at acceptable speeds. No consumer graphics card offers 192 GB of VRAM. The combination of large unified memory and high bandwidth makes Macs an attractive platform for large models if you don’t need the absolute peak performance of an RTX 4090, but rather the ability to run large models at all.

You’ll find more details on this concept in the article about unified memory.

How do I improve my bandwidth?

You can’t double your hardware’s bandwidth through software tuning. It’s a physical property of the memory and memory interface. What you can do is make optimal use of the bandwidth you have and make the right choices when purchasing new hardware.

Use the GPU instead of the CPU. CPU inference on regular RAM is slow because DDR5 only offers 60 to 90 GB/s. Even a cheap graphics card with GDDR6 is three to five times faster. When you use GPU offloading, you load model weights into VRAM and benefit from higher bandwidth.

Choose fast memory. If you rely on CPU inference, use DDR5 instead of DDR4. Dual-channel operation is mandatory, as it doubles bandwidth compared to single-channel. With DDR5-5600 in dual-channel, you’ll reach around 90 GB/s.

Go with Apple Silicon for large models. If you want to run models beyond 24 GB, a Mac with M Pro, M Max, or M Ultra is often the most economical solution. The high bandwidth of unified memory ensures that even large models achieve usable speeds.

Avoid shared-memory APUs for large models. Integrated graphics solutions share RAM with the CPU and offer no dedicated bandwidth. For small models that’s acceptable, but anything from 7B upward becomes uncomfortable. A dedicated graphics card is the better choice here.

Quantize rather than upgrade. If your bandwidth is limited, reduce model size through quantization. A 4-bit model requires a quarter of the bandwidth of a 16-bit model. This is often the easiest path to more tokens per second.

Pay attention to memory interface when buying a GPU. A 128-bit card offers less bandwidth than a 256-bit card with the same memory type. The RTX 4060 with 128-bit reaches 272 GB/s, the RTX 4070 Ti with 192-bit achieves 504 GB/s. The interface is a better indicator than raw VRAM size.

You’ll find more guidance in the buying guide and the article on RAM and VRAM requirements.

Common pitfalls with memory bandwidth

  1. Equating VRAM size with speed. 16 GB of VRAM says nothing about bandwidth. A 16-GB card with a 128-bit interface is significantly slower than a 12-GB card with a 192-bit interface.

  2. Using TFLOPS as a purchasing criterion. Compute power matters for training, but it’s nearly irrelevant for inference. A card with fewer TFLOPS but higher bandwidth is faster for local AI.

  3. Overlooking single-channel RAM. Many PCs ship with only one RAM module. That halves your bandwidth. A second module costs little and doubles performance on CPU inference.

  4. Believing software can replace bandwidth. Optimizations like Flash Attention help, but they can’t increase physical bandwidth. If memory is slow, inference stays slow.

  5. Overestimating APUs as GPU replacements. Integrated graphics use system RAM with DDR bandwidth. That’s not sufficient for serious AI inference.

  6. Comparing Apple Silicon and Nvidia GPUs by VRAM alone. A Mac with 64 GB and 400 GB/s is not directly comparable to an RTX 4090 with 24 GB and 1008 GB/s. Both have different strengths: the Mac fits larger models, the Nvidia card processes smaller ones faster.

  7. Dismissing quantization as quality loss. 4-bit and 8-bit quantization massively reduce bandwidth load with minimal quality loss. Ignoring them wastes performance.

  8. Forgetting the KV cache. With long contexts, the KV cache grows in VRAM and consumes bandwidth. This can noticeably reduce token rate, especially on small graphics cards.

Hardware, cost, and security with memory bandwidth

Bandwidth correlates directly with cost. DDR5 modules are cheap but offer little bandwidth. GDDR6 graphics cards cost several hundred euros, GDDR6X cards like the RTX 4090 over 1500 euros. HBM memory, as in the H100, sits in the five-digit range and is irrelevant for consumers.

Apple Silicon positions itself in between. A Mac Studio with an M2 Max and 64 GB costs significantly less than a workstation PC with comparable VRAM capacity while offering high bandwidth. For users who want to run large models locally without purchasing a professional GPU, that’s an economical alternative.

From a security standpoint, bandwidth changes nothing. Local AI inference is privacy-friendly anyway, since no data is sent to external servers. Higher bandwidth just means inference runs faster. Your data stays on your device, regardless of whether you have 40 GB/s or 1000 GB/s available.

Further Reading and Information on Memory Bandwidth

FAQ: Memory Bandwidth, Common Questions

What is memory bandwidth in simple terms? Memory bandwidth is the amount of data that can move between memory and the processor per second, measured in GB/s. Higher bandwidth means the memory can transport more data in the same time.

Why is memory bandwidth so critical for AI? During inference, all model weights must be read from memory for each token. Bandwidth determines how fast this happens. It’s the primary factor governing token throughput.

Is more VRAM more important than higher bandwidth? No, both matter differently. VRAM determines whether a model can fit into memory at all. Bandwidth determines how fast it runs. Together, they decide overall performance.

What’s the difference between GDDR6 and GDDR6X? GDDR6X is an evolution of GDDR6 with higher data rates. GDDR6 reaches about 768 GB/s, while GDDR6X reaches about 1008 GB/s. GDDR6X is used exclusively by Nvidia.

Why is Apple Silicon so good for local AI? Apple Silicon uses LPDDR5 memory with a wide interface and achieves 100 to 800 GB/s. Combined with large unified memory, Macs can load large models and run them at usable speeds.

Can I increase bandwidth through software? No, bandwidth is a hardware property. Software optimizations can improve the efficiency of existing bandwidth, but they can’t increase bandwidth itself.

What does dual-channel RAM provide? Dual-channel operation uses two memory modules in parallel and doubles bandwidth. With DDR5-5600, bandwidth increases from about 45 GB/s to about 90 GB/s.

How much bandwidth do I need for a 7B model? A 7B model in 16-bit requires 14 GB per token. For 20 tokens per second, you need at least 280 GB/s. With 4-bit quantization, one-third of that is sufficient.

What does the memory interface bit width mean? The memory interface is the width of the data path between memory and GPU. 128-bit, 256-bit, and 384-bit are typical. A wider interface allows more data per clock cycle and increases bandwidth.

Are TFLOPS important for AI inference? TFLOPS play a minor role during inference. Bandwidth is the critical factor. TFLOPS become more important when training models.

What is HBM and why is it so fast? HBM, High Bandwidth Memory, is stacked memory positioned directly next to the processor. Short distances and wide connectivity enable over 3 TB/s. HBM is used in professional GPUs and servers.

Sources and Further Reading

  • Technical datasheets from Nvidia, AMD, and Apple covering respective memory types and bandwidths
  • JEDEC standard specifications for DDR5, GDDR6, and LPDDR5
  • Architecture documentation for Apple Silicon and Nvidia GPUs
  • Real-world tests and benchmarks for token throughput across various hardware configurations
Back to Blog
Share:

Related Posts