GPU-Offloading: Running Models Between RAM and VRAM
What This Article Covers
- What GPU-offloading is and why you need it when VRAM runs out
- How a model gets split layer by layer between GPU and CPU
- Real performance differences between full GPU, partial offloading, and CPU-only
- How to configure GPU-offloading in Ollama and llama.cpp
- Alternative approaches and common pitfalls
Introduction: Understanding GPU-Offloading
Anyone running local AI eventually hits the same wall: your model is bigger than your graphics card’s VRAM. At that point, you have options. You can pick a smaller model, increase quantization, or use GPU-offloading. The last approach means running part of the model on your GPU and part in system RAM on your CPU. It’s not ideal, but often the best compromise between speed and model size.
This article covers GPU-offloading fundamentals. You’ll learn how it works, when it makes sense, how to set it up, and what gotchas to watch for. If you’re already familiar with RAM vs. VRAM and CPU vs. GPU, that helps, but it’s not required. We’ll start from the beginning.
Why Do You Need GPU-Offloading?
Imagine you have a graphics card with 6 GB of VRAM, say an older Nvidia RTX 2060. You want to run a 13B model, that is, one with 13 billion parameters. Quantized to 4-bit, this model needs about 8 GB of memory. Your VRAM only has 6 GB. The model won’t fit.
Without GPU-offloading, you face two outcomes. Either the program crashes saying there’s not enough VRAM, or it loads the model entirely into RAM and runs it on the CPU. On CPU, the model is slow. The CPU has fewer parallel cores than the GPU, and RAM offers lower memory bandwidth than VRAM. In practice, you’d get maybe 5 to 10 tokens per second, if that.
With GPU-offloading, something different happens. The program loads as much of the model as fits into VRAM and keeps the rest in RAM. In our example, about 75 percent of the model fits in the 6 GB VRAM, and the remaining 25 percent stays in RAM. The GPU processes its portion quickly, the CPU its portion more slowly. The result is far faster than CPU-only, since most of the computation happens on the GPU. Instead of 5 tokens per second, you might hit 15 to 20 tokens per second. Not blazing fast, but usable.
That’s where GPU-offloading delivers value. It lets you run models that are technically too large for your graphics card without immediately buying new hardware.
GPU-Offloading in a Nutshell
GPU-offloading means an AI model isn’t fully loaded into VRAM, only part of it is. The rest stays in RAM and gets computed by the CPU. The GPU handles the layers that fit in VRAM, the CPU handles the rest. Data flows back and forth between the two during inference.
Think of it like this: imagine a worker loading boxes onto a conveyor belt. The fast conveyor belt is your GPU, the slow one is your CPU. The fast belt holds only 6 boxes, but there are 8 total. The worker puts 6 on the fast belt and 2 on the slow one. Most boxes arrive quickly, just 2 take longer. Overall, that’s way faster than if all 8 had to go on the slow belt.
That’s roughly how GPU-offloading works. The more of your model that fits on the fast belt (into VRAM), the faster inference runs. The more that has to stay in RAM, the more it bottlenecks overall speed.
Who Should Use GPU-Offloading?
GPU-offloading is for anyone wanting to run a model that doesn’t fit entirely in their graphics card’s VRAM. That mainly includes users with 6 to 8 GB of VRAM wanting to run 13B+ models. Even with a larger 12 or 16 GB card, you’ll hit the wall with 30B or 70B models and need offloading.
If you’re just starting with local AI and only using 7B models, you probably don’t need GPU-offloading yet. Those fit in 8 GB of VRAM. But once you try larger models, the topic becomes relevant. It’s worth understanding the concept before you’re stuck with a slow model and no idea why.
Key Terms Around GPU-Offloading
| Term | Meaning |
|---|---|
| GPU-offloading | Parts of the model are moved to GPU while the rest stays in system RAM on CPU |
| Partial offloading | Only part of the model sits in VRAM, the rest in RAM, computation is split |
| Full offloading | The entire model fits in VRAM, the CPU doesn’t participate in inference |
| Layer | A single level of the neural network; models consist of many layers in sequence |
| VRAM | Memory on the graphics card, exclusive to the GPU |
| RAM | System memory, used by the CPU |
| Memory bandwidth | Amount of data that can flow between memory and processor per second |
| Inference speed | Speed at which the model generates responses, usually measured in tokens per second |
| KV-cache | Cache for already-computed attention values, grows with text length |
| Quantization | Compression of the model by reducing numerical precision, saves VRAM and RAM |
How Does GPU-Offloading Work?
To understand GPU-offloading, you need to know that an AI model isn’t a single block but made up of many layers. A 13B model typically has around 40 layers. Each layer holds part of the model’s weights and is executed in sequence as the model generates a response. Each layer’s output becomes the next layer’s input.
With GPU-offloading, the model gets split at a certain point. The first layers that fit in VRAM load onto the GPU. The remaining layers stay in RAM and run on the CPU. If you load, say, 30 of 40 layers onto the GPU, those 30 run fast on the GPU while 10 run slower on the CPU.
During inference, here’s what happens: your input text runs through the GPU layers first. Then the intermediate result gets copied from VRAM into RAM so the CPU can compute the next layers. The CPU’s output gets copied back into VRAM if more GPU layers follow, or returned as output directly.
This data transfer between VRAM and RAM costs time. The more copying back and forth, the greater the speed penalty. That’s why it’s better to keep as many layers together on the GPU instead of mixing them. In practice, layers load in order anyway, so the split stays contiguous.
The KV-cache, the storage for computed attention values, matters too. It grows with text length and needs extra VRAM. When you calculate how many layers fit in VRAM, you need to account for the KV-cache. Otherwise, the model starts fast but gets slower as text gets longer, because the cache fills up your VRAM.
Speed: Full, Partial, or None
How fast a model runs depends entirely on where it’s computed. Here’s a comparison with real numbers for a 13B model with 4-bit quantization on typical consumer hardware.
| Setup | Where the model runs | Tokens per second | Assessment |
|---|---|---|---|
| Full Offloading | Entirely in VRAM (12 GB or more) | 40 to 60 | Very fast, ideal |
| Partial Offloading, 75% GPU | 30 of 40 layers in VRAM (6 GB) | 15 to 25 | Usable, good for conversation |
| Partial Offloading, 50% GPU | 20 of 40 layers in VRAM (4 GB) | 8 to 15 | Slow, patience required |
| CPU-only | Entirely in RAM | 3 to 8 | Very slow, only for short text |
These numbers are estimates and vary depending on hardware, model, and quantization. The pattern is clear though: the more of the model that sits in VRAM, the faster it runs. The jump from CPU-only to 50% GPU is more dramatic than the jump from 75% to full offloading. Even partial offloading delivers a noticeable improvement.
One key point: speed with partial offloading isn’t linear. Loading 75% of the model onto the GPU doesn’t mean you’ll achieve 75% of full offloading speed. The bottleneck is data transfer between RAM and VRAM, plus CPU computation of the remaining layers. That’s why speed drops disproportionately as more layers stay on the CPU.
Configuration in Ollama and llama.cpp
If you use Ollama, GPU offloading happens automatically. At startup, Ollama detects available VRAM and loads as many layers as possible onto the GPU. You don’t need to configure anything. Still, it helps to know how you can intervene.
In Ollama, control the number of GPU layers with the num_gpu parameter. When starting a model, pass the parameter directly:
ollama run llama3:13b --num-gpu 30
This loads 30 layers onto the GPU and the rest into RAM. If you omit it, Ollama decides on its own. Use the parameter to experiment and find what works best on your hardware.
In llama.cpp, the engine behind Ollama, the parameter is -ngl, short for number of GPU layers. A typical invocation looks like this:
./llama-cli -m modell.gguf -ngl 30 -p "Dein Prompt"
Here too, -ngl 30 loads 30 layers onto the GPU. Use -ngl 0 to run entirely on CPU, or -ngl 99 or any number larger than the layer count to load everything onto the GPU if space permits.
To find out how many layers your model has, watch the output when it loads. Ollama and llama.cpp show how many layers go to the GPU and how many stay in RAM. That tells you whether offloading is working as expected.
Example: 13B Model with 8GB VRAM
Let’s walk through a concrete scenario. You have a graphics card with 8 GB VRAM and want to run a 13B model with 4-bit quantization. The model itself needs about 8 GB for the weights alone. Add the KV-Cache, which needs another 0.5 to 2 GB depending on context length.
Step 1: You start the model with Ollama and no parameters. Ollama detects 8 GB VRAM and tries to load the model fully. The weights need 8 GB, the KV-Cache needs extra space. VRAM isn’t enough for both.
Step 2: Ollama automatically reduces the number of GPU layers. Instead of all 40, it might load 32 layers into VRAM. That’s about 6.4 GB for weights. The remaining VRAM, roughly 1.6 GB, is left for the KV-Cache and overhead.
Step 3: The remaining 8 layers stay in RAM and are computed by the CPU. That’s 20% of the model on CPU, 80% on GPU.
Step 4: Expected speed is around 20 to 30 tokens per second, depending on CPU and RAM speed. That’s much faster than CPU-only but slower than full offloading with 12 GB VRAM.
Step 5: To improve speed, you can shrink the context window to reduce the KV-Cache, or quantize more aggressively, say to 3-bit, to squeeze more layers into VRAM. More on that under Quantization.
This example shows that GPU offloading isn’t a yes-or-no decision but a spectrum. Every layer you can fit into VRAM improves speed. The RAM and VRAM requirements guide will help you work through the numbers for your specific model.
Alternatives to GPU Offloading
GPU offloading isn’t your only option when VRAM is tight. Here are the main alternatives.
Stronger quantization: Dropping a model from 4-bit to 3-bit or 2-bit shrinks memory use significantly. A 13B model that needs 8 GB at 4-bit might fit in 6 GB at 3-bit and then load entirely into your VRAM. The downside is quality loss, noticeable at 3-bit and substantial at 2-bit. See Quantization for details.
Smaller model: Sometimes the simplest answer is to pick a smaller model. A 7B model needs about 4 to 5 GB and fits comfortably in 8 GB VRAM. Quality from good 7B models is often surprisingly close to 13B models. If your use case doesn’t strictly require the larger model, you’ll save yourself a lot of frustration.
More VRAM: The obvious solution is a graphics card with more VRAM. An RTX 3060 with 12 GB is affordable and runs 13B models fully. An RTX 4090 with 24 GB handles even 30B models. The hardware buying guide helps you pick.
Apple Silicon: Macs with Apple Silicon use Unified Memory, where CPU and GPU share one pool. A Mac with 32 GB Unified Memory can offer nearly all of it to the GPU. The small-VRAM problem doesn’t exist the same way here. The tradeoff is lower memory bandwidth than dedicated high-end cards.
Common Pitfalls with GPU Offloading
- Forgetting the KV-Cache: The KV-Cache needs extra VRAM and grows with text length. If you only count model weights, you think everything fits, but longer conversations exhaust VRAM and the model slows down.
- Forcing too many layers onto the GPU: If you manually load more layers than space allows, the program crashes or throws errors. Let Ollama decide, or reduce the count step by step.
- Expecting linear speed gains: 50% GPU offloading doesn’t mean 50% of full speed. Data transfer and CPU computation create disproportionate slowdown. Set expectations accordingly.
- Underestimating RAM speed: CPU computation of remaining layers depends heavily on RAM speed. Slow DDR4 RAM makes partial offloading far slower than fast DDR5.
- Ignoring the PCIe bottleneck: Data moves between RAM and VRAM over PCIe. Old PCIe versions or a card in the wrong slot limit transfer rates and slow offloading.
- Overlooking the CPU as a bottleneck: A weak CPU makes layers on the CPU a bottleneck. A strong GPU helps little if the CPU can’t compute remaining layers fast enough.
- Chasing full offloading when partial suffices: Partial offloading is often enough, especially for 13B models at 75% on the GPU. You don’t need to buy a bigger card just to fit 100% in VRAM.
- Not considering stronger quantization: Before settling for slow offloading, try more aggressive quantization. Often the model then fits entirely in VRAM and runs much faster.
Hardware, Costs, and Security in GPU Offloading
Hardware: GPU offloading requires a dedicated graphics card with at least some VRAM, a reasonably modern CPU, and sufficient RAM. If you want to run a 13B model with 8 GB VRAM using offloading, your system should have at least 16 GB RAM, ideally 32 GB, so the model has enough space and your system stays responsive. Your CPU should be at least a current 6-core processor to prevent CPU layers from becoming a bottleneck.
Costs: GPU offloading itself is free, it’s a software feature. The expense comes from hardware. You can pick up a used RTX 3060 with 12 GB VRAM for around 200 euros, which lets you load 13B models completely without needing offloading at all. If you stick with a smaller card and use offloading, you save money but pay in speed. Do the math yourself to decide if the savings justify slower inference.
Security: GPU offloading doesn’t change the security profile of local AI. The model runs entirely on your machine, whether in VRAM, RAM, or split across both. No request leaves your system, and no data goes to a server. Memory transfers between RAM and VRAM happen internally and never touch the network. If you’re using local AI for privacy reasons, GPU offloading is completely safe.
Further Reading and Resources on GPU Offloading
- AI Hardware Basics for foundational knowledge about hardware for local AI
- CPU vs. GPU if you want to understand why GPUs are faster
- RAM vs. VRAM for the basics of both memory types
- Memory Bandwidth for the most important speed factor
- Quantization if you prefer to shrink models instead of using offloading
- RAM and VRAM Requirements for exact numbers on model sizes
- Ollama for the tool that handles offloading automatically
- GPU Buying Guide if you’re thinking about upgrading your graphics card
FAQ: GPU Offloading - Common Questions
What is GPU offloading in simple terms?
GPU offloading means an AI model runs partly on the GPU in VRAM and partly on the CPU in RAM. This happens when the model doesn’t fit entirely in VRAM. The GPU handles the bulk of the work, and the CPU handles the rest.
When do I need GPU offloading?
When your model needs more memory than your VRAM offers. Typical case: you have 8 GB VRAM and want to run a 13B model that needs 8 GB or more. Without offloading, the model runs entirely on the CPU and is slow.
How much faster is GPU offloading compared to CPU-only?
It depends on how much of the model fits on the GPU. With 75 percent on the GPU, you often get three to five times the CPU-only speed. At 50 percent, it’s considerably less. The gain isn’t linear.
Does GPU offloading make sense for every model?
No. For small models that fit entirely in VRAM, you don’t need offloading. For very large models where only 10 percent fits on the GPU, offloading is often too slow to be practical. The sweet spot is 50 to 80 percent of the model on the GPU.
Can I manually choose how many layers go to the GPU?
Yes. In Ollama, use the num_gpu parameter. In llama.cpp, use -ngl. Both let you set the number of GPU layers. If you omit the parameter, the program decides automatically.
What happens if I load more layers to the GPU than it has space for?
The program will crash or throw an error because VRAM is exhausted. Reduce the number of layers until the model starts. Ollama does this automatically if you don’t set the parameter manually.
What’s the difference between partial and full offloading?
Full offloading means the entire model sits in VRAM and the CPU doesn’t participate in inference. Partial offloading means only part of the model is in VRAM, and the rest is in RAM. Full offloading is faster but requires more VRAM.
Does more RAM help with GPU offloading?
Yes, but only indirectly. More RAM means CPU layers compute faster because there’s enough memory for the model and the system. Speed depends more on RAM speed and CPU performance than on the raw amount.
Do I need GPU offloading with Apple Silicon?
Not in the classical sense. Apple Silicon uses Unified Memory, where CPU and GPU share the same memory pool. There’s no distinction between VRAM and RAM. The model lives in shared memory, and the GPU accesses it directly. The concept of offloading doesn’t exist in the same form here.
Is GPU offloading supported with multiple graphics cards?
Yes. Ollama and llama.cpp can distribute models across multiple GPUs. Each GPU takes some layers. This is one way to keep large models entirely in VRAM without offloading to the CPU.
Does the model lose quality through GPU offloading?
No. GPU offloading doesn’t change the calculations, it just splits them between two processors. The result is identical whether the model runs on the GPU, CPU, or split between them. The only difference is speed.
Sources and Further Reading
- llama.cpp documentation on GPU offloading and layer distribution
- Ollama documentation on GPU parameters
- RAM vs. VRAM as a foundation for understanding memory types
- Memory Bandwidth as a critical factor for speed
- Community experience and benchmarks from the local AI ecosystem


