Skip to content
BotServBotServ
RAMVRAMGPUCPUAI HardwareMemory

RAM vs VRAM

Understand the difference between RAM and VRAM for local AI. Where models run and which is faster.

S

schutzgeist

14 min read
RAM vs VRAM

RAM vs. VRAM

What this article covers

  • The difference between RAM and VRAM and why both matter for local AI
  • How a model travels from storage to RAM to VRAM
  • What happens when VRAM runs out and how GPU offloading works
  • How Apple Silicon and APUs with unified memory change the game
  • Practical recommendations for which setup handles which model size

Introduction: RAM vs. VRAM explained

When you run AI models locally on your own machine, you quickly encounter two types of memory that often get confused: RAM and VRAM. Both are working memory, both feed data to processors, but they live in different places and operate at vastly different speeds. That distinction alone determines whether a model runs smoothly or you wait seconds between words.

If you’re working with local AI, there’s no way around it. Tools like Ollama let you spin up models with a single command, but whether the result feels instant or painfully slow depends on where the model ends up in memory. A solid grasp of RAM and VRAM helps you buy the right hardware, pick models that fit your setup, and skip the frustration.

Why does this matter?

Picture yourself buying a new PC. 32 GB RAM, a decent CPU, a graphics card with 4 GB VRAM. The salesperson assures you it’s perfect for AI work. You install Ollama, download a 13-billion-parameter model, and hit start. Then you wait. Every response takes ten seconds or longer. The computer isn’t slow, the model isn’t broken, but the VRAM is simply too small for that model.

This is a classic mistake many make. RAM and VRAM get lumped together because both are quoted in gigabytes. But 32 GB of RAM helps little if your GPU has only 4 GB of VRAM. The model then runs partly in RAM, which is considerably slower. If you understand how RAM and VRAM work together beforehand, you buy the right hardware and avoid buyer’s remorse.

RAM vs. VRAM - In A Nutshell

RAM is your computer’s main working memory. It’s available to the CPU and every running program. VRAM is the memory on a dedicated graphics card and belongs exclusively to the GPU. For AI computations, VRAM is typically far faster because the GPU has thousands of parallel cores and VRAM offers much higher memory bandwidth.

Think of your computer as a workshop. RAM is the desk where all your papers and materials sit while you work. The bigger the desk, the more materials fit. VRAM is the workbench right next to the machine. The machine is the GPU, and it can only process what sits on the workbench. If the workbench is too small, the machine keeps having to fetch materials from the desk, which costs time.

A model meant to run on the GPU must fit entirely in VRAM. If it doesn’t, parts spill back into RAM. The model slows down because the CPU has far fewer parallel cores than a GPU, and transferring data between RAM and VRAM adds overhead.

Who should read this?

This article is for beginners exploring local AI and wondering what hardware they need. If you’ve never heard of VRAM or aren’t sure why a 7-billion-parameter model runs fast on a GPU but drags on a CPU, you’re in the right place. No prior knowledge required. We’ll step through what RAM and VRAM are, how they interact, and what to consider when shopping.

If you’ve already tinkered with Ollama and puzzled over why some models fly while others crawl, this article helps too. The fundamentals here underpin more advanced topics like RAM and VRAM requirements and quantization.

Key terms for RAM and VRAM

TermDefinition
RAMSystem working memory, used by CPU and all programs
VRAMMemory on a dedicated graphics card, exclusive to the GPU
GPUProcessor on the graphics card with many parallel cores, ideal for AI
CPUMain processor with fewer but more flexible cores
GPU OffloadingParts of the model are moved from the CPU to the GPU
Unified MemoryShared memory for CPU and GPU, for example on Apple Silicon
Shared MemoryGPU uses part of RAM, for instance on APUs without dedicated VRAM
Memory BandwidthAmount of data that can flow between memory and processor per second
APUProcessor with integrated GPU, often shares RAM
Apple SiliconApple chips like M1, M2, M3, M4, which unify RAM and GPU memory

Direct comparison: RAM vs. VRAM

CriterionRAMVRAM
LocationOn the motherboard, inside your computerOn the graphics card, directly next to the GPU
AccessCPU and all programsGPU only
Speed for AIModerate, limited by CPU coresVery high, thanks to thousands of GPU cores
BandwidthTypically 50 to 100 GB/sTypically 500 to 1000 GB/s and higher
LatencyHigherLower, directly attached to GPU
AvailabilityPresent in every computerOnly with a dedicated graphics card
Cost per GBCheap, typically 3 to 5 eurosExpensive, typically 15 to 30 euros and more
UpgradeableEasy, RAM modules can be swappedHardly at all, VRAM is soldered directly to the GPU
Typical sizes16, 32, 64 GB8, 12, 16, 24 GB on consumer GPUs
Primary userCPUGPU
Impact when scarcePrograms slow down or crashModel falls back to CPU, becomes very slow

How RAM and VRAM work together

When you start an AI model with Ollama, this happens behind the scenes. The model sits as a file on your SSD or hard drive. At startup, Ollama loads the file into RAM. RAM is the system’s general working memory, holding everything needed right now.

If you have a dedicated graphics card, Ollama then tries to copy the model from RAM into VRAM. The GPU computes far faster than the CPU because it has thousands of parallel cores. The more of the model that lives in VRAM, the faster inference runs, meaning generating responses.

The flow looks roughly like this:

  1. Model sits as a file on your SSD
  2. Ollama loads the file into RAM
  3. If a GPU is present, Ollama copies the model from RAM to VRAM
  4. The GPU performs calculations and generates the response
  5. The result goes back to the program displaying the answer

When the model fits entirely in VRAM, that’s ideal. The GPU has everything it needs right beside it, and data transfer between RAM and VRAM during inference stays minimal. The model runs at peak speed.

What Happens When VRAM Runs Out?

When a model doesn’t fit entirely into VRAM, GPU offloading kicks in. Ollama loads as much of the model as possible into VRAM and keeps the rest in RAM. The GPU processes the parts stored in VRAM while the CPU handles the parts in RAM. It works, but it’s slower because data constantly moves back and forth between RAM and VRAM.

The less of the model that fits in VRAM, the more work the CPU has to do, and the slower the model becomes. In extreme cases, when there’s no GPU at all or VRAM is tiny, the entire model runs on the CPU in RAM. For small models this is acceptable, but for larger ones it becomes very slow.

One solution is quantization. This compresses the model by reducing the precision of the numbers it uses. A model that originally requires 16 GB VRAM might shrink to 5 GB with 4-bit quantization, making it fit on a smaller graphics card. The tradeoff is slight quality loss, but for most applications it’s barely noticeable.

Another solution is simply getting a graphics card with more VRAM. Consumer Nvidia GPUs come with 8, 12, 16, or 24 GB of VRAM. The buying guide can help you choose.

Apple Silicon and Unified Memory

Apple Silicon, the M1, M2, M3, M4 chips and their Pro and Max variants, work fundamentally differently from computers with separate CPU and GPU. There’s no dedicated VRAM on Apple Silicon. Instead, the CPU and GPU share one pool of memory called Unified Memory.

This means the GPU can access all system RAM. A Mac Studio with 128 GB Unified Memory can theoretically give the GPU nearly all of that. That’s a huge advantage for local AI, because large models that wouldn’t fit in a consumer graphics card with 24 GB VRAM might run fine on a Mac with enough Unified Memory.

The downside is memory bandwidth. Apple Silicon has good bandwidth, but it’s lower than high-end dedicated GPUs like the Nvidia H100. For running local AI at home, Apple Silicon is still an excellent choice because sheer memory capacity is the most important factor, and the bandwidth is sufficient for most tasks.

If you’re considering Apple Silicon, pay attention to which variant you get. Standard M1, M2, M3, and M4 chips have lower memory bandwidth than their Pro, Max, and Ultra counterparts. For AI work, the Pro variants are the minimum worthwhile choice, though Max or Ultra are better.

APUs and Shared Memory

Beyond Apple Silicon, there’s another approach where CPU and GPU share memory: APUs. An APU is a processor with an integrated GPU. Examples include AMD Ryzen AI Max or Intel Core Ultra series. These chips have no dedicated VRAM, so the GPU uses a portion of RAM instead. This is called Shared Memory.

The advantage is that you don’t need a separate graphics card and still benefit from GPU acceleration. The downside is that RAM bandwidth is significantly lower than dedicated VRAM. The GPU is faster than a pure CPU solution but slower than a dedicated graphics card with its own VRAM.

The AMD Ryzen AI Max is particularly interesting because it supports up to 128 GB of RAM and can make a large portion available to the GPU. It’s a genuine alternative to Apple Silicon if you want to build a Windows or Linux machine for running large models locally without buying an expensive dedicated graphics card.

Example: Which Setup for Which Model?

Here are concrete examples to give you a sense of scale. The figures refer to quantized models as typically loaded with Ollama. For exact details, see RAM and VRAM requirements.

7B model (7 billion parameters): A quantized 7B model needs roughly 4 to 5 GB of memory. It fits in any modern graphics card with 8 GB VRAM. On a GPU it runs fast and smooth. On a CPU with 16 GB RAM it works too, but noticeably slower.

13B model (13 billion parameters): A quantized 13B model requires about 8 to 9 GB of memory. It still fits on a 12 GB VRAM card like the Nvidia RTX 3060. With 8 GB VRAM it’s tight, and you’ll need GPU offloading or stronger quantization.

70B model (70 billion parameters): A quantized 70B model needs roughly 40 to 45 GB of memory. This doesn’t fit in any consumer graphics card, since the largest, the RTX 4090, has 24 GB VRAM. You’ll need multiple graphics cards, Apple Silicon with at least 64 GB Unified Memory, or an APU like the Ryzen AI Max with sufficient RAM.

Larger models: Models with 100 billion parameters and beyond require 60 GB of memory and up. For local use, only Macs with 128 GB Unified Memory or multi-GPU setups are feasible.

Common Pitfalls with RAM and VRAM

  • Confusing RAM with VRAM: 32 GB of RAM doesn’t help if your GPU has only 4 GB VRAM. The model will be slow because it won’t fit in VRAM.
  • Buying too small a graphics card: If you buy a GPU with 8 GB VRAM and then try to run a 13B model, you’ll either live with GPU offloading or need stronger quantization.
  • Treating shared memory like dedicated VRAM: An APU uses RAM as VRAM but it’s slower than a dedicated graphics card with its own VRAM because bandwidth is lower.
  • Underestimating Apple Silicon bandwidth: A standard M3 is less suitable for AI than an M3 Max because memory bandwidth is significantly lower.
  • Ignoring quantization: Loading models without quantization needs much more VRAM. With 4-bit quantization, much larger models fit in the same space.
  • Overlooking bandwidth: Even with enough memory, low memory bandwidth can make your model slow. This especially applies with shared memory.
  • Expecting VRAM to be upgradeable: You can’t upgrade VRAM. To get more, you have to buy a new graphics card. RAM, on the other hand, is usually easy to expand.

Hardware, Cost, and Security with RAM and VRAM

Hardware: To get started with local AI, a graphics card with 8 GB VRAM is sufficient if you mainly run 7B models. For 13B models, 12 GB VRAM is recommended; for larger models, 24 GB VRAM or Apple Silicon with adequate Unified Memory. The CPU plays a secondary role as long as it’s modern. The GPU and its VRAM matter most.

Cost: RAM is cheap, typically 3 to 5 euros per GB. VRAM is expensive because it’s soldered to the graphics card and you have to buy the whole card. An Nvidia RTX 4060 with 8 GB VRAM costs around 300 euros; an RTX 4090 with 24 GB VRAM around 2000 euros. Apple Silicon is special because you don’t buy storage separately, you configure the Mac with your desired Unified Memory. A Mac Studio with 128 GB Unified Memory is pricey, but often cheaper than a multi-GPU setup.

Security: When you run models locally, your data stays on your machine. No query goes to a server, no data leaves your system. That’s a major advantage over cloud AI. RAM and VRAM don’t matter here; what matters is that the model runs locally and doesn’t call external APIs. Learn more under What is local AI?.

Further Reading and Resources on RAM and VRAM

FAQ - Common Questions About RAM and VRAM

Can I expand VRAM with RAM?

No. VRAM is soldered directly to the graphics card and cannot be upgraded. If you need more VRAM, you’ll need to buy a new graphics card. RAM, however, can usually be expanded easily by installing additional modules into your motherboard.

Is an 8 GB VRAM GPU worth it?

Yes, for getting started. 8 GB VRAM is plenty for quantized 7B models. With 13B models, you’ll be stretched thin, and you’ll need either GPU offloading or heavier quantization. For larger models, a card with 12 or 16 GB VRAM makes more sense.

Why is Apple Silicon so popular for local AI?

Apple Silicon uses Unified Memory, where the CPU and GPU share a common memory pool. The GPU can access the entire system RAM without the typical VRAM limitations of dedicated graphics cards. A Mac with 64 or 128 GB Unified Memory can load models that wouldn’t fit on any consumer graphics card.

What’s faster: CPU with lots of RAM or GPU with little VRAM?

For AI inference, a GPU is almost always faster, even with limited VRAM, as long as the model fits. GPUs have thousands of parallel cores and far higher memory bandwidth. A CPU with abundant RAM only wins when the model is too large for VRAM and must run on the CPU anyway.

What exactly is GPU offloading?

GPU offloading means parts of a model are moved to the GPU while the rest runs on the CPU using RAM. This happens when the entire model doesn’t fit in VRAM. The more of the model that fits in VRAM, the faster it runs.

Does more RAM help when VRAM is too small?

Somewhat. More RAM lets you load and run larger models on the CPU. But performance stays low because the CPU and RAM are slower than the GPU and VRAM. More RAM doesn’t solve the speed problem, it just makes it possible for the model to run at all.

What’s the difference between Unified Memory and Shared Memory?

Unified Memory on Apple Silicon is a common pool optimized from the start for both CPU and GPU with high bandwidth. Shared Memory on APUs means the integrated GPU uses a portion of regular RAM. Bandwidth is lower than with dedicated VRAM or Unified Memory.

Do I need an Nvidia GPU for local AI?

No, but Nvidia is the easiest choice because most tools and frameworks are optimized for it. AMD GPUs work too, though they sometimes require more configuration. Apple Silicon and APUs are alternatives that don’t require a dedicated graphics card.

How much VRAM do I need for a 7B model?

A quantized 7B model needs roughly 4 to 5 GB VRAM. With 8 GB VRAM, you’re on safe ground. Without quantization, the model requires about 14 GB, which demands a much larger graphics card.

Can I use multiple graphics cards for more VRAM?

Yes, that’s possible. Tools like Ollama support multi-GPU setups where the model is distributed across the VRAM of several cards. It’s a way to run large models locally without Apple Silicon or an APU. The downside is high power consumption and cost.

What happens when VRAM is completely full?

When VRAM fills up and the model isn’t fully loaded, the program crashes or throws an error. Ollama tries to fall back to GPU offloading in such cases, keeping the rest in RAM. If RAM also runs out, the model won’t start at all.

Is integrated GPU VRAM enough?

In most cases, no. Integrated GPUs use Shared Memory from RAM and have low bandwidth. Small 7B models might work, but slowly. For serious local AI, a dedicated graphics card or Apple Silicon is worth it.

Should I upgrade RAM or VRAM if my model is slow?

If the model is slow because it doesn’t fit in VRAM, more RAM won’t help with speed. You’d need a graphics card with more VRAM or stronger quantization. More RAM only helps if the model won’t start because RAM is also too small.

Sources and Further Reading

  • Nvidia memory architecture and VRAM specifications
  • AMD Ryzen AI Max documentation
  • Apple Silicon Unified Memory overview
  • Ollama documentation on GPU offloading
  • Insights from the local AI community
Back to Blog
Share:

Nächster Artikel in AI Hardware

Weiterlesen
AI PC Beginners Guide

Related Posts