Skip to content
BotServBotServ
RAMVRAMAI HardwareModel SizeQuantizationHardware Requirements

RAM and VRAM Requirements for AI Models

How much RAM and VRAM do local AI models need? Overview by model size, quantization, and use case.

S

schutzgeist

18 min read
RAM and VRAM Requirements for AI Models

RAM and VRAM Requirements

What this article covers

  • How much RAM and VRAM you actually need for typical AI models
  • A clear formula to calculate memory requirements yourself
  • Concrete calculations for Llama 3.1 8B, Qwen 2.5 14B, and Llama 3.1 70B
  • Recommended hardware by budget, from under 500 EUR to over 2000 EUR
  • Common pitfalls that prevent your model from running

Introduction: Understanding RAM and VRAM requirements

Before you run an AI model locally, you need to know whether your hardware can handle it. The key question is simple: how much memory does the model need while running? RAM and VRAM play the central role here. The larger the model and the longer the context, the more memory you’ll need.

Why does this matter? Because memory is the primary bottleneck in local AI. You can have the fastest processor and the most expensive GPU, but if you don’t have enough memory, the model won’t start at all or will crash mid-inference. Unlike cloud services where you simply pay for more capacity, local AI is bound by the physical limits of your hardware. This article will help you understand those limits.

If you’re new to local AI, start with What is local AI?. This article builds on those foundations and dives deep into hardware requirements.

Why do you need to know these requirements?

Imagine you buy a computer for 2000 EUR. You’ve consulted with experts, chosen a strong CPU, installed a nice GPU with 8 GB VRAM, and 32 GB of system RAM. You get home and want to run Llama 3.1 70B because you read online that it’s a really capable model. You fire up Ollama, download the model, and then it happens: output becomes sluggish, takes forever, or the model won’t load at all.

What went wrong? The 70B model needs roughly 40 GB of memory even in its most aggressive quantization. Your GPU has 8 GB VRAM, leaving maybe 7 GB after the operating system. The model spills almost entirely into slow system RAM. The CPU handles the computation instead of the GPU, and instead of 30 tokens per second, you get maybe 1 or 2. Your 2000 EUR investment was fine for gaming, but not for a 70B model.

If you’d done the math beforehand, you’d know a 70B model needs a GPU with 24 GB VRAM or more, ideally two of them, or an Apple Silicon Mac with 64 GB Unified Memory. With that knowledge, you’d either plan different hardware or choose a smaller model like Llama 3.1 8B, which runs fine on your card.

That’s the point: knowing the requirements leads to the right purchasing decision. Not knowing them risks wasting thousands on hardware that doesn’t fit your actual needs.

RAM and VRAM requirements at a glance

RAM is your computer’s main memory. VRAM is the video memory on a dedicated graphics card. AI models run on either. GPUs with sufficient VRAM are much faster because they support massive parallel processing. Without enough VRAM, the model falls back to system RAM or the CPU, becoming noticeably slower.

As a rule of thumb: a quantized 7B model needs about 4 to 8 GB of memory. A 13B model needs 8 to 16 GB. A 70B model needs 40 GB or more depending on quantization. Add to that the KV-cache, which grows with context length.

The most important number to remember: take the model size in billions of parameters, divide by 2, and you get the memory requirement in GB for FP16. For Q4_K_M, use roughly one quarter to one third of that. It’s rough, but it gives you a starting point before diving into details.

Who is this article for?

This article is for beginners running an AI model locally for the first time and unsure whether their hardware is sufficient. It’s also useful if you’re planning to buy or upgrade a computer for local AI and want to know what you really need beforehand. You don’t need prior knowledge of AI or hardware, we explain everything step by step.

If you already understand quantization and KV-cache, jump straight to the tables and example calculations. If not, no problem, the next sections will bring you up to speed.

Key terms around RAM and VRAM requirements

TermMeaning
RAMComputer’s main memory, used for CPU-based execution
VRAMMemory on the graphics card, much faster for AI inference
GPUGraphics processor, accelerates AI inference through parallel cores
CPUMain processor of the computer, runs the model when no GPU is available or VRAM is full
Memory requirementRequired RAM or VRAM during model execution
KV-CacheBuffer for previously computed attention values, grows with context length
QuantizationReducing the numerical precision of model weights, decreases memory footprint
GPU OffloadingParts of the model run on GPU, the rest in system RAM
Unified MemoryApple Silicon shares RAM between CPU and GPU, no separate VRAM needed

For more on quantization, see Quantization. For context length and its impact on KV-cache, see Context Length.

How is memory requirement calculated?

Total memory usage when running a model consists of three parts:

  1. Model weights: the actual storage for the model. Depends on parameter count and quantization.
  2. KV-Cache: buffer for attention calculations, grows with context length.
  3. Overhead: space for the operating system, other running programs, and framework internals.

Formula for model weights

The basic formula for model weights is:

Memory (GB) = Parameter count × Bytes per parameter / 1024^3

FP16 uses 2 bytes per parameter, Q8_0 roughly 1 byte, Q5_K_M roughly 0.7 bytes, and Q4_K_M roughly 0.55 bytes.

Example: FP16, 7B model:

7,000,000,000 × 2 / 1024^3 = ~13 GB

Example: Q4_K_M, 7B model:

7,000,000,000 × 0.55 / 1024^3 = ~3.6 GB

Formula for KV-Cache

KV-Cache grows with context length. A rough estimate for a typical model:

KV-Cache (GB) = 2 × Number of layers × Dimension × Context length × 2 bytes / 1024^3

For Llama 3.1 8B with 32 layers, dimension 4096, and 4096 token context:

2 × 32 × 4096 × 4096 × 2 / 1024^3 = ~2 GB

At 32,768 token context, that same cache grows to roughly 16 GB. This shows how critical context length is for memory needs.

Overhead

Budget roughly 2 to 4 GB for the operating system and framework internals. On Windows with a browser and other applications running, it can be 4 to 6 GB. When in doubt, add 4 GB as a safety buffer.

Overall Memory Formula

Total Memory = Model Weights + KV Cache + Overhead

This formula is simplified, but it gives you a solid estimate. In practice, there are minor framework variations, but for a purchasing decision, this is more than sufficient.

Memory Requirements by Model Size

The table below shows memory requirements for model weights alone, without KV cache or overhead. Values are rounded and based on typical GGUF quantizations.

Model SizeQ4_K_MQ5_K_MQ8_0FP16
7B~4 GB~5 GB~7 GB~14 GB
8B~5 GB~6 GB~8 GB~16 GB
13B~8 GB~9 GB~13 GB~26 GB
30B~18 GB~21 GB~31 GB~60 GB
70B~40 GB~45 GB~70 GB~140 GB

These figures cover the model itself only. Plan extra headroom for KV cache and your operating system, as explained in the next section.

What Else Consumes Memory?

Model weights are just one part of the equation. Three additional factors determine total memory usage.

KV Cache

The KV cache stores already-computed attention values so the model doesn’t recalculate the entire context for each new token. It grows linearly with context length. At short contexts of 2048 tokens it’s negligible, but at 32,768 tokens it can reach several gigabytes with larger models.

For typical use with an 8,192 token context, budget an extra 1 to 4 GB depending on model size. At 32,768 tokens, the KV cache for a 70B model can balloon to 20 GB or more.

Operating System and Other Applications

Your operating system itself consumes memory. Windows with a browser and background processes typically uses 4 to 6 GB. Linux gets by on 1 to 2 GB. If you’re coding, editing documents, or keeping other tools open at the same time, the footprint grows.

How Much Buffer Should You Plan?

As a rule of thumb: take the memory requirement for model weights, add the KV cache for your desired context length, then add another 20 percent on top. This ensures the model doesn’t run right at the edge and you have room left for other applications.

Example: An 8B model in Q4_K_M needs 5 GB. KV cache for 8,192 tokens adds roughly 2 GB. Overhead is 4 GB. Total: 11 GB. With a 20 percent buffer you land at about 13 GB. A graphics card with 12 GB VRAM will be tight, one with 16 GB is comfortable.

Example Calculations

Three concrete examples that answer typical questions.

Example 1: I want to run Llama 3.1 8B

Llama 3.1 8B in Q4_K_M requires about 5 GB for model weights. With a context length of 8,192 tokens, add roughly 2 GB for KV cache. Overhead for the OS and framework is about 4 GB.

5 GB (Model) + 2 GB (KV Cache) + 4 GB (Overhead) = 11 GB

With a 20 percent buffer: roughly 13 GB. You need a graphics card with at least 12 GB VRAM, ideally 16 GB, or a machine with at least 16 GB RAM for CPU execution. 16 GB RAM will work, 32 GB RAM puts you on the safe side.

Recommended hardware: RTX 3060 12 GB for GPU execution, or a machine with 32 GB RAM for CPU execution. A Mac Mini M4 with 24 GB Unified Memory handles this without issue.

Example 2: I want to run Qwen 2.5 14B

Qwen 2.5 14B in Q4_K_M requires about 9 GB for model weights. KV cache for 8,192 tokens adds roughly 3 GB. Overhead is 4 GB.

9 GB (Model) + 3 GB (KV Cache) + 4 GB (Overhead) = 16 GB

With a 20 percent buffer: roughly 19 GB. You need a graphics card with at least 24 GB VRAM for pure GPU execution, or you can use GPU offloading with a 16 GB card and spill the rest to system RAM. That requires at least 32 GB of system RAM.

Recommended hardware: RTX 4090 24 GB for pure GPU execution, or RTX 4060 Ti 16 GB with 32 GB system RAM for GPU offloading. A Mac Studio with 64 GB Unified Memory handles this comfortably.

Example 3: I want to run Llama 3.1 70B

Llama 3.1 70B in Q4_K_M requires about 40 GB for model weights. KV cache for 8,192 tokens adds roughly 8 GB. Overhead is 4 GB.

40 GB (Model) + 8 GB (KV Cache) + 4 GB (Overhead) = 52 GB

With a 20 percent buffer: roughly 62 GB. A single consumer GPU cannot handle this. You need either two RTX 4090s with 24 GB VRAM each, or an Apple Silicon Mac with at least 64 GB Unified Memory, preferably 128 GB.

Recommended hardware: Mac Studio M2 Ultra with 128 GB Unified Memory, or two RTX 4090s in a workstation case. Alternatively, a cloud rental setup, though that falls outside local AI.

BudgetWhat’s PossibleRecommended Models
Under 500 EURCPU execution on existing machine, 16 GB RAM7B to 8B in Q4_K_M, slow but usable
500 to 1000 EURMachine with 32 GB RAM or used GPU with 8 to 12 GB VRAM7B to 13B in Q4_K_M, GPU execution possible
1000 to 2000 EURMachine with RTX 3060 12 GB or RTX 4060 Ti 16 GB, 32 GB RAM8B to 14B in Q4_K_M, smooth output
Over 2000 EURRTX 4090 24 GB or Mac Studio with 64 GB Unified Memory30B in Q4_K_M, 70B with offloading possible

This table is a guideline. The exact choice depends on whether you’re buying new or upgrading, and which models you primarily want to run. More details are in the buying guide and the article AI PC for Beginners.

Apple Silicon Recommendations

Apple Silicon has a major advantage for local AI: Unified Memory. CPU and GPU share the same memory pool, there’s no separate VRAM. That means a Mac with 64 GB RAM effectively has 64 GB available for AI models minus the OS. This isn’t achievable with Nvidia consumer GPUs, since the largest consumer GPU tops out at 24 GB VRAM.

Mac Mini M4

The Mac Mini M4 is the most affordable entry point to Apple Silicon for local AI. With 24 GB Unified Memory it runs 7B and 8B models in Q4_K_M without issue, Q8_0 is also possible. 13B models in Q4_K_M work as well. Larger models get tight.

With 32 GB Unified Memory you can comfortably run 14B models in Q4_K_M and 30B models with some offloading. The Mac Mini M4 is the best choice for beginners who don’t want to spend much but still want smooth performance.

Mac Studio

The Mac Studio M2 Max with 64 GB Unified Memory handles 30B models in Q4_K_M easily and 70B models in Q4_K_M with patience. The Mac Studio M2 Ultra with 128 GB Unified Memory runs 70B models in Q5_K_M or Q8_0 and is currently the most powerful consumer solution for local AI.

The downside of Apple Silicon is speed compared to Nvidia GPUs. For models that fit entirely in an RTX 4090’s VRAM, the 4090 is noticeably faster. Only for models that don’t fit in 24 GB VRAM does Apple Silicon shine.

GPU Recommendations

RTX 3060 12 GB

The RTX 3060 with 12 GB VRAM is the price-to-performance winner for beginners. Used, it costs under 300 EUR and runs 7B and 8B models in Q4_K_M and Q5_K_M without issue. 13B models in Q4_K_M also fit, Q8_0 gets tight. It’s not sufficient for 30B and larger.

RTX 4060 Ti 16 GB

The RTX 4060 Ti with 16 GB VRAM is the next step up. New units cost around 450 EUR and provide enough space for 13B and 14B models in Q4_K_M and Q5_K_M. 8B models also run in Q8_0 or FP16. For 30B models, it works with offloading to system RAM, but performance slows noticeably.

RTX 4090 24 GB

The RTX 4090 with 24 GB VRAM is Nvidia’s most powerful consumer GPU. New units cost around 1800 EUR and run 30B models in Q4_K_M entirely within VRAM, enabling very smooth output. 70B models in Q4_K_M require offloading to system RAM, which slows them down but keeps them usable. For pure 70B inference without compromise, you need two 4090s or Apple Silicon with 128 GB.

For more on GPU selection and hardware fundamentals, see KI-Hardware Grundlagen and the article RAM vs. VRAM.

How do I check my current available memory?

Before loading a model, you should know how much memory you have at your disposal. Three commands help with this.

Check VRAM on Nvidia GPUs

Open a terminal and run:

nvidia-smi

This displays a table with all Nvidia GPUs, their total VRAM, and currently used memory. Pay attention to the Memory-Usage column. Whatever is free there is available for your model.

Check running models in Ollama

If you use Ollama, you can see which models are currently loaded and how much memory they occupy:

ollama ps

This shows a table with model name, size, VRAM usage, and status. You can immediately see whether a model fits entirely in VRAM or partially spills into system RAM.

Check system RAM

On Linux:

free -h

On Windows, open Task Manager with Ctrl+Shift+Esc and check the Performance tab for memory usage. On macOS, open Activity Monitor and check the Memory tab.

This way you always know where you stand before starting a model.

Common pitfalls with RAM and VRAM requirements

  1. VRAM budget too tight: You calculate the exact memory needed for a model but forget KV-cache and overhead. The model starts but crashes during longer contexts. Always allocate a 20 percent buffer.

  2. Underestimating context length: A model with 8192 token context uses far less KV-cache than the same model with 32,768 tokens. Processing long documents can double the memory requirement due to KV-cache growth.

  3. Overlooking operating system overhead: On Windows, the OS with browser and background processes easily consumes 4 to 6 GB. If your GPU has 12 GB VRAM, only 6 to 8 GB effectively remains for the model.

  4. Wrong quantization choice: Q8_0 needs twice as much memory as Q4_K_M, often with only marginally better quality. When memory is tight, choose Q4_K_M over Q8_0.

  5. Misunderstanding GPU offloading: When a model doesn’t fit entirely in VRAM, Ollama moves parts to system RAM. It works, but runs slower. You need enough system RAM to hold the remainder.

  6. Using FP16 for everyday use: FP16 requires four times more memory than Q4_K_M without noticeably better quality for most purposes. FP16 matters for training and fine-tuning; for inference, Q4_K_M or Q5_K_M is the better choice.

  7. Misjudging Apple Silicon memory: On Apple Silicon, CPU and GPU share memory. A Mac with 32 GB Unified Memory doesn’t have 32 GB VRAM; it has 32 GB minus the operating system minus other applications. Plan a buffer here too.

Hardware, costs, and security regarding RAM and VRAM requirements

The easiest way to start is with existing hardware. If you already have a machine with 16 GB RAM or more, you can immediately test small quantized models. If you plan to invest in hardware deliberately, keep your target model size and context length in mind.

Costs scale with the GPU. Consumer GPUs with 12 or 24 GB VRAM are inexpensive compared to workstations. High-end GPU workstations or servers with multiple GPUs quickly become costly. Apple Silicon offers good value if you need to run models beyond 24 GB, since Unified Memory is significantly cheaper than multiple Nvidia GPUs.

From a security perspective, the major advantage of local AI is that no data needs to leave your machine, regardless of how much hardware it contains. Your prompts, documents, and results stay on your computer. This is especially important for sensitive professional or research data. Learn more about secure local AI operation in the article LLM lokal betreiben.

Further resources on RAM and VRAM requirements

FAQ - Common questions about RAM and VRAM

Can I run a model if I have less memory than recommended?

Yes, often it works, just slower. Ollama and similar tools offload parts of the model to system RAM when VRAM runs short. For smooth operation, follow the recommendations; otherwise, response generation noticeably slows down.

How much VRAM do I need for Llama 3.1 8B?

For the Q4_K_M variant, around 6 to 8 GB VRAM including KV-cache and buffer. Q5_K_M should have 8 to 10 GB, Q8_0 around 12 GB. With 16 GB VRAM, you’re safe for all quantizations.

Does Ollama run on Apple Silicon?

Yes. Apple Silicon with Unified Memory is popular for local AI because RAM and GPU memory are shared. A Mac with 16 or 24 GB is good for small to medium models; a Mac Studio with 64 GB or more handles large models well.

What’s better: more RAM or more VRAM?

For fast operation, VRAM on a dedicated GPU is better since the GPU processes the model in parallel. More RAM helps when the GPU isn’t sufficient and parts of the model offload to system RAM. Ideally, you have both: enough VRAM for the model and enough RAM as a buffer.

Do I need two GPUs for large models?

Not necessarily. With aggressive quantization, large models fit on a single consumer GPU with 24 GB VRAM. For FP16 or Q8_0 with 70B or larger, you need more VRAM, so multiple GPUs or Apple Silicon with large Unified Memory.

What happens when VRAM is full?

The framework offloads excess model parts to system RAM. The model keeps running, but those parts are computed by the CPU, which is much slower. You’ll notice this as a lower tokens-per-second rate.

How much RAM do I need minimum for local AI?

For small models like 7B in Q4_K_M, 16 GB RAM works in CPU mode. For smooth operation with GPU offloading, you should have 32 GB RAM. For large models starting at 30B, 64 GB RAM or more is recommended.

Is an integrated graphics card enough for local AI?

For very small models and experiments, yes, but performance is significantly worse than with a dedicated GPU. Integrated GPUs share memory with system RAM and have few compute cores. For serious use, a dedicated GPU is recommended.

How does context length affect memory requirements?

The KV-cache grows linearly with context length. A model with 2048 token context needs minimal extra memory; at 32,768 tokens, the KV-cache can be several gigabytes. If processing long documents, plan for more memory accordingly.

Should I choose Q4_K_M or Q5_K_M?

Q4_K_M is the best starting point for tight memory; quality is good enough for most use cases. Q5_K_M needs about 25 percent more memory with minimally better quality. If you have enough memory, use Q5_K_M; if tight, use Q4_K_M.

Can I run multiple models at the same time?

Yes, but each model loads separately into memory. Two 8B models in Q4_K_M together need about 10 GB just for weights, plus KV-cache and overhead. Plan enough memory if running multiple models in parallel.

How much memory does the operating system itself need?

On Linux, 1 to 2 GB is typical. On Windows with browser and background processes, 4 to 6 GB is common. On macOS, allocate about 3 to 4 GB. If unsure, use 4 GB as an OS buffer.

Is FP16 worth it for better quality?

For everyday inference, the quality difference between Q4_K_M and FP16 is usually small, but memory demand is four times higher. FP16 matters for training and fine-tuning. For running models, Q4_K_M or Q5_K_M is the better choice.

What is Unified Memory on Apple Silicon?

Unified Memory means CPU and GPU share the same memory pool. There is no separate VRAM. This is an advantage for local AI: a Mac with 64 GB RAM effectively has 64 GB available for models, minus the OS. Nvidia GPUs cap consumer VRAM at 24 GB.

Sources and Further Reading

  • Ollama Model Documentation
  • Nvidia GPU Specifications
  • Hugging Face Model Cards
  • llama.cpp Quantization Documentation
  • Apple Silicon Specifications
Back to Blog
Share:

Related Posts