Quantization
What This Article Covers
- What quantization is and why it’s essential for local AI
- How to run large models on limited graphics cards using quantization
- Available formats and methods, from GGUF to AWQ to EXL2
- Finding the right quantization level for your hardware
- Common pitfalls and trade-offs to understand
Introduction: Quantization Explained
If you’re exploring local AI, you’ll encounter one term repeatedly: quantization. Modern language models contain billions of parameters, and each parameter takes up storage space. An unquantized model with 13 billion parameters requires roughly 26 GB of memory. Most consumer graphics cards offer only 8 to 12 GB of VRAM. Without quantization, running such models at home simply wouldn’t be feasible.
Quantization is what makes local AI practical for everyday use. It shrinks a model’s storage footprint without retraining it from scratch. This allows large models to fit on smaller GPUs, load faster, and consume less power. In this article, you’ll learn how it works, what methods exist, and which approach suits your use case.
For more context on running models locally, see Running LLMs Locally.
Why Do You Need Quantization?
A concrete example clarifies the problem immediately. Suppose you want to run a 13B model like Llama 2 13B locally. In its original form, it’s stored in FP16, meaning 16 bits per parameter. With 13 billion parameters, that’s roughly 26 GB.
Now look at your graphics card. A typical consumer GPU like the RTX 3060 has 12 GB of VRAM. The unquantized model won’t fit, not by a long shot. You could offload it to system RAM, but execution becomes extremely slow because system RAM has far less bandwidth than a GPU’s VRAM. For details, see RAM vs. VRAM.
Quantizing to Q4, which is 4 bits per parameter, shrinks the model to about 8 GB. That fits comfortably on a 12 GB graphics card, leaving room for the context and intermediate results during text generation. The model runs fast, smoothly, and entirely on the GPU without offloading.
That’s the core point: quantization isn’t an optional optimization detail. It’s the prerequisite for running large models on ordinary hardware at all.
Quantization in a Nutshell
Quantization means representing a model’s numerical values with fewer bits than the original. A model stored in 16-bit floating point (FP16) gets reduced to 8, 5, or 4 bits. Fewer bits mean less memory and faster computation.
Think of it like image compression. A PNG stores every pixel losslessly and takes up significant space. A JPG compresses the data, discarding information, yet looks nearly identical to the human eye. Quantization works similarly: you reduce the precision of stored values, but the model still produces useful, often barely distinguishable results.
A 7B model in FP16 needs about 14 GB. In Q4, it needs only 4 to 5 GB. On an 8 GB graphics card, it runs without problems. The trade-off is a possible, usually minor quality loss, especially at very low levels like Q4 or Q3.
Who Is Quantization For?
Quantization is designed for anyone who wants to run language models locally without investing thousands in server hardware. Specifically, that includes:
- Home users with a consumer graphics card looking to run a capable model
- Developers testing models on workstations before production deployment
- Privacy-conscious users wanting to keep models offline and avoid sending data to cloud services
- Researchers and students experimenting on a limited budget
- Apple Silicon users wanting to efficiently leverage their Mac’s unified memory
In short: if you want to run local AI, you need quantization. There’s practically no way around it.
Key Terms in Quantization
| Term | Meaning |
|---|---|
| FP16 | 16-bit floating point representation, the standard format for trained models |
| INT8 | 8-bit integer representation, halves memory use compared to FP16 |
| INT4 | 4-bit integer representation, reduces memory to one-quarter of FP16 |
| Q4 | Quantization level with 4 bits per weight |
| Q5 | Quantization level with 5 bits per weight |
| Q6 | Quantization level with 6 bits per weight |
| Q8 | Quantization level with 8 bits per weight, close to the original |
| GGUF | File format for quantized models, used by llama.cpp and Ollama |
| K-Quantization | Mixed quantization method storing sensitive weights with more bits |
| Bits | Number of storage bits per parameter value, determines precision and memory use |
| Weights | The learned parameters of the model, adjusted during training |
| Activations | Intermediate values computed during model execution |
How Quantization Works
To understand quantization, you need to know how a model stores numbers. Each parameter, every weight in the model, is a number. Unquantized, this number is stored as a 16-bit floating point value in FP16 format. 16 bits can represent roughly 65,500 different values, allowing high precision.
During quantization, you take these 16-bit numbers and convert them to numbers with fewer bits. At 8 bits, you have 256 possible values. At 4 bits, only 16. This means you group nearby values and assign them a shared, coarser value. A number like 0.123456 becomes 0.125, for instance, because that’s the nearest level representable with 4 bits.
Why does this save memory? Simple: fewer bits per value means fewer bytes overall. A 13B model has 13 billion values. At 16 bits, that’s 2 bytes per value, requiring 26 GB total. At 4 bits, that’s 0.5 bytes per value, requiring only 6.5 GB. Memory use scales proportionally with bit count.
There are two main approaches:
1. Post-Training Quantization (PTQ): The model is trained normally, then quantized afterward. This is the most common case locally. You take a finished FP16 model and convert it to a Q4 or Q8 version. The process takes minutes to hours depending on model size.
2. Quantization-Aware Training (QAT): The model is prepared for quantization during training itself. This yields better results but is labor-intensive and typically done by model developers, not end users.
For local deployment, PTQ is almost always the relevant approach. You download a model already quantized, or quantize it yourself using tools like llama.cpp.
An important distinction is what gets quantized. In most local methods, only the weights are quantized, not the activations. This simplifies the process and suffices for most use cases. During execution, the quantized weights are briefly converted to higher precision for calculation, a step called dequantization.
Types of Quantization
Several methods and formats matter in the local AI space. They differ in compression technique, speed, and compatibility.
GGUF (GPT-Generated Unified Format): The standard format for local models, used by llama.cpp, Ollama and LM Studio. GGUF stores a model in a single file, including metadata and tokenizer. It supports many quantization levels from Q2 to Q8 and K-quantization variants. GGUF is the safest choice for running a model locally because it works on CPU, GPU, and Apple Silicon.
AWQ (Activation-aware Weight Quantization): A method that identifies particularly important weights and stores them with higher precision. AWQ often delivers better quality at 4 bits than simpler approaches, but is mainly designed for GPU operation with vLLM or Text-Generation-Inference. For pure llama.cpp users, GGUF is the better choice.
GPTQ: An older approach that quantizes weights based on their influence on model output. GPTQ produces good results at 4 bits but is increasingly displaced by AWQ and newer methods. It works well with certain frameworks but is less flexible than GGUF.
EXL2 (ExLlamaV2): A format that allows flexible bit rates, for example 4.5 bits or 6.25 bits. EXL2 is optimized for speed with ExLlamaV2 and suits users seeking maximum inference speed on Nvidia GPUs. It’s less common than GGUF but popular among speed enthusiasts.
bitsandbytes: A library that integrates 8-bit and 4-bit quantization directly into PyTorch. It’s often used for training and fine-tuning, for instance with QLoRA. For pure inference, bitsandbytes is less typical, but relevant for developers wanting to adapt models.
Which method works best depends on your setup. For most users, GGUF via Ollama or llama.cpp is the simplest and most compatible path. If you want maximum GPU speed and use ExLlamaV2, choose EXL2. If you’re fine-tuning models, bitsandbytes is essential.
K-Quantization Explained
When browsing quantized models, you’ll encounter labels like Q4_0 and Q4_K_M. The difference matters.
Q4_0 is the straightforward variant: all model weights are uniformly quantized to 4 bits. It’s fast, simple, and saves maximum space, but ignores that not all weights matter equally.
Q4_K_M is a K-quantization variant. The “K” stands for a method that stores different parts of the model with varying precision. Sensitive layers that strongly influence output get more bits, for example 6 bits. Less critical layers get 4 bits or fewer. The “M” stands for “Medium” and represents a compromise between size and quality.
Why is this better? A model consists of many layers, and not every layer matters equally to the final result. If you quantize everything uniformly, you lose too much information in the important layers. K-quantization preserves quality in the critical parts and saves aggressively only where it matters less. The result is a model barely larger than a Q4_0 version but noticeably better answers.
Several K-variants exist:
- Q4_K_S (Small): Slightly more aggressive compression, smaller than Q4_K_M
- Q4_K_M (Medium): The most popular compromise, recommended for most users
- Q5_K_M: 5-bit foundation with mixed quantization, higher quality than Q4
- Q6_K: 6-bit foundation, very close to the original, but larger
As a rule of thumb: if you use Q4, pick Q4_K_M, not Q4_0. The quality difference is noticeable, the storage difference minimal.
Common Quantization Levels Compared
| Level | Bits | Storage (13B model) | Quality | Speed | Best for |
|---|---|---|---|---|---|
| FP16 | 16 | ca. 26 GB | Original | Baseline | Servers, high-end workstations |
| Q8_0 | 8 | ca. 13 GB | Nearly identical | Very fast | GPUs with 16+ GB VRAM |
| Q6_K | 6 | ca. 10 GB | Excellent | Fast | Quality priority |
| Q5_K_M | 5 | ca. 9 GB | Very good | Fast | Solid compromise |
| Q4_K_M | 4 | ca. 8 GB | Good | Very fast | Standard choice for consumer GPUs |
| Q4_0 | 4 | ca. 7 GB | Slightly weaker than Q4_K_M | Very fast | When every GB counts |
| Q3_K_M | 3 | ca. 6 GB | Noticeably weaker | Very fast | Last resort with minimal VRAM |
| Q2_K | 2 | ca. 5 GB | Significantly limited | Very fast | Experimentation only |
Storage figures are estimates and vary slightly depending on model architecture and tokenizer size.
Example: Storage Savings with a 13B Model
Here’s a concrete calculation showing how much you save through quantization. We’ll use a 13B model with 13 billion parameters.
FP16 (Original): 13 billion parameters times 2 bytes (16 bits) equals 26 GB. This doesn’t fit on any typical consumer graphics card.
Q8_0: 13 billion parameters times 1 byte (8 bits) equals 13 GB. This fits on an RTX 4080 with 16 GB VRAM, but tightly, since you still need space for context.
Q5_K_M: 13 billion parameters times 0.625 bytes (5 bits, plus some overhead from mixed quantization) equals around 9 GB. This fits on an RTX 3080 with 10 GB or an RTX 4070 with 12 GB.
Q4_K_M: 13 billion parameters times 0.5 bytes (4 bits, plus overhead) equals around 8 GB. This fits on an RTX 3060 with 12 GB, with enough room for a decent context.
The savings from FP16 to Q4_K_M are enormous: you need only about a third of the memory. Instead of an expensive workstation GPU, a mid-range graphics card suffices.
More on storage requirements for different models is in the article RAM and VRAM Requirements.
Which Quantization Level Should You Choose?
Your choice primarily depends on available VRAM. Here’s a decision guide:
You have 8 GB VRAM or less: Choose Q4_K_M for models up to 7B parameters. For 13B models it gets tight; Q3_K_M or Q2_K are possible but expect quality loss. Better to run a smaller model at higher quality.
You have 12 GB VRAM: Q4_K_M for 13B models is ideal. You still have room for a generous context. Alternatively, Q5_K_M for 7B models if quality matters more than model size.
You have 16 GB VRAM: Q5_K_M or Q6_K for 13B models is a good choice. You get high quality and sufficient context. For 7B models you can use Q8_0 and be close to the original.
You have 24 GB VRAM or more: Q6_K or Q8_0 for 13B models deliver near-original quality. You can also run 30B or 34B models in Q4_K_M.
You use Apple Silicon: Apple Silicon shares RAM between CPU and GPU. A Mac with 16 GB can run a 7B model in Q4_K_M easily. With 32 GB, 13B models in Q4_K_M are realistic. With 64 GB or more, you can use 30B models in Q5 or Q6.
As general guidance: Q4_K_M is the standard choice you’ll rarely regret. If you have more VRAM, step up to Q5_K_M or Q6_K. Use Q8_0 when maximum quality matters more than storage savings. Avoid Q3 and Q2 except in emergencies; quality loss becomes noticeable there.
How to get quantized models
You don’t need to quantize models yourself. Several straightforward options exist for obtaining ready-to-use quantized versions.
Ollama: The simplest approach. Ollama automatically downloads an appropriate quantized version when you install a model. Type ollama run llama3 and Ollama handles the rest. By default, Ollama uses Q4_K_M, which works well for most users.
Hugging Face: The largest model repository. You’ll find quantized GGUF files there, which you can download manually and use with llama.cpp, LM Studio, or other tools. Search for the model name plus “GGUF”, for example “Llama-3-8B-Instruct-GGUF”.
TheBloke: One of the most prolific quantized model providers on Hugging Face. TheBloke has published hundreds of models in various quantization levels from Q2 to Q8. Search for “TheBloke” plus the model name on Hugging Face.
bartowski: Another active provider of quantized GGUF models. If TheBloke doesn’t have a specific model, check bartowski. The quantization quality is comparable.
LM Studio: A desktop application with a graphical interface that downloads models directly from Hugging Face. You search for a model, select the quantization level from a list, and LM Studio downloads the appropriate GGUF file. Ideal for users who prefer avoiding the command line.
If you want to quantize a model yourself, you’ll need the original model in FP16 and a tool like llama.cpp. The process is well documented, but for most users, downloading ready-made GGUF files is simpler.
Common pitfalls with quantization
1. Quantization too aggressive for complex tasks: Q4 works fine for chat and general text. For tasks requiring precise logical reasoning or exact formatting, such as code generation or mathematical calculations, lower levels can introduce errors. Use at least Q5_K_M or Q6_K for these.
2. Forgetting to account for context: A model’s memory requirement isn’t just the model itself. You also need space for the context (the ongoing conversation) and intermediate computation results. Budget roughly 1 to 2 GB extra for context, more for longer conversations. See Context Length for details.
3. Wrong format for your software: GGUF works with llama.cpp and Ollama. But if you’re using vLLM or ExLlamaV2, you need AWQ or EXL2. Check which formats your software supports before downloading a model.
4. CPU quantization versus GPU quantization: Some formats are optimized for CPU operation, others for GPU. GGUF runs on both, but isn’t always fastest on GPU. EXL2 is fast on Nvidia GPUs but won’t run on CPU. Choose the format based on your hardware.
5. Using outdated quantization variants: Q4_0 is older and simpler than Q4_K_M. When you have the choice, always pick the K variant. The quality difference is noticeable, the storage difference minimal.
6. Underestimating VRAM requirements: If a model barely fits in VRAM, longer conversations can cause problems as context grows. Leave yourself a safety margin, especially when working with long texts or large context windows.
7. Apple Silicon shares memory between system and GPU: On Macs with Apple Silicon, the CPU and GPU share the same pool. A 16-GB Mac doesn’t have 16 GB of dedicated VRAM; 16 GB is divided among the operating system, applications, and the model. Plan accordingly.
Hardware, cost, and security with quantization
Hardware: Quantized models dramatically reduce hardware requirements. Instead of an expensive workstation GPU with 24 GB VRAM, a mid-range graphics card with 8 to 12 GB often suffices. This opens local AI to a much broader user base. For hardware fundamentals, see AI Hardware Basics.
Cost: The savings are substantial. An RTX 3060 with 12 GB costs a fraction of an A100 with 80 GB. With quantization, you can run a 13B model on an RTX 3060 that would otherwise require a four-figure GPU.
Security: Quantization doesn’t change your data security. The quantized model is a compressed version of the same model, not a different one. Running it locally means all data stays on your machine. No information is sent to cloud services. This is one of the major advantages of local AI, and quantization doesn’t compromise it.
One thing to keep in mind: quantized models can occasionally behave slightly differently from the original, especially at very low quantization levels. If you’re using a model for security-critical applications, thoroughly test the quantized version before deploying it in production.
Further reading and resources on quantization
- What is local AI? - Fundamentals of running AI locally
- Running LLMs locally - Practical guide to get started
- RAM and VRAM requirements - How much memory you need
- Context length - Why context consumes memory
- Ollama - The easiest way to run local models
- RAM vs. VRAM - Memory type differences
- AI hardware basics - Overview of suitable hardware
FAQ - Common questions about quantization
Do I lose much quality through quantization?
At Q4, the difference is noticeable for many tasks but often still acceptable. At Q5 or Q6, quality loss decreases significantly. Q8 is practically indistinguishable from the original for most applications.
What’s the difference between Q4_0 and Q4_K_M?
Q4_0 quantizes all weights uniformly. Q4_K_M uses mixed quantization, storing more sensitive model parts with more bits, and usually delivers better results at nearly the same storage cost.
Can I quantize a model myself?
Yes, using tools like llama.cpp and corresponding conversion scripts. Most of the time, though, you’ll use already-quantized GGUF files from Hugging Face or Ollama, since that’s faster and simpler.
How much storage do I save with Q4?
About three to four times less than an FP16 model. A 13B model shrinks from roughly 26 GB to about 7 to 8 GB, depending on the variant.
Do I need quantization with Apple Silicon?
Yes. Apple Silicon also benefits greatly from quantized models because its unified memory is limited and faster loading times become possible. A 16-GB Mac can run a 7B model in Q4 comfortably.
What quantization level is best for beginners?
Q4_K_M is the standard recommendation. It offers a good balance between storage savings and quality, and works on most consumer graphics cards. If you’re unsure, start with Q4_K_M.
What does the “K” in Q4_K_M mean?
The “K” stands for the K-quantization method, which stores different model layers with varying precision. The “M” stands for “Medium”, a compromise between size and quality. There’s also “S” for Small and “L” for Large.
Is GGUF better than AWQ or GPTQ?
It depends on your software. GGUF is more flexible and runs on CPU, GPU, and Apple Silicon. AWQ and GPTQ are better suited for GPU operation with specific frameworks like vLLM. For most local users, GGUF is the better choice.
Can I further train a quantized model?
Not directly, but you can use QLoRA, a technique that enables fine-tuning on quantized models. The base model stays quantized, and only small additional modules are trained in higher precision.
What happens if the model doesn’t fit in VRAM?
The software offloads parts of the model to system RAM. This works but is noticeably slower because system RAM has less bandwidth than VRAM. Severe offloading makes text generation sluggish.
Are quantized models secure?
Yes. Quantization is purely a compression technique. The model remains the same; only its storage representation changes. Running it locally keeps your data on your machine.
Does quantization affect speed?
Yes, usually positively. Quantized models need less memory bandwidth, which can speed up execution. On GPU, 4-bit and 8-bit models are often faster than FP16 because less data needs to move.
Can I switch between quantization levels without reloading the model?
No. Each quantization level is a separate file. To change levels, you must load the corresponding model. With Ollama, you can install different levels of the same model and switch between them.
What’s the difference between weight quantization and activation quantization?
Weight quantization reduces only the stored model parameters. Activation quantization also quantizes intermediate values during execution. The latter is more involved and less common in the local space, but can provide additional speedup benefits.
Sources and Further Reading
- llama.cpp quantization documentation on GitHub
- Hugging Face model catalog and documentation
- TheBloke GGUF model collection on Hugging Face
- bartowski GGUF models on Hugging Face
- Ollama documentation for model management
- ExLlamaV2 documentation for the EXL2 format


