Skip to content
BotServBotServ
QuantizationGGUFBitsLocal AIModels

Model Quantization Fundamentals

Understand model quantization: bits, formats, trade-offs and tools for local AI models.

S

schutzgeist

3 min read
Model Quantization Fundamentals

Model Quantization Fundamentals

What this article covers

  • What quantization is and why it matters.
  • Differences between 4-bit, 8-bit, and other formats.
  • How quantization affects quality and speed.
  • Popular formats like GGUF, AWQ, GPTQ, and EXL2.
  • Tools and best practices to get started.

Introduction: Model Quantization Basics

Large language models contain billions of parameters, typically stored as 32-bit floating-point numbers. A 70B model alone requires over 250 GB of memory just for its weights. Quantization reduces the precision of these numbers to dramatically lower model size, RAM usage, and computation. The quality often remains surprisingly high.

Local AI benefits tremendously from quantization. Many models can run on consumer hardware once reduced to 4 bits or less. Understanding the fundamentals helps you choose the right format for your setup.

What is Quantization?

Quantization is the process of reducing the number of bits used to store model weights and activations. Instead of 32 bits per parameter, you might use 8 or 4 bits. This cuts memory requirements and bandwidth, speeds up calculations, and enables deployment on smaller devices.

Example:

  • A 7B model with 32-bit precision: roughly 28 GB.
  • The same model at 4-bit: roughly 4 GB.

Why Use Quantization?

  • Less memory: Models fit on smaller GPUs and CPUs.
  • Faster inference: Fewer data to load and process.
  • Run more models: Keep multiple models in memory simultaneously.
  • Better efficiency: Lower power consumption.
  • Potential quality loss: The more aggressive the quantization, the greater the risk.

Key Terms

  • FP32: 32-bit floating-point, high precision, high memory footprint.
  • FP16: 16-bit floating-point, half the memory cost.
  • INT8: 8-bit integer, good balance.
  • INT4: 4-bit integer, very small, may incur quality loss.
  • GGUF: Format from llama.cpp, widely used for local inference.
  • AWQ: Activation-aware Weight Quantization, often optimized for NVIDIA GPUs.
  • GPTQ: General-purpose Quantization, GPU optimized.
  • EXL2: Format for exllamav2, supports variable bit widths.
  • KV-Cache: Memory for intermediate computations, also quantizable.

Quantization Levels at a Glance

Bit depthMemoryQualitySpeedUse case
32-bit100 percentVery highSlowTraining, research
16-bit50 percentHighMediumGPUs with sufficient VRAM
8-bit25 percentGoodFastConsumer GPUs
4-bit12.5 percentAcceptable to goodVery fastLimited hardware
3-bit / 2-bitMinimalMedium to lowVery fastExperimental

Major Formats

GGUF

The most popular format for local models. Supports many bit depths and runs on both CPU and GPU. Ideal for Ollama, llama.cpp, and KoboldCpp.

Example:

llama-3.1-8b.Q4_K_M.gguf
  • Q4: 4-bit quantization.
  • _K_M: Mixed quantization with higher precision for important weights.

GPTQ

GPU-optimized, compresses weights for NVIDIA hardware. Works well with vLLM and AutoGPTQ.

AWQ

Particularly efficient for fast inference on NVIDIA GPUs.

EXL2

Enables variable bit widths and good performance with exllamav2.

Quantization with Ollama

Ollama ships many models already quantized. You can also load your own GGUF files:

ollama create mein-modell -f Modelfile

A simple Modelfile:

FROM ./llama-3.1-8b.Q4_K_M.gguf

SYSTEM "You are a helpful assistant."

Tools for Quantization

  • llama.cpp: Convert and quantize to GGUF.
  • AutoGPTQ: GPTQ quantization.
  • AutoAWQ: AWQ quantization.
  • exllamav2: EXL2 quantization.
  • Ollama: Works with pre-quantized models.
  • LM Studio: Load and manage quantized models.

Quality vs. Speed

Choosing the right balance depends on your hardware and application:

  • 7B models: Usually 4-bit or 8-bit on consumer hardware.
  • 13B models: 4-bit with Q4_K_M or Q4_K_S.
  • 70B models: Multiple GPUs or powerful CPU, often 4-bit.
  • Coding and reasoning: Prefer higher bit depth since precision matters.
  • Simple chat applications: 4-bit usually suffices.

Common Pitfalls

  • Over-quantizing: 2-bit models lose significant quality.
  • Wrong format: Running GGUF on GPU or GPTQ on CPU can be inefficient.
  • Ignoring KV-Cache: Large contexts consume substantial memory.
  • Focusing only on file size: Q4_K_M and Q4_0 differ in quality.
  • No testing: Always benchmark with your own prompts.

Further Reading and Resources

FAQ: Quantization

Do I lose a lot of quality at 4-bit? With modern methods like Q4_K_M, loss is often minimal. For critical applications, testing is recommended.

Can I quantize models myself? Yes, using llama.cpp, AutoGPTQ, AutoAWQ, or exllamav2.

Which is better: GGUF or GPTQ? GGUF is more versatile; GPTQ is often faster on NVIDIA GPUs.

How much VRAM do I need? Roughly the model size plus KV-Cache. An 8B model in 4-bit typically needs 6 to 8 GB VRAM.

Can I combine quantization methods? Yes, for example 4-bit weights with 8-bit activations.

Sources and Further Reading

Summary: Model Quantization Fundamentals

Quantization makes large AI models usable on modest hardware. By reducing bit depth, you lower memory requirements and computation, while quality remains high with good methods. Formats like GGUF, GPTQ, AWQ, and EXL2 offer different trade-offs. Understanding quantization helps you select the right level for your setup and application.

Back to Blog
Share:

Related Posts