Gemma Models: Google’s Lightweight AI
What This Article Covers
- What Gemma is and how the model family has evolved
- Which Gemma versions exist and what they’re suited for
- How much VRAM you need for different Gemma models
- How to run and use Gemma locally with Ollama
- The strengths and weaknesses of this model family
Introduction
Gemma is Google’s contribution to the open source AI landscape. The name derives from “gem” and stands for the Latin letters G-E-M-M-A, echoing Gemini, Google’s large commercial AI model. The key difference: while Gemini is a cloud service, Gemma models are open source and can run locally on your machine.
Gemma models stand out for their compact size. Even the smallest version with 2 billion parameters delivers surprisingly good results for its scale. The largest version with 27 billion parameters competes with models that have significantly more parameters. This makes Gemma particularly attractive for users with limited hardware who still want quality.
In this article, you’ll learn which Gemma versions exist, how they differ, and how to run them locally with Ollama. If you’re still uncertain which model suits you, start with the model overview and the article Finding the Right Model.
Why Do You Need Gemma?
Picture this: you have a laptop with 16 GB of RAM and want to run a decent AI model locally anyway. Many well-known models like Llama 3 70B or DeepSeek-R1 are too large for that hardware. This is where Gemma comes in: the 2B version runs on nearly any modern machine, and the 9B version needs only about 6 GB VRAM in Q4_K_M.
Gemma becomes particularly interesting if three things matter to you: compact size, good quality per parameter, and a solid foundation for experimentation. Google trained these models using the same research and techniques that power Gemini. That means you get Google-quality performance in a package that runs on consumer hardware.
If you’re already familiar with Llama models or Phi models, Gemma adds another compact option to your toolkit. The 9B version especially is a strong competitor to Llama 3 8B, and the 27B version delivers quality at a level usually found only in 30B or 70B models.
Gemma Models Explained
The Gemma model family has evolved across several generations. Here’s an overview of the key versions:
Gemma 1 was the first generation, available in 2B and 7B sizes. It launched in February 2024 and demonstrated that Google could build compact open source models. For current applications, Gemma 1 is less relevant since successors perform significantly better.
Gemma 2 is the second generation, released in June 2024. It comes in three sizes: 2B, 9B, and 27B. Gemma 2 uses an improved architecture with Grouped-Query Attention and local Sliding-Window Attention, which substantially boosts quality at the same parameter count. The 9B version achieves benchmark results that compete with Llama 3 8B. The 27B version outperforms models twice its size.
Gemma 3 is the newest generation. It expands the family with multimodal capabilities, meaning it can process images, and offers a larger context window. Gemma 3 is available in multiple sizes and targets users seeking a compact all-purpose model with image understanding.
For local deployment, Gemma 2 and Gemma 3 are the most relevant versions. Gemma 1 sees little use anymore since its successors outperform it in every way.
Who This Article Is For
This article is for you if you:
- need a compact model for limited hardware
- want to try Google’s technology without cloud dependency
- plan to run a lightweight model on a laptop or Mac
- wonder whether Gemma is an alternative to Llama or Phi for your needs
- seek a model that performs well on 8 GB of VRAM
You don’t need prior knowledge of Google’s AI research. If you understand what an LLM is and know how to use Ollama, you’re ready to go.
Key Terms
| Term | Definition |
|---|---|
| Gemma | Google’s open source model family, derived from Gemini research |
| Gemini | Google’s commercial AI model line, available only as a cloud service |
| Grouped-Query Attention | Efficient attention method that saves memory and compute |
| Sliding-Window Attention | Attention mechanism that considers only a window of recent tokens |
| Multimodal | A model’s ability to handle multiple input types such as text and images |
| Context window | Maximum number of tokens the model can process at once |
| Parameters | Learnable values in the model that determine its size and capabilities |
| Quantization | Reducing numerical precision to lower memory requirements |
| Tokenizer | Component that breaks text into tokens the model processes |
| Instruction-Tuned | A model trained to follow instructions effectively |
Gemma 2: The Second Generation
Gemma 2 is the version most local users currently run. The three sizes cover different hardware requirements:
Gemma 2 2B is the smallest version and runs on nearly any hardware. At only 2 billion parameters, it requires about 1.5 GB of memory in Q4_K_M. It’s sufficient for basic chat tasks, summarization, and simple text processing. On a laptop with 8 GB of RAM, it runs without issues. Quality is good for the size, but it hits limits quickly on complex tasks.
Gemma 2 9B is the family’s sweet spot. With 9 billion parameters, it delivers quality on par with Llama 3 8B, sometimes exceeding it. In Q4_K_M, it needs about 6 GB of VRAM, which runs smoothly on most modern graphics cards with 8 GB or more. The 9B version is the best choice for users wanting a strong all-around model that doesn’t consume excessive memory.
Gemma 2 27B is the family’s flagship. With 27 billion parameters, it achieves benchmark results that compete with models having 50 to 70 billion parameters. In Q4_K_M, it requires about 16 GB of VRAM. This fits on an RTX 4080 with 16 GB or a Mac with 24 GB Unified Memory. The 27B version is ideal for users who want maximum quality and have the hardware to match.
| Model | Parameters | VRAM (Q4_K_M) | VRAM (Q5_K_M) | VRAM (Q8_0) | Best For |
|---|---|---|---|---|---|
| Gemma 2 2B | 2B | ~1.5 GB | ~2 GB | ~2.5 GB | Laptops, minimal hardware |
| Gemma 2 9B | 9B | ~6 GB | ~7 GB | ~10 GB | All-purpose use on 8 GB+ GPUs |
| Gemma 2 27B | 27B | ~16 GB | ~19 GB | ~28 GB | High-end, maximum quality |
For more on memory requirements, see RAM and VRAM Requirements and Model Size and Storage Needs.
Gemma 3: Multimodal and flexible
Gemma 3 brings several important capabilities to the family. The most significant addition is multimodal support: Gemma 3 can process and describe images, not just text. This opens up applications like image captioning, visual question-answering systems, and diagram analysis.
Beyond that, Gemma 3 offers a substantially larger context window than Gemma 2. While Gemma 2 is limited to 8,192 tokens, Gemma 3 can handle significantly more context. This matters when you need to analyze lengthy documents or sustain extended conversations. For more on context windows, see Kontextlaenge.
Gemma 3 comes in multiple sizes, ranging from very small to medium. The architecture has been further optimized to deliver better results at comparable parameter counts. For most local users, Gemma 3 is the go-to choice when looking for a compact model with image understanding.
Architecture: What makes Gemma special
Gemma 2 and 3 employ several architectural features that set them apart:
Grouped-Query Attention (GQA): This technique reduces KV-cache memory requirements without sacrificing quality. Instead of storing separate values for each attention head, multiple heads share the same values. This makes Gemma especially efficient for longer contexts.
Local Sliding-Window Attention: Gemma 2 combines global and local attention. Local attention only examines a window of recent tokens, saving compute. Global attention is periodically interspersed to maintain awareness of the full context.
Knowledge Distillation: Gemma 2 leverages knowledge distillation from larger models. This means smaller Gemma models learned from the outputs of larger ones. That’s why Gemma 2 9B and 27B punch above their weight in benchmarks.
Together, these techniques explain why Gemma models often outperform larger competitors in benchmarks. Google has essentially translated Gemini research into compact open-source models.
Running Gemma locally with Ollama
The simplest way to run Gemma locally is via Ollama. Here are the essential commands:
# Start Gemma 2 2B
ollama run gemma2:2b
# Start Gemma 2 9B
ollama run gemma2:9b
# Start Gemma 2 27B
ollama run gemma2:27b
# Start Gemma 3 4B
ollama run gemma3:4b
# Start Gemma 3 12B
ollama run gemma3:12b
# Start Gemma 3 27B
ollama run gemma3:27b
Ollama automatically downloads the Q4_K_M variant. If you want a different quantization, you can specify it, for example ollama run gemma2:9b-q5_K_M. See Quantisierung for more on quantization.
Gemma compared to other models
How does Gemma stack up against the competition? Here’s a rough overview:
- Gemma 2 9B vs. Llama 3 8B: Both are roughly equivalent. Gemma 2 9B edges out Llama 3 8B on some benchmarks, while Llama 3 8B wins on others. Both require similar VRAM.
- Gemma 2 9B vs. Phi-3: Phi-3 is smaller and more resource-efficient, but Gemma 2 9B delivers better quality on most tasks.
- Gemma 2 27B vs. Llama 3 70B: Llama 3 70B is stronger overall but demands significantly more VRAM. Gemma 2 27B is the better choice if you don’t have 40 GB VRAM available.
- Gemma 3 vs. Gemma 2: Gemma 3 is the successor and the better option in most cases, especially if you need image understanding or longer context windows.
Your choice ultimately depends on your hardware and use case. With 8 GB VRAM, Gemma 2 9B and Gemma 3 4B or 12B are solid picks. With 16 GB or more, consider Gemma 2 27B or Gemma 3 27B.
Common pitfalls
1. Confusing Gemma with Gemini: Gemini is Google’s cloud model and not available locally. Gemma is the open-source version you can run locally. They’re related but not identical.
2. Using outdated Gemma 1 versions: Gemma 1 is significantly weaker than Gemma 2 and 3. If you use Gemma, ensure you’re running the latest version. In Ollama, Gemma 2 is called gemma2, not gemma.
3. Underestimating Gemma 2’s context window: Gemma 2 has a context window of 8,192 tokens, which is smaller than many competitors. This can be limiting for long documents. Gemma 3 offers a larger window.
4. Running the 27B version on undersized hardware: Gemma 2 27B requires roughly 16 GB VRAM in Q4_K_M. On a 12 GB GPU, it will offload to system RAM and become slow. Use the 9B version instead.
5. German responses weaker than English: Gemma models were trained primarily on English. German works but isn’t as strong as models like Qwen or Llama 3, which are multilingual-optimized.
6. Expecting multimodal capabilities from Gemma 2: Only Gemma 3 can process images. Gemma 2 is text-only. If you need image understanding, choose Gemma 3.
7. Choosing quantization too aggressively: For the 2B version, Q3 or Q2 can cause significant quality loss because the model is already small. Use at least Q4_K_M, ideally Q5_K_M for smaller models.
Hardware, costs, and security
Hardware: Gemma models are particularly hardware-friendly. The 2B version runs on almost any machine with 4 GB RAM or more. The 9B version requires a GPU with 8 GB VRAM or a Mac with 16 GB Unified Memory. The 27B version calls for a GPU with 16 GB VRAM or a Mac with 24 GB or more. An RTX 3060 with 12 GB is ideal for the 9B version.
Costs: All Gemma models are open source and free. Model weights are available on Hugging Face and in the Ollama library. The Gemma license permits commercial use but includes some conditions worth reviewing. For most applications, usage is unproblematic.
Security: As with all local models, your data stays on your machine. This is an advantage over Google’s Gemini cloud service, where data is sent to Google. With Gemma locally, you get Google-grade quality without cloud dependency. Still, check whether your software (Ollama, LM Studio) has telemetry features.
Further reading
- Modell-Übersicht - All model articles at a glance
- Modell finden - How to find the right model
- Llama-Modelle - Meta’s model family compared
- Phi-Modelle - Microsoft’s compact models
- Ollama - The easiest way to run local models
- Quantisierung - How models get smaller
- RAM and VRAM requirements - How much memory you need
- Model size and memory requirements - The relationship between size and memory
FAQ
What’s the difference between Gemma and Gemini?
Gemini is Google’s commercial cloud AI model, available only through an API or web interface. Gemma is the open-source model family that you can run locally on your machine. Both build on similar research, but they’re distinct models.
Are Gemma models free?
Yes, all Gemma models are free to use. Model weights are available on Hugging Face and in the Ollama library. The Gemma license also permits commercial use under certain conditions.
Which Gemma version should I choose?
With 8 GB VRAM, go for Gemma 2 9B or Gemma 3 12B. With 16 GB VRAM or more, pick Gemma 2 27B or Gemma 3 27B. If you only have 4 GB VRAM, use Gemma 2 2B or Gemma 3 4B. For most users, Gemma 2 9B is the best starting point.
Can Gemma process images?
Only Gemma 3 is multimodal and can handle images. Gemma 2 is a text-only model. If you need image understanding, choose Gemma 3.
Is Gemma 2 9B better than Llama 3 8B?
Both models perform roughly on par. In some benchmarks, Gemma 2 9B edges ahead slightly; in others, Llama 3 8B does. They require similar amounts of VRAM. Test both and compare them on your own tasks.
How large is Gemma’s context window?
Gemma 2 has a context window of 8,192 tokens. Gemma 3 offers a significantly larger context window. If you’re processing long documents, Gemma 3 is the better choice.
Can I run Gemma on a Mac with Apple Silicon?
Yes, Gemma runs smoothly on Apple Silicon. A Mac with 16 GB unified memory handles Gemma 2 9B without issue. A Mac with 24 GB or more can also run Gemma 2 27B or Gemma 3 27B.
Do I need an Nvidia GPU for Gemma?
No. Gemma runs on CPU, Apple Silicon, and AMD GPUs. Ollama supports all these platforms. An Nvidia GPU is fastest, but for the smaller variants, a standard laptop processor is sufficient.
Can I use Gemma commercially?
Yes, the Gemma license permits commercial use. You should still review the exact license terms, especially if you’re integrating Gemma into a product. For most use cases, commercial deployment is straightforward.
Where do I find quantized Gemma models?
Ollama automatically downloads a Q4_K_M variant when you start a Gemma model. On Hugging Face, you’ll find additional quantizations from providers like bartowski or TheBloke. Search for “gemma2” plus “GGUF” on Hugging Face.
Is Gemma 3 always better than Gemma 2?
In most cases, yes, especially if you need image understanding or longer context windows. For pure text tasks, the differences are smaller. If you’re already using Gemma 2 and satisfied, switching isn’t essential.
Sources
- Google DeepMind: Gemma model releases and technical reports
- Hugging Face: Gemma model cards and weights
- Ollama model library: Gemma models
- Google AI Blog: Gemma 2 and Gemma 3 announcements
- Gemma license on the official Gemma website


