Ollama Model Recommendations by Use Case
What This Article Covers
- Recommended models for typical use cases.
- Differences between chat, coding, reasoning, and vision models.
- How to choose the right model size.
- Quantization recommendations.
- Hardware considerations.
Introduction: Ollama Model Recommendations by Use Case
Choosing the right model is one of the most important decisions you’ll make with Ollama. Not every model suits every purpose. Some are general-purpose and fast, while others specialize in coding, reasoning, or vision tasks. Your hardware ultimately determines which model sizes and quantization levels are even feasible. Knowing which models fit your needs saves time and frustration.
This article provides practical model recommendations for common use cases.
Key Terms
- 7B, 13B, 70B: Model size in billions of parameters.
- Instruct: Model optimized for following instructions.
- Chat: Model designed for conversation.
- Coder: Model specialized for programming tasks.
- Vision: Model that processes images.
- Reasoning: Model for logical inference.
- Embedding: Model for vector representations.
- Quant: Quantization level.
General Chat
| Model | Size | Use |
|---|---|---|
| llama3.1:8b | 8B | Fast, solid all-around quality |
| qwen2.5:7b | 7B | Fast, good for German |
| mistral:7b | 7B | Quick and compact |
| llama3.1:70b | 70B | High quality, requires significant VRAM |
For everyday questions, 7B to 8B models with Q4-K quantization work well.
Coding
| Model | Use |
|---|---|
| qwen2.5-coder:14b | Code completion |
| codellama:7b | Lighter tasks |
| deepseek-coder:6.7b | Fast and capable |
| qwen2.5-coder:32b | Better quality |
For coding tasks, set num_predict high enough since responses tend to be longer.
Vision and Image Analysis
| Model | Use |
|---|---|
| llava:13b | Image description |
| llava-llama3:8b | Balance between size and quality |
| bakllava | Better vision understanding |
Vision models require additional RAM/VRAM.
Reasoning
| Model | Use |
|---|---|
| qwen2.5:14b | Logic and math |
| deepseek-r1 | Reasoning with detailed explanations |
| phi4 | Smaller reasoning model |
Reasoning models often produce longer responses because they think step by step.
RAG
| Model | Use |
|---|---|
| llama3.1:8b | Standard RAG |
| qwen2.5:7b | Good context utilization |
| mistral:7b | Fast |
For RAG, the embedding model is equally important.
Translation and Summarization
| Model | Use |
|---|---|
| qwen2.5:7b | German/English, summarization |
| mixtral:8x7b | Multilingual tasks |
| gemma2:9b | Quality for mid-range needs |
Embedding Models
For RAG and semantic search:
nomic-embed-textmxbai-embed-largesnowflake-arctic-embed
Choosing Model Size
| VRAM | Model Size |
|---|---|
| 8 GB | 7B Q4 |
| 12 GB | 13B Q4 |
| 24 GB | 13B Q5/Q6 or 70B Q4 with CPU offload |
| 48 GB+ | 70B Q4 or larger |
If a model doesn’t fit in VRAM, it becomes sluggish or crashes.
Choosing Quantization
| Priority | Recommendation |
|---|---|
| Maximum speed | Q4_0 |
| Good balance | Q4_K_M |
| Better quality | Q5_K_M |
| Highest quality | Q8_0 |
Recommendation for Beginners
If you’re new to Ollama, start with llama3.1:8b or qwen2.5:7b. Both are fast, produce good quality, and run on 8 GB VRAM.
Hardware Guidance
- Mini-PC/Laptop: 3B to 7B models.
- AI PC with 16-24 GB VRAM: 7B to 13B.
- Workstation with 48 GB+: 70B models are feasible.
- CPU-only: 3B to 7B with Q4.
Tips
- Start with the model library at Ollama.com.
- Test multiple models.
- Pay attention to quantization.
- Choose the right model for each use case.
- Run benchmarks.
Further Reading and Resources
- BotServ.de Ollama Model Benchmarks
- BotServ.de Ollama Context Length
- BotServ.de Ollama Quantization in Practice
- BotServ.de Ollama Coding
FAQ: Ollama Model Recommendations
What’s the best general-purpose model? Llama 3.1 8B or Qwen 2.5 7B are solid all-rounders.
Which model is best for coding? Qwen 2.5 Coder or DeepSeek Coder.
Can I analyze images with Ollama? Yes, using vision models like Llava.
Do I need 70B models? Only if you want the highest quality and have sufficient VRAM.
Which model works best for RAG? A solid 7B chat model is usually enough.
Sources and Further Reading
- Ollama Library: https://ollama.com/library
- Hugging Face: https://huggingface.co/
- LMSYS Chatbot Arena: https://chat.lmsys.org/
Summary: Ollama Model Recommendations by Use Case
Model selection in Ollama depends on your use case, hardware, and preferred quantization level. Llama 3.1 and Qwen 2.5 are reliable all-rounders, Qwen Coder and DeepSeek Coder excel at programming, and Llava handles vision tasks. For RAG and simple applications, 7B models suffice. By paying attention to quantization and VRAM constraints, you’ll quickly find a model that runs smoothly and delivers good results.


