Skip to content
BotServBotServ
OllamaModelsRecommendationsChatCoding

Ollama Model Recommendations by Use Case

Choose the right Ollama models for chat, coding, vision, RAG, and reasoning. Practical tips and comparisons.

S

schutzgeist

2 min read
Ollama Model Recommendations by Use Case

Ollama Model Recommendations by Use Case

What This Article Covers

  • Recommended models for typical use cases.
  • Differences between chat, coding, reasoning, and vision models.
  • How to choose the right model size.
  • Quantization recommendations.
  • Hardware considerations.

Introduction: Ollama Model Recommendations by Use Case

Choosing the right model is one of the most important decisions you’ll make with Ollama. Not every model suits every purpose. Some are general-purpose and fast, while others specialize in coding, reasoning, or vision tasks. Your hardware ultimately determines which model sizes and quantization levels are even feasible. Knowing which models fit your needs saves time and frustration.

This article provides practical model recommendations for common use cases.

Key Terms

  • 7B, 13B, 70B: Model size in billions of parameters.
  • Instruct: Model optimized for following instructions.
  • Chat: Model designed for conversation.
  • Coder: Model specialized for programming tasks.
  • Vision: Model that processes images.
  • Reasoning: Model for logical inference.
  • Embedding: Model for vector representations.
  • Quant: Quantization level.

General Chat

ModelSizeUse
llama3.1:8b8BFast, solid all-around quality
qwen2.5:7b7BFast, good for German
mistral:7b7BQuick and compact
llama3.1:70b70BHigh quality, requires significant VRAM

For everyday questions, 7B to 8B models with Q4-K quantization work well.

Coding

ModelUse
qwen2.5-coder:14bCode completion
codellama:7bLighter tasks
deepseek-coder:6.7bFast and capable
qwen2.5-coder:32bBetter quality

For coding tasks, set num_predict high enough since responses tend to be longer.

Vision and Image Analysis

ModelUse
llava:13bImage description
llava-llama3:8bBalance between size and quality
bakllavaBetter vision understanding

Vision models require additional RAM/VRAM.

Reasoning

ModelUse
qwen2.5:14bLogic and math
deepseek-r1Reasoning with detailed explanations
phi4Smaller reasoning model

Reasoning models often produce longer responses because they think step by step.

RAG

ModelUse
llama3.1:8bStandard RAG
qwen2.5:7bGood context utilization
mistral:7bFast

For RAG, the embedding model is equally important.

Translation and Summarization

ModelUse
qwen2.5:7bGerman/English, summarization
mixtral:8x7bMultilingual tasks
gemma2:9bQuality for mid-range needs

Embedding Models

For RAG and semantic search:

  • nomic-embed-text
  • mxbai-embed-large
  • snowflake-arctic-embed

Choosing Model Size

VRAMModel Size
8 GB7B Q4
12 GB13B Q4
24 GB13B Q5/Q6 or 70B Q4 with CPU offload
48 GB+70B Q4 or larger

If a model doesn’t fit in VRAM, it becomes sluggish or crashes.

Choosing Quantization

PriorityRecommendation
Maximum speedQ4_0
Good balanceQ4_K_M
Better qualityQ5_K_M
Highest qualityQ8_0

Recommendation for Beginners

If you’re new to Ollama, start with llama3.1:8b or qwen2.5:7b. Both are fast, produce good quality, and run on 8 GB VRAM.

Hardware Guidance

  • Mini-PC/Laptop: 3B to 7B models.
  • AI PC with 16-24 GB VRAM: 7B to 13B.
  • Workstation with 48 GB+: 70B models are feasible.
  • CPU-only: 3B to 7B with Q4.

Tips

  • Start with the model library at Ollama.com.
  • Test multiple models.
  • Pay attention to quantization.
  • Choose the right model for each use case.
  • Run benchmarks.

Further Reading and Resources

FAQ: Ollama Model Recommendations

What’s the best general-purpose model? Llama 3.1 8B or Qwen 2.5 7B are solid all-rounders.

Which model is best for coding? Qwen 2.5 Coder or DeepSeek Coder.

Can I analyze images with Ollama? Yes, using vision models like Llava.

Do I need 70B models? Only if you want the highest quality and have sufficient VRAM.

Which model works best for RAG? A solid 7B chat model is usually enough.

Sources and Further Reading

Summary: Ollama Model Recommendations by Use Case

Model selection in Ollama depends on your use case, hardware, and preferred quantization level. Llama 3.1 and Qwen 2.5 are reliable all-rounders, Qwen Coder and DeepSeek Coder excel at programming, and Llava handles vision tasks. For RAG and simple applications, 7B models suffice. By paying attention to quantization and VRAM constraints, you’ll quickly find a model that runs smoothly and delivers good results.

Back to Blog
Share:

Related Posts