Skip to content
BotServBotServ
OllamaModelsGalleryOverviewRecommendations

Ollama Model Gallery

Popular Ollama models overview. Chat, coding, vision, reasoning, and embedding models.

S

schutzgeist

2 min read
Ollama Model Gallery

Ollama Model Gallery

What this article covers

  • Overview of popular Ollama models.
  • Organized by use case.
  • Model sizes, quantization options, and hardware requirements.
  • Tips for quick selection.

Ollama provides access to a growing library of local language and multimodal models. Chat models like Llama and Qwen, coding models like Qwen Coder, vision models like Llava, and embedding models like nomic-embed-text are just a few examples. The selection expands rapidly, and multiple options exist for nearly every use case.

This article assembles a gallery of the most important Ollama models.

Key terminology

  • 7B, 13B, 70B: Model size in billions of parameters.
  • Instruct: Trained on instructions.
  • Chat: Optimized for dialogue.
  • Coder: Optimized for programming.
  • Vision: Image processing.
  • Embedding: Vector representation.
  • Quantization: Reduced model precision.
  • GGUF: Container format for models.

Chat models

ModelSizeHighlights
llama3.18B, 70BSolid all-rounder
qwen2.57B, 14B, 32BStrong multilingual support
mistral7BFast, compact
gemma29B, 27BGoogle model
phi414BStrong reasoning

Coding models

ModelSizeUse case
qwen2.5-coder1.5B-32BCode completion
codellama7B-70BMeta code model
deepseek-coder1.3B-33BVersatile coder

Vision models

ModelSizeUse case
llava7B, 13BImage description
llava-llama38BCompact image analysis
bakllava7BBetter vision quality
moondream2BVery small, fast

Reasoning models

ModelSizeUse case
deepseek-r11.5B-70BLogical reasoning
qwen2.57B-72BMath and logic
phi414BSmaller reasoning

Embedding models

ModelDimensionUse case
nomic-embed-text768Standard embeddings
mxbai-embed-large1024High quality
snowflake-arctic-embed768/1024Document retrieval
all-minilm384Very small

Specialized models

ModelUse case
nomic-embed-text-v1.5Improved embeddings
granite3.2-visionIBM vision
command-rCohere instruct
ayaMultilingual

Download a model

ollama pull llama3.1
ollama pull qwen2.5-coder:14b
ollama pull llava
ollama pull nomic-embed-text

Choose a model

Use caseModel
General chatllama3.1, qwen2.5
Codingqwen2.5-coder, deepseek-coder
Visionllava, bakllava, moondream
Reasoningdeepseek-r1, qwen2.5
RAGnomic-embed-text, mxbai-embed-large
Germanqwen2.5, aya

Hardware requirements

Model sizeVRAM Q4RAM CPU
3B2-4 GB6-8 GB
7B4-8 GB12-16 GB
13B8-12 GB20-32 GB
30B18-24 GB40-64 GB
70B40-48 GB96-128 GB

Tips

  • Start with 7B models.
  • Pay attention to quantization options.
  • Use different models for chat and RAG.
  • Test against your own requirements.
  • Compare models on Ollama.com.
  • Prefer updated versions.

Further reading and resources

Which model is best for beginners? llama3.1:8b or qwen2.5:7b.

Are there German models? Yes, Qwen 2.5 and Aya have good multilingual support.

What’s the smallest vision model? moondream:2b.

Which model for coding? qwen2.5-coder:14b or deepseek-coder.

Which models for RAG? A solid chat model plus an embedding model.

Sources and further reading

The Ollama library includes many models for chat, coding, vision, reasoning, and embeddings. Llama, Qwen, and Mistral are solid all-rounders; Qwen Coder and DeepSeek Coder excel at programming; Llava and Moondream handle images; nomic-embed-text powers semantic search. The right choice depends on your hardware, use case, and desired quality. Starting with a typical 7B variant usually works well.

Back to Blog
Share:

Related Posts