Local AI Models
What this article covers
- How to find the right model for local AI.
- Which criteria matter when selecting a model.
- References to dedicated model articles.
Introduction
Not every model is suitable for local deployment. Model size, quantization, and licensing all play a role. This section brings together the knowledge you need to choose the right local LLM for your use case.
Content
- Finding a Model - Selection criteria and sources for suitable local models.
- Llama Models - Meta’s open-source lineup: Llama 3.1, 3.2, 3.3.
- Mistral Models - Europe’s alternative: Mistral 7B, Mixtral MoE.
- Qwen Models - Alibaba’s multilingual series with strong coding capabilities.
- Phi Models - Microsoft’s compact SLMs for edge deployment.
- DeepSeek Models - Reasoning and coding with R1 distills.
- Gemma Models - Google’s lightweight models.
- AI Benchmarks - MMLU, HumanEval, GSM8K and model comparisons.
- Quantization Comparison - Q4 vs Q5 vs Q8 in practice.
Key selection criteria
| Criterion | Impact |
|---|---|
| Model size | Determines RAM or VRAM requirements |
| Quantization | Reduces size, affects output quality |
| License | Governs commercial use |
| Task | Coding, chat, reasoning, or multilingual |
FAQ
What model size do I need?
For most text tasks, 7B to 13B parameter models are sufficient. For complex coding or reasoning tasks, 13B to 70B works better.
Where can I get models?
Ollama, Hugging Face, LM Studio, and models from Kaggle or GitHub are common sources.


