Phi Models: Microsoft’s Compact AI
What This Article Covers
- Which Phi models exist and how generations 3, 3.5, and 4 differ
- What a Small Language Model is and why compact models matter
- How much VRAM you need for each Phi model
- How to run Phi models locally with Ollama
- The strengths and weaknesses of compact models
Introduction
Phi is Microsoft’s model series built around a specific philosophy: small but capable. While other organizations chase ever-larger models, Microsoft has taken the opposite path. These models are deliberately compact yet deliver performance you’d expect from much larger counterparts. Microsoft calls them “Small Language Models” or SLMs.
If you want to run local AI on limited hardware, Phi models become particularly attractive. They work on laptops without a powerful GPU, on edge devices, and even on smartphones. This article gives you a complete overview of the series, what makes it special, and how to deploy it locally.
Why Use Phi?
Phi models matter for several reasons. First, they’re extremely compact. The smallest variant, Phi-3 mini, has just 3.8 billion parameters and runs on a GPU with 4 GB VRAM. That makes Phi ideal if you don’t have access to expensive hardware.
Second, performance relative to size is remarkable. Microsoft employs a distinctive training approach called “Textbook Quality.” Rather than training on massive amounts of raw internet data, Microsoft uses carefully filtered, high-quality data structured like textbooks. The result: models that often outperform conventionally trained models of the same size.
Third, Phi models cover a wide range of tasks. They handle chat, basic coding, text summarization, and fundamental reasoning. They’re not your first choice for complex problems, but they’re perfectly adequate for everyday work.
Fourth, edge deployment matters. Phi models are small enough to run on mobile devices, in browsers, or on IoT hardware. This opens use cases that larger models simply cannot support.
Phi at a Glance
Phi is a series of Small Language Models developed by Microsoft Research. The models use the Transformer architecture and are built as decoder-only systems. The core difference from other model families lies in training: Microsoft uses high-quality, filtered data instead of raw internet data.
Model names follow a simple pattern: “Phi” followed by version number and size. Phi-3 mini has 3.8 billion parameters, Phi-3 small has 7.4 billion, Phi-3 medium has 14 billion. Generations Phi-3, Phi-3.5, and Phi-4 build on each other, with improvements in each release.
The models are optimized as Instruct versions for conversation and instruction-following. They’re released under the MIT license, one of the most permissive open-source licenses, allowing unrestricted commercial use.
Who Should Read This?
This article is for:
- Beginners looking for a compact model that runs on modest hardware
- Home users wanting to run AI on a laptop without a powerful GPU
- Developers deploying models on edge devices or in browsers
- Budget-conscious users who can’t afford expensive GPU hardware
- Anyone curious about what Small Language Models are and why they matter
Key Terms
| Term | Definition |
|---|---|
| SLM | Small Language Model, a compact language model |
| Parameters | Learned weights of the model, typically expressed in billions (B) |
| Textbook Quality | Training approach using high-quality, structured data |
| Instruct | Version of a model optimized for instruction-following |
| Edge Deployment | Running models on end devices rather than servers |
| Tokenizer | Splits text into tokens for the model to process |
| Context Length | Maximum number of tokens the model can process at once |
| Quantization | Reducing numerical precision to save VRAM |
| GGUF | File format for quantized models, used by Ollama |
| Fine-Tune | Further training a model on specialized data |
Phi Models Overview
Microsoft has evolved the Phi series through multiple generations. Each brings improvements in performance, efficiency, and capability.
Phi-3
Phi-3 is the third generation, released in April 2024. It comes in three sizes:
- Phi-3 mini (3.8B): The generation’s smallest model with 3.8 billion parameters and a context length of 128,000 tokens. Despite its compact size, it matches the performance of models three times larger across many benchmarks. It runs on a 4 GB VRAM GPU in quantized form.
- Phi-3 small (7.4B): A mid-range model with 7.4 billion parameters. It delivers noticeably better performance than mini and compares well with 7B models from other organizations. Context length reaches 128,000 tokens.
- Phi-3 medium (14B): The largest Phi-3 generation model at 14 billion parameters. It offers the best performance of the generation but requires correspondingly more VRAM.
Phi-3.5
Phi-3.5 is an intermediate generation released in August 2024. It includes two models:
- Phi-3.5 mini (3.8B): An improved version of Phi-3 mini with better multilingual support and longer context. It supports 128,000 tokens context length and is trained in multiple languages including German, French, and Spanish.
- Phi-3.5 MoE (16x3.8B): A Mixture-of-Experts model with 16 experts at 3.8 billion parameters each. Each query activates 2 experts. Total parameters reach around 42 billion, but only 6.6 billion are active per query. This is the most capable Phi model but demands significantly more VRAM.
Phi-4
Phi-4 is the latest generation, released in December 2024. It introduces a key innovation: specialized training for mathematics and reasoning.
- Phi-4 (14B): A 14B model with substantially improved capabilities in mathematics, coding, and reasoning. It achieves performance in these areas comparable to models twice its size. Context length is 16,000 tokens. Phi-4 is released under the MIT license.
Technical Specifications
| Model | Parameters | Context Length | Architecture | License |
|---|---|---|---|---|
| Phi-3 mini | 3.8B | 128K | Dense | MIT |
| Phi-3 small | 7.4B | 128K | Dense | MIT |
| Phi-3 medium | 14B | 128K | Dense | MIT |
| Phi-3.5 mini | 3.8B | 128K | Dense | MIT |
| Phi-3.5 MoE | 42B (6.6B active) | 128K | MoE (16x3.8B) | MIT |
| Phi-4 | 14B | 16K | Dense | MIT |
VRAM Requirements for Phi Models
VRAM usage is one of the biggest strengths of Phi models. Because they’re small, they demand far less VRAM than comparable alternatives. The table below shows typical figures for Q4_K_M.
| Model | VRAM at Q4_K_M | VRAM at Q8 | VRAM at FP16 |
|---|---|---|---|
| Phi-3 mini / Phi-3.5 mini | ~2.5 GB | ~4 GB | ~8 GB |
| Phi-3 small | ~5 GB | ~8 GB | ~15 GB |
| Phi-3 medium / Phi-4 | ~9 GB | ~15 GB | ~28 GB |
| Phi-3.5 MoE | ~24 GB | ~45 GB | ~84 GB |
These are guideline values and vary slightly. You’ll also need roughly 1 to 2 GB for context. For more details, see RAM and VRAM Requirements and Model Size and Storage Needs.
Performance and Capabilities
Phi-3 mini and Phi-3.5 mini punch above their weight. At 3.8 billion parameters, they deliver performance you’d expect from 7B to 10B models. They handle chat, basic coding, summarization, and straightforward questions well. Complex reasoning and demanding coding tasks expose their limits.
Phi-3 small sits alongside other 7B models like Llama 3.1 8B or Mistral 7B. Performance is solid, though not remarkable. The appeal lies in efficiency and the MIT license.
Phi-3 medium and Phi-4 bring substantially more capability. Phi-4 excels at math and coding thanks to specialized training. On math benchmarks, it matches results from models twice its size. For a 14B parameter model, that’s exceptional.
Phi-3.5 MoE is the most powerful Phi model. Its MoE architecture with 16 experts delivers performance on par with 70B models while executing faster. VRAM cost is high, though, since all 42 billion parameters must fit in memory.
Why Small Models Matter
Small Language Models like Phi aren’t just weaker versions of their larger cousins. They solve distinct problems and offer advantages big models can’t match.
Speed: Small models generate tokens faster and respond more quickly. For interactive applications like chat, that’s a genuine win.
Hardware efficiency: Small models consume less VRAM and power. They’re ideal for laptops, mobile devices, and edge deployment. You don’t need an expensive graphics card to run Phi.
Cost: Lower hardware requirements mean lower costs. A graphics card with 4 GB VRAM costs a fraction of one with 24 GB. That democratizes local AI for far more users.
On-device privacy: When a model fits on a smartphone, data stays completely local. No cloud, no transmission, no privacy concerns.
Specialization: Small models can be fine-tuned for specific tasks without prohibitive training costs. Fine-tuning Phi-3 mini is feasible on a single GPU; fine-tuning Llama 70B demands substantial resources.
Use Cases for Phi Models
| Use Case | Recommended Model | Why |
|---|---|---|
| Chat on limited hardware | Phi-3 mini | Runs on 4 GB VRAM |
| Math and coding | Phi-4 | Specialized training in these areas |
| Edge deployment | Phi-3.5 mini | Compact and multilingual |
| Laptop without GPU | Phi-3 mini | Acceptable CPU performance |
| Maximum performance | Phi-3.5 MoE | High-capability MoE architecture |
| Commercial projects | All Phi models | MIT license, unrestricted use |
Unsure which model fits your needs? The article Finding the Right Model can help you decide.
Running Phi Models with Ollama
Ollama is the simplest way to run Phi models locally. One terminal command gets you started.
Launch Phi-3 mini:
ollama run phi3
Launch Phi-3.5 mini:
ollama run phi3.5
Launch Phi-4:
ollama run phi4
Ollama automatically downloads an appropriate quantized version, defaulting to Q4_K_M. To specify a different quantization level:
ollama run phi3:3.8b-q5_K_M
Remove models you no longer need with:
ollama rm phi3
See Managing Models for more details.
Common Pitfalls
1. Overestimating performance expectations. Phi models are small, and size comes with trade-offs. Phi-3 mini handles simple tasks well but struggles with complex reasoning or demanding coding work. If you need more power, choose Phi-3 medium or Phi-4.
2. Underestimating Phi-4’s context limit. Phi-4 maxes out at 16,000 tokens, far below Phi-3’s 128,000. For long documents, Phi-4 isn’t the right fit. Use Phi-3.5 mini if you need extended context.
3. Misjudging Phi-3.5 MoE’s memory footprint. Phi-3.5 MoE has 42 billion parameters total. Even quantized, you’ll need roughly 24 GB VRAM. That’s significantly more than other Phi variants. Check your hardware before loading this model.
4. Overestimating German language quality. Phi models are primarily English-trained. Phi-3.5 mini has improved multilingual support, but German doesn’t match Qwen or Mistral. For German-only applications, other models often perform better.
5. Using outdated versions. Phi-1 and Phi-2 are older and substantially weaker. Always reach for Phi-3, Phi-3.5, or Phi-4.
6. Picking the wrong size for the task. Phi-3 mini is fine for simple work but inadequate for complexity. Phi-4 shines at math but falters with long documents. Match the model to your use case, not just its parameter count.
7. Expecting vision capabilities. Phi models are text-only. They can’t process images. For image analysis, use Llama 3.2 Vision or Qwen VL.
Hardware, Cost, and Security
Hardware: Phi models are remarkably hardware-friendly. Phi-3 mini runs on a graphics card with 4 GB VRAM, even an old GTX 1050. Phi-3 medium and Phi-4 need roughly 9 GB VRAM, fitting comfortably on an RTX 3060 with 12 GB. Phi-3.5 MoE requires around 24 GB VRAM and demands an RTX 3090 or 4090. On a Mac with Apple Silicon, Phi-3 mini works with 8 GB RAM, and Phi-4 with 16 GB RAM.
Cost: The models are free under the MIT license, which permits commercial use without restrictions. Hardware costs are minimal. Phi-3 mini needs no specialized graphics card and runs on almost any computer. That makes Phi the most affordable entry point into local AI.
Security: Phi models run locally, keeping your data on your machine. The MIT license gives you full control without restrictions. Bear in mind Phi models lack built-in safety filters. For production applications, add additional safeguards. The advantage of their small size is they can run on devices with no internet connection at all, ensuring maximum data security.
Further Reading
- Local AI Models - Overview of all model articles
- Finding a Model - Help choosing the right model
- Llama Models - Meta’s open source series compared
- Mistral Models - Europe’s alternative compared
- Qwen Models - Alibaba’s multilingual series compared
- Ollama - The simplest way to run local models
- Managing Models - How to manage models in Ollama
- Quantization - How models are made smaller
- RAM and VRAM Requirements - Understanding memory needs
- Model Size and Storage Requirements - Hardware essentials
- What is Local AI? - Foundational article
FAQ
Which Phi model should I try first?
Phi-3 mini is your best starting point. It runs on almost any hardware with 4 GB VRAM and delivers impressive performance for its size. Start with ollama run phi3.
What is a Small Language Model?
A Small Language Model (SLM) is a language model with only a few billion parameters. Unlike Large Language Models (LLMs) with 70 billion or more parameters, SLMs are compact, fast, and hardware-friendly. The Phi series is the most well-known SLM family.
Is Phi-4 better than Phi-3?
Phi-4 is the latest generation and delivers significantly better performance on math, coding, and reasoning tasks. However, it has a shorter context length of 16,000 tokens compared to 128,000 for Phi-3. Choose Phi-4 for math and coding, or Phi-3 and Phi-3.5 if you need to process longer documents.
Can I run Phi on a laptop without a graphics card?
Yes. Phi-3 mini is small enough to run acceptably on a CPU. Speed is lower than on a GPU, but for simple chat tasks it works fine. With Ollama, it just works out of the box with no special configuration needed.
Are Phi models free?
Yes. All Phi models are released under the MIT License, one of the most permissive open source licenses. You can use, modify, and deploy them commercially without any restrictions.
How good is Phi for German?
Phi models are primarily trained on English. Phi-3.5 mini has improved multilingual support, but German performance doesn’t match Qwen or Mistral. For simple German tasks it’s adequate, but for complex text other models are better choices.
What is Phi-3.5 MoE?
Phi-3.5 MoE is a Mixture-of-Experts model with 16 experts of 3.8 billion parameters each. During each request, 2 experts are activated. It’s the most powerful Phi model, but requires roughly 24 GB VRAM when quantized.
Can Phi understand images?
No. All Phi models are text-only. They cannot process images. If you need image analysis, use Llama 3.2 Vision or Qwen VL instead.
Is Phi better than Llama or Mistral?
It depends on your use case. Phi-3 mini is more efficient than Llama 3.1 8B and runs on weaker hardware. For maximum performance, Llama 3.3 70B or Qwen 2.5 72B are better choices. For edge deployment and limited hardware, Phi is the best option. On math tasks, Phi-4 stands out particularly well.
What does Textbook Quality Training mean?
Microsoft trains Phi models on carefully filtered, high-quality data structured like textbooks. Rather than using massive amounts of raw internet data, Microsoft uses curated datasets with clear explanations and examples. This approach allows smaller models to achieve surprisingly strong performance.
Can I use Phi for commercial projects?
Yes, without any restrictions. The MIT License permits commercial use, modification, and distribution without conditions. There are no constraints on user count or revenue, unlike Meta’s Llama license.
Sources
- Microsoft Research - Phi model overview and documentation
- Ollama model library - Phi models
- Hugging Face - Microsoft Phi model page
- Phi-3, Phi-3.5, and Phi-4 technical reports from Microsoft Research


