DeepSeek Models: Reasoning and Coding from China
What this article covers
- What DeepSeek is and which models make up the family
- How DeepSeek-R1 achieves reasoning on an entirely new level
- Which distilled versions work well for local deployment
- How much VRAM you need for different DeepSeek models
- Where this model family excels and where it falls short
Introduction
DeepSeek is a Chinese AI company that has released some of the most impressive open-source models in a remarkably short time. The model family consists of three main lines: DeepSeek-V3 as a generalist, DeepSeek-R1 as a reasoning specialist, and DeepSeek-Coder for programming tasks. What makes these models special is their architecture as Mixture-of-Experts models (MoE), which combines enormous total parameter counts with efficient execution.
For local deployment, the distilled versions of DeepSeek-R1 are particularly interesting. These are built on smaller base models like Qwen or Llama but have been trained with the reasoning capabilities of DeepSeek-R1. This way you get reasoning quality in sizes that run on consumer hardware. In this article, you’ll learn which DeepSeek models exist, which fit your use case, and how to run them locally with Ollama.
If you’re new to the topic, check out the Model Overview and the Finding Your Model article first.
Why do you need DeepSeek?
Imagine you need to solve a complex math problem or track down a tricky code bug. Standard chat models often give you a quick answer, but for complex tasks it can be wrong or incomplete. DeepSeek-R1 works differently: it thinks out loud, breaks down the problem into intermediate steps, and verifies its own solution before responding. This is called reasoning or chain-of-thought.
If you’ve used Llama models or Qwen models for programming tasks before and found they hit their limits with complex logic, DeepSeek is exactly the next step up. DeepSeek-Coder was trained specifically for coding work and delivers results on many coding benchmarks that compete with commercial models.
For local deployment, the distilled R1 models are especially attractive. A DeepSeek-R1 7B Distill on your graphics card can solve tasks that would otherwise require a 70B model. This saves hardware costs while still delivering better results on logical problems.
DeepSeek models explained briefly
DeepSeek has released several model lines that differ in architecture and purpose. Here’s an overview:
DeepSeek-V3 is the flagship model with 671 billion parameters. It uses a Mixture-of-Experts architecture where only 37 billion parameters are active per token. This makes it more efficient than a comparably large dense model. DeepSeek-V3 is a generalist that shows strong performance across text, code, and mathematics.
DeepSeek-R1 is the reasoning model. It was trained with Reinforcement Learning to solve complex problems step by step. The model outputs a detailed thinking process before providing the final answer. R1 is based on DeepSeek-V3 and achieves results on math and coding benchmarks that rival OpenAI o1.
DeepSeek-R1 Distill are compressed versions where R1’s reasoning capabilities have been transferred to smaller base models. These come in sizes from 1.5B to 70B, based on Qwen and Llama architectures. They are the key models for local deployment.
DeepSeek-Coder is an older model line trained specifically for programming tasks. It comes in sizes from 1.3B to 33B. For new projects, the R1 Distills are usually the better choice, but DeepSeek-Coder remains relevant for pure coding work.
Who this article is for
This article is for you if you:
- are looking for a local model for math or logical tasks
- code and want to run a strong coding model locally
- are interested in reasoning models and want to try chain-of-thought
- need to know which DeepSeek variant will run on your hardware
- already have experience with Ollama and want to try the next level
You don’t need deep prior knowledge. If you know what an LLM is and how to use Ollama, that’s plenty.
Key terms
| Term | Definition |
|---|---|
| MoE | Mixture-of-Experts, an architecture where only a portion of parameters are active per token |
| Reasoning | The ability of a model to solve problems through intermediate steps |
| Chain-of-Thought | A step-by-step thinking process that the model outputs before answering |
| Distill | The transfer of capabilities from a large model to a smaller one |
| Reinforcement Learning | A training method where the model learns through rewards |
| Active Parameters | Parameters that are actually computed per token in an MoE model |
| Dense Model | A model without MoE where all parameters are active at every token |
| Benchmark | A standardized test for comparing model capabilities |
| Token | A text segment processed by the model, roughly a word or part of a word |
| Inference | Running a model to generate text |
DeepSeek-V3: The MoE flagship
DeepSeek-V3 is one of the largest open-source models with 671 billion parameters. The standout feature is its MoE architecture: the model consists of many experts, but only 37 billion parameters are activated per token. This means the compute per token matches a 37B dense model, but the model has the knowledge of a 671B model.
For local deployment on consumer hardware, DeepSeek-V3 isn’t suitable. Even with heavy quantization it needs over 350 GB of storage. That’s only possible on server hardware or Apple Silicon Macs with 128 GB or more Unified Memory. For most users, DeepSeek-V3 is a model to access via API, not locally.
Still, it’s important to understand DeepSeek-V3 because DeepSeek-R1 is built on it. R1’s reasoning capabilities were trained through Reinforcement Learning on the V3 architecture.
DeepSeek-R1: Next-generation reasoning
DeepSeek-R1 is the model that got the AI world’s attention. It solves complex problems not with a quick answer but with a detailed thinking process. The model writes out its thoughts, checks intermediate results, corrects itself, and only then arrives at the final answer.
For example, if you ask R1 how many R’s appear in a word, it doesn’t just return a number. It counts letter by letter, verifies the result, and explains its reasoning. This sounds cumbersome, but it’s exactly why R1 is so strong at math, logic, and coding.
DeepSeek-R1 achieves results on benchmarks like MMLU, HumanEval, and math tasks that compete with commercial models like OpenAI o1. The remarkable part: the model is open source and free to use.
The full R1 model has 671 billion parameters like V3 and is too large for local deployment on consumer hardware. But this is where the distilled versions come in.
DeepSeek-R1 Distill: Reasoning for Consumer Hardware
The distilled versions are the most important part of the DeepSeek family for local users. During distillation, the reasoning capabilities of R1 were transferred to smaller, existing models. The result is compact models with surprisingly good reasoning quality.
The following distill versions are available:
| Model | Base | Parameters | VRAM (Q4_K_M) | Best for |
|---|---|---|---|---|
| DeepSeek-R1 1.5B | Qwen 2.5 | 1.5B | ~1.5 GB | Experiments, very simple tasks |
| DeepSeek-R1 7B | Qwen 2.5 | 7B | ~5 GB | Beginners, simple reasoning tasks |
| DeepSeek-R1 8B | Llama 3.1 | 8B | ~5.5 GB | Beginners, solid all-around quality |
| DeepSeek-R1 14B | Qwen 2.5 | 14B | ~9 GB | Mid-range, solid reasoning |
| DeepSeek-R1 32B | Qwen 2.5 | 32B | ~20 GB | Advanced users, strong reasoning |
| DeepSeek-R1 70B | Llama 3.3 | 70B | ~40 GB | High-end, close to full R1 quality |
The 7B and 8B versions run on GPUs with 8 GB VRAM and offer the best entry point. The 14B version strikes a good balance between performance and hardware requirements. The 32B version needs an RTX 4090 or a Mac with 32 GB unified memory. The 70B version already achieves quality remarkably close to the full R1, but requires appropriate hardware.
For more on memory requirements, see RAM and VRAM Requirements and Model Size and Storage Needs.
DeepSeek-Coder: Programming Locally
DeepSeek-Coder is a dedicated model line trained specifically for programming tasks. Models range from 1.3B to 33B parameters and support numerous programming languages.
For local deployment, the 6.7B and 33B versions are particularly interesting. The 6.7B version runs on an 8 GB GPU and is a solid coding model for everyday tasks. The 33B version requires about 20 GB VRAM in Q4_K_M and delivers significantly better results on complex programming tasks.
If you’re already using R1 distills, however, they cover many coding tasks. DeepSeek-Coder remains relevant if you prefer a pure coding model without R1’s lengthy reasoning output. R1 thinks extensively, which can be unnecessary for quick code questions. DeepSeek-Coder provides more direct answers.
Running DeepSeek Locally with Ollama
The easiest way to run DeepSeek models locally is through Ollama. Here are the essential commands:
# Start DeepSeek-R1 7B Distill
ollama run deepseek-r1:7b
# Start DeepSeek-R1 8B Distill (Llama base)
ollama run deepseek-r1:8b
# Start DeepSeek-R1 14B Distill
ollama run deepseek-r1:14b
# Start DeepSeek-R1 32B Distill
ollama run deepseek-r1:32b
# Start DeepSeek-Coder 6.7B
ollama run deepseek-coder:6.7b
Ollama automatically downloads the Q4_K_M variant, which is the best trade-off between quality and storage for most users. Learn more about quantization in Quantization.
A key characteristic of R1 models: the output includes a thinking block that starts with <think> and ends with </think>. Here the model writes out its reasoning process. The actual answer follows after. In Ollama, you see both parts in the terminal. Some interfaces like Open WebUI display the thinking block separately.
Common Pitfalls
1. Mistaking the thinking output for the answer: DeepSeek-R1 produces a lengthy reasoning process before the answer. If you mistake the thinking block for the answer, you’ll be confused. The actual answer appears after the closing thinking tag.
2. Choosing a model too large for your hardware: The full DeepSeek-R1 with 671B doesn’t run on consumer hardware. Pick a distill version that matches your VRAM instead. An RTX 3060 with 12 GB handles the 7B version with ease.
3. Using reasoning for simple tasks: R1 reasons through every question, even simple ones. If you just want to summarize a short text, a standard model like Llama 3 is faster and more efficient. Use R1 for complex tasks, not everything.
4. Confusing DeepSeek-Coder with R1: DeepSeek-Coder is a separate, older model line. R1 distills are newer and perform better on most tasks. Check which model you actually need.
5. Underestimating context length during reasoning: R1’s thinking process can become very long and consume substantial context. Plan enough VRAM for the KV cache, especially on complex tasks. See more in Context Length.
6. German answers weaker than English: DeepSeek models were trained primarily on English and Chinese. German answers are good but not as strong as models like Qwen, which are optimized for multiple languages. For purely German applications, test Qwen as an alternative.
7. Misunderstanding the MoE architecture: DeepSeek-V3 and R1 have 671B parameters, but only 37B are active per token. This doesn’t mean the model runs as fast as a 37B dense model. All parameters must stay in memory, including inactive ones.
Hardware, Costs, and Security
Hardware: The distilled R1 models are surprisingly frugal. The 7B version runs on an 8 GB GPU, the 14B version needs about 12 GB. For the 32B version, an RTX 4090 with 24 GB or a Mac with 32 GB unified memory is recommended. The 70B version requires 48 GB or more. The full R1 with 671B runs only on server hardware or Macs with 128 GB unified memory.
Costs: All DeepSeek models are open source and free. API usage of the full model is comparatively inexpensive. Local deployment costs only the hardware. An RTX 3060 with 12 GB for under 300 EUR is sufficient for the 7B and 14B distills.
Security: As with all local models, your data stays on your machine. This is especially relevant for DeepSeek because many users have concerns about Chinese models. When running locally, no data is sent to DeepSeek. Still, check whether your software (Ollama, LM Studio) has telemetry features and disable them if needed.
Further Reading
- Model Overview - All model articles at a glance
- Finding Your Model - How to find the right model for you
- Llama Models - Meta’s model family compared
- Qwen Models - Alibaba’s multilingual models
- Ollama - The easiest way to run local models
- Quantization - How models get smaller
- RAM and VRAM Requirements - How much memory you need
- Model Size and Storage Needs - The relationship between size and storage
FAQ
What’s the difference between DeepSeek-R1 and DeepSeek-V3?
DeepSeek-V3 is the base model, a generalist with 671B parameters. DeepSeek-R1 builds on V3 but was trained with Reinforcement Learning for reasoning tasks. R1 works through problems step by step, while V3 answers directly.
Can I run DeepSeek-R1 locally on my PC?
The full R1 with 671B parameters is too large for consumer hardware. The distilled versions from 1.5B to 70B are built for local use and run on standard graphics cards with Ollama.
Which R1 Distill version should I choose?
With 8 GB VRAM, use the 7B version. With 12 GB VRAM, use the 14B version. With 24 GB VRAM, use the 32B version. The 32B version offers the best balance between reasoning quality and hardware requirements.
What does the thinking block in R1 mean?
R1 outputs a reasoning process before the actual answer, wrapped in special tags. The model breaks down the problem into intermediate steps, checks assumptions, and corrects itself. The final answer comes after.
Is DeepSeek-Coder better than DeepSeek-R1 for programming?
For quick coding questions, DeepSeek-Coder is more direct since it skips the long thinking process. For complex programming tasks involving logic, R1 usually performs better. Choose Coder for fast answers and R1 for difficult problems.
Are DeepSeek models free to use?
Yes, all DeepSeek models are open source and free. Model weights are available on Hugging Face. For commercial use, check the respective license; most permit it.
How well do DeepSeek models handle German?
DeepSeek models were trained primarily on English and Chinese. German works but isn’t as strong as models like Qwen, which were specifically optimized for multilingual use. For purely German applications, test Qwen as an alternative.
What does Mixture-of-Experts mean in DeepSeek?
MoE means the model consists of many sub-models (experts), but only the most relevant ones activate per token. DeepSeek-V3 has 671B parameters total, but only 37B are active per token. This saves compute, but all parameters must fit in memory.
Can I use DeepSeek-R1 with LM Studio?
Yes, LM Studio supports DeepSeek models. Search for “deepseek-r1” in LM Studio and select the appropriate Distill version in your desired quantization. LM Studio downloads the GGUF file automatically.
Do I need an Nvidia GPU for DeepSeek?
No. DeepSeek models run on CPU, Apple Silicon, and AMD GPUs. Ollama supports all these platforms. Nvidia GPUs are fastest, but the smaller Distill versions work fine on a Mac with 16 GB Unified Memory.
How do the R1 Distills differ from the base models?
The Distills are built on models like Qwen 2.5 or Llama 3.1 but were fine-tuned with R1’s reasoning capabilities. They behave like R1 with a thinking block, but keep the smaller base model’s architecture and size.
Sources
- DeepSeek official website and model releases
- DeepSeek-R1 Technical Report on arXiv
- DeepSeek-V3 Technical Report on arXiv
- Ollama model library: DeepSeek models
- Hugging Face: DeepSeek model cards and weights
- DeepSeek-Coder GitHub repository


