Reasoning Models Compared
What This Article Covers
- What reasoning models are and how they work.
- DeepSeek-R1, QwQ, and other reasoning models side by side.
- When reasoning pays off and when it doesn’t.
- Hardware requirements and speed.
- Real-world examples for math, logic, and complex problems.
Introduction: Understanding Reasoning Models
Standard models answer directly. Reasoning models think first: they show their work (chain-of-thought), check assumptions, correct themselves, and then deliver an answer. This makes them slower, but significantly better for math, logic, and complex problem-solving.
This article is for anyone who wants to understand and use reasoning models. For foundational concepts, see Text Models and Ollama.
Why Use Reasoning Models?
Imagine asking: “If I have 3 apples, eat 2, and buy 5 more, how many do I have?” A standard model might guess wrong. A reasoning model thinks: “3 - 2 = 1, 1 + 5 = 6. Answer: 6.” It shows the work and arrives at the correct answer more reliably.
Reasoning Models Explained Briefly
Reasoning models use chain-of-thought: they work step by step, show their thinking, and then reach a conclusion. DeepSeek-R1 is the most popular local reasoning model. QwQ (Alibaba) is another option. Both run via Ollama.
The core idea: think before you answer.
Who This Article Is For
- Developers solving complex problems.
- Analysts who need data and logic.
- Researchers wanting to understand reasoning.
- Power users demanding maximum quality.
Some LLM background is helpful.
Key Terms
- Reasoning - thinking before answering. Useful for: complex problems.
- Chain-of-Thought - showing your work. Useful for: transparency.
- Self-Correction - fixing your own mistakes. Useful for: accuracy.
- DeepSeek-R1 - local reasoning model. Useful for: the industry standard.
- QwQ - Alibaba’s reasoning model. Useful for: an alternative.
- Ollama - model server. Useful for: running models.
How Reasoning Works
Standard model:
Question → Answer
Reasoning model:
Question → Thinking process → Answer
Example:
Question: "A train goes 120 km/h. How long for 300 km?"
Thinking process:
- Speed = 120 km/h
- Distance = 300 km
- Time = Distance / Speed
- Time = 300 / 120 = 2.5 hours
- Answer: 2.5 hours
Answer: 2.5 hours
Model Comparison
| Model | Developer | Sizes | Reasoning | Speed | Best For |
|---|---|---|---|---|---|
| DeepSeek-R1 | DeepSeek | 1.5B-70B | Excellent | Slow | Math, logic, code |
| QwQ-32B | Alibaba | 32B | Very good | Medium | Reasoning + agents |
| DeepSeek-R1-Distill | DeepSeek | 1.5B-70B | Good | Faster | Compact reasoning |
| Marco-o1 | Alibaba | 7B | Good | Medium | Simpler reasoning |
| Sky-T1 | NovaSky | 32B | Good | Medium | Open reasoning |
Detailed Breakdown
1. DeepSeek-R1
Strengths:
- Best local reasoning model available
- Displays complete thinking process
- Self-corrects while reasoning
- Excellent for math, logic, and complex problems
- Distill versions for lower VRAM
Weaknesses:
- Slow (thinking takes time)
- Overkill for simple questions
- Large models demand significant VRAM
Recommendation: deepseek-r1:14b for solid balance, deepseek-r1:7b for speed.
2. QwQ-32B (Alibaba)
Strengths:
- Excellent reasoning quality
- Supports tool-calling (for agents)
- 32B size: good balance of quality and efficiency
- Faster than DeepSeek-R1 with comparable output
Weaknesses:
- Only 32B available
- Alibaba model (check data privacy policies)
Recommendation: qwq:32b for reasoning plus agent work.
3. DeepSeek-R1-Distill
Strengths:
- Compact versions (1.5B-70B)
- Faster than the original
- Good reasoning for its size
- Llama/Qwen-based (familiar architecture)
Weaknesses:
- Not as good as original R1
- Smaller versions (1.5B) are limited
Recommendation: deepseek-r1:7b for quick reasoning, deepseek-r1:14b for quality.
When to Use Reasoning, When Not to
| Task | Reasoning Needed? | Recommendation |
|---|---|---|
| Math problems | Yes | deepseek-r1 |
| Logic puzzles | Yes | deepseek-r1 |
| Code debugging | Yes | deepseek-r1 |
| Architecture decisions | Yes | deepseek-r1 |
| Simple questions | No | llama3.1 |
| Chat/assistant | No | mistral-nemo |
| Writing text | No | llama3.1 |
| Quick answers | No | small models |
Speed Comparison
| Model | Thinking Time | Response Time | Total |
|---|---|---|---|
| deepseek-r1:7b | 5-15s | 2-5s | 7-20s |
| deepseek-r1:14b | 10-30s | 5-10s | 15-40s |
| qwq:32b | 15-40s | 5-15s | 20-55s |
| llama3.1:8b | , | 2-5s | 2-5s |
Reasoning models run 3-10 times slower than standard models.
Practical Example: Math Problem
# Ollama API with DeepSeek-R1
import requests
response = requests.post("http://ollama:11434/api/chat", json={
"model": "deepseek-r1:14b",
"messages": [
{"role": "user", "content": "A company has 150 employees. 60% work remotely. Of remote workers, 40% use Linux, 35% use Windows, and 25% use Mac. How many use Mac?"}
],
"stream": False
})
# Response includes thinking process + answer
print(response.json()["message"]["content"])
# Output:
# <think>
# 150 * 0.6 = 90 remote
# 90 * 0.25 = 22.5
# ≈ 23 employees use Mac
# </think>
# Answer: approximately 23 employees
Practical Example: Code Debugging
# DeepSeek-R1 for debugging
response = requests.post("http://ollama:11434/api/chat", json={
"model": "deepseek-r1:14b",
"messages": [
{"role": "user", "content": """Debug this code:
def fibonacci(n):
if n <= 1:
return n
return fibonacci(n-1) + fibonacci(n-2)
print(fibonacci(50))
The code runs very slowly. Why and how do I fix it?"""}
],
"stream": False
})
The model reasons through the exponential time complexity and suggests memoization.
Security Notes
- Exposing the thinking process: The reasoning steps may contain sensitive reasoning. Consider whether you want to display it.
- Slow is not always better: For simple questions, reasoning is wasteful.
- Costs: Longer thinking time means more power consumption. For production, weigh cost against benefit.
- Hallucinations: Reasoning models can still get things wrong. Always validate.
Common Pitfalls
- Using it for simple questions: Reasoning is overkill for “What is your name?” Save it for truly complex problems.
- Impatience: Reasoning takes time. Wait for the thinking process.
- Ignoring the thinking process: The reasoning is the value-add. Read the work, not just the answer.
- Running too large a model: 70B reasoning on 8GB of RAM will crash. Use distill versions.
- No fallback plan: If reasoning takes too long, have a fallback to a standard model.
Related Resources
- Text Models - General-purpose models.
- Coding Models - For programming.
- Ollama - Model server.
- AI Agents - Agents with reasoning.
- Chain-of-Thought - Prompting technique.
- VRAM Calculator - Memory requirements.
Key Takeaways:
- Reasoning models think before answering (chain-of-thought).
- deepseek-r1 is the best local reasoning model.
- QwQ is solid for reasoning plus agents.
- Reasoning is slower but better for math, logic, and complex problems.
- For simple questions, standard models are faster and sufficient.
FAQ
What is a reasoning model?
Which reasoning model can I run locally?
When should I use reasoning?
Why are reasoning models slower?
How much VRAM do I need?
What is the thinking process?
Reasoning or standard model?
Can I use reasoning models for agents?
References and Further Reading
- DeepSeek-R1 - Reasoning model.
- QwQ - Alibaba reasoning.
- Chain-of-Thought - Academic paper.
- Ollama - Model server.


