Reasoning Models for Local AI
What this article covers
- What reasoning models are and how they differ from standard LLMs
- How Chain-of-Thought and test-time compute work
- Which open-source reasoning models are available
- How to run them locally with Ollama and vLLM
- Hardware requirements and common pitfalls
Introduction: Reasoning Models for Local AI
Reasoning models don’t just answer questions, they show their work. They use techniques like Chain-of-Thought to break tasks into intermediate steps. This is especially useful for mathematics, logic, code debugging, and planning tasks. Standard language models output answers directly; reasoning models expose their reasoning process.
Local reasoning appeals to privacy-conscious use cases. Mathematical calculations, internal planning, or complex analysis don’t need to go to the cloud. The tradeoff is that reasoning models demand more compute power because they generate longer outputs.
Why use reasoning models?
Reasoning models help when:
- a problem requires multiple steps to solve
- the solution path matters as much as the final answer
- logical deduction is needed
- mathematics or code debugging is involved
- agents need to develop plans
- decision-making needs to be transparent
Reasoning models explained
Reasoning models are language models trained to think step by step. They articulate intermediate steps, verify their own results, and self-correct. This process is called Chain-of-Thought. Some models also use test-time compute, investing more computation during inference to find better answers.
Key concepts:
- Chain-of-Thought: Step-by-step reasoning.
- Test-Time-Compute: Extra computation during generation.
- Self-Correction: Model fixes its own mistakes.
- Reasoning Trace: The visible thinking process.
- Reward Model: Scores individual steps.
- Inference Scaling: Longer or multiple inference passes.
Who should use reasoning models?
- Developers solving complex problems
- Mathematicians and engineers
- Agent developers needing planning capabilities
- Teams requiring transparent decision-making
- Researchers and hobbyists
Key terms in reasoning
- DeepSeek-R1: Popular open-source reasoning model.
- QwQ: Reasoning model from Alibaba.
- o1: OpenAI’s reasoning model, not available locally.
- STaR: Self-Taught Reasoner, a training approach.
- MCTS: Monte Carlo Tree Search for planning.
- Process Reward Model: Rewards good intermediate steps.
Popular reasoning models
DeepSeek-R1
Well-known and strong in mathematics and code. Many variants and quantizations available. Available locally via Ollama. The 32B variant delivers solid results; 70B+ provides very strong performance.
QwQ
Qwen’s reasoning model. Excellent multilingual capabilities, including German. Available in various sizes. A solid alternative to DeepSeek-R1.
Reflection Models
Some models like Reflection include reasoning components. Quality varies; use only with current versions.
Hardware requirements
Reasoning models tend to be large because they generate extended reasoning chains.
- 7B Distill: Runs on 8 GB VRAM, but with shorter reasoning.
- 14B-32B: 16-24 GB VRAM recommended.
- 70B+: Multiple GPUs or a powerful workstation.
Test-time compute means longer generation times. A fast GPU makes a noticeable difference.
Practical example: DeepSeek-R1 with Ollama
ollama pull deepseek-r1:32b
ollama run deepseek-r1:32b
Example prompt:
A train travels at 60 km/h. A second train starts 2 hours later at 90 km/h.
When does the second train catch up to the first? Think step by step.
The model will show its calculation and answer.
Using reasoning in agents
Reasoning models excel at planning tasks in agents. An agent can break down a goal into individual steps, then verify each one. This reduces errors in complex workflows.
Example workflow:
- Define the goal.
- Model plans the steps.
- Execute and validate each step.
- Replan if errors occur.
- Summarize the final result.
Common pitfalls with reasoning models
- High compute cost: Longer answers take more time.
- Excessive reasoning: Model overthinks simple questions.
- Not always correct: Reasoning improves quality but doesn’t guarantee correctness.
- Large models needed: Smaller distilled versions are weaker.
- Variable German quality: Some models optimize heavily for English.
- Difficult to evaluate: Hard to judge whether the reasoning chain makes sense.
Further reading
FAQ: Reasoning Models
Do I always need a reasoning model? No. For simple tasks, standard text models are faster and sufficient.
Are reasoning models slower? Yes, because they generate longer responses and perform more computation.
Which model is best for mathematics? DeepSeek-R1 and QwQ are good choices.
Can I force reasoning with a normal model? Yes, with prompts like “Think step by step.” But dedicated reasoning models usually perform better.
How much VRAM do I need for DeepSeek-R1 32B? Around 20-24 GB, depending on quantization.
Sources and further reading
- DeepSeek-R1: https://www.deepseek.com/
- QwQ: https://qwenlm.github.io/
- Chain-of-Thought: https://arxiv.org/abs/2201.11903
Summary: Reasoning Models for Local AI
Reasoning models think in steps and excel at mathematics, logic, planning, and complex problems. DeepSeek-R1 and QwQ are leading open-source options. Running them locally keeps sensitive data in-house. Critical factors include sufficient VRAM, patience with longer responses, and realistic expectations: reasoning improves quality but doesn’t guarantee error-free results.


