Qwen3.8-27B Honestly Assessed: Coding Powerhouse, Not an Encyclopedia
What This Article Covers
- The critical point first: where the model’s strength lies and where it failed in community testing.
- What 27 billion parameters can actually do (coding, agentic workflows, vision).
- Comparison with Claude Opus, Sonnet, Haiku, and GPT models: benchmarks and cost.
- When it makes sense to run this locally and when cloud flagships are the better choice.
The Critical Point: Where Does the Focus Lie?
In community tests, Qwen3.8-27B (and its predecessor Qwen3.6-27B) were asked about German federal chancellors. The result: mixed-up order, invented facts, wrong data. Not “I don’t know,” but confidently fabricated answers.
That sounds bad, but it’s only a problem if you misunderstand what the model is designed to do. The 27 billion parameters aren’t focused on world knowledge. Qwen concentrated the model’s capacity almost entirely on coding, agentic work, and multimodal tasks. A dense 27B model simply doesn’t have enough weights to be both a Wikipedia replacement and a coding agent at the same time. Alibaba made a choice: be the best small programmer rather than a mediocre encyclopedia.
If you ask about historical facts, German political chronology, or trivia, you’re testing its weakness. If you give it repository analysis, tool calls, and multi-step code tasks, you’re testing its strength. You should understand this before downloading, because community reports can otherwise be misleading.
What Else the Community Found
The honest list of documented weaknesses:
- Fact hallucination: fabricated historical and political facts, as mentioned above. World knowledge is not the focus.
- Context confabulation: in the Hermes agent test, the model invented an entire conversation history at the start of a new session (documented in GitHub issues).
- Tool-calling loops: reports of repeated fake tool calls in agentic setups, both with official quantizations and community variants.
- Quantization sensitivity: Q4 quantizations degrade noticeably; Qwen models benefit unusually strongly from Q8.
Despite these issues, it sees strong adoption in the community because its coding performance per parameter is hard to beat right now.
What Qwen3.8-27B Actually Is
The facts from the GitHub repo and Hugging Face:
- Dense 27B, natively multimodal (text, images, videos, including long videos)
- 262K token context natively, up to 1M via YaRN
- Apache 2.0: fully open weights, commercially usable
- Thinking mode on by default, with
reasoning_effortandpreserve_thinkingcontrollable - According to the vendor, outperforms the larger Qwen3.7-Plus on coding and office workflows
- Its predecessor Qwen3.6-27B already surpassed its own 397B MoE flagship (Qwen3.5-397B-A17B) on all coding benchmarks, a 27B dense model beating a 15 times larger MoE
The Comparison: Benchmarks Against Claude and Others
Official numbers from Qwen3.6-27B against Claude 4.5 Opus (the 3.8 version supposedly performs even better according to the vendor):
| Benchmark | Qwen3.6-27B | Claude 4.5 Opus | Gemma4-31B |
|---|---|---|---|
| SWE-bench Verified | 77.2 | 80.9 | 52.0 |
| SWE-bench Pro | 53.5 | 57.1 | 35.7 |
| SWE-bench Multilingual | 71.3 | 77.5 | 51.7 |
| Terminal-Bench 2.0 | 59.3 | 59.3 | 42.9 |
| SkillsBench Avg5 | 48.2 | 45.3 | 23.6 |
| LiveCodeBench v6 | 83.9 | 84.8 | 80.0 |
| GPQA Diamond | 87.8 | 87.0 | 84.3 |
| AIME26 (Math) | 94.1 | 95.1 | 89.2 |
Assessment: Opus stays ahead on pure coding (SWE-bench +3.7), but the 27B model ties on Terminal-Bench, beats Opus on SkillsBench (agent tasks) and GPQA (scientific reasoning). For GPT models (GPT-5.2, GPT-4.1), comparable open numbers in the same benchmark framework are missing, but community reports position Qwen3.8-27B in the coding space roughly at GPT-5-mini to GPT-5 level, depending on the task.
Cost: The Real Advantage
| Model | Price (API) | Runs Locally |
|---|---|---|
| Claude Opus 4.5 | ~$15/$75 per Mtok | No |
| Claude Sonnet 4.5 | ~$3/$15 | No |
| Claude Haiku 4.5 | ~$1/$5 | No |
| GPT-5.2 | ~$1.75/$14 | No |
| GPT-5-mini | ~$0.30/$2.40 | No |
| Qwen3.8-27B | ~$0.15-0.60 API, $0 locally | Yes, from ~20 GB VRAM/RAM |
On a Ryzen AI Max+ 395 with 128 GB, it runs comfortably in Q4/Q5; on a mini PC with 64+ GB, the same. That’s the real equation: Opus-level performance on agent tasks, but $0 per token and no data leaving your machine.
When This Model Makes Sense
Choose Qwen3.8-27B if:
- You want a local coding agent (repo-level edits, tool calls, terminal work).
- Privacy matters: your code stays on your hardware.
- You need many tokens: 262K context, free locally.
- Multimodal support: understanding screenshots, diagrams, even videos.
Choose Claude/GPT if:
- You need broad world knowledge (facts, history, politics).
- Peak coding performance matters (Opus still leads SWE-bench).
- You’d rather not run your own hardware.
Real-world tip: most power users combine both. Qwen3.8-27B locally for code work and volume, a cloud model for factual questions and that last 5% of quality. You can configure both in parallel with OpenCode/Aider/OpenChamber, see coding models and OpenChamber.
Further Reading
- IRC-Coding.de: programming tutorials.
- Coding models compared: the full lineup.
- Ollama: run Qwen locally.
- MS-S1 Max: is it worth it: the hardware for 27B models.
- AI for coding projects: agent running continuously.
- ChatGPT prompts: prompting practice.
Key Takeaways:
- Qwen3.8-27B invents facts (chancellor test) because the 27B parameters focus deliberately on coding and agentic work rather than world knowledge. That’s design, not a bug.
- Coding benchmarks: beats its own 397B predecessor, ties with Claude Opus on Terminal-Bench, wins on SkillsBench.
- Apache 2.0, 262K context (1M via YaRN), natively multimodal, runs locally from ~20 GB VRAM.
- Community issues: confabulation and tool loops; higher quantization (Q8) recommended.
- Practice: locally for coding volume and privacy, cloud (Opus/GPT) for world knowledge and peak quality.
FAQ
Why doesn’t Qwen3.8-27B know the German chancellors?
Is Qwen3.8-27B better than Claude?
What hardware does Qwen3.8-27B need locally?
How bad is the hallucination?
How do I run it locally?
Sources and Further Reading
- Qwen3.8-27B on Hugging Face, GitHub
- Qwen3.6-27B Benchmarks (Alibaba Cloud Blog)
- Community: Hallucinations in Qwen3.6-27B (HF Discussion), Hermes Agent Issue
- IRC-Coding.de: programming tutorials.


