Skip to content
BotServBotServ
MistralMixtralMoEMistral 7BOllamalocal AI

Mistral Models: Europe's Open-Source Alternative

Mistral 7B, Mixtral 8x7B, Mistral Large. Architecture, MoE, hardware requirements, local deployment explained.

S

schutzgeist

12 min read
Mistral Models: Europe's Open-Source Alternative

Mistral Models: Europe’s Open-Source Alternative

What this article covers

  • Which Mistral models exist and how they differ
  • What a Mixture-of-Experts architecture is and why Mixtral uses it
  • How much VRAM you need for each Mistral model
  • How to run Mistral models locally with Ollama
  • The strengths and weaknesses of the model series

Introduction

Mistral AI is a French AI company that has been developing one of the world’s most important open-source model series since 2023. Founded by former researchers from Meta and DeepMind, Mistral quickly proved that a European company with a small team could build competitive models. The first model, Mistral 7B, was a milestone: it was small, fast, and yet more capable than many larger competitors’ models.

If you’re interested in local AI, Mistral models are particularly attractive. They’re efficient, well-trained for European languages, and available in multiple sizes. In this article, you’ll get a complete overview of the model series, its architecture, and how to deploy it locally.

Why do you need Mistral?

Mistral models have several qualities that make them especially interesting for local AI. First, they’re extremely efficient. Mistral 7B often delivers better results at the same size compared to equivalent models. This comes down to thoughtful architecture and careful training.

Second, Mistral models excel with European languages. Since Mistral is a French company, the models were trained multilingually from the start. German, French, Spanish, and Italian tend to work better with Mistral than with models primarily designed for English.

Third, Mistral brings a special architecture to the table with Mixtral 8x7B: Mixture of Experts. This architecture combines multiple sub-models and activates only the relevant ones for each query. The result is a model as capable as a large model but faster in practice.

Fourth, the models are available under permissive licenses. Mistral 7B and Mixtral 8x7B are licensed under Apache 2.0, one of the most open-source licenses available. This means you can use them commercially without worrying about restrictions.

Mistral in a nutshell

Mistral is a series of Large Language Models from French company Mistral AI. The models are based on the Transformer architecture but use various optimizations like Grouped-Query Attention and Sliding Window Attention to improve efficiency.

Model names follow a clear convention: “Mistral” followed by the parameter count denotes dense models where all parameters are active for every query. “Mixtral” refers to Mixture-of-Experts models where only a fraction of parameters activate per query. The designation “8x7B” means the model has 8 experts with 7 billion parameters each, of which 2 are active per query.

Beyond freely available models, Mistral also offers commercial models like Mistral Large, which are only available through API. For local deployment, the open-weight models are what matter.

Who is this article for?

This article is for:

  • Beginners who want to understand which Mistral models exist and how they work
  • Home users seeking an efficient model for European languages
  • Developers wanting to integrate Mistral into their applications
  • Anyone interested in MoE architectures without diving through research papers
  • Users looking for alternatives to Llama

Key terminology

TermMeaning
ParameterLearned values in the model, counted in billions (B)
MoEMixture of Experts, architecture with multiple sub-models
ExpertA sub-model within a MoE architecture
RouterComponent that decides which experts to activate
Dense ModelModel where all parameters are active for every query
Sliding Window AttentionTechnique for efficiently processing context
Grouped-Query AttentionOptimization that saves VRAM and compute time
QuantizationReducing numerical precision to save VRAM
GGUFFile format for quantized models, used by Ollama
Context LengthMaximum number of tokens the model can process simultaneously

Mistral models at a glance

Mistral AI has released several models targeting different audiences and use cases. Here are the key models for local deployment.

Mistral 7B

Mistral 7B was Mistral AI’s first model, released in September 2023. It has 7 billion parameters and is a dense model where all parameters are active for every query. Despite its small size, it outperformed the significantly larger Llama 2 13B on many benchmarks.

Mistral 7B’s strengths lie in efficiency and speed. It runs on graphics cards with 8 GB VRAM in quantized form, making it one of the most accessible models for getting started. Multilingual support is solid, especially for European languages.

Mistral 7B is licensed under Apache 2.0, meaning you’re free to use, modify, and commercially deploy it.

Mixtral 8x7B

Mixtral 8x7B was released in December 2023 and is Mistral’s first MoE model. It has 8 experts with 7 billion parameters each. For each query, the router activates 2 of the 8 experts based on the input text. The total parameter count is roughly 47 billion, but only about 13 billion are active per query.

This means: Mixtral 8x7B is as capable as a large model but as fast as a 13B model. In benchmarks, it reaches the performance level of Llama 2 70B while requiring less compute per token.

VRAM requirements are high, however. Since all 8 experts must be held in memory, you need significantly more VRAM than a simple 7B model even in quantized form. Mixtral 8x7B is also available under Apache 2.0.

Mistral Nemo

Mistral Nemo is a further development created in collaboration with Nvidia. It has 12 billion parameters and is a dense model. Nemo offers a context length of 128,000 tokens and is particularly efficient. It’s a good choice if you need a mid-size model that offers more power than 7B but requires less VRAM than Mixtral.

Mistral Large

Mistral Large is Mistral AI’s flagship but is not available as an open-weight model. It can only be accessed through Mistral’s API and is not relevant for local deployment. It’s mentioned here only for completeness.

Mixture of Experts: How MoE works

MoE is an architecture that splits a model into multiple specialized sub-models called experts. Rather than using all parameters for every query, a router selects the most suitable experts for each token.

Think of it as having a team of 8 specialists: one for math, one for programming, one for history, and so on. When you ask a question, a coordinator routes the question to the two specialists best suited to answer it. The other six remain unused and consume no compute. That’s how MoE works.

The advantage is clear: the model has more knowledge and capability overall because it has more parameters. But execution is faster because only a subset of parameters activates. The downside is that all experts must be held in VRAM. Memory requirements depend on the total parameter count, not the active count.

For Mixtral 8x7B, that’s 47 billion parameters total that must fit in memory. Quantized at Q4_K_M, you need about 26 GB of VRAM. That’s significantly more than a 7B model but less than a 70B model.

Technical Specifications

ModelParametersActive ParametersContext LengthArchitectureLicense
Mistral 7B7B7B32KDenseApache 2.0
Mixtral 8x7B47B13B32KMoE (8x7B)Apache 2.0
Mistral Nemo12B12B128KDenseApache 2.0
Mistral Largeproprietaryproprietary128KDensecommercial

VRAM Requirements for Mistral Models

VRAM usage depends on model size and quantization. With MoE models, keep in mind that all experts must fit in memory, even if only a few are active at any time.

ModelVRAM at Q4_K_MVRAM at Q8VRAM at FP16
Mistral 7B~5 GB~8 GB~14 GB
Mistral Nemo 12B~8 GB~13 GB~24 GB
Mixtral 8x7B~26 GB~49 GB~90 GB

These figures are approximate and vary slightly depending on implementation. Beyond the model itself, plan for 1 to 2 GB for context. For more details, see RAM and VRAM Requirements and Model Size and Storage Needs.

Performance and Capabilities

Mistral 7B is remarkably powerful for its size. It outperforms Llama 2 13B on many benchmarks and makes an excellent entry point. It handles chat, text summarization, and straightforward coding tasks well. Complex reasoning tasks expose its limitations, which is expected from a 7-billion-parameter model.

Mixtral 8x7B delivers significantly better performance, matching Llama 2 70B. It tackles complex reasoning, multilingual communication, and sophisticated coding tasks. The MoE architecture makes it faster at inference than an equivalent dense model of the same total parameter count. German language quality is solid, making Mistral a compelling alternative to Llama.

Mistral Nemo sits between 7B and Mixtral. With 12 billion parameters, it offers more capability than Mistral 7B while consuming less VRAM than Mixtral. The 128,000-token context window makes it ideal for processing lengthy documents.

Use Cases for Mistral Models

Use CaseRecommended ModelWhy
Getting started with local AIMistral 7BSmall, fast, Apache 2.0 license
Multilingual chatMistral 7B or NemoStrong European language support
Complex reasoningMixtral 8x7BHigh performance via MoE
Processing long documentsMistral Nemo128K context length
Commercial projectsMistral 7B or MixtralApache 2.0 with no usage restrictions
Resource-constrained deploymentMistral 7BJust 5 GB VRAM at Q4

If you’re unsure which model fits your needs, check out Finding the Right Model for guidance.

Running Mistral Models with Ollama

Ollama is the simplest way to run Mistral models locally. You only need a single command in the terminal.

Start Mistral 7B:

ollama run mistral

Start Mixtral 8x7B:

ollama run mixtral

Start Mistral Nemo:

ollama run mistral-nemo

Ollama automatically downloads a suitable quantized version, defaulting to Q4_K_M. If you want a different quantization level, you can specify it:

ollama run mistral:7b-q5_K_M

To remove models you no longer need, use this command. See Managing Models for more details:

ollama rm mistral

Common Pitfalls

1. MoE models require more VRAM than expected: Mixtral 8x7B has 47 billion parameters total. Even though only 13 billion are active per request, all 47 billion must fit in VRAM. Quantized, you’ll need roughly 26 GB, requiring an RTX 3090 or 4090.

2. Mistral 7B has limited context length: Mistral 7B supports a native context length of 32,000 tokens. Compare this to Llama 3.1 with 128,000 tokens. For very long documents, Mistral Nemo or another model with larger context windows is the better choice.

3. Confusing open-weight with commercial models: Not all Mistral models are freely available. Mistral Large and Mistral Medium are API-only. For local deployment, only Mistral 7B, Mixtral 8x7B, and Mistral Nemo are relevant.

4. MoE inference on CPU is slow: MoE models are optimized for GPUs. Running Mixtral 8x7B on a CPU without GPU acceleration is significantly slower than an equivalent dense model. If you lack a strong GPU, use Mistral 7B instead.

5. Wrong expectations about MoE speed: MoE models are faster than dense models of equal total parameter count, but not faster than small dense models. Mixtral 8x7B is faster than a 47B dense model, but not as fast as Mistral 7B.

6. Using outdated versions: Mistral 7B v0.1 was the first release. Newer versions like v0.2 and v0.3 include improvements. When using Mistral 7B, ensure you have the latest version. Ollama provides the newest version by default.

7. Overlooking license terms for fine-tunes: Mistral 7B and Mixtral use Apache 2.0, but third-party fine-tunes may carry different licenses. Always verify the license of the specific model you download.

Hardware, Costs, and Security

Hardware: Mistral 7B is one of the most hardware-friendly models available. It runs on a GPU with 8 GB VRAM at Q4_K_M. Mixtral 8x7B demands much more: roughly 26 GB VRAM quantized, requiring an RTX 3090 or 4090. Mistral Nemo falls in between at roughly 8 GB VRAM at Q4_K_M. On a Mac with Apple Silicon, Mistral 7B runs smoothly with 16 GB RAM, while Mixtral needs at least 32 GB.

Costs: The open-weight models are free. The expense lies in hardware. For Mistral 7B, a budget card like the RTX 3060 with 12 GB suffices. Mixtral requires high-end hardware in the 1,500 euro range or higher. Mistral Large via API charges per token, which can become costlier than local operation at scale.

Security: Like all local models, your data stays on your machine. Mistral models lack built-in safety filters, so they can produce inappropriate content. The Apache 2.0 license gives you full control over usage. For production systems, implement additional safeguards like input filtering or output validation.

Further Reading

FAQ

Which Mistral model should I try first?

Mistral 7B is the best starting point. It runs on most consumer GPUs with 8 GB VRAM and offers solid performance. Start with ollama run mistral.

What does Mixtral 8x7B mean?

The name indicates the model has 8 experts, each with 7 billion parameters. For each request, 2 of the 8 experts are activated. The total parameter count is roughly 47 billion, but only 13 billion are active per request.

Does Mixtral require more VRAM than a 7B model?

Yes, significantly more. Since all 8 experts must be kept in memory, you’ll need around 26 GB VRAM at Q4_K_M quantization. That’s about five times what Mistral 7B needs.

Are Mistral models free?

Mistral 7B, Mixtral 8x7B, and Mistral Nemo are licensed under Apache 2.0 and are free, including for commercial use. Mistral Large and Mistral Medium are commercial models available only via API.

Is Mistral better than Llama?

It depends on your use case. Mistral 7B is more efficient than Llama 3.1 8B at comparable performance. Mixtral 8x7B is roughly equivalent to Llama 3.1 70B. For European languages, Mistral is often the better choice. If you need maximum performance, Llama 3.3 70B is stronger.

How good is Mistral for German?

Mistral models are well-trained on European languages, including German. German quality often exceeds Llama models, especially for natural phrasing. For German-focused applications, Mistral is a solid choice.

What’s the difference between Dense and MoE?

In a dense model, all parameters are active for every request. In a MoE model, multiple experts exist but only a subset activate per request. MoE models execute faster, but still require the same VRAM as a dense model of equivalent total size.

Can I run Mixtral on an RTX 3060 with 12 GB?

No. Mixtral 8x7B needs around 26 GB VRAM even when quantized. An RTX 3060 with 12 GB is insufficient. You’ll need at least an RTX 3090 or 4090 with 24 GB, and even then it’ll be tight. Alternatively, use Mistral 7B.

Can I use Mistral for commercial projects?

Yes. Mistral 7B, Mixtral 8x7B, and Mistral Nemo are under Apache 2.0, which explicitly permits commercial use. There are no restrictions on user count or revenue, unlike Meta’s Llama license.

What is Mistral Nemo and who is it for?

Mistral Nemo is a 12B model developed in collaboration with Nvidia. It offers a context length of 128,000 tokens and positions itself between Mistral 7B and Mixtral. It’s ideal for users needing more capability than 7B but unwilling to pay the VRAM cost of Mixtral.

Are there Mistral vision models?

Mistral released Pixtral, a vision model that processes images. It’s not as widely available as the text models. For text-only applications, Mistral 7B, Nemo, or Mixtral are the better choice.

Sources

  • Mistral AI - Model overview and documentation
  • Ollama Model library - Mistral models
  • Hugging Face - Mistral AI model page
  • Mistral AI Technical reports on Mistral 7B and Mixtral 8x7B
Back to Blog
Share:

Related Posts