Skip to content
BotServBotServ
LlamaLlama 3.1Llama 3.2MetaOllamalocal AI

Llama Models: Meta's Open-Source Series

Llama 3.1, 3.2, 3.3 overview: sizes, capabilities, hardware requirements, and local use with Ollama.

S

schutzgeist

12 min read
Llama Models: Meta's Open-Source Series

Llama Models: Meta’s Open-Source Series

What This Article Covers

  • Which Llama models exist and how generations 3.1, 3.2, and 3.3 differ
  • The size of each variant and how much VRAM you need for each
  • The strengths and weaknesses of the Llama series
  • How to run Llama models locally with Ollama
  • Common pitfalls when using these models

Introduction

Llama is probably the world’s best-known open-source model series. Developed by Meta, it has fundamentally shaped the market for local AI. Since the first generation in 2023, Meta has continuously released improved versions that have become increasingly powerful and accessible. If you’re interested in local AI, it’s hard to avoid Llama.

The current generations, Llama 3.1, Llama 3.2, and Llama 3.3, cover a broad spectrum: from tiny models for smartphones to massive models that compete with commercial cloud services. This article gives you a complete overview of the series, its variants, and how to deploy them locally.

Why Do You Need Llama?

Llama models are compelling for several reasons. First, they’re freely available for most use cases. You can download them, run them locally, and even use them in commercial projects as long as you respect Meta’s licensing terms. Second, they span an unusually wide range of sizes. From 1 billion parameters up to 405 billion, there’s something for every piece of hardware.

Third, the performance of Llama models is excellent. Llama 3.1 70B and Llama 3.3 70B achieve benchmarks in many tests that rival commercial models like GPT-4. At the same time, Llama 3.2 3B is small enough to run on a laptop. This makes the series one of the most versatile options for local AI.

Fourth, community support plays a major role. Because Llama is so popular, countless quantized versions, fine-tunes, and integration tools exist. If you hit a problem, someone else has probably already solved it and shared their solution.

Llama at a Glance

Llama is a series of Large Language Models developed by Meta and released as open-weight models. “Open weight” means the model weights are freely downloadable, even though the training process itself isn’t fully transparent. The models are built on the Transformer architecture and structured as decoder-only models.

Model names follow a simple scheme: the version number indicates the generation, and the number after the B represents the parameter count in billions. Llama 3.1 8B has 8 billion parameters, Llama 3.1 70B has 70 billion. More parameters generally means better performance, but also higher hardware requirements.

The models are optimized as Instruct versions for dialogue and following instructions. Base versions also exist that simply continue text, but for most users the Instruct versions are what matter.

Who This Article Is For

This article is aimed at:

  • Beginners who want to understand which Llama models exist and which one fits their needs
  • Home users who want to run a Llama model on their own machine
  • Developers who want to integrate Llama into their own applications
  • Decision-makers evaluating Llama as an alternative to cloud AI
  • Anyone seeking a comprehensive overview of the Llama series without wading through technical papers

Key Terminology

TermDefinition
ParameterThe learned values of the model, specified in billions (B)
InstructVersion of a model optimized for following instructions
BaseOriginal version without instruction tuning
TokenizerBreaks text into tokens the model processes
Context LengthMaximum number of tokens the model can process at once
QuantizationReducing numerical precision to save VRAM
GGUFFile format for quantized models, used by Ollama
VRAMGraphics card memory, critical for local AI
Fine-TuneRetraining a model on specialized data
MultimodalModel can process multiple input types, such as text and images

The Llama Generations at a Glance

Meta has developed the Llama series through several iterations. The current generations are Llama 3.1, Llama 3.2, and Llama 3.3. Each generation has its own focus areas and variants.

Llama 3.1

Llama 3.1 is the foundation of the current series and was released in July 2024. It comes in three sizes:

  • Llama 3.1 8B: The smallest model of this generation. It’s fast, uses little VRAM, and is excellent for getting started. With a context length of 128,000 tokens, it can process very long texts.
  • Llama 3.1 70B: The mid-range model. It offers significantly better performance than 8B but requires considerably more hardware. In quantized form, it runs on high-end consumer GPUs or Macs with plenty of RAM.
  • Llama 3.1 405B: The largest model in the series and one of the largest freely available models anywhere. It competes with top commercial models, but running it locally on consumer hardware is unrealistic.

Llama 3.2

Released in September 2024, Llama 3.2 introduces two important innovations: very small models for edge devices and vision capabilities.

  • Llama 3.2 1B: A tiny model with 1 billion parameters. It runs on smartphones and tablets and is designed for simple tasks.
  • Llama 3.2 3B: Slightly larger and more capable than 1B, but still very compact. Ideal for laptops without a strong GPU.
  • Llama 3.2 11B Vision: A multimodal model that can process text and images. It has 11 billion parameters and is the smallest vision model in the series.
  • Llama 3.2 90B Vision: The large vision model. It offers strong image understanding capabilities but requires corresponding hardware.

Llama 3.3

Released in December 2024, Llama 3.3 currently consists of a single variant:

  • Llama 3.3 70B: An improved 70B model that achieves the performance of Llama 3.1 405B with significantly lower hardware requirements. It’s one of the best models in its size class.

Technical Specifications

ModelParametersContext LengthTypeRelease Date
Llama 3.1 8B8B128KTextJuly 2024
Llama 3.1 70B70B128KTextJuly 2024
Llama 3.1 405B405B128KTextJuly 2024
Llama 3.2 1B1B128KTextSeptember 2024
Llama 3.2 3B3B128KTextSeptember 2024
Llama 3.2 11B Vision11B128KMultimodalSeptember 2024
Llama 3.2 90B Vision90B128KMultimodalSeptember 2024
Llama 3.3 70B70B128KTextDecember 2024

All models have a context length of 128,000 tokens, which is more than sufficient for most use cases. For more on why context length matters, see the article on context length.

VRAM Requirements for Llama Models

VRAM usage depends on model size and quantization. The table below shows typical values for the most common quantization level, Q4_K_M.

ModelVRAM at Q4_K_MVRAM at Q8VRAM at FP16
Llama 3.2 1Bapprox. 1 GBapprox. 1.5 GBapprox. 2 GB
Llama 3.2 3Bapprox. 2.5 GBapprox. 4 GBapprox. 6 GB
Llama 3.1 8Bapprox. 5 GBapprox. 9 GBapprox. 16 GB
Llama 3.2 11B Visionapprox. 7 GBapprox. 12 GBapprox. 22 GB
Llama 3.1 70B / 3.3 70Bapprox. 40 GBapprox. 75 GBapprox. 140 GB
Llama 3.2 90B Visionapprox. 52 GBapprox. 95 GBapprox. 180 GB
Llama 3.1 405Bapprox. 230 GBapprox. 430 GBapprox. 810 GB

These figures are estimates and vary slightly depending on the implementation. Beyond the model itself, plan for about 1 to 2 GB for context. For more details, see RAM and VRAM Requirements and Model Size and Storage Needs.

Performance and Speed

Performance varies significantly across Llama model sizes. Here’s a rough breakdown:

Llama 3.2 1B and 3B handle simple tasks well: summarization, straightforward questions, and basic text processing. They’re not suitable for complex reasoning or demanding coding work, but they run almost everywhere.

Llama 3.1 8B is a solid all-purpose model. It manages chat, basic coding, text summarization, and translation. For most home users, this is the best entry point. German quality is solid, though not on par with English output.

Llama 3.2 11B Vision builds on the 8B model by adding image understanding. You can show it a picture and ask questions about it. This works well for descriptions and simple analysis, but struggles with complex visual reasoning tasks.

Llama 3.1 70B and Llama 3.3 70B offer significantly more capability. They excel at complex reasoning, demanding code tasks, and multilingual work. Llama 3.3 70B is the better choice since it achieves the performance of 405B with lower resource requirements.

Llama 3.2 90B Vision is the strongest vision model in the lineup. It can analyze complex images, read charts, and understand visual relationships.

Llama 3.1 405B is the most powerful model in the series, but impractical for local execution on consumer hardware. It’s typically accessed via API or run on server infrastructure.

Use Cases for Llama Models

Use CaseRecommended ModelWhy
Chat on a laptopLlama 3.2 3BSmall, fast, runs everywhere
General-purpose text and codeLlama 3.1 8BBest balance of size and capability
Image analysis and descriptionsLlama 3.2 11B VisionMultimodal, moderate hardware needs
Complex coding and reasoningLlama 3.3 70BVery strong performance with adequate hardware
Multilingual applicationsLlama 3.1 8B or 70BStrong multilingual support from training
Edge deployment on mobileLlama 3.2 1BMinimal footprint

Unsure which model fits your needs? The guide Finding the Right Model can help you decide.

Running Llama Models with Ollama

Ollama is the simplest way to run Llama models locally. You only need a single command in the terminal.

Start Llama 3.1 8B:

ollama run llama3.1

Start Llama 3.2 3B:

ollama run llama3.2:3b

Start Llama 3.2 11B Vision:

ollama run llama3.2-vision

Start Llama 3.3 70B:

ollama run llama3.3

Ollama automatically downloads an appropriate quantized version, defaulting to Q4_K_M. To use a different quantization level, specify it like this:

ollama run llama3.1:8b-q5_K_M

Remove models you no longer need with this command. See Managing Models for more details:

ollama rm llama3.1

Common Pitfalls

1. Choosing a model too large for your hardware: Llama 70B requires about 40 GB VRAM even when quantized. Few consumer GPUs have that much. If you want to run 70B, you’ll need either an RTX 4090 with 24 GB plus system RAM offloading, or a Mac with 64 GB or more.

2. Confusing base and instruct models: Base models continue text but don’t follow instructions. For chat and tasks, you need the instruct version. Ollama uses instruct by default, but when downloading manually, you must select the right variant.

3. Context length exhausts VRAM: 128,000 tokens of context sounds impressive, but fully using that context consumes substantial additional VRAM. With 8B models on 8 GB GPUs, limit context to 4,000 to 8,000 tokens.

4. Overestimating German quality: Llama models are primarily trained on English. German works reasonably well, but complex texts may produce awkward phrasing or anglicisms. For German-focused applications, Mistral or Qwen might be better choices.

5. Testing vision models without images: Llama 3.2 Vision models are optimized for text and images together. Using them for text alone doesn’t improve results over text-only models while consuming more VRAM. Deploy vision models only when you actually process images.

6. Ignoring licensing terms: Llama isn’t entirely free. Meta’s license restricts commercial use, especially at large scale. Review the license before deploying Llama in a commercial product.

7. Using older generations: Llama 2 and Llama 3 (without .1) are outdated and significantly weaker. When running a Llama model, always use the latest generation: 3.1, 3.2, or 3.3.

Hardware, Costs, and Security

Hardware: Needs range dramatically. Llama 3.2 1B runs on nearly any device, Llama 3.1 8B requires a graphics card with 8 GB VRAM or a Mac with 16 GB RAM, and Llama 70B demands high-end hardware. The good news: getting started with an 8B model requires just a consumer GPU like the RTX 3060 with 12 GB.

Costs: The models themselves are free. Costs come from hardware. An RTX 3060 for 8B models costs around 300 euros. Running 70B models requires either high-end hardware in the four-figure range or a Mac Studio with 128 GB RAM. Cloud hosting is an alternative but contradicts the goal of local AI.

Security: Llama models run locally, keeping your data on your machine. This is a major advantage over cloud AI. However, Llama models lack built-in safety filters. They can produce inappropriate content if prompted to do so. For production use, implement additional safety measures.

Further Reading

FAQ

Which Llama model should I try first?

Llama 3.1 8B is the best starting point. It runs on most consumer graphics cards and strikes a good balance between performance and hardware requirements. Start with ollama run llama3.1.

Can I run Llama 70B on my graphics card?

That depends on your GPU. When quantized (Q4_K_M), 70B needs roughly 40 GB VRAM. A single RTX 4090 with 24 GB isn’t enough. You’ll need either two GPUs, one GPU with offloading to system RAM, or a Mac with 64 GB RAM or more.

What’s the difference between Llama 3.1 and Llama 3.3?

Llama 3.3 70B is an improved version of Llama 3.1 70B. It matches the performance of the much larger Llama 3.1 405B while requiring significantly less hardware. If you can run 70B, always choose 3.3 over 3.1.

Are Llama models free?

The model weights are free to download. Meta’s license permits commercial use with restrictions, particularly for companies with over 700 million monthly users. For most users and small businesses, using Llama is free.

Can Llama understand images?

Yes. The Llama 3.2 Vision models (11B and 90B) can process images. You can show them an image and ask questions about it. Pure text models like Llama 3.1 8B cannot process images.

How good is Llama at German?

Llama models handle German well but are trained primarily on English. For simple to moderate tasks, German quality is adequate. Complex texts may produce awkward phrasing. Qwen and Mistral are often better optimized for German.

What does a context length of 128,000 tokens mean?

The model can process up to 128,000 tokens simultaneously, roughly equivalent to 100,000 words. This is sufficient for long documents or extended conversations. Keep in mind that a fully utilized context consumes additional VRAM.

Can I run Llama on a Mac?

Yes. Apple Silicon shares memory between CPU and GPU, which enables running large models. A Mac with 16 GB RAM can run Llama 3.1 8B comfortably. For 70B models, you need at least 64 GB RAM, preferably 128 GB.

What’s the difference between Base and Instruct?

Base models continue text but don’t follow instructions. Instruct models are optimized for dialogue and instruction-following. For chat and most tasks, you need the Instruct version. Ollama uses Instruct by default.

Can I use Llama for my commercial project?

Yes, under certain conditions. Meta’s Llama license permits commercial use but includes restrictions, especially for very large user bases. Check the license terms on Meta’s website before deploying Llama commercially.

Are Llama models better than cloud AI like ChatGPT?

The largest model, Llama 3.1 405B, can match top commercial models. Smaller models are weaker but offer the advantage of local execution, data privacy, and no cost. For most local use cases, 8B or 70B is entirely sufficient.

Sources

  • Meta AI - Llama model overview and documentation
  • Ollama model library - Llama models
  • Hugging Face - Meta Llama model page
  • Meta Llama 3.1, 3.2, and 3.3 technical reports
Back to Blog
Share:

Related Posts