Ollama
What this article covers
- What Ollama is and why you need it.
- Alternative options for running language models locally.
- How to install Ollama and start your first model.
- How to use Ollama from the command line and via API.
- Links to detailed guides.
Introduction
Open language models like Llama, Qwen, and Mistral are freely available. But you still need software to run them. That software is called a runtime or inference engine. Ollama is one such runtime. It downloads model files, manages them, and provides both a chat interface and an API. This means you can chat with models or integrate them into your own applications without relying on cloud services.
Why do you need Ollama?
A large language model is essentially a massive file. To get answers from it, the file must be loaded into whatβs called an inference runner. This runner takes your prompt, computes the appropriate response, and outputs it. Ollama bundles together exactly this runner, model management, and an API into a single, easy-to-use tool.
Alternatives include:
- llama.cpp - Fast C++ runtime, very flexible, but more complex.
- LM Studio - GUI application with a user-friendly interface.
- vLLM - High-performance runtime for server deployments.
- Text Generation Inference (TGI) - From Hugging Face, aimed at production environments.
- KoboldCpp - Focused on creative writing and roleplay.
Ollama stands out for its simplicity. A single command is enough to get a model running.
Installation
On Linux, install Ollama with a script:
curl -fsSL https://ollama.com/install.sh | sh
macOS and Windows have official installers available on the Ollama website. After installation, Ollama runs as a background service and is accessible from the command line.
Getting started
Download and run a model:
ollama pull llama3.1
ollama run llama3.1
pull downloads the model from the Ollama repository. run starts the chat interface in your terminal. You can start chatting immediately. To end your session, press Ctrl + D or type /bye.
Using Ollama as an API
Ollama provides an API designed like the OpenAI API. This makes switching between services straightforward. Example with curl:
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1",
"messages": [
{"role": "user", "content": "Explain Ollama in three sentences."}
]
}'
The response comes back as JSON. Your own applications, agent frameworks, or chat frontends can use this API directly.
What can Ollama do?
| Feature | Use |
|---|---|
| Model management | Download, list, and delete models |
| Chat | Talk to models directly in the terminal |
| API | Connect applications or frameworks |
| Modelfiles | Create custom model configurations |
| Storage | Models remain local on your computer |
When to choose Ollama
- You want a simple solution for local AI.
- You need to connect applications to local models via an API.
- You work with agent frameworks like LangGraph, CrewAI, or AutoGen.
- Privacy and independence from cloud APIs matter to you.
Content and detailed guides
- Installing Ollama - Step-by-step instructions for Linux, macOS, and Windows.
- Ollama API - Endpoints, examples, and integration into your own applications.
- Managing Ollama models - Pull, list, run, rm, and create custom Modelfiles.
- Ollama on Windows - Installation, GPU configuration, and WSL2 setup.
- Ollama on Linux - Installation with systemd, GPU support for Ubuntu, Debian, and Fedora.
- Ollama on macOS - Apple Silicon, Metal framework, and Unified Memory.
- Ollama with Docker - Containerized deployment with GPU support and docker-compose.
- Ollama configuration - Environment variables, Modelfiles, and performance tuning.
- Ollama network access - Remote access, firewalls, reverse proxies, and security.
- Ollama troubleshooting - Common issues and solutions for GPU, OOM, and more.
- Ollama hardware requirements - VRAM, RAM, and GPU recommendations by model size.
FAQ - Frequently asked questions
Is Ollama free?
Yes, Ollama itself is open source and free. Costs only come from hardware and electricity.
Can I use Ollama without a graphics card?
Yes, Ollama runs on CPU as well. Large models will be slower, but smaller 7B models are often fast enough.
Is the Ollama API compatible with OpenAI?
Yes, Ollama provides an endpoint at /v1 thatβs compatible with the OpenAI Chat API. Many tools and frameworks can use it directly.
Which models work with Ollama?
Many open models like Llama 3.1, Qwen 2.5, Mistral, Code Llama, and many others. You can find the full list in the Ollama library.


