Skip to content
BotServBotServ
OllamaLocal AIAPIModel ManagementInstallation

Ollama

Runtime environment for open language models. Download, run, chat, and use models via API.

S

schutzgeist

3 min read
Ollama

Ollama

What this article covers

  • What Ollama is and why you need it.
  • Alternative options for running language models locally.
  • How to install Ollama and start your first model.
  • How to use Ollama from the command line and via API.
  • Links to detailed guides.

Introduction

Open language models like Llama, Qwen, and Mistral are freely available. But you still need software to run them. That software is called a runtime or inference engine. Ollama is one such runtime. It downloads model files, manages them, and provides both a chat interface and an API. This means you can chat with models or integrate them into your own applications without relying on cloud services.

Why do you need Ollama?

A large language model is essentially a massive file. To get answers from it, the file must be loaded into what’s called an inference runner. This runner takes your prompt, computes the appropriate response, and outputs it. Ollama bundles together exactly this runner, model management, and an API into a single, easy-to-use tool.

Alternatives include:

  • llama.cpp - Fast C++ runtime, very flexible, but more complex.
  • LM Studio - GUI application with a user-friendly interface.
  • vLLM - High-performance runtime for server deployments.
  • Text Generation Inference (TGI) - From Hugging Face, aimed at production environments.
  • KoboldCpp - Focused on creative writing and roleplay.

Ollama stands out for its simplicity. A single command is enough to get a model running.

Installation

On Linux, install Ollama with a script:

curl -fsSL https://ollama.com/install.sh | sh

macOS and Windows have official installers available on the Ollama website. After installation, Ollama runs as a background service and is accessible from the command line.

Getting started

Download and run a model:

ollama pull llama3.1
ollama run llama3.1

pull downloads the model from the Ollama repository. run starts the chat interface in your terminal. You can start chatting immediately. To end your session, press Ctrl + D or type /bye.

Using Ollama as an API

Ollama provides an API designed like the OpenAI API. This makes switching between services straightforward. Example with curl:

curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1",
    "messages": [
      {"role": "user", "content": "Explain Ollama in three sentences."}
    ]
  }'

The response comes back as JSON. Your own applications, agent frameworks, or chat frontends can use this API directly.

What can Ollama do?

FeatureUse
Model managementDownload, list, and delete models
ChatTalk to models directly in the terminal
APIConnect applications or frameworks
ModelfilesCreate custom model configurations
StorageModels remain local on your computer

When to choose Ollama

  • You want a simple solution for local AI.
  • You need to connect applications to local models via an API.
  • You work with agent frameworks like LangGraph, CrewAI, or AutoGen.
  • Privacy and independence from cloud APIs matter to you.

Content and detailed guides

FAQ - Frequently asked questions

Is Ollama free?

Yes, Ollama itself is open source and free. Costs only come from hardware and electricity.

Can I use Ollama without a graphics card?

Yes, Ollama runs on CPU as well. Large models will be slower, but smaller 7B models are often fast enough.

Is the Ollama API compatible with OpenAI?

Yes, Ollama provides an endpoint at /v1 that’s compatible with the OpenAI Chat API. Many tools and frameworks can use it directly.

Which models work with Ollama?

Many open models like Llama 3.1, Qwen 2.5, Mistral, Code Llama, and many others. You can find the full list in the Ollama library.

Sources

Back to Blog
Share:

Related Posts