Essential Ollama Commands
What this article covers
- Core commands to get started with Ollama.
- Downloading, running, and removing models.
- Viewing information about active and installed models.
- Creating custom models with Modelfiles.
- Key API endpoints explained.
Introduction: Essential Ollama Commands
Ollama is primarily controlled via the command line. Whether you need to download a new model, customize an existing one, or check status: the CLI provides straightforward commands for all of these tasks. This article covers the most essential Ollama commands you’ll use in daily work.
For any command, you can get full help with ollama --help or ollama <command> --help.
Key Terms
- CLI: Ollama’s command-line interface.
- Model: A downloaded language model.
- Tag: Version or size identifier, such as
:8bor:70b. - Modelfile: Recipe for creating or customizing a model.
- Parameter: Setting like temperature or context length.
- Serve: Start the Ollama server.
Model Management
| Command | Purpose |
|---|---|
ollama pull <model> | Download or update a model. |
ollama run <model> | Start a model and chat interactively. |
ollama list | Show all installed models. |
ollama ps | Show currently loaded models. |
ollama rm <model> | Remove a model. |
ollama cp <source> <target> | Copy a model. |
ollama show <model> | Display model details. |
Examples:
ollama pull llama3.1
ollama pull llama3.1:70b
ollama pull qwen2.5:14b
ollama pull nomic-embed-text
Interactive Chat
ollama run llama3.1
After startup, you can enter prompts directly. Exit the chat with /bye or Ctrl+D.
Single Query
ollama run llama3.1 "Explain quantization."
Model Information
ollama show llama3.1
ollama ps
ollama ps shows which model is loaded and whether it’s running on CPU or GPU.
Modelfiles
| Command | Purpose |
|---|---|
ollama create <name> -f Modelfile | Create a custom model from a Modelfile. |
ollama show <model> --modelfile | Display the Modelfile of an existing model. |
Example Modelfile:
FROM llama3.1
SYSTEM """
You are a precise assistant for local AI.
"""
PARAMETER temperature 0.5
PARAMETER num_ctx 4096
Create:
ollama create my-assistant -f Modelfile
ollama run my-assistant
Server Operation
| Command | Purpose |
|---|---|
ollama serve | Start the Ollama server. |
ollama serve --host 0.0.0.0 | Bind to all interfaces. |
Typically, Ollama runs as a background service. For manual development, start ollama serve in your terminal.
API Endpoints
| Endpoint | Purpose |
|---|---|
POST /api/generate | Single text generation. |
POST /api/chat | Chat with message history. |
POST /api/embed | Generate embeddings. |
GET /api/tags | List installed models. |
POST /v1/chat/completions | OpenAI-compatible chat. |
Example:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1",
"prompt": "Explain Ollama in three sentences."
}'
Useful Options
| Command | Purpose |
|---|---|
ollama run llama3.1 --verbose | Display additional information. |
ollama run llama3.1 --nowordwrap | Disable line wrapping. |
Common Command Sequences
Test a new model
ollama pull qwen2.5:14b
ollama run qwen2.5:14b
Copy and customize a model for RAG
ollama cp llama3.1 my-rag-model
ollama create my-rag-model -f Modelfile
ollama run my-rag-model
Free up memory
ollama ps
ollama rm llama3.1:70b
ollama rm older-model
Usage Tips
- Always use model names with appropriate tags.
- Check whether a model is in active use before updating.
- Use
ollama psto verify GPU or CPU operation. - Create custom assistants with Modelfiles.
Common Pitfalls
- Wrong tag:
llama3.1loads8b, not70bby default. - Model not loaded: Run
ollama pullbeforeollama run. - Insufficient VRAM: Model starts slowly or crashes.
- Server unreachable: Check
OLLAMA_HOSTand service status. - Wrong API URL: OpenAI-compatible endpoint is
/v1/chat/completions.
Further Reading and Resources
- BotServ.de Ollama Features
- BotServ.de Ollama Performance
- BotServ.de Ollama REST API
- BotServ.de Ollama Systemd
- BotServ.de Ollama Model Lifecycle
FAQ: Ollama Commands
How do I see all installed models?
Use ollama list.
How do I start Ollama in the background?
Run ollama serve or set it up as a Systemd service.
How do I create a custom model?
Use a Modelfile and ollama create.
How do I check if a model is running on GPU?
Use ollama ps or nvidia-smi.
What’s the difference between pull and run?
pull downloads the model, run starts it.
Sources and Further Reading
- Ollama Library: https://ollama.com/library
- Ollama Docs: https://github.com/ollama/ollama/blob/main/docs/
- Ollama Modelfile: https://github.com/ollama/ollama/blob/main/docs/modelfile.md
Summary: Essential Ollama Commands
Ollama is controlled through a straightforward CLI. The most important commands are pull, run, list, ps, rm, and create. Modelfiles let you customize models to your needs. The REST API enables integration into your own applications. Master tags, server mode, and API endpoints, and you can use Ollama effectively for chat, RAG, coding, and embeddings.


