Skip to content
BotServBotServ
OllamaCommandsCLILocal AICheat Sheet

Essential Ollama Commands

Quick overview of Ollama CLI. Download, run, manage models and work with the API.

S

schutzgeist

3 min read
Essential Ollama Commands

Essential Ollama Commands

What this article covers

  • Core commands to get started with Ollama.
  • Downloading, running, and removing models.
  • Viewing information about active and installed models.
  • Creating custom models with Modelfiles.
  • Key API endpoints explained.

Introduction: Essential Ollama Commands

Ollama is primarily controlled via the command line. Whether you need to download a new model, customize an existing one, or check status: the CLI provides straightforward commands for all of these tasks. This article covers the most essential Ollama commands you’ll use in daily work.

For any command, you can get full help with ollama --help or ollama <command> --help.

Key Terms

  • CLI: Ollama’s command-line interface.
  • Model: A downloaded language model.
  • Tag: Version or size identifier, such as :8b or :70b.
  • Modelfile: Recipe for creating or customizing a model.
  • Parameter: Setting like temperature or context length.
  • Serve: Start the Ollama server.

Model Management

CommandPurpose
ollama pull <model>Download or update a model.
ollama run <model>Start a model and chat interactively.
ollama listShow all installed models.
ollama psShow currently loaded models.
ollama rm <model>Remove a model.
ollama cp <source> <target>Copy a model.
ollama show <model>Display model details.

Examples:

ollama pull llama3.1
ollama pull llama3.1:70b
ollama pull qwen2.5:14b
ollama pull nomic-embed-text

Interactive Chat

ollama run llama3.1

After startup, you can enter prompts directly. Exit the chat with /bye or Ctrl+D.

Single Query

ollama run llama3.1 "Explain quantization."

Model Information

ollama show llama3.1
ollama ps

ollama ps shows which model is loaded and whether it’s running on CPU or GPU.

Modelfiles

CommandPurpose
ollama create <name> -f ModelfileCreate a custom model from a Modelfile.
ollama show <model> --modelfileDisplay the Modelfile of an existing model.

Example Modelfile:

FROM llama3.1

SYSTEM """
You are a precise assistant for local AI.
"""

PARAMETER temperature 0.5
PARAMETER num_ctx 4096

Create:

ollama create my-assistant -f Modelfile
ollama run my-assistant

Server Operation

CommandPurpose
ollama serveStart the Ollama server.
ollama serve --host 0.0.0.0Bind to all interfaces.

Typically, Ollama runs as a background service. For manual development, start ollama serve in your terminal.

API Endpoints

EndpointPurpose
POST /api/generateSingle text generation.
POST /api/chatChat with message history.
POST /api/embedGenerate embeddings.
GET /api/tagsList installed models.
POST /v1/chat/completionsOpenAI-compatible chat.

Example:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "Explain Ollama in three sentences."
}'

Useful Options

CommandPurpose
ollama run llama3.1 --verboseDisplay additional information.
ollama run llama3.1 --nowordwrapDisable line wrapping.

Common Command Sequences

Test a new model

ollama pull qwen2.5:14b
ollama run qwen2.5:14b

Copy and customize a model for RAG

ollama cp llama3.1 my-rag-model
ollama create my-rag-model -f Modelfile
ollama run my-rag-model

Free up memory

ollama ps
ollama rm llama3.1:70b
ollama rm older-model

Usage Tips

  • Always use model names with appropriate tags.
  • Check whether a model is in active use before updating.
  • Use ollama ps to verify GPU or CPU operation.
  • Create custom assistants with Modelfiles.

Common Pitfalls

  • Wrong tag: llama3.1 loads 8b, not 70b by default.
  • Model not loaded: Run ollama pull before ollama run.
  • Insufficient VRAM: Model starts slowly or crashes.
  • Server unreachable: Check OLLAMA_HOST and service status.
  • Wrong API URL: OpenAI-compatible endpoint is /v1/chat/completions.

Further Reading and Resources

FAQ: Ollama Commands

How do I see all installed models? Use ollama list.

How do I start Ollama in the background? Run ollama serve or set it up as a Systemd service.

How do I create a custom model? Use a Modelfile and ollama create.

How do I check if a model is running on GPU? Use ollama ps or nvidia-smi.

What’s the difference between pull and run? pull downloads the model, run starts it.

Sources and Further Reading

Summary: Essential Ollama Commands

Ollama is controlled through a straightforward CLI. The most important commands are pull, run, list, ps, rm, and create. Modelfiles let you customize models to your needs. The REST API enables integration into your own applications. Master tags, server mode, and API endpoints, and you can use Ollama effectively for chat, RAG, coding, and embeddings.

Back to Blog
Share:

Related Posts