Ollama on Linux: Installation and Setup
What this article covers
- Installing Ollama on Ubuntu, Debian, and Fedora using the official installation script and manually
- Setting up and verifying GPU support for Nvidia CUDA and AMD ROCm
- Understanding systemd services, starting, stopping, and configuring persistent startup
- Downloading your first model and using it via the command line and API
- Identifying and troubleshooting common issues and pitfalls
Introduction: Understanding Ollama on Linux
Ollama is a lightweight tool that lets you run large language models locally on your own machine. On Linux, Ollama runs particularly well and efficiently because it operates as a background service through systemd. Once a model is downloaded, you need no cloud service, no subscription, and no constant internet connection.
This guide walks you through every step required for a working Ollama installation on Linux. I’ve written it with beginners in mind, especially those who’ve never run a local language model before. You’ll get concrete commands, explanations, and guidance whenever something trips you up.
Why do you need this guide?
Imagine wanting to use a language model for private conversations or experimentation without sending your data to a cloud provider. You have a Linux machine with a graphics card and want to get started. Without a structured guide, you’ll quickly face questions like: Where do the models go on disk? How do I start Ollama automatically at boot? And why isn’t the model using the GPU?
That’s exactly what this article addresses. It takes you step by step from installation through GPU setup to API usage, covering every relevant point. By the end, you’ll have a working Ollama installation that starts automatically with your system and is accessible via command line and API.
Ollama on Linux explained
Ollama is an open-source tool that loads, manages, and runs language models locally. On Linux, you install it via a simple shell script that downloads the binary and sets up a systemd service. This service ensures Ollama runs in the background and is accessible through a local HTTP server on port 11434. You interact with Ollama either through the ollama command-line tool or via the REST API.
Models are stored by default in /usr/share/ollama/.ollama/models. You can change this path using the OLLAMA_MODELS environment variable if you want to use a separate disk or partition, for example.
Who this guide is for
This guide is for Linux users installing Ollama for the first time. You should have basic command-line skills, knowing how to open a terminal and run commands. Familiarity with language models or systemd is helpful but not required.
If you’re already using Ollama on Windows or macOS and want to switch to Linux, you’ll find the specific differences and configuration options that Linux offers.
Key terms for Ollama on Linux
| Term | Explanation |
|---|---|
| Ollama | Open-source tool for running language models locally |
| systemd | Init system and service manager on modern Linux distributions |
| curl | Command-line tool for downloading files and making API calls |
| apt | Package manager on Debian and Ubuntu |
| dnf | Package manager on Fedora |
| CUDA | Nvidia platform for GPU-accelerated computing |
| ROCm | AMD platform for GPU-accelerated computing |
| OLLAMA_MODELS | Environment variable for model storage location |
| OLLAMA_HOST | Environment variable for the address Ollama listens on |
| Service | Background process managed by systemd |
| journalctl | Tool for reading systemd logs |
Prerequisites
Before installing Ollama, verify the following:
- Linux distribution: Ubuntu 22.04 or newer, Debian 12 or newer, Fedora 39 or newer. Other distributions work as long as they use systemd.
- RAM: At least 8 GB for small models, 16 GB or more for models with 7 billion parameters or larger.
- Disk space: About 5 GB for Ollama itself, plus 4 to 40 GB per model depending on size.
- GPU drivers: For Nvidia GPUs, you need proprietary drivers installed. For AMD GPUs, you need ROCm.
- Internet connection: Required for installation and initial model download.
Check whether your Nvidia GPU is recognized:
nvidia-smi
If you see a table showing your graphics card and driver version, the driver is properly installed. If you get no output or an error, you need to install Nvidia drivers first.
Step 1: Install Ollama
The easiest method is the official installation script. Open a terminal and run:
curl -fsSL https://ollama.com/install.sh | sh
This script downloads the Ollama binary, sets up the systemd service, and starts Ollama automatically. After installation completes, verify that Ollama is running:
ollama --version
The output shows your installed version. If you get a command not found error, restart your terminal or run source ~/.bashrc.
Manual installation
If you prefer not to use the script or your distribution isn’t supported, you can install Ollama manually:
# Download binary
sudo curl -L https://ollama.com/download/ollama-linux-amd64 -o /usr/local/bin/ollama
sudo chmod +x /usr/local/bin/ollama
# Create user and group
sudo useradd -r -s /bin/false -m -d /usr/share/ollama ollama
After this, you’ll need to create the systemd service file yourself. See the systemd section below for details.
Step 2: Set up GPU support
Nvidia CUDA
If you have an Nvidia GPU and proprietary drivers installed, Ollama detects CUDA automatically. No additional packages needed. After installation, verify that Ollama uses the GPU:
ollama ps
In the output, check the PROCESSOR column for either 100% GPU or a mix of CPU and GPU. If it shows 100% CPU, Ollama isn’t using the GPU. In that case, verify your driver works with nvidia-smi.
AMD ROCm
For AMD graphics cards, Ollama needs ROCm. On Ubuntu and Debian, set up ROCm using AMD’s official installer. Fedora users can find ROCm in the repositories. After installing ROCm, restart the Ollama service:
sudo systemctl restart ollama
Then verify that the GPU is recognized with ollama ps. Not every AMD card is supported by ROCm; older models may only work with CPU.
Step 3: Understanding systemd Services
On Linux, Ollama runs as a systemd service. This means Ollama operates in the background and starts automatically when the system boots. The service file is located at /etc/systemd/system/ollama.service and typically looks like this:
[Unit]
Description=Ollama Service
After=network-online.target
[Service]
ExecStart=/usr/local/bin/ollama serve
User=ollama
Group=ollama
Restart=always
RestartSec=3
Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
[Install]
WantedBy=default.target
The main commands for managing the service are:
# Check service status
sudo systemctl status ollama
# Start the service
sudo systemctl start ollama
# Stop the service
sudo systemctl stop ollama
# Restart the service
sudo systemctl restart ollama
# Enable autostart
sudo systemctl enable ollama
# Disable autostart
sudo systemctl disable ollama
View service logs with journalctl:
# Show the last 50 log entries
sudo journalctl -u ollama -n 50
# Follow logs in real time
sudo journalctl -u ollama -f
Step 4: Running Your First Model
Once installed, you can immediately download and run a model. For getting started, a smaller model like llama3.2 works well:
ollama run llama3.2
Ollama automatically downloads the model before starting it. The download may take several minutes depending on model size and connection speed. Once the model loads, a prompt appears where you can chat directly with it.
Exit the chat by typing /bye. View all downloaded models with:
ollama list
Find additional models in the official registry at ollama.com/library. For more on managing models, see the Model Management guide.
Step 5: Using the API
Ollama exposes a REST API at http://localhost:11434. You can call this API using curl or from any programming language that supports HTTP requests.
Calling the API with curl
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "Erkläre systemd in drei Sätzen.",
"stream": false
}'
The response comes back as a JSON object. Setting "stream": false returns the complete response at once. Without this parameter, the response streams in chunks.
Calling the API with Python
import requests
response = requests.post(
"http://localhost:11434/api/generate",
json={
"model": "llama3.2",
"prompt": "Erkläre systemd in drei Sätzen.",
"stream": False
}
)
print(response.json()["response"])
For detailed documentation of all API endpoints, see the Ollama API guide.
Configuration via Environment Variables
Ollama supports configuration through several environment variables. The most important ones are:
- OLLAMA_MODELS: Sets where models are stored. By default, this is
~/.ollama/modelsfor regular users, or/usr/share/ollama/.ollama/modelsfor the service. - OLLAMA_HOST: Specifies the address and port Ollama listens on. Default is
127.0.0.1:11434. - OLLAMA_ORIGINS: Controls which origins are allowed for CORS, important if you’re calling Ollama from a browser frontend.
Add these variables to the systemd service file by opening it with an editor:
sudo systemctl edit ollama
An empty file opens where you can add override settings. Example:
[Service]
Environment="OLLAMA_MODELS=/mnt/data/ollama-models"
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_ORIGINS=http://localhost:3000,http://localhost:8080"
Save the file and reload systemd:
sudo systemctl daemon-reload
sudo systemctl restart ollama
For comprehensive information on all configuration options, see the Configuration guide.
Ollama on Ubuntu/Debian vs Fedora
The official installation script works the same on Ubuntu, Debian, and Fedora. However, some details differ:
| Aspect | Ubuntu / Debian | Fedora |
|---|---|---|
| Package manager | apt | dnf |
| Nvidia drivers | Via ubuntu-drivers autoinstall or manually | Via dnf install akmod-nvidia |
| ROCm | Via AMD installer | Via dnf install rocm |
| Firewall | ufw (if installed) | firewalld (enabled by default) |
On Fedora, be aware that firewalld may block port 11434 if you want Ollama accessible from the network. Open the port with:
sudo firewall-cmd --permanent --add-port=11434/tcp
sudo firewall-cmd --reload
On Ubuntu with ufw, use:
sudo ufw allow 11434/tcp
For more on Ubuntu basics, see the Ubuntu guide.
Common Pitfalls with Ollama on Linux
-
GPU not detected: Ollama uses only the CPU even though a graphics card is installed. This usually happens due to missing or outdated drivers. Check whether the GPU is recognized by running
nvidia-smiorrocminfo. -
Service won’t start: If
systemctl status ollamashows an error, check the logs withjournalctl -u ollama -n 50. Common causes are incorrect paths in the service file or permission issues. -
Models stored on the wrong disk: By default, Ollama stores models in the service user’s home directory. If space is limited, set
OLLAMA_MODELSto a partition with sufficient storage. -
Port 11434 is already in use: If another service is listening on this port, Ollama won’t start. Change the port via
OLLAMA_HOST, for example to127.0.0.1:11435. -
Permission issues with the service user: The
ollamauser needs read and write permissions in the models directory. If you changeOLLAMA_MODELS, ensure theollamauser has write access. -
CORS errors from the browser: If you’re using a web frontend and need to call Ollama from the browser, set
OLLAMA_ORIGINS. Without this configuration, Ollama blocks browser requests. -
Outdated Ollama version: Older versions may not support certain models. Update Ollama by running the install script again.
For additional solutions to common issues, see the Troubleshooting guide.
Hardware, Costs, and Security with Ollama on Linux
Hardware
For usable performance, a GPU is recommended. An Nvidia GPU with 8 GB VRAM is sufficient for models up to around 7 billion parameters. Larger models require more VRAM, or you can accept slower CPU execution. For more on CPU versus GPU, see the CPU vs. GPU guide.
Costs
Ollama itself is free and open source. Costs come entirely from hardware. If you already have a Linux machine with a suitable graphics card, there are no additional expenses.
Security
By default, Ollama listens only on 127.0.0.1, making it accessible only from your local machine. If you set OLLAMA_HOST to 0.0.0.0, Ollama becomes reachable across your entire network. In that case, you should configure a firewall and restrict access accordingly. Models and conversations remain on your machine; no data is sent to external servers.
For foundational information about running language models locally, see Running LLMs Locally.
Further Reading and Resources for Ollama on Linux
- Ollama Overview - General introduction to Ollama
- Ollama Installation - Installation on other operating systems
- Ollama with Docker - Running Ollama in containers
- Managing Models - Downloading, updating, and deleting models
- Configuration - Complete overview of configuration options
- Troubleshooting - Solutions for common problems
- Ollama API - API endpoints and examples
- Running LLMs Locally - Fundamentals of local language models
- CPU vs. GPU - Hardware basics
- Ubuntu - Ubuntu fundamentals
FAQ: Ollama on Linux - Common Questions
How much RAM do I need for Ollama on Linux?
A minimum of 8 GB for small models like llama3.2 with 3 billion parameters. For models with 7 to 8 billion parameters, 16 GB is recommended. Larger models with 13 billion parameters or more benefit significantly from 32 GB or more.
Where does Ollama store models on Linux?
By default, under /usr/share/ollama/.ollama/models when Ollama runs as a systemd service. You can change the path using the OLLAMA_MODELS environment variable.
Do I need a GPU for Ollama?
No, Ollama runs on CPU as well. However, a GPU significantly accelerates response generation. On CPU, models with 3 to 7 billion parameters remain usable, but larger models become quite slow.
How do I update Ollama on Linux?
Run the install script again: curl -fsSL https://ollama.com/install.sh | sh. Your existing installation will be replaced, but your models and configuration remain intact.
Does Ollama start automatically on boot?
Yes, if the systemd service is enabled. Check with systemctl is-enabled ollama. If it shows enabled, Ollama starts automatically. Otherwise, enable autostart with sudo systemctl enable ollama.
Can I install Ollama without sudo?
The install script requires sudo rights because it copies the binary to /usr/local/bin and creates a systemd service. Manual installation in your home directory without sudo is possible, but you would need to start Ollama manually.
How do I make Ollama accessible across the network?
Set OLLAMA_HOST to 0.0.0.0:11434 in the systemd service file and open the port in your firewall. Keep in mind that Ollama will then be accessible from any machine on your network.
Does Ollama work with AMD graphics cards?
Yes, if ROCm is installed and your graphics card is supported by ROCm. Not every AMD GPU is compatible, and older models may only work on CPU.
How can I see if Ollama is using the GPU?
Start a model and run ollama ps in a second terminal. The PROCESSOR column shows whether the GPU is being used. Alternatively, nvidia-smi displays GPU utilization in real time.
Can I run multiple models simultaneously?
Yes, Ollama can hold multiple models in memory at the same time. Keep in mind that each model consumes VRAM. If VRAM is insufficient, Ollama swaps parts to system RAM, which slows down responses.
How do I uninstall Ollama?
Stop the service, disable it, and remove the files:
sudo systemctl stop ollama
sudo systemctl disable ollama
sudo rm /usr/local/bin/ollama
sudo rm /etc/systemd/system/ollama.service
sudo systemctl daemon-reload
You can also delete the models under /usr/share/ollama/.ollama/models if you no longer need them.
Sources and Further Reading
- Official Ollama documentation:
ollama.com - Ollama GitHub repository:
github.com/ollama/ollama - systemd documentation:
systemd.io - Nvidia CUDA documentation:
docs.nvidia.com/cuda - AMD ROCm documentation:
rocmdocs.amd.com


