Ollama Troubleshooting: Solving Common Problems
What this article covers
- The most common Ollama issues and how to solve them step by step
- Diagnosing GPU, memory, and connection errors with practical commands
- Reading and understanding Ollama logs on Linux, Windows, and macOS
- Tips to help you avoid typical troubleshooting pitfalls
- Answers to key questions about Ollama troubleshooting
Introduction: Ollama troubleshooting explained
Ollama makes it simple to run local models. But when something breaks, many users face cryptic error messages and don’t know where to start. This article walks you through the most common problems and shows you how to solve them systematically.
You don’t need deep system administration knowledge. Each section explains the problem, its cause, and concrete steps to fix it. If you’re new to Ollama, check out the Ollama overview first.
Why do you need troubleshooting?
Imagine you try to load a model with ollama run llama3 and instead of an answer you get an error message. Or your GPU is ignored and inference crawls along on the CPU. These situations are frustrating, but almost always fixable.
Typical scenarios where troubleshooting comes in handy:
- A model won’t load and Ollama throws an Out of Memory error
- The GPU isn’t detected even though drivers are installed
- The API doesn’t respond because the port is blocked
- Ollama won’t start at all after an update
The good news: most problems come down to a handful of root causes. With the right diagnosis, you’ll be up and running again quickly.
How Ollama troubleshooting works
Troubleshooting Ollama follows a simple process:
- Identify the symptom: What’s actually happening, and what should happen?
- Check logs: Ollama writes detailed logs that show you the cause.
- Narrow down the cause: Is it GPU, memory, network, or the model?
- Apply the fix: Follow the relevant steps from this article.
- Verify the result: Test whether the problem is resolved.
This process works for any problem, not just the ones described here.
Who this article is for
This article is for beginners and experienced users running Ollama locally who’ve run into problems. You need basic command-line knowledge, but not system administration expertise. If you haven’t configured Ollama yet, the configuration guide will help.
Key terms for Ollama troubleshooting
| Term | Meaning |
|---|---|
| Log | A file with system messages that help with troubleshooting |
| OOM | Out of Memory; RAM or VRAM is insufficient |
| GPU | Graphics Processing Unit; accelerates inference |
| VRAM | Video RAM; the GPU’s memory for model and context |
| CUDA | NVIDIA’s platform for GPU computing |
| Driver | Software that lets the OS communicate with the GPU |
| OLLAMA_DEBUG | Environment variable for more detailed logs |
| Journalctl | Linux tool for reading systemd service logs |
| Event Viewer | Windows tool for viewing system and application logs |
| Crash | An unexpected program termination |
Problem 1: GPU is not detected
Ollama uses the GPU automatically if drivers and runtime are correctly installed. If not, it falls back to the CPU and inference becomes extremely slow.
Diagnosis
First, check whether your GPU is recognized by the system:
nvidia-smi
If this command works, the system sees your GPU. If not, the driver is missing or broken.
Next, check if Ollama sees the GPU. Start Ollama with debug logs:
OLLAMA_DEBUG=1 ollama serve
In the logs, look for lines like GPU 0 or CUDA. If only CPU appears, Ollama isn’t using the GPU.
Solutions
- Update drivers: Install the latest NVIDIA driver for your operating system.
- Check CUDA: Ollama comes with its own CUDA runtime, but an installed CUDA Toolkit helps with diagnosis.
- Adjust OLLAMA_GPU_OVERHEAD: If Ollama detects the GPU but models won’t load, VRAM might be tight. Set
OLLAMA_GPU_OVERHEADto a higher value:
export OLLAMA_GPU_OVERHEAD=2000000000
This reserves extra VRAM for overhead and prevents crashes.
- Restart Ollama: After driver updates, restart Ollama so it can re-detect the GPU.
Learn more in the CPU vs. GPU article.
Problem 2: Model won’t load or crashes
A common issue: you try to load a model and get an OOM error or Ollama crashes.
Causes
- Insufficient VRAM: The model doesn’t fit in graphics memory.
- Quantization too high: A model with high precision needs more memory.
- Context too long: A large context uses additional VRAM.
Solutions
- Choose a smaller model: If
llama3:70bdoesn’t work, tryllama3:8b. - Reduce quantization: Use a more heavily quantized version. Learn more in the Quantization article.
- Adjust context length: Reduce
num_ctxin the options to save VRAM. Details in the Context Length article. - Check RAM and VRAM: Compare your memory against requirements in RAM and VRAM Requirements.
Example of a model query with reduced context:
ollama run llama3 --num-ctx 2048
Problem 3: Inference is very slow
If inference takes several seconds per token, Ollama is likely running on the CPU instead of the GPU.
Diagnosis
Start Ollama with debug logs and load a model:
OLLAMA_DEBUG=1 ollama serve
Search the logs for library=cuda or library=cpu. If it says cpu, Ollama isn’t using the GPU.
Solutions
- Check GPU offloading: Ollama offloads model layers to the GPU. Verify you have enough VRAM for all layers. Learn more in GPU Offloading.
- Force GPU layers: Use
--num-gputo control how many layers go to the GPU:
ollama run llama3 --num-gpu 35
- Limit bandwidth contention: If other programs use the GPU, close them to free bandwidth.
- Update drivers: Outdated drivers throttle GPU performance.
Problem 4: API connection fails
Ollama serves an API on port 11434. When connections fail, it’s usually due to the port, firewall, or host configuration.
Diagnosis
Test if the API responds locally:
curl http://localhost:11434/api/version
A response with the version number means the API is running. If not, Ollama isn’t started or the port is blocked.
Solutions
- Check the port: Make sure port 11434 is available. See what’s occupying it with:
lsof -i :11434
- Set OLLAMA_HOST: If you’re accessing from another machine, set
OLLAMA_HOSTto0.0.0.0:
export OLLAMA_HOST=0.0.0.0:11434
For details, see Network Access.
- Configure your firewall: Open port 11434 in your firewall if you need remote access.
- Restart the service: Restart Ollama if the API stops responding.
Problem 5: Model doesn’t answer or produces nonsense
Sometimes a model runs but returns incorrect, meaningless, or missing responses.
Causes
- Wrong model: Smaller or poorly trained models hallucinate more often.
- Missing system prompt: Without clear instructions, the model doesn’t know what you expect.
- Context too long: If context exceeds the limit, Ollama drops older messages, confusing the model.
Solutions
- Use a larger model: Switch to one with more parameters if your VRAM allows it.
- Set a system prompt: Give the model a clear role and task description.
- Adjust context length: Reduce
num_ctxif context grows too large. - Redownload the model: If the model file is corrupted, delete and pull it again:
ollama rm llama3
ollama pull llama3
See Managing Models for more.
Problem 6: Disk running out of space
Models are large. A 70B model easily takes 40 GB. Stack multiple models and disk space dwindles fast.
Diagnosis
Check where Ollama stores models. By default:
- Linux:
~/.ollama/models - Windows:
C:\Users\YourName\.ollama\models - macOS:
~/.ollama/models
Solutions
- List and remove models: Show all installed models and delete unused ones:
ollama list
ollama rm unused-model
- Change storage location: Set
OLLAMA_MODELSto a path with more space:
export OLLAMA_MODELS=/mnt/larger_disk/ollama/models
- Monitor disk space: Check regularly, especially if you test new models often.
Problem 7: Ollama won’t start
If Ollama won’t start at all, it’s usually a service, port, or permissions issue.
Diagnosis
Check the service status on Linux:
systemctl status ollama
On Windows and macOS, check the logs for startup errors.
Solutions
- Check for port conflicts: If another program uses port 11434, stop it or change the Ollama port.
- Verify permissions: The Ollama service needs read access to the model directory. On Linux, the service typically runs as the
ollamauser. - Reinstall the service: If the service is broken, reinstall it:
sudo systemctl daemon-reload
sudo systemctl restart ollama
- Check logs: Look at the logs for the exact error. See the next section for details.
Logs and Diagnostics
Logs are your most important troubleshooting tool. Ollama writes detailed messages that show you the root cause.
Linux
Ollama runs as a Systemd service. Read logs with:
journalctl -u ollama -f
The -f flag follows the log in real time. For more detailed output, set OLLAMA_DEBUG=1 in the service file.
Windows
On Windows, open Event Viewer. Look for events with source ollama or Ollama. You’ll also find logs in %LOCALAPPDATA%\Ollama\.
macOS
On macOS, use the Console app. Filter for ollama. Or in the terminal:
log stream --predicate 'process == "ollama"'
Common troubleshooting pitfalls
- Ignoring logs: Many users search blindly instead of reading them. Logs almost always contain the answer.
- Outdated GPU drivers: Without current drivers, Ollama uses only CPU. This is a common reason for slow inference.
- Overestimating VRAM: Even with plenty of VRAM, context needs extra memory. Plan for overhead.
- Wrong environment variables: Variables like
OLLAMA_HOSTorOLLAMA_MODELSmust be set correctly and persist after service restart. - Outdated Ollama version: Bugs in older versions are often fixed. Keep Ollama updated.
- Multiple models running: Running models at the same time overwhelms VRAM and bandwidth. Stop unused models.
- Firewall forgotten: Remote access is often blocked by the firewall. Check it before digging deeper.
Hardware, costs, and security in troubleshooting
Hardware: You don’t need a special machine for troubleshooting. A system with a GPU, current drivers, and enough VRAM suffices. If you test models often, a GPU with 16 GB VRAM or more is worthwhile.
Costs: Ollama itself is free. Costs come only from hardware if you need to upgrade. A used GPU with 12 to 16 GB VRAM usually works fine for 8B and 13B models.
Security: If you expose Ollama on your network, protect the API. Don’t set OLLAMA_HOST to 0.0.0.0 without configuring a firewall, or anyone on your network can access your models. See Network Access for more.
Further reading and resources
- Ollama Overview - Basics and getting started
- Configuration - Set up Ollama correctly
- Network Access - Expose the API securely
- Managing Models - Install and remove models
- Quantization - Shrink models
- RAM and VRAM Requirements - Understand memory needs
- CPU vs. GPU - Why GPU matters
- GPU Offloading - Shift layers to GPU
- Context Length - Keep context and VRAM in check
FAQ: Ollama Troubleshooting - Common Questions
Why isn’t my GPU detected?
Usually the NVIDIA driver is outdated or CUDA isn’t installed correctly. Check with nvidia-smi that the GPU is recognized, and run Ollama with OLLAMA_DEBUG=1 to see GPU detection in the logs.
What does the OOM error mean? OOM stands for Out of Memory. Your VRAM or RAM is insufficient to load the model. Choose a smaller model, reduce quantization, or lower context length.
Why is inference so slow?
Ollama is probably using CPU instead of GPU. Check the logs for library=cuda and make sure the GPU driver is current. If the GPU is detected, check GPU offloading with --num-gpu.
How do I change Ollama’s port?
Set the OLLAMA_HOST environment variable to your desired port, for example OLLAMA_HOST=0.0.0.0:8080. Then restart the Ollama service.
Where does Ollama store models?
By default under ~/.ollama/models on Linux and macOS, under C:\Users\YourName\.ollama\models on Windows. Change the path with the OLLAMA_MODELS variable.
How do I read Ollama logs on Linux?
Use journalctl -u ollama -f to follow logs in real time. For more detailed output, set OLLAMA_DEBUG=1 in the service configuration.
What if Ollama won’t start after an update?
Check the service status with systemctl status ollama and look at the logs. Often it’s a port conflict or changed permissions. Reinstalling the service usually helps.
Why does my model hallucinate? Hallucinations often come from models that are too small, missing system prompts, or context that’s too long. Choose a larger model, give clear instructions, and reduce context length.
Can I run multiple models at once? Yes, but it quickly overwhelms VRAM and bandwidth. Stop unused models and start them only when you need them.
How do I reserve extra VRAM for overhead?
Set OLLAMA_GPU_OVERHEAD to a higher value, for example OLLAMA_GPU_OVERHEAD=2000000000. This prevents crashes when VRAM runs tight.
References and Further Reading
- Ollama GitHub Repository: github.com/ollama/ollama
- Ollama Documentation: ollama.com
- NVIDIA CUDA Toolkit: developer.nvidia.com/cuda-toolkit
- Systemd Journalctl Documentation:
man journalctl - BotServ.de article on Ollama Configuration


