Skip to content
BotServBotServ
OllamatroubleshootingGPUOOMerrorslocal AIdebugging

Ollama Troubleshooting: Solve Common Issues

Fix Ollama problems: GPU not detected, model load failures, out of memory, slow inference. Practical solutions.

S

schutzgeist

9 min read
Ollama Troubleshooting: Solve Common Issues

Ollama Troubleshooting: Solving Common Problems

What this article covers

  • The most common Ollama issues and how to solve them step by step
  • Diagnosing GPU, memory, and connection errors with practical commands
  • Reading and understanding Ollama logs on Linux, Windows, and macOS
  • Tips to help you avoid typical troubleshooting pitfalls
  • Answers to key questions about Ollama troubleshooting

Introduction: Ollama troubleshooting explained

Ollama makes it simple to run local models. But when something breaks, many users face cryptic error messages and don’t know where to start. This article walks you through the most common problems and shows you how to solve them systematically.

You don’t need deep system administration knowledge. Each section explains the problem, its cause, and concrete steps to fix it. If you’re new to Ollama, check out the Ollama overview first.

Why do you need troubleshooting?

Imagine you try to load a model with ollama run llama3 and instead of an answer you get an error message. Or your GPU is ignored and inference crawls along on the CPU. These situations are frustrating, but almost always fixable.

Typical scenarios where troubleshooting comes in handy:

  • A model won’t load and Ollama throws an Out of Memory error
  • The GPU isn’t detected even though drivers are installed
  • The API doesn’t respond because the port is blocked
  • Ollama won’t start at all after an update

The good news: most problems come down to a handful of root causes. With the right diagnosis, you’ll be up and running again quickly.

How Ollama troubleshooting works

Troubleshooting Ollama follows a simple process:

  1. Identify the symptom: What’s actually happening, and what should happen?
  2. Check logs: Ollama writes detailed logs that show you the cause.
  3. Narrow down the cause: Is it GPU, memory, network, or the model?
  4. Apply the fix: Follow the relevant steps from this article.
  5. Verify the result: Test whether the problem is resolved.

This process works for any problem, not just the ones described here.

Who this article is for

This article is for beginners and experienced users running Ollama locally who’ve run into problems. You need basic command-line knowledge, but not system administration expertise. If you haven’t configured Ollama yet, the configuration guide will help.

Key terms for Ollama troubleshooting

TermMeaning
LogA file with system messages that help with troubleshooting
OOMOut of Memory; RAM or VRAM is insufficient
GPUGraphics Processing Unit; accelerates inference
VRAMVideo RAM; the GPU’s memory for model and context
CUDANVIDIA’s platform for GPU computing
DriverSoftware that lets the OS communicate with the GPU
OLLAMA_DEBUGEnvironment variable for more detailed logs
JournalctlLinux tool for reading systemd service logs
Event ViewerWindows tool for viewing system and application logs
CrashAn unexpected program termination

Problem 1: GPU is not detected

Ollama uses the GPU automatically if drivers and runtime are correctly installed. If not, it falls back to the CPU and inference becomes extremely slow.

Diagnosis

First, check whether your GPU is recognized by the system:

nvidia-smi

If this command works, the system sees your GPU. If not, the driver is missing or broken.

Next, check if Ollama sees the GPU. Start Ollama with debug logs:

OLLAMA_DEBUG=1 ollama serve

In the logs, look for lines like GPU 0 or CUDA. If only CPU appears, Ollama isn’t using the GPU.

Solutions

  1. Update drivers: Install the latest NVIDIA driver for your operating system.
  2. Check CUDA: Ollama comes with its own CUDA runtime, but an installed CUDA Toolkit helps with diagnosis.
  3. Adjust OLLAMA_GPU_OVERHEAD: If Ollama detects the GPU but models won’t load, VRAM might be tight. Set OLLAMA_GPU_OVERHEAD to a higher value:
export OLLAMA_GPU_OVERHEAD=2000000000

This reserves extra VRAM for overhead and prevents crashes.

  1. Restart Ollama: After driver updates, restart Ollama so it can re-detect the GPU.

Learn more in the CPU vs. GPU article.

Problem 2: Model won’t load or crashes

A common issue: you try to load a model and get an OOM error or Ollama crashes.

Causes

  • Insufficient VRAM: The model doesn’t fit in graphics memory.
  • Quantization too high: A model with high precision needs more memory.
  • Context too long: A large context uses additional VRAM.

Solutions

  1. Choose a smaller model: If llama3:70b doesn’t work, try llama3:8b.
  2. Reduce quantization: Use a more heavily quantized version. Learn more in the Quantization article.
  3. Adjust context length: Reduce num_ctx in the options to save VRAM. Details in the Context Length article.
  4. Check RAM and VRAM: Compare your memory against requirements in RAM and VRAM Requirements.

Example of a model query with reduced context:

ollama run llama3 --num-ctx 2048

Problem 3: Inference is very slow

If inference takes several seconds per token, Ollama is likely running on the CPU instead of the GPU.

Diagnosis

Start Ollama with debug logs and load a model:

OLLAMA_DEBUG=1 ollama serve

Search the logs for library=cuda or library=cpu. If it says cpu, Ollama isn’t using the GPU.

Solutions

  1. Check GPU offloading: Ollama offloads model layers to the GPU. Verify you have enough VRAM for all layers. Learn more in GPU Offloading.
  2. Force GPU layers: Use --num-gpu to control how many layers go to the GPU:
ollama run llama3 --num-gpu 35
  1. Limit bandwidth contention: If other programs use the GPU, close them to free bandwidth.
  2. Update drivers: Outdated drivers throttle GPU performance.

Problem 4: API connection fails

Ollama serves an API on port 11434. When connections fail, it’s usually due to the port, firewall, or host configuration.

Diagnosis

Test if the API responds locally:

curl http://localhost:11434/api/version

A response with the version number means the API is running. If not, Ollama isn’t started or the port is blocked.

Solutions

  1. Check the port: Make sure port 11434 is available. See what’s occupying it with:
lsof -i :11434
  1. Set OLLAMA_HOST: If you’re accessing from another machine, set OLLAMA_HOST to 0.0.0.0:
export OLLAMA_HOST=0.0.0.0:11434

For details, see Network Access.

  1. Configure your firewall: Open port 11434 in your firewall if you need remote access.
  2. Restart the service: Restart Ollama if the API stops responding.

Problem 5: Model doesn’t answer or produces nonsense

Sometimes a model runs but returns incorrect, meaningless, or missing responses.

Causes

  • Wrong model: Smaller or poorly trained models hallucinate more often.
  • Missing system prompt: Without clear instructions, the model doesn’t know what you expect.
  • Context too long: If context exceeds the limit, Ollama drops older messages, confusing the model.

Solutions

  1. Use a larger model: Switch to one with more parameters if your VRAM allows it.
  2. Set a system prompt: Give the model a clear role and task description.
  3. Adjust context length: Reduce num_ctx if context grows too large.
  4. Redownload the model: If the model file is corrupted, delete and pull it again:
ollama rm llama3
ollama pull llama3

See Managing Models for more.

Problem 6: Disk running out of space

Models are large. A 70B model easily takes 40 GB. Stack multiple models and disk space dwindles fast.

Diagnosis

Check where Ollama stores models. By default:

  • Linux: ~/.ollama/models
  • Windows: C:\Users\YourName\.ollama\models
  • macOS: ~/.ollama/models

Solutions

  1. List and remove models: Show all installed models and delete unused ones:
ollama list
ollama rm unused-model
  1. Change storage location: Set OLLAMA_MODELS to a path with more space:
export OLLAMA_MODELS=/mnt/larger_disk/ollama/models
  1. Monitor disk space: Check regularly, especially if you test new models often.

Problem 7: Ollama won’t start

If Ollama won’t start at all, it’s usually a service, port, or permissions issue.

Diagnosis

Check the service status on Linux:

systemctl status ollama

On Windows and macOS, check the logs for startup errors.

Solutions

  1. Check for port conflicts: If another program uses port 11434, stop it or change the Ollama port.
  2. Verify permissions: The Ollama service needs read access to the model directory. On Linux, the service typically runs as the ollama user.
  3. Reinstall the service: If the service is broken, reinstall it:
sudo systemctl daemon-reload
sudo systemctl restart ollama
  1. Check logs: Look at the logs for the exact error. See the next section for details.

Logs and Diagnostics

Logs are your most important troubleshooting tool. Ollama writes detailed messages that show you the root cause.

Linux

Ollama runs as a Systemd service. Read logs with:

journalctl -u ollama -f

The -f flag follows the log in real time. For more detailed output, set OLLAMA_DEBUG=1 in the service file.

Windows

On Windows, open Event Viewer. Look for events with source ollama or Ollama. You’ll also find logs in %LOCALAPPDATA%\Ollama\.

macOS

On macOS, use the Console app. Filter for ollama. Or in the terminal:

log stream --predicate 'process == "ollama"'

Common troubleshooting pitfalls

  1. Ignoring logs: Many users search blindly instead of reading them. Logs almost always contain the answer.
  2. Outdated GPU drivers: Without current drivers, Ollama uses only CPU. This is a common reason for slow inference.
  3. Overestimating VRAM: Even with plenty of VRAM, context needs extra memory. Plan for overhead.
  4. Wrong environment variables: Variables like OLLAMA_HOST or OLLAMA_MODELS must be set correctly and persist after service restart.
  5. Outdated Ollama version: Bugs in older versions are often fixed. Keep Ollama updated.
  6. Multiple models running: Running models at the same time overwhelms VRAM and bandwidth. Stop unused models.
  7. Firewall forgotten: Remote access is often blocked by the firewall. Check it before digging deeper.

Hardware, costs, and security in troubleshooting

Hardware: You don’t need a special machine for troubleshooting. A system with a GPU, current drivers, and enough VRAM suffices. If you test models often, a GPU with 16 GB VRAM or more is worthwhile.

Costs: Ollama itself is free. Costs come only from hardware if you need to upgrade. A used GPU with 12 to 16 GB VRAM usually works fine for 8B and 13B models.

Security: If you expose Ollama on your network, protect the API. Don’t set OLLAMA_HOST to 0.0.0.0 without configuring a firewall, or anyone on your network can access your models. See Network Access for more.

Further reading and resources

FAQ: Ollama Troubleshooting - Common Questions

Why isn’t my GPU detected? Usually the NVIDIA driver is outdated or CUDA isn’t installed correctly. Check with nvidia-smi that the GPU is recognized, and run Ollama with OLLAMA_DEBUG=1 to see GPU detection in the logs.

What does the OOM error mean? OOM stands for Out of Memory. Your VRAM or RAM is insufficient to load the model. Choose a smaller model, reduce quantization, or lower context length.

Why is inference so slow? Ollama is probably using CPU instead of GPU. Check the logs for library=cuda and make sure the GPU driver is current. If the GPU is detected, check GPU offloading with --num-gpu.

How do I change Ollama’s port? Set the OLLAMA_HOST environment variable to your desired port, for example OLLAMA_HOST=0.0.0.0:8080. Then restart the Ollama service.

Where does Ollama store models? By default under ~/.ollama/models on Linux and macOS, under C:\Users\YourName\.ollama\models on Windows. Change the path with the OLLAMA_MODELS variable.

How do I read Ollama logs on Linux? Use journalctl -u ollama -f to follow logs in real time. For more detailed output, set OLLAMA_DEBUG=1 in the service configuration.

What if Ollama won’t start after an update? Check the service status with systemctl status ollama and look at the logs. Often it’s a port conflict or changed permissions. Reinstalling the service usually helps.

Why does my model hallucinate? Hallucinations often come from models that are too small, missing system prompts, or context that’s too long. Choose a larger model, give clear instructions, and reduce context length.

Can I run multiple models at once? Yes, but it quickly overwhelms VRAM and bandwidth. Stop unused models and start them only when you need them.

How do I reserve extra VRAM for overhead? Set OLLAMA_GPU_OVERHEAD to a higher value, for example OLLAMA_GPU_OVERHEAD=2000000000. This prevents crashes when VRAM runs tight.

References and Further Reading

Back to Blog
Share:

Related Posts