Skip to content
BotServBotServ
OllamaAPIerrorstroubleshootingHTTP

Ollama API Troubleshooting Guide

Identify and fix common Ollama API errors. Curl commands, logs, status codes, and timeout solutions.

S

schutzgeist

3 min read
Ollama API Troubleshooting Guide

Ollama API Troubleshooting

What this article covers

  • Common API errors and their causes.
  • Reading logs.
  • Testing with curl.
  • Status codes and what they mean.
  • Timeouts, connection issues, and model errors.

Introduction: Ollama API troubleshooting

Anyone working with Ollama’s API will eventually hit an error. Connections fail, responses take forever, or the model refuses to load. Error messages tend to be terse and unhelpful. A few targeted diagnostic steps usually narrow down and resolve most issues quickly.

This article collects typical Ollama API errors and shows how to fix them.

Key terms

  • HTTP status code: Response status from the API.
  • Timeout: Request takes longer than allowed.
  • Connection refused: No service listening on that port.
  • CORS: Cross-origin rules for web requests.
  • Model tag: Model identifier with version.
  • Stream: Response delivered incrementally.
  • Rate limit: Request throttling.
  • Log: Service event log file.

Initial diagnosis with curl

curl http://localhost:11434/api/tags

If you get a JSON list, Ollama is running and responding correctly.

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "Hello",
  "stream": false
}'

Status codes

CodeMeaningSolution
200SuccessEverything is working.
400Bad RequestCheck your JSON or parameters.
404Not FoundModel not installed.
408TimeoutIncrease timeout or check the model.
500Server ErrorCheck Ollama logs.
502/503Service unavailableCheck if Ollama is running.

Error: Model not found

{"error": "model 'llama3.1' not found"}

Solution:

ollama pull llama3.1

Error: Connection refused

curl: (7) Failed to connect to localhost port 11434

Causes:

  • Ollama is not running.
  • Wrong host address.
  • Firewall blocks the port.
  • Ollama is bound to a different port.

Solution:

sudo systemctl status ollama
ollama serve

Error: Timeout

curl: (28) Operation timed out

Causes:

  • Prompt is too large.
  • Model is too large for your hardware.
  • GPU not detected.
  • Context window is too long.

Solutions:

  • Increase the timeout:
curl --max-time 300 ...
  • Test with a smaller model.
  • Limit num_predict.
  • Check VRAM:
nvidia-smi

Error: 404 on /api/generate

Verify you’re using the correct endpoint and HTTP method:

curl -X POST http://localhost:11434/api/generate -d '...'

Error: CORS in the browser

Your frontend reports a CORS error. Solutions:

  • Add a proxy with nginx or Traefik.
  • Configure Ollama hosts.
  • Serve frontend and Ollama from the same origin.

Error: Invalid JSON

Response cannot be parsed. Common with streams:

import json
for line in response.iter_lines():
    if line:
        data = json.loads(line)
        print(data["response"], end="")

Error: Model won’t start

Check:

ollama list
ollama ps

If the model isn’t running, verify:

  • Sufficient RAM and VRAM.
  • Model file isn’t corrupted.
  • Ollama is up to date.

Reading logs

Linux:

journalctl -u ollama -f

macOS:

tail -f ~/.ollama/logs/server.log

Windows: Check Event Viewer or the log file at %USERPROFILE%\.ollama\logs.

Error: GPU not being used

ollama run llama3.1

Check the logs. Possible causes:

  • GPU drivers missing or outdated.
  • ROCR_VISIBLE_DEVICES not set on AMD.
  • Container without GPU passthrough.

Error: Slow responses

  • Verify model size.
  • Check quantization.
  • Switch between CPU and GPU.
  • Set num_thread.
  • Adjust batch size.

Troubleshooting tips

  • Always test with curl first.
  • Monitor logs in parallel.
  • Verify model correctness.
  • Mind timeouts and resource limits.
  • Keep Ollama updated.
  • Check firewall and network settings.

Further reading

FAQ: Ollama API troubleshooting

What’s the default port? 11434.

Why doesn’t Ollama respond? Ollama service isn’t running, the port is blocked, or you’re connecting to the wrong host.

How do I check if a model is loaded? Run ollama ps to see active models.

What should I do about timeouts? Increase the timeout, use a smaller model, or check your GPU.

Where are the logs? Linux: journalctl. macOS: ~/.ollama/logs/server.log.

References

Summary: Ollama API troubleshooting

Most Ollama API errors can be isolated using curl, logs, and status codes. Common problems include models failing to load, connection errors, timeouts, and CORS issues. A systematic diagnostic routine, sufficient resources, and an up-to-date Ollama installation resolve the vast majority of issues. By paying attention to logs and HTTP codes, you’ll usually find the root cause quickly.

Back to Blog
Share:

Related Posts