Ollama API Troubleshooting
What this article covers
- Common API errors and their causes.
- Reading logs.
- Testing with
curl. - Status codes and what they mean.
- Timeouts, connection issues, and model errors.
Introduction: Ollama API troubleshooting
Anyone working with Ollama’s API will eventually hit an error. Connections fail, responses take forever, or the model refuses to load. Error messages tend to be terse and unhelpful. A few targeted diagnostic steps usually narrow down and resolve most issues quickly.
This article collects typical Ollama API errors and shows how to fix them.
Key terms
- HTTP status code: Response status from the API.
- Timeout: Request takes longer than allowed.
- Connection refused: No service listening on that port.
- CORS: Cross-origin rules for web requests.
- Model tag: Model identifier with version.
- Stream: Response delivered incrementally.
- Rate limit: Request throttling.
- Log: Service event log file.
Initial diagnosis with curl
curl http://localhost:11434/api/tags
If you get a JSON list, Ollama is running and responding correctly.
curl http://localhost:11434/api/generate -d '{
"model": "llama3.1",
"prompt": "Hello",
"stream": false
}'
Status codes
| Code | Meaning | Solution |
|---|---|---|
| 200 | Success | Everything is working. |
| 400 | Bad Request | Check your JSON or parameters. |
| 404 | Not Found | Model not installed. |
| 408 | Timeout | Increase timeout or check the model. |
| 500 | Server Error | Check Ollama logs. |
| 502/503 | Service unavailable | Check if Ollama is running. |
Error: Model not found
{"error": "model 'llama3.1' not found"}
Solution:
ollama pull llama3.1
Error: Connection refused
curl: (7) Failed to connect to localhost port 11434
Causes:
- Ollama is not running.
- Wrong host address.
- Firewall blocks the port.
- Ollama is bound to a different port.
Solution:
sudo systemctl status ollama
ollama serve
Error: Timeout
curl: (28) Operation timed out
Causes:
- Prompt is too large.
- Model is too large for your hardware.
- GPU not detected.
- Context window is too long.
Solutions:
- Increase the timeout:
curl --max-time 300 ...
- Test with a smaller model.
- Limit
num_predict. - Check VRAM:
nvidia-smi
Error: 404 on /api/generate
Verify you’re using the correct endpoint and HTTP method:
curl -X POST http://localhost:11434/api/generate -d '...'
Error: CORS in the browser
Your frontend reports a CORS error. Solutions:
- Add a proxy with nginx or Traefik.
- Configure Ollama hosts.
- Serve frontend and Ollama from the same origin.
Error: Invalid JSON
Response cannot be parsed. Common with streams:
import json
for line in response.iter_lines():
if line:
data = json.loads(line)
print(data["response"], end="")
Error: Model won’t start
Check:
ollama list
ollama ps
If the model isn’t running, verify:
- Sufficient RAM and VRAM.
- Model file isn’t corrupted.
- Ollama is up to date.
Reading logs
Linux:
journalctl -u ollama -f
macOS:
tail -f ~/.ollama/logs/server.log
Windows: Check Event Viewer or the log file at %USERPROFILE%\.ollama\logs.
Error: GPU not being used
ollama run llama3.1
Check the logs. Possible causes:
- GPU drivers missing or outdated.
ROCR_VISIBLE_DEVICESnot set on AMD.- Container without GPU passthrough.
Error: Slow responses
- Verify model size.
- Check quantization.
- Switch between CPU and GPU.
- Set
num_thread. - Adjust batch size.
Troubleshooting tips
- Always test with
curlfirst. - Monitor logs in parallel.
- Verify model correctness.
- Mind timeouts and resource limits.
- Keep Ollama updated.
- Check firewall and network settings.
Further reading
- BotServ.de Ollama REST API
- BotServ.de Ollama commands
- BotServ.de Ollama performance
- BotServ.de Docker GPU passthrough
FAQ: Ollama API troubleshooting
What’s the default port? 11434.
Why doesn’t Ollama respond? Ollama service isn’t running, the port is blocked, or you’re connecting to the wrong host.
How do I check if a model is loaded?
Run ollama ps to see active models.
What should I do about timeouts? Increase the timeout, use a smaller model, or check your GPU.
Where are the logs?
Linux: journalctl. macOS: ~/.ollama/logs/server.log.
References
- Ollama API: https://github.com/ollama/ollama/blob/main/docs/api.md
- Ollama Troubleshooting: https://github.com/ollama/ollama/blob/main/docs/troubleshooting.md
Summary: Ollama API troubleshooting
Most Ollama API errors can be isolated using curl, logs, and status codes. Common problems include models failing to load, connection errors, timeouts, and CORS issues. A systematic diagnostic routine, sufficient resources, and an up-to-date Ollama installation resolve the vast majority of issues. By paying attention to logs and HTTP codes, you’ll usually find the root cause quickly.


