Managing Ollama in Open WebUI
What this article covers
- Connecting Open WebUI to Ollama.
- Loading, disabling, and prioritizing models.
- Using system prompts, parameters, and model selection.
- Integrating multiple Ollama instances.
- Common pitfalls and best practices.
Introduction: Managing Ollama in Open WebUI
Open WebUI is a user-friendly interface for chatting with local language models. Ollama provides the models and the API. Connect the two together, and you get a simple web interface for accessing all your local models without touching the command line.
Managing Ollama in Open WebUI is straightforward once the connection is established. This article walks you through adding models, adjusting settings, and troubleshooting common issues.
Prerequisites
- Ollama is installed and running.
- Open WebUI is installed, for example with Docker.
- Both services can reach each other.
Setting up the connection
Open WebUI detects Ollama automatically if it runs on the same host. By default, it expects Ollama at:
http://host.docker.internal:11434
If Ollama runs outside Docker, set the environment variable in Open WebUI:
-e OLLAMA_BASE_URL=http://host.docker.internal:11434
For a separate Ollama instance, provide the IP address or hostname:
-e OLLAMA_BASE_URL=http://192.168.1.50:11434
Downloading models
In Open WebUI, you can load models directly from the interface:
- Open Settings.
- Go to
ModelsorAdmin Settings. - Select
Pull a model from Ollama.com. - Enter the model name, for example
llama3.1:8b. - Start the download.
Alternatively, from the terminal:
ollama pull llama3.1:8b
ollama pull mistral
ollama pull qwen2.5:14b
Once the model is available locally, it appears in Open WebUI.
Selecting models in chat
In the chat window, you can select your desired model from the dropdown at the top. Open WebUI displays all locally available Ollama models. You can switch models during a conversation, but keep in mind that the previous context is not automatically carried over.
System prompts and parameters
For each model, you can adjust settings in Open WebUI:
- System Prompt: Instructions that guide the model’s behavior.
- Temperature: Balance between creativity and determinism.
- Top P: Sampling diversity.
- Max Tokens: Maximum response length.
- Context Length: Context window used for the model.
These settings can be configured globally, per model, or per conversation.
Multiple Ollama instances
If you run multiple Ollama servers, Open WebUI can work with them through appropriate configuration or a reverse proxy with multiple backends. For most cases, a single central Ollama instance is sufficient.
Model prioritization and visibility
Administrators can control which models are visible to users in Open WebUI. This is useful when:
- Only specific models should be allowed.
- Large models should be available only to admins.
- Test models need to be hidden before release.
API keys and external providers
Alongside Ollama, Open WebUI can integrate external providers like OpenAI, Anthropic, or OpenRouter. You store API keys in the settings. Locally, Ollama remains the default with no cost per request.
Updating and deleting models
Update to newer tags through Ollama:
ollama pull llama3.1:8b
Remove a model with:
ollama rm llama3.1:8b
In Open WebUI, deleted models no longer appear after a refresh.
Common pitfalls
- Missing Ollama connection: Check
OLLAMA_BASE_URLand connectivity. - Model not visible: Ollama must have loaded the model. Refresh Open WebUI.
- Insufficient memory: The model fails to load, and Ollama exits.
- Wrong model size:
llama3.1without a tag defaults tolatest, often the 70B version. - Network issues: Verify container name resolution.
Further reading
- BotServ.de Open WebUI Features
- BotServ.de Setting Up RAG in Open WebUI
- BotServ.de Open WebUI Tools
- BotServ.de Multiple Users in Open WebUI
FAQ: Ollama in Open WebUI
How do I connect Open WebUI to Ollama?
Usually automatically, if both run on the same host. Otherwise, set OLLAMA_BASE_URL.
Can I download models directly in Open WebUI?
Yes, through the admin panel or ollama pull in the terminal.
How many models can I use simultaneously? As many as fit in your RAM. Each chat loads the selected model.
Do I need API keys? Only if you use external providers in addition to Ollama.
What happens if I run out of RAM? The model fails to load or runs very slowly. In that case, use smaller models or quantized versions.
References and further reading
- Open WebUI Docs: https://docs.openwebui.com/
- Ollama Docs: https://github.com/ollama/ollama/blob/main/docs/
- Open WebUI GitHub: https://github.com/open-webui/open-webui
Summary: Managing Ollama in Open WebUI
Open WebUI is a convenient interface for Ollama models. After connecting via OLLAMA_BASE_URL, you can download models, select them, and run them with custom prompts and parameters. Administrators control visibility and add providers. Keep an eye on connectivity, memory, and model sizes, and you have straightforward access to local AI.


