Ollama Configuration: All Options and Settings
What This Article Covers
- Which environment variables Ollama offers and how to set them
- How to change the model storage location so your disk doesn’t fill up
- How to allow network access from other devices
- How to create custom model variants with Modelfiles
- How to optimize performance through GPU layers, parallelism, and context length
Introduction: Understanding Ollama Configuration
Ollama typically works right after installation. You download a model, ask a question, get an answer. For initial experiments, that’s more than enough. But once you start working seriously with local models, you’ll run into questions that only configuration can solve: Where are models stored? How much VRAM should Ollama use? Can another device on the network access Ollama? How do I keep a model in RAM longer?
This article walks you through all the important settings. You’ll learn which environment variables exist, how to write Modelfiles, and how to configure Ollama on Linux, Windows, and macOS. Step by step, with examples you can use directly.
If you haven’t installed Ollama yet, read the Ollama overview and installation guide first.
Why Do You Need Configuration?
Imagine you’ve installed Ollama on Windows and downloaded a few models. A 7B model takes around 4.5 GB, a 13B model about 7 GB, and larger models quickly become unwieldy. By default, Ollama stores everything on your system drive, usually C:. After three or four models, your SSD fills up and Windows warns you about low disk space.
This isn’t a bug, it’s the default behavior. Ollama doesn’t know you have a large data drive on D:. You need to tell Ollama via the OLLAMA_MODELS environment variable where to store the models.
Here’s another scenario: you want to access Ollama from a second PC or your phone over WiFi. By default, Ollama only listens on localhost, which means it won’t accept connections from outside. Only with OLLAMA_HOST=0.0.0.0 and appropriate OLLAMA_ORIGINS settings do you open the service to your network. More on this in the article Network Access.
Or suppose you’re working with long documents and need a larger context window. By default, Ollama uses 2048 tokens as the context window. For longer texts, you need to increase the value via a Modelfile or API parameter. Read the fundamentals article Context Length for more.
Configuration, then, isn’t optional but essential for adapting Ollama to your hardware and use case.
Ollama Configuration Explained Simply
Ollama is configured mainly through environment variables. These are key-value pairs that Ollama reads on startup. You set them in your operating system, and Ollama applies them the next time it starts.
Additionally, there are Modelfiles. These are text files where you describe how a specific model should behave: which system prompt applies, what parameters govern temperature or context length, and whether the model builds on another model.
The most important environment variables are OLLAMA_MODELS (storage location), OLLAMA_HOST (network address), OLLAMA_ORIGINS (allowed sources), OLLAMA_GPU_LAYERS (GPU layers), OLLAMA_NUM_PARALLEL (parallel requests), and OLLAMA_MAX_LOADED_MODELS (simultaneously loaded models).
Who Is This Configuration For?
This article is for beginners who have installed Ollama and want more control. You don’t need programming knowledge, just basic familiarity with your operating system. Instructions for setting environment variables are provided separately for Linux, Windows, and macOS.
If you’re running Ollama in production, such as a backend for a chat application on your LAN, you’ll also find advanced topics here like performance tuning and systemd configuration.
Key Terms in Ollama Configuration
| Term | Meaning |
|---|---|
| Environment variable | A key-value pair in the operating system that Ollama reads on startup, e.g., OLLAMA_MODELS=/data/ollama |
| Modelfile | A text file describing how a model is configured: base model, system prompt, parameters |
OLLAMA_MODELS | Specifies the directory where Ollama stores and loads models |
OLLAMA_HOST | Determines the address and port Ollama listens on, e.g., 0.0.0.0:11434 |
OLLAMA_ORIGINS | List of allowed origins for CORS, important for web frontends |
OLLAMA_GPU_LAYERS | Number of model layers offloaded to the GPU |
OLLAMA_NUM_PARALLEL | Number of requests Ollama processes simultaneously |
OLLAMA_MAX_LOADED_MODELS | Maximum number of models that remain in memory at the same time |
| Context window | Number of tokens the model considers as context |
| System prompt | An instruction the model receives at the start that controls its behavior |
Environment Variables at a Glance
Ollama recognizes a variety of environment variables. Here are the most important ones with descriptions and defaults:
| Variable | Description | Default |
|---|---|---|
OLLAMA_HOST | Address and port Ollama listens on | 127.0.0.1:11434 |
OLLAMA_MODELS | Directory for model storage | ~/.ollama/models |
OLLAMA_ORIGINS | Allowed origins for CORS requests | * (in newer versions) |
OLLAMA_GPU_LAYERS | Layers offloaded to the GPU | Automatic, depending on VRAM |
OLLAMA_NUM_PARALLEL | Number of parallel requests per model | 1 |
OLLAMA_MAX_LOADED_MODELS | Maximum models loaded simultaneously | 1 (with limited RAM) |
OLLAMA_KEEP_ALIVE | How long a model stays in RAM after the last request | 5m |
OLLAMA_NUM_CTX | Default context length in tokens | 2048 |
OLLAMA_FLASH_ATTENTION | Enables Flash Attention, saves VRAM | 0 (disabled) |
OLLAMA_KV_CACHE_TYPE | Data type for KV cache, e.g., q8_0 for less VRAM | Default data type |
OLLAMA_DEBUG | Enables detailed logging | 0 |
OLLAMA_LLM_LIBRARY | Forces a specific backend library | Automatic |
You don’t need to set every variable. Most defaults make sense. The following sections show when to adjust which variable.
Changing Model Storage (OLLAMA_MODELS)
The most common configuration change is altering where models are stored. By default, Ollama keeps models in your home directory. On Windows, that’s often C:\Users\YourName\.ollama\models, on Linux ~/.ollama/models, on macOS ~/.ollama/models.
If your system drive is small, you’ll want to move models to a different drive.
Linux
Add the variable permanently to your shell configuration:
echo 'export OLLAMA_MODELS=/data/ollama/models' >> ~/.bashrc
source ~/.bashrc
If you’re running Ollama as a systemd service, see the Linux configuration section below.
Windows
- Press the Windows key and search for “Environment Variables”
- Click “Edit the system environment variables”
- Click “Environment Variables”
- Under “User variables” click “New”
- Variable name:
OLLAMA_MODELS, Variable value:D:\ollama\models - Click OK
- Restart Ollama, preferably via Task Manager or a system restart
macOS
On macOS, set the variable in your shell configuration when you start Ollama from the terminal:
echo 'export OLLAMA_MODELS=/Volumes/Data/ollama/models' >> ~/.zshrc
source ~/.zshrc
If you’re using the Ollama app, launch it from the terminal with the variable set, or use launchctl setenv:
launchctl setenv OLLAMA_MODELS /Volumes/Data/ollama/models
Then restart the Ollama app. Move any existing models from the old directory to the new one so you don’t have to re-download them.
Configuring the Network (OLLAMA_HOST, OLLAMA_ORIGINS)
By default, Ollama listens on 127.0.0.1:11434, which means it’s only accessible from your local machine. To access Ollama from another device on your network, you need to change the bind address.
Set OLLAMA_HOST to 0.0.0.0:11434 so Ollama listens on all network interfaces:
export OLLAMA_HOST=0.0.0.0:11434
Now you can reach Ollama from another device using your computer’s IP address, for example http://192.168.1.100:11434.
If you’re using a web frontend like Open WebUI running in the browser and sending requests to Ollama, you’ll also need OLLAMA_ORIGINS. Browsers send an Origin header, and Ollama checks whether that origin is allowed:
export OLLAMA_ORIGINS=http://192.168.1.50:3000,http://localhost:3000
Separate multiple origins with commas. For detailed information, see the article Network Access.
Warning: If you expose Ollama to the internet, anyone can use your models. Restrict access to your local network or place a reverse proxy with authentication in front.
Controlling GPU Layers
Ollama automatically decides how many layers of a model to offload to the GPU. This depends on available VRAM. Sometimes you want to intervene manually, for example if Ollama is consuming too much VRAM and other programs are suffering.
Use OLLAMA_GPU_LAYERS to specify how many layers go to the GPU. The rest runs on the CPU:
export OLLAMA_GPU_LAYERS=20
A higher value means more layers on the GPU and faster responses, but also higher VRAM consumption. A lower value conserves VRAM but makes responses slower.
You can also set the value per model using a Modelfile, see the next section. For more background on GPU offloading, check out the article GPU Offloading.
Modelfiles: Creating Custom Models
Modelfiles are text files where you describe how a model is configured. You can set a system prompt, tune parameters like temperature and context length, and build on an existing model.
Modelfile Structure
A Modelfile consists of instructions, one per line. The main ones are:
FROM: The base model your model builds onSYSTEM: The system prompt that controls model behaviorPARAMETER: Model parameters like temperature, context length, top-kTEMPLATE: The prompt template defining how inputs are formatted
A Simple Example
Create a file called Modelfile:
FROM llama3.1:8b
SYSTEM "Du bist ein hilfreicher Assistent, der auf Deutsch antwortet. Du erklärst Dinge einfach und verständlich."
PARAMETER temperature 0.7
PARAMETER num_ctx 4096
PARAMETER top_p 0.9
Create the model with:
ollama create mein-assistent -f Modelfile
Then start it like any other model:
ollama run mein-assistent
Parameters at a Glance
| Parameter | Purpose | Typical Value |
|---|---|---|
temperature | Controls creativity, lower = more deterministic | 0.7 to 0.9 |
num_ctx | Context length in tokens | 2048 to 8192 |
top_p | Nucleus sampling, limits selection probability | 0.9 |
top_k | Limits selection to the k most likely tokens | 40 |
repeat_penalty | Penalizes repetitions | 1.1 |
num_gpu | Number of GPU layers for this model | Automatic or e.g. 20 |
stop | Sequences where the model stops generating | e.g. "\nUser:" |
Modelfiles are the simplest way to adapt a model to your use case without modifying the base model. For more on managing models, see the article Managing Models.
Performance Tuning
If you’re using Ollama more intensively, such as as a backend for multiple users or applications, performance questions arise. These variables help you address them.
OLLAMA_NUM_PARALLEL
By default, Ollama processes one request per model at a time. If you want to handle multiple requests in parallel, increase the value:
export OLLAMA_NUM_PARALLEL=2
This means Ollama processes two requests simultaneously. This costs more VRAM because the KV cache is held for both requests. Check that your VRAM is sufficient.
OLLAMA_MAX_LOADED_MODELS
This variable sets how many models stay in memory at once. The default is often 1, meaning Ollama unloads one model before loading the next. If you frequently switch between models, increase the value:
export OLLAMA_MAX_LOADED_MODELS=2
This costs RAM and VRAM because both models are kept in memory simultaneously. Use this only if you have enough memory.
Adjusting Context Length
Context length determines how many tokens the model considers as history. More context means you can process longer documents or conversations, but it costs more VRAM. Set the context length via a Modelfile with PARAMETER num_ctx 8192, or via the API with the num_ctx parameter.
For background on context length and VRAM consumption, see the article Context Length.
Controlling keep_alive
With OLLAMA_KEEP_ALIVE, you set how long a model remains in RAM after its last request. The default is 5m, or five minutes. If you make frequent requests, you can increase the value to save load times:
export OLLAMA_KEEP_ALIVE=30m
If you want to conserve RAM, set the value lower or to 0 so the model is unloaded immediately.
Flash Attention
Enable Flash Attention with OLLAMA_FLASH_ATTENTION=1. This can reduce VRAM consumption with long contexts and increase speed. Not all models support it, so experiment and compare results.
Configuration on Linux (systemd)
On Linux, Ollama typically runs as a systemd service. Set environment variables in the service file.
- Open the service file:
sudo systemctl edit ollama.service
- Add the variables to the override:
[Service]
Environment="OLLAMA_MODELS=/data/ollama/models"
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_ORIGINS=http://192.168.1.50:3000"
Environment="OLLAMA_NUM_PARALLEL=2"
- Save and reload systemd:
sudo systemctl daemon-reload
sudo systemctl restart ollama
- Verify that the variables are active:
systemctl show ollama --property=Environment
This approach is permanent and survives restarts. Never edit the original service file directly. Always use systemctl edit so your changes aren’t lost during updates.
Configuration on Windows
On Windows, use the graphical interface to set environment variables.
- Press the Windows key and type “environment variables”
- Select “Edit system environment variables”
- Click “Environment Variables”
- Under “User variables”, click “New”
- Enter the variable name and value, for example:
- Name:
OLLAMA_HOST, Value:0.0.0.0:11434 - Name:
OLLAMA_MODELS, Value:D:\ollama\models
- Name:
- Click OK
- Exit Ollama from the system tray (bottom right of the taskbar)
- Restart Ollama
Environment variables take effect after restarting Ollama. If you run Ollama as a Windows service, set variables under “System variables” instead of “User variables” so the service can access them.
Configuration on macOS
On macOS, the method depends on how you launch Ollama.
Ollama App
For the graphical app, use launchctl setenv:
launchctl setenv OLLAMA_MODELS /Volumes/Data/ollama/models
launchctl setenv OLLAMA_HOST 0.0.0.0:11434
Restart the Ollama app. These variables persist until your next reboot. For a permanent solution, create a plist file in ~/Library/LaunchAgents/.
Terminal
If you launch Ollama from the terminal, set variables in your shell configuration:
echo 'export OLLAMA_HOST=0.0.0.0:11434' >> ~/.zshrc
source ~/.zshrc
Then start Ollama with ollama serve in the terminal.
Common Configuration Pitfalls
-
Variables not taking effect: You must completely restart Ollama, not just reload a model. On Linux with systemd:
sudo systemctl restart ollama. On Windows: quit Ollama from the system tray and restart it. -
Models disappear after changing storage location: When you change
OLLAMA_MODELS, Ollama no longer recognizes the old models. Move the contents of the old directory to the new one, or re-download the models. -
CORS errors in the browser: If a web frontend calls Ollama and you see “CORS policy” errors, the origin is missing from
OLLAMA_ORIGINS. Add the frontend address, including the port. -
Insufficient VRAM with large context: A context length of 8192 tokens requires significantly more VRAM than 2048. If Ollama crashes or becomes extremely slow, reduce
num_ctxor enable Flash Attention. -
Port conflicts: If port 11434 is already in use, Ollama won’t start. Change the port via
OLLAMA_HOST=0.0.0.0:11435or find and kill the process holding the port. -
systemd override being ignored: Check with
systemctl show ollama --property=Environmentto verify the variables are actually loaded. A missingsudo systemctl daemon-reloadis often the culprit. -
Security with network sharing:
OLLAMA_HOST=0.0.0.0exposes Ollama to every device on your network. Without authentication, anyone can use your models. Put a reverse proxy with password protection in front, or restrict access via firewall. -
Environment variables in the wrong scope on Windows: If Ollama runs as a service, variables must be under “System variables”, not “User variables”. Otherwise the service won’t find them.
For further help with issues, see the Troubleshooting article.
Hardware, Costs, and Security in Configuration
Configuration depends heavily on your hardware. With lots of VRAM, you can offload more layers to the GPU, use larger contexts, and load multiple models at once. With limited VRAM, you’ll need to make tradeoffs.
VRAM and Layers
An 8B model with 4-bit quantization requires around 5 GB VRAM for weights. The KV cache for context adds more on top. At 4096 tokens of context, that’s another 1 to 2 GB; at 8192 tokens, correspondingly more. Read more about quantization and how it reduces memory requirements in Quantization.
RAM and CPU
If not all layers fit on the GPU, the remainder runs on the CPU. It’s slower, but it works. For CPU-only systems, choose smaller models and shorter contexts.
Costs
Ollama itself is free and open source. Costs come only from hardware: a GPU with enough VRAM, sufficient RAM, and storage for models. If you’re just trying Ollama, a regular PC will do. For serious use with large models, a dedicated GPU makes sense. Learn more in Running LLMs Locally.
Security
If you use Ollama only locally, there are no security concerns. As soon as you allow network access, avoid exposing Ollama directly to the internet. Use a reverse proxy with authentication, such as Nginx or Caddy, and restrict access to your LAN. Set OLLAMA_ORIGINS as restrictively as possible.
Further Reading and Configuration Resources
- Ollama Overview - Getting started with Ollama
- Ollama Installation - Step-by-step setup
- Ollama API - API reference
- Managing Models - Load, delete, and update models
- Network Access - Make Ollama available on your LAN
- Troubleshooting - Solve common issues
- Context Length - How context windows work
- Quantization - Reduce model size
- GPU Offloading - Offload layers to the GPU
FAQ: Ollama Configuration - Common Questions
Where does Ollama store models by default?
In your home directory under ~/.ollama/models on Linux and macOS, and C:\Users\YourName\.ollama\models on Windows. Change this location with OLLAMA_MODELS.
How do I change Ollama’s port?
Set OLLAMA_HOST=127.0.0.1:11435 to use port 11435. Restart Ollama afterward.
What is a Modelfile?
A text file where you specify which base model your model is built on, what system prompt to use, and which parameters apply. Create one with ollama create.
How do I allow access from the network?
Set OLLAMA_HOST=0.0.0.0:11434 so Ollama listens on all network interfaces. Also set OLLAMA_ORIGINS for CORS if you’re using a web frontend.
How do I increase the context length?
Either via a Modelfile with PARAMETER num_ctx 8192, or through the API using the num_ctx parameter in options. Longer context requires more VRAM.
What does OLLAMA_KEEP_ALIVE do?
It determines how long a model stays in RAM after the last request. Default is five minutes. Higher values save load time but consume more RAM.
How do I see which variables are active?
On Linux with systemd: systemctl show ollama --property=Environment. On Windows and macOS, check environment variables in system settings or run echo $OLLAMA_HOST in the terminal.
What is Flash Attention?
A technique that reduces VRAM usage with long contexts and increases speed. Enable it with OLLAMA_FLASH_ATTENTION=1. Not all models support it.
Can I load multiple models simultaneously?
Yes, with OLLAMA_MAX_LOADED_MODELS=2 or higher. This consumes RAM and VRAM since both models are kept in memory.
How many parallel requests does Ollama support?
By default, one per model. With OLLAMA_NUM_PARALLEL=2 or higher, Ollama processes multiple requests concurrently. VRAM usage increases because the KV cache is maintained for each request.
How do I debug Ollama?
Set OLLAMA_DEBUG=1 to enable detailed logging. On Linux, view logs with journalctl -u ollama. On Windows and macOS, check the console where you started Ollama.
Do I need to restart Ollama after setting a variable?
Yes, Ollama reads environment variables only at startup. A service or app restart is required.
References and Further Reading
- Official Ollama documentation: GitHub repository at
ollama/ollama - Ollama FAQ and Issues on GitHub
- Article GPU Offloading on BotServ.de
- Article Context Length on BotServ.de
- Article Quantization on BotServ.de
- Article Running LLMs Locally on BotServ.de


