Ollama on Windows: Installation and Setup
What This Article Covers
- How to install Ollama on Windows 10 and Windows 11.
- How to verify your GPU is detected and properly configured.
- How to run your first model and use the API from PowerShell.
- How to set up Ollama in WSL2 and access your Windows GPU.
- Common issues and how to troubleshoot them.
Introduction: Ollama on Windows Explained
Ollama is a runtime environment for open-source language models. It downloads models, manages them locally, and exposes a local API. On Windows, an official installer sets up Ollama as a background service. You don’t need Linux or a complicated development setup to get started. This article walks you through installation, GPU configuration, and running your first models step by step. For more background on Ollama, see the Ollama overview.
Why You Need This Guide
Picture this: you download Ollama, install it, and launch your first model. Instead of smooth responses, everything lags because your GPU isn’t recognized. Or you read somewhere that you need WSL2 but aren’t sure if it’s relevant for your use case. Maybe you’ve tried calling the API from PowerShell and hit errors because the syntax was wrong.
This article tackles exactly these situations. You’ll get a clear, practical guide to install Ollama on Windows, configure your GPU, and run your first models. If you’re new to local AI, start with What is local AI?.
Ollama on Windows: The Essentials
Ollama comes as a ready-made installer for Windows. Download it from the official website, run it, and Ollama installs itself as a background service. Then open PowerShell or Command Prompt and use commands like ollama run llama3.1 to start models. If you have an Nvidia GPU, Ollama detects it automatically as long as your drivers are current. AMD and Intel GPUs are also supported, though they may require extra setup. Access the API at http://localhost:11434.
Who This Guide Is For
This article is for Windows users installing Ollama for the first time. You don’t need programming experience, but you should know how to open PowerShell or Command Prompt. If you’re already familiar with Ollama and just need to solve a specific problem, jump straight to the troubleshooting section or common issues.
Key Terms
| Term | Meaning |
|---|---|
| Ollama | Runtime environment for local language models |
| WSL2 | Windows Subsystem for Linux, lets you run Linux programs on Windows |
| GPU | Graphics card that accelerates computations |
| CUDA | Nvidia platform for GPU computing |
| Driver | Software that connects your operating system to the GPU |
| PowerShell | Modern command line for Windows |
| CMD | Classic Windows command prompt |
| Docker Desktop | GUI for Docker on Windows |
| PATH | Environment variable that tells the system where to find programs |
| Service | Background process that runs automatically |
Prerequisites
Before installing Ollama, check the following:
- Operating System: Windows 10 (version 2004, Build 19041 or later) or Windows 11. Both are officially supported.
- GPU Driver: A current graphics driver is important. For Nvidia, use driver version 525 or newer. AMD requires Adrenalin 23.11.1 or newer. Intel Arc series GPUs are also supported.
- RAM: At least 8 GB for small models. For 7B models like Llama 3.1, 16 GB is recommended. Larger models need more.
- Storage: Each model takes several gigabytes. Plan for at least 20 GB of free space if you want to test multiple models.
- PowerShell or CMD: Know how to open a command line. Press
Windows key + R, typepowershell, and press Enter.
For more on the fundamentals, see Running LLMs Locally.
Step 1: Download and Install Ollama
Get the installer from the official website:
- Open
https://ollama.com/downloadin your browser. - Select the Windows download. You’ll get a file called
OllamaSetup.exe. - Double-click the file to run it. Windows may show a security warning. Confirm that you want to run the app.
- The installer runs automatically. No path adjustments needed.
- When done, Ollama runs as a background service. You’ll see a small icon in the system tray, bottom right of the taskbar.
Now open PowerShell and verify Ollama is available:
ollama --version
If you see a version number, installation was successful. If the command isn’t found, restart PowerShell or check whether Ollama is in your PATH. For more on the general installation process, see Ollama Installation.
Step 2: Check Your GPU Driver
Ollama uses your GPU automatically if drivers are correctly installed. Here’s how to verify.
Nvidia
Open PowerShell and enter:
nvidia-smi
You’ll see a table showing your GPU, driver version, and CUDA version. If the command works, your GPU is ready for Ollama. If you get an error, install the latest driver from Nvidia’s website or via GeForce Experience.
AMD
AMD GPUs need the latest Adrenalin driver. Ollama uses the ROCm interface, supported from driver version 23.11.1 onward. Check in PowerShell:
wmic path win32_VideoController get name
If the output confirms your AMD GPU, Ollama should recognize it after installation. If not, update the driver via AMD Software.
Intel
Intel Arc series GPUs are supported. Make sure your Intel Arc driver is current. Older integrated GPUs like Intel UHD are sometimes recognized but offer limited performance. For more on CPUs versus GPUs, see CPU vs. GPU.
Step 3: Run Your First Model
After installation and GPU verification, download your first model. Open PowerShell and enter:
ollama run llama3.1
This command automatically downloads the model if you don’t already have it locally. The download takes a few minutes depending on your connection. Then the chat starts right in the terminal. You can ask a question immediately:
>>> Explain local AI in three sentences.
The model responds directly in the terminal. End the session with /bye or press Ctrl + D. Find more models in Ollama Model Management.
Step 4: Using the API
Ollama exposes an HTTP API at http://localhost:11434. You can call it from PowerShell, Python, or other programming languages.
API call from PowerShell
$body = @{
model = "llama3.1"
prompt = "Was ist lokale KI?"
stream = $false
} | ConvertTo-Json
Invoke-RestMethod -Uri "http://localhost:11434/api/generate" -Method Post -Body $body -ContentType "application/json"
The response comes back as a JSON object. The response field contains the generated text.
API call with Python
import requests
response = requests.post(
"http://localhost:11434/api/generate",
json={
"model": "llama3.1",
"prompt": "Erkläre Quantisierung in zwei Saetzen.",
"stream": False
}
)
print(response.json()["response"])
More examples and endpoints are available in the Ollama API guide.
Ollama with WSL2
WSL2 lets you run Linux programs on Windows. This is useful if you want to use Ollama in a Linux environment, for instance because you have Linux-based tools or scripts. There are two scenarios to consider.
Scenario 1: Ollama runs on Windows, you use it from WSL2
If you’ve installed Ollama on Windows, the service runs in the background. From WSL2, you can reach the API via the Windows IP address. Open WSL2 and use the API:
curl http://host.docker.internal:11434/api/generate -d '{
"model": "llama3.1",
"prompt": "Hallo aus WSL2",
"stream": false
}'
The address host.docker.internal points to your Windows host. This way you can run Ollama on Windows while still accessing it from WSL2.
Scenario 2: Install Ollama directly in WSL2
Alternatively, install Ollama directly in WSL2 using the Linux installation script:
curl -fsSL https://ollama.com/install.sh | sh
In this case, Ollama uses the GPU via WSL2’s GPU passthrough feature. WSL2 needs to be configured with GPU support for this to work. It works best with Nvidia GPUs. If Ollama doesn’t automatically detect your GPU, install the CUDA Toolkit in WSL2.
Make sure you don’t run both variants simultaneously. If Ollama is running on both Windows and WSL2, you’ll encounter port conflicts since both use port 11434.
Ollama as a Windows Service
After installation with the official installer, Ollama runs automatically as a background service. This means you don’t need to start Ollama manually. The service runs in the background after Windows boots and provides the API.
You can check the service status in PowerShell:
Get-Service ollama
If the status shows Running, the service is active. Start it manually with:
Start-Service ollama
Stop it using:
Stop-Service ollama
The service starts automatically on system boot. If you want to change this behavior, you can adjust the startup type:
Set-Service ollama -StartupType Manual
Configuration: Changing the Storage Location
By default, Ollama stores models in your user directory at C:\Users\YourName\.ollama\models. If your system drive is running low on space, you can change the storage location using an environment variable.
Set the OLLAMA_MODELS variable to a different path:
[System.Environment]::SetEnvironmentVariable("OLLAMA_MODELS", "D:\ollama-models", "User")
After that, restart the Ollama service for the change to take effect:
Restart-Service ollama
Newly downloaded models will go to the new directory from now on. Models you’ve already downloaded need to be moved manually or re-downloaded. For more configuration options, see the Ollama Configuration guide.
Common Gotchas with Ollama on Windows
1. GPU is not detected
Ollama falls back to the CPU if your GPU drivers are outdated or missing. Check whether your GPU is available using nvidia-smi. Update your drivers and restart the Ollama service.
2. Command ollama not found
After installation, the command is sometimes not immediately available. Close PowerShell and open it again. If that doesn’t help, verify that Ollama is in your PATH. The default path is C:\Users\YourName\AppData\Local\Programs\Ollama.
3. Port 11434 is in use
If another program is using the port, Ollama won’t start correctly. Check which process is blocking the port with netstat -ano | findstr 11434. Either close that process or change the port using the OLLAMA_HOST environment variable.
4. Models load slowly
If Ollama runs without GPU acceleration, loading large models takes time. Verify GPU detection and make sure you have enough VRAM. A 7B model typically needs 4 to 5 GB of VRAM.
5. WSL2 and Windows Ollama simultaneously
If you’ve installed Ollama on both Windows and WSL2, you’ll have port conflicts. Choose one approach. Either use the Windows installation and access the API from WSL2, or install Ollama only in WSL2.
6. Antivirus blocks Ollama
Some antivirus programs block model downloads or the service itself. Add Ollama to your antivirus exclusion list if you encounter issues.
7. Special characters in PowerShell
PowerShell sometimes has trouble with special characters in prompts. Use ae, oe, and ue instead, or use the API with Python where encoding works more cleanly.
More solutions are in the Ollama Troubleshooting guide.
Hardware, Costs, and Security with Ollama on Windows
Hardware
Performance depends heavily on your hardware. An Nvidia GPU with 8 GB of VRAM is sufficient for 7B models like Llama 3.1. For 13B models, you need at least 12 GB of VRAM. Without a GPU, Ollama runs on the CPU, which is significantly slower. More details in CPU vs. GPU.
Costs
Ollama is open source and free. Costs arise only from hardware and electricity. Downloading a 4 GB model costs nothing beyond your internet bandwidth.
Security
Ollama runs locally by default. The API is accessible only on localhost, not from the outside. If you want to expose the API to your network, set OLLAMA_HOST=0.0.0.0:11434. In that case, pay attention to firewall rules and authentication. All models and data stay on your machine. Nothing is sent to cloud services.
Further Links and Resources for Ollama on Windows
- Ollama Overview - Everything about Ollama at a glance.
- Ollama Installation - Installation on Linux, macOS, and Docker.
- Ollama API - Endpoints, curl, Python, and JavaScript.
- Managing Ollama Models - Download, list, and delete models.
- Ollama Configuration - Environment variables and settings.
- Ollama Troubleshooting - Solutions for common issues.
- Ollama with Docker - Running Ollama as a container.
- Running LLMs Locally - Basics of local inference.
- What is Local AI? - Introduction to the topic.
- CPU vs. GPU - When to use which hardware.
FAQ: Ollama on Windows - Common Questions
Do I need WSL2 to run Ollama on Windows?
No. The official Windows installer is sufficient. WSL2 is only necessary if you specifically want a Linux environment.
Does Ollama work without a GPU?
Yes, Ollama runs on CPU. Performance is noticeably lower, but adequate for small models and testing.
Where does Ollama store models on Windows?
By default, in C:\Users\YourName\.ollama\models. You can change the path using the OLLAMA_MODELS environment variable.
How do I verify my GPU is recognized?
For Nvidia GPUs, run nvidia-smi in PowerShell. If a table with GPU information appears, your GPU is ready.
Can I use Ollama across multiple Windows user accounts?
Yes, the service runs system-wide. However, models are stored in the user directory. Each user has their own model folder unless you set OLLAMA_MODELS to a shared path.
Does Ollama start automatically when Windows boots?
Yes, the installer registers Ollama as a service with automatic startup. You can change this in Windows Services.
Can I run Ollama and Docker Desktop simultaneously?
Yes, as long as there are no port conflicts. If you want to run Ollama in Docker, use a different port or stop the Windows service. See Ollama with Docker for details.
How do I change Ollama’s port?
Set the OLLAMA_HOST environment variable to something like 127.0.0.1:11435 and restart the service.
Are AMD GPUs supported?
Yes, starting with driver version 23.11.1 with ROCm support. Update your Adrenalin driver and verify that Ollama detects the GPU.
Is Ollama free on Windows?
Yes, Ollama is open source and free. Costs apply only to hardware and electricity.
Can I use Ollama from PowerShell?
Yes, all commands like ollama run, ollama pull, and ollama list work in PowerShell and the classic Command Prompt.
References and Further Reading
- Ollama website
- Ollama downloads
- Ollama GitHub repository
- Microsoft WSL2 documentation
- Nvidia CUDA Toolkit
- AMD ROCm documentation


