Skip to content
BotServBotServ
OllamaWindowsInstallationGPUWSL2local AI

Ollama on Windows: Installation and Setup

Install and configure Ollama on Windows. Step-by-step guide with GPU setup, WSL2, and troubleshooting.

S

schutzgeist

10 min read
Ollama on Windows: Installation and Setup

Ollama on Windows: Installation and Setup

What This Article Covers

  • How to install Ollama on Windows 10 and Windows 11.
  • How to verify your GPU is detected and properly configured.
  • How to run your first model and use the API from PowerShell.
  • How to set up Ollama in WSL2 and access your Windows GPU.
  • Common issues and how to troubleshoot them.

Introduction: Ollama on Windows Explained

Ollama is a runtime environment for open-source language models. It downloads models, manages them locally, and exposes a local API. On Windows, an official installer sets up Ollama as a background service. You don’t need Linux or a complicated development setup to get started. This article walks you through installation, GPU configuration, and running your first models step by step. For more background on Ollama, see the Ollama overview.

Why You Need This Guide

Picture this: you download Ollama, install it, and launch your first model. Instead of smooth responses, everything lags because your GPU isn’t recognized. Or you read somewhere that you need WSL2 but aren’t sure if it’s relevant for your use case. Maybe you’ve tried calling the API from PowerShell and hit errors because the syntax was wrong.

This article tackles exactly these situations. You’ll get a clear, practical guide to install Ollama on Windows, configure your GPU, and run your first models. If you’re new to local AI, start with What is local AI?.

Ollama on Windows: The Essentials

Ollama comes as a ready-made installer for Windows. Download it from the official website, run it, and Ollama installs itself as a background service. Then open PowerShell or Command Prompt and use commands like ollama run llama3.1 to start models. If you have an Nvidia GPU, Ollama detects it automatically as long as your drivers are current. AMD and Intel GPUs are also supported, though they may require extra setup. Access the API at http://localhost:11434.

Who This Guide Is For

This article is for Windows users installing Ollama for the first time. You don’t need programming experience, but you should know how to open PowerShell or Command Prompt. If you’re already familiar with Ollama and just need to solve a specific problem, jump straight to the troubleshooting section or common issues.

Key Terms

TermMeaning
OllamaRuntime environment for local language models
WSL2Windows Subsystem for Linux, lets you run Linux programs on Windows
GPUGraphics card that accelerates computations
CUDANvidia platform for GPU computing
DriverSoftware that connects your operating system to the GPU
PowerShellModern command line for Windows
CMDClassic Windows command prompt
Docker DesktopGUI for Docker on Windows
PATHEnvironment variable that tells the system where to find programs
ServiceBackground process that runs automatically

Prerequisites

Before installing Ollama, check the following:

  • Operating System: Windows 10 (version 2004, Build 19041 or later) or Windows 11. Both are officially supported.
  • GPU Driver: A current graphics driver is important. For Nvidia, use driver version 525 or newer. AMD requires Adrenalin 23.11.1 or newer. Intel Arc series GPUs are also supported.
  • RAM: At least 8 GB for small models. For 7B models like Llama 3.1, 16 GB is recommended. Larger models need more.
  • Storage: Each model takes several gigabytes. Plan for at least 20 GB of free space if you want to test multiple models.
  • PowerShell or CMD: Know how to open a command line. Press Windows key + R, type powershell, and press Enter.

For more on the fundamentals, see Running LLMs Locally.

Step 1: Download and Install Ollama

Get the installer from the official website:

  1. Open https://ollama.com/download in your browser.
  2. Select the Windows download. You’ll get a file called OllamaSetup.exe.
  3. Double-click the file to run it. Windows may show a security warning. Confirm that you want to run the app.
  4. The installer runs automatically. No path adjustments needed.
  5. When done, Ollama runs as a background service. You’ll see a small icon in the system tray, bottom right of the taskbar.

Now open PowerShell and verify Ollama is available:

ollama --version

If you see a version number, installation was successful. If the command isn’t found, restart PowerShell or check whether Ollama is in your PATH. For more on the general installation process, see Ollama Installation.

Step 2: Check Your GPU Driver

Ollama uses your GPU automatically if drivers are correctly installed. Here’s how to verify.

Nvidia

Open PowerShell and enter:

nvidia-smi

You’ll see a table showing your GPU, driver version, and CUDA version. If the command works, your GPU is ready for Ollama. If you get an error, install the latest driver from Nvidia’s website or via GeForce Experience.

AMD

AMD GPUs need the latest Adrenalin driver. Ollama uses the ROCm interface, supported from driver version 23.11.1 onward. Check in PowerShell:

wmic path win32_VideoController get name

If the output confirms your AMD GPU, Ollama should recognize it after installation. If not, update the driver via AMD Software.

Intel

Intel Arc series GPUs are supported. Make sure your Intel Arc driver is current. Older integrated GPUs like Intel UHD are sometimes recognized but offer limited performance. For more on CPUs versus GPUs, see CPU vs. GPU.

Step 3: Run Your First Model

After installation and GPU verification, download your first model. Open PowerShell and enter:

ollama run llama3.1

This command automatically downloads the model if you don’t already have it locally. The download takes a few minutes depending on your connection. Then the chat starts right in the terminal. You can ask a question immediately:

>>> Explain local AI in three sentences.

The model responds directly in the terminal. End the session with /bye or press Ctrl + D. Find more models in Ollama Model Management.

Step 4: Using the API

Ollama exposes an HTTP API at http://localhost:11434. You can call it from PowerShell, Python, or other programming languages.

API call from PowerShell

$body = @{
    model = "llama3.1"
    prompt = "Was ist lokale KI?"
    stream = $false
} | ConvertTo-Json

Invoke-RestMethod -Uri "http://localhost:11434/api/generate" -Method Post -Body $body -ContentType "application/json"

The response comes back as a JSON object. The response field contains the generated text.

API call with Python

import requests

response = requests.post(
    "http://localhost:11434/api/generate",
    json={
        "model": "llama3.1",
        "prompt": "Erkläre Quantisierung in zwei Saetzen.",
        "stream": False
    }
)
print(response.json()["response"])

More examples and endpoints are available in the Ollama API guide.

Ollama with WSL2

WSL2 lets you run Linux programs on Windows. This is useful if you want to use Ollama in a Linux environment, for instance because you have Linux-based tools or scripts. There are two scenarios to consider.

Scenario 1: Ollama runs on Windows, you use it from WSL2

If you’ve installed Ollama on Windows, the service runs in the background. From WSL2, you can reach the API via the Windows IP address. Open WSL2 and use the API:

curl http://host.docker.internal:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "Hallo aus WSL2",
  "stream": false
}'

The address host.docker.internal points to your Windows host. This way you can run Ollama on Windows while still accessing it from WSL2.

Scenario 2: Install Ollama directly in WSL2

Alternatively, install Ollama directly in WSL2 using the Linux installation script:

curl -fsSL https://ollama.com/install.sh | sh

In this case, Ollama uses the GPU via WSL2’s GPU passthrough feature. WSL2 needs to be configured with GPU support for this to work. It works best with Nvidia GPUs. If Ollama doesn’t automatically detect your GPU, install the CUDA Toolkit in WSL2.

Make sure you don’t run both variants simultaneously. If Ollama is running on both Windows and WSL2, you’ll encounter port conflicts since both use port 11434.

Ollama as a Windows Service

After installation with the official installer, Ollama runs automatically as a background service. This means you don’t need to start Ollama manually. The service runs in the background after Windows boots and provides the API.

You can check the service status in PowerShell:

Get-Service ollama

If the status shows Running, the service is active. Start it manually with:

Start-Service ollama

Stop it using:

Stop-Service ollama

The service starts automatically on system boot. If you want to change this behavior, you can adjust the startup type:

Set-Service ollama -StartupType Manual

Configuration: Changing the Storage Location

By default, Ollama stores models in your user directory at C:\Users\YourName\.ollama\models. If your system drive is running low on space, you can change the storage location using an environment variable.

Set the OLLAMA_MODELS variable to a different path:

[System.Environment]::SetEnvironmentVariable("OLLAMA_MODELS", "D:\ollama-models", "User")

After that, restart the Ollama service for the change to take effect:

Restart-Service ollama

Newly downloaded models will go to the new directory from now on. Models you’ve already downloaded need to be moved manually or re-downloaded. For more configuration options, see the Ollama Configuration guide.

Common Gotchas with Ollama on Windows

1. GPU is not detected

Ollama falls back to the CPU if your GPU drivers are outdated or missing. Check whether your GPU is available using nvidia-smi. Update your drivers and restart the Ollama service.

2. Command ollama not found

After installation, the command is sometimes not immediately available. Close PowerShell and open it again. If that doesn’t help, verify that Ollama is in your PATH. The default path is C:\Users\YourName\AppData\Local\Programs\Ollama.

3. Port 11434 is in use

If another program is using the port, Ollama won’t start correctly. Check which process is blocking the port with netstat -ano | findstr 11434. Either close that process or change the port using the OLLAMA_HOST environment variable.

4. Models load slowly

If Ollama runs without GPU acceleration, loading large models takes time. Verify GPU detection and make sure you have enough VRAM. A 7B model typically needs 4 to 5 GB of VRAM.

5. WSL2 and Windows Ollama simultaneously

If you’ve installed Ollama on both Windows and WSL2, you’ll have port conflicts. Choose one approach. Either use the Windows installation and access the API from WSL2, or install Ollama only in WSL2.

6. Antivirus blocks Ollama

Some antivirus programs block model downloads or the service itself. Add Ollama to your antivirus exclusion list if you encounter issues.

7. Special characters in PowerShell

PowerShell sometimes has trouble with special characters in prompts. Use ae, oe, and ue instead, or use the API with Python where encoding works more cleanly.

More solutions are in the Ollama Troubleshooting guide.

Hardware, Costs, and Security with Ollama on Windows

Hardware

Performance depends heavily on your hardware. An Nvidia GPU with 8 GB of VRAM is sufficient for 7B models like Llama 3.1. For 13B models, you need at least 12 GB of VRAM. Without a GPU, Ollama runs on the CPU, which is significantly slower. More details in CPU vs. GPU.

Costs

Ollama is open source and free. Costs arise only from hardware and electricity. Downloading a 4 GB model costs nothing beyond your internet bandwidth.

Security

Ollama runs locally by default. The API is accessible only on localhost, not from the outside. If you want to expose the API to your network, set OLLAMA_HOST=0.0.0.0:11434. In that case, pay attention to firewall rules and authentication. All models and data stay on your machine. Nothing is sent to cloud services.

FAQ: Ollama on Windows - Common Questions

Do I need WSL2 to run Ollama on Windows?

No. The official Windows installer is sufficient. WSL2 is only necessary if you specifically want a Linux environment.

Does Ollama work without a GPU?

Yes, Ollama runs on CPU. Performance is noticeably lower, but adequate for small models and testing.

Where does Ollama store models on Windows?

By default, in C:\Users\YourName\.ollama\models. You can change the path using the OLLAMA_MODELS environment variable.

How do I verify my GPU is recognized?

For Nvidia GPUs, run nvidia-smi in PowerShell. If a table with GPU information appears, your GPU is ready.

Can I use Ollama across multiple Windows user accounts?

Yes, the service runs system-wide. However, models are stored in the user directory. Each user has their own model folder unless you set OLLAMA_MODELS to a shared path.

Does Ollama start automatically when Windows boots?

Yes, the installer registers Ollama as a service with automatic startup. You can change this in Windows Services.

Can I run Ollama and Docker Desktop simultaneously?

Yes, as long as there are no port conflicts. If you want to run Ollama in Docker, use a different port or stop the Windows service. See Ollama with Docker for details.

How do I change Ollama’s port?

Set the OLLAMA_HOST environment variable to something like 127.0.0.1:11435 and restart the service.

Are AMD GPUs supported?

Yes, starting with driver version 23.11.1 with ROCm support. Update your Adrenalin driver and verify that Ollama detects the GPU.

Is Ollama free on Windows?

Yes, Ollama is open source and free. Costs apply only to hardware and electricity.

Can I use Ollama from PowerShell?

Yes, all commands like ollama run, ollama pull, and ollama list work in PowerShell and the classic Command Prompt.

References and Further Reading

Back to Blog
Share:

Related Posts