Skip to content
BotServBotServ
LLMlocal AIOllamaAI installationmodelself-hosting

Run LLMs Locally on Your Hardware

Start a Large Language Model on your own hardware. Requirements, Ollama setup, model selection, and first commands.

S

schutzgeist

11 min read
Run LLMs Locally on Your Hardware

Running LLMs Locally

What this article covers

  • What hardware and software you need to run an LLM locally
  • How to install Ollama on Linux, macOS, and Windows
  • How to choose and start a suitable starter model
  • How to interact with the Ollama API using curl and Python
  • Alternative approaches like LM Studio and Docker

Introduction: Running LLMs locally explained

A Large Language Model, or LLM, powers chatbots, text assistants, and AI agents. Normally you access these models through an API, paying per token and sending your data to a cloud provider. It works, costs money, and raises privacy concerns.

With the right software and hardware, you can run an LLM on your own machine. The model runs locally, you keep control of your data, and you pay no API fees. This article shows you how to do this with Ollama, the simplest entry point. For more background, see What is local AI?.

Why run an LLM locally?

Four concrete reasons make local operation worthwhile:

  1. No API costs: You pay nothing per request. Once installed, the model runs without ongoing fees. This makes it attractive for experimentation, prototypes, and regular use.
  2. Data privacy: Sensitive documents, internal code, or personal text stays on your machine. Nothing leaves your network as long as you run Ollama locally.
  3. No rate limits: Cloud providers restrict how many requests you can send per minute. Locally, there are no such limits. You can submit hundreds of prompts in succession.
  4. Offline capability: The model works without an internet connection. This is useful while traveling, on secure networks, or during outages.

For a detailed comparison, see Local AI vs. Cloud AI.

Running an LLM locally in a nutshell

Running an LLM locally means the model file and inference runtime both run on your machine. You provide a prompt, your computer processes it, and returns an answer. The process requires three things: suitable hardware, a runtime environment, and a model in a compatible format.

Ollama is a popular entry point because it bundles model management, an API, and a chat interface in one package. To get started, you need a single command like ollama run llama3.1. Ollama downloads the model, unpacks it, and starts a local server that you can access from the terminal, a browser, or your own application.

Who is this for?

Local LLM operation targets several audiences:

  • Developers who want to test models without API keys and integrate them into their own applications
  • Privacy-conscious organizations that process sensitive documents without sending data externally
  • AI enthusiasts who want to experiment with custom workflows and understand how models work
  • Self-hosters who control their own infrastructure and want independence from cloud providers

Command-line experience is helpful but not required. Ollama reduces the entry barrier to just a few commands.

Key terms

TermMeaning
LLMLarge Language Model, a language model with many parameters
InferenceProcessing a prompt to generate a response
OllamaRuntime environment for downloading and running local models
GGUFA common format for quantized local models
ParametersModel weights, typically measured in billions
TokenA text chunk the model processes, roughly a word or part of a word
QuantizationReducing model size by lowering the bit depth per parameter
ModelfileA configuration file for Ollama that specifies the model, parameters, and system prompt
Context WindowThe maximum amount of text the model can process at once

For more on quantization, see Quantization. For details on context windows, see Context Length.

Prerequisites for running an LLM locally

Hardware

Your hardware determines which models you can run and how fast the responses come. Here’s an overview of typical starter setups:

SetupRAM / VRAMSuitable ModelsSpeed
Minimum8 GBphi3 (3.8B), small modelsSlow, CPU only
Standard16 GBllama3.1 (8B), qwen2.5 (7B)Acceptable, faster with GPU
Comfortable32 GBmistral (7B) quantized, llama3.1 (8B) smoothFast with GPU
High-performance64 GB+Larger models up to 70B quantizedVery fast with strong GPU

A GPU significantly accelerates inference. Nvidia cards have the best support, while Apple Silicon (M1 to M4) also performs very well thanks to Unified Memory. Without a GPU, everything runs on the CPU, which is still acceptable for small models but becomes frustratingly slow for larger ones.

See RAM and VRAM requirements and AI hardware basics for detailed specifications.

Software

  • Operating system: Linux, macOS, or Windows 10/11
  • Ollama: The runtime environment, free to download
  • Optional: Docker, if you want to run Ollama containerized
  • Optional: Python 3.8 or later, if you want to access the API via script

Step 1: Check your hardware

Before installing Ollama, verify what your machine can handle. This prevents disappointment later when a model doesn’t run smoothly.

Check RAM

On Linux and macOS:

free -h

On macOS, you can also open Apple menu “About This Mac” to find your memory.

On Windows:

wmic memorychip get capacity

Or open Task Manager, go to the “Performance” tab, and check used and total RAM.

Check VRAM

If you have an Nvidia GPU, use nvidia-smi:

nvidia-smi

This command shows the GPU name, VRAM, and current load. For AMD GPUs, try rocm-smi. For Apple Silicon, system_profiler SPDisplaysDataType shows the details.

Disk space

Models require storage. An 8B model in quantized form takes about 4 to 6 GB. Plan for at least 20 GB of free space if you want to try multiple models.

Step 2: Install Ollama

Ollama is available for Linux, macOS, and Windows. Installation varies slightly by platform. See Ollama installation for a complete guide.

Linux

The fastest way is the official installation script:

curl -fsSL https://ollama.com/install.sh | sh

The script downloads the latest version, sets up the service, and starts it automatically. Verify Ollama is running:

ollama --version

If you prefer manual installation, download the binary from the Ollama website, move it to /usr/local/bin/, and set up a systemd service.

macOS

Download the installer from ollama.com, open the .dmg file, and drag Ollama into the Applications folder. Launch Ollama once, and it will run in the background. Check in the terminal:

ollama --version

macOS with Apple Silicon automatically uses Unified Memory. No extra configuration needed.

Windows

Download the installer from ollama.com and run it. Ollama sets itself up as a background service. Open Command Prompt or PowerShell and verify:

ollama --version

If you use WSL2, you can also install Ollama inside WSL2, as described above under Linux.

Step 3: Choose a Model

Your hardware determines which model to pick. Start with a small model that runs smoothly and responds quickly. For a full list of available models, see Model Management.

ModelParametersRAM Requirement (quantized)Strengths
llama3.18B~5 GBAll-rounder, strong German support
qwen2.57B~4.5 GBExcellent at code and logic
mistral7B~4.5 GBCompact, fast, reliable
phi33.8B~2.5 GBVery small, runs on limited hardware

Begin with llama3.1 or phi3 if your machine has limited RAM. You can always install and compare multiple models later.

Step 4: Run a Model

Once Ollama is installed, start a model with a single command:

ollama run llama3.1

On first run, Ollama downloads the model. This takes a few minutes depending on your internet connection. Then an interactive chat opens in the terminal. Enter a prompt, press Enter, and wait for the response.

Example prompt:

Erkläre mir in drei Sätzen, was ein LLM ist.

To end the chat, type /bye or press Ctrl+D.

If you want to download a model without starting it immediately:

ollama pull llama3.1

This is useful when you want to load multiple models in advance.

Step 5: Use the API

Ollama automatically starts a local server at http://localhost:11434. You can call the model from your own applications through this endpoint. For more details, see Ollama API.

curl Example

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "Was ist lokale KI?",
  "stream": false
}'

With "stream": false, you get the complete response at once. Without it, Ollama streams the response in small JSON chunks, which is useful for live output.

Python Example

import requests

response = requests.post(
    "http://localhost:11434/api/generate",
    json={
        "model": "llama3.1",
        "prompt": "Was ist lokale KI?",
        "stream": False
    }
)

print(response.json()["response"])

For more convenient access, use the official Python library ollama:

pip install ollama
import ollama

response = ollama.generate(
    model="llama3.1",
    prompt="Was ist lokale KI?"
)

print(response["response"])

Alternative: LM Studio

LM Studio is a graphical alternative to Ollama. It provides a desktop interface for Windows, macOS, and Linux where you download models with a click, start them, and test them in a chat.

Advantages of LM Studio:

  • Intuitive interface without the command line
  • Model search built into the app via Hugging Face
  • Integrated chat with Markdown support
  • Compatible API server you can use like Ollama

Disadvantages:

  • Closed source, unlike Ollama
  • Less control over model parameters
  • Harder to automate

For beginners uncomfortable with the command line, LM Studio is a good choice. If you plan to write scripts and automation, Ollama is the better fit.

Alternative: Docker

If you want to run Ollama containerized, such as on a server or NAS, use the official Docker image:

docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

Then pull a model into the container:

docker exec -it ollama ollama pull llama3.1

And run it:

docker exec -it ollama ollama run llama3.1

For GPU support, add --gpus all:

docker run -d --gpus all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

Docker is especially useful if you want to run Ollama permanently on a server or already have other services containerized.

Ollama Commands at a Glance

CommandDescription
ollama run modelDownload and run a model interactively
ollama pull modelDownload a model without starting it
ollama listShow all installed models
ollama rm modelRemove a model from your system
ollama show modelShow details about a model
ollama psShow running models
ollama serveStart the API server
ollama createCreate a custom model from a Modelfile
ollama stop modelStop a running model
ollama --versionShow the installed version

For a complete overview, see Ollama Overview.

Common Pitfalls When Running Local LLMs

  1. Insufficient RAM: An 8B model needs at least 5 GB of free RAM. If your system already uses most of its memory, inference will crash or become extremely slow. Close other programs before starting a model.
  2. Wrong model size: Beginners often grab large models like 70B, which don’t run smoothly on normal hardware. Start small and work your way up.
  3. No GPU support in Docker: Without the --gpus all parameter, Ollama in a container runs only on CPU. Check that the Nvidia Container Toolkit is installed.
  4. Confusing pull and run: ollama pull only downloads, ollama run starts the chat. Running pull and waiting for a response won’t work.
  5. API unreachable: If the server isn’t running, every API call fails. Start Ollama with ollama serve or ensure the background service is active.
  6. Disk full: Models are large. Multiple 8B models quickly take up 20 GB or more. Delete old models with ollama rm.
  7. Streaming confuses beginners: The API sends many small JSON chunks by default. Set "stream": false if you need the complete response at once.

Hardware, Cost, and Security for Running Local LLMs

Hardware

Getting started usually doesn’t require new hardware. Many standard computers with 16 GB of RAM or a Mac with M-series processors work fine. If you regularly use larger models, invest in a dedicated GPU. An Nvidia card with 12 GB VRAM, such as an RTX 3060, is a solid middle ground.

Cost

Expenses stay reasonable. Ollama is free, as are most open-source models. The main cost factors are electricity and potentially new hardware. A GPU draws significantly more power under load, which adds up with continuous operation.

Security

From a security perspective, all data stays local as long as you run Ollama in your own network. No prompt or response leaves your machine. This is a significant advantage over cloud providers. If you run Ollama in a Docker container and expose the port externally, secure the connection with a firewall or reverse proxy.

Further Reading and Resources on Local LLM Operation

FAQ - Common Questions About Local LLM Operation

Can I have multiple models installed?

Yes. Ollama manages any number of models. You switch between them with ollama run modellname. Use ollama list to see all installed models.

Do I have to download each time?

No. A model stays stored locally after the first download. Only updates or new models need to be loaded again.

Does Ollama run without a graphics card?

Yes, on the CPU. Response speed is considerably slower than on a GPU. For small models like phi3 it’s still acceptable, but larger models become frustrating.

Where do I find suitable models?

Ollama offers a model library on its website. Alternatively, you can find models on Hugging Face, often in GGUF format. Download them directly with ollama pull modellname.

Can I use Ollama in Docker?

Yes. Official Docker images are available and particularly useful for server or NAS environments. For GPU support, you need the Nvidia Container Toolkit and the —gpus all parameter.

How much RAM do I need at minimum?

Small models like phi3 work with 8 GB. For 8B models like llama3.1 you should have 16 GB so the operating system and model can run simultaneously.

What’s the difference between pull and run?

ollama pull downloads the model only. ollama run downloads it if not already present and starts the interactive chat directly.

Can I access Ollama from other devices on the network?

Yes. By default Ollama listens on localhost. You can set the OLLAMA_HOST environment variable to 0.0.0.0:11434 to make Ollama reachable from the network. Secure access with a firewall.

How do I update a model?

Run ollama pull modellname again. Ollama downloads the latest version if available. The old version gets overwritten.

What is quantization and do I need it?

Quantization shrinks a model by reducing the storage depth per parameter. Ollama uses quantized models by default, so you don’t need to worry about it. Learn more in the Quantization article.

Can I create my own models?

Yes. With a Modelfile you define the base model, system prompt, and parameters. Create a file named Modelfile and run ollama create meinmodell -f Modelfile.

How do I stop a running model?

Use ollama stop modellname to unload the model from memory. It stays saved on disk and can be restarted anytime.

Is my data sent to a server?

No. All processing happens locally on your machine. Neither prompts nor responses leave your system as long as you don’t deliberately expose Ollama on the network.

Can I use Ollama with an AMD GPU?

Yes, with limitations. Ollama supports AMD GPUs through ROCm, primarily on Linux. Windows support is limited. Check the current documentation for compatible models.

Sources and Further Reading

Back to Blog
Share:

Nächster Artikel in Local AI

Weiterlesen
What is Local AI?

Related Posts