Skip to content
BotServBotServ
OllamamacOSApple SiliconM1M2M3M4MetalUnified Memorylocal AI

Ollama on macOS: Installation and Setup

Install Ollama on macOS and Apple Silicon. Step-by-step guide with Unified Memory, Metal framework, and Mac tips.

S

schutzgeist

11 min read
Ollama on macOS: Installation and Setup

Ollama on macOS: Installation and Setup

What this article covers

  • Installing Ollama on macOS via DMG and Homebrew
  • Using Apple’s Metal framework on Apple Silicon for GPU-accelerated inference
  • Unified Memory on M-series chips and RAM requirements for different models
  • Configuring model storage location and autostart via LaunchAgent
  • Common pitfalls and solutions for Mac users

Introduction: Ollama on macOS explained

Ollama is a runtime environment that lets you run Large Language Models locally on your Mac. It works especially well on macOS because Apple Silicon features Unified Memory. This means CPU and GPU share the same RAM pool. You don’t need a separate graphics card with its own VRAM to load models with GPU acceleration. Ollama uses Apple’s Metal framework to automatically tap into the GPU capabilities of M-series chips. This guide walks you through installation, configuration, and first use on your Mac.

Why do you need this guide?

Mac users have a particular advantage when running AI models locally: Unified Memory. A Mac Studio with 64 GB RAM can load a 34-billion-parameter model entirely into shared memory and run it on the GPU. On a traditional PC, you’d need an expensive graphics card with sufficient VRAM. Many Mac users don’t realize how to harness this potential. While Ollama detects Apple Silicon automatically, you still need to know how much RAM your models require, how to adjust storage location, and how to verify the GPU is actually being used. That’s what this guide addresses.

Ollama on macOS at a glance

Ollama runs on macOS as a background process and provides a local API at http://localhost:11434. After installation, you download a model with ollama pull and start it with ollama run. On Apple Silicon, Ollama automatically uses the Metal framework for GPU acceleration. On Intel Macs, inference runs entirely on the CPU. You can install Ollama either via the official DMG download or through Homebrew.

Who this guide is for

This guide is aimed at Mac users installing and setting up Ollama for the first time. You don’t need prior experience with local AI. If you’ve already installed Ollama on another system, you’ll find macOS-specific details here. For platform-agnostic installation instructions, see Installing Ollama.

Key terms for Ollama on macOS

TermDefinition
OllamaRuntime environment for local language models
Apple SiliconApple’s custom chip family starting with M1, ARM-based architecture
Unified MemoryShared RAM pool for CPU and GPU on Apple Silicon
MetalApple’s graphics and compute framework for GPU acceleration
M-seriesApple’s chip designation starting with M1, including M2, M3, and M4
GPU TilesPhysical GPU units within an Apple Silicon chip
HomebrewPackage manager for macOS for installing software via terminal
LaunchAgentmacOS mechanism for automatically starting processes at login
OLLAMA_MODELSEnvironment variable for setting model storage location
Activity MonitormacOS tool for monitoring CPU, GPU, and RAM usage

Requirements

Before installing Ollama, verify the following:

  • macOS version: Ollama requires macOS 13.6 (Ventura) or later. Older versions don’t work reliably.
  • Apple Silicon vs. Intel: Ollama runs on both architectures. Apple Silicon gets GPU acceleration via Metal. Intel Macs use CPU only and are significantly slower.
  • RAM: Small models like Llama 3.2 with 3B parameters need 8 GB. For 8B parameter models, 16 GB is recommended. Larger models with 70B parameters require 64 GB or more. For detailed RAM requirements, see Unified Memory.
  • Storage: Each model takes disk space. An 8B model in Q4 quantization uses about 4 to 5 GB. Plan adequate storage, especially if you’re keeping multiple models.

Step 1: Download and install Ollama

The easiest way is to download the official app:

  1. Open ollama.com in your browser.
  2. Click the download button and select macOS.
  3. Download the .zip file and unzip it.
  4. Drag the Ollama app to the Applications folder.
  5. Start Ollama with a double-click.

On first launch, Ollama sets up the background service. You’ll see a small Ollama icon in the menu bar. The service now runs in the background and the API is accessible at http://localhost:11434.

Open Terminal and verify the installation:

ollama --version

If a version number appears, Ollama is ready to use.

Step 2: Install via Homebrew

Alternatively, you can use Homebrew. Homebrew is a package manager for macOS that lets you install software via Terminal. If you don’t have Homebrew yet, install it first:

/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"

Then install Ollama with a single command:

brew install ollama

Homebrew downloads the binaries and sets everything up. After installation, start the service manually:

ollama serve

The difference from the DMG install: With Homebrew, Ollama doesn’t automatically start as a background app. You either start the service manually or set up a LaunchAgent, which is covered below.

Step 3: Apple Silicon and Metal

On Apple Silicon, Ollama automatically uses the Metal framework. Metal is Apple’s programming interface for GPU compute and graphics. Ollama doesn’t need extra configuration to use the GPU. When you start each model, Ollama checks for a Metal-compatible GPU and loads the model into Unified Memory.

You can verify Metal is active by starting a model and watching the output. Ollama indicates whether the GPU is being used during load. Alternatively, check the Activity Monitor in the GPU tab to see GPU Tiles activity. If you see activity there while querying a model, Ollama is using Metal.

Learn more about the hardware foundation at Apple Silicon.

Step 4: Start your first model

After installation, download your first model. Llama 3.1 with 8B parameters is a good starting point:

ollama pull llama3.1

The download takes a few minutes depending on your connection. Then start the model:

ollama run llama3.1

Ollama opens a chat in your Terminal. You can ask questions immediately. Type /bye to end the session. For managing additional models, see Managing Ollama Models.

Step 5: Using the API

Ollama exposes a local HTTP API at http://localhost:11434. You can call Ollama directly from your own applications, scripts, or automation workflows.

Here’s a simple call with curl:

curl http://localhost:11434/api/generate -d '{
  "model": "llama3.1",
  "prompt": "Was ist lokale KI?",
  "stream": false
}'

In Python, it looks like this:

import requests

response = requests.post(
    "http://localhost:11434/api/generate",
    json={
        "model": "llama3.1",
        "prompt": "Erkläre Unified Memory in zwei Saetzen.",
        "stream": False
    }
)
print(response.json()["response"])

For detailed examples and a full list of endpoints, see Ollama API nutzen.

Apple Silicon Specifics

Apple Silicon has a feature that makes it particularly attractive for local AI: Unified Memory. The CPU and GPU access the same physical memory pool. There’s no copying of data between system RAM and VRAM, as you’d have with traditional GPUs.

For Ollama, this means the entire model must fit within your available RAM. A Mac with 16 GB of RAM can easily load an 8B model in Q4 quantization, since it occupies about 5 GB. A 70B model, however, requires around 40 GB and only runs smoothly on Macs with 64 GB of RAM or more.

To check GPU utilization:

  1. Open Activity Monitor via Spotlight or the Applications folder.
  2. Switch to the “GPU” tab.
  3. Submit a request to Ollama.
  4. Watch the GPU tile utilization climb.

If GPU utilization increases during inference, Ollama is using the Metal framework correctly. If only the CPU is busy, you have a configuration problem. Learn more about memory architecture at Unified Memory and general information under LLM lokal betreiben.

Intel Macs

Ollama works on Intel Macs, but without GPU acceleration. Inference runs exclusively on the CPU, which means significantly slower response times. An 8B model that answers in seconds on Apple Silicon might take 30 seconds or longer on an Intel Mac.

Additional limitations on Intel Macs:

  • No Metal support, so no GPU offloading
  • Higher power consumption during inference
  • Models compete with the operating system for RAM
  • Large models with 13B+ parameters are practically unusable

For Intel Macs, stick with smaller models like Llama 3.2 with 3B parameters. If you want to work seriously with local AI, upgrading to an Apple Silicon Mac makes sense. The Mac Mini is an affordable entry point, while the Mac Studio offers more RAM options for larger models.

Configuration: Storage Location and Autostart

By default, Ollama stores models in ~/.ollama/models. If your internal SSD has limited space, you can change the storage location using the OLLAMA_MODELS environment variable:

export OLLAMA_MODELS=/Volumes/External/ollama-models

To make this setting permanent, add it to your shell configuration, for example in ~/.zshrc:

echo 'export OLLAMA_MODELS=/Volumes/External/ollama-models' >> ~/.zshrc
source ~/.zshrc

To start Ollama automatically at login, set up a LaunchAgent. Create a property list file at ~/Library/LaunchAgents/com.ollama.plist:

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
    <key>Label</key>
    <string>com.ollama</string>
    <key>ProgramArguments</key>
    <array>
        <string>/usr/local/bin/ollama</string>
        <string>serve</string>
    </array>
    <key>RunAtLoad</key>
    <true/>
    <key>KeepAlive</key>
    <true/>
</dict>
</plist>

Then load the LaunchAgent:

launchctl load ~/Library/LaunchAgents/com.ollama.plist

Ollama will now start automatically on every login. Additional configuration options like port changes and network access are covered in Ollama Konfiguration.

Common Pitfalls with Ollama on macOS

  1. Not enough RAM for your chosen model: Ollama will load the model anyway but then uses swap memory. Inference becomes painfully slow. Check beforehand whether your RAM is sufficient for the model size.

  2. Metal is not being used: If Ollama only uses the CPU, the cause might be an outdated macOS version. Metal requires at least macOS 13.6. Update your system and restart Ollama.

  3. Model storage location on external SSD is unreachable: If you set OLLAMA_MODELS to an external drive that isn’t connected when Ollama starts, the service fails. Make sure the drive is mounted before Ollama launches.

  4. Port 11434 is already in use: If another process is using the port, the Ollama server won’t start. Check with lsof -i :11434 to see which process is blocking it, then stop it or change the Ollama port via OLLAMA_HOST.

  5. Homebrew path differs between Intel and Apple Silicon: On Intel Macs, Homebrew is at /usr/local/bin, while on Apple Silicon it’s at /opt/homebrew/bin. Adjust the path in your LaunchAgent accordingly, otherwise the service won’t find the Ollama binary.

  6. App sandbox blocks API access: If you want to call Ollama from another app and macOS blocks the connection, grant Ollama local network access in System Settings under “Privacy & Security”.

  7. Activity Monitor shows no GPU activity: Sometimes Activity Monitor doesn’t update the GPU display in real time. Use the terminal tools asitop or powermetrics instead to monitor GPU utilization more accurately.

Hardware, Costs, and Security with Ollama on macOS

Hardware: Ollama’s performance on macOS depends directly on your chip and RAM. An M2 with 16 GB of RAM works well for models up to 8B parameters. An M4 Max with 64 GB of RAM can run models up to 70B parameters smoothly. The number of GPU cores affects inference speed. More cores mean faster token generation.

Costs: Ollama itself is free and open source. You pay no licensing fees and no API costs. Your only expenses are hardware and power consumption. Apple Silicon is very efficient, with low power draw during inference compared to dedicated graphics cards.

Security: Ollama runs locally by default and is not accessible from the network. If you want to expose Ollama to other devices on your network, set OLLAMA_HOST=0.0.0.0:11434. In that case, use a firewall and authentication, since the endpoint would otherwise be unprotected. Models are downloaded by Ollama and stored locally. Your prompts and responses never leave your Mac.

Further Reading and Resources for Ollama on macOS

FAQ: Ollama on macOS, Common Questions

Does Ollama work on every Mac?

Ollama runs on Macs with macOS 13.6 or later. Apple Silicon gets GPU acceleration, while Intel Macs use the CPU. Both architectures are supported.

Do I need a separate graphics card?

No. On Apple Silicon, Ollama uses the integrated GPU through the Metal framework. Unified Memory replaces the separate VRAM of a dedicated graphics card.

How much RAM do I need for Ollama on macOS?

A 3B model runs on 8 GB. For 8B models, 16 GB is recommended. For 70B models, you’ll want at least 64 GB. A model’s size in Q4 quantization roughly equals 60 percent of its parameter count in GB.

How can I tell if the GPU is being used?

Open Activity Monitor and switch to the “GPU” tab. If you see activity during an Ollama request, Ollama is using Metal. Alternatively, you can use the terminal tool asitop.

Can I install Ollama via Homebrew instead of the app?

Yes. Run brew install ollama to install via terminal. The difference is that you’ll need to start the service manually with ollama serve or set up a LaunchAgent.

Where does Ollama store models on macOS?

By default, in ~/.ollama/models. You can change the location using the OLLAMA_MODELS environment variable, such as pointing to an external SSD.

Does Ollama start automatically on login?

If you use the official app, yes. With Homebrew, you’ll need to set up a LaunchAgent for automatic startup.

Does Ollama work on Intel Macs?

Yes, but only on the CPU. Inference is significantly slower than on Apple Silicon. For Intel Macs, stick with smaller models like Llama 3.2 with 3B parameters.

Can I make Ollama accessible from the network?

Yes, using the environment variable OLLAMA_HOST=0.0.0.0:11434. Be mindful of firewall settings and authentication, since the endpoint would otherwise be exposed.

Does Ollama consume a lot of power on Mac?

Apple Silicon is very efficient. Power consumption during inference is low compared to dedicated graphics cards. When idle, Ollama uses minimal resources as long as no model is loaded.

Can I load multiple models at the same time?

Yes, as long as you have enough RAM. Each loaded model occupies memory. If RAM runs out, macOS uses swap memory and performance drops sharply.

References and Further Reading

  • Ollama official website and download: ollama.com
  • Ollama GitHub repository with documentation
  • Apple Developer documentation for the Metal framework
  • Homebrew official website: brew.sh
  • Apple Silicon technical overviews of M-series chips
Back to Blog
Share:

Related Posts