Skip to content
BotServBotServ
vLLMPagedAttentionContinuous BatchingInferenceServerOpenAI APILocal AI

vLLM: High-Performance LLM Inference

What is vLLM? Fast inference engine for production with PagedAttention, Continuous Batching, and OpenAI-compatible API.

S

schutzgeist

10 min read
vLLM: High-Performance LLM Inference

vLLM: High-Performance Inference for Servers

What this article covers

  • What vLLM is and when to use it instead of Ollama or llama.cpp.
  • How PagedAttention and Continuous Batching boost performance.
  • Step-by-step installation and starting a vLLM server.
  • Calling the OpenAI-compatible API with practical examples.
  • Performance tuning, hardware requirements, and common pitfalls.

Introduction: vLLM explained

Running a local LLM that needs to serve multiple users simultaneously? You’ll quickly outgrow tools optimized for single requests. vLLM fills that gap. It’s an inference engine built specifically for high throughput and low latency on server hardware.

vLLM was introduced in 2023 by researchers at UC Berkeley and quickly became the standard for production LLM serving. The core idea: use expensive GPU VRAM as efficiently as possible to handle many requests in parallel.

Why do you need vLLM?

Imagine running an internal chatbot for a 50-person team. Employees ask questions, often simultaneously. With a simple solution like Ollama, requests are processed sequentially. With 10 concurrent users, the last ones wait minutes for their answer.

vLLM solves this. It processes dozens of requests in parallel on a single GPU without each one reserving the entire KV-Cache. The result: higher throughput, shorter wait times, happier users.

Typical vLLM scenarios:

  • Production chatbots for teams or customers.
  • API backend serving multiple applications at once.
  • Batch processing large volumes of text.
  • Load distribution across multiple GPUs with Tensor Parallelism.

If you’re just experimenting locally with a model, llama.cpp or Ollama are better choices. vLLM shines under load.

vLLM at a glance

vLLM is a Python library and inference engine that runs large language models on GPUs. It provides an OpenAI-compatible API, so you can use existing clients and libraries without modification. The focus is efficiency: more tokens per second with the same hardware budget.

Who is vLLM for?

vLLM targets developers and operators bringing LLMs to production. If you have a GPU in a server and need to serve multiple users or applications, vLLM is the right choice. For beginners experimenting locally, it’s overkill and more complex to set up than Ollama.

Quick overview of prerequisites:

  • An NVIDIA or AMD GPU with sufficient VRAM.
  • Linux as the operating system, Windows only via WSL2.
  • Basic Python and command-line skills.
  • Understanding of GPU memory and quantization.

Key terms around vLLM

TermExplanation
vLLMInference engine for high throughput on GPUs
PagedAttentionMemory management for the KV-Cache, similar to virtual memory in operating systems
Continuous BatchingRequests are dynamically added to running batches without waiting for the entire batch to complete
ThroughputNumber of tokens processed per second across all requests
LatencyTime until the first tokens of a single request return
OpenAI APICompatible HTTP interface matching OpenAI’s endpoints
Tensor ParallelismDistributing a model across multiple GPUs
QuantizationReducing the precision of weights to save VRAM
KV-CacheBuffer for keys and values during text generation
QPSQueries Per Second, a measure of processed requests per second

What makes vLLM special?

PagedAttention for efficient memory

The KV-Cache is the biggest memory consumer during inference. Without optimization, vLLM reserves a contiguous memory block for each request, sized for the maximum possible sequence length. This leads to enormous waste since most requests are much shorter.

PagedAttention solves this by dividing the KV-Cache into small blocks (pages), similar to how an operating system manages virtual memory. A request only occupies as many pages as it actually needs. As the sequence grows, new pages are allocated. This reduces memory waste to under 4 percent.

Continuous Batching for high throughput

Traditional batching waits for all requests in a batch to finish before starting the next one. A long request blocks all shorter ones in the same batch. Continuous Batching dynamically adds new requests at each step and removes completed ones immediately. The GPU stays continuously busy.

Together, these techniques allow vLLM to increase throughput many times over compared to naive inference. Benchmarks often show 2x to 4x higher throughput versus HuggingFace Transformers on the same model and hardware.

Installation

vLLM requires a CUDA-capable NVIDIA GPU or a ROCm-capable AMD GPU. Details on choosing hardware can be found in the article on CPU vs. GPU.

Prerequisites:

  • Linux (Ubuntu 20.04 or newer recommended)
  • NVIDIA GPU with Compute Capability 7.0 or higher
  • CUDA 12.1 or newer
  • Python 3.9 to 3.12

Installation via pip:

pip install vllm

For AMD GPUs, use the additional index:

pip install vllm --extra-index-url https://download.pytorch.org/whl/rocm6.2

After installation, verify that vLLM detects your GPU:

python -c "import vllm; print(vllm.__version__)"

Starting the server

vLLM includes a built-in server that provides an OpenAI-compatible API. Start it via a Python module:

python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --host 0.0.0.0 \
  --port 8000 \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192

Key arguments:

ArgumentPurpose
--modelModel name on HuggingFace or local path
--hostNetwork interface, 0.0.0.0 for external access
--portServer port, default 8000
--gpu-memory-utilizationFraction of VRAM vLLM can use, default 0.9
--max-model-lenMaximum context length in tokens
--tensor-parallel-sizeNumber of GPUs for Tensor Parallelism
--quantizationQuantization method, e.g. awq or gptq

On first startup, the model loads into VRAM and initializes the KV-Cache. This takes anywhere from seconds to minutes depending on model size.

Using the API

Once the server is running, call it like the OpenAI API. A simple chat request with curl:

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer token-abc123" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [
      {"role": "user", "content": "Explain PagedAttention in three sentences."}
    ],
    "temperature": 0.7
  }'

The response follows the familiar OpenAI format:

{
  "id": "chat-abc123",
  "object": "chat.completion",
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "PagedAttention divides the KV-Cache into small blocks..."
      }
    }
  ]
}

The completions endpoint is also supported:

curl http://localhost:8000/v1/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "prompt": "The capital of France is",
    "max_tokens": 5
  }'

List available models:

curl http://localhost:8000/v1/models

In Python, use the official openai library with an adjusted base URL:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="token-abc123"
)

response = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "What is Continuous Batching?"}]
)
print(response.choices[0].message.content)

Performance Tuning

Tensor Parallelism

When models don’t fit on a single GPU, you can distribute them across multiple GPUs. Set --tensor-parallel-size to the number of available GPUs:

python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-70B-Instruct \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.9

Your GPUs need strong interconnect via NVLink or PCIe, otherwise communication becomes a bottleneck. Learn more about connection speed in the article on memory bandwidth.

GPU Memory Utilization

The --gpu-memory-utilization parameter controls how much VRAM vLLM can reserve. The default value of 0.9 means 90 percent. On a server with 24 GB VRAM, that’s roughly 21.6 GB for model weights and KV-cache.

Lower this value if other processes need VRAM. Only increase it when vLLM is the sole GPU application running.

Max Model Length

--max-model-len caps the maximum token count per request. A higher value enables longer contexts but reserves more memory for the KV-cache. Choose the smallest value your use case actually needs.

Quantization

Quantization reduces model size and VRAM requirements significantly. vLLM supports AWQ, GPTQ, and other methods:

python -m vllm.entrypoints.openai.api_server \
  --model TheBloke/Llama-3.1-8B-Instruct-AWQ \
  --quantization awq

An 8B model in FP16 needs roughly 16 GB VRAM. With AWQ quantization, that drops to around 6 GB, making it feasible on smaller GPUs.

vLLM vs. Ollama vs. llama.cpp

FeaturevLLMOllamallama.cpp
Primary useProduction, multi-userLocal experimentationResource-efficient inference
Multi-user performanceVery highModerateLow to moderate
Setup easeModerateVery easyEasy
GPU supportNVIDIA, AMDNVIDIA, AMD, MacNVIDIA, AMD, CPU
Continuous batchingYesLimitedNo
OpenAI APIYes, nativeYes, compatibleVia external server
CPU inferenceNoYesYes
Best forServers, teamsSingle workstationLow-end hardware

The right choice depends on your use case. For a production server serving multiple users, vLLM is the best option. For quick local tests, try Ollama. For CPU-only systems, use llama.cpp. See the running LLMs locally article for a deeper comparison.

Example: vLLM Server for a Team

A practical setup for a small team with one NVIDIA GPU (24 GB VRAM).

Step 1: Create a virtual environment.

python -m venv vllm-env
source vllm-env/bin/activate
pip install vllm

Step 2: Pick a quantized model that fits in your VRAM. For 24 GB, an 8B model in AWQ works well.

Step 3: Start the server with sensible defaults.

python -m vllm.entrypoints.openai.api_server \
  --model casperhansen/llama-3-8b-instruct-awq \
  --quantization awq \
  --host 0.0.0.0 \
  --port 8000 \
  --gpu-memory-utilization 0.85 \
  --max-model-len 4096 \
  --api-key token-team-2026

Step 4: Set up the server as a systemd service so it starts automatically after reboot.

# /etc/systemd/system/vllm.service
[Unit]
Description=vLLM Server
After=network.target

[Service]
Type=simple
User=ubuntu
WorkingDirectory=/home/ubuntu
ExecStart=/home/ubuntu/vllm-env/bin/python -m vllm.entrypoints.openai.api_server --model casperhansen/llama-3-8b-instruct-awq --quantization awq --host 0.0.0.0 --port 8000 --gpu-memory-utilization 0.85 --max-model-len 4096 --api-key token-team-2026
Restart=on-failure

[Install]
WantedBy=multi-user.target

Enable and start:

sudo systemctl daemon-reload
sudo systemctl enable vllm
sudo systemctl start vllm

Step 5: Distribute the API endpoint to your team. Everyone uses the base URL http://server-ip:8000/v1 with the API key you’ve assigned.

Common vLLM Pitfalls

  1. Insufficient VRAM. A 70B model in FP16 requires over 140 GB VRAM. Check whether the model fits in your available memory before starting, and use quantization if needed.

  2. Wrong CUDA version. vLLM requires CUDA 12.1 or newer. An outdated CUDA installation leads to cryptic import errors. Verify with nvcc --version.

  3. Windows without WSL2. vLLM does not run natively on Windows. Use WSL2 with Ubuntu, or installation will fail.

  4. Tensor parallelism without NVLink. Multiple GPUs connected via slow interconnects severely degrade performance. Check your link with nvidia-smi nvlink -s.

  5. Forgotten API key. Without --api-key, the server accepts requests without authentication. In production, that’s a security risk.

  6. Max model length set too high. A value of 128000 reserves enormous KV-cache pools and leaves little room for model weights. Lower it to what you actually need.

  7. Base model without instruction tuning. Untuned base models produce poor results on the chat endpoint. Always pick instruct or chat versions.

  8. Downloads block startup. On first run, vLLM downloads the model from HuggingFace. Slow connections can take a long time. Pre-download the model with huggingface-cli and pass the local path instead.

Hardware, Costs, and Security with vLLM

Hardware

vLLM absolutely requires a GPU. For entry-level setups, an NVIDIA RTX 4090 with 24 GB VRAM suffices. For larger models, you need server-grade GPUs like the A100 or H100 with 40 to 80 GB VRAM, possibly multiple for tensor parallelism.

Costs

The GPU is the biggest cost factor. A used RTX 4090 runs around 1000 euros. A cloud server with an A100 costs several hundred euros monthly. Factor in power consumption, a RTX 4090 draws roughly 450 watts under load. Read more about hardware selection in What is local AI?.

Security

Always set an API key with --api-key. Don’t expose the server to the internet without firewall rules. Use a reverse proxy like Nginx with TLS if you make the endpoint publicly accessible. Log access to detect abuse.

Further Reading and vLLM Resources

FAQ: vLLM - Common Questions

What is vLLM?

vLLM is an inference engine for large language models on GPUs. It optimizes throughput and memory efficiency through PagedAttention and continuous batching.

Is vLLM free?

Yes, vLLM is open source under the Apache 2.0 license and free. Costs only arise from hardware and electricity.

Can I use vLLM on CPU?

No, vLLM requires a GPU with CUDA or ROCm. For CPU inference, use llama.cpp.

Do I need multiple GPUs for vLLM?

No, one GPU is enough for models that fit in its VRAM. You only need multiple GPUs for large models or higher throughput.

What’s the difference between vLLM and Ollama?

vLLM is optimized for production and high throughput, while Ollama focuses on easy local use. For multi-user scenarios, vLLM is the better choice.

Does vLLM support the OpenAI API?

Yes, vLLM ships with a server that natively implements OpenAI endpoints for chat completions and completions.

Which models work with vLLM?

Any model in HuggingFace format that vLLM supports. This includes Llama, Mistral, Qwen, Phi, and many others. Quantized AWQ and GPTQ variants are also supported.

How much VRAM do I need for vLLM?

At least as much as the model weights require, plus space for the KV-cache. An 8B model in FP16 needs roughly 16 GB, around 6 GB with AWQ. Always reserve headroom for the KV-cache.

Can I use vLLM on Windows?

Only via WSL2 with Ubuntu. Native Windows support is not available.

What is PagedAttention?

PagedAttention is vLLM’s memory management strategy for the KV-cache. It divides the cache into small blocks, similar to virtual memory in operating systems, and drastically reduces waste.

Does vLLM work with AMD GPUs?

Yes, vLLM supports AMD GPUs via ROCm. Installation requires adding the ROCm index to pip.

How do I start vLLM automatically after reboot?

Set up a systemd service that runs the vLLM server on system startup. You’ll find an example in the team setup section above.

References and Further Reading

  • Kwon, W. et al. (2023): Efficient Memory Management for Large Language Model Serving with PagedAttention. arXiv:2309.06180.
  • vLLM Project: Official documentation, GitHub repository.
  • NVIDIA: CUDA Toolkit Documentation.
  • HuggingFace: Model Hub and Transformers documentation.
Back to Blog
Share:

Related Posts