vLLM: High-Performance Inference for Servers
What this article covers
- What vLLM is and when to use it instead of Ollama or llama.cpp.
- How PagedAttention and Continuous Batching boost performance.
- Step-by-step installation and starting a vLLM server.
- Calling the OpenAI-compatible API with practical examples.
- Performance tuning, hardware requirements, and common pitfalls.
Introduction: vLLM explained
Running a local LLM that needs to serve multiple users simultaneously? You’ll quickly outgrow tools optimized for single requests. vLLM fills that gap. It’s an inference engine built specifically for high throughput and low latency on server hardware.
vLLM was introduced in 2023 by researchers at UC Berkeley and quickly became the standard for production LLM serving. The core idea: use expensive GPU VRAM as efficiently as possible to handle many requests in parallel.
Why do you need vLLM?
Imagine running an internal chatbot for a 50-person team. Employees ask questions, often simultaneously. With a simple solution like Ollama, requests are processed sequentially. With 10 concurrent users, the last ones wait minutes for their answer.
vLLM solves this. It processes dozens of requests in parallel on a single GPU without each one reserving the entire KV-Cache. The result: higher throughput, shorter wait times, happier users.
Typical vLLM scenarios:
- Production chatbots for teams or customers.
- API backend serving multiple applications at once.
- Batch processing large volumes of text.
- Load distribution across multiple GPUs with Tensor Parallelism.
If you’re just experimenting locally with a model, llama.cpp or Ollama are better choices. vLLM shines under load.
vLLM at a glance
vLLM is a Python library and inference engine that runs large language models on GPUs. It provides an OpenAI-compatible API, so you can use existing clients and libraries without modification. The focus is efficiency: more tokens per second with the same hardware budget.
Who is vLLM for?
vLLM targets developers and operators bringing LLMs to production. If you have a GPU in a server and need to serve multiple users or applications, vLLM is the right choice. For beginners experimenting locally, it’s overkill and more complex to set up than Ollama.
Quick overview of prerequisites:
- An NVIDIA or AMD GPU with sufficient VRAM.
- Linux as the operating system, Windows only via WSL2.
- Basic Python and command-line skills.
- Understanding of GPU memory and quantization.
Key terms around vLLM
| Term | Explanation |
|---|---|
| vLLM | Inference engine for high throughput on GPUs |
| PagedAttention | Memory management for the KV-Cache, similar to virtual memory in operating systems |
| Continuous Batching | Requests are dynamically added to running batches without waiting for the entire batch to complete |
| Throughput | Number of tokens processed per second across all requests |
| Latency | Time until the first tokens of a single request return |
| OpenAI API | Compatible HTTP interface matching OpenAI’s endpoints |
| Tensor Parallelism | Distributing a model across multiple GPUs |
| Quantization | Reducing the precision of weights to save VRAM |
| KV-Cache | Buffer for keys and values during text generation |
| QPS | Queries Per Second, a measure of processed requests per second |
What makes vLLM special?
PagedAttention for efficient memory
The KV-Cache is the biggest memory consumer during inference. Without optimization, vLLM reserves a contiguous memory block for each request, sized for the maximum possible sequence length. This leads to enormous waste since most requests are much shorter.
PagedAttention solves this by dividing the KV-Cache into small blocks (pages), similar to how an operating system manages virtual memory. A request only occupies as many pages as it actually needs. As the sequence grows, new pages are allocated. This reduces memory waste to under 4 percent.
Continuous Batching for high throughput
Traditional batching waits for all requests in a batch to finish before starting the next one. A long request blocks all shorter ones in the same batch. Continuous Batching dynamically adds new requests at each step and removes completed ones immediately. The GPU stays continuously busy.
Together, these techniques allow vLLM to increase throughput many times over compared to naive inference. Benchmarks often show 2x to 4x higher throughput versus HuggingFace Transformers on the same model and hardware.
Installation
vLLM requires a CUDA-capable NVIDIA GPU or a ROCm-capable AMD GPU. Details on choosing hardware can be found in the article on CPU vs. GPU.
Prerequisites:
- Linux (Ubuntu 20.04 or newer recommended)
- NVIDIA GPU with Compute Capability 7.0 or higher
- CUDA 12.1 or newer
- Python 3.9 to 3.12
Installation via pip:
pip install vllm
For AMD GPUs, use the additional index:
pip install vllm --extra-index-url https://download.pytorch.org/whl/rocm6.2
After installation, verify that vLLM detects your GPU:
python -c "import vllm; print(vllm.__version__)"
Starting the server
vLLM includes a built-in server that provides an OpenAI-compatible API. Start it via a Python module:
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 8000 \
--gpu-memory-utilization 0.9 \
--max-model-len 8192
Key arguments:
| Argument | Purpose |
|---|---|
--model | Model name on HuggingFace or local path |
--host | Network interface, 0.0.0.0 for external access |
--port | Server port, default 8000 |
--gpu-memory-utilization | Fraction of VRAM vLLM can use, default 0.9 |
--max-model-len | Maximum context length in tokens |
--tensor-parallel-size | Number of GPUs for Tensor Parallelism |
--quantization | Quantization method, e.g. awq or gptq |
On first startup, the model loads into VRAM and initializes the KV-Cache. This takes anywhere from seconds to minutes depending on model size.
Using the API
Once the server is running, call it like the OpenAI API. A simple chat request with curl:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer token-abc123" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [
{"role": "user", "content": "Explain PagedAttention in three sentences."}
],
"temperature": 0.7
}'
The response follows the familiar OpenAI format:
{
"id": "chat-abc123",
"object": "chat.completion",
"choices": [
{
"message": {
"role": "assistant",
"content": "PagedAttention divides the KV-Cache into small blocks..."
}
}
]
}
The completions endpoint is also supported:
curl http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"prompt": "The capital of France is",
"max_tokens": 5
}'
List available models:
curl http://localhost:8000/v1/models
In Python, use the official openai library with an adjusted base URL:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="token-abc123"
)
response = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "What is Continuous Batching?"}]
)
print(response.choices[0].message.content)
Performance Tuning
Tensor Parallelism
When models don’t fit on a single GPU, you can distribute them across multiple GPUs. Set --tensor-parallel-size to the number of available GPUs:
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-70B-Instruct \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.9
Your GPUs need strong interconnect via NVLink or PCIe, otherwise communication becomes a bottleneck. Learn more about connection speed in the article on memory bandwidth.
GPU Memory Utilization
The --gpu-memory-utilization parameter controls how much VRAM vLLM can reserve. The default value of 0.9 means 90 percent. On a server with 24 GB VRAM, that’s roughly 21.6 GB for model weights and KV-cache.
Lower this value if other processes need VRAM. Only increase it when vLLM is the sole GPU application running.
Max Model Length
--max-model-len caps the maximum token count per request. A higher value enables longer contexts but reserves more memory for the KV-cache. Choose the smallest value your use case actually needs.
Quantization
Quantization reduces model size and VRAM requirements significantly. vLLM supports AWQ, GPTQ, and other methods:
python -m vllm.entrypoints.openai.api_server \
--model TheBloke/Llama-3.1-8B-Instruct-AWQ \
--quantization awq
An 8B model in FP16 needs roughly 16 GB VRAM. With AWQ quantization, that drops to around 6 GB, making it feasible on smaller GPUs.
vLLM vs. Ollama vs. llama.cpp
| Feature | vLLM | Ollama | llama.cpp |
|---|---|---|---|
| Primary use | Production, multi-user | Local experimentation | Resource-efficient inference |
| Multi-user performance | Very high | Moderate | Low to moderate |
| Setup ease | Moderate | Very easy | Easy |
| GPU support | NVIDIA, AMD | NVIDIA, AMD, Mac | NVIDIA, AMD, CPU |
| Continuous batching | Yes | Limited | No |
| OpenAI API | Yes, native | Yes, compatible | Via external server |
| CPU inference | No | Yes | Yes |
| Best for | Servers, teams | Single workstation | Low-end hardware |
The right choice depends on your use case. For a production server serving multiple users, vLLM is the best option. For quick local tests, try Ollama. For CPU-only systems, use llama.cpp. See the running LLMs locally article for a deeper comparison.
Example: vLLM Server for a Team
A practical setup for a small team with one NVIDIA GPU (24 GB VRAM).
Step 1: Create a virtual environment.
python -m venv vllm-env
source vllm-env/bin/activate
pip install vllm
Step 2: Pick a quantized model that fits in your VRAM. For 24 GB, an 8B model in AWQ works well.
Step 3: Start the server with sensible defaults.
python -m vllm.entrypoints.openai.api_server \
--model casperhansen/llama-3-8b-instruct-awq \
--quantization awq \
--host 0.0.0.0 \
--port 8000 \
--gpu-memory-utilization 0.85 \
--max-model-len 4096 \
--api-key token-team-2026
Step 4: Set up the server as a systemd service so it starts automatically after reboot.
# /etc/systemd/system/vllm.service
[Unit]
Description=vLLM Server
After=network.target
[Service]
Type=simple
User=ubuntu
WorkingDirectory=/home/ubuntu
ExecStart=/home/ubuntu/vllm-env/bin/python -m vllm.entrypoints.openai.api_server --model casperhansen/llama-3-8b-instruct-awq --quantization awq --host 0.0.0.0 --port 8000 --gpu-memory-utilization 0.85 --max-model-len 4096 --api-key token-team-2026
Restart=on-failure
[Install]
WantedBy=multi-user.target
Enable and start:
sudo systemctl daemon-reload
sudo systemctl enable vllm
sudo systemctl start vllm
Step 5: Distribute the API endpoint to your team. Everyone uses the base URL http://server-ip:8000/v1 with the API key you’ve assigned.
Common vLLM Pitfalls
-
Insufficient VRAM. A 70B model in FP16 requires over 140 GB VRAM. Check whether the model fits in your available memory before starting, and use quantization if needed.
-
Wrong CUDA version. vLLM requires CUDA 12.1 or newer. An outdated CUDA installation leads to cryptic import errors. Verify with
nvcc --version. -
Windows without WSL2. vLLM does not run natively on Windows. Use WSL2 with Ubuntu, or installation will fail.
-
Tensor parallelism without NVLink. Multiple GPUs connected via slow interconnects severely degrade performance. Check your link with
nvidia-smi nvlink -s. -
Forgotten API key. Without
--api-key, the server accepts requests without authentication. In production, that’s a security risk. -
Max model length set too high. A value of 128000 reserves enormous KV-cache pools and leaves little room for model weights. Lower it to what you actually need.
-
Base model without instruction tuning. Untuned base models produce poor results on the chat endpoint. Always pick instruct or chat versions.
-
Downloads block startup. On first run, vLLM downloads the model from HuggingFace. Slow connections can take a long time. Pre-download the model with
huggingface-cliand pass the local path instead.
Hardware, Costs, and Security with vLLM
Hardware
vLLM absolutely requires a GPU. For entry-level setups, an NVIDIA RTX 4090 with 24 GB VRAM suffices. For larger models, you need server-grade GPUs like the A100 or H100 with 40 to 80 GB VRAM, possibly multiple for tensor parallelism.
Costs
The GPU is the biggest cost factor. A used RTX 4090 runs around 1000 euros. A cloud server with an A100 costs several hundred euros monthly. Factor in power consumption, a RTX 4090 draws roughly 450 watts under load. Read more about hardware selection in What is local AI?.
Security
Always set an API key with --api-key. Don’t expose the server to the internet without firewall rules. Use a reverse proxy like Nginx with TLS if you make the endpoint publicly accessible. Log access to detect abuse.
Further Reading and vLLM Resources
- vLLM GitHub repository with documentation and examples.
- vLLM Paper: Efficient Memory Management for LLM Serving with PagedAttention for the underlying research.
- HuggingFace Model Hub to find compatible models.
- Local AI software overview on BotServ.de.
FAQ: vLLM - Common Questions
What is vLLM?
vLLM is an inference engine for large language models on GPUs. It optimizes throughput and memory efficiency through PagedAttention and continuous batching.
Is vLLM free?
Yes, vLLM is open source under the Apache 2.0 license and free. Costs only arise from hardware and electricity.
Can I use vLLM on CPU?
No, vLLM requires a GPU with CUDA or ROCm. For CPU inference, use llama.cpp.
Do I need multiple GPUs for vLLM?
No, one GPU is enough for models that fit in its VRAM. You only need multiple GPUs for large models or higher throughput.
What’s the difference between vLLM and Ollama?
vLLM is optimized for production and high throughput, while Ollama focuses on easy local use. For multi-user scenarios, vLLM is the better choice.
Does vLLM support the OpenAI API?
Yes, vLLM ships with a server that natively implements OpenAI endpoints for chat completions and completions.
Which models work with vLLM?
Any model in HuggingFace format that vLLM supports. This includes Llama, Mistral, Qwen, Phi, and many others. Quantized AWQ and GPTQ variants are also supported.
How much VRAM do I need for vLLM?
At least as much as the model weights require, plus space for the KV-cache. An 8B model in FP16 needs roughly 16 GB, around 6 GB with AWQ. Always reserve headroom for the KV-cache.
Can I use vLLM on Windows?
Only via WSL2 with Ubuntu. Native Windows support is not available.
What is PagedAttention?
PagedAttention is vLLM’s memory management strategy for the KV-cache. It divides the cache into small blocks, similar to virtual memory in operating systems, and drastically reduces waste.
Does vLLM work with AMD GPUs?
Yes, vLLM supports AMD GPUs via ROCm. Installation requires adding the ROCm index to pip.
How do I start vLLM automatically after reboot?
Set up a systemd service that runs the vLLM server on system startup. You’ll find an example in the team setup section above.
References and Further Reading
- Kwon, W. et al. (2023): Efficient Memory Management for Large Language Model Serving with PagedAttention. arXiv:2309.06180.
- vLLM Project: Official documentation, GitHub repository.
- NVIDIA: CUDA Toolkit Documentation.
- HuggingFace: Model Hub and Transformers documentation.


