Running AI Agents with vLLM
What this article covers
- How to run AI agents with vLLM.
- Why vLLM is better for production and many parallel requests.
- Setup, model selection, and OpenAI-compatible API.
- Real-world examples for high-performance agents.
- Best practices for batch processing and scaling.
Introduction: vLLM for agents explained
vLLM is a high-performance inference server for LLMs. It uses PagedAttention for efficient memory management and supports batch processing to handle many parallel requests. For agents, this means faster responses, more concurrent agents, and better resource utilization.
This article is intended for advanced users who want to run agents with vLLM. You can find foundational information in vLLM and Running AI agents locally.
Why do I need vLLM for agents?
Imagine you have 10 agents working simultaneously. Ollama processes requests sequentially, which is slow. vLLM batches requests: all 10 agents get answers at the same time. For production environments with many agents, vLLM is the better choice.
vLLM for agents in a nutshell
vLLM = high-performance LLM server with PagedAttention and batch processing. OpenAI-compatible API for agent frameworks. For many parallel agents and production use.
The core idea: faster inference, more concurrent agents, better GPU utilization.
Who is this article for?
- Production teams running many agents in parallel.
- Performance-focused developers who want to maximize GPU usage.
- Scalers deploying agents for many users.
- Advanced users leveraging vLLM for agent applications.
Key concepts
- vLLM - High-performance inference. Useful for: production deployments.
- PagedAttention - Memory optimization. Useful for: efficiency.
- Batch processing - Multiple requests handled simultaneously. Useful for: parallel agents.
- Ollama - Alternative option. Useful for: simple setup.
- OpenAI API - Standard interface. Useful for: frameworks.
Setup
1. Install vLLM
pip install vllm
2. Start the server
# Single model
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-8B-Instruct \
--port 8000
# With GPU
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-8B-Instruct \
--tensor-parallel-size 2 \
--port 8000
3. Connect with an agent framework
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed"
)
# Agent loop (same as LM Studio)
response = client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Task"}],
tools=tools,
tool_choice="auto"
)
Real-world example: Parallel agents
import asyncio
from openai import AsyncOpenAI
client = AsyncOpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed"
)
async def run_agent(agent_id, task):
"""Single agent"""
response = await client.chat.completions.create(
model="meta-llama/Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": task}],
tools=tools,
tool_choice="auto"
)
return response.choices[0].message
async def run_parallel_agents():
"""10 agents in parallel"""
tasks = [f"Task {i}" for i in range(10)]
agents = [run_agent(i, task) for i, task in enumerate(tasks)]
results = await asyncio.gather(*agents)
return results
# All agents run in parallel
results = asyncio.run(run_parallel_agents())
vLLM vs. Ollama
| Aspect | vLLM | Ollama |
|---|---|---|
| Performance | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ |
| Batch Processing | ⭐⭐⭐⭐⭐ | ⭐⭐ |
| Setup | ⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
| GUI | ❌ | ❌ (CLI) |
| Multi-model | ⭐⭐⭐ | ⭐⭐⭐⭐ |
| Production | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ |
| Best for | Many parallel requests | Simple setup |
Security considerations
- API not authenticated: The OpenAI API is unprotected. Use a firewall to restrict external access.
- GPU memory: vLLM uses significant GPU memory. Monitoring is essential.
- Model security: Models can execute prompt injection attacks. See Prompt injection.
Common pitfalls
- More complex setup: vLLM is more technical than Ollama. For simple setups, use Ollama.
- GPU memory: vLLM requires substantial VRAM for batch processing. Size correctly.
- Model format: vLLM needs HuggingFace format, not GGUF. Different models than Ollama.
- No GUI: vLLM is CLI and API only. For a GUI, use LM Studio.
- Overkill for small deployments: For 1-2 agents, Ollama is simpler.
Further reading
- vLLM - vLLM in detail.
- Ollama - Alternative option.
- Running AI agents locally - Overview.
- Multiple agents - Parallel agents.
- Hardware requirements - Hardware.
Key takeaways
- vLLM = high-performance inference with PagedAttention and batch processing.
- Better than Ollama for many parallel agents and production use.
- OpenAI-compatible API for agent frameworks.
- More complex setup, but better performance.
- For 1-2 agents, Ollama is simpler.
FAQ
What is vLLM?
vLLM or Ollama?
How much faster is vLLM?
Which models can I use?
How much VRAM do I need?
Does vLLM support tool calling?
Is my data private?
How do I scale vLLM?
Sources and further reading
- vLLM - GitHub.
- vLLM Docs - Documentation.
- PagedAttention - Paper.


