Skip to content
BotServBotServ
vLLMAI AgentHigh-PerformanceBatch ProcessingOpenAI API

Running AI Agents with vLLM

Run AI agents with vLLM. High-performance inference, batch processing, OpenAI-compatible API.

S

schutzgeist

4 min read
Running AI Agents with vLLM

Running AI Agents with vLLM

What this article covers

  • How to run AI agents with vLLM.
  • Why vLLM is better for production and many parallel requests.
  • Setup, model selection, and OpenAI-compatible API.
  • Real-world examples for high-performance agents.
  • Best practices for batch processing and scaling.

Introduction: vLLM for agents explained

vLLM is a high-performance inference server for LLMs. It uses PagedAttention for efficient memory management and supports batch processing to handle many parallel requests. For agents, this means faster responses, more concurrent agents, and better resource utilization.

This article is intended for advanced users who want to run agents with vLLM. You can find foundational information in vLLM and Running AI agents locally.

Why do I need vLLM for agents?

Imagine you have 10 agents working simultaneously. Ollama processes requests sequentially, which is slow. vLLM batches requests: all 10 agents get answers at the same time. For production environments with many agents, vLLM is the better choice.

vLLM for agents in a nutshell

vLLM = high-performance LLM server with PagedAttention and batch processing. OpenAI-compatible API for agent frameworks. For many parallel agents and production use.

The core idea: faster inference, more concurrent agents, better GPU utilization.

Who is this article for?

  • Production teams running many agents in parallel.
  • Performance-focused developers who want to maximize GPU usage.
  • Scalers deploying agents for many users.
  • Advanced users leveraging vLLM for agent applications.

Key concepts

  • vLLM - High-performance inference. Useful for: production deployments.
  • PagedAttention - Memory optimization. Useful for: efficiency.
  • Batch processing - Multiple requests handled simultaneously. Useful for: parallel agents.
  • Ollama - Alternative option. Useful for: simple setup.
  • OpenAI API - Standard interface. Useful for: frameworks.

Setup

1. Install vLLM

pip install vllm

2. Start the server

# Single model
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.1-8B-Instruct \
    --port 8000

# With GPU
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.1-8B-Instruct \
    --tensor-parallel-size 2 \
    --port 8000

3. Connect with an agent framework

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"
)

# Agent loop (same as LM Studio)
response = client.chat.completions.create(
    model="meta-llama/Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "Task"}],
    tools=tools,
    tool_choice="auto"
)

Real-world example: Parallel agents

import asyncio
from openai import AsyncOpenAI

client = AsyncOpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed"
)

async def run_agent(agent_id, task):
    """Single agent"""
    response = await client.chat.completions.create(
        model="meta-llama/Llama-3.1-8B-Instruct",
        messages=[{"role": "user", "content": task}],
        tools=tools,
        tool_choice="auto"
    )
    return response.choices[0].message

async def run_parallel_agents():
    """10 agents in parallel"""
    tasks = [f"Task {i}" for i in range(10)]
    agents = [run_agent(i, task) for i, task in enumerate(tasks)]
    results = await asyncio.gather(*agents)
    return results

# All agents run in parallel
results = asyncio.run(run_parallel_agents())

vLLM vs. Ollama

AspectvLLMOllama
Performance⭐⭐⭐⭐⭐⭐⭐⭐
Batch Processing⭐⭐⭐⭐⭐⭐⭐
Setup⭐⭐⭐⭐⭐⭐⭐⭐
GUI❌❌ (CLI)
Multi-model⭐⭐⭐⭐⭐⭐⭐
Production⭐⭐⭐⭐⭐⭐⭐⭐
Best forMany parallel requestsSimple setup

Security considerations

  • API not authenticated: The OpenAI API is unprotected. Use a firewall to restrict external access.
  • GPU memory: vLLM uses significant GPU memory. Monitoring is essential.
  • Model security: Models can execute prompt injection attacks. See Prompt injection.

Common pitfalls

  • More complex setup: vLLM is more technical than Ollama. For simple setups, use Ollama.
  • GPU memory: vLLM requires substantial VRAM for batch processing. Size correctly.
  • Model format: vLLM needs HuggingFace format, not GGUF. Different models than Ollama.
  • No GUI: vLLM is CLI and API only. For a GUI, use LM Studio.
  • Overkill for small deployments: For 1-2 agents, Ollama is simpler.

Further reading

Key takeaways

  • vLLM = high-performance inference with PagedAttention and batch processing.
  • Better than Ollama for many parallel agents and production use.
  • OpenAI-compatible API for agent frameworks.
  • More complex setup, but better performance.
  • For 1-2 agents, Ollama is simpler.

FAQ

What is vLLM?

A high-performance inference server for LLMs with PagedAttention and batch processing. For many parallel requests and production deployments.

vLLM or Ollama?

vLLM for production and many parallel agents. Ollama for simple setup and few agents. vLLM is faster, Ollama is easier.

How much faster is vLLM?

2-10x faster for parallel requests through batch processing. For single requests, similar to Ollama.

Which models can I use?

All HuggingFace models in HuggingFace format. llama3.1, qwen2.5, mistral, and others. Not GGUF format (that’s for Ollama and LM Studio).

How much VRAM do I need?

More than Ollama for batch processing. llama3.1:8b requires around 10 GB for single requests, 20 GB for batching. For many parallel agents, you need more VRAM.

Does vLLM support tool calling?

Yes, if the model supports tool calling. The OpenAI-compatible API supports tool_calls like OpenAI.

Is my data private?

Yes, completely. vLLM processes everything locally. No data is sent to external servers.

How do I scale vLLM?

Tensor parallelism for multiple GPUs, multiple servers for horizontal scaling. Kubernetes for orchestration. For many agents, vLLM is the best choice.

Sources and further reading

Back to Blog
Share:

Related Posts