Skip to content
BotServBotServ
LM StudioAI AgentLocal AITool-CallingSetup

Running AI Agents with LM Studio

Run AI agents locally with LM Studio. Setup, tool-calling, local models and practical examples.

S

schutzgeist

4 min read
Running AI Agents with LM Studio

Running AI Agents with LM Studio

What this article covers

  • How to run AI agents with LM Studio.
  • Setup, model selection, and tool calling with LM Studio.
  • Connecting LM Studio to agent frameworks.
  • Practical examples for local agents.
  • Best practices for performance and privacy.

Introduction: LM Studio for agents explained

LM Studio is a desktop app for running local LLMs. It provides a straightforward UI, a local API, and OpenAI-compatible endpoints. For agents, this means easy setup, local models, and no cloud dependency.

This article is for users who want to run agents with LM Studio. For background, see LM Studio and Running AI agents locally.

Why use LM Studio for agents?

Imagine wanting to run an agent locally, but Ollama feels too technical. LM Studio gives you a GUI: load a model, start the server, use the API. For agent frameworks, the API is OpenAI-compatible.

LM Studio for agents in a nutshell

LM Studio = GUI for local LLMs plus a local API (OpenAI-compatible). Agent frameworks like LangChain and CrewAI connect via the API. Everything stays local, no cloud.

The core idea: simple setup for local agents.

Who should read this?

  • Beginners who want to run agents without the terminal.
  • Developers wanting to quickly test local agents.
  • Privacy-conscious users who prefer GUI over CLI.
  • Prototypers building agents fast.

Key concepts

  • LM Studio - GUI for local LLMs. Use when: you want simple setup.
  • Ollama - CLI for local LLMs. Use when: you need an alternative.
  • OpenAI-compatible API - Standard API. Use when: integrating with agent frameworks.
  • Tool calling - Invoking tools. Use when: building agents.
  • LangChain - Agent framework. Use when: building complex agents.

Setup

1. Install LM Studio

# Download from lmstudio.ai
# Or: Flatpak, AppImage, etc.

# Start
lm-studio

2. Load a model

  1. Open LM Studio
  2. Search for a model (e.g., llama3.1, qwen2.5)
  3. Start the download
  4. Load the model

3. Start the server

  1. Click the “Server” tab
  2. Port: 1234 (default)
  3. Click “Start Server”
  4. API runs at http://localhost:1234/v1

4. Connect to your agent framework

from openai import OpenAI

# LM Studio API (OpenAI-compatible)
client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="not-needed"  # LM Studio doesn't require an API key
)

# Agent loop
response = client.chat.completions.create(
    model="local-model",
    messages=[
        {"role": "system", "content": "You are an agent with tools."},
        {"role": "user", "content": "What is the current temperature?"}
    ],
    tools=tools,
    tool_choice="auto"
)

Practical example: Agent with LM Studio

import json
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="not-needed"
)

class LMStudioAgent:
    def __init__(self):
        self.tools = {
            "get_weather": self.get_weather,
            "get_time": self.get_time,
            "calculate": self.calculate
        }

    def run(self, task):
        """Agent loop"""
        messages = [{"role": "user", "content": task}]

        for _ in range(10):  # Max iterations
            response = client.chat.completions.create(
                model="local-model",
                messages=messages,
                tools=self.get_tool_definitions(),
                tool_choice="auto"
            )

            message = response.choices[0].message
            messages.append(message)

            # Process tool calls
            if message.tool_calls:
                for tool_call in message.tool_calls:
                    result = self.execute_tool(tool_call)
                    messages.append({
                        "role": "tool",
                        "content": json.dumps(result),
                        "tool_call_id": tool_call.id
                    })
            else:
                return message.content

        return "Max iterations reached"

    def execute_tool(self, tool_call):
        """Execute a tool"""
        name = tool_call.function.name
        args = json.loads(tool_call.function.arguments)
        return self.tools[name](**args)

LM Studio vs. Ollama

AspectLM StudioOllama
UIGUICLI
SetupSimplerMore technical
APIOpenAI-compatibleCustom API
ModelsHuggingFaceOllama library
ServerIntegratedSeparate
Best forDesktop, GUIServer, CLI

Security considerations

  • Fully local: LM Studio doesn’t send data anywhere.
  • Model choice: Verify whether the model includes telemetry.
  • API security: The local API has no authentication. Use a firewall to block external access.
  • Model storage: Models can be large. Check available disk space.

Common pitfalls

  • Model too large: Large models (>30B) demand significant RAM or VRAM.
  • API unreachable: Make sure the server is running and check the port.
  • Tool calling unsupported: Not all models support tool calling. Verify compatibility.
  • Slow performance: LM Studio can be slower than Ollama for agents.
  • Single model only: LM Studio loads one model at a time. Switch models or run multiple instances.

Further reading

Key takeaways:

  • LM Studio = GUI for local LLMs plus OpenAI-compatible API.
  • Simple setup for agents without the terminal.
  • Ideal for desktop users and prototyping.
  • OpenAI-compatible API works with agent frameworks.
  • For production: Ollama or vLLM offer better performance.

FAQ

What is LM Studio?

A desktop application for running local LLMs: a GUI for loading and running models, plus a local OpenAI-compatible API for agent frameworks.

LM Studio or Ollama?

Choose LM Studio for a GUI and simple setup. Choose Ollama for server deployments, CLI work, and production. Both can run local models.

Does LM Studio support tool calling?

Yes, if your model supports tool calling (e.g., llama3.1, qwen2.5). The API is OpenAI-compatible and supports tool_calls.

How do I connect agent frameworks?

Use the OpenAI-compatible API: base_url=“http://localhost:1234/v1”. LangChain, CrewAI, and other frameworks can connect directly.

Which models can I use?

Any HuggingFace model in GGUF format: llama3.1, qwen2.5, mistral, and others. For agents, choose models with tool-calling support.

Is my data private?

Yes, completely. LM Studio processes everything locally. No data is sent to external servers.

Is LM Studio performant?

Good for desktop use and prototyping. For production and high concurrency, Ollama or vLLM perform better.

References and further reading

Back to Blog
Share:

Related Posts