Skip to content
BotServBotServ
Agent SecuritySandboxGuardrailsHuman-in-the-LoopTool-CallingSecurityAI Agent

AI Agent Security: Safeguarding Intelligent Agents

Secure AI agents with sandboxes, tool permissions, human oversight, and guardrails. Practical guide for safe AI agent deployment.

S

schutzgeist

11 min read
AI Agent Security: Safeguarding Intelligent Agents

Agent Security: Securing AI Agents

What this article covers

  • The risks that AI agents pose and how to minimize them
  • How sandboxes, tool permissions, and guardrails work together
  • When human-in-the-loop approvals make sense and how to implement them
  • How to make all agent actions auditable through audit logs
  • A complete practical example showing you how to secure an agent step by step

Introduction: Understanding agent security

AI agents are powerful. They can read files, send API requests, execute code, and communicate with external systems. This power comes with risk. An agent that acts autonomously can take unintended actions, send data to the wrong place, or run in endless loops that drain your budget.

Agent security means systematically reducing these risks. You give the agent only the tools it actually needs. You monitor its actions. You build approval processes before critical operations run. This way, you benefit from the agent’s autonomy without losing control.

This article is part of the Secure Operations series and builds on the foundations covered in Agent Systems.

Why do you need agent security?

Imagine you have an AI agent that manages files on your system. Its job: clean up and archive old log files. You give it access to your entire filesystem.

The agent interprets its task broadly. It spots files with the log prefix and deletes them. Unfortunately, a project directory contained a file called logic.js, which the agent mistook for a log file. It’s gone. Or the agent sends an email to all customers because it thinks that’s part of archiving.

These aren’t theoretical scenarios. Agents operate on probability, not human common sense. If you give them unfettered access, mistakes happen. Agent security is the safeguard that prevents small errors from becoming big disasters.

Agent security explained

Agent security encompasses all measures that constrain an AI agent so it can complete its task without causing harm. The core principle: minimal permissions, maximum control.

This includes:

  • Sandbox: The agent runs in an isolated environment with no access to your main system.
  • Tool permissions: Each tool the agent can use is explicitly allowed and can be restricted.
  • Human-in-the-loop: Critical actions require human approval before execution.
  • Guardrails: Rules that prevent the agent from taking undesired or dangerous actions.
  • Audit logging: Every agent action is logged and traceable.

Who this article is for

This article is for developers who deploy AI agents in production or plan to. You should have a basic understanding of how tool calling works and what defines agent systems. Security background is helpful but not required.

Key terms

TermExplanation
SandboxIsolated runtime environment where the agent has no impact on the host system
GuardrailsSecurity rules that keep the agent within defined boundaries
Tool callingMechanism by which an agent invokes external tools and functions
PermissionAuthorization that determines whether an agent can use a particular tool
Human-in-the-loopPrinciple where a human must confirm critical decisions
Rate limitingRestricting the number of actions per time unit
Audit logRecord of all actions an agent has executed
Least privilegePrinciple of granting an agent only the minimal necessary permissions
VetoAbility to reject a planned agent action
EscalationForwarding a decision to a higher authority, usually a human

Risks from AI agents

Unintended actions

Agents don’t always interpret instructions the way you expect. An agent allowed to execute code might start a script that alters your system. An agent with write access to a database might delete records because it classifies that as cleanup.

Data leaks

If an agent has access to sensitive data and also communicates with external APIs, it can send data where it shouldn’t go. An agent that reads customer data and calls a webhook could inadvertently transmit that data to a third-party service.

Infinite loops

Agents can get stuck in loops. The agent calls a tool, gets a result that causes it to call the tool again. Without a stopping condition, this loop runs forever. It consumes compute, API quotas, and money.

Cost explosion

Every tool call, every API request, and every inference step has a cost. An agent operating without limits can rack up significant expenses quickly. With cloud-based models like GPT-4, these costs add up fast.

Security measures

Sandbox

A sandbox isolates the agent from the rest of your system. The agent can only access resources you explicitly provide within the sandbox. Files outside the sandbox are invisible, network access is restricted, system commands only affect the sandbox.

Tool permissions

Not every tool that’s theoretically available should be accessible. Define exactly which tools each agent needs. An agent that only analyzes data doesn’t need an email-sending tool. An agent that generates reports shouldn’t have delete access to the database.

Approval gates

Approval gates are authorization workflows for critical actions. Before the agent executes an action flagged as critical, it pauses and waits for your approval. More on this in the Human-in-the-loop section.

Logging

Every agent action is logged: which tool was called, with what parameters, when, and what the result was. This lets you trace what happened and identify the cause of failures.

Rate limits

Rate limits cap how many times an agent can perform an action within a given time window. For example, a maximum of 10 tool calls per minute or 100 API requests per hour. This prevents infinite loops and cost explosions.

Sandbox environments

Docker

Docker is the most common way to isolate agents. You create a container image with exactly the tools and permissions the agent needs. The container has no access to your host filesystem unless you explicitly mount volumes. You can control network access through Docker network rules. You’ll find the basics in Docker Fundamentals.

Virtual Machines

For even stronger isolation, virtual machines offer a solid option. A VM has its own kernel and is completely isolated from the host system. This requires more overhead than Docker, but provides stronger security, especially when the agent executes code that could potentially be harmful.

Restricted Filesystem

Independent of Docker or VM, you can restrict the filesystem. The agent operates within a specific directory and cannot navigate beyond it. Combined with Ollama and local execution, this setup covers many practical use cases.

Tool Permissions

Which tools require approval

Tools with irreversible or external effects always need approval:

  • Deleting or overwriting files
  • Sending emails
  • API requests to external services
  • Modifying or deleting database entries
  • Executing code
  • Triggering payments

Which tools are safe

Tools that only read or compute are considered safe:

  • Reading files
  • Database queries (SELECT)
  • Search requests
  • Calculations and transformations
  • Caching information

The classification depends on context. Read access to sensitive customer data isn’t safe if the agent can simultaneously call external APIs. Always consider the combination of all tools, not each tool in isolation.

Human-in-the-Loop

Human-in-the-Loop means a person confirms critical decisions before the agent executes them. The agent plans an action, submits it for approval, and pauses. Only when you confirm does the action run.

This operates at several levels:

  • Full approval: Every action requires confirmation. This is secure but slow.
  • Categorical approval: Only actions in specific categories, such as email sending or file deletion, need approval.
  • Threshold approval: Actions below a certain risk score run automatically; those above it require approval.

You’ll find detailed information in the article Human Approvals.

Audit Logging

An audit log records every action the agent takes. At minimum, each entry captures:

  • Timestamp
  • Tool called
  • Tool call parameters
  • Result of the action
  • Status: successful, failed, or aborted

With a complete audit log, you can later reconstruct what the agent did, why it did it, and where something went wrong. This matters not only for debugging but also for compliance and accountability.

Store audit logs somewhere the agent itself cannot write to. Otherwise, a faulty agent can cover its tracks.

Example: Securing an agent

Imagine you’re building an agent that summarizes emails and forwards important messages to a Slack channel. Here’s how to secure it step by step.

Step 1: Set up the sandbox

The agent runs in a Docker container. The container image includes only the necessary tools: an email client, Slack API client, and the AI model. No shell access, no extra programs.

Step 2: Define tools

The agent receives exactly three tools:

  • read_emails: Reads emails from the inbox
  • summarize_email: Summarizes an email
  • send_slack_message: Sends a message to Slack

Nothing else. The agent cannot delete files, execute code, or send emails.

Step 3: Set permissions

read_emails and summarize_email run without approval. send_slack_message requires approval before the message is sent. The agent proposes a message, you confirm it, then it gets sent.

Step 4: Set rate limits

Read a maximum of 20 emails per hour, send a maximum of 10 Slack messages per hour. This prevents the agent from flooding the Slack channel in case of a loop.

Step 5: Enable audit logging

Every tool call is logged to a database the agent only has read access to. You can always see which emails were read and which Slack messages were sent.

Step 6: Define guardrails

The agent cannot summarize emails marked as confidential. It cannot send Slack messages to channels not on a whitelist. These rules are embedded in the system prompt and tool logic.

Common Pitfalls

  1. Too many permissions: You give the agent “everything” because it’s easier. It works until it doesn’t. Start with minimal permissions and expand only when necessary.

  2. No rate limits: Without them, an agent can execute hundreds of API calls in minutes if caught in a loop. Always set limits, even if you’re confident no loop will occur.

  3. Audit logs in the same directory: If the agent can write its own logs and has access to the log directory, it can delete traces after an error. Store logs outside the agent’s reach.

  4. Bypassing approvals: When approvals are requested too often, users tend to confirm blindly. This makes Human-in-the-Loop pointless. Use approvals selectively, only for truly critical actions.

  5. Guardrails only in the prompt: Guardrails that exist only in the system prompt aren’t real guardrails. The agent can ignore them. Add technical restrictions to the tool logic as well.

  6. No tests for failure cases: Test what happens when the agent calls a tool incorrectly, when an API is unreachable, or when the agent enters a loop. Failures are exactly when security matters most.

  7. Forgotten cleanup actions: An agent that creates temporary files should remove them too. If that’s not part of its job, files accumulate over time, filling the sandbox and slowing the system.

Hardware, Cost, and Security

Security comes at a cost. Docker containers consume extra memory. VMs need their own resources. Audit logs need storage. Rate limits can slow agents because they must wait.

These costs are small compared to what happens when an unsecured agent causes damage. A deleted database, a data leak, or runaway costs at a cloud provider far exceed the infrastructure investment for secure agent execution.

If you work locally with Ollama, API costs disappear, but the security measures remain the same. Sandboxing, permissions, and logging are independent of where the model runs.

For more on how agents plan actions and reflect on their own steps, see the guide Planning and Reflection.

Further Reading

FAQ

What is agent security?

Agent security encompasses all measures that constrain an AI agent so it can accomplish its task without causing harm. This includes sandbox environments, tool permissions, guardrails, human-in-the-loop approvals, and audit logging.

Do I need a sandbox for every agent?

Yes. Even a simple agent can perform unintended actions. A sandbox requires minimal overhead and prevents errors from affecting your main system.

What’s the difference between guardrails and tool permissions?

Tool permissions determine whether an agent can use a tool at all. Guardrails determine how that tool can be used. Permissions are a binary switch; guardrails are rules within that switch.

When do I need human-in-the-loop?

Human-in-the-loop makes sense for any action that is irreversible, has external consequences, or incurs costs. Sending emails, deleting files, calling external APIs, and processing payments are typical candidates.

How do I prevent infinite loops?

Combine rate limits with a maximum step count per task. The agent terminates once it reaches a certain number of tool calls, regardless of outcome.

Where should I store audit logs?

Audit logs belong in a location where the agent has no write access. This could be a separate database, an external logging system, or a read-only directory.

What does agent security cost?

The infrastructure for security, such as Docker containers and logging, consumes compute resources and storage. Compared to the cost of an unsecured agent causing damage, these expenses are negligible.

Can I secure an agent without Docker?

Yes. You can restrict the filesystem, define tool permissions, and enable audit logging without Docker. Docker provides additional isolation, but it’s not the only approach.

What is least privilege?

Least privilege is the principle of granting an agent only the minimum permissions it needs. If the agent only needs to read, it gets no write access. If it needs one file, it doesn’t get access to the entire directory.

How do I test agent security?

Simulate failure scenarios. Have the agent call a tool incorrectly, disconnect from an API, trigger a loop. Verify that your security measures activate and the agent behaves in a controlled manner.

Are local models safer than cloud models?

Local models like Ollama reduce the risk of data leaks because no data is sent to a cloud provider. Agent security, meaning control over the agent’s actions, is independent of this distinction and necessary in both cases.

Sources

  • OWASP: Top 10 for Large Language Model Applications
  • NIST: AI Risk Management Framework
  • Docker: Container Security Documentation
  • Anthropic: Responsible Scaling Policies
Back to Blog
Share:

Related Posts