Agent Security: Securing AI Agents
What this article covers
- The risks that AI agents pose and how to minimize them
- How sandboxes, tool permissions, and guardrails work together
- When human-in-the-loop approvals make sense and how to implement them
- How to make all agent actions auditable through audit logs
- A complete practical example showing you how to secure an agent step by step
Introduction: Understanding agent security
AI agents are powerful. They can read files, send API requests, execute code, and communicate with external systems. This power comes with risk. An agent that acts autonomously can take unintended actions, send data to the wrong place, or run in endless loops that drain your budget.
Agent security means systematically reducing these risks. You give the agent only the tools it actually needs. You monitor its actions. You build approval processes before critical operations run. This way, you benefit from the agent’s autonomy without losing control.
This article is part of the Secure Operations series and builds on the foundations covered in Agent Systems.
Why do you need agent security?
Imagine you have an AI agent that manages files on your system. Its job: clean up and archive old log files. You give it access to your entire filesystem.
The agent interprets its task broadly. It spots files with the log prefix and deletes them. Unfortunately, a project directory contained a file called logic.js, which the agent mistook for a log file. It’s gone. Or the agent sends an email to all customers because it thinks that’s part of archiving.
These aren’t theoretical scenarios. Agents operate on probability, not human common sense. If you give them unfettered access, mistakes happen. Agent security is the safeguard that prevents small errors from becoming big disasters.
Agent security explained
Agent security encompasses all measures that constrain an AI agent so it can complete its task without causing harm. The core principle: minimal permissions, maximum control.
This includes:
- Sandbox: The agent runs in an isolated environment with no access to your main system.
- Tool permissions: Each tool the agent can use is explicitly allowed and can be restricted.
- Human-in-the-loop: Critical actions require human approval before execution.
- Guardrails: Rules that prevent the agent from taking undesired or dangerous actions.
- Audit logging: Every agent action is logged and traceable.
Who this article is for
This article is for developers who deploy AI agents in production or plan to. You should have a basic understanding of how tool calling works and what defines agent systems. Security background is helpful but not required.
Key terms
| Term | Explanation |
|---|---|
| Sandbox | Isolated runtime environment where the agent has no impact on the host system |
| Guardrails | Security rules that keep the agent within defined boundaries |
| Tool calling | Mechanism by which an agent invokes external tools and functions |
| Permission | Authorization that determines whether an agent can use a particular tool |
| Human-in-the-loop | Principle where a human must confirm critical decisions |
| Rate limiting | Restricting the number of actions per time unit |
| Audit log | Record of all actions an agent has executed |
| Least privilege | Principle of granting an agent only the minimal necessary permissions |
| Veto | Ability to reject a planned agent action |
| Escalation | Forwarding a decision to a higher authority, usually a human |
Risks from AI agents
Unintended actions
Agents don’t always interpret instructions the way you expect. An agent allowed to execute code might start a script that alters your system. An agent with write access to a database might delete records because it classifies that as cleanup.
Data leaks
If an agent has access to sensitive data and also communicates with external APIs, it can send data where it shouldn’t go. An agent that reads customer data and calls a webhook could inadvertently transmit that data to a third-party service.
Infinite loops
Agents can get stuck in loops. The agent calls a tool, gets a result that causes it to call the tool again. Without a stopping condition, this loop runs forever. It consumes compute, API quotas, and money.
Cost explosion
Every tool call, every API request, and every inference step has a cost. An agent operating without limits can rack up significant expenses quickly. With cloud-based models like GPT-4, these costs add up fast.
Security measures
Sandbox
A sandbox isolates the agent from the rest of your system. The agent can only access resources you explicitly provide within the sandbox. Files outside the sandbox are invisible, network access is restricted, system commands only affect the sandbox.
Tool permissions
Not every tool that’s theoretically available should be accessible. Define exactly which tools each agent needs. An agent that only analyzes data doesn’t need an email-sending tool. An agent that generates reports shouldn’t have delete access to the database.
Approval gates
Approval gates are authorization workflows for critical actions. Before the agent executes an action flagged as critical, it pauses and waits for your approval. More on this in the Human-in-the-loop section.
Logging
Every agent action is logged: which tool was called, with what parameters, when, and what the result was. This lets you trace what happened and identify the cause of failures.
Rate limits
Rate limits cap how many times an agent can perform an action within a given time window. For example, a maximum of 10 tool calls per minute or 100 API requests per hour. This prevents infinite loops and cost explosions.
Sandbox environments
Docker
Docker is the most common way to isolate agents. You create a container image with exactly the tools and permissions the agent needs. The container has no access to your host filesystem unless you explicitly mount volumes. You can control network access through Docker network rules. You’ll find the basics in Docker Fundamentals.
Virtual Machines
For even stronger isolation, virtual machines offer a solid option. A VM has its own kernel and is completely isolated from the host system. This requires more overhead than Docker, but provides stronger security, especially when the agent executes code that could potentially be harmful.
Restricted Filesystem
Independent of Docker or VM, you can restrict the filesystem. The agent operates within a specific directory and cannot navigate beyond it. Combined with Ollama and local execution, this setup covers many practical use cases.
Tool Permissions
Which tools require approval
Tools with irreversible or external effects always need approval:
- Deleting or overwriting files
- Sending emails
- API requests to external services
- Modifying or deleting database entries
- Executing code
- Triggering payments
Which tools are safe
Tools that only read or compute are considered safe:
- Reading files
- Database queries (SELECT)
- Search requests
- Calculations and transformations
- Caching information
The classification depends on context. Read access to sensitive customer data isn’t safe if the agent can simultaneously call external APIs. Always consider the combination of all tools, not each tool in isolation.
Human-in-the-Loop
Human-in-the-Loop means a person confirms critical decisions before the agent executes them. The agent plans an action, submits it for approval, and pauses. Only when you confirm does the action run.
This operates at several levels:
- Full approval: Every action requires confirmation. This is secure but slow.
- Categorical approval: Only actions in specific categories, such as email sending or file deletion, need approval.
- Threshold approval: Actions below a certain risk score run automatically; those above it require approval.
You’ll find detailed information in the article Human Approvals.
Audit Logging
An audit log records every action the agent takes. At minimum, each entry captures:
- Timestamp
- Tool called
- Tool call parameters
- Result of the action
- Status: successful, failed, or aborted
With a complete audit log, you can later reconstruct what the agent did, why it did it, and where something went wrong. This matters not only for debugging but also for compliance and accountability.
Store audit logs somewhere the agent itself cannot write to. Otherwise, a faulty agent can cover its tracks.
Example: Securing an agent
Imagine you’re building an agent that summarizes emails and forwards important messages to a Slack channel. Here’s how to secure it step by step.
Step 1: Set up the sandbox
The agent runs in a Docker container. The container image includes only the necessary tools: an email client, Slack API client, and the AI model. No shell access, no extra programs.
Step 2: Define tools
The agent receives exactly three tools:
read_emails: Reads emails from the inboxsummarize_email: Summarizes an emailsend_slack_message: Sends a message to Slack
Nothing else. The agent cannot delete files, execute code, or send emails.
Step 3: Set permissions
read_emails and summarize_email run without approval. send_slack_message requires approval before the message is sent. The agent proposes a message, you confirm it, then it gets sent.
Step 4: Set rate limits
Read a maximum of 20 emails per hour, send a maximum of 10 Slack messages per hour. This prevents the agent from flooding the Slack channel in case of a loop.
Step 5: Enable audit logging
Every tool call is logged to a database the agent only has read access to. You can always see which emails were read and which Slack messages were sent.
Step 6: Define guardrails
The agent cannot summarize emails marked as confidential. It cannot send Slack messages to channels not on a whitelist. These rules are embedded in the system prompt and tool logic.
Common Pitfalls
-
Too many permissions: You give the agent “everything” because it’s easier. It works until it doesn’t. Start with minimal permissions and expand only when necessary.
-
No rate limits: Without them, an agent can execute hundreds of API calls in minutes if caught in a loop. Always set limits, even if you’re confident no loop will occur.
-
Audit logs in the same directory: If the agent can write its own logs and has access to the log directory, it can delete traces after an error. Store logs outside the agent’s reach.
-
Bypassing approvals: When approvals are requested too often, users tend to confirm blindly. This makes Human-in-the-Loop pointless. Use approvals selectively, only for truly critical actions.
-
Guardrails only in the prompt: Guardrails that exist only in the system prompt aren’t real guardrails. The agent can ignore them. Add technical restrictions to the tool logic as well.
-
No tests for failure cases: Test what happens when the agent calls a tool incorrectly, when an API is unreachable, or when the agent enters a loop. Failures are exactly when security matters most.
-
Forgotten cleanup actions: An agent that creates temporary files should remove them too. If that’s not part of its job, files accumulate over time, filling the sandbox and slowing the system.
Hardware, Cost, and Security
Security comes at a cost. Docker containers consume extra memory. VMs need their own resources. Audit logs need storage. Rate limits can slow agents because they must wait.
These costs are small compared to what happens when an unsecured agent causes damage. A deleted database, a data leak, or runaway costs at a cloud provider far exceed the infrastructure investment for secure agent execution.
If you work locally with Ollama, API costs disappear, but the security measures remain the same. Sandboxing, permissions, and logging are independent of where the model runs.
For more on how agents plan actions and reflect on their own steps, see the guide Planning and Reflection.
Further Reading
- Secure Operations - Overview of all articles on secure AI system operation
- Human Approval in Security - Approval gates and veto rights
- Configuring Guardrails - Rules and filters for agents
- Prompt Injection Protection - Defending against attacks
- Sandbox for AI Agents - Docker, VM, and filesystem isolation
- Audit Logging - Track agent actions
- Human Approvals - Detailed guide to Human-in-the-Loop
- Tool Calling - Basics of calling external tools
- Agent Systems - Architecture of agent systems
- Planning and Reflection - How agents plan and reflect
- Docker Basics - Isolation with containers
- Ollama - Run local AI models
FAQ
What is agent security?
Agent security encompasses all measures that constrain an AI agent so it can accomplish its task without causing harm. This includes sandbox environments, tool permissions, guardrails, human-in-the-loop approvals, and audit logging.
Do I need a sandbox for every agent?
Yes. Even a simple agent can perform unintended actions. A sandbox requires minimal overhead and prevents errors from affecting your main system.
What’s the difference between guardrails and tool permissions?
Tool permissions determine whether an agent can use a tool at all. Guardrails determine how that tool can be used. Permissions are a binary switch; guardrails are rules within that switch.
When do I need human-in-the-loop?
Human-in-the-loop makes sense for any action that is irreversible, has external consequences, or incurs costs. Sending emails, deleting files, calling external APIs, and processing payments are typical candidates.
How do I prevent infinite loops?
Combine rate limits with a maximum step count per task. The agent terminates once it reaches a certain number of tool calls, regardless of outcome.
Where should I store audit logs?
Audit logs belong in a location where the agent has no write access. This could be a separate database, an external logging system, or a read-only directory.
What does agent security cost?
The infrastructure for security, such as Docker containers and logging, consumes compute resources and storage. Compared to the cost of an unsecured agent causing damage, these expenses are negligible.
Can I secure an agent without Docker?
Yes. You can restrict the filesystem, define tool permissions, and enable audit logging without Docker. Docker provides additional isolation, but it’s not the only approach.
What is least privilege?
Least privilege is the principle of granting an agent only the minimum permissions it needs. If the agent only needs to read, it gets no write access. If it needs one file, it doesn’t get access to the entire directory.
How do I test agent security?
Simulate failure scenarios. Have the agent call a tool incorrectly, disconnect from an API, trigger a loop. Verify that your security measures activate and the agent behaves in a controlled manner.
Are local models safer than cloud models?
Local models like Ollama reduce the risk of data leaks because no data is sent to a cloud provider. Agent security, meaning control over the agent’s actions, is independent of this distinction and necessary in both cases.
Sources
- OWASP: Top 10 for Large Language Model Applications
- NIST: AI Risk Management Framework
- Docker: Container Security Documentation
- Anthropic: Responsible Scaling Policies


