Coding Agents with Local AI
What this article covers
- What coding agents are and how they take on simple development tasks.
- How tool-calling lets agents execute code, read files, and run tests.
- Frameworks like LangChain, CrewAI, and Autogen that support agent workflows.
- Building workflows for refactoring, documentation, and testing.
- Security, costs, and common pitfalls.
Introduction: Coding Agents with Local AI
Coding agents go beyond simple code completion. They can read multiple files, execute commands, interpret results, and plan their own steps. This makes it possible to automate tasks like refactoring, documentation, testing, or small feature implementations. Running them locally keeps sensitive codebases in your own network.
A coding agent needs more than just a language model. It needs tools: file system access, shell, Git, test runners, linters. It needs memory for the current plan and a feedback loop to verify results. This is what we call an agent system or tool-calling.
Why do I need coding agents?
Much of software development consists of repetitive tasks:
- Refactoring old code
- Creating tests
- Writing documentation
- Finding outdated patterns
- Simple bug fixes
- Implementing minor features
An agent can prepare these tasks or execute them entirely. You review and approve the work. This saves time and reduces tedious labor.
Coding agents explained
An agent consists of several components:
- Language model: Plans steps and generates code.
- Tools: Execute actions like reading files or running commands.
- Memory: Remembers the conversation history so far.
- Planning: Breaks tasks into individual steps.
- Reflection: Checks results and corrects itself.
Key terminology:
- Tool-calling: Model invokes defined functions.
- ReAct: Reasoning and Acting, a pattern for agents.
- Observability: Understanding what the agent has done.
- Sandbox: Isolated environment for unsafe operations.
- MCP: Model Context Protocol for tool integration.
Who should use coding agents?
- Developers who want to automate repetitive tasks.
- Teams with large, well-structured codebases.
- Tech leads looking to improve code quality and documentation.
- Anyone wanting to experiment with tool-calling locally.
Key concepts in coding agents
- LangChain: Framework for agents and workflows.
- LangGraph: Extension for complex agent flows.
- CrewAI: Multi-agent framework with roles.
- Autogen: Microsoft framework for agent conversations.
- OpenAI Functions: The model that inspired tool-calling.
- MCP-Server: Standardized tool integration.
Building a local coding agent
1. Define tools
The agent needs clearly defined functions:
read_file(path)write_file(path, content)run_command(command)run_tests()search_files(query)
2. Structure the prompt
You are a coding assistant. Read the file main.py.
Identify duplicate code. Refactor it into a helper function.
Write tests for the new function.
Run the tests.
3. Agent loop
- Model receives task and available tools.
- Model decides which tool to call.
- Tool is executed.
- Result returned to model.
- Repeat until task is complete.
4. Security
- Restrict file access.
- Execute commands in a sandbox.
- No critical operations without approval.
- Review changes before applying them.
Practical example: Refactoring with Python
import ollama
tools = [
{
'type': 'function',
'function': {
'name': 'read_file',
'description': 'Reads a file',
'parameters': {
'type': 'object',
'properties': {
'path': {'type': 'string'}
},
'required': ['path']
}
}
}
]
response = ollama.chat(
model='llama3.1:8b',
messages=[{
'role': 'user',
'content': 'Read main.py and summarize its structure.'
}],
tools=tools
)
Important: Not all local models support tool-calling well. Qwen and some Llama variants work better than others.
Frameworks compared
LangChain + LangGraph
Very flexible. Many examples and integrations. Adds some overhead. Good for complex, multi-step workflows.
CrewAI
Easy to get started. Agents get roles and tasks. Works well for clearly structured team scenarios.
Autogen
Conversation-focused. Agents talk to each other until a result emerges. Interesting for special workflows, but more involved.
Security for coding agents
- Sandbox: Run all commands in an isolated environment.
- Read-only first: Allow only read operations initially.
- Human approval: Write operations must be confirmed.
- No secrets in prompts: Filter keys and passwords.
- Auditing: Log all actions.
Common pitfalls with coding agents
- Poor tool-calling ability: Not every model handles tools well.
- Too many permissions: Agent deletes or overwrites files.
- Hallucinations: Invents files or APIs.
- Infinite loops: Agent repeats itself without progress.
- Context window too small: Large codebases don’t fit.
- No testing: Results aren’t verified.
Further reading and resources
FAQ: Coding agents
Can a coding agent take over my entire project? No. It helps with individual tasks; human review remains necessary.
Which model is best for tool-calling? Qwen, Llama 3.1, and some fine-tuned models.
Are coding agents secure? Only in sandboxes with restricted permissions.
Can I fine-tune my own codebase? Yes, through training or RAG using your own files.
Do I need a framework? For simple agents, Python with Ollama is enough. Frameworks help with complex workflows.
Sources and further reading
- LangChain: https://www.langchain.com/
- LangGraph: https://www.langchain.com/langgraph
- CrewAI: https://www.crewai.com/
- Autogen: https://microsoft.github.io/autogen/
Summary: Coding Agents with Local AI
Coding agents can read files, execute commands, and independently handle small development tasks. They rely on tool-calling, planning, and feedback. Running them locally protects sensitive codebases. Key factors are good tools, sandbox security, human approval, and choosing the right model. Frameworks like LangChain, CrewAI, and Autogen make complex workflows easier but are not strictly necessary.


