Skip to content
BotServBotServ
Coding AgentOpenHandsAiderOllamaAI CodingAutomationSelf-HostingPractical ProjectContinuous Operation

Local Coding Agent: Automated Refactoring Overnight

Run a local AI coding agent (Ollama + OpenHands/Aider) nightly on your repo for refactoring, tests, and docs.

S

schutzgeist

4 min read
Local Coding Agent: Automated Refactoring Overnight

Project: Coding Agent on Continuous Duty, Local AI Agent Refactors Your Repo Overnight

What This Project Does

A coding agent runs on your server at night: it reads your repo, identifies improvements, writes refactorings or missing tests, opens pull requests, and you review them over coffee in the morning. Completely local, your code never leaves the company.

The stack:

Git Repo (Gitea/GitLab self-hosted)
    └── Agent (OpenHands or Aider + Ollama)
            ├── Coding Model: qwen2.5-coder:32b or similar
            ├── Sandbox: Docker container per run
            └── Output: branch + PR + report to Mattermost

Prerequisites: server with 64+ GB RAM (or MS-S1 Max with 128 GB for larger coding models), Docker, self-hosted Git server. About 3-4 hours setup.

Why This Architecture

  • OpenHands/Aider instead of GitHub Copilot: no cloud dependency, no subscription fees, code stays in-house, the only viable path for IP-sensitive organizations.
  • Continuous operation: the agent runs as a Cron/systemd task, tackling a different job each night (refactoring Monday, tests Tuesday, docs Wednesday).
  • Sandbox enforcement: agent code executes in a disposable container, never directly on the host. See Sandboxing.
  • PR-based workflow: the agent never pushes to main; it creates branches and PRs for you to review. Human-in-the-loop by design.

Step 1: Local Coding Model

ollama pull qwen2.5-coder:32b    # best local coding model ~20GB
# or deepseek-coder-v2 for larger context window
ollama pull deepseek-coder-v2

Model selection details in Coding Models Comparison; for continuous operation, you need strong tool-calling ability and a large context window.

Step 2: Git Server and Repo Setup

Self-host Gitea (Docker) or use an existing GitLab instance. Give the agent its own bot user with push rights only to agent/* branches:

Repo Rules:
- Agent may push to: agent/**
- Agent may NOT push to: main, develop
- PRs require human review

Step 3: The Nightly Agent Run

# docker-compose.yml, Agent service
services:
  coding-agent:
    image: docker.all-hands.dev/all-hands-ai/openhands:latest
    environment:
      - LLM_MODEL=ollama/qwen2.5-coder:32b
      - LLM_BASE_URL=http://host.docker.internal:11434
      - SANDBOX_RUNTIME_CONTAINER_IMAGE=nikolaik/python-nodejs:python3.12-nodejs22
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock
      - ./workspace:/opt/workspace_base
    # No persistent port needed, runs headless via Cron
#!/bin/bash
# nightly-agent.sh, Cron: every day at 2 AM
TASK=$(cat /opt/agent-tasks/$(date +%A).txt)   # Monday: "refactor", Tuesday: "tests"...

cd /opt/workspace_base/myrepo
git checkout main && git pull

docker run --rm coding-agent \
    --task "$TASK" \
    --headless

git checkout -b "agent/$(date +%Y%m%d)-nightly"
git push origin "agent/$(date +%Y%m%d)-nightly"
# Gitea API: create PR + post report to Mattermost

Task rotation (agent-tasks/):

  • Monday.txt: “Find and refactor duplicated code in src/”
  • Tuesday.txt: “Write missing unit tests for utils/”
  • Wednesday.txt: “Update outdated docstrings”
  • Thursday.txt: “Find potential bugs, write findings report”

Step 4: The Morning Report

After each run, the agent posts to Mattermost (Mattermost Bots):

[BOT] Nightly Report, Repo: backend
Branch: agent/20260922-nightly | PR: #47
Changed: 3 files | Tests: 12 new, all passing
Summary: consolidated duplicated validation in validators/
Review needed: yes, 1 heuristic change in auth.py

You click the PR, review it, merge it, or let the branch expire. Cost of failure: zero.

Step 5: Security, What the Agent Cannot Do

  • No direct push to main: enforce at the Git server level, not just by convention.
  • No secrets access: block .env, secrets/ via .agentignore; see Secrets.
  • No internet access: run the sandbox container with --network=none (except for the Ollama host connection).
  • Budget limits: set maximum runtime and token budget per night, otherwise it may loop endlessly.
  • Audit log: record every agent action in a log; see Audit Logging.

Extensions

  • Multi-repo: one agent, multiple repos, rotating across the week.
  • Issue triage: agent reads open issues, auto-creates fix branches.
  • Hermes variant: Hermes Agent as orchestrator, delegates coding tasks to subagents and reports via Telegram.
  • Quality gate: agent only merges after passing pytest + ruff in a sandbox run.
  • Cluster: two MS-S1 Max servers each running an agent in parallel for repos A and B simultaneously.

What You’ll Learn

  • Headless agent operation (no chat interface, pure task execution)
  • Git workflows for non-human contributors
  • Sandbox and permission design patterns
  • Where local coding models excel (refactoring, tests, documentation) and where they still fall short (complex architectural decisions)

Further Reading

Key Takeaways:

  • Local coding agent on continuous duty: Ollama + OpenHands/Aider + dedicated Git user + sandbox.
  • PR-based workflow: agent never pushes to main, humans review in the morning.
  • Nightly task rotation: refactoring, tests, documentation, bug hunting.
  • Security: no secrets access, no internet, token budgets, audit logs.
  • Local + continuous operation: zero API costs, code stays in-house, essential for IP-sensitive organizations.

FAQ

How good are local coding agents really?

Good for mechanical tasks: refactoring, writing tests, docstrings, finding duplicates. Weak on architectural decisions and subtle bugs, which is why you use PR-based review instead of auto-merge.

How much RAM does the agent need?

For qwen2.5-coder:32b, expect ~20-24 GB model plus overhead, so 64 GB is comfortable. On 128 GB (MS-S1 Max), you can run the 70B+ coding models with large context windows, which matters for bigger codebases.

Can the agent break the repo?

Not the main repo; it works on separate branches, main stays untouched. Worst case is a bad PR that you close. Sandbox isolation and branch-level permissions are your safeguards.

Why not Copilot/Codeium?

Copilot is interactive assistance in the editor; this agent works autonomously on your entire repo overnight. Plus: no code goes to Microsoft or OpenAI. For IP-sensitive codebases, it’s often the only viable option.

OpenHands or Aider?

OpenHands for autonomous multi-step tasks (agent capabilities, browser control, terminal access). Aider for more targeted edit workflows and is lighter-weight. For continuous operation, OpenHands is the more complete platform; Aider is faster for individual edits.

Sources and Further Reading

Back to Blog
Share:

Related Posts