This manual sets up an agent that fixes production errors on its own. A new error in Bugsink starts Claude Code on the production host. The agent reads the error, the trace, the logs and the database. It writes a fix with a regression test, pushes it, watches the deploy, and resolves the issue. Nobody reviews the change before it ships.
This guide is for a prototype, where the owner accepts that risk. The bounds are the CI gate, hard limits, and the skill's rules, not the model's judgement. Read Observability for AI-native apps first. The agent is only as good as the stacks and traces it reads.
web ──exception──▶ Bugsink ──webhook──▶ agent container (same host) ──claude -p──▶ fix
▲ │ ssh host, prod.sh, psql │
└───── resolve ──────┘ │ git push
production ◀── deploy workflow (full gate) ◀─────┘
| Part | What it is |
|---|---|
| Trigger | Bugsink's custom webhook alert: new, regressed, or unmuted issue |
| Agent | A compose service: a small receiver (queue, limits) plus the Claude Code CLI, logged in to a Max subscription |
| Access | SSH to its own host as the operator user, a write deploy key, the internet |
| Test database | A sidecar Postgres on tmpfs, only for tests |
| Instructions | A repository skill: investigate, decide, fix, verify, push, watch the deploy, resolve |
| Gate | CI runs the full check before deploy; an optional guarded branch route lands small fixes |
A common first attempt is a Claude Code routine in Anthropic's cloud, fed a scrubbed brief by a relay that polled Bugsink. It is safe and nearly blind. It sees a stack and nothing else: no trace, no logs, no database. It can fix only what the stack alone explained.
The cloud cannot give it more access. Cloud sessions reach the internet through an HTTP(S) proxy, so SSH and Postgres do not work. Self-hosted cloud environments exist only on Team and Enterprise plans. Allowlisting Anthropic's egress range (160.79.104.0/21) opens the port to every Claude user, so it filters noise but does not authenticate anyone.
A container on the production host needs none of that. No port opens and no key leaves the machine. On a Max plan the CLI runs headless on the subscription.
| Cloud routine | Agent on the host | |
|---|---|---|
| Sees | A scrubbed stack | Everything the operator sees |
| Network to prod | None (HTTP proxy only) | Local: SSH to the host, internal Docker networks |
| Data sent to the model provider | Only an allowlisted brief | Whatever it reads becomes model context, the same as an operator's own Claude session |
| Cost of mistakes | A bad commit | A bad commit, or a bad action on production |
| Setup | Routine plus environment plus relay | One compose service plus a login |
Choose the host only when you accept the second column. For a product with real users and regulated data, keep the cloud routine and its allowlisted brief. Observability, practice 10, covers that variant.
The agent shares the host with production. Before building anything, run a throwaway container with the limits you plan to give it. Inside it, run a clone, an install, one test file, a typecheck and a lint. Meanwhile, probe the site's latency.
A two-core, 8 GB host with cpus: 1, mem_limit: 3g and cpu_shares: 256 gave numbers like these:
| Step | Result |
|---|---|
| Install | 47 s |
| One test file | 6 s |
tsc |
19 s |
| Peak memory | 1.84 GB |
| Site latency, mean | 69 → 88 ms, no errors |
cpu_shares below the web service's default of 1024 makes the agent yield to users. Forbid next build and full suites on the host. CI runs them.
Run one idempotent script on the host (sudo bash -s < provision-host.sh). It creates /etc/app/agent/, a root-owned directory whose private files only the agent's uid can read:
-
host_keygives SSH to the host as the operator user. Itsauthorized_keysline is limited to Docker bridges, so the key opens nothing from outside the machine:from="172.16.0.0/12",no-agent-forwarding,no-X11-forwarding ssh-ed25519 AAAA… app-agent -
known_hostspins the host's own key forhost.docker.internal. -
github_keyis added as a write deploy key:gh repo deploy-key add key.pub --allow-write. -
hook_secretis the path secret for the webhook, because Bugsink signs nothing.
The container gets extra_hosts: host.docker.internal:host-gateway and an ssh_config that makes ssh prod mean the host. Your operator scripts (prod.sh status|logs|sql|oo|bugsink) then work unchanged inside the container. Traffic from the container to the host's sshd takes the host's INPUT path, so no firewall rule changes.
Revoking the agent takes two steps: delete the authorized_keys line and the deploy key.
FROM node:24-bookworm-slim@sha256:… # same base as web: layers already on the host
RUN apt-get install -y git openssh-client python3 make g++ jq curl ripgrep procps
ENV COREPACK_HOME=/usr/local/share/corepack
RUN corepack enable && corepack prepare pnpm@11 --activate
ENV CLAUDE_CONFIG_DIR=/work/claude DISABLE_AUTOUPDATER=1
RUN npm install -g @anthropic-ai/claude-code@<pinned>
RUN curl -fsS https://api.github.com/meta | jq -r '.ssh_keys[] | "github.com " + .' > /etc/ssh/ssh_known_hosts
RUN useradd --uid 10002 -m agent && mkdir /work && chown agent /work
USER agent
CMD ["node", "main.ts"]- Pin the CLI and turn off auto-update. An agent that deploys without review should not change under itself.
- Take GitHub's host keys from its API over TLS, not from
ssh-keyscan, which trusts on first use. - Put
CLAUDE_CONFIG_DIRon the work volume, so the login survives new images.
agent:
build: { context: ./agent }
cpus: "1.0"
cpu_shares: 256
mem_limit: 3g
environment:
AGENT_MODEL: ${AGENT_MODEL:-opus}
CLAUDE_CODE_EFFORT_LEVEL: ${AGENT_EFFORT:-medium}
TEST_PG_URL: postgres://test:test@agent-postgres:5432/postgres
expose: ["8080"]
extra_hosts: ["host.docker.internal:host-gateway"]
networks: { agent: {}, egress: {} }
volumes: ["agentwork:/work", "/etc/app/agent:/run/agent:ro"]
agent-postgres:
image: postgres:18-alpine
command: [postgres, -c, fsync=off, -c, full_page_writes=off, -c, synchronous_commit=off, -c, max_connections=400]
tmpfs: ["/var/lib/postgresql:size=1g"]
networks: { agent: {} }
bugsink:
environment:
ALERTS_WEBHOOK_OUTBOUND_MODE: allowlist_only
ALERTS_WEBHOOK_ALLOW_LIST: agent
ALERTS_WEBHOOK_DENY_NON_GLOBAL: "false"
networks: { agent: {}, back: {} }
networks:
agent: { internal: true }- Pitfall: testcontainers do not work from the agent's network. Tests start Postgres on Docker's default bridge. Docker isolates bridges from each other, so the agent cannot reach the mapped ports and fails with
Failed to connect to Reaper. Compose cannot attach a service to the default bridge either (network-scoped alias is supported only for containers in user defined networks). Give the agent a sidecar cluster, and let the test global setup take an external URL. When the URL is set, the setup drops old test databases, rebuilds the seed, migrates it, and creates the template from it. - Pitfall: Bugsink 2.6's allowlist does not lift the non-global block. An allowlisted internal host is still refused as
non-global IP.allowlist_onlytogether withDENY_NON_GLOBAL=falsekeeps every alert confined toagent.
About 250 lines of dependency-free Node, run with type stripping. It is split into pure decisions (dispatch.ts, unit-tested) and I/O (main.ts).
POST /hook/<secret>takes a timing-safe secret comparison. The body is Bugsink's issue serializer plusalert_reason. Validateidas a UUID, and answer the reasonTESTwithout running anything.- Queue, one run at a time, persisted to
/work/state.json. A duplicate is an issue already queued or running. A naive version misses the running case and fires a second run the moment the first ends. - Limits: runs per rolling 24 hours (they share the owner's subscription), two attempts per issue, and a wall-time cap enforced with
spawn(..., { signal: AbortSignal.timeout(ms) }). A run with no outcome at startup is markedinterrupted. - Before each run:
git fetch,git checkout -f -B autofix origin/main,git clean -fd(not-x, sonode_modulessurvives), thenpnpm install --frozen-lockfile. Use asyncspawn, so the webhook keeps answering during the install. - Run:
claude -p <prompt> --model opus --output-format stream-json --verbose --dangerously-skip-permissions, with the stream written to/work/runs/<time>-<issue>.jsonl. That file is the whole audit trail. - Prompt: the instruction comes first, then the issue title inside
<issue_title>tags, stated to be data. Any signed-in user can choose an exception message.
The receiver only says which issue to work on. A skill in the repository says how:
- Investigate.
bugsink event <id>gives the stack, the request and the trace id. Then pull the trace's spans and logs from OpenObserve, the error history, and the recent deploys. Read the git history of the files involved. - Decline when the cause is upstream or transient, when the fix is a product, legal or money decision, when the root cause cannot be stated in two sentences with evidence, or when the only fix hides the error. Declining is a good outcome.
- Fix. Write the regression test first and make the smallest root-cause change.
- Verify only the changed files: lint, types, their tests with one worker.
- Push. Small fixes in
src/andtests/go to a guarded branch that CI checks and fast-forwards. Anything else goes tomain, where the deploy gate stands before production. - Watch the deploy until the new commit is live, then resolve the issue.
- Report the root cause, the change, the checks run, and the push target.
Rules that have no exception:
- Users' data never goes into commits, fixtures or reports. Tests use synthetic data.
- Never print secrets. Read env keys, never values.
- Dump before any database write. Never delete rows without a counted
where. - Never make paid model calls.
- Never restart a service to hide an error.
Pitfall: Bugsink alerts on an issue only once. It fires when an issue is new, regressed (it came back after being resolved), or unmuted. An open issue that keeps failing stays silent. So the agent must resolve the issue after its fix is live. Then a recurrence alerts again as a regression and uses the second attempt. A declined issue stays open and silent until a person looks at it.
ssh -t host 'sudo docker exec -it -u agent <agent-container> claude' # then /login, subscriptionPitfall: a narrow terminal wraps the OAuth URL, and copying it breaks the link. Copy the full URL on one line, open it, and paste the returned code, including the part after #. The owner types the code; an assistant never does.
The login lands in /work/claude and survives new images. Check it:
claude -p "Reply with the single word ready." --output-format jsonThe result should have "is_error": false.
In Bugsink's shell, once per project:
MessagingServiceConfig.objects.update_or_create(
project=Project.objects.get(name="web"), display_name="autofix agent",
defaults={"kind": "custom", "config": json.dumps({"webhook_url": HOOK_URL})})
config.get_backend().send_test_message() # the agent should log admission "test"The project's alert_on_new_issue, alert_on_regression and alert_on_unmute flags are on by default.
Test each link where it lives, in this order. Each test is cheap and catches its own class of failure.
| Check | Proves |
|---|---|
| Resource trial on the host, with latency probed | The agent will not starve production |
From the container: ssh prod, prod.sh status, sql, oo, bugsink |
The access works without improvised keys |
Clone plus git push --dry-run through the deploy key |
The agent can push, and nothing else can |
| Server tests against the sidecar | The agent can verify database code |
| Receiver: wrong secret, test payload, bad body, new alert, duplicate | Limits and queue behave before a real run spends an attempt |
claude -p before the login |
The pipeline reaches the model; it stops at "Not logged in" |
claude -p after the login, with a read-only task |
It runs on the subscription with the intended model |
| Bugsink's test message | Webhook delivery through the allowlist |
- Pitfall: a green deploy workflow is not a running container. The platform accepts the webhook and restarts containers minutes later. Wait for the new container to be healthy before testing it.
- Pitfall: do not enable the webhook before the login works. Every alert before that ends as a failed run and spends an attempt.
| Need | Command |
|---|---|
| Watch | prod.sh logs agent: alerts, runs started and ended, exit codes |
| Read a run | The JSONL in /work/runs/ |
| See queue and attempts | /work/state.json |
| Pause | docker stop the agent. Alerts that arrive meanwhile are lost, not replayed. |
| Stop for good | Remove the Bugsink messaging service and delete /work/claude |
| Revoke access | Delete the authorized_keys line and the deploy key |
| Budget | Runs share the owner's Max plan. The daily cap is a plan-budget choice, not a technical limit. |
When you replace a cloud routine, clean up behind it:
- disable or delete the routine and its cloud environment;
- delete the relay's Bugsink token;
- remove the routine's env keys from the platform.
- The owner has accepted, in writing, an unattended agent with production access, and that data it reads becomes model context.
- A resource trial ran on the real host, with the site's latency measured.
- Keys were generated on the host. The SSH key is
from=-restricted, the deploy key is write, and the hook secret is in the URL path. - The CLI is pinned with auto-update off, and the login lives on a volume.
- Tests run against a sidecar database, not testcontainers.
- The receiver enforces the daily cap, attempts, one run at a time, duplicate-while-running, and a timeout.
- The skill investigates through traces, declines freely, dumps before writes, and resolves after deploy.
- CI's full gate stands between every push and production.
- Every link was verified in order, and the webhook was enabled last.
- Claude Code cloud environments: security proxy, network levels, API credentials
- Self-hosted environments: Team and Enterprise only
- Routines:
claude/branch pushes - Anthropic IP addresses
- Bugsink source, tag 2.6.1:
alerts/service_backends/custom.pyandwebhook_security.py