An operating system for a team of one

How I work with several AI coding agents and keep one record of what was decided. The rules, the files, and the one decision that made the rest possible.

When I build my own tools, I am a team of one. For most of 2026 that has meant me plus several AI coding agents.

The second one arrived the day I hit the usage limit on the first subscription, with work still to finish. It stayed because I could not fully rely on one agent’s account of its own work.

Every session started from zero. An agent would reopen a decision I had settled the week before. The record of what had happened lived in chat that nobody, including me, was going to read again.

The fix was one decision and an operating model built on it: a written contract that every agent, in every tool, reads before it does anything. I call it the Agent Operating System, and this is what it changed.

Files remember; chat only discusses

Everything follows from one decision: files over chat.

Task state, session state, decisions, open arguments, hand-off notes: each lives in a file with a known name at a known path. Each coding tool’s instruction file points at that shared record. Chat is where work is discussed. Files are where work is remembered.

Once that is true, a session in any tool starts the same way: it reads the record and checks the repository before it acts. I call that ss. Its mirror, sss, closes a session by writing the record back. The next session, in any tool, picks up from files rather than from memory.

The agent stops at a written line

Once the record lives outside chat, the model draws lines around decisions.

On my side of the line: I own direction, scope, releases, priorities, trade-offs, and the final call when a debate is exhausted. Agents own technical analysis, sequencing, execution planning, and routine implementation decisions inside a plan I have approved.

A third list matters most: things an agent must never decide alone. Scope changes. Deployment and infrastructure changes with operational risk. Security posture. Irreversible data changes. Anything that overrides a decision already made.

When work touches that list, the agent stops and says so. I wrote the boundary down to cover changes with consequences beyond the current task: unreviewed changes to live infrastructure, a scope quietly doubled, a settled decision reversed in passing.

Inside the approved plan the agent acts. Outside it, everything is a question first, including the typo in a file the plan leaves alone. “Obvious” gets the same question. Asking costs one line; a well-meant fix in the wrong place costs days.

One thing at a time

The same line shapes how the agents talk to me. The communication rules came from use, and I wrote them for myself first. One question in front of me, not four. The result in two sentences; the audit trail goes in the log file. Nothing re-asked that I already approved. No urgency language unless something is on fire.

The rule governs questions and decisions, not execution. Agents work in parallel, each on a narrow packet, with an orchestrator holding the wider context.

Plan first, with an exit

The default workflow is strict: a written plan, reviewed and approved, before anything is executed, and the record updated after.

That rigidity broke in practice. Much of this work is small and reversible, and a written plan there is ceremony.

So the model defines a lighter mode, and defines it precisely: on a low-risk, reversible task, when I ask for it, the written plan, the review, and the formal approval collapse into a one-line stated intent and a go-ahead. Inspecting first, verifying after, and updating the record never drop. Naming the exception beat pretending it did not exist.

The model had other holes too. Two other models reviewed the public edition and found five: overlapping edits, runaway cost, instructions hidden in data, no rollback, and rules that went stale. Each now has a rule in the repo.

The lock that could never open

The rules were tested on a real fault. While building Mano, an offline app, one agent implemented the app lock and another audited security. The auditor found that the lock could never succeed on Android: the activity type was wrong.

It did not edit the file, because the first agent had claimed it. It wrote the finding, with evidence, into the discussion file and opened a task. The first agent picked the task up at its next checkpoint, wrote the revert step, and made the change.

Then the change broke the debug build: a guard I had asked for fired at configuration time for every build. The verification failed, which is a rollback trigger. Revert, record, fix properly in a second reviewed commit. After three review rounds the lock prompted correctly on a device, the automated flow test passed on both platforms, and the change was ready for the app.

The system works because the record survives the session, even when the agent does not.