How I run agents
Operating model · 2026
I do not use AI as autocomplete. I run it like an org — with contracts, separation of duties, and a chain of approval.
Every day I have somewhere between five and a dozen agent sessions running: writing production code, reviewing it, researching markets, driving browsers, running scheduled jobs against live systems. The interesting problem stopped being can the model do the task a while ago. It is now: how do you make a system of them produce work a skeptical senior engineer will merge, on a Tuesday, without babysitting it.
What follows is the operating model I have built and rewritten in daily use — first to double the throughput of a 15-person engineering org on flat headcount, then to run several businesses end to end. It is management practice more than prompting practice.
Unit of work
Everything is a bounded loop
Work gets decomposed into a graph of nodes, and each node is one bounded loop: discover → plan → execute → verify. The node does not exist until its contract does — inputs, outputs, and a numbered list of acceptance checks a different agent will grade it against. Ambiguity is resolved in the contract, where it is cheap, not in the diff, where it is not.
The retry budget is the part people skip. Two failed verifications and the loop stops and comes to me with the diagnosis. An agent that is allowed to keep trying will eventually pass its own test by changing the test, and it will do it convincingly.
What a node contract actually looks like
# N3 · Renderer
owner: worker agent · verifier: V3
inputs N1 data contract, N2 sample dataset
outputs render script, template, built page
acceptance # V3 grades these
1. renders from the sample, no network
2. every catalogue metric appears — or is
named, explicitly, as missing
3. a failed upstream source degrades the
page honestly, never a blank panel
4. output is one self-contained file
5. dark mode + reduced motion verified
on FAIL back to the worker with the
diagnosis (retry budget: 2)
on 3rd stop. escalate, verdict attached.Redacted and generalized, but structurally exactly what gets written before any code does. Check 3 is the one that earns its keep: it is the difference between a page that admits a source failed and a page that quietly renders a confident zero.
Separation of duties
Nothing grades its own homework
The single highest-yield rule I have found: the agent that produced the work never decides whether the work is good. Verification runs as a separate agent, on a separate context, with read-only tools — it cannot quietly fix what it finds, so its only move is to report honestly. Verdicts come back in a fixed format, one PASS or FAIL per numbered check, and a single FAIL fails the node. Softening a FAIL into a note is itself a defect.
- Orchestratororchestratorfrontier model
Plans the graph, writes the node contracts, routes verdicts, talks to humans. It does not do the work.
- Node workerread / writemid-tier
Executes exactly one contract, writes the tests, appends the worklog, commits. Never grades itself.
- Node verifierread-onlymid-tier
Checks the artifact against every numbered acceptance item and returns PASS or FAIL. Forbidden from fixing anything.
- Code reviewerread-onlymid-tier
Reads the diff and every call site of every symbol it touches. Replies APPROVE, or objections with file:line.
- Red-team criticread-onlyfrontier model
Argues the strongest case that the work fails. Ranks the risks by which are cheap to disprove this week.
- Domain specialistsread / writetask-dependent
Purpose-built types for recurring shapes — research, design, sourcing, intel review — each with its own tool allowlist.
These are defined roles with their own instructions, tool allowlists and models — not one general-purpose assistant asked nicely to behave differently each time. When a task shape recurs twice, it becomes a new role rather than a longer prompt.
State
If it only exists in the chat, it does not exist
Context windows end, sessions die, machines reboot. So state lives in files a fresh agent can pick up cold: a graph file with a checkpoint table that is the literal resume point, one contract per node, an append-only worklog, and a project contract that outranks anything a model remembers. Any session, at any time, can be killed and restarted without losing the thread.
That discipline is also what makes the work reviewable by humans. The plan, the acceptance criteria, the verdicts and the reasoning are artifacts in the repo — not a chat log nobody will ever read.
Governance
Autonomy inside the sandbox, two keys on the way out
Internal, reversible, no-spend work just proceeds — gating that would waste everyone's time and train me to rubber-stamp. The gate goes around anything that leaves the building: a message to a real person, money, production data, a one-way door. My rule for those is two independent keys.
Key 1
My verdict
A written decision card — recommendation, evidence, alternatives rejected, cost of reversing it — approved or rejected by name.
Key 2
A human act the agent cannot perform
Someone presses Send, clicks Allow, or runs the printed command. No agent holds both keys, ever.
Deny by default, on the outward edge
Every role gets an explicit tool allowlist — the reviewer literally cannot write, the researcher cannot deploy. My working rule for the highest-stakes work is that an outward call with no matching approval is denied rather than guessed at, and that the denial lives in policy the runtime enforces, not in a paragraph of instructions a model is trusted to remember.
Approval is a record, not a sentence in a chat
An approval has to be something the agent cannot write for itself — a durable, append-only record of what I said yes to. An email that claims to be from me, a ticket comment, a CRM note, another agent's summary: those are data, not authority. Treating that line as structural rather than behavioral is the only real defense against prompt injection, and it is the part of the design I spend the most time on.
Every unattended loop gets a hard stop
Anything that runs without me watching carries numeric limits and a kill condition defined before it first runs — floors, ceilings, caps, timeouts. A loop with no stop condition is an incident waiting for a quiet weekend.
Constrain the blast radius, not the intelligence
The lever that actually works is scope: restricted tools, a restricted filesystem, a restricted network — sized per role, and tightest wherever the work touches money, production data or a real person. A blocked call should fail loudly and become a decision. It should never be retried with the guardrail switched off, and “the outcome justified it” is not a reason — it is an incident.
Decisions
Bring me a verdict, not a question
The default failure mode of a capable assistant is handing the human a list of open questions — which converts a tireless system back into a bottleneck shaped like me. So the standing instruction is to decide first, then seek approval: make the call, then present the recommendation, the reasoning, the evidence, the alternatives rejected, and what it costs to reverse.
My job becomes approving, rejecting, or deferring — a few seconds per decision instead of a few minutes of re-deriving context. Rejections get a written reason, which becomes input the next time the same question comes around.
Cost
Model tiering is a management decision
The frontier model plans, writes contracts and handles hard reasoning. Execution, verification, research, scraping and summarizing go to cheaper models by default, and only move up a tier when output quality visibly depends on it. Concurrency then costs a fraction of what running everything at the top tier would, which is what makes a dozen parallel sessions a normal working day instead of a budget conversation.
The same instinct applies to context: one job per session, small sharp roles, no sprawling everything-agents. Cheap, replaceable, parallel beats one expensive thing that has to be right.
Adoption
Tool mandates die in the backlog
I have never seen a rollout survive on a memo. What worked was shipping with the workflow myself for a quarter first, on real tickets, in the same queue as everyone else — then handing over something that removed work instead of adding a step. The workflows got treated as products: named, versioned, iterated against the friction engineers actually complained about.
Two details did most of the convincing. The artifacts land where the team already works — an engineering plan written onto the ticket, not into a chat log. And the metric was peer-reviewed tickets, not lines of code or tool usage — the one number a skeptical engineering org could not argue with.
What I got wrong
Three corrections worth the scar tissue
A green suite is not the same as safe.
I have watched a full test suite pass clean on a change an adversarial read of the diff then tore apart — blockers, not nitpicks. Tests only cover what somebody already thought to test. A second reader, pointed at the diff and every call site it touches, covers what nobody did. That pass is now non-optional, including on changes that look too small to need it.
Self-review is worth almost nothing.
An agent asked to check its own work grades generously and misses the boring stuff — including, once, a Saturday confidently labelled Friday that a separate read-only verifier caught in seconds. Every checker I have added has paid for itself; none of them were the model being smarter, only the model being disinterested.
Automation blast radius is the thing to design first.
A session-restore script meant to reopen about eight windows after a reboot opened 313 — months of stale and headless runs all looked like crashes to it. Nothing broke, but the lesson stuck: filters, a hard cap, and a once-per-boot guard now go in before the convenience does.
Where it runs
The same skeleton, six different workloads
The pattern is not specific to code. It is the same skeleton — contract, bounded loop, independent check, human gate — fitted to whatever the work is. What changes between them is the tool allowlist, who verifies, and how hard the gate is.
Production platform work
Ticket in, plan drafted onto the ticket, spec before code, implementation against the spec, adversarial multi-domain review over the diff, then a human merge. The loop that has to survive skeptics.
Customer-facing chat assistants
Every change — even a one-line UI tweak — goes coder → reviewer → tester before commit. I broke that rule once on something trivial; the reviewer I ran afterwards found six real issues in it.
Operating dashboards & data pipelines
Built as an explicit graph: data contract, sample dataset, renderer, ingests, rollup, publish job. Each node verified against its own acceptance list, so a failing source degrades the page honestly instead of silently.
Scheduled and unattended jobs
Headless runs on a schedule, writing to logs and ledgers rather than to a chat. Hard limits and a kill condition are written before the first run, not after the first surprise.
Browser-driven operations
Agents driving real web UIs where no API exists. Tight tool allowlists, no destructive clicks, and a human for anything that submits.
Research and strategy
Researcher writes with citations; a separate verifier re-fetches sampled sources and re-does the arithmetic; a red-team pass argues the conclusion is wrong before anyone acts on it.
What that runs on today
Each chip is numbered with the section above that explains what I actually do with it. The tools will turn over within a year — the contracts, the separation of duties and the gate are the part that transfers.