Agent Output Is Not the Product Metric
When agents make production cheap, the builder’s job shifts upstream: choose fewer valuable problems, write a decision-grade task packet, and verify customer impact. Pull requests per day is a capacity measure; accepted product change is the outcome.
- 01Claire Vo interviews Ryan Carson on the $20k month, 15-agent queue, 40 daily PRs, customer visit, and why output is the wrong goalindependent analysis · Aug 24, 2026↗
- 02Devin documentation on managed parallel sessions, monitoring, compute limits, and reusable work patternsprimary↗
- 03Devin documentation on turning repeated successful tasks into reusable playbooksprimary↗
- Ryan Carson says he spent $20,000 on Devin in one month while operating 15 concurrent agents across engineering, customer success, and investor work.
- His system is deliberately plain: folders, P0 threads, and a handwritten list coordinate the queue instead of a large orchestration dashboard.
- A reusable “LAN PR” workflow reportedly closes the loop on about 40 pull requests a day using review loops, a video walkthrough, and automatic merge rules rather than a standing QA team.
- The interview’s strongest counterweight to the throughput story is that one in-person customer visit changed the product direction—and both operators argue that producing more AI output is the wrong goal.
The headline number is $20,000 of Devin usage in a month, but the operating system underneath it is more revealing. Carson keeps as many as 15 cloud agents moving through a short priority queue, turns repeated work into playbooks, and uses automated review paths to process roughly 40 pull requests a day. That creates extraordinary build capacity for one founder. It does not create automatic product judgement. In the same conversation, Carson describes leaving the computer to meet a customer and changing Untangle’s direction because of what he learned. Put those facts together and the constraint is clear: once implementation supply expands, the scarce work becomes deciding what deserves to enter the queue, exposing assumptions in the task packet, and checking whether shipped work changes a customer outcome.
Teams that celebrate agent count, tokens spent, or pull requests merged can become extremely efficient at compounding the wrong bet. The failure is not theoretical: a weak priority becomes fifteen concurrent implementations; a vague spec creates forty review decisions; and a missing customer signal lets polished output masquerade as progress. The useful metric is therefore not agent throughput. It is the share of changes that clear verification, reach a real user, and produce the evidence the original decision asked for. Agent capacity should buy shorter learning loops—not a larger backlog in motion.
Put a work-in-progress limit in front of the agent queue. For every proposed run, require a five-line task packet: customer or operator problem, evidence that it is worth solving now, smallest acceptable change, verifier, and the decision the result will unlock. Allow only three active bets per builder even if the tooling can run thirty sessions. Route low-risk changes through automated tests and screenshots; route behaviour-changing or irreversible work through a named human review. Review accepted changes weekly using four numbers: reached users, verifier pass rate, reversions, and evidence that changed the next decision. If a run cannot name its downstream decision, do not start it.
Put it to work 30 minutes
An agent work-in-progress ledger with ten classified changes and one decision-grade task packet
Audit the last ten agent-generated changes in your product. Mark each one “accepted,” “reworked,” “reverted,” or “never reached a user.” Then choose one live task and rewrite it as a five-line decision packet with a verifier and a work-in-progress slot. The artifact is useful if another builder can tell why the task matters, what good looks like, and what you will decide when it ships.
Your turn
Use the task above. Record the result and anything you still need to check.
Your draft stays in this browser. No account needed.
For a bounded migration, test-coverage push, or other work with objective acceptance checks, raw throughput can be a useful leading measure. The warning applies when the work changes product behaviour or depends on uncertain customer demand.
The operating numbers and workflow details come from one founder’s account, so they are not a general benchmark. Devin’s product documentation independently confirms the underlying mechanics—parallel managed sessions and reusable playbooks—but not Carson’s reported outcomes.
Did limiting work in progress increase the percentage of agent-generated changes that reached users and produced decision-changing evidence?
Watch: accepted-change rate · human rework minutes · reversion rate · changes reaching users · decisions changed by evidenceKeep Going