How do you know your AI operator is working? The first review.
The first routine is handed over, the operator has been running for two weeks, and the question underneath is still open: is this actually going well? The usual answer is to read everything. That is understandable and it is the wrong instrument.
In short: Reading everything is not governance, it is the same work again. The first review measures a number instead, FTT: the share of runs that pass without rework — in twenty minutes, once a week, against three questions.
How do you know your AI operator is working? The first review.
You have handed over a routine. The operator has been running for two weeks, the output looks reasonable, and a question sits underneath it that most people do not say out loud: how do I actually know this is going well?
The most common answer is to read every single output. That feels responsible, and for the first few days it is right. As a permanent arrangement it does the work a second time, more slowly, and it produces nothing you can use next week.
Reading everything is not governance. It is the same work again, at higher cost and lower insight.
Checking AI output means measuring the process, not the single result
The distinction sounds academic and it is the whole point. Whether one piece came out well is something you can see by reading it. Whether the role reliably produces good pieces is not visible that way. That takes several runs and a number.
The number is FTT (First Time Through), the share of runs that pass the check on the first attempt. Not the share that was usable in the end, because in the end almost everything is usable once enough hands have touched it. The share that gets through without rework.
It is the right number because it measures what you actually care about: whether this work is leaving your head. A role whose output gets reworked every time has taken nothing off you, however good the final version looks.
Twenty minutes, three questions
The first review needs no format and no meeting. It needs a fixed slot, once a week, and three questions in this order.
Do not confuse it with the Weekly Operating Review. That one looks at the company, what moved and what is stuck. This one looks at a single role, and its result eventually becomes one of the items that shows up there.
1. How many runs passed on the first attempt? Count them. A weekly routine gives you six data points in six weeks, which is already enough to tell a trend from a bad day.
2. Where did the others fail, and was it the same place? This is the more useful question. Five different failures are noise. The same failure five times is a signal, and usually a signal about your standard rather than about the operator.
Count a third group while you are at it: the runs nobody checked. A week was tight, a case fell outside the standard, someone waved it through. If those quietly join the passes, you are measuring your own review capacity and calling it quality. Not checked is a result of its own, not a pass.
3. What changes in the contract as a result? If the answer is "nothing", it was not a review, it was a status update. A learning only counts when it changes an artifact, and in a Role Contract that is precisely what the upgrade rules are for.
A high score early is not good news
This surprises people. If your first measurement comes in near a hundred percent, the likeliest explanation is not that the role is running beautifully. It is that your check is too soft.
A standard nothing fails against is not measuring anything. It describes your expectation rather than a requirement, and the number it produces tells you nothing about quality that you did not already believe.
If nothing ever fails, you are not checking. You are confirming.
A low first score is uncomfortable and useful. It shows you exactly where your definition of good has been living in someone's head instead of in a standard, and those are places you can fix independently of any AI.
When the same thing breaks twice
There is a rule for recurring failures and it is older than any of this technology. When the same failure happens a second time, do not add another check. Change the process so it cannot happen.
That is Poka Yoke, and the difference is practical. An extra check costs you time every week and catches the failure only after it has occurred. A change to the contract, the template or the sequence costs you time once and removes the opportunity.
Stacking review steps instead builds you, over a few months, exactly the control loops you wanted an operator to get rid of.
When the role earns the next stage
The number also answers the question that comes next. A role starts at Shadow: it drafts, humans execute. Moving to Copilot belongs tied to evidence rather than to a date. Several consecutive runs where the score holds, and no recurring failure point still open.
The ladder runs both ways. If the score drops, the role goes back a stage, and that is not a setback. It is the proof that the stages hang on measurement rather than on confidence.
Company 0
For us the recurring failure was not on the AI operator's side, it was on ours. The LinkedIn posts that accompany the weekly article are meant to run at least 1,100 characters, a standard set long before this particular routine existed. Even so, drafts kept arriving under that line for several weeks running, and each time it got fixed after the fact.
The same place kept failing. Instead of adding another check, the character count became a fixed condition inside the brief itself, verified before anything is handed over. The mistake has not recurred since, not because anyone is watching more closely, but because the gap it used to slip through is no longer part of the process.
What this means for you
Put the slot in the calendar before you need it. Twenty minutes, weekly, from the very first run. A review you schedule because something caught your eye is not a review, it is a reaction.
And expect an uncomfortable number in the first few weeks. It is the most useful thing this role can give you early on, because it shows you where your standard has so far existed only in your head.
Frequently Asked Questions
What is FTT and what does it stand for?
First Time Through: the share of runs that pass the check on the first attempt, without rework. Not the share that was usable in the end, but the share that got through without a second pass.
How often and how long does the first review take?
Twenty minutes, once a week, on a fixed slot from the first run onward. A review you only schedule when something catches your eye is not a review, it is a reaction.
Is a near-100% pass rate a good sign?
Usually not. It more often means the check is too soft than that the role is running beautifully. A bar nothing fails against is not measuring anything, it is only describing what you already expected.
What do I do when the same mistake happens twice?
Do not add another check. Change the process, the contract, or the template so the mistake cannot happen again. That is Poka Yoke, and it costs less than any additional control loop.
Want to build this in your own company?
Rocket Routine OS is the operating system behind these articles, and it is not open yet. Join the waitlist and you hear first, get a fortnightly honest account of what is working and what is not, and move further up the list with every referral.
Join the waitlistWant to understand the whole system?
The entire architecture of Rocket Routine OS as a PDF — 20 pages, freely available, no form required.
Download whitepaper (PDF)