I built a supervisor for my coding agents, measured it, and deleted it
I built a supervisor loop for my coding agents, measured it against no supervisor, and removed it: identical success, two percent worse cost.
I built a supervisor loop for my coding agents and wrote down the keep rule first: it stays only if the numbers improve. Against no supervisor, both arms completed three of three tasks, and the supervised arm cost two percent more per task. It did not improve, so it left the default path.
The supervisor was the last thing I built and the first thing I removed. It completed three tasks out of three. Running with no supervisor at all also completed three out of three. The version without it cost $0.1604 per completed task. The version with it cost $0.1635. Two percent worse, on a sample far too small for two percent to mean anything, except for the direction it points in.
So I took it out of the default path. Not because two percent is expensive, but because of a rule I had written down before the trial started: the supervisor stays only if the numbers improve.
That rule, and not the supervisor, is the subject of this post. The component is unremarkable and the measurement is small. What is uncommon is deciding in advance what evidence would make you delete your own work, and then reading that evidence honestly when it turns up.
What the supervisor actually was
It was a loop inside a coordinator I built for running several AI coding agents against one codebase at the same time. The loop is the shape everyone reaches for: plan, work, check, retry, all of it inside caps.
Three settings were required by the design. The stop check has to be a real check and never the agent’s opinion of its own work. There is a maximum number of tries and a maximum spend, and whichever one is hit first stops the loop and escalates to a human. And each try carries forward through a shared variable, so try four reads what try three wrote instead of starting again from the original instruction.
Nothing exotic there. It is a sensible loop, and the reason it exists is easy to state: an agent gets something wrong, a check catches it, and the agent gets another attempt with the failure in hand rather than a blank slate. Every part of that is defensible on paper.
Which is exactly what makes it dangerous. A component nobody can argue with is a component nobody measures.
A component nobody can argue with is a component nobody measures.
The rule came before the build
I wrote the criterion down first: this stays only if the numbers improve. That ordering is the whole trick, and it is uncomfortable precisely because it is cheap.
Write the criterion afterwards and you will be holding a finished component when you judge it. At that point the question quietly changes from “did this help” to “is this reasonable”, and reasonable always wins. You have the code, the design reads well, the argument for it is sound, and there is no number in the room strong enough to overrule all of that. Deciding beforehand is the only way to make the number louder than the sunk effort.
The trial
Three identical synthetic tasks per arm, run before the live trials of the wider mechanism. Each task was verified by a check the agent could not touch: the coordinator holds its own frozen copy of the check from the moment the work is handed out, and grades against that copy. Agents hunt for whatever grades them, not from malice but because editing the check is the shortest path to passing it. Without the frozen copy, the arm with the supervisor would have been free to mark its own homework.
| Success | Cost per completed task | |
|---|---|---|
| No supervisor | 3 of 3 | $0.1604 |
| Supervisor | 3 of 3 | $0.1635 |
Read that precisely, because the temptation is to read more into it. Success was identical. Cost was two percent worse, which at three tasks per arm is noise.
I am not claiming a supervisor makes agents more expensive. The claim is narrower and duller: after building it, I had no evidence it helped, and a small amount of evidence pointing the other way. The rule said improve. It did not improve. Out of the default path it went.
The caveat travels with the result
None of the six runs failed on first attempt. The retry machinery never fired once.
That matters more than the cost figure. In this trial the supervisor was overhead by construction: it planned, it checked, and it had nothing to retry, because nothing needed retrying. The two percent is roughly the price of the extra planning and checking turns, spent on tasks that were going to pass anyway.
So the result refutes one sentence and leaves another standing. It refutes “a supervisor helps on tasks that already succeed”. It does not touch “a supervisor helps on tasks that sometimes fail”. That second trial has not been run, and I am not going to pretend the first one answered it.
Attaching the caveat is not modesty. A finding quoted without its limits gets reused as if it had none, and the person who reuses it will not be there when it fails. If you publish a negative result, publish the exact shape of the hole in it.
It refutes “a supervisor helps on tasks that already succeed”. It does not touch “a supervisor helps on tasks that sometimes fail”.
Removed, not destroyed
One detail keeps this from being melodrama. The supervisor is one loop definition, not special code. It sits in the same loop machinery every other repeated step uses.
So taking it out of the default path costs nothing, and putting it back when there is evidence for it costs nothing either. That is worth engineering for. When removing a component means unpicking it from four other places, the honest verdict becomes expensive to act on, and expensive verdicts get softened into “keep it for now”. Build things so that a negative result is cheap to obey.
Measuring your own component is the rare part
Most people ship the component and then write the post about why it is clever. The incentives run entirely that way. The post about the thing you built is easier to write, it reads as competence, and it has a diagram. The post about the thing you removed reads, on the surface, like admitting you wasted a week.
It is not the same thing. The useful output of that trial was never the supervisor. It was that the coordinator now records cost per completed task including retries, which means the next component I add has a number to beat before anyone has an opinion about it. Once you can measure one component honestly, the rest of the system stops being a matter of taste.
There is a version of engineering judgement that is only ever about what to build. The harder half is deciding what to keep, and you cannot decide that with an argument. You need a measurement, taken against the version of the system that does not have your component in it at all.
That comparison is the part people skip. Running the arm without your own work in it feels like a formality right up until it wins.
The paper
This trial, the mechanism it sits inside, and the three things that mechanism could not see are written up in a paper: 10.5281/zenodo.22670723. It is licensed CC BY 4.0, so use it, quote it, and argue with it freely.
The next post in this series goes to the thing the coordinator was built for in the first place: moving the control point for write collisions between parallel agents from after the write to before the process ever starts, and what that caught when it ran on real repositories.
Written by Sagar Thakkar, AI systems architect specialising in large-scale data processing, cost-optimised cloud-native systems, and reliable production infrastructure. More at sagarthakkar.com.