All posts

The Price of a Bug Is Set by Where You Catch It

The same one line defect costs twenty seconds at the keystroke and a week in production. Adding another gate does not change that number. Moving the check inward does.

Serkan Ozal10 min read

Pick a bug your team shipped last quarter. Not the dramatic one. An ordinary one, the kind that turned out to be a line or two.

Ask a different question about it than the one you asked at the time. Not why it happened, and not who missed it. Ask what it cost. Then ask what it would have cost if something had caught it one step earlier.

For most bugs those two numbers are not close, and the difference between them has nothing to do with the bug.

Here is an ordinary shape. Somebody removes a limit from a query, because the list it fetches is small. It is small in the test fixtures. It is small in staging. It is small in every environment the team owns. It is not small for the largest customer, and no environment the team owns contains the largest customer.

Nothing in that change is wrong. It is correct code, it is less code than what it replaced, and it gets approved on the merits, because on the merits it is fine. It becomes a defect at the moment it meets data that nobody had.

Now put that same defect in five different places and watch the price move.

The fix column never changes. The one on the right does.

The same bug, five prices

Caught at the agent turn, while the change is still on the screen, the fix is one line and there is nothing else. Nobody was interrupted. Nothing was queued. No one else knows it happened.

Caught by a local run, the fix is one line, plus the two minutes you spent finding out.

Caught in CI, the fix is one line, plus a red build, plus whatever is queued behind it, plus the time it takes to load the problem back into your head, because by then you had moved on to something else.

Caught in staging, the fix is one line, plus somebody else’s afternoon, because the person who finds it is usually not the person who wrote it, and before anyone can fix anything they have to work out which change in the batch caused it.

Caught in production, the fix is one line, plus an incident, plus people who cannot reproduce it because none of their accounts are big enough, plus a search through everything that shipped that week, plus a rollback that breaks something else, plus a customer who now has an opinion about your reliability.

Same defect. Same one line fix. The spread between the cheapest version and the most expensive one is several orders of magnitude, and none of it is in the code.

A bug does not have a fixed price. The price is set by how far it travels before something notices.

Quality lives in rings

Think about where a defect can be caught. Not as a list of tools. As a set of rings around the moment the code came into being.

The innermost ring is the agent turn, or your own hand on the keyboard. The code exists, nothing else has happened yet, and everything about why it exists is still loaded. The next ring out is a local run. Then the pull request and CI. Then staging, or whatever you call the place where things get checked together. Then production, where the customer is.

Every ring is further from the keystroke, and every ring multiplies the cost of the same defect.

Not because the fix gets harder. The fix is usually identical. One line, twenty seconds, whichever ring you are standing in. What multiplies is everything around the fix.

Three things that grow with every ring

The first is context. At the inner ring, the person who can fix this fastest has the whole problem in their head already. Two rings out they have moved on to something else and have to come back. Four rings out, three people who have never touched this code are in a channel at nine in the morning trying to reproduce something they cannot reproduce.

The second is batching. At the keystroke your change is alone. By CI it is one of nine in the queue. By production it is one of forty in a release. When something breaks at the inner ring you learn this change is wrong. When it breaks at an outer ring you learn something in this batch is wrong, and the first hours go to working out which thing, a search you would never have had to run.

The third is blast radius. Inside your editor, a bug costs you a minute. In CI it costs the queue, and everybody standing behind you in it. In staging it costs a team an afternoon and usually a meeting. In production it costs a customer first and you second, and the second part is the cheaper part.

None of these is about how serious the bug is. A missing limit on a query and a data corrupting race condition follow the same curve. What sets the price is not what the bug is. It is where it gets caught.

The same check finds the same bug in any ring. Only the bill changes.

What teams do when quality hurts

Here is the move almost everyone makes.

Something slips through and hurts. So you add a check. A required review. A new suite in CI. A staging soak before release. A sign off. Each one is defensible on its own, and each one is proposed by someone who is right about the incident that caused it.

And every one of them gets added at an outer ring.

That is not a coincidence. Outer rings are where you have control. There is a pipeline config to edit and a policy to enforce, and you can make a rule that nobody can get around. The inner ring is somebody’s laptop, and you cannot legislate a laptop.

I want to be fair to those gates, because I am not arguing against them. Every one of them catches real things. A bug caught in CI is a good outcome. A bug caught in staging is a good outcome. In the example above, a staging environment holding one realistic account would have caught it, and that would have been a cheap and completely respectable way to find it.

But notice what adding a gate does not do. It does not lower the price of anything that gets past it. And it makes the loop longer for every single change that goes through it, including all the ones that were fine.

So you end up with a longer loop and the same economics. The team feels slower and more careful, and the bugs that still escape cost exactly what they cost before.

Adding a gate changes what gets caught. Moving one changes what it costs.

Both are worth doing. Only one of them changes the price of the bugs you miss.

Move the check, do not add one

The way out is not a bigger pile of checks. Most teams already have plenty. The way out is to take the checks you already trust and run them earlier, where they are cheap.

Same check. Same signal. Different ring.

That sounds obvious written down, and it almost never happens, because moving a check inward is real work and adding one outward is a config change. Moving inward means the check has to become fast enough that nobody minds waiting for it. It has to be automatic, because a check that needs somebody to remember it is not at the inner ring, it is nowhere.

And then the hard part. The inner ring has to be able to see what only the outer rings could see. In the example, production was the only place with a customer big enough to expose the problem, which means moving that check inward is not a matter of running the same test sooner. It means giving the inner ring realistic data, a real dependency, a real running system to work against.

That is engineering, not policy. Which is exactly why it gets skipped, and why the gate gets added instead.

But look at what you get when you do it. A check that used to fire a week after a change now fires while the change is still open on the screen. The context has not been lost, because nobody left. The change has not been batched with forty others, so the answer is about this change specifically. Nobody else is standing in the blast radius. The same defect, caught by the same logic, at a fraction of the price.

The agent turn made this urgent

For most of software’s history the innermost ring was a person typing, and that ring had a natural speed limit. You could not generate defects faster than a human could produce lines.

That is over.

A coding agent now writes in ninety seconds what used to take an afternoon. But look at what happens next. The agent finishes, and the answer to did that work still comes from the same rings we built for the old world. It comes from CI, forty minutes later, batched with everything else. Or it comes from staging next week. Or, if you are unlucky, from a customer.

The generation ring got a hundred times faster. The verification rings did not move at all. So the ratio between them, which was always bad, is now absurd.

And it hits the agent harder than it hits us. When an answer takes forty minutes a human goes and does something else. An agent cannot. It either sits there, or it does the thing agents do, which is keep going, writing more code on top of code that nothing has confirmed. Every additional turn without feedback is another turn of work that might be built on something wrong.

An agent working without a check at its own ring is not fast. It is accumulating unverified work quickly.

The write step collapsed. The answer step did not move.

Fewer checks, closer in

I want to be careful, because this is not an argument for skipping verification. It is an argument about where you spend it.

If you move your checks inward you may end up running fewer of them, and that is fine. A defect caught at the agent turn never reaches CI, so it never consumes a CI slot, a reviewer’s attention, a staging afternoon, or an incident channel. The outer rings do not get weaker. They get quieter, and the things that do reach them are more interesting, because the cheap failures stopped arriving.

That is the version of this I would want if I were running the team. Keep the gates. Keep CI, keep staging, keep the release checks. Then start moving the checks you rely on most one ring inward, one at a time, and watch what stops showing up at the outer rings.

The teams that feel fast are almost never the teams with the fewest checks. They are the teams whose checks fire early. Their pipelines are not shorter because they cut corners. They are shorter because most of what would have failed there already failed somewhere cheaper, in front of the person, or the agent, who could fix it in twenty seconds.

That is the whole game. Not verify more. Verify further in.

The cheapest bug you will ever fix is the one that never got out of the room where it was made. Everything else is a tax on distance, and you have been paying it every day without seeing the bill.

WHERE I’M COMING FROM

This is the exact problem we are building IronBee to solve, verification that keeps up with the code your agents write. It is public now.

See how it works → ironbee.ai

Start using it → console.ironbee.ai

Now generally available

Ready to ship
AI-generated code with confidence?

Create your account and start verifying what your agents ship. Catch regressions before they reach production.

No credit card required.