All posts

Some bugs do not exist until you run the code

Static analysis and AI review can only judge how a change looks. The bugs that cost the most do not exist until the code runs.

Serkan Ozal9 min read
Every check was green. The bug shipped anyway.

Years before any of this, I build an observability product.

The job is simple to describe. Programs send us small records, which are called span, of what they did. We collect them, put them back together into one picture of one request, and store it so somebody can look at it later, when something has gone wrong. That picture is called a trace.

The hard part is that the pieces do not arrive together. One request touches six services. Six services on six machines, each with its own buffer, its own network, its own bad day. The pieces show up whenever they show up.

So we had to decide something. When is a trace finished?

For a while we just waited. Every piece that arrived pushed the deadline back, and only after a full minute of quiet did we call a trace finished. It worked. It also meant we were sitting on millions of unfinished traces in memory at any moment, and it meant your data showed up a minute after your request did.

Then we made it smarter (at least we tried).

The last part of a request to finish is the outermost part. The root. Everything else is nested inside it, so everything else ends before it does. So when the root arrives, the request is over. There is nothing left to wait for. Close the trace, index it, move on. No timer, no minute of waiting, no pile of half traces sitting in memory.

Read that again. It sounds right. It sounded right to all of us. It went through review and nobody argued, because there was nothing to argue with. The reasoning was clean. The code was cleaner. It was less code than the thing it replaced.

The root finishes last. It does not arrive last.

Nothing failed

An agent on a loaded machine flushes its buffer a few seconds late. A service having a bad minute holds its batch while it deals with something worse. One upload gets retried and lands well behind the rest. The root would arrive, we would close the book, and then a straggler would show up for a trace that no longer existed.

Nothing failed.

That is the part I want you to sit with. No error. No exception. No red line on any dashboard. Our own numbers said we were taking in everything, because from the inside we were. Every piece that arrived in time was stored perfectly. The late ones arrived for a trace that had already been filed, and we dropped them the way you drop mail for someone who moved out.

And the late pieces came from the services that were struggling.

So the traces missing data were the traces of the slow requests. The exact requests our customers had opened our product to understand. The healthier your system was, the more complete your traces looked. The worse your day was going, the less was there.

Nobody found it by reading

I know that, because we read that file many times.

Good engineers read it. They read it while chasing something else, and later they read it while chasing this. It survived every reading, and it survived for the same reason every time. There was no bad line in it. Every statement did the correct thing. The function did what its name said. No missing null check, no swallowed exception, no lock taken in the wrong order. The linter had nothing to say. The type checker had nothing to say. The reviewer had nothing to say, because a reviewer reads code, and the code was fine.

It was found by watching.

Somebody stopped looking at the logic and started looking at the clock. They took one real trace and lined up the arrival times of its pieces. Root at 4.1 seconds. Flush. A child at 11.3.

One picture, and it was over. What had been invisible in the text for months was obvious in the timing on sight.

One note on the details. The diff and the timings in this post are rebuilt from memory. The names are changed and the code is cut down to the part that matters. The shape of what happened is exact.

The same change, seen two ways. One view has nothing to say about it. The other one ends the argument.

So I gave it to a coding agent

Not long ago I pulled that change back up and put it in front of a coding agent. I asked it to review the change the way it would review something going to production. Be strict, I said. Assume there is a bug.

It gave me a good review. Careful, well organized, the kind of review I would be happy to get from a person. It liked that the change removed a timer. It suggested a clearer name for one variable. It asked whether we had a metric for flush latency, which was a fair question.

Then it approved it. It called the reasoning sound. And it offered to add a comment explaining why the arrival of the root span means the trace is complete, so the next engineer would understand the intent.

It would have helped us document the bug.

I want to be careful here, because this is not a story about a weak model. That review was better than most human reviews I have received. The model was not confused and it was not making things up. It read the change and it correctly described what the change said.

That is the whole problem. It read.

The ceiling

Here is what I think we get wrong about static analysis and AI code review.

We talk about them as if they are on a ramp. Today they catch some things, tomorrow they catch more, and one day they catch everything. Give the model a bigger context window. Let it see the whole repository. Train it on more bugs. It will get there.

It will get better. It will not get there. There is a ceiling, and the ceiling is not about intelligence. It is about what kind of thing code is.

Code is a description of a program. It is not the program.

When you read code you build a model of it in your head, and that model has to make assumptions to stay small enough to hold. You assume this call is fast. You assume this message arrives once. You assume the clock on this machine matches the clock on that one. You assume the root arrives last.

Every one of those assumptions is invisible in the text. That is what makes them assumptions.

A running program has none of that. It has real timing, real ordering, real load, real machines with real clocks that drift apart. The bug we shipped did not live in the file. It lived in the space between two machines, and that space only exists while both of them are running.

You cannot review your way into a space that is not in the file.

A better reader gets better at the top half. The bottom half is not a reading problem.

It is not only races

Race conditions are the easy example, so people treat them as the whole category. They are not. Once you start looking, the class is much bigger, and it is where most of the expensive incidents live.

Four bugs with no bad line in them.

None of these has a bad line. All of them have a bad moment.

Review was built for the first column. Most of your incidents come from the second.

What actually finds them

The bug in our pipeline was found by a timeline.

Not by a smarter reader. Not by a bigger context window. By running the thing and looking at what happened, in order, with real times attached.

That is worth saying plainly, because it should change what you build. If the bugs that cost you the most only exist while the program runs, then the only verification that can reach them has to run the program. Everything else is a guess about the program, and a guess is what we already had.

This is where I think AI review is pointed in the wrong direction. We took the most capable readers ever built and asked them to read harder. We should be handing them a running system instead.

An agent that can start your service, exercise it, watch what actually happened at runtime, and compare that against what was supposed to happen is doing something different in kind from an agent reading a diff. Not a better opinion. A measurement.

And this is possible now in a way it was not five years ago. The instrumentation problem is largely solved. You can attach to a running process and get spans, timings, ordering and errors out of it without touching the source. What was missing was something patient enough to run the software again and again, watch it closely, and notice that a child arrived eleven milliseconds after the door closed.

That was always a machine’s job. We just did not have the machine.

Two kinds of green

I still like code review. I still run static analysis. They catch real things, they are cheap, and I am not asking anyone to turn them off.

But I have stopped treating their approval as the same kind of fact.

When a linter passes, I know something about the text. When a review passes, I know a competent reader had no objection to the text. Both are true and both are useful. Neither one tells me the program works, because neither one has seen the program work.

When something runs the code, exercises it, and shows me what happened, I know a different thing. Not that it looks right. That it ran, and here is what it did.

Those are two kinds of green, and for years we let the cheap one stand in for the expensive one, because the expensive one was a person and people do not scale.

That excuse is gone. We have machines patient enough to run the software, watch it, and tell us what actually happened. Pointing them at the text instead is the waste of this era.

WHERE I’M COMING FROM

This is the exact problem we are building IronBee to solve, verification that keeps up with the code your agents write. It is public now.

See how it works → ironbee.ai

Start using it → console.ironbee.ai

Now generally available

Ready to ship
AI-generated code with confidence?

Create your account and start verifying what your agents ship. Catch regressions before they reach production.

No credit card required.