Nobody Writes a Test for the Bug They Did Not Think Of
Ask a model to write tests and it grades its own homework, on the questions it chose. Ask it to use the software and it has to meet an answer it did not write.

A few weeks ago I ran a small experiment on our own console, because I wanted to look at this failure instead of arguing about it.
I handed a coding agent one sentence, lifted straight out of our own API docs. A user may hold at most 10 access tokens per account. Implement the limit, I said, and write tests for it.
It wrote the obvious thing. Count the caller’s tokens for the current account, and if there are already ten, refuse the eleventh with a 409.
Now read that sentence again. It has two readings, and they are not the same system.
One is a per person limit. Every member of an account may hold ten of their own. The other is a per account limit. The account holds ten in total, across everybody in it. The first bounds how many keys one engineer can leave lying around. The second bounds how much of your product a single compromised account can reach. Ours exists for the second reason. The agent implemented the first.
Then it wrote the tests.
Six of them. Create ten, expect the eleventh to be refused. The boundary at nine. Revoke one and confirm you can create again. A token belonging to a different account does not count toward yours. Bad input. Careful, well named, every one of them green.
And not one of them has a second person in it.
So the account limit is now ten times the number of members, and nothing anywhere reports a problem. The endpoint behaves exactly as its tests describe. You would notice this the first time you invited somebody and looked at what the account actually held. You would never notice it by reading a green suite.
The model was not being careless. It picked one reading of an ambiguous sentence, which is exactly what a person does. Then it wrote the exam for its own reading, and marked it correct.
That is the part worth staring at. Not the mistake. The grading.
You can reproduce this on your own code in five minutes. Find a sentence in your system with two plausible readings, ask for the implementation and the test suite in the same breath, and watch the suite agree with whichever reading the implementation picked.
A test is a written down expectation
A test is not an experiment. It is an expectation, written down in advance. Somebody decided what should happen, and the run either matches that decision or it does not.
Which means everything a suite can tell you was already in somebody’s head before the suite ran. The set of bugs a suite can catch is not the set of bugs that exist. It is the set of failures that occurred to someone.
That gap is where incidents come from. Nobody sat down, imagined the exact way their system would fall over, wrote a test for it, and then shipped the bug anyway. The failure was outside the list, because if it had been on the list it would have been fixed instead of tested.
And here is what does not fix that. Writing more tests from the same understanding. Ten more assertions about the things you already worried about do not move the boundary of what you worried about. They make the inside of the circle denser. The bug is outside the circle.

The author grades their own homework
Now add the part that is new.
For most of software’s history the person writing the tests was at least a slightly different mind from the person writing the code, and often a completely different one. Even when it was the same engineer, a week had passed, or a reviewer was involved, or QA existed. Independence was not guaranteed, but it was not zero either.
That margin is where the ambiguous sentence used to get caught. Somebody who had not made the assumption looked at it and asked the annoying question.
When one agent writes the code and then writes the tests, that margin is gone. Same model, same context window, same reading of the ticket, same wrong assumption if there is one. In the experiment, the function and the tests were not two things checking each other. They were one reading of one sentence, written down twice.
The suite is not checking the code. The suite and the code are the same claim, written twice.

What tests are actually good at
I am not building to the conclusion that tests are a waste. They are not, and I would not work on a system without them.
But it is worth being precise about the job they do well.
A test suite is memory. Once you know that a behavior matters, a test pins it in place so nobody breaks it again by accident. That is enormously valuable and there is no substitute for it. Every bug you have already paid for should end its life as a test, so that you never pay for it twice.
What a suite is not is a way of learning something new. It cannot be, because every question in it was written by someone who already had the answer.
So the mistake is not writing tests. The mistake is asking memory to do discovery, and then feeling verified when it comes back green.
Using the software is a different question
There is another thing you can ask an agent to do, and it is not a better version of writing tests. It is a different activity.
Instead of asking for assertions, ask it to use the application. Sign in. Create the thing. Edit it. Delete it. Do it again with an account that has fifty thousand records instead of four. Do it with the network made slow. Go through the flow the way somebody who has never seen the product would go through it, including the wrong turns.
And then read what the system emitted while that was happening. The requests it made. The queries it ran and how many. The errors it logged even when the screen looked fine. The latencies. What the database looked like before and after.
The difference is not effort. It is the direction the information flows.
A test asks whether the program matches a prediction. A run tells you what the program did, including all the parts nobody thought to predict.
But what says it is wrong?
The obvious objection is that if nothing asserted anything, nothing can fail. Somebody has to decide what counts as broken.
That is true, and the answer is the interesting part.
Some things are wrong no matter what the feature was supposed to do. You do not have to know the refund rules to know that the request came back with a 200 and an error object inside it. You do not have to know the product to know that loading one page issued 1,847 queries, or that a delete left a row pointing at a parent that no longer exists, or that the second call took eight times longer than the first, or that a customer’s email address ended up in a log line.
None of those judgements require knowing what the author intended. That is the whole point. The examiner does not have to share the author’s understanding, because it is not grading against the author’s understanding. It is grading against things that are true of working software in general.
That is what independence means here, and it is the thing a generated test suite structurally cannot have.

What this finds that a suite never will
The failures that come out of actually driving software are a recognizable family, and none of them had a test that was going to catch them.
The feature works and the screen is correct, and underneath it the page made one query per row, which is fine at forty rows and fatal at forty thousand. The call succeeds and returns a helpful looking response, and in the logs there is a stack trace that somebody caught and swallowed two years ago. The write lands and the second write does not, and now there is a record with no owner, which nobody will notice for a month. Everything is fast the first time and slow every time after, because the first request populated a cache and the second one queued behind a lock. The flow completes and leaves the user on a screen with no way forward, which is not an error in any technical sense and is a complete failure in every other sense.
Every one of those is visible in about ten seconds if something exercises the software and looks at what happened. Not one of them is visible in a test file.

Keep both. Know which is which.
So the practical version of all this is not complicated.
Keep the suite, and be honest about its job. It is there so that things you already learned stay learned. When something does get found, that is when to write the test, because now you know the question is worth asking.
And stop asking the thing that wrote the code to also write the exam. Ask it to use the software instead, at the moment the change is made, and to bring back what the system actually did rather than what it expected the system to do.
That is a different kind of answer, and it is the one you were looking for when you asked for tests in the first place.
An agent that writes the code and the tests has answered its own question. An agent that runs the software has to meet an answer it did not write.
WHERE I’M COMING FROM
This is the exact problem we are building IronBee to solve, verification that keeps up with the code your agents write. It is public now.
See how it works → ironbee.ai
Start using it → console.ironbee.ai
