What a Verification Loop Adds to a Coding Agent: A First Look
This is the opening post in an ongoing series. We start with one model pair on one project, and the analysis will continue across more models and more datasets. We are sharing early on purpose, and we will keep sharing as we go.

Introduction
A coding agent can produce a lot of code quickly. What it cannot do, on its own, is know whether that code works. The model writes something plausible, the run moves on, and any mistake travels with it. On real multi-step projects this compounds: one unverified error early on can sink everything built on top of it. The fix human teams use is old and boring. Someone checks the work, against the running application, before anyone builds on it.
This post asks one question: how much does adding that check change what a coding agent can finish? Our hypothesis going in was that a closed verification loop (run the app, exercise the change, find the failure, drive the fix, then prove the repair) matters more than raw model strength for a large class of failures. Code is not done until it is verified.
To test this we needed a concrete verification layer, and we used our own. IronBee is a verification and fix layer for AI coding agents: when an agent writes code, IronBee opens the app in a browser, exercises the change, and looks for problems. When it finds one, it analyzes the failure, finds the root cause, drives the fix, and then re-verifies, so a fix counts only once it is proven. We built IronBee, so read this as a disclosed-interest experiment with an open method. In this post IronBee is the instrument, not the subject: the conclusions we care about are about verification loops in general, and the setup is public so anyone can rerun or challenge it.
We spent a week building a careful test, and we are sharing the opening results now, with more to come. We took real web development work and ran two models, one low-cost and open, one a commercial frontier model. Setting this up well was harder than it sounds, and the next section explains why. We treat this as a beginning, not a verdict. The point of the series is to see how the picture changes as we add more models and more kinds of work.
Methodology
Choosing the right dataset. We wanted work that looks real, not toy problems. We picked Web-Bench, an open benchmark from ByteDance. It is public, so anyone can check our setup: the code is on GitHub, the dataset is on Hugging Face, and the design is described in the Web-Bench paper. Web-Bench has 50 projects. Each project is a single web app assembled over 20 tasks that must be done in order, so each task builds on the one before it. That sequential structure is exactly why we chose it over a benchmark of independent, one-off tasks: a verification loop earns its keep when work builds on earlier work, so an unverified error compounds instead of staying local. Every task has a hidden test that checks the result, and the agent never sees it. If a task fails, the run stops there. This keeps the benchmark honest: an agent cannot skip ahead or hide a broken step.
For this first work we use one of these projects, the survey app, which builds a form with questions, required fields, and a preview. It spans the difficulty range we cared about, from a simple form to logic that has to survive the round trip between the design page and the preview. It is also a good match for what we set out to measure. A verification loop pays off while the code is being written, on exactly this kind of failure: code that reads as correct on its own but only breaks when the running app is exercised. So the lift we see here reflects what verification actually adds during real development, not an artifact of an easy or contrived task. Adding more projects is one of the first things on the list.
Choosing the models. We started with two. DeepSeek (deepseek-v4-pro) is a strong model with a low price, so it stands in for the low-cost coder we want to lift. Opus (claude-opus-4-8) is a frontier model, and it sets the quality bar. Putting them side by side lets us ask a clear question: can a verification loop bring a low-cost model up to the level of a frontier model? We built the harness so we can run several models through the same path, and this pair is where we begin. One note on the setup. In this first run DeepSeek is text-only, so it cannot see the screen. That is fine here, because the checks read the page as structured text, through the accessibility tree, the DOM, and the console, not as an image. The model reaches its result on structured feedback, not by looking at pixels.
What the verifier can and cannot see. Two rules keep the comparison fair, and they are worth stating plainly because they are the first things a skeptical reader will ask. First, the hidden Web-Bench tests stay hidden from everyone: the coding model never sees them, and neither does the verification layer. The tests live only in Web-Bench’s evaluator, which runs after the agent returns its files; IronBee derives its checks from the task description and the running application, the accessibility tree, the DOM, the console, not from the benchmark’s test files. Second, the loop is not free, and we count it: the verification layer makes model calls of its own, and those calls are included in every cost figure for the DeepSeek + IronBee arm in this post. And one more thing worth stating: the verification layer runs on the same provider as the coder, not a stronger one. Its analysis and fix steps use DeepSeek’s smaller, cheaper model, deepseek-v4-flash, while the code itself is written by deepseek-v4-pro. So the lift is not a more capable model slipped in through the verifier.
Setting the retry level. The agent gets a task, writes code, and a hidden test checks it. We give each task a small retry budget, and the retry level matters more than it looks. Some tasks are almost impossible to pass on the first try, by design: the hidden test checks a precise mechanism that the description states loosely, or even in a misleading way.
Task-9 asks for a checkbox to set the number of stars, but the test needs a control that can hold a number like 5 or 10, which a checkbox cannot.Task-15 asks to make questions required, where the test only accepts the browser’s built-in required marking, not hand-written checks that merely look right. A model that follows the wording literally picks the wrong mechanism and fails, and it cannot know until it sees the error. Because it does not always make the same choice, the same task can pass on one run and fail on the next, and the second try, now armed with the error, often recovers it. That is the point: the benchmark asks whether a model can read a failure and repair it, not just guess the intent blind. So one try mostly measures blind guessing. With three tries most agents get there in the end, so extra help is hard to see. With two tries real differences show up, so we run with two.
Is the loop just more attempts? One objection deserves an answer up front. A verification loop spends extra inference per task, so is the lift just a bigger compute budget in disguise? Partly it has to be. The loop does more work. But the retries the benchmark grants are blind: the model sees a failure signal and guesses again. The loop’s iterations are guided by evidence from the running application, which is a different kind of attempt, not just another one. Whether guided iteration beats an equal budget of blind retries at matched cost is exactly the ablation this framing demands, and it is planned for a future post: DeepSeek alone with a larger retry budget, against DeepSeek with the loop, dollar for dollar. Until that runs, read the results below with this open question in mind.
Measuring cost honestly. The cost turned out to be subtle. The built-in cost number is not reliable for a third-party model, because it prices every model as if it were the frontier model. The gap is large: applying frontier prices to DeepSeek overstated its cost by roughly twenty times. So for the low-cost model we did not trust the built-in number. We computed the cost from the provider’s own published token prices instead, and every cost figure in this post is built that way.

Table 1. Price Catalog for Opus and DeepSeek models.
Trusting the numbers. The same model can finish very different amounts of work from one run to the next. A single run tells you little. So we run each setup five times and read the averages. We also pinned one published version of IronBee, so every run used the same tool, and we kept parallel runs fully separate so they could not interfere with each other.
Scoring by difficulty. Web-Bench labels every task easy, moderate, or challenging, and the tasks are sequential. So we turn how far a run gets into a single score from 0 to 100, weighted by difficulty. The base plus all 20 survey tasks add up to exactly 100. This weighting is our own reading aid; the native Web-Bench metric is a simple pass rate that does not weigh by difficulty.

Table 2. The difficulty-weighted score. The base and the 20 survey tasks sum to exactly 100.
Results
Before the numbers, here is how we read them. The question of the post, how much does closing the loop change what an agent can finish, splits into two comparisons. First we compare a model against itself, with and without the loop. Then we compare across models: a low-cost model inside the loop against a frontier model without it, on both score and price. We take them in that order.
Analysis 1: How far each arm gets
Here we compare DeepSeek on its own against DeepSeek with IronBee.

Table 3. DeepSeek on its own. Score is out of 100, weighted by difficulty.

Table 4. The same model, with IronBee checking and fixing the work.
DeepSeek on its own clears the easy tasks and then stalls at the first moderate one in four runs out of five; one run reaches task 12. That is an average of 6.8 tasks finished out of 20, an average score of 20.4, and a median of 11. With the loop, the same model pushes through the moderate band and well into the challenging band: 17 tasks finished on average, a score of 80.6, a median of 86, and one run that completed the entire project. The worst run with the loop (62) beats the best run without it (48). In effect, IronBee provided 4x the problem solving intelligence compared to the model on its own, utilizing the same model and prompts.
The difficulty labels make this easy to picture. The baseline hits a wall at the line between easy and moderate tasks, and it stops. The loop carries the same model past that wall and into the hard tasks. The key point is what did not change. We did not switch to a bigger model or write a better prompt. The model is the same. What changed is that the agent now gets a second, evidence-based look at its own work. The checker exercises the change in a real browser, reports what is wrong, and the agent fixes it before it moves on.
The mechanism is worth dwelling on, because it explains why the lift is this large. In this project most of the hard bugs hide in the round-trip from the design page to the preview page: the code for each page can look correct in isolation while the hand-off between them is broken. That is exactly the kind of failure a first pass tends to miss and a check against the running application tends to catch. The wall in the baseline is not a lack of coding ability. It is the absence of any ground truth about whether the code works.

Figure 1. How far each arm gets across the 20 sequential tasks. Cells are shaded by task difficulty; solid cells are the tasks an arm completes on average. DeepSeek on its own stops at the easy-to-moderate wall, while DeepSeek with IronBee and Opus both push into the challenging tier.
Analysis 2: Same result, different price
Now we compare Opus on its own against DeepSeek with IronBee. First, Opus’s own runs and their cost; then the two arms side by side.

Table 5. Opus’s runs and their cost, with the loop arm’s cost alongside. Score range: 21-100. Cost is per run, at the published Opus prices from Table 1.

Table 6. The two arms land in about the same place on score, at roughly one-seventh of the cost per run. The DeepSeek + IronBee cost includes the verification loop’s own model calls. Per-run DeepSeek + IronBee costs: 2.27, 2.75, 1.77, 2.27, 2.81.
These two land in about the same place. Opus baseline scores 82.8. DeepSeek with IronBee scores 80.6, from Table 4 above. The roughly 4x that IronBee produced in the previous section is what brought a low-cost model level with a frontier one. And it did so cheaply: Opus baseline costs near 15 dollars per run, DeepSeek with IronBee near 2.4 dollars, about 7x cheaper. So verification, not a bigger model, reached frontier-level intelligence for a fraction of the price.
Two honest footnotes to that. First, the medians differ more than the averages: Opus, when it works, tends to finish the whole project, while the loop arm more often lands just short. Second, look at Opus’s k=1: the frontier model also had one run collapse at task 7. Open-loop failure is not a weak-model problem; it just costs weak models more often. That is exactly why the next arm we want to run is Opus with the loop, to see whether verification adds reliability at the top end too.
This is the practical payoff we wanted to test first. When quality matters, teams often reach for the most capable model by default. This first data point suggests another path: for some workloads, a low-cost model with a verification loop lands in the same place for far less. Whether that generalizes beyond one project and one loop implementation is what the rest of the series is for.
Closing
So what does this first look suggest? Closing the loop did not just polish the output. It changed how far a low-cost model could go, from a wall at the first moderate task to matching a frontier model’s average on this project, and it did so at a fraction of the frontier model’s cost. Those are two clear early signals, and they point at a simple idea: a lot of what separates a weak run from a strong one is not raw model power. It is whether anything checks the work, against the running application, before the agent moves on.
We also want to be honest about the limits, because this is a starting point. It is one project, one pair of models, and one implementation of a verification loop, ours. The same model can swing a lot from run to run, so we read these as early signals, not final numbers. The gains are not automatic either: verification helps most when the model can understand and act on what the check reports, so the size of the lift can vary from one model to the next. And the ablation that would separate guided iteration from a plain bigger retry budget has not run yet.
That last point is exactly why the series continues from here. The most interesting question is not whether verification helped one low-cost model on one project. It is how much it helps across the board, and where it stops helping. So next we are adding more projects, so the result does not rest on a single app. We are running the matched-cost retry ablation described in the methodology. We are adding an Opus-with-loop arm, to see whether verification raises reliability for frontier models too. We are moving to harder problems, where a first attempt is more likely to be wrong and a verification loop has more to catch. And we are adding more models, including multimodal ones, the text-only setup here is a lower bar, since the model reached this result without ever seeing the screen. We will keep the method and the numbers open as we go.
If you work on AI coding quality, we would love to hear how you measure it, and what you would want to see us test next. This is only the first post, and more is coming soon.
References
- GitHub Web-Bench repository: github.com/bytedance/web-bench
- Hugging Face: huggingface.co/datasets/bytedance-research/Web-Bench
- Web-Bench paper: arxiv.org/abs/2505.07473
- Claude Pricing: https://platform.claude.com/docs/en/about-claude/pricing
- Deepseek Pricing: https://api-docs.deepseek.com/quick_start/pricing