Your AI QA Engineer Now Runs on Every Vercel Preview
IronBee opens the preview deployment nobody opens, uses your app the way a person would, and comes back on the pull request with the root cause. It runs inside CI too, with nothing deployed at all.
Every pull request on Vercel comes with a complete, running, production shaped copy of your application. It builds in under a minute. It is isolated, it is per change, and it exists for exactly as long as the review is open.
And most of the time, nothing happens to it. Someone reads the diff, approves, merges. The best QA environment our industry has ever handed out for free mostly sits there unopened.
Today we are shipping the part that was missing. IronBee is an AI QA engineer that opens that preview and uses your product, on every pull request, before anyone merges anything.

What is new
Three things went live.
A Vercel integration. Connect your project and IronBee picks up each preview the moment it is ready, tests the application, and reports back on the pull request.
A GitHub App, with CI support. Install it on a repository and IronBee runs on your pull requests. When there is no deployment at all, it connects into your GitHub Actions job through a secure tunnel and tests the app running inside the runner.
A new place for the QA agent to run. Every run gets its own isolated, disposable microVM. Nothing to provision, nothing shared between runs, nothing left behind afterwards.
On Vercel
Connect a project once. After that, every pull request works like this.
The preview finishes building. IronBee picks up the URL, reads the change in the pull request, and works out what a person would actually do differently because of that change. Then it goes and does it. It signs in. It fills the forms. It clicks the wrong thing on purpose. It walks the flow the way somebody who has never seen your product would walk it.
There is no script to write and no selector to maintain. Nothing breaks the next time a button moves.
Then it comments on the pull request, usually before a reviewer has finished reading the diff.
The reason this fits Vercel so well is that a preview is already what QA always wanted and never got. A clean environment, one per change, with nothing else in it. No queue, no shared staging to coordinate over, no thread asking who broke the environment.
We were not going to find a better place to put a QA engineer.
In CI, with nothing deployed
Plenty of what you ship never gets a URL. Backend services. Workers. A package inside a monorepo. Anything where the build produces an artifact instead of a deployment.
So IronBee also runs inside your CI job.

A GitHub Actions run already starts your application. That is what your integration tests do. IronBee connects into that job through a secure tunnel and drives the application running right there in the runner. Nothing has to be deployed anywhere, and nobody has to stand up another environment to own.
If it can run in CI, IronBee can use it in CI.

Where the QA agent runs
This is usually the next question, and it is the right one. Where does this thing execute, and what does it get to see?
Every run gets its own isolated microVM. It is created for that run, it holds only that run, and it is destroyed when the run ends. Nothing is shared between runs and nothing survives one. There is no long lived machine sitting around with your tokens on it, no browser grid to maintain, no runner pool to keep warm and patched. There is nothing for you to provision at all.
That matters more than it sounds. A QA engineer needs real access to be useful. It signs in as a real user, it talks to real services, it sees real responses. The only sane way to hand something that much access is to give it a machine that exists for a few minutes and then stops existing.
It does not stop at “it failed”
This is where we differ most from the rest of this category, so let me be specific.
Plenty of tools can tell you that a flow broke. You get a red X, a screenshot of a spinner, and a suggestion to retry. Which leaves the actual work exactly where it was. Somebody still has to reproduce it, find the cause, and decide what to change.
IronBee collects the evidence while it is driving your app, not after.
If you already emit OpenTelemetry, we read it. If you install our SDK, we get more: runtime introspection into the application while the QA agent is using it, so we can see what the code was doing at the moment the screen went wrong. Requests, traces, queries, logs, timings and state, all lined up against the exact user action that produced them.
So the report is not “checkout is broken”. It is the request that failed, the service that was slow, the query that came back empty, the exception that got swallowed, and the line in this pull request that caused it.
And when the fix is clear, IronBee will write it and open a pull request.
That is the difference between a QA engineer who files a ticket saying “broken” and one who walks over and tells you what is wrong. Every team has worked with both.

It does not take the agent’s word for it
Here is the second difference, and it is the one I would want to hear about if I were reading somebody else’s announcement.
You hand an AI a changeset and tell it to test the change. It goes off, does some things, and comes back saying everything looks good.
How do you know it actually tested your change?
You do not. Not from the answer. A model that got confused, or ran out of room, or quietly decided a flow was too hard to reach will produce the same cheerful summary as a model that did the work. And because the summary is well written and specific, it reads like evidence. It is not evidence. It is a claim.
So IronBee does not trust it. While the QA agent is using your application, we trace what actually executes inside it. Then we take the changeset from the pull request and check it mechanically. Did this function run during the test. Did this branch get taken. Did this handler get called.
If part of your change never executed, we do not pass it. We tell you which part was not reached, and why the run never got there.
An agent saying it tested something is a claim. An execution trace is a fact.
Those are not the same kind of green, and treating them as the same is how a team ends up trusting a QA agent that has quietly stopped doing its job.

Getting started
Install the GitHub App. Connect your Vercel project. That is the setup.
After that, open a pull request. The preview builds, IronBee uses the application, and you get a comment with what it tried, what broke, the root cause with the evidence behind it, and which parts of your change actually ran. If nothing broke and the whole change was covered, you get a short green comment and you get on with your day.
Nobody on your team has to write a test, maintain a selector, or remember to click the preview link.
Why now
Coding agents already write more code than anyone can carefully review. That is not a complaint, it is just the shape of the work now. The bottleneck moved from writing code to knowing whether the code works.
Meanwhile Vercel quietly built the best possible place to answer that question and gave it to everyone. Every pull request already comes with its own live application. The gap was never the environment. The gap was that using it was somebody’s job, and nobody had time.
Now it is somebody’s job again. Your AI QA engineer, on every preview and every CI run, catching bugs before you ship.
TRY IT ON YOUR NEXT PULL REQUEST
IronBee is public. The Vercel integration and the GitHub App are live, and the CI path works today through GitHub Actions. If you try it on a real pull request and it misses something, I would rather hear about that than the compliment.
See how it works → ironbee.ai
Start using it → console.ironbee.ai
