Nobody Can Run Your App but You
Every testing tool that asks for your repository is promising to rebuild an environment it has never seen. It will get most of it right. The part it gets wrong is where your bugs live.
Somebody offers to test your application for you. All they need is access to your repository.
It sounds generous. It is worth stopping for a minute and writing down what they just promised.
They did not promise to read your code. They promised to reproduce the conditions your code runs under, well enough that what happens on their machine tells you something true about what will happen on yours.
Now go and look at what that actually involves.
A repository is the part that fits in a box
Start with the build. Which package manager, which lockfile, which runtime version, which native dependencies, which flags, in which order.
Then the runtime. Environment variables, most of which are not in the repository, because that is the entire point of environment variables. Feature flags, which usually live in a service somewhere and differ per environment. Config assembled at boot from three places.
Then the things your application talks to. A database at a particular migration. A cache. A queue with consumers. Two or three internal services. An identity provider. A payment sandbox with rate limits. An object store.
Then the data. An empty database reproduces nothing. Half of what makes an application interesting is the shape of what is already inside it.
Then the parts nobody writes down anywhere. How many CPUs. How much memory. Which of these things share a host. What the network between them looks like. What the clocks say.
None of that is in your repository. That is not an oversight. That is the definition.
A repository is the part of your system that fits in a box.

The environment problem has no bottom
Here is the part that does not get easier with effort.
Every customer is a different environment, and not slightly different. One is a single app with a managed database. One is fourteen services and a message bus. One has a dependency licensed to a specific machine. One is not legally allowed to let its data leave a country.
A vendor who runs your application has to hold all of that, and the list grows every time they sign somebody new.
There are exactly two ways out, and both are bad.
The first is to support everything. That is not a roadmap, it is an unbounded surface, and the failure mode is predictable: the customers who do not fit are the interesting ones.
The second is to enforce a shape. Put your app in this layout. Use this compose file. Expose this entrypoint. Make your seed data look like this.
Watch what that does. You change your application so that a testing product can hold it. The thing being accommodated is the tool. And whatever you changed to make it fit is, by definition, not the thing you ship.

Some systems do not fit in anybody’s sandbox
Some of this is not a matter of effort at all.
A dependency licensed per machine. A third party sandbox with a quota that three parallel runs will exhaust. Hardware. Data that is not allowed to be copied anywhere. A service mesh whose behavior is most of the product. Twelve services with real state between them, where standing up the twelfth one is a week of somebody’s life.
“Works on my machine” is a joke because environments are hard to reproduce even for the person who built them. A vendor promising to reproduce everybody’s machine is the same joke at a larger scale, told with a straight face.
What you hand over to make it possible
To boot your application somewhere else, that somewhere else needs whatever booting your application needs.
Your source code. All of it, not just the diff.
Your secrets. Database credentials, API keys, the signing key, the identity provider client secret. Not test values, because test values do not boot the application. The real ones, or shadow copies that are just as real.
And usually a slice of your data, because an empty database does not reproduce anything worth reproducing.
So in order to find out whether your code is safe, you begin by making it less safe. And this is not a one time upload with a delete button at the end of it. It is a standing arrangement. Every run needs it again.
I am not saying any particular vendor is careless with it. I am saying you have added a permanent copy of your most sensitive material to a system you do not operate, in order to answer a question about software you do operate.
Even a perfect rebuild is a different program
Now assume the impossible. Assume they got all of it. Every service, every variable, every row.
It is still not your system.
Concurrency is different, because the machine shape is different. Latency between services is different, because the network is different. Cache warmth is different. Lock contention is different. Which of two events arrives first is different, because arrival order is a property of the network and not of the code. Clocks drift differently.
That reads like a list of small things. It is a list of exactly where the expensive bugs live.
Nobody has a serious incident because a function returned the wrong string. They have incidents because two correct things met in the wrong order, or because a call that took four milliseconds in testing takes two hundred under real load, or because a queue that had never backed up backed up.
A rebuild randomizes the one dimension you most needed to hold still.
And notice which way the failure goes. It does not fail loudly. It passes. You get a green result from a system that was never the system you were asking about, and green is the answer that stops people looking.

Parity rots
Suppose they get it right on day one.
You add a service. You bump a runtime. Somebody changes a default in a config chart. A new feature flag appears. The vendor’s copy has to follow every one of those, forever, or the thing being tested drifts away from the thing being shipped.
That is a second environment to maintain. Teams already have one of those and they already resent it. The difference is that this one lives somewhere you cannot see, and you find out it has drifted the way people always find out, which is afterwards.
The other direction
There is another way to get a test and a running application into the same place. Almost nobody starts with it, because it is harder to build and harder to sell.
Do not move the application. Move whatever is doing the testing.
Your application already runs in several places that are genuinely yours. It runs in your CI job, at exactly the commit of the pull request, with the services your own compose file starts. It runs on a preview deployment, if it is the kind of thing that gets a URL. It runs on your laptop while you are still writing it.
Every one of those is the real thing. Nobody had to guess at any of it, because nobody rebuilt anything.
So the question stops being how do we reproduce their environment. It becomes how do we reach into it.
How we ended up doing this
This is the shape we landed on at IronBee. I will keep it short, because the argument matters more here than the product.
Something small starts inside your environment. Your CI job, or your own machine. It opens a connection outward and holds it open. Our QA agent connects back through it and uses the application that is already running there, the way a person would.
Nothing gets deployed. Nothing is exposed to the internet. There is no inbound firewall rule and no public URL. The connection is opened from your side, it belongs to that one run, and it disappears when the run ends.
The agent itself runs in a secure, isolated, disposable microVM. It is created for that one run and destroyed when the run ends, so nothing is shared between runs and nothing survives one.
The application under test is the one your pull request built, in the place it already runs, with the services it already has. Your source code and your secrets never move. What leaves your environment is observations, which is to say what happened while the agent was using the product.
Where the application does have a URL, like a Vercel or Netlify preview, we simply use the URL, because a preview is already a real deployment of your app. The tunnel is for everything that never gets one, which for most teams is most of what they change.


Six questions
If you take one thing from this, take the questions. They are short, and they sort tools quickly.
- Where does my application run while you are testing it?
- Does my source code leave my environment?
- Which of my secrets do you need, and where do they live afterwards?
- What happens the week I add a new service?
- Is the thing you tested the artifact I am going to ship?
- When your copy of my environment drifts from mine, who notices, and how?
A vendor who runs your application somewhere else needs a good answer to all six, and most of the honest answers are uncomfortable. A vendor who tests your application where it already runs mostly does not need answers, because the questions stop applying.

Your environment is not a configuration you can hand to somebody. It is a place. And the only useful thing to do with a place is go there.
WHERE I’M COMING FROM
This is the problem we are building IronBee to solve. An AI QA engineer that reaches your application where it already runs, in CI, on a preview, or on your own machine, and reports what actually happened. It is public now.
See how it works → ironbee.ai
Start using it → console.ironbee.ai
