Developer at a desk at sunset between two robots: one on the left assembling code and app screens on blue holographic displays, one on the right holding a magnifying glass over a monitor showing red error warnings, illustrating a builder agent and a separate testing agent verifying AI-written software.

AI coding agents build faster, and they ship bugs at the same speed. The code isn’t the problem. Testing habits built for a slower pace are. Four things help: write the acceptance criteria before any code exists so the agent can’t grade its own homework, back the demo video with a Playwright run that checks the database and the network instead of just the clicks, give a second agent the job of breaking what the first one built, and test what the user actually gets rather than that the Save button exists. “Implementation complete” isn’t an answer anymore. Most of this verification can be done by AI too, so quality doesn’t have to cost speed. Writing code used to be the expensive part, and it isn’t now.

Most of what I build at Idea Warehouse now goes through an AI coding agent, and the speed is real. An idea that would have taken me a few weeks to get working can sometimes be running in a few days. The trouble is that the bugs show up at the same speed.

I see it in my own testing and I see it in products other people are shipping. Features land faster, and so do broken workflows and edge cases that should never have reached production. I don’t think the code these agents write is the problem. I think our testing habits were built for a slower pace and haven’t caught up.

Show me

My habit is simple. When an agent tells me a feature is done, I ask it for a video of the app actually working, clicking through the screens and finishing the workflow it just built. Then I go through the same workflow by hand.

That stops the agent from calling something finished because the code compiled or a unit test went green. But the more I do it, the more I think the video is only one piece of what’s needed.

Decide what working means first

If the same agent builds a feature and then decides whether it works, don’t be surprised when it decides it works. I’d rather write the acceptance criteria before any code exists, covering the happy path and the obvious ways it can go wrong, like bad input or a user without the right permissions. Where I can, those criteria become tests that fail first, and the agent’s job is to make them pass. It’s much harder for an agent to write tests that happen to approve whatever it already built when the tests were there before the code.

Make the video a real test

I still want the video. It’s the fastest way for me to see if something behaves the way I expected. But a video is a recording of someone clicking around, and it can look fine while the database is empty.

For web apps, Playwright can run the same workflow and check what actually happened. Was the record created, and does the database hold the right value? Did any network requests fail, or did the browser throw errors along the way? The agent can record the video while that test runs, so I end up with something I can watch and something a machine can verify.

Get a second agent to break it

This is probably the biggest change. The agent that wrote the feature shouldn’t be the only one testing it. I want a second agent whose instructions are close to the opposite: you didn’t build this, now try to break it. Paste in absurdly long text. Hit refresh halfway through a transaction. The builder is motivated to show the software works, so the tester needs to be motivated to show it doesn’t.

You can go a step further and let an agent loose on the app like a new user, trying whatever it can find. Regular automated tests are good at confirming that a workflow we already know about still works, and bad at finding out what happens when someone uses the app in a way nobody planned for. An agent exploring on its own can turn up a confusing screen or state that gets out of sync. When it finds something real, I have it reproduce the bug and write a regression test before anyone fixes it. The exploring finds the surprise and the regression test keeps it from coming back.

Test what the user gets

AI is very good at writing tests that cover everything and prove nothing. A test checks that the Save button exists, that clicking it fires an event and that a success message appears. It all passes, and the data was never saved.

The test I care about is closer to this: someone creates a project, closes the app, comes back later and their work is still there. I don’t much care how the agent built it. I care whether the person using it gets what they came for.

The same goes for how the screen looks. A button can pass every functional test while sitting hidden behind another element. On the screens that matter, comparing each new build to a screenshot I know was good will catch the CSS change that shoves half the interface out of place.

What done looks like now

“Implementation complete” doesn’t tell me much anymore. What I want back is evidence that the tests we agreed on up front passed and that a second agent tried and failed to break it, along with the trace and the video. Then I test it myself.

That sounds like a lot more process, but most of it can be done by AI as well, so we don’t have to trade quality for speed. We can take some of the time AI saves us in building and spend it on far more testing than we ever used to do.

Our development processes were designed around the idea that writing code was the expensive part. That isn’t true anymore. When another feature costs almost nothing to build, the hard part becomes proving it does what we meant, and proving that yesterday’s features survived today’s change.

I think verification is the next bottleneck, and the tools for it need to get as good as the tools we now use to write the code.

Frequently Asked Questions

Why shouldn’t the agent that built the feature also test it?

Because it has already decided the thing works. Ask it to prove that and it will write tests that pass. A second agent given the opposite instruction, you didn’t build this and your job is to break it, goes looking for the cases the builder never considered. It is the same reason you don’t review your own pull request.

What should an agent hand back instead of “implementation complete”?

Four things. The acceptance criteria you agreed on before any code existed, with results. A test run that checks real state, meaning the database holds what it should and nothing failed in the network tab. A report from the second agent on what it tried and what survived. And the video, so you can watch the workflow yourself. Then you run through it by hand anyway.

Isn’t a demo video enough proof?

No, because a video only shows the interface responding. It can look completely normal while nothing was written to the database. Run a Playwright test alongside it so you get both: a recording you can watch and assertions a machine checked. The video tells you it looks right, the test tells you it is right.

Why write acceptance criteria before the code?

It settles what working means while nobody has an answer to defend. Criteria written afterward tend to describe whatever got built. Turn them into tests that fail first and the agent’s job becomes making them pass, which is a much harder thing to fake than writing a test around finished code.

Doesn’t all this verification cancel out the speed gain?

Not if the agents do it. Writing acceptance tests, exploring the app, filing reproductions and keeping regression suites current are all things AI handles well. The point isn’t to slow building back down. It’s to spend some of the time AI gives you on far more testing than you could previously afford.