Skip to main content

Jev made our live tests 38% faster (and caught more bugs)

One of our favorite uses of AI is to run automated, variable testing. As an example, we have a live-test workflow that runs before we ship any changes. It spawns several agents to run tests in a real browser and CLI under varying conditions. It clicks through the product like a customer, checks each screen against what should happen, and even acts irrationally at times. After, it writes up what broke, if anything, with a suggested fix.

But over time, those runs became slow. Most of the time goes to the model deciding what to click next: about 18 seconds per decision, or 23 minutes of a 26-minute run.

So with the announcement of Jev, a fast decision model from TypeSafe AI, we were curious to see if this new primitive could help speed things up.

Since Jev can answer small questions in about half a second (“which button next?”, “is this step done?”), we decided to let it drive the browser. Claude (Opus 5.5) still does the heavy lifting: it orchestrates, writes the test plan, reviews the findings, and writes the final report.

Who does what. Before: Claude plans, drives the browser at about 18 seconds per decision, judges, and reports. Now: Claude plans, Jev drives the browser at about 0.5 seconds per decision, and Claude reviews and reports. If Jev isn’t sure what to do next, it hands that one step back to Claude.

To test it, we had a separate agent plant 10 bugs in a branch. Then we ran each setup 3 times and had a blind judge score what each run found. When Claude did all steps alone, it found 7.3 bugs per run. With Jev driving and Claude reviewing, that number improved to 8 out of 10.

It took 38% less time, a third of the cost, and was consistently better at finding bugs (8 of 10 in every run).

One run, 10 planted bugs. Time per run: 17.7 minutes for Claude Opus 5.5 alone, 10.9 minutes with Jev driving and Claude reviewing. Claude cost per run: $7.42 against $2.50. Bugs found: 7.3 against 8.0 of 10. Averages of 3 runs per setup (median time and cost), scored by a blind judge. The Claude-alone numbers come from an earlier round with a different set of 10 planted bugs.

For the curious: Jev on its own found only 3 of 10. It’s fast at picking the next step and weak at deciding whether a page looks wrong. So our current optimal setup is to let Jev drive and Claude judge. That split is now the default for our live tests.

It’s early, and 5 rounds of planted bugs is a small sample. But a fast, cheap model that makes one decision at a time is a useful primitive, like an if-statement with brains, and we’re already looking for the next workflow to plug it into.

By the way, if you’re excited about this new way of writing software, and like building with the latest tech, we’re hiring.