30 of 30 tasks completed
Qwen3-8B generated 50 action proposals. The v2 guard passed 45 directly and blocked or replaced 5 stale-row clicks before execution.
Inspect the 77 KB JSON artifactAdversarial CI for browser agents
Mutant Web runs your agent's action policy against hostile interface states: stale rows, duplicate identities, delayed confirmations, hidden totals, irreversible controls, and locale drift. The free runner works locally. A changing Team feed keeps the test from becoming memorized and stale.
No account · No prompt upload · Zero runtime dependencies
Free and local-first
Export a propose(context) function from your harness. Mutant Web supplies task, state, and allowed actions, then checks the proposal against the expected action and case policy.
- uses: MutantWeb/ci@v1
with:
adapter: test/agent-adapter.mjsAvailable now as MutantWeb/ci@v1 on GitHub Marketplace.
Tested, broken, corrected, rerun
The controlled evidence uses one browser environment, one model family, and thirty mutation cases. Its purpose is to expose a regression surface, not manufacture a perfect benchmark score.
Qwen3-8B generated 50 action proposals. The v2 guard passed 45 directly and blocked or replaced 5 stale-row clicks before execution.
Inspect the 77 KB JSON artifactThe locked offline matrix scored 25/30 with reasoning off and 24/30 with reasoning on. The five-case gap is the initial product surface.
The first guard allowed a Closed-row click while the task required Open items. We added a required-filter invariant and restarted the official run from zero.
The recurring product
A static test pack gets learned, patched around, and forgotten. Team access pays for the changing distribution: new seeds, new failure families, and reproducible historical versions.
TEAM
Or €1,990/year. VAT may apply.
Request founding access Founding access is €499 for the first year while the feed is in release-candidate status. No call required.Already purchased? Activate your license →Honest limits
Regression tests should travel with the code