Reset
Start each task from the same declared precondition.
Design-partner validation
Runstance is validating deterministic regression evidence for browser agents: reset the task, repeat the run, grade the resulting state, and gate the release.
Not a finished product. No credentials or production access.
Agent reported success, but address.postcode was not persisted in 2 of 5 runs.
For teams shipping
Browser agentsComputer useWorkflow agentsAgent QAA confident completion message can hide a wrong cart, a stale record, or an action that only worked once. Runstance is designed around the environment state that remains after the agent stops.
Start each task from the same declared precondition.
Expose flakes and drift across versions and multiple runs.
Check the observable postcondition, not the agent’s claim.
Ship, investigate, or block with evidence attached.
The validation targets the gap between task completion and dependable release evidence.
The trace looks plausible, but the environment says otherwise.
A new model or prompt improves the average and breaks a critical task.
One lucky pass hides instability that appears only under repetition.
Two paid commitments required to proceed
We are looking for teams with a real missed regression and a safe way to reproduce it. The design-partner pilot only starts after scope, payment, and data handling are agreed.
Discuss a qualifying taskRefundable before kickoff. No promised outcome, certification, or production integration.
Only a redacted trace or isolated sandbox is in scope.
Technical artifacts are deleted within 30 days or sooner on request.
No model training, public benchmark, logo, or case study without separate written permission.
One regression is enough to start