Test Healing

It happens that not the site broke but the test itself: the capture declaration is outdated, a step waits for a button that was renamed, the run is unstable. Fixing the interface is not needed in that case — the test needs a correction. This is test healing: the agent prepares a change to the test file, but applies it only with a human’s consent.

How It Starts

Healing cannot always be started. It is opened by a reviewer’s verdict on a failure cluster — “Test maintenance” or “Flaky test”. The “Real regression” verdict closes the road: the interface needs fixing, and changing the test would only mask the breakage. For environment problems and unworked clusters, healing is unavailable too.

A reviewer or a person with a higher role can start healing. If a campaign is already running for a cluster, a second one is not started — instead, the current one opens.

Propose → Approve → Verify → Revert

The healing path is built so that a human always stands between the proposal and the file:

  1. Proposal. The agent studies the failures and prepares an exact change to the test file. Nothing is written into the file at this point — the proposal simply awaits a decision, with an explanation beside it of why the agent suggests exactly this.
  2. Approval. A human reads the change and approves it. On approval a reason can be left; on rejection it is mandatory.
  3. Write and verify. The file is rewritten in a verified way, and its previous content is preserved — down to the last byte. Then a verify build starts, and it decides the heal’s fate: if it passes, the heal is accepted.
  4. Revert. If the verify build does not pass, the heal stays in the applied-and-awaiting state. A human either reverts it or approves it again. There is no automatic revert — the decision is always human.

Along the way the campaign’s states are visible: “Requested”, “Proposal”, “Proposal awaiting review”, “Applied — verify running”, “Verified heal”, “Reverted”. There are also outcomes without a heal: “No heal needed”, “Budget exhausted”, “Stopped — no progress”, “Round failed”. Earlier campaigns for this test stay in the list, so the history of decisions is not lost.

Revert Restores the Exact Previous File

Reverting restores the test file to exactly what it was before the heal was applied — from that very saved content. If other edits made it into the file after the heal, they will be overwritten, and the product warns about this right in the revert dialog. The revert’s reason lands in the audit log, as does the application itself: from the log you can understand what the agent proposed, who approved it, and why it was reverted.

Budgets and Circuit Breakers

Heals are the most expensive part of the agent’s work, so they are limited by hard circuit breakers. There is no way around them:

  • Token ceiling per task — beyond this volume the agent will not spend on a single task.
  • Heal attempts per test — how many times one and the same test may be treated since its last successful run.
  • Heal attempts per build — a shared limit for one build.
  • Concurrent heal campaigns — how many heals may run in the project at once.
  • Stop on the same error — if an error repeats and a previous heal already tried to fix it, the campaign does not start.

A budget denial never disappears silently: it is recorded with a reason. The reasons are: the task’s token ceiling is exhausted, the attempts for the test or for the build are exhausted, another campaign is already running, the error repeats, visual AI is off for the project. Such a record can explain why the agent “did nothing” and help you decide whether to change the limits. A reminder: a budget only stops spending — it cannot approve a heal.

Next: how to see how many tasks the AI completed, what they cost, and why something never started — in AI Log.

← Back to the documentation index