Build Failure Triage

When a build fails, the important thing is not to dissect the same thing ten times over. If the same query broke on fifty pages, that is one cause, not fifty problems. So failures are first gathered into clusters by cause, and each cluster is then worked through once. In the interface this section is called triage — failure analysis.

Clusters by Cause

A cluster is a set of failures with a common cause. The causes are:

  • Same rule failure — one and the same check failed everywhere it occurred.
  • Same step difference — a specific test step fails, not the whole interface.
  • Environment or network failure — the stand did not come up, the connection dropped, the wait timed out.
  • DOM change — the page’s structure changed, and the declared elements are no longer found.
  • Mixed causes — the cluster gathered different things; working through it takes more care.
  • Unclassified cause — there was not enough data.

A cluster shows how many members it has and which snapshots and steps are inside. The ordering is set by an importance estimate: it adds up severity, coverage (how much is affected), novelty, and which interface layer is hit. There is also a plain ordering — by member count. Clusters can be rebuilt when new data appears; this takes the member role.

From a cluster you can create an issue — the product will attach the difference images, their snapshots, and direct links. If this cause has been seen before, no new issue is opened: the old one is simply strengthened with new evidence. Rule-failure clusters have no visual evidence, and they are already tracked on the Issues tab — creating an issue for them again is unnecessary.

The Narrative

The triage carries a short explanation: what exactly happened and why it turned out this way. The AI writes it — but only from the build’s facts, not from guesses about your code.

If the AI is unavailable or off, the product does not invent an explanation but shows dry statistics: how many steps passed of how many, how many checks failed, how many snapshots changed, appeared, and disappeared. At the bottom it says honestly which one you are reading — “AI narrative” or “Dry statistics”. Precise numbers beat a plausible invention.

An AI Suggestion and a Human Verdict Are Different Things

Every cluster has two separate fields: Suggested (the AI’s hint) and Reviewer verdict (the human’s decision). They are stored separately and never substitute for each other. The AI can suggest “intentional change”, “real regression”, or “unsure”. The human writes their verdict beside it, and you can see who recorded it and when.

The point of the separation is simple: the hint helps you understand what is happening faster, but responsibility for the conclusion stays with the human. Even if the AI confidently said “regression”, it remains just a hint.

Cluster Verdicts

The verdict vocabulary is the same for the suggestion and the decision:

  • Real regression — the interface is genuinely broken. The interface itself needs fixing.
  • Flaky test — the test sometimes passes, sometimes fails; the site is fine.
  • Test maintenance — the capture declaration or the step is outdated; the test needs updating.
  • Environment — the stand or the network is at fault, not the code and not the test.
  • Intentional change — the interface was changed on purpose; this is how it should be.
  • Unclassified — it could not be worked out.

Verdicts here are not just a label: two of them open the road to test healing — “Test maintenance” and “Flaky test”. The “Real regression” verdict closes that road: the interface needs fixing, not the test, so healing is unavailable in that case. It is just as unavailable for environment problems and unworked clusters.

Budgets

Failure analysis costs money — the owner pays for every AI request. So the spend is visible beside it: how many tokens went in a day, how many heal campaigns are running right now, and how many heal attempts have been spent since the last successful run. If the “stop on identical repeating errors” breaker is on, it is honestly marked: stopped — the same error repeats.

It is important to understand the budget’s role: it stops spending but never approves anything. An exhausted budget is not “consent by default” and not “skip the check” — it is simply the absence of new hints. When the AI is off, repeated runs are free, and the product says so directly.

Next: how the test healing a reviewer’s verdict opens is arranged — in Test Healing.

← Back to the documentation index