A run is one execution of a test on the real site. Its main property: after a run, evidence remains — a snapshot of every step and the result of every check. When something went wrong, you can see where exactly and what it looked like. Below is how a run goes and what its result means.
When a test declares snapshots for visual checks, the same run is called a build: it has a branch (which version of the site is checked), a mode — full, smart or comparative — and a human review state. Everything works the same, plus the snapshots are compared against baselines.
How a run starts
The test lands in a queue on the server. From there it is claimed by an executor — a Chrome with the addon installed and runner mode enabled. The executor checks the queue itself, once a minute, so nothing extra is needed from you.
The server itself can be the executor too. When all of a test’s steps are scripts, the server claims such a test from the queue at once, without a browser, and the runner line says “server”. Server-side checks this way stop depending on whether the computer with Chrome is on.
One browser executes one run at a time. This is not an oversight but a deliberate decision: a test occupies a tab and repeats a person’s actions, and two people cannot work in one tab at the same time. When runs are many, add several executors or a dedicated computer for them.
A build can be started with a branch right away (which version of the site to check) and a mode: a “full build” checks everything, a “smart build” takes only the affected places, a “comparative build” stands two environments side by side. The review state — awaiting decision, approved, rejected — a build gathers as it goes. All of this is taken apart in the Build Verification section.
What happens along the way
- The executor opens a separate tab and walks the test’s steps in order.
- Every step leaves a screenshot and, where applicable, check results.
- A step with a script is executed on the server, and its text output is visible in that step’s line.
- During a build, visual snapshots are additionally captured and uploaded to the server.
- The “Setup” and “Cleanup” phases run as separate lines — when a phase fails, the error names it.
- In the app you can watch the steps light up as they execute — a run can be observed live.
- At the end the tab closes itself.
During a run, the browser tab becomes active: the snapshot is taken from the page that is visible on screen. So a run is noticeable to whoever sits at that computer — one more reason to keep a separate machine for executors.
The run’s result
- Passed — all steps executed and all checks passed. The part of the interface under test works as expected.
- Failed — the steps went through, but some check did not pass. One is enough for the run to count as failed: this is exactly how the product notices breakages.
- Error — the test could not reach the end: a step failed. This happens both because of a breakage on the site and because of a stale action. When a step has “continue on error” on, the test goes on; otherwise it stops there.
- Visual — a separate result line: “Visual: pass”, “Visual: review required”, “Visual: fail” or “Visual: no data”. It is about comparing snapshots against baselines and does not replace the verdict on steps and checks.
Checks and steps are different things, and worth keeping apart. A failed check means “there is a problem on the page”, but the test continues and collects the remaining results. A step error means “cannot go further” — the test stops. Visual differences do not collapse into a step error: a snapshot differs — the build collects the differences and waits for a human decision.
Evidence
A run does not just say “failed” — it shows. In the run’s history lie the snapshots of all steps and the full list of checks, with the problem spots leading. You can open the step you need and see with your own eyes what the page looked like at that moment. That is the main difference from eyeball checks: a month later you can still explain what happened. A step with a script has no snapshot — it has no page; its evidence is the text output and the check records, and a script’s failed check files an issue exactly like a rule’s failed check.
When the executor disappears
The computer was switched off, the browser closed, the connection dropped — the run stays unfinished. After half an hour the product marks it with the error “runner disconnected”: it cannot hang in progress forever. Such a run can be started again, and it begins from the beginning — an unfinished one does not continue by itself, because the snapshots and results cannot be reassembled. The same rule covers server-side runs: if the server died in the middle of one, after half an hour it is closed with the “runner disconnected” error, so no runs hang around.
That is why a permanently powered-on computer with an open Chrome is kept as the executor for schedules. A regular work browser fits poorly: it gets closed at the end of the day.
A runner can be paused — for repairs or an update, for example. Before that it finishes the current build and stops taking new ones; when work resumes, the pause is lifted and the queue flows again. The same action in the Operations & Health section is called “drain the runner”.
Next: how to run tests on a schedule — in the Schedules section.