Visual Testing

Rules check specific things: whether the page has a button, whether its text is right. A shifted layout, a changed color, or a block that disappeared — you will not catch those this way, but they are immediately visible when you compare pictures. That is what visual checking is built on: a snapshot of the page is compared against a sample a human once approved.

The Step Snapshot

A test step declares what exactly to capture. There are four kinds of image:

  • Viewport — what is visible without scrolling. The fastest and most predictable kind.
  • Full page — including everything below the fold. Useful for long pages.
  • Component — a large block in full: a login form, an orders table, the header.
  • Element — a single detail: a button, a field, an icon.

There are also snapshots without an image — text ones. They store the page’s structure, its styles, the accessibility tree (what screen readers “see”), and the sequence of addresses. Such snapshots catch what an image cannot show: for example, the button stays in place, but the link under it now points to a different address.

The Baseline

A snapshot is not compared with the previous snapshot but with a baseline — a sample a human once approved. The baseline has a version history: you can see who approved it and when, and which build it came from.

The first build creates baselines on its own — otherwise there would be nothing to compare against. That is why the first build of a new step does not count as a failure: it simply remembers how everything looks now. Details — in Baselines.

Cheap First, AI Last

The comparison runs in order — from the cheapest method to the most expensive:

  1. Content fingerprint. A short code is computed from the picture; if the codes match, the content is identical. The comparison ends there: no pixels are counted, no AI is called, no money is spent.
  2. Strict comparison. The codes did not match — the pictures are compared pixel by pixel, and text snapshots line by line.
  3. Script classifier. If one is wired, every ambiguous case goes first to your script, and only what it could not resolve reaches the AI. The script may only raise the importance of a difference or close an ambiguous case — it cannot lower it and approves nothing. Its work spends no AI budget, and on failure the move simply passes to the AI as if the script did not exist.
  4. AI analysis. Only on the remaining unclear cases does the AI join in: it looks at the difference and tries to explain in words what changed.

Your script can do the comparing, too. If one is wired, the product stops comparing the pictures itself and hands the job to the script: it receives both snapshots — or both text files — as files in its working folder and returns a difference percentage. That percentage is then judged by the same thresholds as an ordinary comparison. Identical snapshots still pass by fingerprint without ever reaching the script, and if the script breaks, the product honestly records a comparison error — not “no changes”.

The last word still belongs to a human. Both the script and the AI may raise the importance of a difference, but they can never lower it and can never approve a snapshot for you.

Hence the economics: repeated runs cost almost nothing — identical pictures are filtered out by fingerprint, and you pay only for analyzing new unclear cases.

How to wire your own compare engine and script classifier — in Scripts & Extensibility and in Visual Testing Settings.

What a Build Is

A run with visual checks is called a build. It has a branch (the label baselines are tied to), a mode, and a review state — what a human has already looked at and what still awaits a decision. There are three modes: a full build runs every declared check, a smart one only the affected checks, and a comparative one compares two environments against each other. Details — in Build Verification.

What a Human Sees

A build has a one-line verdict. Looking at it is simpler than parsing the list of differences:

  • “Visual: pass” — all differences are resolved, nothing remains.
  • “Visual: review required” — there are differences a human must look at.
  • “Visual: fail” — something broke: checks failed or the build fell over.
  • “Visual: no data” — there are no snapshots at all. Usually this means the test’s steps do not yet declare what to capture.

Honest Limits

What visual checking does not do — so expectations stay realistic:

  • It does not check business logic. The correctness of calculations, stock balances, and access rights is not visible in a picture.
  • A difference does not explain the cause. It shows that the page looks different, but not why it happened.
  • A picture and accessibility catch different things. A snapshot sees colors and layout, but will not notice that a button became unreadable for a screen reader.
  • These are proofs, not a release permit. The release decision is still made by a human.
  • The first two or three weeks go into tuning. Until the regions that are not compared (dates, banners, random reviews) are marked, there will be many false differences.
  • Your own compare engine or classifier is Python code in the project folder. It runs with the same rights as the server, there is no sandbox, and it approves nothing: baseline approval stays with the human.

Next: how the Verify tab is arranged — the build list, the modes, and the review states — in Build Verification.

← Back to the documentation index