Visual Testing Settings

Visual checks are a delicate matter: somewhere every detail matters, and somewhere an animation gets in the way. So they have settings of their own. The good news: every setting has a sensible default, which is why a project without tuning works exactly like a project with the recommended values.

The settings live in the project file, in the visual: block. You can edit them right in the interface, or view the result and copy it as ready-made text.

Application presets

The easiest start is a ready profile: eight presets describe the typical cases. One choice replaces thirty fields.

  • SaaS dashboard — data-dense internal screens: strict pixel compare with masked dynamic regions.
  • Marketing site — public pages with animated banners: full-page captures and relaxed thresholds.
  • Canvas / WebGL — canvas applications: layout-only compare plus masks over the drawing areas.
  • E-commerce — catalogs and checkout: the text layer is enforced, to catch price and copy drift.
  • Documentation — help and blogs: full-page captures tolerant of content churn.
  • Mobile-first — products where the phone matters more: 390×844 is captured before the desktop size.
  • SPA — single-page applications: an extra serial frame surfaces render flake.
  • CMS — editing interfaces: the text and DOM layers are logged.

An honest caveat: applying a preset overwrites the whole visual configuration wholesale, including your edits. It can only be undone by a new edit, so pick the preset before fine tuning.

Capture

The Capture section answers the question of what gets captured and how. The default capture type is set here: the visible area, the full page, a component, or a single element.

Then comes the viewport matrix: one capture per size. This way the same page is checked both on a wide monitor and on a phone.

Masks are named areas painted over before the comparison. They hold what changes by itself and must not count as a difference: clocks, “2 minutes ago” labels, avatars, ad blocks. A mask is defined either by an element selector rule or by a rectangle with coordinates; it has a name and a color. Ignored elements are a similar device: they are painted black and shown without a name or border.

Stabilization

Before taking a shot, the product calms the page down — otherwise two captures of the same screen would differ all by themselves. Every technique is on by default, because a shaky capture is worse than a strict one.

  • Freeze time — ticking clocks and countdown timers stop showing a new instant on every capture.
  • Seed random — random numbers and shuffled lists stop changing the layout.
  • Disable animations — the shot does not catch a half-faded element mid-motion.
  • Hide caret — the blinking input bar creates no difference.
  • Normalize scrollbar — it renders the same way across systems.
  • Wait for network idle — late responses do not repaint the page after the shot.
  • Wait for DOM settle — the application stops changing the markup.
  • Wait for fonts — a font substitution does not shift all the sizes.
  • Block third-party domains — ads and counters do not show a different creative; the domains you need can be allow-listed.
  • Serial frames — when more than one, several frames are taken and compared with each other: momentary flicker shows up as “unstable” instead of landing in the diff.

Check layers and thresholds

A check consists of several layers, and each has its own mode: “enforce” — a failure fails the build, “log” — recorded without failing, “disable” — the layer is skipped entirely. The layers are: visual, text, network, console, accessibility, design tokens, performance, and addresses.

The comparison engine is chosen in the project settings — there are three:

  • Pixel compare — compares images pixel by pixel and counts how many diverged; the default value.
  • Structural compare — looks at the image as a whole and estimates how similar it is to the baseline.
  • Script compare — your own python script does the comparing. It is named in the project settings: the key and static arguments go into the visual.diff.script block. Without a named script the product will not let you enable this engine.

An honest caveat about the script engine: its failure is a comparison error, not “no differences”. A broken engine never shows a green result. Details — in Scripts & Extensibility.

The visual layer’s thresholds define what counts as a difference:

  • Max diff pixels — an absolute limit; an empty value turns the limit off.
  • Max diff pixel ratio from 0 to 1 — also switched off by an empty value.
  • Color sensitivity — pixels with a smaller delta do not count.
  • Color tolerance — colors distinguishable by less than this threshold are indistinguishable to most people.
  • Require same dimensions, layout only, and detect vertical shift — three separate rules for special cases.

So you do not have to pick the numbers by hand, there are three ready sensitivity sets: strict (threshold 0.1), recommended (keeps your numbers; the default threshold is 0.2), and relaxed (threshold 0.3).

Branches

Your application can have branches — versions of the code being worked on. The settings say how to behave with each:

  • Default branch — what a build is compared against when its own branch has no approved baseline. This is the fallback.
  • Auto-approve branches — builds on matching branches may accept provably safe differences on their own. An empty value means nowhere.
  • Mandatory review branches — builds on matching branches always require a human decision. An empty value means no such branches.

Branch patterns are written the usual way: * — within one path segment, ** — across several. For example, release/* or feature/**.

Baselines are bound to the branch with a fallback to the default branch: a build is first compared with its own baseline, and when there is none, with the main branch’s baseline. The product does not read the tested application’s repository history, so there is no complicated common-ancestor arithmetic here.

AI and budgets

Artificial intelligence is plugged in only where it is hard without it. It has a master switch, and two tasks are on by default: ambiguous diff analysis (the last step of the comparison, where simple rules could not decide) and failed build analysis (grouping and explanation). A model can be set separately; when it is not, the project agent’s model is used.

There is exactly one AI provider here — the connected Xedant Agent; there are no direct calls to anyone else’s models. Replays cost nothing: the plain comparison answers first, and only the ambiguous goes to the model.

Before the AI, your own classifier script can do a pass: it is wired in the project settings, in the visual.ai.classifyScript block, and handles only ambiguous differences — before any call to the model. The script may only raise a difference’s importance or settle an ambiguous case; it cannot lower the importance. Its work consumes no AI budget, and on failure the turn simply passes to the AI. The picker has a “None” option — then this step simply does not exist.

AI calls cost money, so there are ceilings — circuit breakers: tokens per task and per campaign, the number of heals per test and per build, how many heal campaigns run at once, and a stop when the same error keeps repeating without a change. A campaign’s ceiling is checked at its phase boundaries, not in the middle of the work. What matters: the circuit breakers stop the spend but approve nothing — neither a plan nor a test.

Notifications

The product can knock on the outside world: you set a notification URL where a short message about an event goes, and there is a “Send test webhook” button to check the connection. Notifications are sent for two events — “review required” and “build failed”.

An event has a threshold, so a notification arrives only when the threshold is crossed: a one-pixel difference will not wake anyone.

The effective configuration

The bottom of the settings shows the effective configuration — the result with every default substituted, exactly what would be written on saving. It can be copied as ready text, and beside it sits the list of validation errors.

Both script wirings are visible here too: the chosen comparison engine (diff.engine: script), the diff script itself (diff.script), and the classifier script (ai.classifyScript). It is always clear what is in force right now.

A useful detail: a broken setting does not prevent reading the data — the default values simply take effect, and the error itself is shown honestly. The product will not accept an invalid edit and will not touch the file on disk.

Next: the step-by-step launch order — in Setup Checklist.

← Back to the documentation index