Scripts & Extensibility

The product covers the typical interface checks out of the box, and everything else can be added by yourself. A script is a plain file with Python code that lives in the project folder: it can become a test step, the capture comparison engine, a difference classifier, or a handler for a finished run. Nothing has to be rebuilt or updated for that — that is the whole point of extensibility. The branch is optional: without scripts everything works as before.

What a script is

A script is a file with Python code. It lives in the project folder beside the rules and actions, and it is stored with the project in git — so it gets the same version history and passes the same review as the rest of the project’s files.

Before every run the product reads the file anew. That is why an edit takes effect immediately, without a server restart: save the file — and the next run already executes the new code. Why this matters: the product stops being a set of ready-made blanks. A new kind of check is born right where the project lives — a check of an API address, a counter in a database, or a needed file on disk.

Where a script lives

Scripts live in the scripts/ folder inside the project. There are three things in it:

  • scripts.yml — the manifest: the script’s title, what it is for, whether it is enabled, its own timeout, the allowed server variables, and its own variables.
  • <id>.py — the code itself. The file name is the script’s name in the manifest.
  • requirements.txt — an optional list of the libraries the script needs.

Everything else a script needs for work that is not its code — the Python environment and the working directories for intermediate files — the product keeps in its own data, not in the project folder.

Where a script can be wired in

The same file plugs into five points of the pipeline’s work. What each one is good for:

  • A test step. A check of an API address, a database, or a file right inside a run — beside the steps that work in the browser. The details are in Tests.
  • A whole test made of scripts. When every step of a test is a script, no browser is needed: the test executes on the server, while the queue, the schedule, and external access work as usual. The runner line of such a run says server — you can see it in Runs.
  • The capture comparison engine. The script compares two captures itself and answers by how many percent they differ: your own algorithm instead of the built-in one.
  • The difference classifier. The script looks at disputed differences before the model takes over, and decides whether it is an important change or accidental noise. It may only raise the importance, never lower it, and it approves nothing.
  • The handler after a run. When a run finishes, the product hands its summary to external code — that is how notifications, reports, and exports are made.

Besides that, any script can be run manually — with any arguments, to check it before wiring. The comparison engine and the classifier are described in Visual Testing.

The Scripts tab

The most convenient place to work with scripts is the separate Scripts tab in the project window. It shows:

  • The list of scripts with “wired into” badges — it is immediately clear which tests and settings reference the script.
  • The code editor with Python highlighting and a “why this script exists” note field. The note helps recall the purpose half a year later.
  • The script’s settings: its own timeout, the allowed server variables, and its own variables.
  • A trial run with arguments as JSON — that is how a script is checked before it gets wired into tests.
  • The run journal — every run of this script on one screen.

At the top the tab warns honestly about whatever is in the way: Python is not found, the environment is still being prepared, or script execution is disabled. These warnings are better not ignored — they are exactly why a script may fail to start, and the reason is visible at once.

What a script gets and returns

The contract is simple. On the way in, the script receives one JSON object: which point called it, what run is going on, which wiring arguments are set, and the point’s own facts — the capture file names, for example. On the way out goes the last JSON object the script printed; lines printed earlier do not interfere.

What exactly to return depends on the point. A test step answers “worked or not” and may return check rows; the comparison engine — the difference percent; the classifier — the difference’s class; the run handler — anything. Here is a short script that checks an API address:

import json, sys, urllib.request

data = json.loads(sys.stdin.read())
url = data.get("args", {}).get("url")

with urllib.request.urlopen(url, timeout=20) as resp:
    ok = resp.status == 200

print(json.dumps({"ok": ok, "output": "the address answers" if ok else "the address does not answer"}))

All the other details of the contract — how the input object is built and what every point expects on the way out — are gathered in External API Access.

Variables and secrets

A script does not see the whole server environment — only what was allowed for it: the list of allowed server variables and its own variables. A value of the form $NAME is substituted from the environment at run time and is stored nowhere. Values under “secret” names — token, secret, password — are read back as ***, so a password cannot be overwritten by accident.

Why it is built this way: project files must never end up holding passwords — they live in variables, exactly as with actions. How to create the variables and credentials themselves is described in Environments, Credentials & Variables.

What is visible about every run

Every run leaves exactly one journal record — a successful one and a failed one alike. The record shows the wiring point, the source (manual, a schedule, external access, a run), the status, the duration, and the tail of the error when there was one. Failures are named in words, not codes: execution disabled, no Python, timed out, the code exited with an error, the script file disappeared, unclear output, too much output, run canceled.

The records are kept for 30 days by default. The journal is the first place to look when a script behaves unexpectedly: it shows whether the run ever started and what the script actually answered.

The runtime and libraries

A private Python environment is created per project, on the first run. The libraries from requirements.txt are installed once and reinstalled only when the file has changed — no needless work happens on every run.

When Python is not found or a library install failed, the product says so directly in the warning on the tab and writes the reason into the journal. There is no silent skip where a script seems to exist but never works.

Circuit breakers

Scripts have limits — they can be changed in the project settings but not removed:

  • One run’s time. 60 seconds by default, 600 seconds at the ceiling. When the time runs out, the whole process tree is killed, children included — so a hung script cannot keep working quietly.
  • Concurrent runs. Two by default; the rest wait for their turn.
  • Output volume. Capped at about a megabyte: print the result, and put bulky data into files in the working directory.
  • Script count and code size. Up to 200 scripts per project and up to 256 KB of code each.
  • Journal retention. 30 days by default.

There are also two switches: the server-wide switch TEST_AGENT_SCRIPTS_ENABLED and the project’s own switch in its settings. A disabled script never goes silently green: the launch attempt lands in the journal with the reason — “execution disabled”.

Honest boundaries

  • There is no sandbox. A script runs as the server’s own user, with its access to files and the network. This is a trusted person’s code stored with the project — so never run strangers’ scripts. The limits above are protection from an accidental error, not isolation.
  • Python only. No other language is supported — a deliberate restriction, not an unfinished feature.
  • No permanently running services. Every run is a separate process that ends together with the work.
  • A script approves nothing. Even the classifier may only raise a difference’s importance. Baseline approval and the verdict on differences stay with the human.

Next: the same scripts over an access key, for machines and continuous-integration pipelines — in External API Access.

← Back to the documentation index