Getting Started

Test Agent is a program for checking your website, app or internal system. You open a page in Chrome, capture the screen and explain in words what is wrong — artificial intelligence analyzes the comment once and turns it into a check that from then on works by itself, without calling the model. The second layer is visual testing: the product compares every build against an approved reference image of the page and shows the differences as a picture, while the verdict is always delivered by a human. Below is what this product is, who it fits and where to start.

The main idea

AI is needed only at the learning stage. You show the page and explain the problem in words — the model writes a rule: a small check that looks for this problem in the browser. From then on the rule works without the model: fast, free and on all similar pages at once. The more rules accumulate, the less AI is left in the product — and the cheaper every check becomes.

Compare that with the usual approach, where AI “looks” at every page on every release: there the model works constantly, and every run costs money. Here the model does its job once per problem, and from then on plain code does its work.

Visual checks follow the same logic. A comparison is almost always resolved without the model: the snapshot is matched by its fingerprint, and if the fingerprint differs — by thresholds and by the areas you excluded. A matching snapshot costs nothing at all, no comparison line appears, and repeats are free. The model joins only on ambiguous cases and never delivers a verdict in place of a human: it can only suggest, while approving or rejecting a change is something a person does.

Who it fits

  • Online stores — the cart, payment, the account: checks catch breakages before they reach customers.
  • Sites and landing pages — request forms, buttons, text readability after every layout edit.
  • Internal systems — the working programs your staff uses every day: checks catch both inconveniences and real errors.
  • Anywhere the interface changes — if edits ship often, manual “eyeball” checking sooner or later stops keeping up.

What the product consists of

Nine words that keep coming up below. All of them are names of sections in the interface itself.

  • Project — one site or app under test. Each project has its own rules, actions and tests.
  • Screen — a snapshot of the page and its description: the address, the title, the structure. Plus comments in words about what is wrong on this page.
  • Rule — a check that looks for one specific problem. Written once, works forever.
  • Action — a recorded scenario of working with the interface: log in, open the payment page, fill in a form.
  • Test — a sequence of steps: open an address, run an action, check the rules.
  • Run — one execution of a test. Every step leaves a snapshot and check results — that is evidence, not just a green checkmark.
  • Build — a run with visual checks: the product captures the test’s steps and compares them against the baseline. A build has a branch, a review state and a failure reason.
  • Baseline — a human-approved reference image of the page that snapshots are compared against. It is kept with a version history, so you can see who approved it and when.
  • Review — a human’s examination of the differences: approve, request changes or reject. Only a human can approve a change — neither the model nor an external program does that.

Around this set work a few more things: the schedule (when tests run by themselves), issues (identical errors merged into one line), auto-fix (the agent repairs a broken check) and the summary on the main screen, which answers the question “what needs me right now”.

The product covers typical interface checks out of the box, and you can write the unusual ones yourself. A script is a file with Python code that lives in the project folder and plugs into the product’s work: that is how an API address, a database or a file gets checked — things that are not visible on the page. The branch is optional — without scripts everything works as before, and the details are collected in the Scripts & Extensibility section.

The visual part has sections of its own: the “Verify” tab with the list of builds, “Baselines” with their versions, the “Review board” with differences waiting for a decision, “Triage” with the reasons builds fail and “Coverage” with an inventory of what is protected by checks.

First sign-in

Installation is described in the Installation section — there it is one command and one settings page. After startup, open the server address in a browser: if no users exist yet, the app offers to register. The first person to register becomes the administrator — other users are added later. The sign-in lasts 90 days, so you will not re-enter the password every day.

Then the administrator adds the remaining users and assigns them roles: Viewer, Reviewer, Member, Admin and Owner. The design approver is marked separately — only that person approves differences in appearance. API keys are issued in the same place: they are needed so that checks can be started from your continuous integration system.

What the window looks like

  • On the left — the list of sections: “Projects”, and for the administrator also “Ops” (runners, queue, quotas). At the bottom — “Prompts” and “Chats”. The “Diagnostics” screen opens from the health warning strip about the state of checks.
  • In the center — “Projects”: cards of your projects and the numbers on top showing what needs attention.
  • Inside a project — tabs: “Screens”, “Rules”, “Actions”, “Scripts”, “Tests”, “Coverage”, “Agents”, “Verify”, “Issues”, “Settings”. The “Scripts” tab is for those who extend the product with their own checks; you do not have to open the other tabs for that.
  • From a project — “Baselines”, the “Review board”, the “Audit trail”, the “AI log” and the visual-testing setup “Checklist”.

First steps

  1. Create a project — give a name and your site’s address. The product immediately creates the project folder with all its settings.
  2. Install the Chrome addon — it captures screens and runs checks right in the browser.
  3. Capture a screen and write a comment — open the page you need, turn on the panel and explain in your own words what is wrong.
  4. Press “Call agent” — a ready report about the screen is placed into the chat. Add your request and send it: the model returns a rule, new knowledge about your interface and a suggested screen type.
  5. Verify the rule — run it on the page and press “Verify”. Verification means “the check runs without errors”, not “the page is fine”: the verdict about the page itself comes later, during runs.
  6. Assemble a test — a sequence of steps from the actions and rules that belong together.
  7. Mark capture on steps — note in the test which steps need capturing. The snapshot is taken by the Chrome addon.
  8. Run the first build — it will create baselines for all the snapshots by itself: there is nothing to compare against yet.
  9. Approve the baselines — look at the first snapshots on the review board and approve them. From that moment, a difference from these snapshots counts as a change.
  10. Run the second build and look at the verdict — if the page has not changed, the comparison comes out “no changes”; if it has changed, you will see the differences and decide whether to update the baseline or fix the page.
  11. Set a schedule — for example, every morning at 9:00. The test will run by itself, and you will see the result in the summary.

The step-by-step checklist — eight steps — opens by itself for a new project, and you can always come back to it later: Setup Checklist.

What you will need

The server the product runs on, and a Chrome browser with the addon installed — it is called the executor: it is the one that executes checks and runs tests. Everything lives on your side: neither screenshots, nor rules, nor tests are sent anywhere, apart from the calls to the AI agent that you start yourself. For visual builds Chrome must be connected all the time: only the addon takes snapshots — the product has no separate browser on the server and no cloud re-rendering of pages.

Two limitations, honestly, from the start. The addon is private: it is not in the Chrome Web Store and is installed manually from a folder the product prepares itself. And the schedule fires only when the executor computer is on — a missed moment is not caught up.

Next: how projects are arranged and what lives in their folders — in the Projects & Settings section.

← Back to the documentation index