A screenshot is a photograph: it shows how everything looked at the moment of capture, but you cannot press a button and see what happens. The live page solves that task: the agent gets access to a real open tab in your browser and works with it directly. Below is what this takes and what the trade-offs are.
Why this is needed
When the agent writes an action or repairs a check, it helps to see the page as a whole: expand a menu, look at what is hidden, check how a form behaves. From a single screenshot that is guesswork; from the live page it is working with facts. That is why a screen’s report gets access to the real tab added to it, when that is configured.
What it takes
- The server address reachable by the agent — a separate setting at installation. When it is empty, the live page is not connected, and everything else works as usual.
- An open tab with the panel enabled — access exists only to the tabs where you turned the addon panel on yourself. Closing the tab closes the access.
Hence a simple rule: the fewer tabs you open for the agent, the less it sees. Consent here is given per page — turning the panel on means allowing.
Two ways of access
- The command set on a dedicated address — the program asks the page for what it needs: show the description, a snapshot, the text; click, type a value, navigate. The list of commands with examples is printed right at that address — no separate documentation is needed.
- The MCP server — MCP is a common language of browser commands that modern models understand: a page snapshot, clicks, typing, working with tabs, history, reload. When the agent can speak this language, it works with the page using commands familiar to it, with no special preparation.
The MCP connection is made in the agent’s settings, in the MCP Servers section: you enter the Test Agent server address and the access key. The key is the same one the Chrome addon uses; for external clients a long-lived API key fits too — they are described in the API Keys section.
The visual tools
Besides the browser commands, the MCP server has nine tools for visual checks: run a build, get the list of what awaits a decision, look at a difference, re-compare, propose a decision for a group of failures, start a test healing, publish evidence, get help — and one more, about approval, described below.
The answers are deliberately short: no page markup, no console stream, no pictures inside the reply. When a picture is needed, the tool returns a link instead of embedding the image — this way the model does not spend room on megabytes. And honestly: the tools cannot approve a baseline — none of the nine. The tool named “approve” always answers with a refusal and a link to the review screen: it exists so that the model learns the rule at once instead of trying again.
Tool groups
The tools are split into two groups: the browser commands and the visual checks. At connection, a client can ask for only one of the groups, or for both. What the client did not request, it will not see in the list and cannot call — the attempt gets a refusal with an explanation. When no group is specified, all tools are available, as before.
Working with scripts has not been brought into the MCP tools yet: scripts are created and edited on the “Scripts” tab or over external access — where they can also be read, deleted, trial-run and their run journal browsed (see Scripts & Extensibility and External API Access). So machines do have access to scripts — there just are no MCP tools for them yet.
How it looks from your side
The agent can open a tab of its own, work in it and come back to yours. When you are sitting at the computer at that moment, you will see the tabs switch — the work happens in your usual Chrome, not somewhere on the side. That can be inconvenient, so it is better not to do this in the middle of a busy workday.
An important rule: the agent cannot close your tabs. It can close only the ones it opened itself. Your open pages stay in place, whatever the agent is doing.
Honestly about security
- The access key travels to the model. For the agent to reach the page, the server address and the key get into the text of the request to it. This means trust in the model — and in where it runs — inevitably becomes part of the scheme.
- The screenshot link works without sign-in. That is done on purpose, so the model can fetch the image with a plain request. The flip side: anyone who can reach the server can open a screenshot by its number.
- So keep the server in a closed perimeter — inside your own network or behind protection, not on the open internet. Enable the live page when it is really needed.
When the live page is not needed at all, simply leave the server address unset at installation: then neither the key nor the links go anywhere, and all checks keep working from snapshots.
Next: Getting Started — the route has come full circle, and can be reread with the details in mind.