The owner pays for difference and failure analysis, so it must be visible what exactly is being paid for. The AI log shows three things: which tasks the AI completed, what they cost, and why some task never started at all. Along the way it answers the main question — “are we really not paying on every run?”.
No, we are not. Repeated runs compare pictures by fingerprint and pixel by pixel, deterministically and for free. The agent is asked only when something new appeared or broke. That is why an empty log is a good sign: it means the test suite is stable and there were no extra calls — “AI has never been called — that is the target state for a stable suite”.
The Task Table
Every AI request is one row. It shows the task kind, its status, the target (which snapshot, build, or test), the time and latency, the tokens spent — input and output separately — and the cost. For tasks where the AI judged importance, its confidence in the answer stands beside it.
The task kinds in the log are:
- Difference classification — understand what this is: an error, noise, or an intentional change.
- Build triage and build narrative — gather failures into clusters and explain them in words.
- Failure-cluster refinement — clarify the cause inside an already-gathered cluster.
- Coverage generation — propose a test for an uncovered spot.
- Test healing — prepare a change to the test file.
- Comment analysis, action suggestion, rule or action fixes — the work on screens and comments.
A row can be filtered by kind and expanded: you see what the agent read and what it answered. Input bodies are visible to admins only — your pages’ text can get into them. The full request lives in the chat with the agent; a service copy is kept here.
Totals and Budgets
Above the table is the summary: total spend, a breakdown by task kind, average latency, and the share of builds where the AI was called at all. Another useful number is the spend per found defect — “Tokens per issue”: only defects a human confirmed from failure clusters are counted. This shows the real price of the value, not just the token bill.
The budget has a ladder of states, and it honestly names its position:
- Under half spent — everything is fine.
- Half to 80% spent — a warning.
- 80 to 100% spent — critical; time to look at the spend.
- Limit exceeded — new AI tasks are refused until the limits change.
If no limit is set at all, the product says so — there is no limit, rather than substituting zero or infinity. The limits are edited right here: the token ceiling per task, heal attempts per test and per build, the number of concurrent campaigns, and the repeating-error breaker. These are the same circuit breakers that stop test healing — and just the same, they approve nothing, they only limit spending.
Model Tools
The work of the external tools the model uses is shown on a separate line: how many calls there were, how many tool descriptions were sent to it, and how many answers it received. This is an overhead that is usually invisible: tool descriptions also take space in the input and cost money. If they start to outweigh the task itself, you notice it here, not in the general bill.
Why a Task Was Refused
It happens that a task never started. This is not an error but a tripped circuit breaker, and the reason is recorded in the log. The reasons are:
- The task’s token ceiling is exhausted.
- This test’s heal budget is exhausted since its last successful run.
- This build’s heal budget is exhausted.
- Another heal campaign is already running in the project.
- The error repeats — a previous heal already tried to fix it.
- Visual AI is disabled for this project.
Every refusal stays in the log with its reason — it does not vanish silently. From it you can see whether the limits should be raised, or whether the agent is genuinely not needed.
Next: how to see which places of the application are not protected by tests — in Coverage.