Agent Campaigns

A campaign is turnkey work: you tell the agent what must be covered, and it walks the whole path by itself — from scouting the site to ready checks. A person steps in not at every step but only where a decision is needed: approve the plan, answer a question, accept a test.

Why this is needed: otherwise you would have to decide manually which checks to create, write them, and keep their order in order. The campaign takes that routine on itself but leaves the big decisions to the human — deliberately.

Phases

A campaign moves through phases, and the interface shows which one it is in:

  1. Preflight — preparation: it is checked that the campaign can be started at all.
  2. Scout — the agent finds out what the application has: it walks the site through the paired browser.
  3. Plan — the agent decides which checks are needed to close the coverage gaps.
  4. Plan review — the campaign stops and waits for a human.
  5. Generate — tests are prepared by the approved plan, each as a proposal.
  6. Run — the tests are executed.
  7. Heal — if something broke, a fix is proposed.
  8. Summary — a wrap-up: what is done, what remains, and what awaits a decision.

Along the way the campaign can be paused — the browser is then released — or stopped completely; a stop requires a reason. Only one campaign runs in a project at a time: first deal with the live one instead of starting another.

The stop at plan review

The campaign’s main seam is the plan review. When the plan is ready, the campaign does not move on by itself: it holds the browser and waits until a person approves the plan, sends it back for another pass with a note, or rejects it. A rejection requires an explanation; a rework note lands in the next pass.

While the campaign waits, the browser window is occupied — that is by design: the person sees the same thing the agent sees and can check its conclusions. A pause, on the contrary, releases the browser. The campaign does not pass this point without a human — a protection against the agent inventing what and how to test and then approving it itself.

A generated test is never approved on its own

Every generated test lands not in the working set but in review — as a proposal. The test file is written only after a person accepts it; until then nothing in the project changes. The human decision is recorded in the audit log: you can see who accepted the test and when.

The campaign itself cannot accept a test for a person, and neither can external tools: acceptance is available to humans only.

Campaign budgets

The expensive part of a campaign is the AI calls, so it has a spending ceiling. The ceiling is checked at phase boundaries — before planning, generation and healing — not in the middle of the work. When the limit is exhausted, the campaign stops at the boundary and asks a question: raise the ceiling and continue, or close the campaign.

Other limits are set in the project: heal attempts per test and per build, and how many heal campaigns run at once. They all work the same way: they cap the spend but approve nothing — neither a plan nor a test.

Open questions

Sometimes the agent lacks information: it needs access or a decision about a disputed place. It then asks a question right in the campaign and waits for an answer. Such questions do not get lost: they are visible in the “what needs attention” summary and in the list of questions waiting for a person; overdue ones are marked.

An answer returns the campaign to work, and the answer history is kept — it shows which decisions have already been made.

The leaderboard

There is an optional participant leaderboard. It is off by default and turned on by a separate switch in the project settings. When on, it counts only what brought value: confirmed defects and approved changes. Tokens, test count and activity never enter the rating — otherwise it would pay to chase volume instead of finding real problems.

Only humans earn points: they are the ones who confirm defects and deliver verdicts. A season is a rolling 90 days, so old merits gradually leave the count.

Next: how to run the same scenarios on different stands without keeping passwords in project files — in Environments, Credentials & Variables.

← Back to the documentation index