App Map & Crawling

The app map shows which addresses your site consists of: which sections exist, what lives at each address, and what covers it. The map is filled by a browser with the installed extension — there is no other way, because the addresses must actually be opened and looked at to see what lives on them.

This is a view from the other side than the coverage inventory. There the story is about buttons and forms, and here about pages: you can see which sections of the application nobody tests at all.

The route map

The map is a tree of addresses. Every address has a card: what is on the page — the components noticed there and the request endpoints the page calls — and which tests cover this address. All of it is collected along the way; nothing needs to be described by hand.

Similar addresses collapse into one route: /users/123 and /users/456 are the same route, because the number in the address is replaced with {id}. Query parameters are dropped at the same time: ?page=2 and ?page=3 are also one address. Otherwise the map would bloat with every single user or list page, and nothing could be made out in it.

Crawling

A crawl is when the browser walks the site by itself and records what it sees. A crawl has settings:

  • Start URLs — where to begin; when left empty, the project base URL is used.
  • Max depth — how many steps away from the start URLs to go.
  • Max nodes — a safety fuse so the crawl cannot run away into infinity.
  • Same origin only — do not leave the project’s boundaries.
  • Respect robots.txt — honor the restrictions the site itself declares for automatic crawlers.
  • Allow production hosts — off by default.

The crawl opens its own tab and closes it when finished. It does not touch the tabs you are working in and does not switch them — you can keep working calmly.

A pause instead of an error

If no paired browser is connected right now, the crawl does not fall over and does not write an error. It honestly shows “Waiting for a paired browser” and continues from the same place as soon as the browser appears. This matters: people close their browsers for the night and for the weekend, and the crawl must not be considered broken because of that.

A crawl can also be stopped manually — and then a reason is mandatory. If the crawl stopped at a login page, the product says so plainly: addresses behind the login may be unreachable.

Production protection

The crawl does not touch the production site without explicit permission. The restrictions are checked before navigating to a page, not after: both robots.txt and the production host list cut an address off before the browser opens it. A skipped address does not disappear silently — the reason is visible in the crawl notes.

Hence an honest caveat: the crawl cannot accidentally harm a production site, but it will not see anything useful there either until you allow it yourself. Respect robots.txt is also on by default.

Coverage over the map

The map can be colored by actual coverage: you see where the tests really go, what is touched partially, and what is not touched at all. It is the same view as in the inventory, but from the address side: not “which rows are closed” but “which pages the tests walk”. This is a convenient way to find whole sections that nobody tests.

Unimportant routes

A route can be marked unimportant — for example, a utility page or a technical address. A reason is mandatory. An important consequence: no tests are created for unimportant routes — so automation does not breed checks where you do not want them. If the decision changes, one action returns the route to the map.

The map also shows who answers for a route — an owner can be assigned or changed. And one more useful detail: the map has a screenshot gallery, so what is in front of you is not only addresses but also how those pages look.

Next: how the agent walks the path from scouting the site to ready checks — in Agent Campaigns.

← Back to the documentation index