Site health crawls
What the bounded daily crawl checks, how ownership is verified, and the limits that protect your site.
Updated
What it is
Site health is a bounded crawl of a site you own. Once a day, the worker fetches your pages the way a cautious crawler would and reports problems that uptime checks cannot see: broken internal links, server errors on deep pages, redirect chains, http resources on https pages, pages that unexpectedly carry noindex, missing titles, and pages that returned 200 on the previous crawl but fail now. It is not a ranking tracker and it does not use third-party data.
Verifying ownership
Crawls start only after you prove control of the origin. Add a site under Site health and complete either option:
- DNS: create a TXT record at _uptimeguard.<host> with the value shown.
- File: serve the shown token at https://<host>/.well-known/uptimeguard-verification.txt.
Then press "Check verification". DNS lookups use a public resolver; the file is fetched through the same destination policy as monitors, so private addresses are rejected.
Bounds
- Same origin only. Links to other hosts, including subdomains, are recorded as findings but never followed.
- robots.txt is honoured for the user agent UptimeMonitor360-SiteHealth. A root Disallow produces a critical finding and stops the crawl.
- One scheduled crawl per site per 24 hours. "Crawl now" is allowed once per hour.
- Page budget per plan: Starter 50, Business 100, Agency 150, Scale 250 pages per crawl. Discovered links beyond the budget are counted, not fetched.
- 256 KiB of HTML per page, 10 seconds per page, 10 minutes per crawl, 250 ms pause between requests, at most 5 redirects.
- Findings are capped at 500 per crawl and history is kept for 90 days.
Findings
Severities are critical (server errors, unreachable home page, regressions, robots blocking the root), warning (4xx pages, unexpected noindex, mixed content, slow first byte, protocol downgrades, missing titles) and info (missing descriptions, duplicate titles, redirect chains, missing sitemap, missing h1, viewport or lang). Mark intentional noindex paths under Expectations so they stop appearing.
Limits of the method
Pages are parsed with a lightweight HTML scanner, not a browser. Content injected by JavaScript is not seen, and link discovery follows <a href> only. Treat findings as hints to investigate, not as a complete audit.