Why a clean sitemap matters
A sitemap is a promise to search engines: these are the pages I want in your index, at the address where each one lives. Every entry that breaks that promise, whether it redirects, returns a 404, carries a noindex tag or names a different canonical, costs a crawl and teaches Google that the file is not worth trusting. Once that happens, the sitemap stops doing its one job, which is helping your new and updated pages get found quickly.
Most sitemaps drift without anyone noticing, because plugins and frameworks generate them automatically and nobody reads the output. Pages get redirected, old posts get noindexed, and the sitemap keeps listing them anyway. This checker reads the file the way a search engine does, so you can see that drift and fix it.
What it checks
It starts from your robots.txt, reads every sitemap named there (falling back to the usual addresses such as /sitemap.xml when none is named), and follows sitemap indexes down to every child file, including gzip files. For each sitemap file it checks that:
- the file loads with a 200 status at its own address, without a redirect, and is served as XML;
- it is a real
<urlset>or<sitemapindex>within the protocol's limits of 50,000 URLs and 50 MB; - an index lists only sitemaps, never another index, and robots.txt names the sitemap so every search engine can find it.
For every URL listed, it checks that the URL appears only once, sits on your own host, and carries a lastmod in the right format that is neither in the future nor identical to almost every other date in the file, since that usually means the date records when the sitemap was built rather than when the page changed.
Then it visits the URLs themselves, up to the limit you choose, taking them from every sitemap in turn so each section of your site is sampled. Each one must be allowed by robots.txt, answer 200 without redirecting, carry no noindex in its robots meta tag or X-Robots-Tag header, and name itself as its canonical. Finally it works the other way round, checking the pages your home page links to and listing any that load, can be indexed and still appear in no sitemap.
What you get
The report is a set of fix cards, most urgent first, and each card says what to change in words you can hand to a developer: remove these URLs from this file, replace these with the address they redirect to, or fix the canonical on these pages. Every card lists the URLs it covers, and the whole report downloads as a CSV with one row per URL, ready to paste into a ticket or a spreadsheet.
How the crawl works
Pages are fetched by EghosaBot, which obeys robots.txt and its Crawl-delay setting, so the check never fetches anything your site has asked crawlers to leave alone. It reads only the head of each page, where the noindex and canonical tags live, which keeps the check light on your server. A run reads up to 50 sitemap files and 20,000 listed URLs, and visits up to 500 of them, which keeps each check quick and light on your server.
Your report and your data
The check is free. It asks for your email address before it runs, and you only join the mailing list if you tick the box. Each report has its own private link, which works for 90 days and is then deleted along with everything the check gathered. The privacy policy explains the rest.