25 August 2026
An XML sitemap search engines actually follow
Huge sitemaps full of 301s teach Google your list is sloppy. Small and clean gets followed more often.
A sitemap says: these URLs matter. If half of them are 301 or 404, the crawler learns it can ignore your list. I have seen 80,000-line indexes with a third of them parameter junk. Nobody follows those faithfully. A plugin that dumped “everything”, thank-you pages included.
Canonical URLs only, status 200, same host as live. lastmod only if you actually touch it, otherwise you lie every day. Split above 50,000. No search results, no cart, no old pagination. And no URL that robots.txt Disallows. That is arguing with yourself.
Put the URL in robots.txt and in Search Console. Only one of the two is how it goes missing. Two sitemaps that overlap a bit is fine. Two that disagree is not. Index now, index later: pick.
A sitemap URL that 404s is not a “watch out”. It is a hard fail. After a content release, check the new pieces are in and the old ones are out. A sitemap does not need more drama than that. Small, clean, same host. That is the whole trick.
I crawl the sitemap URL itself, not only the pages inside it. If that XML 301s to another host, or skips gzip and weighs 8 MB, nobody follows it faithfully. One index of a few hundred clean URLs beats three indexes full of filters. Dull. Also the only version that keeps working.
Want this on your site today?
First report free. Then auto-scan, a PDF in your inbox and alerts when a score drops.
Start a free report