Free sitemap checker and URL extractor
Every sitemap a site declares, not just /sitemap.xml
A site announces its sitemaps in robots.txt, and large sites rarely have one file. This reads what is declared, follows a sitemap index down to its children, and unpacks gzipped files, then lists the URLs it found with the section each belongs to.
- Sitemap indexes: a file of files, which is how any site past 50,000 URLs has to do it.
- Gzipped sitemaps: .xml.gz, read without you downloading anything.
- News, image and video sitemaps: the specialised namespaces, recognised for what they are.
- Sections: URLs grouped by the first part of their path, so the shape of the site is visible at a glance.
What a sitemap check tells you
| What you see | What it usually means |
|---|---|
| A file that 404s or 403s | robots.txt points at a sitemap that is not there — search engines are following that too |
| Far fewer URLs than pages | Whole sections are missing from the sitemap, so they rely on being linked to |
| Far more URLs than pages | Stale entries: pages removed from the site but never from the file |
| A section you did not expect | Usually a tag or filter tree generating thousands of thin URLs |
| No lastmod anywhere | Nothing can crawl you incrementally — every run has to re-read everything |
A sitemap is how you crawl part of a site
Once a site’s sitemap files are known, a crawl can take one section of it — a guidance library, a product tree, a newsroom — and leave the rest alone. That is the difference between reading 400 pages a week and 400,000, and it is why the first thing a project does is read what the site declares about itself.
Frequently asked questions
How do I find a website’s sitemap?
Enter the domain above. It reads robots.txt for the sitemaps the site declares and falls back to /sitemap.xml, then follows any index down to the files underneath.
How do I extract all URLs from a sitemap?
Run the check and copy the list. The free tool returns the first 500 URLs; a crawl on an account reads every one and can export them as CSV or JSONL.
What if a site has no sitemap?
The tool says so, which is itself worth knowing. A crawl can still map the site by following its links — it just cannot tell an orphan page from a linked one without a sitemap to compare against.
How many URLs can one sitemap file hold?
Fifty thousand, and 50 MB uncompressed. Past that a site needs a sitemap index pointing at several files, which is why reading only /sitemap.xml misses most of a large site.
Does it check gzipped and news sitemaps?
Yes — .xml.gz is unpacked, and news, image and video sitemaps are recognised as what they are rather than skipped.
See what a site declares
Then crawl the part of it you actually want.