Sitemaps and robots.txt

Two plain files at the site root complete the picture. /sitemap.xml lists the canonical URLs you want crawled, each with an optional <lastmod> date; one file may hold at most 50,000 URLs or 50 MB uncompressed, and Google ignores <priority> and <changefreq>. /robots.txt, standardized as RFC 9309 in September 2022, tells crawlers which paths not to fetch (User-agent: * then Disallow: /cart/) and where the sitemap lives (Sitemap: https://example.com/sitemap.xml). It controls crawling, not indexing: a disallowed URL can still appear in results, so use noindex (Meta Tags That Matter) to hide a page. Sitemap and Sitemap and robots.txt Tools build and test both files, and llms.txt and AI Crawlers covers rules for AI crawlers.