Sitemaps

Sitemaps in Crawling and Discovery

A sitemap lists a site's URLs for crawlers: url entries with loc and optional lastmod in the namespace http://www.sitemaps.org/schemas/sitemap/0.9 5,294 , at most 50,000 URLs and 50 MB per file. Teams generate them from catalogs and read other sites' sitemaps to plan crawls. The protocol's XSD checks this one:

sitemap.xsl: a sitemap from the catalogXML
<xsl:stylesheet version="3.0" xmlns:xsl="http://www.w3.org/1999/XSL/Transform"
    xmlns="http://www.sitemaps.org/schemas/sitemap/0.9" expand-text="yes">
  <xsl:output indent="yes"/>
  <xsl:template match="/catalog">
    <urlset>
      <xsl:for-each select="book[supply/availability = 'in-stock']">
        <url>
          <loc>https://booknest.example.com/books/{identifier}</loc>
          <lastmod>{substring(../header/sent, 1, 10)}</lastmod>
        </url>
      </xsl:for-each>
    </urlset>
  </xsl:template>
</xsl:stylesheet>
Generating and validating the sitemapShell
curl -fsS -o sitemap.xsd https://www.sitemaps.org/schemas/sitemap/0.9/sitemap.xsd
xslt3 -s:booknest-catalog.xml -xsl:sitemap.xsl -o:sitemap.xml
sed -n '3,5p' sitemap.xml
xmllint --noout --schema sitemap.xsd sitemap.xml
Output
   <url>
      <loc>https://booknest.example.com/books/BN-0001</loc>
      <lastmod>2026-10-01</lastmod>
sitemap.xml validates

Queries on the result need the sitemap namespace (Default Namespaces and Scoping), or //url matches nothing.