Canonical tag adoption, 2025: 68% of desktop pages set one
The 2025 Web Almanac SEO chapter puts rel=canonical on 68% of desktop and 67% of mobile pages, up from 65%, with mismatches under 1%.
Indexing is the stage between crawling and ranking where a search engine decides whether to store a page and under which canonical URL. A page can be crawled successfully and still never be indexed, usually because it duplicates another URL or is judged too thin to be worth storing.
The 2025 Web Almanac SEO chapter puts rel=canonical on 68% of desktop and 67% of mobile pages, up from 65%, with mismatches under 1%.
Bing's own study of sitemaps found 18% of lastmod values set incorrectly, usually to the generation date, and Google says it only trusts accurate ones.
A JetOctopus case study: log files showed Googlebot spending crawl on filter, sort and pagination URLs, and closing them shrank the site to 730K pages.
Build a sitemap with xhtml:link hreflang annotations for localized pages, with every alternate listed on every URL as Google's documentation requires.
Paste status codes and see how Google's crawlers document handling each one, plus a redirect hop count against Google's 10-hop default.
Validate robots meta tags and X-Robots-Tag headers against the rules Google documents, and see the combined result for Googlebot.
Normalize a list of URLs by RFC 3986 rules and your own policy, then see which URLs collapse into the same page and need one canonical.
We read the robots.txt of 60 publishers and SEO vendors. News sites block AI crawlers almost universally. The SEO industry does not.
Webmasters say the indexed pages report for their news sitemap has not moved since August 28, even though new articles are indexed and visible.
Build correct <link rel="alternate" hreflang> tags for a set of localized URLs and check them against Google's reciprocity and x-default rules.
Tracing Google's mobile-first indexing announcements from 2016 to the 2023 completion notice, and its documentation on mobile-desktop content parity.