
Key takeaways
- JetOctopus's case study reports a jewelry online shop with over 16 million pages but only 342,000 monthly visits, an extreme mismatch.
- Profitable pages on that site were not indexed for months while Googlebot spent its resources on canonical-tagged filter and pagination URLs.
- The recommendation was to close category, filter, pagination and sort combinations, which cut the site to 730,000 pages.
- The same article shows a real estate site where 23.5 million of 38 million monthly Googlebot requests went to AJAX content.
- The client is unnamed and JetOctopus sells the log analyzer used, so treat this as the vendor's own account, not an audited result.
Server logs record what Googlebot actually requested, not what a crawler simulation guesses. JetOctopus, a log and crawl analysis vendor, published three short case studies from that data in one article. This piece covers the one with the largest measured change, and notes what the article does not say.
What did the jewelry site’s logs show?
The site had more than 16 million pages and 342,000 monthly visits. One part of it ranked at the top of the results, while another part was not indexed for months, and those were relevant, profitable pages rather than outdated ones. JetOctopus crawled 2 million pages, found almost 900 internal links on each, and estimated around 16 million pages in total. It then analyzed the logs and found Googlebot spending resources on pages carrying rel=canonical tags, because the site had placed a canonical or self-canonical on every item in every category, filter and paginated page.
What was changed and what was the result?
The recommendation was to close the combinations of category, filter, pagination and sort that together generated the bulk of those pages. The source’s stated result is that the site’s size decreased to 730,000 pages, so the bot could crawl the most profitable pages. The article gives no crawl-request counts, indexed-page counts or traffic figures for the site after the change, and no timeframe. We therefore report only the page-count change.
Does the article say robots.txt was the fix?
Not for this site. The recommendation says “close” the combinations without naming the mechanism. The article’s takeaway does discuss trade-offs: Google recommends noindex over blocking, but on very large e-commerce sites the bot still spends crawl budget fetching a page to read that directive, whereas robots.txt can stop crawling of some sections entirely.
What do the other two cases in the article show?
On a marketplace with 226,000 pages, filter URLs were disallowed with a Disallow: /*? rule, yet duplicates stayed in Google’s index. The cause was a conflicting Allow: /product rule, and Google follows the least restrictive matching rule. On a real estate site with over 1 million pages, logs showed 38 million Googlebot requests in a month, of which 23.5 million went to AJAX content, and the recommended fix was blocking valueless AJAX in robots.txt. Neither case reports an after figure.
The evidence
- Subject
- Jewelry e-commerce site crawl waste found in server logs
- Timeframe
- Not stated by the source
- Verified
- Yes, against first-party data
| Metric | Before | After |
|---|---|---|
| Crawlable pages on the site | 16000000pages | 730000pages |
Sources
- 1.How Googlebot Crawls Your Pages. Logs Insights - JetOctopus, May 27, 2020Primary
About the author

Founder
Founder, UpgradIQ, Inc.
Adam Hafez works on technical SEO and search measurement: how pages get crawled, indexed, ranked and now quoted by answer engines. He founded UpgradIQ, which reads Google Search Console and GA4 to tie ranking movement back to the changes that caused it. He publishes what the data supports and states the limits of it.
- Technical SEO
- Search Console and GA4 measurement
- Answer engine optimization
- Structured data
Related reading
Canonical tag adoption, 2025: 68% of desktop pages set one
The 2025 Web Almanac SEO chapter puts rel=canonical on 68% of desktop and 67% of mobile pages, up from 65%, with mismatches under 1%.
How accurate are sitemap lastmod values? Bing's study
Bing's own study of sitemaps found 18% of lastmod values set incorrectly, usually to the generation date, and Google says it only trusts accurate ones.
Hreflang XML sitemap generator
Build a sitemap with xhtml:link hreflang annotations for localized pages, with every alternate listed on every URL as Google's documentation requires.
HTTP status code crawl handling checker
Paste status codes and see how Google's crawlers document handling each one, plus a redirect hop count against Google's 10-hop default.



