Skip to content
SEO Madmanby Adam Hafez

NavBoost: what the DOJ trial and the API leak actually confirmed

Court testimony and a leaked Google document, read together, describe a click-based re-ranking system Google spent years declining to confirm in public.

Published: · Read time: 3 minutes

Written byAdam Hafez
Share
Plain Markdown
A statue of Lady Justice holding scales atop a stone courthouse building

Key takeaways

  • Google VP of Search Pandu Nayak testified under oath that NavBoost is "one of the important signals" used to re-rank search results.
  • Nayak described NavBoost as memorizing clicks over a rolling window, reported by Search Engine Land as 13 months and 18 months before 2017.
  • A leaked internal document, independently obtained in May 2024, names click categories including goodClicks, badClicks and lastLongestClicks.
  • The leak's own first reporter cautioned that a named feature is not proof it currently affects ranking or carries meaningful weight.
  • For years Google's public statements treated click data as unconfirmed for ranking, a position the trial testimony sits in tension with.

Two disclosures in 2024 landed on the same subject from opposite directions. In November 2023, Google VP of Search Pandu Nayak took the stand in United States v. Google and described a system called NavBoost. Six months later, an internal Google document reached SparkToro’s Rand Fishkin from an anonymous source and named click categories that line up with what Nayak described under oath. Neither event was staged to corroborate the other. Read together, they are the closest the public has to independent confirmation of how click data moves through Google’s ranking pipeline.

What Nayak testified to

Google’s VP of Search told the court that NavBoost is “one of the important signals” Google uses, not the retrieval system itself. Search Engine Land’s trial coverage quotes him describing it as a memorization system that works from clicks on a query going back roughly 13 months, a window that was 18 months before a change in 2017.

Its job in the pipeline is culling, not fine-tuning. Nayak’s account, as reported, places NavBoost early: an initial retrieval stage returns a large candidate set of documents, and NavBoost uses historical click behavior to cut that set down before later ranking stages run. That is a different claim than “clicks nudge the final order,” and it is worth keeping the two apart when the testimony gets summarized.

It cannot help pages with no click history. By Nayak’s own description, a new or rarely-clicked page gets nothing from NavBoost, because the system has no aggregated behavior to memorize for it. That limitation was part of the testimony, not an inference added afterward.

What the leaked document named

The Content Warehouse API leak is a separate chain of custody entirely. An anonymous source sent internal documentation to Erfan Azimi, who passed it to Rand Fishkin at SparkToro, who published his findings in May 2024 alongside analysis from iPullRank’s Mike King. Search Engine Land’s coverage of the same leak confirms the same field names independently: badClicks, goodClicks, lastLongestClicks and unsquashedClicks appear in the documentation as distinct click-related attributes.

The document names features; it does not state their weight. Both outlets are explicit on this point. Search Engine Land reports that the leak “did not specify how any of the ranking features are weighted, just that they exist.” Fishkin went further in his own post, warning readers directly not to treat a named API field as proof that Google currently uses it the way the name implies - a field can be retired, reserved for testing, or never shipped to production.

Treat the exact terms as leaked shorthand, not a public spec. goodClicks, badClicks and lastLongestClicks read as plainly as the terms in this piece describe them, but they come from internal engineering documentation that Google has never authenticated in public. That caution applies to the names themselves, not just to their weighting.

The tension with Google’s own public position

For roughly two decades, Google’s public statements treated a direct ranking role for click data as something between unconfirmed and denied. A 2019 Google response reported by Search Engine Land said interactions are used “for personalization, evaluation purposes and training data,” a sentence built to avoid the word ranking rather than address it directly.

This piece states the tension; it does not resolve it. Sworn testimony describing NavBoost as an important signal sits uneasily next to years of public phrasing that stopped short of confirming any such role. That is not the same as proof that every earlier public statement was false: Google’s representatives may have been describing a narrower system, working from incomplete internal knowledge, or using “ranking signal” more narrowly than NavBoost’s actual function. What can be said without overreaching is that the public framing and the sworn testimony do not match, and neither this trial nor this leak is the last word on exactly how today’s ranking pipeline weighs a click.

Why we care

If a page has real search demand but no click history, in Nayak’s own account of NavBoost, this particular signal has nothing to offer it - which puts the weight back on other ranking inputs and on technical SEO fundamentals that do not depend on accumulated behavior data. Neither the testimony nor the leak is grounds for chasing a specific click-pattern tactic; both sources here explicitly warn against reading more precision into named fields or a single number than the disclosure supports. Read the primary coverage linked above before repeating a specific figure, and hold the 13-month window and the exact leak terminology as reported claims, not settled specification.

The evidence

Sample
2 disclosures, checked across 3 independent outlets

Hypothesis: Public court testimony and a leaked internal document, read together and checked against each other, describe click-based re-ranking behavior that Google's public statements had for years declined to confirm.

Method: Cross-referencing DOJ v. Google trial testimony from Pandu Nayak against independent reporting on the May 2024 Content Warehouse API leak, treating the two as separate disclosure events and looking for corroboration between them rather than relying on either alone. Each specific claim below was checked against at least two independent outlets before inclusion; where sources diverged or a term appeared in only one account, that is stated rather than smoothed over.

Findings

  • Pandu Nayak testified that NavBoost is one of the important ranking signals Google uses, not the sole retrieval system.
  • Search Engine Land's trial coverage quotes Nayak describing NavBoost as memorizing clicks on queries from the past 13 months, down from an 18-month window before 2017.
  • Nayak described NavBoost's role as culling a large candidate set of documents down to a much smaller set before final ranking, rather than making a final adjustment at the end.
  • The May 2024 Content Warehouse API leak, first reported by SparkToro after an anonymous source shared it with Rand Fishkin, named click-related features including goodClicks, badClicks and lastLongestClicks.
  • Google's own public statements on click data as a ranking input were vague for years; a 2019 Google response to Search Engine Land said interactions are used "for personalization, evaluation purposes and training data" without confirming or denying a ranking role.

Limitations: A leaked internal document cannot be verified against Google's own confirmation; Google has never confirmed the authenticity of specific field names in the leak, only that some of the documentation appears real. Field and variable names in leaked source documentation may be internal shorthand, deprecated, or used only in testing, not a description of what ranking does today - the leak's own first reporter said as much. Trial testimony is elicited in an adversarial proceeding and emphasized by the party asking the questions; a plaintiff's antitrust case has an interest in NavBoost sounding central to ranking. This piece describes what was documented and testified to, not a settled or current account of exactly how Google ranks results.

Sources

  1. 1.How Google Search and ranking works, according to Google's Pandu Nayak - Search Engine Land, December 5, 2023Primary
  2. 2.An anonymous source shared thousands of leaked Google Search API documents with me; everyone in SEO should see them - SparkToro, May 28, 2024Primary
  3. 3.HUGE Google Search document leak reveals inner workings of ranking algorithm - Search Engine Land, May 28, 2024
  4. 4.Google's CTR answer just what you'd expect, and this is why SEOs go bananas - Search Engine Land, February 28, 2019

Frequently asked questions

Does this prove Google uses click-through rate to rank pages today?

It proves Google built a system, NavBoost, that a VP of Search testified under oath uses aggregated click history to re-rank results, and that internal documentation names specific click categories. It does not prove the exact weight of any single click category today, and the leak's own first reporter warned against treating a named field as proof of its current use.

Why does the click-window figure vary between 13 and 18 months in different write-ups?

Search Engine Land's trial coverage reports Nayak's testimony as a 13-month rolling window, down from 18 months before a 2017 change. Some secondary write-ups repeat only one of the two numbers, which is why this piece states both and attributes them to the testimony rather than treating either as a permanent constant.

Is the leaked document itself public?

The underlying Content API Warehouse documentation was shared with and analyzed by SparkToro and iPullRank, who published excerpts and analysis rather than the raw internal codebase. Google has not published or authenticated the document itself.

About the author

Adam Hafez
Adam Hafez

Founder

Founder, UpgradIQ, Inc.

Adam Hafez works on technical SEO and search measurement: how pages get crawled, indexed, ranked and now quoted by answer engines. He founded UpgradIQ, which reads Google Search Console and GA4 to tie ranking movement back to the changes that caused it. He publishes what the data supports and states the limits of it.

  • Technical SEO
  • Search Console and GA4 measurement
  • Answer engine optimization
  • Structured data

The briefing

One email when something in search actually changes. No digest padding.

Subscribe