Search

Are Google's 'Leaked' AI Ranking Signals Testable?

Audit the viral '7 leaked Google ranking signals': what the public API documents, what console labels do not prove, and why consumer AI Mode remains unverified.

Mohit7 min read
A viral '7 leaked ranking signals' claim passes through a receipt gate and splits into a documented, testable API field set and an unverifiable consumer 'AI Mode' claim.

Verdict. The viral “7 leaked Google AI ranking signals” claim is not a leak. It combines a documented enterprise search API, console labels, and one name collision, then drops the original hedge. Three claims are documented and testable. Two named signals do not exist in Google’s public API. The claim that consumer AI Mode uses these signals is unverifiable from Google’s public material. Test the fields, keep the pipeline description, and drop the “leak” label.

The operator and the claim

You run a site that you want AI search to cite, or you need to decide whether an “AI SEO” claim is worth a sprint. A post with ~167K views says Google “leaked” seven AI Mode ranking signals: Base Ranking, Gecko, Jetstream, BM25, a three-tier PCTR system, Freshness, and Boost/Bury. It also names a “500-token chunking” constant and three structured-data flags.

The decision is which claims have receipts and which you can test. That is what this Workflow Decision Lab checks. The post is wrong in specific, operationally relevant ways.

How a hedge became a leak

The viral post is secondhand. Its primary source is a researcher’s November 2025 writeup on Google’s Discovery Engine (a.k.a. Vertex AI Search, now “Agent Search”), an enterprise search product with public docs and a public API. The author set a clear scope:

“I’m not saying AI Mode uses identical code. … Is this exactly how AI Mode works? I can’t say for certain.”

The viral post removed that hedge and used “leak” instead. That changes a documented, configurable enterprise product into a supposed secret ranking stack. The later claims inherit that scope error.

What Google’s public API exposes

Google’s Discovery Engine returns a rankSignals object for each result. The public API reference lists these fields:

rankSignals field What it measures
default_rank Base rank from the standard algorithm
semantic_similarity_score Query/document embedding similarity
relevance_score “Deep-relevance model” query-document interaction
keyword_similarity_score Keyword matching, documented as BM25
pctr_rank Predicted click-through / conversion
topicality_rank Keyword-similarity adjustment
document_age Age of the document in hours
boosting_factor Combined custom boosts you applied

That is eight fields plus custom_signals. The public schema has no gecko or jetstream field. Two names in the viral seven are absent from the product schema.

The field set is real, documented, and testable. Its viral name is wrong.

Where the leak claim fails

Gecko is a console label, not a signal

The source author notes a deprecated embedding model named Gecko, while Cloud Console labels one column “Gecko score.” That UI label sits over semantic_similarity_score. Calling Gecko a ranking signal mistakes a dashboard caption for an API field.

Jetstream is a name collision

Google’s JetStream is a throughput-optimized TPU LLM-inference engine (an Apache-2.0 repo being archived in February 2026). It is unrelated to relevance ranking. The viral post calls Jetstream “a cross-attention model that handles negation better than embeddings.” The public API calls its deep-relevance signal relevance_score, a model that “handles complex query-document interactions.” Same claimed job. Different, correct name.

500 tokens is a default you can change

Agent Search’s chunkSize setting “defaults to 500” with “supported values: 100–500.” includeAncestorHeadings defaults to false. These are ingestion settings in your data store. If you can change a “constant” in your console, it is documented configuration, not leaked internal data.

Testability ledger

The useful artifact is a ledger that holds each claim against a primary source and a test.

Claim (as it went viral) Primary source says Testable?
“7 named ranking signals” Public rankSignals: 8 fields + custom Yes, read the API
“Gecko = embedding signal” No gecko field; deprecated model, console label Yes, grep the docs
“Jetstream = relevance model” JetStream = TPU inference engine, unrelated Yes, look it up
“500-token chunk constant” chunkSize default 500, range 100–500 Yes, configure it
“BM25 still ranks” keyword_similarity_score = BM25 Yes, documented
“PCTR 3-tier, 100K unlock” Only pctr_rank is public; tiers are a console note No, no public API
“AI Mode runs this pipeline” Enterprise product docs, not consumer No, not observable
“Serve stage uses Gemini 2.5 Flash” No public model disclosure No, unverified

Rows 1–5 are documented and can be checked against the public API in an afternoon. Rows 6–8 are where the “leak” framing earns clicks, but they lack a receipt. The three structured-data flags (“searchable / indexable / retrievable”) sit between those groups: enterprise search schemas do support field-level control, but the specific three-flag consumer mapping remains unverified until the schema reference confirms it.

Reusable test protocol

Use this five-step check for any named-signal claim from your team, a client, or a viral post:

  1. Find the primary source. Record the publisher, credential, and any hedge. A secondhand post dropping a hedge is a red flag.
  2. Map every named signal to a public API field. A missing name points to a console label, a name collision, or fabrication.
  3. Check product scope. Enterprise documentation does not prove consumer behavior. Agent Search does not prove AI Mode.
  4. Write one test and one falsifier per claim. “Testable” means you can name both the query and the result that would prove the claim wrong.
  5. Publish the ledger, not the leak. The primary source supports the ledger; it does not support inflated framing.

You do not need Google’s internal system. You need the public API, the primary source, and a table. That is the same discipline behind a local LLM inference field guide built from one machine’s measurements instead of vendor claims.

Failure modes

  • A console label gets treated as an API field. Gecko and Jetstream show why you check the schema before building on a name.
  • Enterprise behavior gets presented as consumer behavior. Documented does not mean “AI Mode does this.”
  • A configurable default gets presented as a fixed constant. chunkSize is configuration.
  • A secondhand post drops the source hedge. Read the source when the repost is louder than it is.
  • A documented field gets treated as proof of a consumer ranking effect. Those are different claims.

If agents already spend money against assumptions, this has the same failure shape as a limit that does not actually stop spending: a control appears present, but it is not the control you assumed.

When this protocol does not apply

This protocol is for secondhand, named-signal claims. Use a different standard when the claim has a primary source you can access and matching product scope, such as a real consumer experiment, an official Google engineering post, or a dated API response. In that case, “leak” can be evidence.

Raw artifacts also change the evidence: API responses, screenshots, and dated logs can support a claim when the author limits it to that exact scope. Ask two questions: does it have a receipt, and does that receipt match the claim’s scope?

Bottom line

The “leak” combines a documented enterprise API, a console label, and a name collision after removing the hedge. The real parts are public: the rankSignals field set, BM25 ranking, and configurable 500-token chunking. The three-tier PCTR unlock, the named answer model, and “therefore AI Mode does this” remain untestable. Test the field set. Keep the pipeline. Stop calling product docs a leak.