Google spent years denying that it uses domain authority. The denials, John Mueller’s included, are accurate: Google does not read Moz’s DA metric, Ahrefs’ Domain Rating, or any other third-party score. At the same time, Google’s own leaked documentation from 2024 contains a field called siteAuthority applied in Q*, a query-independent, site-level, largely static authority evaluation.
In this article I go through what site authority actually is, how you build it, and why it moves in steps rather than weeks. The deep evidence comes last: the field-by-field table from the leaked schema, the patent trail, and why the state barely moves between updates.
TL;DR
- siteAuthority is Google’s own field, not a vendor’s. Documented in the 2024 leak, it shares nothing with Moz’s DA, Ahrefs’ DR, or Semrush’s AS beyond the name.
- It’s largely static. The state is retuned offline and pushed out with broad core updates, so it moves in steps, not week to week.
- You build it by funding branded demand and fixing behavior on ranking pages, not by chasing a vendor score or a DR-led link placement.
- Trust and authority are two separate gates. Trust decides whether you’re eligible to appear at all; authority decides where you land once you are.
- Expect results in quarters, not weeks, arriving as step changes around core updates, and know what to stop funding while you wait.
One caveat frames everything below. The leak is a snapshot of Google’s documentation, not a live view of its production systems. Nobody outside Google can tell you whether this exact field runs this exact way today, or what weight it carried in the past.
Here comes a hard-to-swallow pill: in SEO you rarely get proof on this level. You get a model that either predicts what you observe or does not, and this one does to some extent. Read every plain statement here as my working conclusion from connecting documents, testimony, exploit data and client work.
If you own the budget rather than the keyword list, here is what this evidence changes. Site authority is a capital-allocation problem: branded demand is an input you can fund, vendor scores are not a KPI worth reporting upward, and returns arrive in steps around core updates rather than weeks, so the horizon belongs in the plan, not in the apology.
The dashboard and the per-business-model table below turn those three sentences into something you can run a quarterly review on. The deep evidence at the end of this article is what backs them up.
What is siteAuthority in Google’s leaked documents?
siteAuthority is a site-level score in Google’s leaked Content Warehouse documentation, described as “converted from quality_nsr.SiteAuthority, applied in Qstar”. It is not Moz’s Domain Authority (DA), Semrush’s Authority Score (AS), or any other third-party metric. The description discloses no scale and no inputs. What is documented: the field exists and feeds Q*, Google’s site quality scoring.
The naming collision fuels real confusion. I’ve had a client ask me how we could improve their DR (Ahrefs’ Domain Rating), convinced the metric itself is connected to Google’s systems and triggers rankings. Actually, the logic is reversed: a vendor score at best echoes authority, it does not create it.
If you want the full evidence chain, with grades for every claim, I built it in my Google Leak ranking signals pillar.
Fund branded demand. Earn links from sources that sit in the upper index tiers, carry real search traffic and match your topic, at a velocity that reads as reputation, not a campaign. Fix behavior on pages that already rank, keep the territory coherent, remove the liabilities, and hold through an update cycle.
When a CMO asks me “so what is our authority”, my answer has three eras. It used to be measured by keyword rankings. Then by the traffic particular URLs collect, and its quality. Now it is measured by retrieval eligibility and branded search, shaped by AI Overviews and the LLM discovery stage.
The only authority that counts, in every era, is results. If you want a dashboard for site-level authority, watch those: are you eligible for the answer surfaces where you compete, does your visibility hold its level through core updates, is your branded demand growing, and is your brand cited by name in AI answers.
The same watchlist splits by business model. Here is how I read it across the four I work with most:
| Business model | Where authority comes from | First three moves | What to measure first | What to expect (est.) |
|---|---|---|---|---|
| Ecommerce | Behavior plus brand demand: commercial pages that satisfy intent, and people searching for the store by name. Rankings tend to follow branded demand down the same way they followed it up | Fund demand for the store name, not third-party scores; internal links from search-click pages into revenue categories; fix intent on ranking category pages before adding filter-page sprawl | Branded searches for the store; click share on ranking categories; whether priority categories hold level through core updates | Internal-link changes can show in 1 quarter; intent fixes in 1-2 quarters; brand demand is a 4+ quarter build; site-level lifts arrive as steps at core updates |
| SaaS (B2B) | Brand-entity demand plus links in context. Authority rises when the company and its experts are cited by name in category conversations and AI answers | Original research that earns mentions and links from one asset; named experts with a visible body of work, not house bylines; clean up weak informational pages before scaling output | Named mentions in AI answers; branded search for the product; top-10 entry rate for demo and sign-up pathways | AI-answer citations can start moving in 1-2 quarters once coverage builds; territory concentration reads in 2-4 quarters; brand demand is a 4+ quarter build |
| B2B expert services (legal, financial, consulting) | The named expert: appearing alongside respected names, publishing original viewpoints, staying in a tight territory | Talks, podcasts and publications that place the expert next to known names; positions and client cases, not glossary definitions; keep the territory narrow enough to own | Citations of the named expert in answer surfaces; branded search for the expert and firm; lead pages holding position through updates | Territory concentration reads in 2-4 quarters; expert recognition is a 4+ quarter build; when trust crosses the gate first, answer surfaces open before classic rankings move |
| B2C mass services (banking, insurance) | Offline brand demand usually already exists; the SEO job is converting it into query-level trust and strong click behavior | Title and intent discipline across templates; repair pages that rank but lose the click or the lead; prune thin sprawl, never trust, legal or core service pages | Click and lead rate on ranking service pages; answer-surface eligibility on priority queries; branded search holding or growing | Title fixes can lift immediately; behavior-confirmed gains show in 1-2 quarters and settle over about a year (the 13-month window); the offline-brand edge is defended, not built |
These horizons are case observations from my own accounts, not documented constants. Page-level fixes can show earlier; site-level authority usually re-rates in steps around core updates.
And the other side of the allocation, what to stop funding:
- Ecommerce: thin new categories, filter-page sprawl, and DR-led placements that do nothing for store demand.
- SaaS: anonymous top-of-funnel output at scale before the weak informational pages are fixed.
- B2B expert services: generic glossary pages with no point of view and no named expert behind them.
- B2C mass services: template churn, legal-page pruning, and content audits that ignore click and lead behavior.
The same reading works outward: when you evaluate a domain you might get a link from, look for the shadow of the stored state rather than the score on a badge. The vendor scores, DA, DR, TF, AS, are vanity metrics if you treat them literally as indicators, and they disagree with each other on the same domain because each vendor runs its own formula over its own crawl index.
I showed at BrightonSEO how heavily they can be manipulated without any impact on a page’s quality, rankings, or ability to pass authority. If you buy placements, start from my guest post buying guide instead of a DR filter. Site-level authority is also one of the first things I assess in a strategic SEO audit: as eligibility, update behavior and branded demand read against your competitive set, not as a single score.
Every account we have says Google’s site-level quality is largely static: retuned offline against live experiments and rater feedback, then pushed out in broad core updates. That is why a site’s authority moves in steps around updates rather than drifting week by week. It is a level you sit at, not a dial you nudge.
The court record carries this directly. In the DOJ trial exhibits, Google’s Hyung-Jin Kim put it plainly: quality is “largely static and largely related to the site rather than the query”. Erfan Azimi, who sourced the leaked documents, walked me through the NSR mechanism behind it: collected from live experiments, retuned with rater feedback, then “pushed out via a broad core update” so it lands on the entire site at once. His full account, and my reasons for trusting it, are in the pillar I linked earlier.
This mechanism explains two things every experienced SEO has watched. First, sites that jump or drop as a step exactly around a core update, then hold the new level for months. Second, the frustrating giveback pattern, where gains earned between updates get repriced at the next one.
You can’t measure siteAuthority from outside. No tool reads Google’s internal store, so I won’t pretend to show you an isolated proof. What I can show you is the shadow this mechanism casts, watched over three years. We run SEO for a dietary catering delivery service, a brand far less recognized than its competitors, who sponsor top-tier sports teams and run offline shops.
We built broad thematic coverage, a handful of links, and better UX with real trust signals. After about two years the site reached top 5 on the most profitable commercial queries in its industry and the highest search visibility among its competitors, and it has survived every core update since the first half of 2023. Consistent, update-resistant visibility from a weaker brand is exactly what a slowly earned site-level authority state should look like from the outside.
It cuts both ways. This summer my own guest post buying guide stepped down in a single update window: “buy guest post links” went from position one to four in the US, and about twenty long-tail terms dropped out, with nothing changed on the page. The level moved, and that is the point.
One distinction I keep coming back to in my talks: trust and authority are not the same thing. Trust decides eligibility at the retrieval stage, whether you enter the set of candidates at all. Authority decides ordering among the similarly trusted. Most authority arguments happen one stage too late.
Independent probing of Google’s systems, Mark Williams-Cook’s among it, keeps finding hard eligibility lines: below some quality level a site simply stops appearing in features like featured snippets and People Also Ask. I unpacked the split at the Chiang Mai SEO Conference (slides here).
What else does Google store about a whole site?
The full siteAuthority entry sits in the CompressedQualitySignals module of the leaked API documentation: one line, nothing more. It confirms the concept existed inside Google, at least when that documentation was written. It does not give us the recipe. Whoever tells you they know exactly how siteAuthority is calculated is guessing.
Q* is Google’s own name, not an analyst’s invention. Eleven field descriptions across the leaked quality models reference it, including a deprecated field described as “NSR override bid, used in Q* for emergency overrides”. siteAuthority is not one number in isolation. The leaked schema stores site-level signals we can divide into three families: link graph inheritance, behavioral confirmation, and brand demand. Several of them sit in PerDocData, the record Google keeps per indexed URL. Your site’s reputation travels with every page you publish.
PerDocData is a per-document record, and yet it carries (or carried) site-level context: the PageRank of your homepage, site signals, legacy toolbar PageRank. Google stamps the site onto the page. This also works in the other direction: a good external link does not only pass authority, it improves the linked page’s own per-document record and its discoverability for crawlers.
There’s a lot of buzz and debate about links vs. mentions as authority-building factors. Links and mentions play different roles here, and I keep them deliberately separate (more on that split in my off-site SEO in the age of AI writeup). Long story short: links aim mostly at documents, while mentions at entities. They impact trust and authority in different ways.
Here is the palette of site-level signals, grouped by family. Quoted lines are verbatim from the leaked documents; the rows tagged patent or source account come from outside the leak, and I keep those evidence grades separate:
| Field | Family | What the evidence says | What plausibly moves it |
|---|---|---|---|
| homepagePagerankNs | Link graph | “The page-rank of the homepage of the site”, copied onto every document | Link equity of your homepage |
| pagerankNs (seed distance) | Link graph | Patent US9953049B1: ranks pages by shortest distance to a set of specially selected, high-quality seed pages in the link graph | Proximity to institutions, large publishers, trusted sources |
| toolbarPagerank | Link graph | Still stored, in an internal directory literally named “fakepr” | Legacy; nothing you should chase |
| onsiteProminence | Behavioral | “computed by propagating simulated traffic from the homepage and high craps click pages” | Internal links from pages that earn real search clicks |
| NSR | Behavioral + brand | Retuned offline, pushed out with broad core updates; inputs include links, PageRank, clicks, impressions, Chrome traffic (source account) | Sustained behavior and demand, not weekly tweaks |
| predictedDefaultNsr | Behavioral + brand | “Predicted default NSR score computed in Goldmine via the NSR default predictor” | A starting estimate before behavior accumulates |
| siteAuthority | Aggregate | “Converted from quality_nsr.SiteAuthority, applied in Qstar” | The families above, consolidated |
| fireflySiteSignal | Behavioral (contested reading) | “Contains Site signal information for Firefly ranking change”; a companion module lists 15 site metrics, only 3 with documented meanings | On the field names: clicks, good clicks, impressions, publishing cadence |
One precision note on onsiteProminence, because it changes daily practice. The description says simulated traffic propagates from the homepage and from “high craps click pages”; craps is Google’s click data system, so this means pages with real search clicks, not pages with any traffic.
Build your internal link source lists from Search Console, not from your analytics. A page with heavy social or direct traffic and no search clicks doesn’t lend prominence on this reading.
predictedDefaultNsr has a patent twin too. Predicting Site Quality (US20140280011A1) describes scoring a brand-new site from a phrase model, before any behavior exists, and feeding the predicted score straight to the ranking engine. The leak field and the patent describe the same job from two sides: a starting estimate that behavior later replaces.
A second note on fireflySiteSignal. Its companion module tracks daily clicks, good clicks, impressions and article counts in 30-day windows, yet twelve of its fifteen fields carry no documented description. Shaun Anderson mapped the module in depth; read any field-by-field interpretation there, his or mine, against that gap.
The documents call Firefly a ranking change and never once a demotion. The tempting reading ties it to the March 2024 update, when Google formalized scaled content abuse as a spam policy aimed at mass-produced content regardless of how it was made. The documents never state that connection; what they do document is a module counting articles per 30-day window, exactly the measurement such a policy would need.
Add the three families together: link graph inheritance, behavioral confirmation, brand demand. Court testimony backs the first two families; the brand-demand family has its own patent trail, covered below. Notice what sits behind all three families: the brand entity.
Google works to understand your business the way it understands other entities: through co-occurrence, attributes, and demand. That stored, site-level state is largely static and largely about the site. In my opinion that is the strongest board-level case for funding brand that SEO can make.
I wrote more about the site-scope side of this in siteFocus, siteRadius and topical authority, and about why this favors recognized brands in why Google favors big sites.
What feeds siteAuthority and Q*?
Nothing in the documents gives a formula, so read this as my working map of the inputs. Five groups keep showing up across the leak, the patents, the testimony and the source account: links that count, behavior with its history, brand demand, content quality with coherence, and the negatives.
Links: the source’s standing, not the count
The documents describe an index kept in tiers. scaledSelectionTierRank is “a language normalized score ranging from 0-32767 over the serving tier (Base, Zeppelins, Landfills)”. Google’s own storage names a Landfills tier, and a link from a page parked there is not the link you were sold. So the questions that matter about a source are business questions: is it actually indexed and served, does it carry real search traffic, is it topically related to you, and what do its anchors say.
Velocity is measured too. phraseAnchorSpamDays tracks “over how many days 80% of these phrases were discovered”, next to counts of spam phrases across unique domains, a combined penalty, and a tally of demoted anchors. A burst of same-anchor links has a timestamped signature, and the schema ships with the penalty fields attached. Link growth that reads as reputation is spread across time, sources and anchors; growth that reads as a campaign is exactly what those fields exist to catch.
Behavior, traffic history and momentum
The source account behind NSR lists its components plainly: links, PageRank, clicks, impressions, Chrome traffic, and spam factors. Behavior is not one input among many; on that account it is the retuning loop itself.
History counts as well. hostAge stores the earliest date a host was seen, and its description says Google uses it “to sandbox fresh spam in serving time”. A new domain earns history before it earns the benefit of the doubt, and a domain with years of clean traffic behind it carries that record into every update.
Brand demand
The patent trail is the strongest here. Site quality score (Navneet Panda and April Lehman, filed 2012) computes site quality as a ratio: queries that name your site or brand directly, divided by all the queries where someone clicked through to your site at all. The more your traffic is people asking for you by name, relative to everyone who lands on you for any reason, the higher the score.
Mark Williams-Cook’s independent read of Google scoring data points the same way: branded search and mentions, user interactions as click probability, anchor relevance. Until demand and behavior accumulate, the predicted default described earlier stands in for the site.
Different decades, different authors, different reasons, and the same shape keeps appearing, with brand demand in most versions of it. Convergence is not proof. It is, however, enough to be the working model I run client strategy on.
Content quality and coherence
Per-URL effort and originality scoring (contentEffort, OriginalContentScore) aggregates into how the site reads as a whole, and the site-scope fields siteFocus and siteRadius measure whether the domain holds one recognizable subject or drifts. The Firefly module’s article counts per 30-day window add publishing cadence to the picture. A site that publishes a lot of thin, off-territory material is feeding several of these inputs at once, in the wrong direction.
The negatives: what drags the state down
These are not adjacent systems; their own descriptions say applied in Qstar. scamness (“Scam model score. Used as one of the web page quality qstar signals”), unauthoritativeScore (same phrasing), lowQuality, serpDemotion, and clutterScore, a “delta site-level signal in Q* penalizing sites with a large number of distracting/annoying resources”.
The quality score has a subtraction side, and shady link or content tactics write directly into it. One may argue a spam-heavy profile or thin-content footprint is also what drags a site below eligibility lines; the documents do not show the threshold mechanics, but the direction is written into the field names.
FAQ
Is siteAuthority the same as Moz’s Domain Authority?
No. Moz’s Domain Authority is a vendor’s prediction of ranking likelihood, built from its own link index. siteAuthority is an internal Google field from the leaked schema, converted from a site-quality model (quality_nsr) and applied in Q*. The names collide, but the two numbers come from entirely different systems, built on different data.
Can you check your siteAuthority score?
No. No tool reads Google’s internal systems; the field is only known from the 2024 leak. You can watch the shadow it casts: eligibility for snippets, People Also Ask and AI Overviews on the trust side, and step changes around core updates plus your branded-search trend on the authority side, per Williams-Cook’s probing and my own tracking.
Did Google confirm the leaked documents are real?
Yes. Google confirmed the documents’ authenticity in a statement to The Verge on May 29, 2024, while cautioning against out-of-context conclusions drawn from internal documentation. I documented the full provenance chain, including that statement and the leak’s origin, in the Google Leak pillar linked above, along with why I find the source credible.
Direction and strategy: Szymon Slowik. Research and drafting support: LLM. Reviewed, corrected and finalized: Szymon Slowik.
Created by Szymon, edited with an LLM… see any difference? 🙂