I audited a regional information portal that launched with 750 pages: SSR configured correctly, valid sitemaps, canonical tags in place, and nothing in the stack blocking a crawler. Google indexed the homepage and nothing else. The barrier here was never technical. It was economic.
TL;DR
- A crawlable site can still get almost nothing indexed. The barrier here wasn’t crawlability, it was Google’s cost-to-value calculation on the domain.
- “Discovered” and “Crawled” – currently not indexed are two different problems. One is a crawl-priority issue, the other is a content-quality rejection, and they need different fixes.
- 96% of this portal’s unindexed pages were never even fetched. That ratio pointed straight at a crawl-priority problem, not a content problem.
- Four forces compound: zero information gain, low crawl demand, high processing cost, no offsite signals. None of them acts alone.
- Fix priority matters: information gain and offsite demand first, cost reduction second, quality fixes last. Reversing that order wastes budget on pages Google still won’t fetch.
Google’s retrieval system runs a cost-to-value calculation on every URL, fed by crawl demand, information gain, and domain authority signals. When all of those come back negative, it withholds crawl resources entirely, however technically sound the site is. I will walk the diagnostic, then the four mechanisms behind it, then the order I would spend money in. The documented evidence for why Google’s systems work this way waits until the end.
The pattern is not specific to this portal. Any company launching a new content vertical hits the same retrieval economics, whether that is a SaaS knowledge base, a marketplace category expansion, or an aggregator entering a new market. The technical audit passes, the pages look right, and Google ignores them anyway. Seen it on a launch of your own?
One note before the diagnosis. Cost-to-value is my own working model, built to connect what Google documents (crawl budget guidance, DOJ testimony, the leaked schema) with a pattern I have watched repeat across sites. Where a source states something directly, I say so and point to it. The rest is me reasoning from those pieces to explain this one case, which makes it my read and not a disclosed Google formula.
The site and its indexation problem
The site aggregates cultural events, places, and local news. A place page carries contact details, opening hours, location data with Google Maps integration, public transport stops, and a 12-month event calendar. Under it sits Nuxt 3 with Server-Side Rendering, confirmed twice: once by reading the source, once through Sitebulb’s Response vs Render analysis, which scored 93/100 with zero meta robots changes across 757 URLs. The server delivers fully pre-rendered HTML with structured data in JSON-LD.
For weeks after launch, only the homepage entered Google’s index. Every subpage stayed outside it: places, events, news, categories. Search Console showed no manual actions and no security issues, and the sitemaps processed successfully.
The instinct in most SEO diagnostics is to hunt for a technical block: a misconfigured robots.txt, a noindex tag hiding in the rendered DOM, a JavaScript failure the crawler cannot parse. None of that applied. The server returned status 200 with correct content-type headers and self-referencing canonicals checked out. URL Inspection’s live test confirmed Google could access, render, and read every page I pointed it at.
Most audits stop here. That is exactly where the diagnostic starts.
What the crawl data revealed
A full crawl with JavaScript rendering enabled showed almost no gap between server response and rendered DOM. The canonical delta was near-zero: 755 URLs unchanged, 2 modified, 1 created. So SSR was doing its job and the rendering layer was not the problem.
The Pages report in Search Console said something else. Its “Why pages aren’t indexed” breakdown split like this:
| Status | URL Count |
|---|---|
| Discovered – currently not indexed | 342 |
| Crawled – currently not indexed | 13 |
| Redirect error | 5 |
| Page with redirect | 3 |
Six sitemaps had been submitted, declaring over 750 URLs between them: news, events, event categories, places, static pages. All of them returned “Success” status, so Google read and processed every one. Not a single subpage reached the index.
The link profile filled in the rest. Ahrefs put the Domain Rating at 9, with 228 referring domains and 1,156 backlinks, against zero organic traffic and zero organic keywords. A site-level search excluding the domain itself returned only the homepage. Not one of the institutions whose profiles the portal hosted linked back to its own page there, and no subpage was mentioned anywhere else on the web.
Two more problems turned up in the code. The JSON-LD on place pages carried an AI prompt artifact in the description field, a string fragment clearly generated during development and never cleaned out. Elsewhere the Schema.org data disagreed with what the page showed: mismatched URLs, phone numbers, email addresses. The structured data declared isAccessibleForFree: false while a visible badge read “Free admission.”
Outside the structured data, breadcrumbs on place pages pointed at categories under /events/ instead of /places/. The main descriptive section had no H2, and rel="nofollow" sat on links to the institutions’ own official websites.
Those quality issues matter, but the crawl data points at something structural underneath them.
Discovered vs. Crawled: two different problems
These two GSC statuses look similar but usually point toward different fixes: “Discovered” means Google knows a URL exists but has not crawled it yet. “Crawled” means Google fetched and read the page, then chose not to index it. Neither status on its own proves the cause, though they do narrow where to look first. Treating them as interchangeable is a common mistake in this kind of diagnosis.
| Status | What it means | What it tells you | The fix |
|---|---|---|---|
| Discovered – currently not indexed | Google knows the URL exists (via sitemap or link) but hasn’t crawled it yet | Points toward a crawl-priority or domain-signal problem, though the report alone doesn’t prove the cause | Domain-level: information gain, offsite demand signals, crawl cost |
| Crawled – currently not indexed | Google crawled the page but does not currently index it | Points toward a content or technical issue on the page itself (quality, duplication, canonicalization) | Page-level: content quality, uniqueness, structure |
The 342:13 ratio
This ratio is the diagnostic key. The 342 URLs parked in “Discovered – currently not indexed” are URLs Google knows about and has never fetched. Its crawl scheduler weighed the predicted value against the cost of fetching them and passed. That is not a judgment on content quality, because Google has not seen the content.
The 13 URLs in “Crawled – currently not indexed” are the opposite case. Google fetched those pages and decided, after reading them, that they did not belong in the index. That is a quality rejection, a different animal from never being fetched at all.
Do the arithmetic on the 342:13 split. 96% of the 355 URLs in the two “not indexed” buckets were never fetched at all, which puts the barrier at crawl scheduling rather than content evaluation. Most SEO practitioners meet an indexation problem by improving page quality, and here that would fix 13 URLs and leave 342 untouched. Those 342 need the domain to earn its crawl investment first, before page-level quality is even in play.
Say you run a B2B SaaS company and you have just launched a documentation hub or a resource center. The same split decides which of two organic-growth budgets you should be writing: one for content quality, one for domain-level authority and external demand signals. Pages that get crawled and rejected want the first. Pages that get discovered and never fetched want the second, because content improvements on pages Google has not read are money burned.
A negative cost-to-value loop
Four forces are at work, and no one of them explains the result on its own.
The one that matters most is zero information gain. And I mean zero. Every data point on the portal’s place pages already sits in Google’s index, taken from more authoritative sources: addresses, opening hours, contact details, transit stops, institutional descriptions. The official institution websites carry them, and so do Google Maps, Google Business Profiles, Facebook pages, and the public transit authority feeds.
The portal collects all of that and adds nothing that would widen what the index covers. For the retrieval system, indexing these pages buys no marginal value.
On top of that sits extremely low crawl demand, which compounds the information gain problem instead of merely adding to it. DR 9, zero organic traffic, zero keywords in the index. No external links to any subpage, and no branded queries aimed at specific URLs.
The sitemaps declare 750 URLs, but a sitemap creates supply, not demand. With no behavioral signals and nothing external vouching for the site, Google’s crawl scheduler files these pages at lowest priority.
Then there is the cost side. Each page serves a heavy DOM: a 12-month calendar rendered twice, once for mobile and once for desktop, contact sections duplicated for responsive variants, thousands of lines of inline SVG for decorative icons. Add extensive navigation markup on top. For a domain with near-zero crawl demand, fetching and processing pages like that costs Google more than the information it gets back.
The fourth force closes the loop: no fresh offsite signals. Not one of the institutions hosted on the portal links back, there are no social media citations, and nothing in other publications mentions it. So the cycle feeds itself. Without external signals there is no crawl demand, without crawling no indexation, without indexation no traffic, and without traffic nothing generates external signals again.
These four factors compound rather than acting independently. A site with low domain authority but genuinely unique content can still earn indexation, because information gain supplies the incentive on its own. Flip it: aggregate existing data, but arrive backed by strong external signals, and the crawl investment still comes, because demand overrides the low-gain assessment. Take information gain and crawl demand away at the same time, and the retrieval system has no economic reason to act.
Before you apply this: check where you actually stand
Pull your own GSC Pages report before assuming this playbook applies to you. Look at the “Why pages aren’t indexed” breakdown. If most of your unindexed URLs sit in “Crawled – currently not indexed,” you probably have a content or technical problem to work through page by page. That is a page-level job, not the domain-level economics this case turned on.
If most sit in “Discovered – currently not indexed” the way they did here, a 342:13-style split, treat this framework as a reasonable hypothesis to start from rather than an automatic diagnosis. Sample a handful of the affected URLs and see whether the pattern holds across your templates. Read Crawl Stats and the server logs (yes, the logs) next to the GSC report before you commit budget to a fix.
Shifting the economics: where capital should go
To break the loop you have to change the cost-to-value calculation at every stage, in order of impact.
Highest impact first: genuine information gain. The site has to produce something Google does not already have from a more authoritative source. Original reporting with a point of view would count. So would thematic guides connecting several places and events within one cultural district, post-event reviews, interviews with the organizers.
I use “information gain” descriptively here. It is a working diagnostic, not a documented Google crawl-allocation score. What matters is content that widens what the index already holds, which is a different thing from more content.
A B2B SaaS knowledge base fails the same test. If your documentation restates what the official API docs already say, the retrieval system has no reason to index a second copy. Whatever you add has to be additive: unique analysis, integration examples, practitioner perspective.
Next: offsite demand signals. Links from the institutions whose profiles the site hosts would be the most targeted signal available, an authoritative source saying the page about their own institution is worth something. Directory, industry-platform, and social-profile links do the duller job of setting a baseline.
As a rule of thumb from my own cases, the first $5,000 in this vertical should go entirely toward those institutional backlinks and credible external references, not toward producing more aggregated pages.
Third: the cost side. Move the 12-month calendar to client-side lazy loading instead of rendering it server-side in two responsive variants. Kill the DOM duplication between mobile and desktop, swap inline SVGs for a sprite, drop the excessive preload directives, and page weight falls a long way. None of that touches the value side, but it improves the ratio: each page gets cheaper for Google to process when crawl demand does arrive.
Once the economics have moved, the quality issues are worth the time: the AI artifact in the structured data, Schema.org markup that disagrees with the visible page, the broken breadcrumb hierarchy, the nofollow on institution links, the missing heading structure. Fix them on their own, though, and indexation does not follow. You would have a technically cleaner site that Google still ignores, which is the most expensive kind of clean.
The sequence matters: information gain and offsite signals first, cost reduction second, quality fixes third. That is the order the retrieval pipeline evaluates in, and it is the order your budget should follow.
How you’ll know it’s working
Watch the same GSC Pages report you used to diagnose the problem. The signal is the “Discovered” count falling while the indexed count rises. Not traffic, which lags indexation by definition. In my experience the first movement in crawl stats shows up within a few crawl cycles of new offsite signals landing, which for a domain at this scale means weeks rather than days.
If the Discovered count has not moved after a full cycle of those fixes, go back to the content. Ask whether it genuinely adds information gain, or whether it still overlaps what Google has from more authoritative sources.
The retrieval systems behind this decision
The diagnosis maps onto documented Google systems. See the scale those systems run at and the cost-to-value logic becomes unavoidable.
Google crawls trillions of pages. It indexes a fraction of them. Pandu Nayak, Google’s VP of Search, testified at the US vs Google antitrust trial that the index held about 400 billion documents as of 2020. That is a curated selection, not a record of the web.
He laid out the economics under oath. Page sizes are growing. Metadata per document is growing too, and with storage capacity fixed, each increase means fewer documents can be indexed: a tradeoff between the volume of data, its diminishing returns, and the cost of processing it. That establishes the resource constraint Google works under; it does not disclose a URL-level formula, and it does not prove why this portal landed where it did.
A new regional portal with 750 pages of aggregated data, all of which Google already holds from better sources, sits on the wrong side of that equation. At planetary scale every URL competes for retrieval resources against hundreds of billions of alternatives. The system is deprioritizing this site rather than ignoring it, and from where it sits the call is rational.
Google’s crawl budget documentation names URL popularity and staleness as factors in crawl demand. It does not say a never-indexed URL automatically scores zero on both. But with no subpage links and no search history to draw on, low demand was the reasonable read here, not a documented certainty.
Navboost, confirmed through DOJ trial documents, uses click and interaction data to modulate ranking, not crawling directly. A domain with zero organic traffic and zero clicks generates none of the behavioral signal Navboost feeds on. I take that missing proof to compound the same way the other signals do: nothing here tells the system this domain deserves attention.
The Google Leak and the DOJ trial documents both point to quality scores that attach to a domain, not only to individual pages. The leaked siteAuthority field is one documented example, and its description ties it to ranking (Q*) rather than to crawl scheduling directly.
From there I am reasoning, not reporting. A new domain with no ranking history and no authoritative links likely scores low on that kind of domain-level signal, and a low domain-level signal is a plausible contributor to low crawl priority. Nobody has confirmed the mechanism connecting the two.
Helpful content signals add another layer. Google folded its former standalone Helpful Content classifier into its core ranking systems back in March 2024, but the question behind it survives: was this made primarily for users, or to attract search traffic. A large run of pages with repetitive structure and minimal unique content is exactly what that question is built to catch.
This case cannot prove that signal caused the 13 “Crawled – currently not indexed” exclusions. But look at the place page template: identical UI structure, identical calendar and map components, one paragraph of unique description varying between pages. It fits the pattern closely.
One thing the URL Inspection documentation says out loud and most practitioners skip: the “What isn’t tested” section warns that pages must meet quality and security guidelines the tool never checks. A passing live test (“Google can access,” “Page is indexable”) is a necessary condition for indexation, not a sufficient one. Those green checkmarks buy false confidence when the actual barrier is economic rather than technical.
Good luck with the waiting part. Pozdrawiam serdecznie!
FAQ
Does “Discovered – currently not indexed” always mean low crawl demand?
In most cases, yes. This status indicates Google found the URL through a sitemap or internal link, but the crawl scheduler hasn’t allocated resources to fetch it. The common causes are low domain authority, absence of external demand signals, and predicted low information gain. It’s a crawl-priority decision, not a content-quality decision, because Google hasn’t seen the content yet.
Can submitting more sitemaps or using URL Inspection fix this?
Neither addresses the underlying economics. Sitemaps declare URL supply, they don’t create demand. URL Inspection’s “Request Indexing” can accelerate processing for individual URLs, but at scale (342 URLs in this case), it doesn’t change the crawl scheduler’s domain-level priority assessment. The fix is structural: improve the signals that drive crawl demand.
How does this relate to the Helpful Content system?
Helpful Content evaluates whether pages are made for users or for search engines. Structurally identical template pages with minimal unique content can trigger negative signals here. In this case, it’s a secondary factor: crawl demand is the primary barrier, but the template affects quality evaluation for the 13 URLs Google fetched and rejected.
How long does it take to fix a “Discovered – currently not indexed” problem?
There’s no fixed timeline, but the sequence matters more than the calendar. First movement in the GSC Pages report typically shows up within a few crawl cycles of genuine offsite signals landing, weeks rather than days for most domains. Quality-only fixes, without addressing information gain and crawl demand first, often show no movement at all.
Direction and strategy: Szymon Slowik. Research and drafting support: LLM. Reviewed, corrected and finalized: Szymon Slowik.
Created by Szymon, edited with an LLM… see any difference? 🙂