Guide · Indexing
Understanding Indexing and Crawling: How Your Page Gets into the Google Index
Three stages, three tools, one checkpoint
Google processes your page in three stages: crawling, rendering, indexing. Whether a URL is in the index, you check in under a minute with the URL inspection of the Google Search Console. The three most common blockers are noindex, faulty canonicals and robots.txt blocks – three tools that get mixed up constantly.
Crawling, rendering, indexing: the three stages
Crawling: the Googlebot retrieves the URL. It finds it via links, your XML sitemap or a re-check of already known URLs.
Rendering: Google executes the page’s JavaScript – with a Chromium-based rendering service. Only then does Google see the page as your visitor sees it.
Indexing: Google evaluates the rendered content, chooses the canonical version and adds the page to the index. Or not.
The point where most misunderstandings arise: crawled doesn’t mean indexed. Google crawls significantly more URLs than it indexes. And even indexing guarantees no visitors – an analysis of around 14 billion pages shows that 96.55 percent of all examined pages get zero traffic from Google search (Ahrefs, 2023). Without the index, no ranking, no matter how good the text is.
Checking in the Search Console whether a URL is in the index
Forget the site: query – it’s a rough estimate, not a report. The reliable source is the Google Search Console, and it costs 0 euros.
Single URL: enter the address into the URL inspection. You’ll see in seconds: is the URL on Google? When was it last crawled? Which canonical URL did Google choose?
Whole website: the “Indexing → Pages” report lists all non-indexed URLs with reasons. Two statuses come up constantly: “Crawled – currently not indexed” (a quality or duplicate problem) and “Discovered – currently not indexed” (not yet crawled, often a question of crawl prioritisation).
robots.txt, noindex, canonical: three tools, three tasks
robots.txt controls crawling
A disallow rule in robots.txt forbids retrieval. But it doesn’t prevent indexing: if the URL is linked, it can still end up in the index – as an empty entry without content.
noindex controls indexing
The meta robots tag noindex says: “Don’t put this page in the index.” For Google to read that, the page has to be crawlable. From this follows mistake number 1 in almost every audit: blocking a page via robots.txt and setting it to noindex at the same time. Google never sees the noindex. If you want a page out: set noindex, allow crawling.
Canonical is a hint, not an instruction
The canonical tag recommends a preferred version to Google for duplicates. Google usually follows – but not always. It’s no good as a blocking tool.
Crawl budget: only a topic from five-figure URL counts
The most overrated topic in technical SEO. Websites under about 10,000 URLs practically don’t need to worry. It becomes relevant from around 1 million URLs – or from about 10,000 URLs with daily changing content, for example large shops with faceted navigation.
If your website has 300 pages and one of them isn’t indexed, that’s not down to crawl budget. It’s blocked, too thin or a duplicate.
JavaScript rendering: when content only appears in the browser
If your CMS delivers content via client-side JavaScript, Google may see an almost empty page on first retrieval. Here’s how to check that in 2 minutes: open the URL inspection in the Search Console, start the live test, look at the rendered HTML. If texts or products are missing there, you have a rendering problem.
The most robust solution: server-side rendering – the normal case with WordPress, Shopify and Shopware, the most common problem area with headless setups.
More system – status codes, redirects, sitemaps – in the technical SEO guide.
Indexing signals at a glance
| Signal | Controls | Effect |
|---|---|---|
| noindex | Indexing | Page stays out of the index – requires it to be crawlable |
| Canonical | Duplicates | Recommendation to Google, not a block |
| robots.txt (disallow) | Crawling | Prevents retrieval, but not necessarily indexing |
| XML sitemap | Discovery | Should only contain index-worthy URLs with status code 200 |
The four signals solve different tasks. Whoever mixes them – for example using robots.txt as a deindexing tool – produces exactly the mistakes from the next section.
A Shopware shop with a crawling problem
“On a Shopware shop with around 60,000 URLs, about two thirds of all crawls went into filter URLs without any search demand. We blocked the parameter patterns via robots.txt, cleaned up the sitemap and concentrated internal linking on categories. After eight weeks, new products appeared in the index after two to three days instead of around two weeks.”
— Viktor Pásztor, SEO freelancer
The example shows why crawl budget only becomes a topic with large, dynamic websites: it’s not the number of pages alone that decides, but how much of it Google wastes on irrelevant URLs.
Common mistakes with indexing and crawling
robots.txt block combined with noindex
Google never sees the noindex if the page is simultaneously blocked via robots.txt. The page remains in the index as an empty entry.
Canonical misunderstood as deindexing
It’s a recommendation for duplicates, not an eviction. Google usually follows the canonical – but without guarantee.
noindex taken live from staging
Check after every relaunch: a forgotten noindex from the staging environment takes the complete live website out of the index.
Sitemap full of junk, crawl budget panic on small sites
Only index-worthy URLs with status code 200 belong in the sitemap. And under 10,000 URLs, you solve index problems through quality and internal linking, not through crawl budget.
Frequently asked questions about indexing
How long does it take Google to index a new page?
Between a few hours and several weeks. On established websites with a clean sitemap and internal linking, 1 to 7 days are realistic.
Why is my page crawled but not indexed?
Google currently doesn’t consider the page index-worthy – typical with duplicates, very thin content or weak internal linking. There is no entitlement to indexing.
How do I quickly remove a page from the Google index?
For the immediate effect, use the removals tool in the Search Console – it hides the URL for around 6 months. Only noindex on a crawlable page or status code 410 works permanently.
Does the Indexing API help with normal pages?
No. Google officially supports the Indexing API for only two content types: job postings and livestream videos.
Viktor Pásztor
Viktor Pásztor is an SEO freelancer in Berlin, has been in digital marketing for over 15 years and manages around 15 client projects in parallel, mostly e-commerce and B2B. He has been working 100 % remotely for more than five years – with WordPress, Shopify, Shopware and TYPO3. More about Viktor Pásztor
Keep reading
Technical SEO
What a crawl finds and what happens next.
On-page optimisation
How content is prepared for ranking and indexing.
SEO pricing
What working with me costs.
Indexing problems on your website?
An SEO audit checks crawling, rendering and indexing point by point – including a look into your Search Console. Reply within 24 hours on business days.
→ Request an SEO audit now