Emulent Let’s Talk
Skip links
Emulent Diagram

How Google Ranks Pages

From a page existing on the web to appearing in a results position. Four stages, each feeding the next: crawling and indexing run continuously in the background, ranking and serving happen fresh for every query.

01

Crawling

Before Google can rank a page it has to find it and read it. Googlebot works through an enormous, constantly refreshed list of URLs, fetching each page the way a mobile browser would.

URL discovery

New URLs are found by following links from pages Google already knows, reading submitted XML sitemaps, and revisiting URLs from earlier crawls. A page nothing links to and no sitemap lists may never be discovered at all.

Crawl scheduling

An algorithm decides which URLs to crawl, how often, and how hard to hit each site. This crawl budget depends on how important the pages seem, how often they change, and how fast the server responds without straining.

Fetching

Googlebot requests the page mobile-first, using a current Chrome rendering engine. It records the HTTP response: 200 gets processed, redirects are followed, 404s and server errors are noted and retried on a decaying schedule.

Crawl controls

Site owners steer the crawler: robots.txt blocks paths from being fetched, sitemaps suggest priorities, and Search Console reports what was crawled. Blocking a page here stops crawling, not necessarily indexing of its URL.

02

Indexing

A fetched page is not yet in Google. Indexing turns the raw HTML into a structured entry in the search index, and plenty of pages are crawled but never make it in.

Rendering

Pages enter a render queue where JavaScript is executed, so content that only appears after scripts run can still be indexed. Rendering can lag the initial crawl by hours or days, which is why JS-heavy sites index more slowly.

Content analysis

Google parses the text, title tag, headings, links, images, and video, works out what the page is about, and detects its language and regional relevance. Structured data markup is read here and can qualify the page for rich results.

Canonicalization

Duplicate and near-duplicate pages are clustered, and one canonical URL is chosen to represent the cluster in results. Signals include rel=canonical tags, redirects, sitemap URLs, and internal linking. Only the canonical gathers full ranking credit.

Storage in the index

The processed page is written to the index, a database of hundreds of billions of pages organized like a book’s index: every significant word maps to the pages that contain it. A noindex tag found at this stage keeps the page out entirely.

03

Ranking

Ranking happens at query time. Language models such as BERT and MUM interpret intent before any matching happens: correcting spelling, expanding synonyms, recognizing entities, and classifying whether the searcher wants to learn, buy, navigate, or find something local. Candidate pages are then scored against these families of signals.

Relevance of content

The most basic signal is whether the page contains the query’s terms and concepts, weighted by where they appear: titles and headings count more than body text. Beyond keywords, systems judge topical depth and whether the page’s format matches the intent, such as a tutorial for a how-to query.

Quality of content

Among relevant pages, Google favors those that demonstrate experience, expertise, authoritativeness, and trust (E-E-A-T). Links from prominent sites remain a core proxy, descended from the original PageRank algorithm, alongside helpful-content signals that reward pages written for people rather than for search engines.

Usability of pages

When competing pages are similar in relevance and quality, page experience breaks ties: Core Web Vitals (loading speed, interactivity, layout stability), mobile-friendliness, HTTPS, and the absence of intrusive interstitials that block the content.

Context and settings

The same query returns different results depending on location, language, device, and recent search activity. Freshness systems also weigh in: “football scores” surfaces pages from the last hour, while “how to tie a tie” can happily return a page from years ago.

Named systems doing this workPageRankBERTMUMRankBrainHelpful content systemReviews systemSpamBrainCore updates
04

Serving results

Scored pages are assembled into a results page in well under a second. What the searcher sees is more than ten blue links: the layout itself is chosen per query.

Ranked web results

The classic list, ordered by final score. Titles and snippets are generated per query, often rewritten from the page’s own text to highlight the part that answers the search.

AI Overviews and snippets

For queries where a synthesized answer helps, an AI Overview or featured snippet sits above the results, built from and linking to highly ranked pages. Ranking well is what earns a page a place inside these features.

Result features

Depending on intent, the page mixes in a local pack with a map, image and video blocks, top stories, shopping results, and knowledge panels. Each feature runs its own ranking over its own corpus.

Continuous re-evaluation

Nothing is final. Pages are re-crawled and re-scored as the web changes, and Google ships thousands of ranking changes a year, including broad core updates that can reshuffle results across every topic.

Crawling and indexing happen continuously, ahead of any search. Query understanding, ranking, and serving happen fresh for every query, in under a second. The loop never stops: pages are re-crawled, re-indexed, and re-scored as the web and Google’s systems change. Companion diagrams: how Gemini chooses sources and how ChatGPT chooses sources.

Want your pages to be the ones that get found and cited?We build SEO, AI search optimization and content strategy into every site we design.

Get a Free Quote