Technical SEO

A complete guide to technical SEO — a plain-language Beginner's Guide and a systems-level Advanced Guide to crawling, rendering, indexing, and ranking.

First published: Jun 25, 2026 · Last updated: Jul 18, 2026 · Advanced
demand #53 in Technical SEO#73 on the site

Two guides in one. The Beginner's Guide explains technical SEO from scratch — the crawl → index → rank pipeline, the handful of foundations every site needs, how to check your own site, and which myths to ignore. The Advanced Guide goes systems-deep: crawl budget, the rendering decision, canonicalization's ~40 signals, internal linking, Core Web Vitals as three separate problems, ongoing monitoring, migrations, and AI search. The throughline is the one I always come back to — technical SEO is the most important part of SEO until it isn't. It's the foundation that lets content and links rank, not a ranking trick of its own. You can't rank a page Google won't index, so the highest-value work is usually the most boring.

Explore this topic

Loading local progress…

Technical SEO Foundations

The essential crawl, index, and diagnose sequence.

View the full learning path
Follow the guided pathStart with the first guide →
  1. How Search WorksHow Google and Bing turn the open web into an answer — the crawl → index → serve pipeline as three gates, with rendering, canonicalization, and ranking systems explained.
  2. CrawlingHow search engines discover and download the web — Googlebot and Bingbot, URL discovery, the crawl scheduler, rendering, and how crawling differs from indexing and ranking. The hub for everything crawl-related.
  3. Robots.txtWhat robots.txt actually does — it controls crawling, not indexing — plus the exact syntax, how Google handles it under the hood, and the mistakes that break sites.
  4. SitemapsWhat a sitemap is, whether you need one, the types (XML, HTML, image, video, news), how to find, create, and submit one, and why sitemaps help discovery but never guarantee indexing.
  5. IndexingHow search engines store and organize pages so they can rank — content analysis, canonicalization, why crawled isn't indexed, and reading the GSC Page indexing report.
  6. CanonicalizationHow search engines pick one canonical URL among duplicates and consolidate ranking signals onto it — why rel=canonical is a hint, not a rule, and how to align every signal.
  7. Canonical TagHow to implement rel=canonical correctly — the HTML link element, the HTTP Link header for PDFs, absolute URLs, one per page — and where people go wrong.
  8. Mobile-First IndexingWhat mobile-first indexing actually is — Google using your mobile HTML for indexing and ranking — why content parity is rule
  9. Duplicate ContentThere's no duplicate content penalty. What duplicate content really costs you, how Google clusters duplicates and picks a canonical, and how to fix it.
  10. On-Page SEO: Complete GuideA practical map of the page-level signals you control, what each one can influence, and where to find the site's detailed implementation guides.
  11. Meta Tags for SEOWhich HTML head elements actually matter for SEO in 2026 — the title link, meta description, and robots meta tag — plus X-Robots-Tag, rel attributes, and the dead tags to drop.
  12. Title TagWhat a title tag is, why it's a light ranking factor and the main source of the SERP title link, why Google rewrites titles since 2021, and how to stop it.
  13. Meta DescriptionThe meta description suggests your SERP snippet but is not a ranking factor. See Google's query-dependent selection and a 62.78% difference observed in one Ahrefs dataset.
  14. H1 TagWhat an H1 tag is, how many you really need per page, whether it's a ranking factor, how it differs from the title tag, and how to audit yours.
  15. Open Graph Tags: Not a Ranking Factor, Still an SEO DeliverableWhat Open Graph tags are, why they're not a Google ranking factor, what Google actually does with og:title/og:image/og:site_name, documented per-platform og:image sizing, per-platform fallback and caching behavior, and how to force a re-scrape.
  16. Multiple H1 Tags: Does It Hurt SEO?Does having more than one H1 tag hurt your SEO? Google's John Mueller says no — and the HTML spec allows multiple sibling h1s. The catch is what the 2022 spec change actually changed, and why accessibility, not ranking, is the real reason to keep one clear H1.
  17. Social Sharing Images (og:image and twitter:image): Sizes, Limits, and FallbacksThe image spec sheet for og:image and twitter:image — the safe 1200×630 (1.91:1) default, per-platform dimensions and file-size caps, the absolute-URL rule, Google's own 2026 thumbnail requirements, and why images fail vs. fall back.
  18. Search Engine Tools (Webmaster Tools)The free, first-party consoles search engines give you — Google Search Console and Bing Webmaster Tools — what they show, how they differ, and why every site should connect both.
  19. Google Search Console (GSC)What Google Search Console is, who needs it, how to set it up, and which report answers your question — the hub above every GSC report deep dive.
  20. URL Inspection ToolHow Google's URL Inspection tool reports a single URL — the indexed snapshot vs the live test, reading the coverage panel, canonicals, Request Indexing, and the API.
  21. Core Web VitalsGoogle's three real-user UX metrics — LCP, INP, and CLS — their "good" thresholds, field vs lab data, how much they matter for ranking, and the tools that measure them.
  22. Critical Rendering PathThe browser's step-by-step pipeline from bytes to visible pixels — DOM, CSSOM, render tree, layout, and paint — why CSS and JavaScript block rendering, the three optimization levers, and how it drives FCP, LCP, and Googlebot rendering. The hub for render-blocking resources.
  23. HTTPS for SEOHow much HTTPS actually helps rankings, why it matters far more for trust and browser features, and how to migrate HTTP→HTTPS without losing traffic — redirects, mixed content, and HSTS.
  24. Critical CSSCritical CSS — extracting the above-the-fold styles, inlining them, and deferring the rest to speed up first paint. Why Google calls it advanced and optional, the real tradeoffs (lost caching, maintenance risk, race conditions), and how to tell whether CSS is even your bottleneck.
  25. HSTS: HTTP Strict Transport Security for SEOWhat HSTS actually does, the Strict-Transport-Security header syntax (max-age, includeSubDomains, preload), the browser-only internal redirect crawlers never see, why it doesn't replace your 301s, and how preload can lock you in — from Patrick Stox.
  26. HTTP to HTTPS MigrationThe step-by-step HTTP→HTTPS migration playbook — pre-migration audit, certificate selection, staging tests, redirect mapping at scale, canonical/sitemap/hreflang updates, getting Search Console coverage right, monitoring windows, launch-day regressions, and a rollback plan.
  27. Mixed ContentWhat mixed content is, why active mixed content gets blocked while passive gets warned about, and how to detect and fix insecure sub-resources at scale — the browser console, CSP reporting, upgrade-insecure-requests, block-all-mixed-content, and how CMSes and ad tech reintroduce it.
  28. SSL/TLS CertificatesDV vs OV vs EV, wildcard vs SAN, Let's Encrypt and free automated issuance, certificate chain failures, expiration and auto-renewal, and what breaks for users vs crawlers when a certificate is invalid — the certificate-level deep dive under the HTTPS hub.

Every published guide in this hub, grouped by the site taxonomy.

HTTP Status Codes30

On-Page80

How Search Works64

Mobile SEO10

Website Structure19

Platform SEO57

Tools26

Web Performance31

Ecommerce SEO10

JavaScript SEO6

Edge SEO5

Site Migrations9

News & Discover SEO5

International SEO6

HTTPS5

Pillar overview1

Video SEO2

How the concepts connect

See which concepts make later topics easier to understand.

Loading local progress…

Start on the left. Follow the arrows right. Open the preview to explore individual guides.

Required first Useful next Related
Technical SEO content dependency graph Arrows run from a prerequisite concept to the topic it unlocks. A semantic nested list follows this graph. Required prerequisite: Technical SEO to How Search WorksRequired prerequisite: How Search Works to DiscoveryRequired prerequisite: Discovery to CrawlingRequired prerequisite: Crawling to IndexingRequired prerequisite: Indexing to CanonicalizationRequired prerequisite: Canonicalization to Canonical TagRecommended next topic: Crawling to Robots.txtRecommended next topic: Crawling to SitemapsRecommended next topic: Crawling to RenderingRequired prerequisite: Rendering to JavaScript SEORelated topic: Crawling to Meta Robots TagRecommended next topic: Indexing to URL Inspection Tool START HERE Technical SEO Foundation guide How Search Works Open guide Discovery Open guide Crawling Open guide Robots.txt Open guide Sitemaps Open guide Indexing Open guide Meta Robots Tag Open guide Canonicalization Open guide Canonical Tag Open guide Rendering Open guide JavaScript SEO Open guide URL Inspection Tool Open guide
Expanded map

How the concepts connect

Prefer an outline? View the required path

This simpler hierarchy shows only the concepts that should be understood first.

  1. Technical SEO Platform scope: generic-http-html
    1. How Search Works required Platform scope: generic-http-html
      1. Discovery required Platform scope: generic-http-html
        1. Crawling required Platform scope: generic-http-html
          1. Indexing required Platform scope: generic-http-html
            1. Canonicalization required Platform scope: generic-http-html
              1. Canonical Tag required Platform scope: generic-http-html, nextjs-vercel, wordpress, shopify
  2. Robots.txt Platform scope: generic-http-html, apache, nginx, cloudflare-workers-rules
  3. Sitemaps Platform scope: generic-http-html
  4. Meta Robots Tag Platform scope: generic-http-html, nextjs-vercel, wordpress, shopify
  5. Rendering Platform scope: generic-http-html, nextjs-vercel
    1. JavaScript SEO required Platform scope: nextjs-vercel, wordpress, shopify
  6. URL Inspection Tool Platform scope: platform-managed

Recommended and related connections

TL;DR — Technical SEO is the same crawl → render → index → serve pipeline on every site — there’s no separate “technical SEO algorithm” — and it’s a foundation, not a ranking factor of its own. Treat the pipeline as a series of gates and diagnose which one a page is stuck at before changing anything. The leverage is mostly negative (not losing what you’ve earned), so the boring structural work — canonicalization, redirects, internal links — pays best, and it compounds at scale. Most sites don’t need to manage crawl budget; rendering is a separate step that can lag; Core Web Vitals are three distinct problems and a minor ranking lever; canonicalization is a weighted decision across ~40 signals; and from 2025 on, AI search gates eligibility on clean technical signals before it ranks or cites you. The highest skill here is prioritization — knowing what to ignore.

Technical SEO decides eligibility, not position

Technical SEO is the one part of SEO whose payoff is almost entirely negative: its job is to keep you from losing rankings, not to win them. Google doesn’t hand out positions for clean plumbing. The same crawling, indexing, and ranking systems run whether your site is pristine or a disaster — there’s no separate “technical SEO algorithm” sitting behind them. What technical SEO actually decides is whether your pages can enter those systems at all, and whether the engine understands them correctly once they’re in.

So the right mental model isn’t “do technical SEO to rank.” It’s “do technical SEO so your content and links are allowed to rank.” That inversion is the whole reason the unglamorous work — canonicalization, redirects, internal links — is the highest-value work, and why the single most useful skill in this discipline is prioritization: knowing what to fix and, just as often, what to leave alone.

The pipeline, as gates

Everything hangs off one pipeline, and the operative clause from Google is “not all pages make it through each stage.” Evidence for this claim Google Search describes crawling, indexing, and serving as three stages; discovery is part of the crawling stage, and not every page advances through each stage. Scope: Google Search documentation; conceptual explanation, not a promise of ranking outcomes. Confidence: high · Verified: Google: How Search Works Don’t picture a conveyor belt that carries every page to the finish. Picture a series of gates, each with its own pass/fail:

  • Crawl — discovery (links + sitemaps + push protocols) plus the fetch. A page nothing links to, or one disallowed in robots.txt, may never arrive.
  • Render — Google runs your JavaScript in a recent headless Chrome (the Web Rendering Service) before it can fully understand the page. This is a separate step from the fetch, it’s stateless, and it can lag.
  • Index — the engine processes the page, picks a canonical among duplicates, and decides whether to store it. “Indexing isn’t guaranteed” even when crawling and rendering succeed.
  • Serve — query understanding, then ranking across many automated systems, then the search features layered on top.

Keep crawl ≠ render ≠ index ≠ rank separate in your head and most of technical SEO stops being mysterious. When a page underperforms, you don’t guess and you don’t change ten things — you find which gate it failed and fix that stage.

One honest caveat before you treat any pipeline description as gospel, including mine: it’s a model, not the source code. How Search Works is a talk I give at conferences that walks through this whole pipeline (slides on SlideShare), and I open it with a warning I’ll repeat here: “this is my understanding of systems… not going to be 100% complete or accurate.” Hold it loosely and use it to reason about problems.

Who actually does the crawling

“Googlebot” sounds like one program. It’s a family — desktop, mobile (the one that matters, since indexing is mobile-first), image, news, video, and ads crawlers — all drawing from the same crawl budget pool, which is why a runaway image or parameter crawl can starve the crawling of your real content.

And it’s not just search engines anymore. When I analyzed Cloudflare Radar crawl data (an Ahrefs piece I wrote on the new wave of bots), search-engine crawlers still crawled the most, but AI bots were firmly in second place and on track to overtake them. If you read your logs, the cast of characters has changed — and managing it (which AI crawlers you allow, and confirming the ones hitting you are who they claim) is now part of the job.

Crawl budget: when it matters, and when it doesn’t

Google defines crawl budget as “the set of URLs that Google can and wants to crawl,” set by crawl capacity (your server’s health) and crawl demand (popularity and staleness). You raise effective budget two ways: give bots more capacity, or — far more often — stop wasting it. Consolidate duplicates, block low-value spaces, return 404/410 for permanently gone pages, fix soft 404s, keep sitemaps current with accurate lastmod, and avoid long redirect chains.

The reassuring part, and I’ll keep saying it: most sites don’t need to worry about crawl budget. Google itself tells you that if your pages are generally crawled the same day they’re published, “you don’t need to read this guide.” It starts to bite around 1M+ pages, or 10k+ pages that change rapidly. Below that, spend your energy elsewhere. Evidence for this claim Google says crawl-budget guidance is mainly relevant to very large sites, including sites with over one million unique pages or over ten thousand rapidly changing pages. Scope: Google Search guidance; the page-count examples are diagnostic starting points, not hard eligibility thresholds. Confidence: high · Verified: Google: Large site crawl budget guide Bing’s Fabrice Canel frames the same idea more bluntly: less is more — fewer URLs to crawl is better for SEO.

robots.txt: crawl control, not index control

The most important distinction in this whole file: robots.txt controls crawling, not indexing. Disallowing a URL stops bots from fetching it — it does not keep it out of the index. A disallowed page can still be indexed (URL only, no content) if other pages link to it, and worse, if you disallow a page you also stop Google from ever seeing a noindex tag on it.

So the rules are:

  • Want a page gone from search? Allow crawling and add noindex. Never use robots.txt to deindex.
  • Want bots to skip a low-value URL space (internal search, infinite faceted combinations) and don’t care about indexing? robots.txt disallow is correct.
  • Managing AI crawlers? This is also where you allow or block GPTBot, ClaudeBot, PerplexityBot, CCBot and friends — a strategic decision, not a default.

Canonicalization: a weighted decision, not a command

Canonicalization is where a lot of advanced technical SEO lives, and it’s widely misunderstood. rel="canonical" is a hint, not a directive. Google weighs it against many other signals — redirects, internal links, sitemap inclusion, HTTPS, URL structure — when it picks the representative URL. My deep dive on canonicalization puts the count around 40 signals that feed canonical selection, which is why you sometimes see “Duplicate, Google chose different canonical than user” in Search Console: your tag was outvoted.

The practical implications:

  • Don’t send conflicting signals. I spent years on enterprise sites (I ran technical SEO in-house at IBM), and in a talk I give called Enterprise SEO Chaos I show real pages that “redirected to one version, canonicaled to a second, and internally linked to a third.” Pick one URL and make every signal agree.
  • Signal strength roughly ranks redirect > rel="canonical" > internal links > sitemap. A 301 is a much stronger statement than a canonical tag.
  • Duplicate content isn’t a penalty. Google’s Gary Illyes has said roughly 60% of the web is duplicate content, and Google treats some of it as normal — not a spam violation. The cost is split signals and wasted crawling, not a punishment. The fix is consolidation, not panic.

And a note on JavaScript: I once ran a test — injecting a rel="canonical" via JavaScript on a page that had none in the HTML — and Google honored it, even though it had publicly said it wouldn’t. After that surfaced, Google updated its JavaScript SEO documentation. The lesson isn’t “use JS canonicals”; it’s that this stuff is testable, and the docs aren’t always the last word.

The rendering decision

Rendering is the step most overviews skip, and it’s where JavaScript sites get into trouble. “During the crawl, Google renders the page and runs any JavaScript it finds using a recent version of Chrome.” Evidence for this claim Google processes JavaScript pages in crawling, rendering, and indexing phases and uses a recent version of Chrome for rendering. Scope: Google Search JavaScript processing; rendering and indexing remain subject to technical and quality constraints. Confidence: high · Verified: Google: JavaScript SEO basics It’s a separate, stateless service, it can cache resources for weeks, and it can lag the initial fetch — so a JS-dependent change can take a while to be reflected.

JavaScript isn’t the enemy here. As I put it in my guide to JavaScript SEO, JavaScript is not bad for SEO, and it’s not evil — it’s just different from what many SEOs are used to. The real decision is how you render:

  • Server-side rendering (SSR) — safest for SEO; the HTML arrives complete.
  • Static generation (SSG/pre-rendering) — best of both worlds for content that doesn’t change per request.
  • Client-side rendering (CSR) — highest risk; the content only exists after JS runs, so you’re betting on the render step.
  • Dynamic rendering — Google calls it a workaround, not a recommendation; Bing is more favorable. Treat it as a bridge, not a destination.

Two traps to know cold. First, lazy-loading: Googlebot doesn’t scroll or click, so content that only loads on interaction can stay invisible — make sure it loads when it’s in the viewport. Second, links: Google can only follow a link that’s a real <a href> element. A routerLink or a click handler with no href is not a crawlable link. Verify the rendered output against the raw HTML with the URL Inspection tool whenever you suspect a gap.

Site architecture and internal linking

Internal links do three jobs at once: they help bots discover pages, they distribute PageRank, and they pass topical context through anchor text. John Mueller has called internal linking “super critical for SEO” and one of the biggest levers you have on your own site — and I agree. It’s one of the highest-ROI things you control directly.

A few systems-level points:

  • Orphan pages — pages nothing links to — are the first thing to hunt for. If it’s not linked, it’s barely discoverable and gets almost no equity.
  • Architecture is crawl-funnel management. Important pages belong close to the home page; deep, click-distant pages get crawled less and rank worse.
  • PageRank sculpting with nofollow is dead (since 2009). Nofollowing internal links makes that equity evaporate rather than redistribute. Manage flow with real architecture, not nofollow tricks.

Core Web Vitals: three problems, not one

The biggest practitioner error with page experience is treating it as a single “make the site faster” problem. Core Web Vitals are three distinct problems with different root causes and different fixes:

  • LCP (Largest Contentful Paint) — loading. Driven by server response time, render-blocking resources, and how fast the main content asset loads. Target under 2.5 seconds.
  • INP (Interaction to Next Paint) — interactivity. Driven by JavaScript execution blocking the main thread. Target under 200 milliseconds. (INP replaced FID in 2024 — if you still see FID anywhere, the advice is stale.)
  • CLS (Cumulative Layout Shift) — visual stability. Driven by images without dimensions, late-loading fonts, and injected content. Target under 0.1.

Two things matter beyond the definitions. Field data, not lab data: Google ranks on real-user CrUX data, not your Lighthouse score, so a Lighthouse 65 with good field data beats a Lighthouse 100 with bad field data. And proportion: I’ll be honest — I don’t think Core Web Vitals have much impact on SEO, and unless a site is extremely slow I generally won’t prioritize fixing them for rankings. Do the work for users and conversions; just don’t oversell it as a ranking lever.

Structured data: signals for search and AI

Structured data (use JSON-LD) doesn’t rank you, but it makes pages eligible for rich results and increasingly helps AI systems parse your content for citation. It’s genuinely useful — and genuinely oversold as a ranking signal. My honest framing: most of SEO is doing the basics well, and content and links move the needle more than schema does. Implement it where it unlocks a rich result or clarifies an entity; don’t expect it to lift rankings on its own. (And note: schema markup URLs are not crawlable internal links — Mueller has confirmed this.)

International, briefly

If you serve multiple languages or regions, use distinct URLs per version and hreflang annotations to map them, and prefer ccTLDs or subdirectories over URL parameters. Don’t auto-redirect by IP — Google explicitly warns against it and it breaks crawling. International SEO is deep enough to be its own pillar; this is just the technical handshake.

Technical SEO is an ongoing system, not a one-time audit

The framing every competitor guide gets wrong: technical SEO is not a checklist you complete once. Sites change constantly — deploys break canonical tags, a release slips a noindex into a template, a new ad script tanks INP, redirect chains accumulate. The mature practice is monitoring and regression detection:

  • Watch GSC Page Indexing for sudden changes in indexed counts and excluded statuses.
  • Watch Crawl Stats and your logs for response-code spikes and crawl-pattern shifts.
  • Re-validate crawl, render, and redirects after every significant deployment.

On log files specifically: I used to treat them as a once-every-few-years troubleshooting tool. That’s changed. Logs are now the clearest place to see which AI crawlers are actually hitting you and how often — something no other tool shows you as directly — so for anyone who cares about AI search, they’ve become a lot more useful than they were.

Site migrations: the highest-stakes event

A migration — new domain, HTTP to HTTPS, a replatform, a URL restructure — is the single highest-stakes technical event, because it touches every URL at once. Map old to new 1:1, use 301/308 permanent redirects, keep them in place indefinitely (I would not rush to remove them — a couple of redirect hops is nothing to worry about), and use the GSC Change of Address tool where it applies. Migrations can be complex and involve a lot of people, but don’t panic — you can fix almost anything that goes wrong. There’s a full Site Migrations cluster under this pillar.

The modern shift, and it cuts against the lazy “technical SEO is dead” take: from 2025 on, AI search systems decide eligibility before they ever rank or cite. To be quoted in an AI answer your page generally has to be cleanly canonicalized, fast enough, renderable without heroics, and structured enough to be parsed with confidence. Messy signals don’t just lower a ranking now — they can remove you from the answer entirely. Because Bing’s index feeds many LLM answers, Bing Webmaster Tools and IndexNow matter more than Bing’s search share suggests. Technical hygiene matters more in the AI era, not less.

Where the leverage actually is

If you take one thing from this guide, make it prioritization. Spend your time on indexing, canonicalization, internal links, and clean migrations — the work that decides whether pages exist in search and consolidate their equity. Don’t lose sleep over crawl budget, Core Web Vitals, duplicate content, or short redirect chains unless you have a specific, diagnosed problem. And don’t chase perfection — I doubt there’s a major website that’s technically perfect, and if there were, I’d worry they were wasting resources on things that don’t matter instead of things that do.

This hub maps the rest of the pillar: How Search Works, Site Migrations, On-Page, Search Engine Tools, and JavaScript SEO. Start wherever your site is breaking — the pipeline tells you which gate to look at first.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin an expert quote first.