What is technical SEO (and why it decides your rankings)

Alexandre Silva Anjos
Alexandre Silva Anjos · Founder of SEOWrench

Seventeen years building software, now applied to SEOWrench.

Most of what gets written about SEO is about content and links. Both matter, but they come later. Before any page can compete for a position, the search engine has to reach it, read what it says and decide to keep it in the index. If any of those three steps fails, the best article in the world is shown to nobody.

Technical SEO is the work of clearing the obstacles out of those three steps. It is not about writing better. It is about making sure that what is already written can be found.

This guide goes through each part in the order a search engine meets your site. Every section ends with a "how to check", which tells you where to look on your own site, with a free tool.

Contents

  1. How a search engine reaches your page
  2. Crawling
  3. robots.txt
  4. XML sitemap
  5. Indexing and noindex
  6. Status codes
  7. Redirects
  8. Canonical
  9. Title, description and headings
  10. HTTPS and mixed content
  11. Mobile
  12. Speed and Core Web Vitals
  13. Structured data
  14. hreflang, for sites in more than one language
  15. JavaScript and rendering
  16. Where to start

How a search engine reaches your page

Google describes its own work in three stages.

  • Crawling. A bot downloads the page, reads its links and notes the new addresses to visit later.
  • Indexing. The engine tries to understand what the page is about and decides whether it goes into the index, the collection that results are drawn from.
  • Serving. When someone searches, the engine picks, among the indexed pages, the ones that answer best.

Technical SEO lives almost entirely in the first two. The third depends on content, relevance and reputation, but it only exists for a page that made it through the other two.

One thing many people find out late: each stage has its own way of failing silently. A page blocked from crawling shows no error on screen. A page left out of the index still opens normally for anyone who has the link. Nothing warns you, which is why technical work starts with checking, not with fixing.

Crawling

A search bot discovers a new page in two ways: by following links from pages it already knows, and by reading the sitemap you give it. A page with no links pointing to it, and missing from the sitemap, is invisible. It exists on the server and does not exist for the search engine.

Crawling also has limits. On a small site, Google usually gets through everything. On a large site, with thousands of URLs, the bot chooses what to visit, and every useless address (a repeated filter, a tracking parameter, a duplicate page) competes for attention with the pages that matter.

The most common crawling problems are three:

  • Orphan pages, which no other page links to.
  • Links the bot cannot follow, like a button that only navigates through JavaScript, with no real <a href> in the HTML.
  • Broken internal links, which lead to a page that no longer exists.

How internal links are organized is a big enough topic to have its own guide: the one on site architecture for SEO covers hierarchy, navigation and links, with a checklist.

How to check. In Search Console, the Pages report, under Indexing, lists what Google found and did not index, with the reason. URL Inspection tells you whether a specific page has been crawled and when. With a SEOWrench account, the site scan crawls your pages by following links, the way a search engine would, and points out the ones that do not load.

robots.txt

robots.txt is a text file at the root of the site (yoursite.com/robots.txt) that tells bots where they can and cannot go. It is the first thing a well-behaved bot reads.

The most expensive mistake with it is misunderstanding what it does. robots.txt controls crawling, not indexing. A page blocked there can still show up on Google if other sites link to it. It just shows up without a description, because the bot could not read the content. To take a page out of the index, the tool is noindex, covered in the next section.

The other common mistakes:

  • A forgotten Disallow: / after the site leaves its staging environment. It blocks everything.
  • Blocking CSS and JavaScript. The search engine needs them to see the page the way a person sees it.
  • Blocking the AI search bots without noticing. Many rules copied from the internet block AI bots as a group, and with them go the ones that serve ChatGPT and Perplexity answers. The post on what GEO is explains the difference between a training bot and a search bot.

How to check. Open yoursite.com/robots.txt in the browser and read it. Search Console has a robots.txt report, under Settings, that shows the version Google read and any parsing errors. The SEOWrench free check reads your domain's robots.txt and tells you which AI search bots are blocked.

XML sitemap

A sitemap is the list of pages you want the search engine to know about, in an XML file. It does not guarantee indexing and it does not replace internal links. It helps the bot find what is new and what changed, especially on a large site or for pages buried deep.

A good sitemap has three qualities:

  • Only URLs that exist and that you want indexed. No pages that redirect, that return an error or that carry noindex.
  • The final, canonical URL. If the page answers at /product/ and the sitemap says /product, the search engine gets two messages.
  • An honest modification date. The lastmod field only helps if it changes when the content changes. A date that changes on every deploy teaches the search engine to ignore the field.

How to check. The Sitemaps report in Search Console shows whether the file was read, how many URLs it lists and how many were indexed. A large gap between the two counts is the signal to dig in.

Indexing and noindex

Crawling is not indexing. After reading the page, the search engine still decides whether it deserves a place in the index. It may leave it out for duplicate content, for thin content, or because you asked it to.

The request is noindex, which goes in a meta tag in the HTML or in an HTTP header:

<meta name="robots" content="noindex">

It is the right tool for pages that should not appear in search: logged-in areas, the cart, internal search results, thank-you pages. It is also behind some of the worst SEO scares, when it reaches production by mistake. The classic case is a new site that launches with the staging noindex still switched on.

One detail decides whether noindex works at all: the bot has to be able to read the page to see the tag. If robots.txt blocks the same page, the search engine never sees the request, and the page may stay in the index.

How to check. URL Inspection in Search Console tells you whether the page is indexed and, if not, why. The free check warns you when the page you pasted is marked noindex.

Status codes

Every time someone, person or bot, requests a page, the server answers with a three-digit number before sending the content. That number tells the search engine what happened.

Code What it means What the search engine does
200 it worked reads it and may index it
301, 308 moved permanently follows it and adopts the new address
302, 307 moved for now follows it, but tends to keep the old address
404, 410 does not exist drops it from the index over time
500 to 599 the server failed tries again later, and backs off

Two problems show up again and again. The first is the broken page nobody noticed, still linked from the menu or from an old post. The second is the disguised error: the page says "not found" to the reader but answers 200 to the bot. Google calls this a soft 404, and it wastes crawling on an empty page.

How to check. The Pages report in Search Console splits URLs by reason, including "not found (404)" and "soft 404". The free check shows the code the page returned, and the account scan points out the pages on the site that do not load.

Redirects

A redirect sends whoever requests one address to another. It is the right way to handle a page that changed URL, the move from http to https, and the version with and without www.

What goes wrong:

  • A temporary redirect where it should be permanent. If it moved for good, it is a 301 or 308.
  • A chain. Page A sends to B, which sends to C. Every hop costs time, and the bot may give up halfway.
  • A loop. A sends to B, which sends back to A. The page never opens.
  • An internal link pointing at the old address. The redirect works, but every click inside your own site makes an extra hop for no reason. The real cost of that, and what it does not cost, is in the post on internal links that redirect.

How to check. The free check tells you whether the page you pasted redirected and where to. The account scan lists the internal links on the site that land on a redirect, with the page each one sits on.

Canonical

When the same content answers at more than one address, the search engine has to pick one. It happens more than it seems: with and without a trailing slash, with a campaign tracking parameter, with a store filter, in a print version.

The canonical tag is how you say which address is the main one:

<link rel="canonical" href="https://yoursite.com/product/">

It is a strong hint, not an order. Google may choose another address if other signals say otherwise, like the sitemap or internal links pointing at the other version. That is why the canonical works best when everything else agrees with it.

Common mistakes: a missing canonical, a canonical pointing at a page that redirects, and the same canonical copied onto every page of a template, which tells the search engine the whole site is a single page.

How to check. URL Inspection shows the canonical you declared and the one Google chose. When they differ, that is where the problem lives. The free check warns you when the page declares no canonical or when it points to another domain.

Title, description and headings

These three elements tell the search engine, and whoever sees the result, what the page is about.

  • Title. It is the blue headline in the search result. One per page, different on every page, with the subject at the start. A title that is too long gets cut on screen, and the part that gets cut is often the best one.
  • Meta description. It is the text under the headline. It does not change the position, but it changes the click. Google may replace it with whatever passage of the page it prefers, and often does when the description is generic.
  • Headings. One <h1> that states the subject of the page, and <h2> and <h3> in order, like the chapters of a book. Jumping from <h1> to <h4> breaks nothing visible, but it scrambles the structure the bot uses to understand the text.

How to check. The free check measures the length of the title and the description, and checks that the page has a single <h1> and that the headings follow their order. To see how Google is actually showing your pages, search site:yoursite.com and read the results.

HTTPS and mixed content

Every site should answer over HTTPS, with the padlock. Google treats it as part of page experience, and browsers warn visitors when a site does not have it.

The problem left over after the migration is mixed content: the page opens over HTTPS but loads an image, a script or a font over http://. The browser blocks part of it, and what it does not block weakens the padlock. It happens a lot on older sites, with image addresses typed by hand into the body text.

How to check. The free check tells you whether the page answers over HTTPS and lists the resources loaded over http://. In the browser, the developer tools console also flags every mixed resource.

Mobile

Google indexes the mobile version of your site. It calls this mobile-first indexing. If content exists on desktop and disappears on the phone, for the search engine it does not exist.

The mobile basics are:

  • The viewport meta tag, which tells the browser to use the width of the screen:

    <meta name="viewport" content="width=device-width, initial-scale=1">
    
  • The same content, the same structured data and the same links in both versions.

  • Text you can read without zooming, and buttons you can tap without hitting the one next to them.

How to check. The free check confirms that the page declares a viewport. For the rest, open the page on a phone, or in the device mode of the browser tools, and actually use it. PageSpeed Insights also runs its analysis on the mobile version.

Speed and Core Web Vitals

Speed counts twice. For the search engine, which uses page experience as one of its signals, and for the person, who gives up on a slow page before reading it.

Google measures experience with three numbers, the Core Web Vitals:

Metric What it measures Good
LCP how long the largest element on screen takes to appear up to 2.5 seconds
INP how long the page takes to react to a click or tap up to 200 milliseconds
CLS how much the page jumps around while loading up to 0.1

"Good" applies to 75% of real visits, not to a single test. INP replaced FID in March 2024, so older material that talks about FID is out of date.

The usual suspects: large uncompressed images, fonts that load late and push the text around, heavy third-party scripts at the top of the page, and ads or banners that arrive late and shift everything.

How to check. PageSpeed Insights shows the three numbers with real-visit data, when the site has enough traffic, and with lab data always. The Core Web Vitals report in Search Console groups pages by status. The SEOWrench free check measures server response time and HTML size, which are part of the picture, but it does not measure Core Web Vitals. For those, PageSpeed Insights is the source.

Structured data

Structured data is a description of the page in a format a machine reads without guessing. Instead of the search engine inferring that a piece of text is a price, you state that it is a price. The most common format is JSON-LD, a block of code inside the HTML, using the schema.org vocabulary.

It does two jobs. It can make a page eligible for rich results on Google, like stars, FAQs and breadcrumbs. And it helps any machine understand the page, including AI search engines, which need to know who wrote it, when, and what it is about.

What goes wrong is declaring what the page does not show. Marking up a rating that does not exist, or an FAQ that is not in the text, is the kind of thing Google penalizes.

How to check. Google's Rich Results Test tells you whether a page's structured data is valid and which rich results it is eligible for. The free check warns you when the page has no structured data at all.

hreflang, for sites in more than one language

If your site has the same page in English and in Portuguese, hreflang tells the search engine which version to show to whom. Without it, a Brazilian reader may land on the English version, and the two versions compete for the same search.

The rules that break most often:

  • It has to be reciprocal. If the English page points to the Portuguese one, the Portuguese page has to point back. Without the return link, Google ignores the pair.
  • Each version also points to itself.
  • The language code has to be valid, like en and pt-BR.
  • A default version, the x-default, for visitors who fit none of them.

How to check. This is one item the free check does not cover. Search Console no longer has a dedicated report for it either. The way to check is to open the page source of one page in each language and read the tags by hand, or to use a third-party tool that crawls both versions. Google's documentation on localized versions has the examples.

JavaScript and rendering

Many modern sites build the page in the browser, with JavaScript. The HTML that arrives from the server is nearly empty, and the content appears after the script runs.

Google can run JavaScript, but it does so in a second step, which can take a while, and not everything survives it. And most AI search bots do not run JavaScript at all. For them, the page is the HTML that arrived, and if the text is not there, it does not exist. A page can rank on Google and be invisible to ChatGPT. The post on how SEOWrench measures GEO shows what each engine can read.

The fix is to deliver the main content in the HTML that comes from the server, with server-side rendering or prerendering. JavaScript stays for interaction, but the text no longer depends on it.

How to check. URL Inspection in Search Console has an option to test the live URL and see the HTML Google rendered. To see what a bot without JavaScript receives, open the page source (not the inspector, which shows the page already built) and look for the main text. The free check does this for you and tells you whether the page content is readable without JavaScript.

Where to start

The list above is long, and not everything weighs the same. The order that pays off most is the search engine's own: first what stops it from arriving, then what stops it from reading, and only then what improves.

  1. What stops it from arriving. A noindex by mistake, a robots.txt blocking what it should not, an important page returning an error.
  2. What stops it from reading. Content that only appears with JavaScript, a canonical pointing to the wrong place, redirect chains.
  3. What improves. Title and description, headings, structured data, speed.

The first group is usually small and cheap to fix, and it is the one that changes the outcome the most. A single forgotten noindex tag is worth more than a month of fine-tuning titles.

The SEOWrench free check reviews a page in seconds, with no account, and sorts what it finds by severity: what is critical, what is a warning and what is an opportunity. For the whole site, the free account audits up to 5 pages per scan. What it does not cover, like Core Web Vitals and hreflang, is stated in the sections above, along with the Google tool that does.

Sources

  1. Google Search Central: how Google Search works
  2. Google Search Central: introduction to robots.txt
  3. Google Search Central: sitemaps overview
  4. Google Search Central: block indexing with noindex
  5. Google Search Central: HTTP status codes and network errors
  6. Google Search Central: redirects and Google Search
  7. Google Search Central: consolidate duplicate URLs
  8. Google Search Central: mobile-first indexing
  9. Google Search Central: page experience
  10. web.dev: Web Vitals
  11. Google Search Central: introduction to structured data
  12. Google Search Central: localized versions of your page
  13. Google Search Central: JavaScript SEO basics
Keep reading
Audit your page for freeAnalyze →