Site structure for SEO: pages, URLs and internal links

Alexandre Silva Anjos
Alexandre Silva Anjos · Founder of SEOWrench

Seventeen years building software, now applied to SEOWrench.

A page nobody links to is, as far as a search engine is concerned, a page that does not exist. It can be published, well written and fast. If no path leads to it, the crawler never arrives, and what is not crawled is not indexed.

Site structure decides three things: which of your pages a search engine gets to see, how often it comes back to them, and which page it treats as the main one for each topic. It is organisation work, not design work, and it is usually the cheapest fix available, because it does not require writing new content.

The free check counts how many internal and external links one of your pages has, in seconds, with no sign-up.

This guide is practical. It applies to a company site, a blog and a store, and it ends with a checklist you can run today.

What site structure is

Site structure, also called site architecture, is how your pages are organised and connected. It has three layers, and almost every problem lives in one of them:

  • Hierarchy. Which pages exist and what sits under what. The org chart of the site.
  • Navigation. How a person reaches each page. Menu, footer, internal search, breadcrumb.
  • Internal links. How pages cite each other inside the text, outside the navigation.

One detail trips a lot of people up: the URL hierarchy and the navigation hierarchy do not have to be identical, and often are not. What matters is that both tell the same story.

Site structure is not how the menu looks, how many pages you have, or which CMS you run. You can have excellent structure in plain HTML and terrible structure in an expensive site.

Why structure decides what gets indexed

Googlebot discovers URLs in two ways: following links and reading the sitemap. Both are necessary, and they do not do the same job.

The sitemap is a list. It says "these URLs exist". A link is a path, and it says a lot more: it says another page on your site found that page relevant enough to cite, and it says, through the anchor text, what the page is about.

That is why a sitemap does not fix bad structure. A page that only exists in the sitemap tends to be crawled, indexed with low priority and forgotten.

When a site earns links from elsewhere, that authority does not stay stuck on the homepage. It flows through the site along internal links. A page one click from the homepage, cited by ten other pages, gets a far bigger share than a page that only shows up at the end of a pagination.

In practice this gives you real control over which page ranks. You do not have to wait for an external link to promote your most profitable service page: your own relevant pages pointing at it, with descriptive anchors, already do a lot.

Crawling has a cost

If your site has fifty pages, crawl budget is not your problem, and ignoring the subject is the right call. If it has fifty thousand, with filters generating combinations, the crawler starts spending time on URLs that do not matter and takes longer to come back to the ones that do.

The arithmetic is simple: every crawlable address that should not exist eats the time of one that should.

The shape: flat and deep

The best structure for search is usually flat in navigation and deep in content. Few layers between the homepage and any page, and plenty of material inside each topic.

FLAT: 3 layers Home /shoes /shoes/running

DEEP: 5 layers to the same page Home /products /footwear /shoes /shoes/running

The same destination page, reached two ways. Every extra layer is extra dilution.

Each extra layer does two bad things at once. It pushes the page away from the authority that enters through the homepage, and it adds one more point where the crawler can simply stop.

The three-click rule is a symptom, not a law

You will hear that no page should be more than three clicks from the homepage. There is no rule in that shape, and there is no Google documentation establishing that number.

The number started as a web design guideline in the early 2000s, in the spirit of "visitors give up if it takes too long", and moved into SEO writing by repetition. Later usability testing found no relationship between the number of clicks and whether people succeed.

What is still useful is the symptom. When a page sits six clicks away, it is almost always because it got buried in a badly planned hierarchy or at the end of a long pagination. The burial is the problem, and the click count is only the thermometer.

Use the number as a question, not as a target: if this page is far away, why is it far away? Sometimes the answer is fine, and the page really is secondary. Usually the answer is that a link is missing.

Silos: grouping by topic

A silo means grouping the pages about one topic and making them link to each other, with a main page that represents the theme.

On a blog, a silo is a theme with several posts and a central guide. In a store, it is the category with its products. The idea is that a search engine can see the set and understand that you have depth on that topic, instead of finding five loose texts.

⚠️ The common mistake is the rigid silo: forbidding a page from linking outside its group. That is an imitation of an old theory and it makes the experience worse. If a post about site structure needs to cite what technical SEO is, it should, even if the two sit in different groups. The reader comes before the geometry.

URLs: the pattern that ages well

A URL is one of the few SEO elements that is expensive to change later. It is worth getting right early.

Avoid Prefer
/p?id=4482 /shoes/running
/2024/03/12/how-to-choose-running-shoes /guides/how-to-choose-running-shoes
/Products/Running_Shoes /shoes/running
/category/sub/sub/sub/product /shoes/running/model-x
/page-1, /page-2 words that describe the page

The rules that settle most cases:

  • Lowercase always, with hyphens between words. Underscores work worse, and a capital letter creates two URLs where one should exist.
  • Words instead of numbers. An ID says nothing to a reader and nothing to a search engine.
  • No dates, unless it is news. A date in the URL makes a 2024 guide look old in 2026, even after you update it.
  • Shallow. If you can cut a level without losing clarity, cut it.
  • Stable. A good URL is one that does not need to change when the menu changes.

Changing a URL almost never pays off

When a URL changes, you lose part of the history it accumulated, even with a correct 301 redirect. External links keep pointing at the old address, and the search engine takes time to move the signals across.

That was exactly the decision we made about the post you are reading. Its address says site-architecture and the text says site structure. Changing it would make the URL prettier and throw away the impressions it had already earned. It stayed as it was.

Change a URL when the address is genuinely wrong: a topic that no longer matches the content, a typo, the wrong language. Do not change it for taste.

The menu is not the index of the site. It is the path to what most people are looking for, and stuffing it with everything that exists makes navigation worse for everyone.

What tends to work:

  • A short menu, with the main categories and nothing else.
  • A footer for the institutional pages, and for the topics that did not fit in the menu.
  • A real category page for each group, with its own text, not just a list of links.
  • Breadcrumbs on every internal page, marked up with BreadcrumbList in structured data. They help a person move up one level and they tell a search engine where the page lives.

Internal site search deserves a warning: it is great for visitors and terrible for the index. Search result URLs should not be crawlable, because they are infinite and have no content of their own.

The clickable text is what tells a search engine what the destination is about. "Click here" says nothing, and wastes the link.

Weak anchor Anchor that works
click here site structure guide
learn more how to choose running shoes
this article technical SEO checklist
read free technical SEO test

Three habits that solve most of it:

  1. Every new post cites at least two older posts. The easiest one to remember.
  2. Every relevant older post gets a link to the new one. Almost nobody does this, and it is what pulls new content out of isolation.
  3. The most important page on the site gets links from the most pages. If your service page is only cited by the menu, it is being treated as a footer item.

There is no magic number of links per page. There is judgement: one link that makes sense mid sentence is worth more than ten stacked at the bottom.

E-commerce: categories, filters and faceted navigation

Stores are where site structure breaks fastest, because the catalogue multiplies URLs on its own.

A category is a page, and it deserves its own content. /shoes/running should have text that explains the category, not just a product grid. That page is the one competing for "running shoes", not the individual product.

A filter is usually not a page. Colour, size and price range create combinations that multiply. Thirty combinable filters turn into thousands of addresses with the same content reordered, and that is how a small store shows up in the coverage report with tens of thousands of URLs nobody created on purpose.

The practical path:

  • Decide which facets deserve to be indexable by looking at real search demand. "Women's running shoes" usually has demand and deserves a page. "Blue running shoes size 41" almost never does.
  • For the ones that deserve it, give them a clean URL, their own title and their own text.
  • For the rest, use rel="canonical" pointing at the category and avoid linking them in a crawlable way.
  • Pagination with real links to pages 2, 3 and 4, so the crawler can reach the end of the catalogue. Infinite scroll with no links is the same as hiding half the store.
  • An out of stock product should not become a 404 if it is coming back. Keep the page, explain the state and offer alternatives.

AI crawlers follow the same logic

The bots that feed AI answers discover content the same way: following links and reading sitemaps. The difference is that they are less patient with pages that need JavaScript to show their text.

So the structure that helps Google helps that other layer of discovery too. It is worth checking whether they can reach your content: when we tested the biggest sites in the world, 29% were blocking those bots without realising it.

Checklist to run today

  • Is any important page far from the homepage in number of clicks?
  • Is there a page that is only reachable through the sitemap?
  • Do the URLs use words, lowercase and hyphens?
  • Does any URL carry a date, an ID or an unnecessary parameter?
  • Does each category have its own text, not just a list?
  • Are there breadcrumbs on internal pages, with structured data?
  • Do internal link anchors describe the destination?
  • Do older posts link to newer ones, and not only the other way round?
  • Is internal site search out of the index?
  • Does pagination have real links to the next pages?
  • Do filters without search demand point a canonical at the category?
  • Does the sitemap list the URLs that exist, and only those?

How to check this on your site

Much of that checklist can be done by hand on a small site. Past a few dozen pages, it cannot.

The free SEOWrench check crawls the site and returns, page by page, what is broken in title, meta description, H1, canonical, indexability and images without alternative text. On structure, it counts how many internal and external links each page has, and it flags a page with no internal links at all, which is the most common symptom of a forgotten page.

It does not, today, measure click depth or build a map of orphan pages across the site. If anyone promises that here, they are promising what the tool does not do.

Sources

  1. Google Search Central: managing crawl budget for large sites
  2. Google Search Central: consolidate duplicate URLs with canonical
  3. Google Search Central: breadcrumb structured data
Keep reading
Audit your page for freeAnalyze →