Skip to content

What Language Models Actually Read on a Web Page

A regional dental chain runs forty near-identical pages, one per clinic. Every page carries the same header — full service list, a locations dropdown, a phone number — and the same three-column footer, repeated word for word from page to page. Strip that shared scaffolding away and what’s left, on more than a few of the forty, is an address and a sentence about parking. To a visitor scrolling past the header on the way to what they came for, the page reads as normal. To something that has to decide what’s actually distinct about clinic #23 versus clinic #24, the header and footer are most of what it has to work with, and they’re identical everywhere.

What does “content” mean to something reading a page automatically?

A browser doesn’t need to ask this question. It renders the whole document and a person’s eye does the filtering, skipping the header and footer without being told to. Something parsing a page for its distinct information has no eye to do that filtering for it. It has to decide, from the HTML structure alone, which blocks are the site’s furniture — present on every page, and therefore uninformative about this particular one — and which blocks are the substance a person actually came for.

That separation has a name in information retrieval: boilerplate removal, or main-content extraction, and it’s been a research problem for more than twenty years, refined across a documented line of algorithms built specifically to draw that line reliably across arbitrary page templates. It’s not a solved problem so much as a well-studied one — different tools draw the boundary slightly differently, and a page whose main content sits inside an unusual markup pattern can still confuse a parser tuned on more conventional layouts.

How much of a typical page is furniture, not content?

Enough that it’s worth measuring rather than assuming. One audit check built to test AI visibility does exactly the separation described above, on every page it crawls: it strips the header, navigation, footer and sidebar regions, counts the words left over, and compares that figure against how many words sat inside the parts it just removed. When the repeated chrome adds up to more than roughly sixty percent of everything on the page, the check flags it — not because navigation is bad, but because a block of text that’s identical across forty pages can’t be what makes any one of those forty worth reading over the others.

Building all forty pages off one shared template is the obviously sensible choice — nobody hand-writes a new header for every clinic, and a visitor benefits from finding the same menu in the same place regardless of which location page they landed on. Run the check against those pages anyway, and most of the forty would likely fail it: a couple of sentences of real, page-specific text sitting inside several hundred words of header and footer that never change. The consistency that makes the template a good idea for a visitor is the same thing that makes the forty pages read as close to identical at the level of raw text, each with a small, page-specific footnote at the end.

Does a search engine have the same blind spot?

Less of one, and the reason is worth being specific about. A search engine indexes a whole site over time and can lean on signals no single fetch provides on its own — how a page’s structure compares to the rest of the domain, how its links behave across the crawl, patterns learned across millions of other sites with the same templating problem. A tool composing an answer from a handful of freshly fetched pages doesn’t get that aggregate history. It’s closer to working page by page, with whatever separation its own parser manages to draw on that one document, in that one pass.

This is inference, not something publicly documented by any AI vendor — none of the major providers publish exactly how their systems weigh boilerplate against main content once a page has been fetched. What’s measurable is the input each kind of system gets, not the internal weight it assigns once it has that input.

Only if what’s underneath grows to match. A page with three menu items and two sentences of substance still has almost nothing distinct to offer — cutting the menu down doesn’t manufacture the missing content share, it just makes a thin page look slightly less thin by percentage. The fix that actually moves the number is the harder one: writing more of what makes that specific page different from its forty siblings, not less of what makes them the same. A duplicate-content check catches the sibling problem from another angle — pages that restate the same paragraph under a different heading compete with each other for the same citation regardless of how their navigation is trimmed.

What does this look like on a real site, and where does it stop being measurable for free?

The check itself needs pages to run against, and how many pages it sees depends on what kind of scan requested it. A free scan reads the homepage; a full crawl across a templated set of pages — the case where this problem actually shows up — is part of the paid tier. A single homepage rarely has the repeated-template problem in the first place, since there’s nothing else on the site to repeat against; it’s the location pages, the category pages, the city-by-city variants that this measurement is built to catch, and catching them means crawling more than one URL. The audit methodology lists what each tier actually fetches. A page-level extractability check covers a related question — whether a single page’s paragraphs are shaped to be quoted — without needing the rest of the site.

None of this proves what a given model does with the ratio once it has the page. What’s checkable is the shape of the input a template-heavy site hands over: mostly identical scaffolding, with a page-specific remainder that, on badly templated sites, turns out to be thinner than it looks in a browser.

Frequently asked questions

What's the difference between how a browser shows a page and how something reading it automatically sees it?

A browser presents a page as one visual unit — header, navigation, sidebar, footer and body copy all in their place, so the eye skips straight past the repeated parts. Something parsing that same page for its distinct content has no visual layout to lean on. It has to work out, from the markup alone, which text is furniture repeated on every page and which text is actually about this one.

How much of a typical page is navigation and footer rather than content?

It varies by site, but on pages built from a shared template — location pages, category pages, product listings — the repeated header and footer can outweigh the unique text by a wide margin. One audit check separates the two and flags a page once the repeated chrome makes up more than roughly sixty percent of everything on it.

Does a traditional search engine have the same blind spot?

Less of one. Separating a page's main content from its boilerplate has been a studied problem in information retrieval for over two decades, and search engines have had that long to build aggregate signals — across a whole crawl, a whole site, a whole link graph — that dilute a single page's noise. A tool composing an answer from one fetched page at a time doesn't get the same aggregate cushion.

Does trimming navigation text fix this by itself?

Not if the content underneath it stays thin. A page with a short, tight menu and two sentences of substance still has almost nothing distinct to offer. The fix that actually moves the ratio is writing more of what makes each page different, not just less of what makes them the same.

Can this be checked without paying for a full audit?

Partially. The underlying check runs against whatever pages are available — a full site crawl on a paid audit, or just the homepage on a free one. A free scan gives a read on the homepage alone; seeing the pattern across a whole set of templated pages needs the full crawl.