Field note SEO
SEO architecture at national scale
Geography, intent and templates as one information architecture.
At national scale, search stops being a content problem and becomes an information architecture problem that happens to be judged by a search engine. Sites at this size rarely fail because the writing is poor. They fail because they generate thousands of pages that answer the same question.
The grid, and its trap
The obvious structure for a service operating across a country is a grid: service by location. Multiply a handful of services by thousands of municipalities and you have a large site quickly.
The trap is that most cells of the grid contain nothing distinct. If the page for a service in one town differs from the next only by a name substituted into three sentences, the two pages are the same page, and generating five thousand of them creates a large set of near duplicates competing with each other.
The discipline is to only create a cell when something real distinguishes it: providers actually present, prices that genuinely differ, constraints specific to the place, demand that exists in the data. A grid filled where the data is real is smaller and works. A grid filled everywhere is larger and does not.
Intent is a second axis, and it is the more valuable one
Location answers where. Intent answers what the person is trying to do, and it segments demand far more usefully: comparing options, understanding a price, checking a requirement, solving a constraint, or ready to request a quote now.
Pages built on intent carry information that a template cannot fake, which is why they hold position and attract links. They also convert differently, and treating them identically in measurement hides which part of the architecture is working.
A practical structure: intent pages carry the substance and the authority, location pages carry the local specificity and the conversion path, and the internal linking between them is designed rather than emergent.
Templates as a data contract
A template at scale is a promise about what data every page of that type will have. If the promise is not enforced, pages ship with empty sections, placeholder phrasing, or silently omitted blocks, and the average quality of the type falls below the threshold where any of them rank.
Make the contract explicit. Each template declares the fields it requires. A page without them is not published rather than published thin. Publishing thin is the single most common cause of a large site plateauing, and it is invisible in aggregate metrics because it is distributed across thousands of pages.
Crawl budget is a real constraint at this size
Beyond a certain number of URLs, the practical question is which of your pages get crawled, how often, and whether the crawler spends its time on the ones that matter. Faceted navigation, sort parameters, session artefacts and pagination can generate far more URLs than the site has content, and every one of them consumes attention.
The controls are unglamorous and effective: canonical relationships that reflect the actual duplication structure, parameters excluded deliberately, sitemaps segmented by type so indexation can be measured per type rather than in aggregate, and internal linking that puts important pages within a short click depth.
Segmented sitemaps are the diagnostic that pays for itself. When indexation is measured per type, a failing template is visible in a week instead of a quarter.
Writing for retrieval, not only for ranking
A growing share of the traffic that matters arrives through systems that read a page to answer a question rather than to rank it. Those systems reward structure that a machine can extract: explicit entities, stated relationships, structured data that agrees with the visible text, and answers that are actually present rather than implied.
The overlap with good practice is large. Say the thing plainly, near the top, in the words the reader used. Mark up what the page is about. Keep the structured data honest, because a mismatch between markup and content is now checkable by the systems that consume both.
Measure the architecture, not the site
Aggregate traffic hides everything at this scale. The useful reporting is per template: how many pages of this type exist, how many are indexed, how many receive impressions, how many receive clicks, and what they convert at.
Read that way, decisions become obvious. A type with high indexation and no clicks has a relevance problem. A type with content and no indexation has a discovery problem. A type with clicks and no conversions is attracting the wrong intent. Three different diagnoses, three different fixes, all invisible in a single traffic curve.
Contact
Working on something in this territory?
Financial infrastructure, regulated systems, AI in controlled environments, cryptography, platforms at scale.