Between the 15th and the 18th of September I crawled 44 startup websites from the outside and counted the most common technical SEO issues on them. No Search Console, no analytics, no login: just the pages anyone can fetch. It came to 14,023 pages and 9,705 internal links.
Most SEO advice aimed at companies this size is about strategy: thin content, the wrong keywords, a topic map, the sort of thing that fills a deck.
That is not what I found broken.
What is broken is plumbing. Links that point at pages that no longer exist. Page templates that forgot the heading. Descriptions and titles copied from page to page, hundreds of times on some sites. Four sites out of 42 came through every check clean. The median site had four different problems at once.
None of it is hard to fix. Almost all of it is invisible unless somebody goes looking.
What I measured, and what I could not
Every site got the same treatment: read the sitemap, fetch the pages, parse each one, then follow every internal link and record where it actually ended up.
From the outside you can check a lot: status codes, titles, descriptions, headings, canonicals, robots directives, structured data, word counts, link targets, redirect chains. You cannot check which canonical Google actually picked, what a page ranks for, or real-world loading data from actual visitors. Those need the site owner's own tools.
So this post is about the half you can verify without asking anyone's permission — which happens to be the half that is usually broken.
Two of the 44 sites returned too little to judge: one had no usable sitemap, one gave me a single page. They stay in the data as excluded rather than quietly disappearing. Every percentage below is out of the remaining 42.
What was actually wrong
Each bar is the share of sites where the problem appeared at least once. Two more sites were crawled but returned too little to judge.
Nothing here is exotic. These are the first things any audit should look at, and most of them are one template defect repeated across hundreds of pages.
Links that go nowhere, and links that take the long way round
26 of 42 sites had at least one broken internal link. Not a broken link somewhere on the internet — a link on their own site, pointing at their own site, landing on an error. The median site had 7 of them. Four sites had 50 or more.
32 of 42 had internal links going through a redirect. That is the single most common finding in the whole study. The median affected site had 16 redirecting link targets, and three sites had 100 or more, which almost always means a migration whose links were never updated.
A redirect is not a crisis. But a link that redirects today is a link that 404s after the next site rebuild, and when the link lives in a navigation bar or a footer, one template puts it on every page you have.
Pages that forgot they were pages
24 of 42 sites had pages with no H1 at all. The median was 8 pages. Then look at the spread.
The median site
8 pages with no H1. A handful of templates that somebody forgot, usually a landing page or a jobs listing.
The most affected sites
Five sites had 100 or more pages with no H1, and on four of them that was most of the pages crawled. Not hundreds of mistakes. One template, shipped hundreds of times.
One page template, hundreds of pages. That is why these fixes are cheap.
That gap, between 8 pages and hundreds, is the most useful thing in this study. These are not hundreds of separate mistakes made by hundreds of careless people. It is one template — a jobs listing, a customer story, a docs layout — shipped hundreds of times with the heading missing.
Which is exactly why this work is worth doing. A finding that affects hundreds of pages usually has a one-line fix. I found the same shape on my own site: one missing line in one template left 17 of 18 pages without a share image.
The same pattern runs through the rest: 25 of 42 had duplicate titles, 26 of 42 had duplicate meta descriptions. On eight sites, 100 or more pages shared their description with another page, and on five the same was true of titles. 25 of 42 had pages under 150 words.
Canonicals
18 of 42 sites had pages with no canonical tag at all. 12 had canonicals pointing at a different page — I compared with trailing slashes, http versus https and www normalised away, so /about and /about/ do not count as a mismatch.
This one is worth caring about more than the headings. A canonical pointing at the wrong page is an instruction to Google to index something else instead of this page. When it is wrong, it is wrong loudly.
Structured data
Three sites had no JSON-LD anywhere. Five had JSON-LD that does not parse.
The five are the interesting ones. Invalid markup is not partially read — it is discarded. Those sites did the work, shipped the code, and get exactly nothing for it, while the report says "structured data: present". A single unescaped line break inside an FAQ block is enough.
A further 24 sites had structured data on some pages but not others, which is usually one page type that was never included.
The number I nearly published
Here is the part I would rather not write, and the reason I am writing it anyway.
My first pass through this data said half the sites had a sitemap listing URLs that do not return 200. It was a good line. It was in my draft.
Then I did the thing I make myself do before anything goes out: I picked a sample of findings and re-fetched them live, days after the original crawl. Of 40 sampled sitemap URLs, 29 came back perfectly fine.
The crawler had filed every sitemap URL whose first response was not a 200 under one heading — which lumps a permanent redirect in with a dead page. Most of them were redirects. Real, worth reporting, but a completely different finding from "your sitemap lists dead pages".
Split properly, the numbers are: 19% of sites list a URL that ends in an error, and 45% list a URL that redirects. The 50% I almost published would have been wrong.
The crawls ran between 15 and 18 September. I re-fetched a sample on 20 September, because a finding that has already been fixed is not a finding.
The one broken link that came back 200 had been fixed in the four days between the crawl and the re-check.
I also re-derived every number in this post a second way, from the raw per-page records rather than the crawler's own summary. That comparison covered 484 numbers and 23 disagreed — mostly one field that counts images on some sites and pages on others, which I dropped and recomputed.
None of that makes me clever. It makes the number safe to show you, which is not the same thing. Re-running the same script would have told me nothing; it agreed with itself the first time.
The sites that came through clean
Six of the 42 had none of: broken links, dead sitemap URLs, missing H1s, missing titles, missing canonicals or invalid structured data. Widen it to all nine checks — adding duplicate titles, missing descriptions and near-empty pages — and it drops to four.
That drop is worth a sentence, because it depends on which checks you count. A site can pass all of the first six and still have a large number of pages whose canonical points somewhere else, which none of those six looks at. Clean is a narrower word than it sounds.
The four are not the biggest sites in the study, and they are not all developer tools. What they have in common is duller than good SEO: their page templates are consistent. Every page type carries a heading, a description and a canonical pointing at itself, because the template guarantees it and not because somebody remembered.
That is the actual lesson. Nobody hand-fixes hundreds of pages. They fix the template, and the pages follow.
Check your own site in five minutes
Three checks, in the order I would run them:
1. Do your internal links land? Open your main navigation and footer in a browser with the network tab open, and click through. Any 301 hop or 404 in there is on every page of your site.
2. Does every page type have an H1? Pick one URL of each kind — homepage, product, pricing, blog post, docs page, careers listing, customer story — open it in Chrome and press Ctrl+U (Cmd+Option+U on a Mac) to see the raw HTML. Then press Ctrl+F (Cmd+F on a Mac) and search for <h1.
No match means that whole template is missing one. One is what you want. More than one is worth a look.
3. Does the canonical point at itself? In the same view, search for rel="canonical" and read the address after href.
If it names a different page, you are telling Google to index that page instead.
If any of those surprise you, the same thing is probably true across a whole template rather than on that one page.
What I would not do with these numbers
I would not put "62% of sites have broken links" in a pitch deck as an industry statistic. It is not one. It is what 44 particular sites looked like on four particular days: heavily developer tools and SaaS, founder-led, mostly on Next.js, Webflow or Framer.
Two more limits worth stating. Most of the crawls were capped at 250, 500 or 800 URLs, so every "share of pages" figure is a share of what I crawled, not of the whole site. And I fetched pages without running JavaScript — which is the right way to see what a crawler receives, and the subject of the last thing I published, but it is not what a visitor sees.
None of this says these sites rank badly. I have no ranking or traffic data for any of them and I am not claiming any.
What it does say is that the boring checks are where the findings are. Not the strategy. The plumbing.
The data
The site-by-site data behind this post is not public, on purpose.
Every row carries exact counts, and exact counts can make a site recognisable to the people who run it, even without its name. What is published is the method, written out above, and the live re-check in the chart.
If a number here looks wrong to you, tell me, and I will re-check it the same way.
If you want to know what your own site looks like from the outside, send me the URL. You get the findings whether or not you ever work with me, and each one comes with a way to check it yourself.
Edited 21 September 2026: figures that described a single site have been replaced with counts across several sites, so no single site can be picked out. No percentage changed.
Frequently asked questions
Are these percentages typical of all websites?
No, and I would not quote them that way. This is what 44 specific sites looked like: mostly developer tools and SaaS, founder-led, on Next.js, Webflow or Framer, because that is who I audit. Read them as what to check first, not as an industry average.
Can you really audit a site without Search Console access?
You can check a great deal of it. Everything in this post came from public pages. What you cannot see from outside is which canonical Google actually chose, what queries a page ranks for, and real-world Core Web Vitals. Those need the client's own data, and a good audit says so rather than guessing.
Is a redirecting internal link really a problem?
It is a small one, and it is the most common thing I find. Each hop is a wasted request, and a link that redirects today is often a link that 404s after the next migration. It matters most in navigation and footers, where one template puts the same link on every page.
Does a missing H1 hurt rankings?
On its own, probably not much, and I have no ranking data here that would let me claim otherwise. I report it because pages without an H1 are usually pages nobody owns: they tend to be missing the description and the canonical too. It is a symptom worth following rather than a crisis in itself.
Why not name the sites?
Because the point is the pattern, not the company. Every figure here is a share of sites, a median or a count across several sites, so no single site can be picked out, named or not.
