Fixed on this site

When to noindex a page: my résumé failed my own test

On the day I published a study of pages crawlers cannot read, the page failing it was mine.

On 20 September I published a study of 50 AI-built websites, measuring how much text each one sends to a crawler that does not run JavaScript. Under 50 words, I called a page invisible.

The same day, an audit of my own site found a page sending 18. It was my interactive résumé, the one made to look like an old desktop computer.

Before and after

This site, audited 20 September 2026

Before
18 words to a reader that does not run JavaScript: my name, a tagline, a loading message and a clock. The page could be indexed, had no canonical tag, and was linked from my About page.
The line I had published
Under 50 words counts as invisible in my own 50-site study.
After, checked 21 September
The page carries noindex, follow, and a canonical tag pointing at itself.

How a page can look full and send almost nothing

The résumé is an experience: windows, icons, a start menu. All of it is drawn by JavaScript after the page arrives. Before that script runs, the page is a loading screen.

Google runs JavaScript, so this was never a simple indexing block. But Google is not the only reader, and its own documentation says so:

Keep in mind that server-side or pre-rendering is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript.

Google Search Central, JavaScript SEO basics

A crawler that does not run JavaScript gets the loading screen and nothing else.

The 50-site study: what AI crawlers actually get from AI-built sites

Two honest fixes, and why I picked noindex

There were two ways to close this. One was to put the résumé's text into the page itself, so it arrives before any script runs. The other was to ask search engines not to index the page at all.

I chose noindex. The page is an experience, not something anyone needs to land on from a search. What noindex does, in Google's words:

noindex is a rule set with either a <meta> tag or HTTP response header and is used to prevent indexing content by search engines that support the noindex rule, such as Google.

Google Search Central, blocking indexing

The page also got the canonical tag it had been missing, pointing at itself.

The trap: noindex and robots.txt together

The most common way to get this wrong is to block the page in robots.txt as well, to be safe. That undoes the noindex, because Google has to be able to read the page to see the rule:

For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.

Google Search Central, blocking indexing

So my résumé is noindexed and not blocked. Crawlers can reach it, and Google is asked not to list it.

When a page should be noindexed

My rule of thumb, which is judgement rather than a Google rule: noindex a page when being found in search would not help the person who lands on it. Thank-you pages, internal search results, test pages, and experiences like mine that make no sense as a first visit.

Do not noindex a page you want found just because it is thin. Fix the page instead.

Check my fix yourself

  1. Open my interactive résumé. The desktop résumé
  2. Press Ctrl+U (Cmd+Option+U on a Mac) to see the page source.
  3. Press Ctrl+F and search for noindex. You will find content="noindex, follow".

Check your own site, no code needed

  1. List the pages on your site that are experiences rather than reading: quizzes, calculators, animations, booking widgets, logged-in areas.
  2. Open each one and press Ctrl+U to see the source, then Ctrl+F to search for a sentence you can see on screen. If it is not in the source, a crawler that does not run JavaScript never sees it.
  3. For each of those pages, decide whether people should find it from search. If yes, its text needs to be in the page source. If no, it is a noindex candidate.
  4. If you have Search Console, the Pages report under Indexing shows which pages Google has indexed. Any page there you did not expect is worth the same question.

What to ask your developer

  • Which of our pages load their text with JavaScript after the page opens?
  • For each of those, can the main text be in the HTML from the start, or should the page carry noindex?
  • Is any noindexed page also blocked in robots.txt? It should not be.

What this does not show

noindex is a rule for search engines that support it, and Google says it does. I have not tested which AI crawlers follow it, so I am not claiming it keeps this page out of AI answers.

The 18 words are what arrives before any script runs. Google, which does run JavaScript, may have seen far more.

Want your site checked the same way?

Send a URL. Three to five verified findings come back within two working days, free, each with a way to check it yourself.

Send me your URL
Human Machine