Case notes

How to get cited by AI search: I tested the advice against 603 real citations

What actually goes with being cited by AI search, measured rather than asserted, and why the usual comparison makes useless features look brilliant.

Chart showing a page ranked first is cited 63% of the time and a page ranked ninth 7%
On this page9 sections

Every second post about AI search tells you to add FAQ schema, write answer-first paragraphs and put a key-takeaways box at the top. I wanted to know whether any of that is really how you get cited by AI search, so I went and counted.

I asked Perplexity 30 business and technical questions, twice each, and recorded every source it cited. That came to 60 answers and 603 citations. Then I measured 438 pages — the ones that got cited and the ones that ranked but did not — with the same script, looking for the features everyone recommends.

The measurement in three numbers
603citations collected by hand from 60 Perplexity answers
438pages measured, cited and uncited, with the same script
82%of everything cited was a company blog or resource page

Two things predicted citation. Most of the advice did not. And one popular tactic was, if anything, slightly negative.

How I collected this

There is no public API for "who gets cited", so the citations were collected by hand.

I asked Perplexity the same 30 business and technical questions — what is kubernetes, sales pipeline stages, OKR examples, that sort of thing — logged out, with no personalisation, and copied every source from the answer's own source list. Then I asked all 30 again, minutes later, to see whether the answers were stable. That is the 60 answers and 603 citations.

For the comparison group I took the ordinary web-search top ten for those same questions. Then every page — cited and uncited — went through one measuring script that records the same features on each: table of contents, visible dates, author bylines, FAQ sections, schema types, word counts, where the question's words appear, and so on.

Which means the comparison is like for like. Nothing below rests on my impression of a page.

The biggest lever is the boring one

How often a page was cited, by its position in web search

30 questions, each asked twice. For every page in the web-search top ten, how often the AI answer cited it.

Position 1cited in 63% of answers
63%
Position 2
41%
Position 3
38%
Position 5
31%
Position 7
17%
Position 9cited in 7% of answers
7%

Ranking first does not guarantee a citation, and ranking ninth does not rule one out. But the gradient is steep, and it is the biggest single effect in the whole study.

A page ranked first in web search was cited 63% of the time. A page ranked ninth, 7%.

That is the strongest single effect in the study, and it is not a formatting trick. It is ordinary search ranking, which means the work that earns AI citations is mostly the work that earns rankings.

But ranking is not the whole story, and the exception is large: three-quarters of the pages Perplexity cited were not in the web-search top ten for that question at all. They were long guides — a median of 2,924 words — on sites that rank for the subject generally, even if that exact page did not rank for that exact query.

So both things are true. Position inside the top ten is a steep gradient. And most citations come from outside the top ten entirely.

One more number that changes who should bother: 82% of everything cited was a company blog or resource page. Not Wikipedia, not news. Wikipedia was 1%. If you sell something and write about it properly, you are the kind of source these engines are reaching for.

The mistake that makes bad advice look good

Here is how "add FAQ schema, get cited more" gets published in good faith.

The same data, compared two ways

Cited pages vs everyone else

Cited pages have an FAQ section 43% of the time. Uncited ones, 27%. Written up as a finding, that becomes: add an FAQ and get cited more.

the comparison that sells courses

Pages at the same rank

Hold ranking position constant and the FAQ advantage vanishes. So does most of the rest. Two things survive.

the comparison that survives

Compare winners with losers and everything the winners do looks like a cause. Compare pages that rank the same, and most of it disappears.

Compare the pages that got cited with the pages that did not, and cited pages look better at almost everything. They have a table of contents twice as often. They have FAQ sections more often. They are longer, fresher, better linked.

The problem is that cited pages also rank better — and ranking is what is driving most of the gap. You are comparing winners with losers and then crediting the winners' habits.

The fix is to compare pages that rank in the same position. Ask not "do cited pages have FAQs" but "between two pages sitting at position four, does the one with an FAQ get cited more often?"

Do that, and most of the advice evaporates.

What survived

What survived the controlled comparison

Odds of being cited, among pages at the same web-search position. Above 1 means more likely; the two in accent colour are the only ones that held up statistically.

Table of contentsheld up
3.5x
Visible "Updated on" dateheld up
3.1x
Named author bylinedirectional, just missed
2.8x
Query words in the first 120 wordsdirectional
2.5x
FAQ sectionno effect
1.7x
Word countno effect
0.9x
FAQPage schemano effect, if anything negative
0.7x

A table of contents and a visible updated date are cheap and one-time. Nothing here is worth buying as a package.

A table of contents. Pages at the same rank were 3.5 times more likely to be cited when they had one. It was the strongest page-level effect measured.

A visible "Updated on" date. 3.1 times. And it has to be real: 75% of cited pages had actually been modified within the last year.

Two more were suggestive but did not clear the bar once I corrected for testing many features at once: a named author byline (2.8×) and the question's own words in the first 120 words (2.5×).

Everything else — word count, images, reading time, key-takeaway boxes, question-shaped headings, author bio blocks, Article schema — came out at roughly no effect.

And FAQPage schema came out at 0.74. Below 1. Slightly negative, not statistically significant, which in plain English means: no evidence it helps, and certainly no evidence worth selling.

I want to be careful about what that does and does not mean. It does not mean FAQ sections are bad. Answer questions your customers actually ask. It means FAQ markup is not a lever on AI citation, and anyone charging for it as one is charging for a result this data does not support.

What the engines themselves say

Worth reading alongside the numbers, because it lines up better than you would expect.

Google's own documentation on AI features says there are no additional requirements to appear in AI Overviews or AI Mode, and no special optimisations necessary. To be eligible, a page must be indexed and eligible to appear with a snippet. That is it.

Perplexity's documentation says PerplexityBot exists to surface and link sites in its search results, and that it is not the crawler used to collect training data. Which means the robots.txt decision that keeps you out of AI answers is not the one most people worry about: blocking a training bot does not remove you from AI search. Blocking a search bot does.

Google also describes AI Mode issuing several related sub-queries for a single question. That rewards a page that answers a whole question space rather than one narrow query — which is consistent with the table-of-contents result, since a page big enough to need a table of contents is usually a page covering the whole space.

None of this contradicts the measurement. All of it points the same way: be reachable, rank, and be the complete current answer.

What I would actually do

In order, because the order is the advice:

  1. Check that AI search crawlers can reach you. This is free and it is the one that can be fatally wrong. Check your robots.txt for OAI-SearchBot, PerplexityBot and Claude-SearchBot by name — and check that your pages contain their text before any JavaScript runs, which is its own problem on modern sites.
  2. Pick the questions you should own and give each one page, using the question's own wording in the title, the H1 and the opening paragraph.
  3. Make that page the complete answer — and note that this is not the same as making it long. Cited pages were longer on average, but once I held ranking position constant, word count showed no effect at all. Length is what completeness looks like from outside, not the thing itself. Write until the question is answered and then stop.
  4. Add a table of contents. One-time, cheap, and the strongest thing measured here.
  5. Put a named author and a real "Updated on" date on it — and actually update the page. A stamped date on an unchanged page is a lie your readers can check.
  6. Do not pay for FAQ schema as an AI-citation service.

Where this could be wrong

I would rather say this myself than have it said to me.

It is correlation, not cause. Nobody added a table of contents to these pages to see what happened. A table of contents may simply be what big, maintained, well-resourced pages have. The honest claim is "goes with", not "causes".

One engine, one day. Perplexity, logged out, no personalisation, on the 19th of September. ChatGPT and Google AI Mode may weight sources differently. I measured whether answers are stable minute to minute — they are, 25 of 30 questions returned the same sources — but not whether they are stable week to week.

No authority measure. I run on free tools, so I have no domain-authority metric to control for. Ranking position partly stands in for it, imperfectly.

Thirty English questions, business and technical. Nothing here covers local search, India-specific queries, health or finance, or online shopping.

My table-of-contents detection undercounts the ones built by JavaScript — it missed about one in eight when I tested it. That biases toward missing the effect, so the real one is probably no weaker than 3.5×.

The data

All of it is public: github.com/ram15126/ai-search-citation-study

Every one of the 603 citations as collected, the measurement of all 438 pages, and the scripts. Two of them run on the published data alone, so you can check any table in this post in one command without refetching a thing.

The citation list was hashed inside the browser as I collected it and the saved file re-hashed on disk before I used it, so I know nothing was lost in transcription. If you re-run it and get different numbers, open an issue.

The part I keep coming back to

The advice that survived contact with the data is dull: be reachable, rank, be complete, be current. The advice that did not survive is the advice that comes in a package with a price.

That is not a coincidence. "Add this markup" is sellable. "Earn the ranking and keep the page current" is not, because it is work.

If you want to know whether AI search can actually read and cite your site, send me the URL. You get the findings either way, each with a way to check it yourself.

Frequently asked questions

Does FAQ schema help you get cited by AI search?

Not in this measurement. Among pages ranking at the same position, pages with FAQPage schema were cited slightly less often, and the difference was not significant. FAQ sections are worth writing when people actually ask those questions. They are not an AI-citation lever, and I would not sell them as one.

Does this mean ranking is all that matters?

Ranking is the biggest lever measured here, but it is not everything: three-quarters of the pages Perplexity cited were not in the web-search top ten for that question at all. They were long guides on sites that rank for the topic generally. Rank matters; being the complete answer matters too.

Is this true for ChatGPT and Google AI Mode too?

Unknown, and I will not pretend otherwise. This is Perplexity, logged out, on one day. Different engines pick sources differently. What is likely to carry across is the dull part: be crawlable, rank, and be the current complete answer.

Do AI answers change every time you ask?

Less than people assume. Asking each question twice, minutes apart, returned the same set of sources 25 times out of 30. So if a client's citations change for a fixed question, that is probably a real change rather than noise.

Does a table of contents cause citations?

I cannot say that, and the difference matters. Nobody added a table of contents to these pages to see what happened. It may simply mark a big maintained guide, and that may be what is really being cited. It is cheap enough to add anyway.

  • #ai search
  • #aeo
  • #perplexity
  • #technical seo

Want the findings for your site?

Send a URL. Three to five verified findings come back within two working days, free, each with a way to check it yourself.

Send me your URL