Knowledge base / Start here
How the engines decide who to cite, and how much of it is actually known.
Some of this is documented by the companies themselves. Some is inferred from watching thousands of answers. Some is guesswork being sold as fact. This page separates the three.
Key takeaways
- Do the engines publish how they choose sources? Partly. Google documents what makes a site eligible for AI features and says there is no separate ranking system for them. The others describe their retrieval in general terms and publish no selection criteria.
- What is the one thing they share? They retrieve before they generate. Each one searches an index, pulls a handful of pages and writes from those, so being in the index being searched is the first requirement.
- Which index does each one use? Google’s AI features use Google’s own. ChatGPT’s web results lean on Bing. Copilot is Bing. Perplexity runs its own crawler alongside third-party sources.
- What is still unknown? The weighting. Nobody outside these companies knows how much brand mentions, freshness, schema or domain history each count for, and anyone quoting you a percentage is estimating.
What do all four engines have in common?
They retrieve before they generate. Each one runs a search against an index, pulls a small set of pages, and writes an answer from what it pulled. The model is not recalling your website from training. It is reading it at the moment of the question.
That single fact decides most of the practical work. If your pages are not in the index being searched, nothing about your writing matters. If they are in it but the answer is buried in paragraph four, the engine summarises someone whose answer was in paragraph one.
Which index does each engine search?
| Engine | Where its web results come from | What that means for you |
|---|---|---|
| Google AI Overviews and Gemini | Google’s own index. | Your Google ranking and indexing work carries straight over. See AI Overviews. |
| ChatGPT | Bing, for its web browsing results. | Bing Webmaster Tools verification and a submitted sitemap are the entry ticket. See ChatGPT SEO. |
| Copilot | Bing. | Same work as ChatGPT, with Microsoft’s own surfaces on top. See Copilot and Bing. |
| Perplexity | Its own crawler, alongside third-party search sources. | Make sure your robots.txt is not blocking crawlers you meant to allow. See Perplexity SEO. |
This is the layer worth acting on, because it is verifiable. You can check whether your pages are in each index yourself today.
What is documented, and by whom?
Google publishes guidance for its AI features in Search. The part worth reading twice is that there is no separate set of ranking signals for them: the same work that makes a page eligible for ordinary results makes it eligible here. Google asks for unique, satisfying content, working technical basics, and markup that reflects the page.
The others publish far less. OpenAI, Microsoft and Perplexity document how their crawlers identify themselves and how to allow or block them, which is operationally useful and says nothing about selection. None of them publishes criteria for which of ten retrieved pages gets named.
So the honest state of knowledge is: the eligibility layer is documented, the selection layer is not.
What is inferred rather than known
These are patterns many people have observed, including us, which no engine has confirmed. We act on them because they are cheap and consistent with the documented layer, and we label them as inference rather than fact.
- Pages that already rank well are quoted more often than pages that do not. Consistent with retrieval-then-generation, and with Google saying its AI features use the same systems.
- Self-contained paragraphs get quoted more than paragraphs that depend on the one above them. This is what you would expect of any extraction step.
- A business described the same way across the web gets described more accurately than one described six different ways. See NAP consistency.
- Pages with accurate structured data do not automatically win, but pages whose markup contradicts their visible text seem to do worse.
What nobody outside these companies knows is the weighting. If you are shown a figure for how much any single factor counts, ask how it was measured. Usually it is a correlation across a sample of answers, which is interesting and is not the same thing.
What we do with the uncertainty
We work the documented layer hard and treat the rest as probability. Rank the page, get it indexed in Google and Bing, put the answer first, name the entities, keep the schema honest. Every one of those helps under any weighting, and none of them is wasted if the systems change.
We also tell you when a result is a sample rather than a measurement. Ask an assistant the same question twice and you may get two different source lists, so a single screenshot proves very little. That is why our free score reports what it read on your pages rather than claiming a position in a system we cannot see, and marks anything it could not measure as not measured.
Questions about citation
Not in the organic citations. Paid placements inside AI products exist and are labelled as ads where they appear. Anyone offering to get your brand into the organic answer for a fee is either describing ordinary optimisation work or selling something that does not exist.
Usually the index. A site well indexed in Google and thinly indexed in Bing will show up in AI Overviews and not in ChatGPT, which is the single most common split we see. Checking site: counts in both engines takes two minutes and explains most cases.
No. Ask the same question twice and you can get different sources, because these systems retrieve fresh each time and some of them personalise by location and account. Treat any single result as one sample, not a ranking.
They appear to matter, and that is an honest “appear”. A model that has seen your business described consistently across the web has more to draw on than one that has only seen your own site. How much weight it carries is not documented by anyone.
Several exist, and the thing to understand before buying one is that they all work by asking the engines questions and recording the answers, which is what you can do manually. The value is in doing it consistently against the same question set, not in any privileged access.
Work the parts that are documented and cheap: rank, get indexed in both engines, answer the question in the first paragraph, name your entities, keep your schema accurate. Those help whatever the weighting turns out to be. Skip anything whose only justification is a confident claim about a system nobody outside the company can see.
Find out whether the engines can read your pages at all.
Before worrying about weighting, check the parts that are documented. The free AI Visibility Score reports eight of them on your own site.
Get your AI Visibility Score