How AI Engines Pick Which Brands to Mention
Five signals decide which brands ChatGPT, Gemini, Claude, and Perplexity name in their answers: source pool, recency, citation density, schema clarity, and off-site presence. Here is the stack, ranked by leverage.

A buyer types "best small business CRM" into ChatGPT. The answer names three brands. There are forty real brands in that category, all serving small businesses well enough. Why these three, and not the other thirty seven? The mechanic is not random and it is not a popularity contest. It is a stack of signals AI engines weigh on the fly, every time, for every prompt. This post walks through the five signals that determine which brands get cited, with public data where it exists and clearly marked observations where it does not.
Key Takeaways
- The signal stack is five things, weighted by the engine: source pool composition, recency, citation density, schema and entity clarity, off-site presence.
- 76% of top AI citations in 2026 come from content updated within 30 days. A 2019 evergreen guide loses its citation slot to a 2026 refresh on the same topic (Omnibound, 2026).
- 68% of AI citations originate off-site. Reddit threads, YouTube transcripts, and review platforms often outweigh the brand's own blog (Omnibound, 2026).
- The signals compound. A brand winning on three of five signals will out-cite a brand winning on only one.
The buyer types a prompt. What does the engine actually do?
In the second or two between the user pressing enter and the answer appearing, the engine does four things in sequence.
First, it interprets the prompt, expanding the explicit request into the implicit one. "Best small business CRM" expands into something like "list 3 to 5 CRM products well-rated for businesses with 1 to 50 employees, with a one-sentence reason per pick." Second, it retrieves sources, pulling from the engine's index (Bing for ChatGPT and Copilot, Google for Gemini and AI Overviews, the engine's own index for Claude and Perplexity). The retrieval is broader than a Google search; an engine may pull 30 to 100 candidate sources before scoring them. Third, it ranks the sources, weighing recency, authority, citation density, and how directly the source answers the prompt. Fourth, it generates the answer, lifting passages, paraphrasing claims, and naming brands. Citation, if any, attaches at this last step.
Each of those four steps is where a signal can move your brand from "not retrieved" to "retrieved but not ranked" to "ranked and quoted." The optimization is not about gaming the answer. It is about being one of the sources the engine reaches for in steps two and three, in enough places that step four cannot avoid naming you.
Different engines weight the four steps differently. Perplexity makes the retrieval step deeply visible (it shows its sources). ChatGPT mostly hides it. Google AI Overviews surfaces sources only when there is one strong primary source. The optimization is the same regardless of how visible the chain is.
Signal 1: the source pool the engine drew from
The source pool is the universe of pages the engine considers candidates for a prompt. It is not the same as Google's top 100 results, and it changes engine to engine.
ChatGPT's pool leans heavily on Reddit, Wikipedia, YouTube transcripts, and a long tail of niche forums. Reddit grew especially important after the 2024 partnership that pushed Reddit content into ChatGPT's retrieval. The Princeton GEO paper found that tactics like quoting evidence and citing sources both lifted visibility 10 to 40 percent in their benchmark (arXiv:2311.09735), which maps directly onto the kind of content Reddit and Wikipedia produce: evidence-led, source-cited.
Gemini's pool leans on Google's index, with extra weight on Google's own knowledge graph and structured data. A brand with clean Organization schema and consistent Wikidata entries earns more retrieval weight here than on ChatGPT.
Perplexity's pool is the broadest of the named engines because its job is explicitly to surface citations. Brand mentions in technical documentation, GitHub README files, and Stack Overflow threads cite well on Perplexity in a way they often do not on ChatGPT.
Claude's pool is the smallest of the four, partly because Claude historically routed retrieval through tool calls rather than a default index. That changed in 2025 with the web search tool. The pool is still smaller and more enterprise-flavored.
The implication for a brand: optimization starts with auditing the pool, not with writing more blog posts. Pull the engines' citation lists for your category's top 30 prompts. If 60% of the citations come from Reddit and your brand is in 2 of those threads, the highest-leverage hour you can spend is on the 18 Reddit threads where you are missing.
Signal 2: how recent the content is
AI engines reward recency more aggressively than SEO ever did. 76% of top AI citations in 2026 came from content updated within 30 days, per Omnibound's tracking (Omnibound, 2026). That is a sharp departure from Google's evergreen pattern. A guide written in 2019 that ranked at position 1 for five years can hold its rank while losing its AI citation slot to a freshly published 2026 piece on the same topic.
The mechanic behind this is straightforward. AI engines are trained on internet snapshots that go stale fast, especially in categories where the facts shift (pricing, feature lists, regulatory landscape, vendor selection). When the engine is uncertain whether a source is current, it down-weights it. The retrieval step prefers sources with date stamps the engine can see; the ranking step penalizes sources where the dates conflict with what the engine knows from elsewhere.
The practical fix is a recency cadence, not a one-time refresh. The pages that matter most (pillar guides, category comparisons, product pages, FAQ pages) need an honest 30-day review rhythm. Honest matters here. Bumping a date stamp without editing the content is detectable; engines cross-reference content fingerprints against indexed history. The update has to be real: a current statistic, a corrected number, a new example, a removed deprecated reference.
For an SEO team transitioning into AEO, the recency cadence is the single biggest cultural shift. The mental model of "write it once, let it rank for years" stops working. The new model is "publish once, refresh on a schedule, retire when the topic drifts."
[INTERNAL-LINK: For the CTR side of the recency story, see our AEO vs SEO comparison]
Signal 3: citation density on the topic
Citation density is the number of times your brand is named across the source pool for a given topic. It is not the same as backlink count, which counts hyperlinks from external pages. Citation density counts brand mentions, with or without links. A brand mentioned in 4 of 10 relevant Reddit threads has higher citation density than one mentioned in 1 of 10, regardless of how many of those mentions hyperlinked.
The dynamic is power-law. Engines cross-reference, so each new citation makes the next citation more likely. A brand mentioned in three Reddit threads, a YouTube transcript, and a category review article looks like a consistent presence to the engine. A brand mentioned only in its own marketing site, regardless of how often, looks like a single source making claims about itself.
There are two ways to lift citation density: build presence where the source pool already is, and earn organic mentions in the places the source pool reaches. Building presence is direct: a Wikipedia stub on the company, a founder profile on a category review platform, a sourced answer on the top-cited Reddit thread for your category. Earning organic mentions is slower but compounds: a launch that gets covered, a research piece that gets cited, a customer story that gets shared. The mistake teams make is treating these as PR campaigns. Citation density is built one entry at a time, over months.
The brands that win citation density consistently are the ones with a named human (a founder, a researcher, an analyst) publishing on the topic in their own name. Author identity carries citation weight that brand identity does not. AI engines weight named human authors slightly higher than generic Team bylines, mirroring Google's E-E-A-T stance.
Signal 4: schema and entity clarity
Schema markup does not pick the brand for the engine. It labels the brand so the engine knows what it is. Without schema, the engine still reads the page; with schema, the engine reads the page with confidence about the page's type, the author, the publication date, and the brand's relationships to other entities.
The four schema types that move the needle on AEO are simple:
- Article schema with a named author, publication date, and update date. Tells the engine the page is editorial content, dated, attributable to a person.
- FAQ schema where the page answers recurring questions. Tells the engine the questions and answers map to user queries, ready to be lifted.
- HowTo schema where the page documents a process. Tells the engine the page is a step-by-step.
- Organization schema with consistent name, logo, sameAs entries (Wikidata, LinkedIn, Twitter), and contact details. Tells the engine the brand is one entity with consistent identity across the web.
Entity clarity beyond schema means the brand's name is consistent everywhere. The same name on Wikipedia, Wikidata, LinkedIn, Crunchbase, Twitter, GitHub. Conflicting names create entity ambiguity, which engines resolve by lowering confidence and citing less.
The highest-leverage hour a technical SEO can spend on a page that already ranks is adding Article schema with a named author. The lift shows up on both Google and the AI surfaces.
Signal 5: off-site presence
68% of AI citations originate off-site, per Omnibound's 2026 tracking (Omnibound, 2026). This is the signal most marketing teams under-invest in because the muscle memory from SEO says optimize your own site first. The engine math says otherwise.
The off-site presence map for a typical category has six layers.
Reddit and category forums. The single largest off-site source for ChatGPT. Top threads on category queries cite 3 to 8 brands by name in the top comments.
YouTube and podcast transcripts. Strong on long-tail technical queries. Engines pull from transcripts even when the video itself was never optimized for SEO.
Wikipedia and Wikidata. Disproportionately strong for entity recognition. A brand with a Wikipedia stub is unambiguous to every engine.
Vertical review platforms. G2, Capterra, TrustRadius for B2B SaaS; equivalents in adjacent categories. Engines cite these when prompts include comparison language.
News and trade outlets. Weighted heavily on recency. A brand with a tier-1 news mention this month carries more weight than a tier-1 mention from two years ago.
Niche forums and community spaces. Stack Overflow, Hacker News, indexable Discord and Slack communities for technical brands.
The fix is to map the layers, find the gaps, and fill them. Not by spamming. By being a useful presence: a real answer in the Reddit thread, a real Wikipedia stub with sourced facts, a real comparison entry on the right vertical review platform. The work is slower than SEO link-building because each entry requires substance. The payoff is durable because the engines weight presence more than backlinks.
What this means for a brand starting from zero
If you are reading this and your brand has zero presence in any of the five signal layers, the temptation is to do them all at once. Do not. The signals compound, but they also have a sequence.
Start with citation density on the highest-traffic prompts in your category. Pick 20 prompts. Find where your brand is missing in the existing source pool. The first month of work is closing those gaps. It is unglamorous and it does not show up in any AEO dashboard until the engines refresh, which is six to eight weeks.
Once citation density is moving, attack recency. Pick the five most-cited pages on your own site and put them on a 30-day review schedule. The schedule needs to be calendared, owned, and visible to the team. Without that structure, recency drifts.
Schema and entity clarity is third. Add Article schema with named authors first; FAQ and HowTo schema follow on the pages where they apply. Audit the brand's name across Wikipedia, LinkedIn, Crunchbase, GitHub. Make them all match.
Off-site presence is the fourth and longest investment. Six months of consistent presence-building outperforms a year of one-off PR campaigns. The brands that finish 2026 with strong off-site presence are the ones that started in Q1 and never stopped.
The fifth signal, source-pool composition, is something you read, not something you write. It tells you which of the four levers above to pull next.
[INTERNAL-LINK: For the definition of AEO and GEO that frames this stack, see our GEO explainer]