Guide

How ChatGPT Chooses Sources

In grounded mode, ChatGPT chooses sources through a retrieval-then-synthesis pipeline: it searches the web for candidate pages, ranks them by relevance and trust, then composes its answer around the passages it can extract and verify. Understanding each stage tells you exactly where brands get filtered out.

  • how ChatGPT chooses sources
  • ChatGPT citations
  • Marketing teams

Last updated

Stage 1: Retrieval

ChatGPT issues search queries derived from the user's prompt and collects candidate pages. Filtered here: sites that block AI crawlers, pages that don't match the reformulated queries, and content too new or too stale for the index it searches. If you never appear in AI answers, check crawlability first — it's the cheapest fix in the pipeline.

Stage 2: Source ranking

Candidates are weighted by relevance (does the page answer this specific question), authority (domain trust, external corroboration, entity clarity), freshness (recent updates win ties, especially for comparative and pricing questions), and consistency (claims that agree with other trusted sources rank; contradictions get dropped).

Stage 3: Extraction and synthesis

The model quotes passages it can cleanly lift: direct claims, tables, definitions, steps. This is where structure decides between two equally authoritative pages — the one with a boundable, attributable answer gets cited; the one with the answer buried in narrative doesn't. Ungrounded mode skips retrieval entirely and answers from training knowledge, which is why brands present in one mode and absent in the other need different fixes.

Becoming a chosen source

Work the stages in order: let the crawlers in, cover the questions, build corroborated authority, structure for extraction, stay fresh. Then verify — Seeqly's citation audit shows which of your pages ChatGPT actually fetched and cited, which turns this guide from theory into a scoreboard.

Frequently asked questions

Does ChatGPT prefer certain websites?

It favors sources that are relevant, corroborated, fresh, and extractable — in practice: established publications, active communities like Reddit, review platforms, and well-structured brand pages.

Why does ChatGPT cite my competitor but not me?

At one of the three stages you're losing: not retrieved (crawlability/relevance), outranked (authority/freshness), or unquotable (structure). Prompt traces show which.

Do other engines choose sources the same way?

The pipeline is similar across grounded engines, but weights differ — Perplexity surfaces more sources visibly, Gemini leans on Google's index. Per-engine measurement catches the differences.

Related reading