← All reports
AEO · Research report

How Answer Engines Select Sources

A retrieval-led model for understanding eligibility, relevance, supporting evidence, and the limits of what a citation can prove.

First Brand Research3 min read4 primary sources
Executive summary

A source can be accessible without being selected, and selected without supporting every sentence around its citation. Understanding these distinctions helps a brand improve the material it publishes without pretending to know a platform’s private ranking formula. The most useful working model follows a sequence: access, retrieval, synthesis, and attribution. Each stage creates a different question for the publisher.

01

Eligibility is a prerequisite, not a promise of inclusion.

02

Publish material that answers a specific decision with usable evidence.

03

Evaluate the claim beside a citation, not the link alone.

From available page to attributed answer
AccessCan the relevant system reach the page?
A conceptual model of source use. Platforms implement these stages differently.

First, the source must be available

Before discussing content strategy, confirm that important pages are accessible through the appropriate search systems. Check whether the page can be found, read, and understood in its intended form. A missing page, an access restriction, or an unclear page purpose needs a different remedy from weak supporting evidence.

Google states that supporting pages for its AI features must be indexed and eligible for a search snippet; meeting requirements does not guarantee inclusion. OpenAI’s ChatGPT search guidance identifies allowing OAI-Searchbot as an eligibility step. These are platform-specific access conditions, not a universal optimization checklist.

Then, the system needs relevant material

The foundational retrieval-augmented generation paper by Lewis and colleagues combines a language model with retrieved passages. It demonstrates a general architectural idea: generation can use external material rather than relying only on information stored in model parameters. It does not reveal the complete source-selection formula of today’s consumer search products.

For a publisher, our practical interpretation is to make each important page useful for an identifiable need. Explain the relevant service, constraint, process, or comparison directly. Generic promotional language gives a reader less material for answering a precise question.

One answer may depend on several questions

A person asking for an appointment platform may also need evidence about calendar compatibility, permissions, reminders, and migration. Google describes query fan-out across related topics and sources in its AI search documentation. This supports thinking about the decision’s supporting questions rather than optimizing only for its headline wording.

Build connected pages where the topic warrants them. A concise overview can link to detailed specifications and policies. The objective is a coherent information structure: broad enough to explain the decision, but specific enough that a supporting statement can be checked without interpreting an entire marketing site.

A citation is a connection that needs checking

The ALCE research benchmark treats factual correctness and citation quality as separate evaluation problems. That distinction matters in a brand audit. Open the linked source and compare its actual claim, conditions, and evidence with what the answer says.

Record whether the source directly supports the statement, supports only part of it, or does not establish it. Also distinguish a company’s own page from an independent publisher discussing that company. These observations explain the evidence path; they should not be collapsed into a single count of links.

Improve the inputs you can inspect

Use a source audit to identify concrete issues: inaccessible pages, missing explanations, unsupported claims, outdated policies, or contradictory identity information. Fix these in an order that reflects customer importance. Avoid claiming that a particular heading pattern or markup addition guarantees recommendation.

Test the changes across a consistent set of questions, keeping the complete answers and citations. If inclusion improves, report the observation and its limits. The durable goal is to build a body of material that deserves to be used: clear enough to retrieve, specific enough to help, and well-supported enough for a person to verify.

Sources & methodology

01  Google Search Central — AI features and your website ↗

Explains search eligibility, query fan-out, and variation across AI surfaces.

03  Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks ↗

Foundational original research combining retrieved passages with language generation.

04  Gao et al. — Enabling Large Language Models to Generate Text with Citations ↗

Original research separating answer correctness from citation quality; not a benchmark of current consumer products.

About this report

Desk research combining platform documentation with foundational retrieval and citation-evaluation papers. The four-stage model is an explanatory framework, not a reconstruction of any provider’s proprietary ranking system.

Make your next move a clear one.

Make your next move a clear one.

Make your next move a clear one.