Loading...

ChatGPT doesn't recommend brands at random—it weighs a combination of training-data exposure, real-time retrieval from indexed web sources, and citation-worthiness signals such as structured content, third-party validation, and topical consistency. In practice, this means brands with clear, well-structured, frequently referenced content across the web get surfaced more often than brands with thin or inconsistent digital footprints, regardless of ad spend.
I've spent the last several months at MentionOwl analysing exactly how this plays out across thousands of tracked prompts for SaaS and DTC clients. What I've found consistently contradicts the assumption that AI recommendations are some kind of black-box lottery. They're not. There's a discoverable logic underneath, and in this piece I'll explain how ChatGPT citations work and, more importantly, which levers product marketers can pull.
Before I go further, I want to be transparent about something important: OpenAI does not publish a complete, page-level ranking formula for ChatGPT citations, and it explicitly states there's no way to guarantee top placement in ChatGPT Search. What follows is built from OpenAI's own documentation, published generative engine optimisation research, and pattern analysis across the AI visibility data we track daily at MentionOwl. Where I'm citing observed patterns rather than confirmed mechanics, I'll say so.
To understand why ChatGPT mentions one brand over another, you first need to understand that a ChatGPT response isn't drawing from a single, unified database. It's assembled from up to three distinct layers, and which layers are activated depends heavily on the product mode and the query itself.
The first layer is pre-training data—the static corpus the model was trained on, which OpenAI describes as a mix of publicly available information, licensed or partnership data, and content generated by human trainers and researchers, all captured up to a training cutoff date. Critically, OpenAI doesn't publish a complete inventory of exactly which pages or domains fed into this corpus, which means you can't audit your presence in it directly. You can only infer it from how the model discusses your brand in non-browsing conversations.
The second layer is the browsing and retrieval layer, which activates when ChatGPT Search is enabled, including in Plus, Team and increasingly free-tier contexts. Here, the model issues live retrieval calls against an indexed web corpus, pulling current pages, comparison articles and review content. It may then attach inline citations and a visible Sources panel to the answer. This is the layer where most practical AI SEO and generative engine optimisation work has leverage because it queries a web index that updates continuously rather than a frozen training snapshot.
The third layer involves licensed content partnerships. OpenAI has struck deals with a number of publishers, and while the company hasn't disclosed a public weighting system for how licensed sources are treated relative to organically indexed pages, it's reasonable to assume—based on how retrieval systems generally work—that trusted, licensed sources carry different retrieval priority from an arbitrary blog post.
Here's why this three-layer structure matters in practice: a claim on your own marketing page (“the #1 rated tool for freelancers”) sits entirely inside your controlled narrative. Retrieval systems generally treat self-published claims as lower-corroboration evidence than the same claim appearing in an independent comparison article or review site. I've seen this play out repeatedly in our client data—brands with strong owned-site copy but weak third-party presence get mentioned far less than their market share would suggest.

Once you understand the source layers, the next question is how ChatGPT decides what to surface and in what order. Based on OpenAI's documentation and the generative engine optimisation research coming out of Princeton, including the GEO benchmark study that evaluated 10,000 queries across multiple domains, the process runs through roughly five stages:
Query interpretation — ChatGPT parses the intent behind a prompt such as “best project management tool for freelancers” and determines whether it can answer confidently from training data alone or whether it needs to trigger live retrieval for current, verifiable information.
Source retrieval — For retrieval-augmented queries, the model pulls a shortlist of candidate documents. OpenAI's own guidance indicates this shortlist is shaped by relevance, reliability and usefulness to the specific query—not by domain authority alone, and critically, not by any mechanism related to advertising spend.
Source scoring — Documents get implicitly evaluated against factors including topical relevance, structural clarity, clean headings, lists, schema markup and corroboration—whether multiple independent sources say the same thing. The GEO research found that content featuring clear explanations, statistics, direct quotations and technically precise language tended to perform better in generative visibility tests, sometimes by margins as large as 40% in the study's benchmark conditions. However, that's an academic result under controlled test conditions, not a guaranteed real-world uplift.
Synthesis and citation surfacing — The model composes an answer and, depending on the surface—ChatGPT Search versus standard chat—may attach inline citations. Brands referenced across multiple corroborating sources tend to surface earlier and more consistently than brands appearing in just one high-authority post.
Repetition effect — This is one of the more counterintuitive findings from our monitoring work: a brand that appears modestly across ten independently authored sources often builds stronger “citation gravity” than a brand that lands one glowing feature in a single prestigious publication. Corroboration density, not peak authority, seems to be the deciding factor.
It's worth being precise about terminology here too: a citation is not an endorsement. ChatGPT can cite your page as one piece of evidence while still concluding that a competitor is the better fit for the user's stated criteria. This is why I always encourage clients to separate mention rate, citation rate and recommendation rate as three distinct metrics rather than collapsing them into one vague “visibility” number.

When we run comparative audits across SaaS clients at MentionOwl, the pattern that separates high-citation brands from low-citation brands is remarkably consistent. It rarely comes down to product quality differences—it comes down to how discoverable and corroborated the information around the product actually is.
| Factor | Low-Citation Brand | High-Citation Brand |
|---|---|---|
| Content depth | Vague marketing copy, generic feature lists | Detailed comparison pages, explicit use-case breakdowns, transparent pricing |
| Third-party corroboration | Relies almost entirely on owned-site claims | Present in G2/Capterra reviews, independent roundups, niche blog coverage |
| Structured data | Minimal or missing schema | FAQ, product and organisation schema implemented and maintained |
| Domain authority | Generalist site touching many unrelated topics | Consistent topical coverage of a defined niche, such as “CRM for agencies” |
| Freshness | Published once, rarely revisited | Comparison and pricing pages updated on a regular cycle |
The freshness point deserves emphasis because it's the one teams most often neglect. Pricing, integrations, security certifications and feature sets change constantly in SaaS, and retrieval systems appear to favour recently updated content when synthesising answers. A comparison page that hasn't been touched in eighteen months isn't just stale—it risks actively feeding the model outdated or inaccurate information about your product, which can suppress recommendations or, worse, produce a citation that misrepresents your current offering.

Beyond the structural factors above, there's a set of trust signals that consistently correlate with higher citation frequency and more favourable framing in the data we track:
Given everything above, here's the practical sequence I recommend to product marketing teams starting this work:

I want to be direct about something: there is no universal, OpenAI-sanctioned “AI visibility score”. It's a measurement framework the industry has built to make an otherwise invisible, probabilistic system trackable. At MentionOwl, we've built ours around four components that I think give the most honest picture of a brand's standing:
The reason we run this daily rather than weekly or monthly comes down to volatility. Our data shows ChatGPT's cited sources can shift meaningfully week to week as the underlying web index refreshes—a comparison article gets updated, a new review lands on G2, or a competitor publishes fresh pricing. A single manual prompt check, however carefully done, is really just one frame from a moving picture. That's the entire premise behind building continuous LLM monitoring rather than treating AI visibility as a quarterly audit item.

It draws from two layers: its static pre-training corpus, which includes a broad snapshot of web content up to a cutoff date, and a live retrieval layer that pulls from an indexed web corpus plus select licensed publisher content when a query requires current information. Brands cited across multiple independent sources in that retrieval layer tend to surface more consistently than those relying solely on owned-site content.
No—there's no evidence, and no disclosed mechanism, by which ad spend directly buys placement in ChatGPT's answers. What ad spend can influence indirectly is traffic and brand searches, which occasionally correlate with increased third-party coverage. However, the actual citation logic runs on retrieval relevance and source authority, not media budget.
In nearly every case I've analysed, it comes down to corroboration density: the competitor appears across more independent sources—review sites, comparison blogs and community forums—that the model's retrieval layer treats as trustworthy. It's rarely about domain authority alone; it's about how many separate places on the web are saying the same thing about them.
The live retrieval layer refreshes continuously because it's querying a web index in near real time, but the static training data only updates with major model releases, which happen on a much longer cycle. This is exactly why relying on a one-time manual check of ChatGPT's answers gives you a misleading picture—ongoing monitoring is needed to catch genuine shifts.
Generally yes, because both can serve as high-corroboration sources that many retrieval systems weight heavily. That said, for most SaaS and DTC brands, consistent presence in review platforms and niche industry publications will move the needle faster than chasing a Wikipedia entry, which has strict notability requirements.

AI shopping assistants like ChatGPT and Gemini now curate product shortlists before shoppers see a search results page. Here's how UK e-commerce brand

AI SEO and traditional SEO aren't the same game. Here's the side-by-side comparison marketing teams need to report share-of-voice and AI visibility da

Learn how ChatGPT citations affect local UK businesses, how to test AI search visibility, and practical ways to help customers find you instead of riv