How ChatGPT Actually Decides Which Brands to Recommend (And What Solo Founders Can Do About It)

ChatGPT citations: how brand recommendations actually work, and what UK founders can influence
Meta description: Understand how ChatGPT citations work, where recommendations get their information, and what solo founders can influence to improve AI visibility.
ChatGPT citations aren't governed by one public ranking score, and there's no single switch you can flip to guarantee your brand shows up. That's the honest answer, and I know it's not the satisfying one. When a solo founder asks me why a competitor keeps appearing in ChatGPT answers and they don't, what they usually want is a hack. What I can actually offer, after running daily query monitoring across client accounts at MentionOwl, is something more useful: a clear picture of which signals are documented, which are patterns I've observed in that data, and which levers are worth your limited time.
Here's the structure of this piece. ChatGPT recommendations can draw on model training knowledge, live web retrieval, or both, and the exact mix varies by product, model version, search mode, and even the specific query you type. I'll walk through how ChatGPT citations work, where the information comes from, why some small brands out-cite larger competitors, what you can realistically influence, and how to measure whether it's working. I'll flag clearly where a claim comes from OpenAI's own documentation or published research, and where it's a pattern I've observed in client data rather than a universal law.
How ChatGPT citations and AI-generated recommendations work
A ChatGPT answer isn't one process. It's several layers working together. OpenAI's documentation on ChatGPT Search (openai.com/index/introducing-chatgpt-search) describes a system that can rewrite a user's question into one or more search queries, retrieve information from the web, rank that retrieved material, and then generate an answer with inline citations and a sources panel. That's meaningfully different from the base model simply recalling facts memorised during training, and it's worth sitting with that distinction before doing anything else.
This is also where a mention and a citation get confused constantly, so let me make it concrete. Imagine someone asks ChatGPT for a cheap accounting tool for a small UK consultancy. One possible answer might say: "Several freelancers find [Brand A] good value for sole traders." That's a mention: Brand A's name appears, warmly, but there's no link and no traceable source. A different answer might say: "According to a comparison on [Review Site], [Brand B] offers VAT-ready invoicing from £12/month [1]," with a clickable source in the panel. That's a citation: a specific source, attributed and linkable. OpenAI's own guidance is blunt that a citation reflects evidence the system used, not necessarily an endorsement of quality.
For measurement purposes, this distinction matters enormously: a brand can be mentioned with zero citation, and a brand can be cited without being the one actively recommended.
Share of voice builds on this distinction, and it's the metric I track most closely in client reporting. It isn't about appearing once in a lucky response. It's about appearing consistently across the realistic range of questions a prospective customer might type: "best accounting software for a five-person consultancy," "cheapest alternative to Xero," "privacy-focused invoicing tool for freelancers." Princeton's GEO (Generative Engine Optimization) research, published via arXiv in 2023 (arxiv.org/abs/2311.09735), found that the specific wording of a request changes which ranking criteria the model applies, meaning content that explicitly connects to use cases, constraints, and trade-offs performs differently from generic product claims. One strong citation on one query tells you very little. Consistent citation across a representative set of queries tells you a great deal.

Where ChatGPT gets information for brand recommendations
The question I'm asked most often, almost always phrased the same way: does ChatGPT crawl the web live, or is it working from old training data? The honest answer is "it depends on the product and settings in use," and that ambiguity is itself useful: it tells you to plan for both layers rather than betting on one.
Three sources feed into any given answer, and it's worth being precise about what each one actually is:
- Training data: the static knowledge built into the model during training, with a fixed cutoff date that doesn't update in real time. OpenAI doesn't publish a complete, searchable list of training sources, so you genuinely cannot verify whether your brand's content made it into any particular training snapshot.
- Live retrieval: when ChatGPT Search or web-browsing features are active, the system fetches current web pages for that specific prompt, via indexes and retrieval infrastructure rather than a live crawl of the entire internet in real time. This is the layer where fresh content and recently published comparison pages have a realistic chance of being picked up, though I'd treat any specific timeframe (days versus weeks) as variable rather than guaranteed, since it depends on indexing partners and query type.
- Third-party aggregation: review sites, directories, comparison articles, and forum discussions that retrieval systems appear to treat as credible intermediaries. Both Microsoft's Bing documentation and Google's own guidance on how search works describe answer systems blending brand-owned content with independent third-party evidence rather than relying on either exclusively. Neither source confirms the exact weighting ChatGPT applies.
What this means practically: your AI visibility strategy can't be a single bet. Optimise only your own site, and you're ignoring retrieval's apparent reliance on third-party corroboration. Chase only reviews and directory listings, and you risk a site that's never technically legible enough to enter the retrieval set in the first place, however good the business is offline.

Why some small brands get more ChatGPT citations than bigger brands
This genuinely surprises most solo founders, and it's one of the more encouraging patterns in the research I've reviewed. There's no publicly documented, permanent "brand trust score." When retrieval is active, the system pulls pages for a specific prompt and builds an answer from that retrieved material. That means a smaller brand can out-cite a larger competitor if its page is more relevant, more specific, more current, or simply easier for a retrieval system to parse cleanly.
Three patterns show up repeatedly in the client data I review, though I'd frame these as observed correlations rather than confirmed universal ranking factors:
Site structure appears to matter more than domain authority alone. A sprawling enterprise site with ten layers of navigation and content locked behind heavy JavaScript rendering can be harder for a retrieval system to extract a clean answer from than a tightly structured small-business page with clear headings and direct answers. I run structural legibility checks as part of client audits, covering things like heading hierarchy, structured data, and crawlability, precisely because legibility and authority are different axes, and conflating them leads founders to chase the wrong fix.
Specificity seems to beat breadth. The Princeton GEO paper found generative systems favour content that answers a narrow question precisely over generic "About Us" pages that talk around a topic. A page documenting exactly how your product integrates with a specific platform, with dated screenshots and troubleshooting steps, gives a model concrete, extractable material. A vague mission statement doesn't.
Consistency correlates with citation likelihood. If a brand is described the same way on its own site, in independent reviews, and on relevant directories, using the same category language, the same pricing framing, the same differentiators, that repetition across independent sources appears to track with higher citation rates in the queries I monitor. Conflicting information across sources seems to push models toward caution or omission instead. I'd note this is a pattern observed in my own monitoring data rather than something OpenAI has confirmed as a mechanism.
In direct terms: citation likelihood, from what I've tracked across client queries, correlates with how clearly a page answers a specific question, how well-structured it is for machine parsing, and whether the same information is corroborated elsewhere on the web, not simply domain size or age.
What UK solo founders can influence in ChatGPT citations
I want to be precise here, because a lot of AI SEO advice overpromises control that doesn't exist. You cannot edit the model. You cannot guarantee a citation on any single query. What you can do is improve the underlying signals consistently, over time. Here's a practical 30-day starting point for a UK-based solo founder with limited budget:
Week 1: audit and identify. List 10 real customer questions, not keywords, actual phrasing, like "best bookkeeping software for a sole trader doing VAT returns" or "cheapest invoicing tool for a UK freelancer." Check whether your site currently has a page that directly answers any of them. Run a basic crawlability check: can the page's core content be read without JavaScript rendering, does it have a clear heading structure, and is pricing (in pounds sterling) current and dated?
Week 2: fix or build one page. Take the weakest query from your list and either fix the existing page or write a new one. Compare a weak version ("We're the best accounting software for small businesses") against a stronger one: "Pricing from £15/month for sole traders, VAT-registered support, Making Tax Digital compatible, last updated [date], compared against [two named alternatives] on cost and feature trade-offs." The second version gives a model usable, specific language. The first gives it nothing to extract.
Week 3: build third-party footprint deliberately. Get listed in two relevant UK directories or review platforms for your category, and pursue one genuine comparison mention: a podcast, a roundup post, a Companies House-adjacent business directory, whatever fits your niche. This is slower and less glamorous than writing content, but corroboration across independent sources is the pattern that shows up most consistently in the data I review.
Week 4: establish a baseline and measure. Before changing anything else, record where you currently stand across a representative query set. You need a baseline before you can tell whether weeks five through twelve actually moved anything.
Accept that influence here is indirect. You can affect which sources get cited, over weeks and months, by improving legibility, specificity, and consistency, not by trying to manipulate any single output. There's no switch. There's a slow accumulation of legible, corroborated, specific evidence.

How to measure AI visibility and ChatGPT citation performance
Manually typing questions into ChatGPT every few days gives you an incomplete, sometimes misleading, picture of your actual AI visibility. Prompt-level results vary by wording, location, model version, and search mode. A favourable answer on a Tuesday afternoon tells you almost nothing about what a prospective customer sees on Thursday morning, phrased slightly differently, possibly from a different region if you're serving UK versus wider English-speaking markets.
A reproducible measurement framework needs four elements: a defined query set (ideally 15–30 realistic customer questions, refreshed periodically), a recorded baseline before you change anything, consistent repeat testing on a fixed schedule rather than ad hoc checks, and a simple log of citations versus mentions versus absence for each query over time. Four metrics are worth tracking specifically: query coverage (the percentage of your query set where your brand appears at all), position-weighted citations (whether you're the first source cited or the fourth), share of voice relative to named competitors, and soft mentions without citation. None of these numbers are perfectly comparable across tools, and none are claimed as an industry standard. They're a practical way to turn noisy, one-off snapshots into a trend.
I built a product, MentionOwl, specifically because running this manually across multiple AI platforms daily isn't realistic for a solo founder. It automates query generation from your site content and runs it across ChatGPT, Claude, Gemini, Copilot, and Perplexity on a schedule. I mention that as a disclosed example of one way to automate the process, not as proof that automation itself guarantees better citations. The underlying levers in the section above matter regardless of which tool, if any, you use to measure them.
On timelines: from monitoring daily query runs across client accounts, meaningful shifts in citation frequency typically show up over several weeks rather than days, since both training-data patterns and retrieval indexing need time to reflect new content. That's why I'd push founders toward weekly trend review over daily checking. Daily snapshots are noisy by nature, and treating noise as signal is how people end up making content decisions based on random variance rather than an actual trend.

Frequently asked questions about ChatGPT citations
Does ChatGPT crawl the web live or use training data to recommend brands?
Both are possible, and which one applies depends on the product configuration. Base responses lean on training data with a fixed cutoff; browsing-enabled modes and search integrations can pull current pages. Results also vary by region and query, so don't assume a result from one location or phrasing generalises everywhere.
Why does ChatGPT cite some sources and not others?
In client data I've tracked, citation likelihood correlates with how clearly a page answers a specific question, how well-structured it is for parsing, and whether the information is corroborated elsewhere, not simply domain size or age. This is an observed pattern, not a confirmed ranking formula.
Can I influence which sources ChatGPT trusts?
Indirectly. You can't edit the model, but you can improve the legibility and consistency of the signals it draws on: clean site structure, direct answers to real customer questions, and a credible third-party footprint. Think of this as generative engine optimisation rather than traditional SEO, and expect gradual rather than immediate change.
How long does it take for changes to affect ChatGPT citations?
From monitoring daily query runs, meaningful shifts typically show up over several weeks rather than days, since training data and retrieval indexing both need time to reflect new content. Track weekly trends against a fixed baseline rather than reacting to single-day snapshots.
[1] Example figures for illustration. Always verify current pricing directly with providers.
Related Articles

The AI Legibility Checklist Every E-commerce Site Needs (And Why Most Fail It)
Discover why AI shopping assistants skip most e-commerce sites. This AI legibility checklist walks UK online retailers through 16 technical checks to

Why Agencies Should Offer AI Visibility Monitoring in 2025: The Business Case for a New Retainer
Learn how agencies can turn brand monitoring into an AI visibility retainer with practical deliverables, pricing, tools, and a plan to win clients ear
Create content like this automatically
Scribe uses AI to generate high-quality blog posts that engage your audience and drive traffic.
