Loading...

If you've noticed organic traffic behaving strangely over the past year (branded search dipping, direct visits flatlining, demo requests down even though your rankings look fine), I'd start by asking a different question than the one your SEO dashboard is answering. The question isn't "are we ranking?" It's "are we being recommended?"
Here's the short version. SaaS buyers are increasingly asking ChatGPT, Perplexity, and Gemini for software recommendations before they ever hit a review site, and most vendors have no visibility into what these engines are saying about them. If your product isn't showing up in those AI-generated shortlists, there's a real risk you're being filtered out of the funnel before a human ever sees your homepage. Traditional SEO metrics won't reliably tell you this is happening. Google Search Console can't see a conversation that never generated a click, and Google Analytics can't attribute a deal that started with someone typing "best project management tool for a 20-person agency" into ChatGPT and never touching your domain until they were already sold. I want to be upfront about a limitation here: declining branded and direct traffic has several possible causes, including seasonality, reduced marketing spend, SERP layout changes, and shifting demand, so AI exclusion is a hypothesis worth testing with real data, not an automatic conclusion.
I've spent a lot of time in the data behind this shift, and I want to walk through how it's happening, how AI engines actually decide which tools to mention, the warning signs that you're being quietly excluded, and, most importantly, what to do about it, including a measurement framework you can actually run this week. This discipline is now commonly called AI SEO or generative engine optimization (GEO), and for SaaS and DTC brands competing on comparison and "best for" queries, it's becoming difficult to ignore, in the UK market as much as anywhere else.
The traditional SaaS buying journey looked something like this: keyword search, a scroll through blue links, a click into a review site like G2 or Capterra, a comparison article, and eventually a demo request. That journey is being compressed for a meaningful share of buyers, and in some cases short-circuited entirely, by a single conversational interaction.
Google began rolling out AI Overviews in the United States in May 2024 and expanded the feature to more than 100 countries and territories, including the UK, by the end of that year, reaching over a billion people according to the company's own reporting. OpenAI launched ChatGPT search in October 2024, giving users cited answers inside the chat interface rather than a list of links to evaluate themselves. Most of the hard usage data on this shift is still US-centric: Bain & Company's 2024 research on AI-assisted buying found that a majority of consumers now rely on AI-generated results for a meaningful share of their searches, and Adobe reported that over a third of US consumers had already used generative AI for online shopping research, with more than half saying they intended to. Worth flagging clearly: these are US and global figures, not UK-specific ones. UK-specific B2B research on this exact behaviour is still thin, though Ofcom's Online Nation tracking has consistently shown that a substantial and growing minority of UK adults have used a generative AI chatbot, with adoption climbing year over year. Treat the direction of travel as reliable and the precise percentages as provisional until UK-specific B2B studies catch up.
What's happening structurally is this: the buyer's consideration set, the shortlist of two or three vendors they'll seriously evaluate, is increasingly being formed inside a chat window rather than assembled by the buyer across multiple tabs. Picture a UK-based ops lead at a 20-person agency in Manchester, researching project management tools on a Tuesday afternoon. She asks ChatGPT which tool suits a small agency team, gets three names with reasoning attached, and books a demo with one of them without ever running a Google search. If your product isn't part of that synthesis, you don't get a chance to make your case in round two of a comparison. You're not in round one.
This dynamic matters more for mid-market and SMB SaaS than for enterprise, where procurement processes, RFPs, and existing vendor relationships still dominate the buying motion regardless of what an AI assistant says. It doesn't yet replace review sites or procurement outright. It sits earlier in the funnel, shaping which vendors even get considered. But for the segment where a founder, ops lead, or marketing manager is doing their own research, the AI-assisted shortlist is becoming a real first filter, and it's precisely this segment where most SaaS companies have the least visibility into what's happening, because their analytics stack was built to measure search-engine referrals, not chat-based research sessions that leave no referrer at all.

This is the blind spot I keep coming back to: your web analytics can tell you that traffic is down, but it can't tell you whether an AI engine chose your competitor over you in a query that never resulted in a visit at all. That's a different visibility problem than anything traditional SEO reporting was built to solve, and it's the gap that AI visibility monitoring exists to close.
Precision matters here, because Google AI Overviews, ChatGPT Search, Perplexity, Gemini, Copilot, and Claude do not all work the same way, and treating them as one system leads to bad optimisation decisions. Some of these are retrieval-augmented: they run a live web search at query time and generate an answer grounded in what they find, which is how Perplexity, Google AI Overviews, and ChatGPT's search mode largely operate. Others lean more heavily on what a model learned during training, which updates on a much slower cycle and can lag behind your current site by months. A brand can rank well in conventional Google search and still be absent from an AI Overview or a ChatGPT answer for a functionally identical query, because retrieval, indexing, and citation selection are governed by different mechanics than classic ranking.
There isn't one universal ranking formula, and none of the platforms have published a definitive scoring algorithm. Based on available research, platform documentation, and the patterns we track, the following factors are best understood as measurement heuristics, things that correlate with inclusion, rather than confirmed, guaranteed ranking signals:

Researchers at Princeton studying generative engine optimization (Aggarwal et al., "GEO: Generative Engine Optimization," 2024) found that techniques like adding citations, quotations, and more authoritative, evidence-backed language improved a source's visibility in generated answers by up to 40% in their experimental evaluation. The effect varied significantly by domain, query, and model, and their study focused on a specific set of synthetic query sets rather than the full range of commercial SaaS buying queries. GEO isn't a guaranteed formula. The honest takeaway is that specific, verifiable, well-corroborated claims consistently outperform generic positioning language across the systems studied so far.
Before treating any of the following as evidence of AI exclusion, it's worth ruling out the more mundane explanations first: seasonal demand swings, a paused ad campaign, a competitor's funding announcement or PR push, tracking or tagging changes on your site, and normal SERP volatility. AI visibility problems tend to show up alongside these signals, not instead of them, so isolate the pattern before you act on it.
With that caveat in place, here's what's worth watching for deliberately, using a repeated sample of at least ten to fifteen prompts per engine rather than a single check:
Any one of these in isolation might be noise. Two or three together, sustained over several weeks of repeated testing, is a pattern worth investigating with actual logged data rather than a single manual prompt.
Once you suspect AI visibility is a gap, here's a sequence with rough ownership and effort attached, so it's actually executable rather than aspirational:

Here's a mistake I see constantly: a founder runs three prompts through ChatGPT, feels reassured because their brand showed up once, and moves on. That's a false sense of security. AI outputs are probabilistic and shift with prompt wording, model version, location, and available citations at the moment of the query. One favourable answer today tells you very little about your standing next week. It's also worth being precise that "changes daily" doesn't mean every model retrains daily. Retrieval-based answers can shift quickly because the underlying web index changes, while training-based knowledge shifts more slowly, on the vendor's own release schedule.
What works better is a repeated, logged set of buyer-relevant prompts run against named competitors on a consistent cadence: weekly for most teams, daily around specific events. The comparison below is where the practical difference shows up:
| One-off manual prompting | Repeated, logged tracking | |
|---|---|---|
| Sample size | 1–3 queries, once | 25–50+ queries, on a fixed schedule |
| Detects competitor gains | No — depends on remembering to re-check | Yes — flags share-of-voice shifts against a baseline |
| Sentiment tracking | Not captured | Tracked over time, tied to specific content or review changes where possible |
| Actionability | Anecdotal, reactive | Logged data that can brief content and PR before deals are lost |
| Coverage across engines | Usually just ChatGPT | Multiple engines, weighted by where your buyers actually research |
Definitions worth being explicit about when you build this: mention rate is the share of your logged prompts where your brand appears at all; first-position rate is the share where you're named first or most prominently; citation rate is how often a specific page of yours is linked as a source; share of voice is your mention count relative to named competitors across the same prompt set; sentiment is a simple positive/neutral/outdated tag applied manually or by a second model pass, since none of this is yet standardised across the industry.

The real value of repeated tracking is timing. When a competitor publishes a new comparison page or lands a PR placement that starts shifting citations in their favour, you want to know within a week or two, not discover it a quarter later when pipeline has already softened. You can build the tracking spreadsheet described above by hand, or use a dedicated tool. Disclosure: I work with MentionOwl, which automates this exact prompt-and-log process across major AI platforms and rolls the metrics above into a single 0–100 visibility score. I mention it here as one example of how the mechanics work in practice, not as a claim that it's the only way to do this. The principle holds regardless of which tool you use to run it: a monitoring cadence that matches how fast these outputs actually change, tied to a defined prompt set and named competitors.
A growing share do, based on the US and global adoption data cited above, though robust UK-specific B2B research on this exact behaviour is still limited. Directionally, buyers appear to be using AI assistants to shortlist tools the way they once used Google "best X software" searches, sometimes skipping the click-through and acting on the AI's answer directly. This varies by category, deal size, and how habituated a given buyer is to using AI tools at all.
Usually it's some combination of stronger query coverage on comparison content, more consistent and recent third-party mentions across review sites and forums, and cleaner machine-readable structure on their site. These are correlations we've observed across the accounts we track, not a confirmed formula any platform has published, so the honest answer is "probably several of these factors together," and an audit is the only way to know which ones matter most for your specific category.
You can't directly edit a model's output, but you can influence the inputs it draws on: your website's clarity and structure, the comparison content you publish, and the third-party sources that mention you. This is the core practice behind generative engine optimisation and AI SEO, and it works on a lag. Expect weeks to months before changes show up consistently in outputs, not days.
Worth being precise: no monitoring approach can prove that a specific AI mention caused a specific closed deal, in the same way attribution has always been imperfect for dark social and word-of-mouth. What repeated tracking can do is show you directional share-of-voice trends against named competitors over time, which is still more actionable than a single manual check. Given how quickly retrieval-based outputs can shift, sometimes within days of a re-crawl or a competitor's new content, a one-time check tells you very little on its own.

Mention counts tell leadership nothing about brand perception. Learn how sentiment analysis across ChatGPT, Gemini, and Perplexity reveals whether AI

A real-world case study on AI visibility decline: how a UK small business lost customers to a competitor in ChatGPT and Gemini answers—and the data th

A plain-language breakdown of the AI visibility score—query coverage, position-weighted citations, and share of voice—so founders know exactly what to