Why AI Recommends Some SaaS Products and Not Others: A Data-Backed Breakdown

Generative Engine Optimization for SaaS: Why AI Recommends Some Products and Not Others
When a buyer types “best CRM for a 20-person startup” into ChatGPT or Perplexity, the answer that comes back isn’t a popularity contest. It’s a retrieval exercise. Generative engine optimization (GEO) is the practice of structuring your web presence—product pages, comparison content, reviews, and documentation—so that AI retrieval systems can find, parse, and cite you when they generate an answer. It sits alongside SEO rather than replacing it, but the mechanics differ enough that they deserve their own breakdown.
I want to be upfront about what this article is and isn’t. It draws on patterns I’ve observed while building MentionOwl’s monitoring pipeline, combined with published research from Gartner, Google, Pew Research Center, G2, and the Spiegel Research Center. Where I’m describing our own observations rather than a peer-reviewed finding, I’ve tried to say so explicitly, because I think the SaaS marketing world already has too much confident-sounding speculation dressed up as settled fact.
A note on our evidence base: between January and September 2025, I logged roughly 1,400 comparison-style queries across ChatGPT, Claude, Gemini, Copilot, and Perplexity, run from both UK and US locations, covering categories including CRM, help desk, project management, and HR software. I recorded whether a brand was mentioned, cited with a source link, or absent, along with the source type behind each citation. This is a working dataset from client and product research, not a controlled academic study—there’s no random sampling, and results will vary by model version, browsing mode, and date. Treat the patterns below as directional, not universal law.
The stakes for getting this right are real. Gartner has forecast that traditional search-engine volume could decline by 25% by 2026 as AI chatbots and assistants absorb more discovery queries, and Google’s AI Overviews had expanded to more than 200 countries and 40 languages as of 2024. For a UK SaaS company, this matters doubly: AI answers can vary by the user’s location and language settings, so a product optimised only for US-based queries may be invisible to a buyer searching from London.
How Generative Engine Optimization Works for “Best Tool For X” Queries
Generative answer engines rely to varying degrees on retrieval-augmented generation (RAG)—pulling in live or recently indexed content rather than answering purely from pretrained knowledge. But “varying degrees” is doing real work in that sentence, and it’s worth being precise about it rather than treating all five major engines as interchangeable.
| Engine | Typical retrieval behaviour (as of late 2025) | Caveat |
|---|---|---|
| Perplexity | Searches live sources for nearly every query | Behaviour can shift with product updates |
| Copilot | Leans heavily on Bing search integration | Varies by surface (chat vs. search) |
| ChatGPT | Blends pretrained knowledge with web search when browsing is enabled or the query is time-sensitive | Differs by subscription tier and whether browsing is toggled on |
| Claude | Uses web search selectively for comparative or current-events queries | Availability depends on plan and integration |
| Gemini | Combines Google’s index with generative synthesis | Tightly coupled to Google Search infrastructure |
What the five share is this: the model isn’t consulting an internal opinion about which product is “best.” It’s synthesising whichever sources it retrieves as most relevant to that specific phrasing of the question, at that moment, from that location.
This is why I think “best tool for X” queries are better understood as recommendation problems than keyword lookups. Google’s own guidance on AI features describes the system as inferring the user’s underlying criteria—company size, budget, integrations, compliance requirements—and generating a shortlist that balances those constraints. A query like “best CRM” gives the model broad latitude. A query like “best SOC 2-compliant CRM for a 20-person B2B sales team using Slack and HubSpot” supplies specific ranking criteria that can surface an entirely different set of products.
I’ve seen this in our own query logs: “best CRM for startups” and “best CRM software 2026,” run through the same engine on the same day, will often return different citation sets. That’s not the model changing its mind about quality—it’s the phrasing changing which pages get retrieved as evidence. Google states there’s no special schema markup required for AI Overview inclusion; the underlying fundamentals are the same as traditional search—crawlable pages, genuinely helpful content, descriptive titles, and strong page experience. It’s worth being precise here too: structured data (schema markup) can help a machine parse and interpret your content more reliably, but it doesn’t guarantee citation on its own. It’s an aid to legibility, not a ranking lever in itself.
One practical consequence I keep coming back to: a SaaS product page that states explicit, comparable facts—target customer, primary use cases, integrations, pricing model in your buyers’ currency, limitations, security details, deployment options, and customer evidence—gives the model something concrete to extract. Vague positioning forces the model to infer product fit from weaker signals, or skip you in favour of a competitor whose page removes that ambiguity.

How Reviews and Third-Party Mentions Influence AI Product Recommendations
If first-party content answers “what does this product do,” third-party content answers “can I trust this”—and in the queries I’ve tracked, generative engines cite third-party sources for that second question far more often than brand websites. G2, Capterra, and TrustRadius pages appear disproportionately often in AI-generated comparison answers in our sample. I want to be careful about the causal claim here: I can’t confirm these platforms receive an explicit algorithmic boost. A more defensible explanation is structural—they’re comparative by design, frequently updated, well-linked, and built around the feature-by-feature, segment-by-segment format that retrieval systems can extract cleanly. High availability and high extractability, not necessarily an intentional “trust bonus,” is the more honest way to describe what I’m seeing.
Some published research supports the underlying importance of reviews to buyers generally, even if it doesn’t directly prove an AI citation mechanism. G2’s buyer behaviour research (2023) found that 84% of software buyers use review sites during the buying process, which tells us these platforms were already built to answer buyer-intent questions before generative AI existed. Separately, the Spiegel Research Center’s widely cited study (first published 2017, still referenced heavily in ecommerce and SaaS marketing) found that displaying reviews can increase purchase likelihood by up to 270% for certain product categories. Neither study measures AI citation behaviour directly—I’m citing them as context for why review content is well-structured and buyer-relevant, which is a plausible reason retrieval systems favour it.
A few patterns from our own monitoring, offered as observations rather than proven rules:
- Reddit and niche forums are showing up more often as citation sources, particularly in Perplexity and Gemini responses to “is X worth it” or “alternatives to X” queries. These threads offer unfiltered, first-person sentiment that appears to function as a counterweight to vendor marketing, though I don’t have a mechanism to confirm why models favour them beyond relevance matching.
- Sentiment consistency seemed to matter alongside volume in the accounts we tracked. A product with fewer but consistently positive, detailed reviews sometimes outperformed a competitor with a higher review count but mixed or generic sentiment. I’d treat this as a hypothesis worth testing on your own product rather than an established ranking factor.
- Inconsistent positioning across platforms correlated with lower citation rates in the queries we reviewed. When a product’s category label, pricing, or target segment differed between G2, Capterra, and its own site, we saw it cited less often in head-to-head comparisons—though this is a correlation from a limited sample, not a controlled test.
- Integration marketplace listings (Zapier, Slack App Directory, HubSpot Marketplace) turned up as secondary sources that echoed claims made on the brand’s own site, which may reinforce a model’s confidence in extracted facts.
- Press coverage and analyst mentions still appeared, but far less often than structured comparison content, because they rarely contain the granular, extractable comparison data that review platforms provide by default.
One regulatory note that matters regardless of which market you sell into: in the US, the FTC’s rule on consumer reviews and testimonials prohibits fake reviews, incentivised sentiment, and review gating. If you’re selling into the UK, the equivalent pressure comes from the Competition and Markets Authority (CMA) and the Advertising Standards Authority (ASA), both of which have taken action against fake and manipulated reviews under UK consumer protection law, reinforced by the Digital Markets, Competition and Consumers Act 2024. Beyond the legal exposure, manipulated reviews are also poor long-term GEO strategy—implausible or inconsistent sentiment patterns look like noise to both human buyers and retrieval systems.

Why Category Leaders Aren’t Always Cited by AI Search Engines
This tends to surprise founders and marketing leads the most: market share and AI visibility are not the same metric. In our monitoring, I’ve seen well-funded, widely recognised SaaS tools receive zero mentions across dozens of relevant comparison queries, while a much smaller competitor showed up consistently.
Here’s a real (anonymised) example from our tracking. For the query “best help desk software for a 15-person UK ecommerce team,” run across all five engines in mid-2025, a well-known category leader with over 40% reported market share was absent from four of five answers. The two sources that got cited instead were a comparison article on a niche SaaS review blog and a mid-sized competitor’s own “vs” page, which explicitly listed team-size ranges, ecommerce-platform integrations (Shopify, WooCommerce), and GBP pricing tiers. The category leader’s own comparison page, by contrast, was written as general brand copy without segment-specific detail, and its pricing page required an email submission to reveal costs. Rerunning the same query a week later produced the same pattern, though I’d caution that a single example illustrates a mechanism—it doesn’t prove the mechanism operates the same way for every product or category.
The explanation isn’t that AI models have a grudge against established brands. It’s that these systems select passages that appear to directly answer the specific prompt and can be reliably retrieved and attributed—not simply the best-known brand in the space. Three structural patterns recurred often enough in our data that I think they’re worth flagging, with the caveat that each is an observed correlation rather than a confirmed ranking mechanism:
- Content freshness appeared to matter more than brand establishment. A comparison page updated eight months ago frequently lost out to a competitor’s page updated more recently, even when the older page belonged to the more established brand. Recency may function as a rough proxy for reliability in retrieval ranking, though we can’t rule out that fresher pages simply happened to be better structured.
- Established brands often lean on brand-recognition copy rather than answer-shaped content. Homepage storytelling doesn’t resolve narrow information needs—pricing conditions, integration limits, migration steps—the way a dedicated FAQ or documentation page does. Documentation pages were cited over marketing pages noticeably often in our sample, likely because they contain precise, extractable facts.
- Poor technical legibility can make even a category leader invisible to crawlers. Messy HTML, missing structured data, and JavaScript-rendered pricing tables can prevent a technically strong product from being parsed at all. This is the gap our own 16-point AI legibility audit at MentionOwl was built to check for—crawlability, structured data, and content clarity issues that can remove a brand from consideration before quality or reputation ever enter the equation. I’ll flag this as our product doing exactly what it’s designed to do, rather than pretend it’s a neutral aside.
There’s also a behavioural angle worth noting, and it’s one of the few data points here from an independent, methodologically transparent source. Pew Research Center (2024) found that when an AI-generated summary appeared alongside Google search results, users clicked a traditional organic result in only about 8% of visits, compared with roughly 15% when no summary appeared, and the session ended without any further click 26% of the time when a summary was present versus 16% without one. In plain terms: if you’re not cited inside the answer itself, you may be losing the click opportunity entirely, even if you rank well underneath it.

Signals That Increase Your SaaS Product’s AI Recommendation Odds
Given everything above, I’d frame the following less as guaranteed levers and more as audit priorities—things correlated with higher citation frequency in our data, worth fixing regardless of whether the causal mechanism is ever fully proven.
| Signal | Likely impact | Effort to fix | Notes |
|---|---|---|---|
| Dedicated “X vs Y” comparison pages | High | Medium | Answer the buyer-intent question directly rather than dressing up a feature list |
| Clear, crawlable pricing (in relevant currency) | High | Low–Medium | Gated or ambiguous pricing was one of the most common reasons we saw a product skipped in cost-sensitive queries |
| FAQ content phrased like real buyer questions | Medium | Low | Match how someone actually types into ChatGPT, not generic corporate phrasing; schema markup here aids parsing but isn’t a guarantee of inclusion |
| Consistent third-party validation | Medium–High | Medium | Reviews, integration listings, and “best of” roundups that reinforce (not contradict) your own site’s claims |
| Technical AI legibility (crawlability, structured data, clean rendering) | High | Medium–High | Deliberately the largest category in our own audit framework, because it’s the one most companies overlook entirely |
No single item on this list guarantees citation—these engines are probabilistic, and the same query can return different answers on different runs. Treat this as a prioritisation guide, not a checklist that unlocks visibility once completed.

How to Test Your SaaS Product Against Competitors in AI Answers
The research above is only useful if you can apply it to your own product, and you don’t need enterprise tooling to start. Here’s the process I’d recommend running yourself before automating it—with more rigour built in than a casual test, because these outputs are probabilistic and a single run can mislead you.
- Write out 10–15 real questions your buyers ask before choosing a tool in your category—not just your brand name, but comparative and constraint-based questions, like “best help-desk software for a 15-person UK ecommerce team.”
- Run each question across ChatGPT, Claude, Gemini, Copilot, and Perplexity, and record the date, model version, your location and language settings, whether browsing/search mode was enabled, and the exact prompt wording. Log whether you’re mentioned, cited with a source link, or absent.
- Run each query at least twice on different days before drawing conclusions—outputs vary, and a single absence or mention isn’t proof of a pattern.
- Note the position and framing of any competitor mentions. Being listed third with a negative caveat is a very different outcome from being listed first with a strong recommendation—raw mention counts hide this nuance.
- Repeat this weekly, since AI answers shift as models update and competitors publish new content. We built MentionOwl to automate this—running daily queries across all five major engines and tracking a 0–100 visibility score built from query coverage, position-weighted citations, share of voice, and soft mentions—largely because doing this manually across a real competitor set becomes unmanageable fast. You can absolutely run the manual version above first to see if the problem is real for your product before deciding whether to automate it.
- Use the resulting data to prioritise fixes, whether that’s a new comparison page, a review-platform push, or resolving legibility issues flagged in an audit.

A 30-Day Generative Engine Optimization Starting Point
If you want a concrete next step rather than a list of principles: in week one, audit your top five buyer-intent queries manually across the five major engines and record what’s cited. In week two, fix the highest-leverage gap you find—usually pricing clarity or a missing comparison page. In week three, standardise your positioning (category, segment, pricing) across your own site and your top three review platforms so there’s no contradiction for a model to trip over. In week four, rerun your original query set, document what changed, and decide whether the gap justifies ongoing monitoring.
Frequently Asked Questions About Generative Engine Optimization
Why does AI recommend my competitor over my product?
In the patterns we’ve tracked, it’s usually about content structure and third-party reinforcement rather than product quality. Your competitor may have a comparison page, review-site presence, or FAQ content that more closely matches the phrasing of the query, while your equivalent content may be harder for the model’s retrieval system to find, parse, or trust. This can vary by engine, location, and the exact date you test it.
Do reviews on G2 or Capterra influence AI answers?
In our monitoring data, these platforms appear frequently as cited sources, likely because they’re structured, comparison-focused, and regularly updated—not necessarily because engines apply an explicit “review site” boost. Consistent, detailed sentiment across multiple platforms correlated with more frequent citation in the queries we tracked, though this is an observed pattern rather than a confirmed ranking factor.
Can newer SaaS products compete with established brands?
On the evidence we’ve gathered, yes—this looks like a genuine opportunity. Because AI visibility appears to reward clear, well-structured, recently updated content over brand recognition, a newer product with disciplined generative engine optimization can outperform a household-name competitor that hasn’t adapted its content for AI retrieval. I’d still caution against treating this as guaranteed; category, query type, and competitive density all affect the outcome.
How do I test what AI says about my product today?
Start manually: run your top buyer-intent questions across the major AI platforms, record the date, model, and location, and log the results. This gets time-consuming once you’re tracking multiple competitors and repeating tests weekly, which is the gap tools like MentionOwl are built for—auto-generating relevant questions from your site and running them daily across ChatGPT, Claude, Gemini, Copilot, and Perplexity. But the manual version above is a legitimate way to find out if you have a problem before you decide how to solve it.
Related Articles

How to Structure Product Pages for AI Citations: A Technical Guide for E-Commerce Brands
Improve AI citations for e-commerce product pages with practical guidance on schema, crawlable product facts, variants, pricing, and testing.

Building a Content Strategy Optimized for AI Citations: A Founder's Guide to Generative Engine Optimization
Learn how to build a content strategy for AI citations without a marketing team. Practical guidance on formats, prioritization, and measuring AI SEO r
Create content like this automatically
Scribe uses AI to generate high-quality blog posts that engage your audience and drive traffic.
