Loading...

Here's the workflow I recommend to every agency that asks me about AI brand monitoring: build a representative prompt set from real customer language, run it consistently across the AI search platforms that matter for that client's category, and convert the results into caveated, client-ready actions rather than raw scores. That's the whole operational answer, and the rest of this piece unpacks each of those three steps in enough detail that you could implement them next week rather than next quarter.
I've spent a fair amount of time recently pulling apart how agencies actually operationalise AI brand monitoring, and the honest answer is that most teams are still doing it manually — typing questions into ChatGPT one at a time, screenshotting the answer, and hoping nobody asks about Perplexity. That gap is closing. The agencies that build a repeatable process now will be answering client questions with data rather than guesswork the next time someone asks, "Why does it keep recommending [competitor]?"
The shift in how people research purchases isn't theoretical, though I want to be careful about which numbers are UK-specific and which aren't, because most of the robust data currently available is US-led.
Pew Research Center published a study in March 2025 examining Google search behaviour among a panel of US adults, finding that AI Overviews appeared in 18% of the sampled queries, and that users clicked through to a traditional web result in just 8% of those AI-Overview visits, compared with 15% on searches without an AI summary present (Pew Research Center, "Google users are less likely to click on links when an AI summary appears in the results,” March 2025 — search this exact title on pewresearch.org, as Pew's URLs restructure over time). That's roughly a halving of click-through behaviour whenever an AI-generated answer sits above the fold. It's a single study, on a US sample, at one point in time, so treat it as a strong directional signal rather than a UK constant — but the direction matters, because it describes exactly the mechanism by which a brand can lose visibility invisibly.
Gartner's public commentary, cited widely since 2024 (Gartner press release, "Gartner Predicts Search Engine Volume Will Drop 25% by 2026 Due to AI Chatbots and Other Virtual Agents"), forecasts that traditional search engine volume could fall 25% by 2026 as generative AI tools absorb query volume. This is a forecast, not a measured outcome, and Gartner itself frames it as a prediction rather than an observed trend — I'd cite it as context for why agencies are moving now, not as evidence of where things already stand.
On actual measured behaviour, Adobe Analytics' 2024 US holiday shopping data (Adobe, "Adobe Analytics: US Online Holiday Shopping Set to Reach New Heights,” November 2024) found generative-AI referral traffic to US retail sites rose more than 1,000% year-over-year during that period, with those visitors browsing more pages per session than visitors from most other referral sources. That's concrete, but it's US retail traffic, not UK, and not necessarily representative of other sectors like B2B services, finance or healthcare.
The UK-specific evidence is thinner than I'd like, and I think it's more honest to say that plainly than to paper over it with a vague Ofcom mention. Ofcom's Online Nation report tracks UK adult internet behaviour annually, and its most recent published editions have shown a rising share of UK adults reporting at least occasional generative AI chatbot use — but because Ofcom updates this report on its own cycle and revises methodology between waves, I'd point you to Ofcom's current Online Nation release directly (ofcom.org.uk, search "Online Nation") rather than repeat a number here that may already be out of date by the time you're reading this. What I can say with more confidence, from working with UK agencies directly, is that client-side questions about AI visibility have gone from essentially zero eighteen months ago to a recurring theme in new-business pitches and quarterly reviews across the accounts I've been close to — that's a qualitative observation from direct client work, not a survey finding, and I'm flagging it as such.
What this means practically is that clients are starting to notice something unsettling before you've had the chance to raise it yourself. They ask ChatGPT or Perplexity a question about their own category, and a competitor's name comes back instead of theirs. Account teams I've worked alongside have had this exact scenario raised in client meetings with no answer prepared, and that's an uncomfortable position to be in with a paying client.
The reassuring part is that this isn't a separate discipline you need to sell clients on from scratch. It's a natural extension of the SEO and PR reporting conversations you're already having. Clients understand visibility, sentiment and share of voice conceptually — AI brand monitoring applies those same concepts to a new set of AI search platforms. Google's own Search Liaison team has stated publicly that foundational SEO practices — crawlability, structured data, authoritative third-party references — remain relevant inputs to how AI systems select and cite sources (Google Search Central Blog, "AI features in Search," blog.google/products/search), which means you're extending an existing service line rather than inventing a new one.
The risk cuts the other way too. If you don't raise AI search visibility with a client, another agency pitching against you probably will, and it can meaningfully differentiate a pitch — particularly when the competing agency hasn't thought to bring it up at all.
Single-platform monitoring gives an incomplete, sometimes misleading, picture of AI visibility, because each platform sources, weights and cites information differently. A brand can rank well in ChatGPT's answers while being nearly absent from Perplexity's citation list, or vice versa. Before I get to specific platforms, a caveat on the user figures below: these come from different companies, measured over different time windows (weekly versus monthly active users), self-reported in press releases rather than audited third-party data, and they are not directly comparable to each other. I'm including them for scale context only, not as a ranking.
| Platform | Reported users | Metric type | Date | Source |
|---|---|---|---|---|
| ChatGPT | 400M+ | Weekly active users | Feb 2025 | OpenAI, company statement |
| Gemini Apps | 400M+ | Monthly active users | May 2025 | Google, I/O 2025 keynote |
| Microsoft Copilot | 100M+ | Monthly active users | 2024 | Microsoft, company statement |
| Perplexity | Not directly comparable — no consistent MAU/WAU disclosure | — | — | — |
| Claude | Not directly comparable — no consistent MAU/WAU disclosure | — | — | — |
Agencies that only check ChatGPT are reporting a fraction of the picture. Genuine AI share-of-voice measurement means testing the same prompt set across all five surfaces, because a competitor might dominate Perplexity's citations while being nearly invisible on Copilot.
A practical note on UK-specific testing: setting a platform to "UK/en-GB" is not one switch. Depending on the platform, it can mean the account's registered region, the browser or device's language and locale settings, a UK-based IP address (relevant if you're running automation from a US-hosted server), whether the session is logged in or anonymous, and even which app store the mobile app was downloaded from. If you're not controlling for these consistently, you may be comparing runs that aren't actually comparable to each other.

Suggested alt text: "Grid comparing ChatGPT, Perplexity, Gemini, Copilot and Claude by primary use case for agency monitoring."
Before trusting any dashboard number, it's worth being explicit about what these tools are doing, because the methodology determines how much weight a client should put on the output — and this is the section most agencies skip, to their own detriment later.
Prompt generation isn't the same as customer research. Crawling a client's website and generating plausible questions from its content gives you a starting list, but it is not a substitute for knowing what real customers actually ask. Treat automated prompt generation as one input among several: pair it with Google Search Console query data, on-site search logs, PPC search term reports, CRM and sales-call language, and review content. Then have a human review, deduplicate and categorise the resulting list before it goes into daily monitoring. Skipping that review step is the most common way agencies end up tracking prompts nobody actually types.
Aim for a tight, deduplicated set of 15–30 prompts per client rather than 100 loosely related ones — duplicate or near-duplicate prompts inflate apparent coverage without adding real signal, and a smaller, well-curated set is easier to defend to a client who asks, "Where did these questions come from?"
The glossary terms below are useful, but they only become defensible reporting once you attach an explicit method to each one. Here's the model I'd suggest as a baseline, adjustable to your own tooling:
A sample prompt-level record, which is roughly what your underlying data should look like before it gets rolled up into a dashboard score:
| Field | Example |
|---|---|
| Prompt | "best vitamin C serum for sensitive skin" |
| Platform | Perplexity |
| Run date | 2025-03-14 |
| Model/version (if exposed) | Perplexity, default model |
| Brand appears? | Yes, position 3 |
| Competitor(s) appearing | Competitor A (position 1), Competitor B (position 2) |
| Citation or claim | Citation — linked to client's ingredient page |
| Sentiment | Neutral-positive |
| Reviewer | Analyst initials, for QA traceability |
Without a record structure like this, a "visibility score" is just a number with no audit trail, and I'd treat any tool or process that can't produce this level of detail on request with some scepticism.
None of this means the data is useless. It means a single day's snapshot is a sample, not a census, and the value comes from the trend line built over several weeks, not any one answer. A visibility score of 34/100 is an index for a specific prompt set on specific platforms over a specific window — it is not the brand's objective standing in some universal AI market, and I'd say that explicitly in every client report rather than let the number imply more precision than it has.
Here's the arithmetic that makes the case for automation. Assume ten core prompts, tested across five platforms, with one response logged per prompt per platform — that's 50 individual query-and-log actions per client per week (10 × 5 = 50, run once weekly as a baseline; daily monitoring multiplies this by seven). Each query-and-log action — reading the response, screenshotting it, recording sentiment, and noting which citations or claims appear — takes roughly three to four minutes for an experienced analyst. That's 2.5 to 3.5 hours of pure collection per client, per week, before any competitor analysis or interpretation happens. Across a portfolio of fifteen to twenty clients, that's close to a full-time role dedicated purely to collection, with zero strategic output — and that's before you've added competitor prompts, which roughly double the workload if you're tracking two named competitors per client.
| Manual Monitoring | Automated Monitoring | |
|---|---|---|
| Time per client/week | 2.5–4+ hours | Minutes (review only) |
| Platforms covered | Usually 1–2 (time-limited) | All 5 major platforms |
| Consistency | Single snapshot, session-dependent | Repeated daily runs |
| Sentiment tracking | Manual, subjective | Logged against a rubric |
| Competitor visibility | Ad hoc | Continuous, comparative |
| Scalability across clients | Poor | Built for it |
Daily automated runs matter for a reason beyond time saved: they turn a single anecdotal snapshot into a trend line, which is what lets you distinguish a genuine sentiment shift from ordinary model variability. Running daily doesn't eliminate sampling bias — it reduces the chance that one unrepresentative answer gets reported as fact, but the underlying limitations covered above still apply. It also means you're more likely to catch a negative shift or a new competitor mention before the client stumbles across it themselves, which is a better position to report from than reacting after the fact.
Scaling past a handful of accounts requires treating AI brand monitoring as a repeatable operational pipeline, not a bespoke task per client. In practice, that means:
Yes, largely, though "fully automated" oversells it slightly. Automation tools handle the repetitive part — running the same prompt set against multiple platforms on a schedule and logging the raw responses — typically via approved business APIs where available (OpenAI and Microsoft both offer commercial API access with different terms to their consumer apps) or browser-based automation where no API exists, which is more fragile and more exposed to platform terms-of-service changes. What doesn't automate cleanly yet is sentiment nuance, citation-accuracy checking, and strategic interpretation — those still need a human reviewer, which is why I'd frame this as "automated collection, human-reviewed output" rather than a fully hands-off system.

Suggested alt text: "Infographic contrasting manual chatbot checking with automated daily AI brand monitoring workflows."
Collecting the data is only half the job — the value an agency adds is packaging it into something a client can act on. Here's a worked, illustrative example (hypothetical, built from patterns I've seen across real accounts, not a single verified case study): imagine a mid-sized UK skincare brand where we tracked 24 deduplicated prompts across all five platforms daily for four weeks. Using the formula above (appearances ÷ total prompt-platform-day combinations × 100), the baseline visibility score came in at 34/100, against a named competitor at 61/100 — largely because that competitor's ingredient glossary pages were being cited repeatedly by Perplexity and ChatGPT's Search mode. After the client fixed structured data and added clearer product comparison pages — an AI legibility fix, not a guaranteed visibility fix — visibility rose to 48/100 over six weeks, and citations pointing to the client's own blog increased from two to eleven.
I'd flag that as a plausible, illustrative pattern rather than a guaranteed result. Technical fixes are correlated with improved citation rates in the accounts I've reviewed, but AI systems don't publish their ranking logic, so any causal claim here should be treated with real caution — "correlated with" is doing a lot of work in that sentence, deliberately.

Suggested alt text: "Mockup of an AI brand monitoring dashboard showing visibility score, sentiment trend and share of voice."
Most UK agencies I've spoken with land on one of two structures: a flat monthly add-on fee per client, or tiered pricing based on the number of tracked prompts and competitors — similar in logic to how keyword-tracking tiers are already priced in SEO retainers. The ranges below are illustrative starting points based on conversations with several UK agencies, not a benchmark survey, and your actual pricing should reflect your own tool costs, review time and client complexity.
| Tier | Prompts | Platforms | Reporting | Illustrative price/month | Setup fee |
|---|---|---|---|---|---|
| Starter | Up to 15 | All 5 | Monthly | £250–£400 | £150–£300 one-off |
| Growth | Up to 30 | All 5 | Weekly digest + competitor benchmarking | £600–£900 | £300–£500 one-off |
| Enterprise | Custom | All 5 + custom cohorts | Dashboard/API integration | £1,200+ | Scoped individually |
A worked margin example: say your monitoring tool costs £80/client/month once scaled across a portfolio, and analyst review time — reviewing flagged sentiment, checking citations, writing the digest commentary — runs 1.5 hours/week at a £40/hour blended rate, or roughly £240/month. Your all-in cost for a Growth-tier client is around £320/month. Priced at £750/month, that's a gross margin of roughly 57%, which is broadly in line with what agencies typically target on retained analytics services. If your own maths gives you a lower margin than your SEO or PR retainers for comparable effort, you're probably underpricing it — recalculate rather than guess.
Minimum contract periods matter here too: because the value case builds over four to six weeks of trend data, I'd avoid selling this on a month-to-month basis and instead set a minimum three-month term, which also protects your margin against the setup time.
Given that entry-level monitoring tools often offer low-cost trials, there's little reason not to pilot this with one or two willing clients before a portfolio-wide rollout. That lets you test your reporting format, refine your prompt sets, and build a case study before pricing it confidently across the board. I'd resist bundling it in for free as a value-add — clients place real perceived value on "AI reputation" right now precisely because it feels new to them, which supports a premium price point rather than a freebie.

Suggested alt text: "Comparison of Starter, Growth and Enterprise pricing tiers for AI brand monitoring as an agency service."
At minimum, track ChatGPT, Perplexity, Google Gemini and Microsoft Copilot, adding Claude where budget allows or where the client's audience skews toward research-heavy decision-making. Gemini's integration with Google Search makes it increasingly relevant to UK organic visibility conversations given Google's search share here, while Copilot matters most for B2B clients whose buyers work inside Microsoft-heavy organisations. I'd avoid assuming any single platform dominates UK consumer research without checking usage patterns specific to your client's sector, since this varies significantly by category.
Standardise the prompt-building template across accounts, centralise scheduling in one tool rather than running queries per analyst, set clear alert thresholds that trigger human review, and assign one named reviewer per account for sign-off. Automation removes the collection bottleneck; it doesn't remove the need for a consistent process across your whole portfolio, so build that process once and apply it everywhere rather than customising it per client.
Mostly, yes. Scheduled automation — via approved business APIs where available, or browser automation where they aren't — can handle running your prompt set across platforms and logging raw responses. What still needs a human is sentiment nuance, citation-accuracy checking against the source, and strategic interpretation of what the trend actually means for the client. Treat it as automated collection with human-reviewed output, not a fully hands-off system.
Treat any single query as a sample, not a definitive answer. Model versions update without notice, responses vary by wording and location, and platforms personalise by region — all of which means a trend built from daily runs over several weeks is far more trustworthy than any one screenshot. Report visibility scores with their denominator (number of prompts × platforms × days) attached, so clients understand they're looking at a directional signal for a defined prompt set, not a fixed market ranking.
Monitoring publicly available AI outputs about a client's own brand is generally a lower-risk activity than collecting personal data about individuals, but it's not risk-free, and "get consent" isn't automatically the right UK GDPR answer — consent is one lawful basis among several, and it may not even be the appropriate one here. What matters more in practice: document what you're logging (queries, responses, sentiment tags), where it's stored, how long you retain it, who has access, and whether any captured response contains personal data about named individuals (a competitor's staff member, a reviewer's name) that would trigger separate obligations. Build this into your service agreement as a data-processing clause rather than relying on an assumed consent, and take specific legal advice if a client's category is likely to surface personal data regularly — this article isn't a substitute for that advice.
Translate it into reporting language clients already know from SEO: a visibility score (0–100, always reported with its prompt/platform/date scope attached), share of voice against named competitors, and a sentiment trend over time. Pair this with concrete technical fixes — like AI legibility improvements — framed as correlated actions rather than guaranteed outcomes, so clients see AI visibility as an extension of existing marketing work rather than a confusing new discipline.
Daily runs give the most reliable trend line, aggregated into weekly reporting for clients. For UK accounts, set location, language and account region to UK/en-GB wherever the platform supports it, and check IP location if you're running automation from cloud infrastructure — several tools personalise citations by region, and a US-default query can return a meaningfully different picture than what a UK customer would actually see.
Rather than rolling this out across your whole portfolio at once, I'd run a contained pilot first:
Whichever monitoring tool you choose to support this, check that it covers all five major AI search platforms, refreshes at least daily, explains its scoring methodology in plain terms rather than a black-box number, supports UK/regional query settings, and offers API or export access so the data isn't trapped in someone else's dashboard.😊