The AI Legibility Checklist Every E-commerce Site Needs (And Why Most Fail It)

AI Legibility for E-commerce: The 16-Point Checklist for UK Retailers
Meta description: Use this UK e-commerce AI legibility checklist to help ChatGPT, Gemini and Perplexity find, understand, trust and cite your product pages.

AI legibility measures how easily generative AI systems such as ChatGPT, Gemini and Perplexity can crawl, parse and accurately extract information from your e-commerce website. For UK retailers, it is becoming an important part of AI SEO, alongside traditional search engine optimisation.
In the technical audits I run across online stores through MentionOwl — a rolling sample of roughly 140 UK and EU retail sites audited between late 2023 and mid-2024, using a fixed 16-point checklist rather than a scientific random sample — a clear majority fail at least five of those sixteen checks. I want to be upfront that this is an internal observation from my own client and audit base, not a peer-reviewed study of the wider market. AI visibility is also difficult to measure consistently because every provider's retrieval system behaves differently and changes without notice. Even with that caveat, the pattern is consistent enough across the sites I see that I think it is worth treating as a serious operational risk rather than a curiosity.
When I say a site “fails” a check, I mean something concrete: an AI system attempting to answer a shopper's question — something as mundane as “which running shoe is best for wide feet under £100?” — either cannot locate the retailer's product, pulls stale pricing, or states an attribute that is not supported anywhere on the page. These are not hypothetical failure modes. They are common problems for the UK online retailers I have audited, and they matter because AI-assisted shopping is a growing discovery channel, even though reliable, independent UK-market sizing for exactly how much commercial traffic it currently represents is still limited.
What Is AI Legibility for E-commerce Product Pages?
It helps to separate AI legibility from traditional SEO crawlability right away, because conflating the two is where many retailers go wrong. Traditional SEO is largely concerned with whether Googlebot can index a URL and rank it favourably against a query. AI legibility is more granular: can a generative system extract a discrete, verifiable fact from your page and use it confidently in a generated answer?
This distinction matters because of how many AI systems consume content, although it is important not to describe one architecture as universal across every provider. Many retrieval-augmented systems chunk content into smaller passages, convert those passages into embeddings, and retrieve the most semantically relevant chunks when constructing a response. However, the exact mechanics differ between ChatGPT, Gemini, Claude and Perplexity, and none of the major providers publish full technical detail on how they weight or verify sources.
What I can say with more confidence, because it is observable rather than inferred, is the outcome: a page that ranks brilliantly in Google because of strong backlinks and page authority can still be functionally invisible to an AI assistant if the actual product facts — price, availability, materials and dimensions — are buried in a JavaScript-rendered widget, scattered across vague marketing copy, or contradicted elsewhere on the same domain.
Google's documentation on product structured data describes its purpose as helping systems reliably identify the product, understand its attributes and commercial terms, verify those facts from authoritative sources, and present them without having to infer missing information. That guidance is written primarily with Google Search and Shopping surfaces in mind, not ChatGPT or Perplexity directly, but the underlying principle transfers well: inference is the enemy. Every time a system has to guess a fact because your page did not state it plainly, you introduce the risk of a misquote, an outdated price, or a competitor being cited instead.
In the visibility scoring work we do at MentionOwl, I have observed a correlation between a site's technical legibility score and how often it gets cited across ChatGPT, Gemini, Claude and Perplexity for relevant purchase queries. I am describing a correlation from my own client data, not a causal mechanism I can prove. I do not have access to any provider's internal ranking logic, and neither does anyone outside those companies.
What I can say is that sites passing most of the sixteen checks tend to show up with specific, accurate details — the correct price, the right stock status and genuine review counts — while sites failing the majority either vanish from the answer or appear with garbled, outdated or fabricated specifics. This is why a page can rank number one on Google and still never be mentioned by an AI assistant: ranking reflects authority and relevance signals built over time, while AI citation reflects whether the retrieval system could extract a clean fact at query time.

The 16-Point AI Legibility Checklist for UK Retailers
This is the checklist I run during a MentionOwl audit, grouped into five categories. I would treat the first two categories — access and product identity — as the foundation, because nothing downstream matters if an AI crawler cannot reach or correctly identify your product in the first place.
| # | Check | Category | Why It Matters | How to Test |
|---|---|---|---|---|
| 1 | Robots.txt allows AI crawlers | Access | Blocking GPTBot, ClaudeBot, PerplexityBot or Google-Extended removes you from that provider's training or retrieval access entirely | Read robots.txt directly; check for disallow rules under these user agents |
| 2 | Correct HTTP status codes | Access | Soft 404s or redirect chains waste crawl budget and can cause a product to be dropped silently | Crawl your top URLs with a status-code checker |
| 3 | Core facts present in raw HTML | Access | Many, although not all, AI crawlers do not execute JavaScript as a browser does, so client-rendered facts may not exist for them | Fetch the page with JavaScript disabled and read the raw response |
| 4 | Rendered DOM matches raw HTML | Access | A mismatch means you cannot be sure which version a crawler actually sees | Compare “view source” output with the browser-rendered DOM |
| 5 | Canonical URLs set correctly | Product Identity | Variant and filter URLs without canonicals create ambiguity about which page is authoritative | Check canonical tags across colour and size variants |
| 6 | Stable SKU/GTIN in schema | Product Identity | Gives the system an unambiguous entity to anchor facts to | Validate Product schema and confirm SKU/GTIN populate |
| 7 | Consistent brand and product naming | Product Identity | Name mismatches between the page, schema and feed can fragment the entity into several “different” products | Compare the on-page title, schema name and feed title |
| 8 | Price in static HTML and Offer schema | Commercial Facts | This is the most common failure I see and one of the most likely to cause an incorrectly quoted price | Fetch with JavaScript disabled; check for a numeric price and matching Offer schema |
| 9 | Availability and stock in static HTML and Offer schema | Commercial Facts | An AI system citing “in stock” for a sold-out item is a direct trust failure for the shopper | Use a JavaScript-disabled fetch; check the availability field |
| 10 | Price and stock match your product feed | Commercial Facts | Contradictions between schema, page and Merchant Center feed are themselves a legibility failure | Spot-check a sample against your live feed export |
| 11 | Distinctive, non-templated copy | Content | Identical templated descriptions give a system no reason to prefer your listing over a similar competitor | Read five of your pages against five competitor pages |
| 12 | Specifications available as text | Content | A dimension chart saved as a JPEG is invisible to text extraction, regardless of how clear it is to a human | Confirm specification tables exist in HTML, not only as images |
| 13 | FAQ content answers real pre-purchase questions | Content | This maps well to conversational queries such as “will this sofa fit through a narrow doorway?” | Compare your FAQ with actual customer service queries |
| 14 | Genuine reviews present in static HTML | Trust Signals | Review evidence loaded only through client-side JavaScript may not be seen by every crawler | Fetch with JavaScript disabled and check whether review text appears |
| 15 | Review counts and ratings are consistent | Trust Signals | A mismatch between schema rating and visible rating is a contradiction, not a feature | Compare the AggregateRating schema with the number shown on-page |
| 16 | Structured data validates and matches visible content | Trust Signals | Markup that contradicts the page is treated as unreliable, not as a shortcut | Run Schema.org's validator, then manually compare each field with the page |
From the audits behind these sixteen checks, I would estimate that items 3, 8, 9 and 14 — all variants of “the fact exists but only after JavaScript runs” — account for a large share of the failures I find. That is a pattern I am confident in because it is directly observable in raw HTML, not something inferred from AI output.
Structured Data and AI Crawlers: What Product Schema Can Do
Structured data is where I see the widest gap between what retailers believe they have implemented and what is actually functioning. Google's documentation is clear that Product, Offer and AggregateRating schema are the fields most likely to support enhanced product understanding on Google's own surfaces. They can communicate price, currency, availability, condition, SKU, GTIN and review aggregates.
What is less clear is exactly how much weight ChatGPT, Gemini, Claude or Perplexity individually give to this markup when generating an answer, since none of those providers publish that detail. Schema is a well-documented signal for search and shopping surfaces. Its direct influence on generative citations is something I can observe correlating with better outcomes in my own tracking, but it is not something I can prove mechanistically for every provider.
| Schema Type | Documented Search Use | Possible AI Benefit | Caution |
|---|---|---|---|
| Product (name, SKU, GTIN, brand) | Strong | Establishes a stable entity identity that a system can anchor facts to | Must match visible page content and feed exactly |
| Offer (price, currency, availability) | Strong | Provides an exact commercial fact for shopper queries | Stale schema prices are worse than no schema at all |
| AggregateRating / Review | Strong, when genuine | Supplies quotable, specific social proof | UK advertising rules, including ASA/CAP requirements, require review data to be genuine and current |
| FAQ schema | Moderate | Conversational structure may align with question-style queries | The underlying FAQ copy matters more than the markup wrapper |
| HowTo schema | Moderate for instructional queries | Step structure provides a ready-made format to cite | Only add it where genuine step-by-step content exists |
| Generic Article schema on category pages | Weak | Too broad to extract a single product fact from | Rarely worth prioritising over Product and Offer schema |
| Decorative or visual-only markup | None | No machine-readable fact is attached | Purely presentational; it does not aid legibility |
FAQ and HowTo content are, in my experience, underused by UK e-commerce sites, regardless of whether the schema wrapper is present. Shoppers increasingly phrase queries conversationally — “will this sofa fit through a narrow doorway?” or “is this moisturiser safe for sensitive skin?” — and a page that answers those questions in plain, specific text gives an AI system something concrete to retrieve.
I would separate the two things clearly: the content answering the question is doing the real work; the schema is a bonus layer, not a substitute for writing the answer in the first place.
It is also worth being direct about what schema cannot do. Google states plainly that structured data does not guarantee a page will receive rich results. The same caution applies to AI citation: adding a large JSON-LD block is not a shortcut. The markup still has to match the visible page, and it still has to agree with your Merchant Center feed if you run one for Google Shopping in the UK.
A contradiction between your schema, rendered HTML and feed is itself a legibility failure, not a clever workaround. It is one of the more common problems I find in checks 10, 15 and 16 above.

Common E-commerce Mistakes That Block AI Access
The sixteen-point checklist covers what to test. In practice, I find the failures cluster around a small number of root causes, repeated across almost every audit:
- JavaScript-rendered pricing and stock status. This is consistently the most damaging pattern I find. Not every AI crawler fails to execute JavaScript, and provider capabilities change, but a meaningful share currently relies on the initial server response rather than a fully rendered page. If a fact only exists after a client-side API call fires, you are relying on crawler behaviour you do not control.
- Descriptions hidden behind “read more” toggles or image-only specification sheets. A dimension chart saved as a JPEG is invisible to text extraction, no matter how clear it looks to a human shopper.
- Missing or inconsistent canonical URLs. This is particularly common where UK retailers generate dozens of near-duplicate variant URLs for colour and size without canonical tags.
- Robots.txt or bot-management rules accidentally blocking AI user agents. I regularly find sites that have blocked GPTBot or PerplexityBot through an overly broad rule, often without the marketing team knowing. Remember that confirming access in robots.txt means a crawler is permitted; it does not prove that it is visiting or that a visit led to a citation.
- Thin or templated copy. Generic product descriptions give an AI system little reason to prefer your listing over a near-identical competitor.
- Review and rating data rendered only through JavaScript. This can remove genuine social proof from the version of the page a crawler receives.
None of these are exotic problems. They are the product of storefronts built primarily for human browsing convenience, with machine readability treated as incidental.
How to Run an AI Legibility Audit: A Six-Step Workflow
The sixteen checks above explain what to test. This is the workflow I would use to apply them, with results recorded in a simple spreadsheet containing the URL, check number, pass or fail status and fix owner.
- Fetch your top 10 product pages with JavaScript disabled. This covers checks 3, 4, 8, 9 and 14. Use a browser extension or simple fetch tool to see exactly what is in the raw HTML response, then compare it with what renders in a normal browser.
- Validate your structured data. Use Schema.org's validator, then manually cross-check the fields against checks 5, 6, 10, 15 and 16. The schema must match the visible page and your feed, not merely pass validation syntactically.
- Ask ChatGPT, Gemini and Perplexity direct product questions. Use at least ten realistic, non-branded queries that a shopper might type. Log whether your brand appears, is cited correctly or is misquoted. Treat this as a citation spot-check, not proof of crawl access: a wrong answer could mean the page was not retrieved, was retrieved and misread, or was correctly read and inaccurately summarised.
- Check robots.txt and, if accessible, server logs. Look for GPTBot, ClaudeBot, PerplexityBot and Google-Extended activity. Confirm access is allowed and that visits are occurring, while remembering that a logged visit does not guarantee the content was used in an AI-generated answer.
- Audit your product copy for distinctiveness. Read five competitor pages alongside your own and ask whether an AI system would have a reason to quote you specifically rather than paraphrase a rival.
- Run the complete 16-point AI legibility audit. Use the checklist above manually or through an automated tool to identify structural contradictions between your feed, schema and rendered page.

How to Monitor AI Visibility and Generative Engine Optimisation Improvements
A one-off audit tells you where you stand today. It tells you nothing about tomorrow, which is a significant limitation given how frequently AI providers can change their retrieval behaviour.
A model provider can change how it weights sources, adjust its chunking strategy or alter which crawler feeds its index. Your visibility can therefore move without any change on your site. This is why I built MentionOwl's scoring around continuous measurement rather than a single snapshot.
Four metrics are worth tracking separately because they answer different questions:
- Query coverage: how many realistic shopper questions surface your brand at all.
- Position-weighted citations: whether you are the primary source or an afterthought in an answer that mentions you.
- Share of voice: how you compare with named competitors across the same queries.
- Sentiment: whether the AI represents your brand favourably when it mentions you.
Avoid over-reading single-day movements. Query phrasing varies, models are updated without notice and a small sample can look noisy even when nothing has changed. I generally trust a trend across a week or two more than a single day's result, and I would recommend the same for anyone monitoring AI visibility manually.
Structural fixes can move faster than people expect. I have seen clients expose price and stock in static HTML and observe citation accuracy shift within roughly one to two weeks of re-testing. This is not universal or guaranteed, but it is frequent enough that I would not treat these improvements as a slow-burn SEO-only investment.
Competitor tracking matters too. If a rival improves its AI legibility while you are not monitoring the market, its share of voice can grow at your expense even if your own site has not changed. That is the main argument for reviewing AI visibility weekly rather than quarterly.

Frequently Asked Questions About AI Legibility
What is AI legibility exactly?
AI legibility is how easily a generative AI system can access, parse and accurately extract facts from your website, particularly pricing, availability and specifications. It is distinct from traditional SEO because it is less about ranking and more about whether a system can confidently cite your content when answering a shopper's question.
A quick test is to fetch a product page with JavaScript disabled and check whether price and stock are still visible in the raw response.
Does product schema markup matter for AI SEO?
Product schema is a well-documented signal for search and shopping surfaces. In the audits I have run, sites with complete Product, Offer and AggregateRating schema tend to be cited more consistently by AI shopping assistants than sites relying on plain text alone.
I would describe this as a strong correlation observed in my own data rather than a guaranteed mechanism, since no AI provider publishes exactly how it weights schema. In all cases, schema must match the visible page and your product feed. Mismatched schema is worse than no schema.
How do I check if AI can read my e-commerce site?
Fetch your pages with JavaScript disabled and check whether price, stock and specifications survive in the raw HTML. A meaningful share of AI crawlers relies on that initial response rather than a fully rendered page.
Then check robots.txt for anything blocking GPTBot, ClaudeBot or PerplexityBot, and review server logs to confirm whether they are visiting. This tells you about access and raw content, not whether a specific visit led to a citation. To assess that, test realistic shopper queries directly in the AI tools.
What is the quickest fix for an AI legibility problem?
In most of the audits I have run, the highest-impact fix is moving pricing and stock availability out of JavaScript-only rendering and into static HTML with accurate, matching Offer schema.
I have seen this move sites from invisible to cited within one to two weeks of re-testing in several cases. Results vary by site, however, and the timeline is not guaranteed because it depends on how often the relevant AI provider re-crawls and re-indexes your pages.
Related Articles

Why Agencies Should Offer AI Visibility Monitoring in 2025: The Business Case for a New Retainer
Learn how agencies can turn brand monitoring into an AI visibility retainer with practical deliverables, pricing, tools, and a plan to win clients ear

Turning AI Visibility Data Into a Winning QBR Presentation: A Marketing Team's Guide
Turn AI visibility and share of voice data into a leadership-ready QBR narrative with metrics, competitor benchmarks, and next-quarter goals.
Create content like this automatically
Scribe uses AI to generate high-quality blog posts that engage your audience and drive traffic.
