What this blog covers

What the published evidence actually says about schema markup and AI citation, why platform statements and independent studies disagree, and how to decide what schema is worth deploying when the answer is genuinely uncertain.

What is schema markup, and what does it do for AI?

Schema markup is structured data added to a page in a standardised vocabulary that tells machines what the content is, who produced it, and when it was verified. It is written in JSON-LD in almost all modern implementations, and it describes things a human reader infers from context but a machine cannot reliably parse from prose alone.

For AI answer engines specifically, the claimed benefit is that structured data reduces the interpretive work a retrieval system has to do. A page that declares itself an article with a named author, a publication date and a set of question-and-answer pairs has told the system what it is, rather than requiring the system to infer it. Whether that declaration measurably increases citation rates is precisely the point in dispute. :contentReference[oaicite:0]{index=0}

The evidence points in two directions

Take the platform statements first, because they carry real weight even though they are not disinterested.

Google is more careful here than the industry usually reports. Its AI features documentation states plainly that you do not need to create new machine-readable files or markup to appear in those features, and that there is no special schema.org structured data to add. Its AI optimisation guide goes further: structured data is not required for generative AI search, and is worth continuing as part of overall SEO because it supports rich-result eligibility. Microsoft’s Bing team has taken a warmer position, saying schema helps large language models understand content in the context of Copilot. Both companies operate retrieval systems and both have reasons to shape publisher behaviour, which is worth noting without dismissing either statement.

Against that sits the independent evidence, and it is thinner than the confidence of the debate suggests. The Search/Atlas study of December 2024 found no correlation between schema coverage and citation rates. As Aimee Jurenka set out in Search Engine Land in March 2026, no peer-reviewed studies exist on the question, and the two companies whose engines drive much of the AI citation conversation – OpenAI and Perplexity – have never said publicly whether their pipelines preserve schema or discard it during ingestion.

There is also a live complication. Google withdrew FAQPage rich results from search on 7 May 2026, and the supporting documentation was retired the following month. A significant number of brands read that as instruction to remove FAQ markup entirely. The visual treatment ended; the markup remains valid, and it continues to describe a page’s question-and-answer structure to any system that chooses to read it. Removing it on the strength of a rich-result deprecation confuses a search feature decision with a machine-readability decision, and it is worth revisiting any older guidance on FAQ pages written before that change.

Jurenka’s framing is the one we work to: schema is infrastructure rather than a lever. Infrastructure does not produce visible results on its own, and its absence produces failures that are hard to attribute.

What schema does that is not disputed

Setting the citation question aside, several effects of structured data are not seriously contested, and they are sufficient to justify a baseline deployment on their own.

Schema establishes entity identity. Organisation markup states what the brand is, what it is called, where it operates and how it connects to its other properties. For any system building an entity graph – which includes Google’s, and which feeds its AI surfaces – that is the difference between a stable object to attach authority to and a set of loosely related pages. This is the least glamorous and most consequential use of markup.

Schema communicates authorship and recency. Article markup declares who wrote a piece and when it was last genuinely revised. Given that retrieval systems demonstrably deprioritise stale content, and given that expertise signals are weighted in quality assessment, declaring both explicitly removes ambiguity that would otherwise have to be inferred.

Schema still drives rich results for the types Google supports. That list changes – Practice Problem support was removed in January 2026, FAQPage rich results in May – so it needs periodic review rather than a one-time build.

And schema is close to free once the pattern is established. Deployment is a one-off engineering cost with minimal maintenance. When an intervention is cheap, reversible and recommended by the platform that owns the largest answer surface even as it declines to promise a citation benefit, the burden of proof for skipping it is higher than the burden for doing it.

Framework: The Schema Value Ladder

Our Search Intelligence practice separates schema work into three tiers by how well the evidence supports it, and funds them differently. We call it the Schema Value Ladder.

Tier What it covers Evidence position How to fund it
Tier 1 – Established Organisation, Product, Article, Breadcrumb. Entity identity, authorship, recency, site structure. Strong. Supported by platform statements, by rich result mechanics, and by entity graph behaviour independent of AI. Deploy across all key pages as baseline hygiene. Not a discretionary line.
Tier 2 – Plausible FAQPage, HowTo, Speakable, Dataset, Person for author credentials. Mixed. Bing says schema helps AI understanding; Google says it is not required but worth keeping; one independent study found no citation correlation. Cheap to deploy. Deploy where the content genuinely has that structure. Do not restructure content to fit a schema type.
Tier 3 – Unsupported llms.txt and similar manifest files, schema volume as a target, marking up content that does not have the declared structure. Weak to negative. A seven-month server log study across around 900 domains recorded no requests for llms.txt from GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Do not fund as a deliverable. Deploy llms.txt if you wish – it costs nothing – but do not count it as work.

The ladder exists to stop two opposite failures: treating schema as a growth lever that will produce citations on its own, and treating one negative study as grounds to abandon markup entirely.

The ladder explained

Tier one is not really an AI decision: Organisation and Product markup earn their place through entity clarity regardless of what any answer engine does with them, which is why they should be deployed before the citation question is even asked. In audits we find brands with sophisticated content programmes and no Organisation schema anywhere, which means every system reading the site has to infer the brand’s identity from prose.

Tier two is where honest uncertainty lives: FAQPage markup is the clearest case. It no longer produces a rich result in Google search. Microsoft says structured data helps AI systems understand content; Google says it is not required for its AI features while recommending you keep it for rich-result eligibility; one independent study found no citation correlation. The markup costs almost nothing to maintain. On that balance we keep it, and we tell clients plainly that we are keeping it on a judgement about asymmetric cost rather than on proof of effect. The discipline that matters in tier two is honesty about the content: mark up question-and-answer pairs where the page genuinely contains them, and do not manufacture a FAQ block to justify the markup.

Tier three is where budget gets wasted: The clearest example is llms.txt, a proposed manifest file intended to tell AI systems what a site contains. A seven-month server log study across around 900 domains, published in August 2026, recorded 1,227 requests for llms.txt files – of which none came from GPTBot, ClaudeBot, PerplexityBot or Google-Extended. Nearly two thirds came from a single commercial data aggregator. John Mueller of Google has said a self-reported manifest cannot function as a differentiator between sites, and that Search ignores it. Google’s own AI features documentation says the same thing in different words: no AI text files are needed. One vendor study claims a twenty-four per cent citation lift from a valid llms.txt file, but it is correlational with undisclosed methodology, and the log evidence is materially stronger. Deploy the file if you like, because it costs nothing. It should not appear on an invoice.

The same tier-three logic applies to schema volume as a metric. Marking up more page types does not increase citation probability; it increases the surface area for markup errors and, where the declared structure does not match the content, introduces a trust problem rather than solving one.

Real-world scenario: Kurlon and Converse

Two programmes are worth reading together, because neither isolates schema as a variable and it would be dishonest to present either as proof that markup alone produces citations.

Kurlon carried none of the structured data signals – Organisation, Product, FAQPage, BlogPosting – that let a model parse and trust a brand. Full schema deployment was one of four moves, alongside keyword mapping across more than two hundred terms to remove cannibalisation, intent-driven FAQ modules and comparison content, and activation of reviews, ratings and diversified authoritative backlinks. Across seven months the programme produced 765% growth in AI Overview visibility, 600% growth in brand mentions across AI surfaces, and a 67% rise in search impressions.

Converse presented a narrower and more instructive technical problem. A backend configuration was intercepting meta descriptions before they reached the page source, so search engines were writing their own summaries for every page, and identical metadata ran across the entire product catalogue. Alongside fixing that, L&F deployed Organisation, Product, FAQ and Website schema explicitly to qualify pages for rich results and AI Overview inclusion. Across seven months: 35% growth in organic clicks, 47% growth in transactions, 32% growth in organic revenue, and average position improving from 14.7 to 9.3.

What these show is schema working as part of a system where the other parts were also fixed. What they cannot show, and what no available public evidence shows, is schema producing citations on a page whose content, entity data and external credibility were unchanged. That distinction is worth keeping in any proposal that promises otherwise.

Read the full Kurlon case study and the Converse programme.

Kurlon x Lyxel&Flamingo, March to October 2025. Converse x Lyxel&Flamingo, June 2025 to January 2026.

Going deeper: the schema decision list

Nine checks that separate markup worth funding from markup worth skipping.

  • Is Organisation schema deployed and validating on your homepage and key pages? If not, start here regardless of your view on AI citation.
  • Does your brand name, legal entity and contact information match exactly across your site, your business profiles and major third-party listings?
  • Is Article schema declaring a named author and a genuine last-modified date on every editorial page?
  • Do your FAQ blocks contain questions your buyers actually ask, or questions written to justify the markup?
  • Have you kept FAQPage markup after the May 2026 rich result withdrawal? The markup remains valid and readable.
  • Is Product schema accurate on price and availability, particularly for the Indian market rather than a global default?
  • Are you funding llms.txt as a deliverable? Stop. Deploy it if you want; do not pay for it.
  • Does your schema declare a structure the page genuinely has? Mismatched markup is a trust problem, not a neutral one.
  • When did you last review which schema types Google still supports? The list changed twice in the first half of 2026.

Key takeaways

  • Google documents that structured data is not required for its AI features and that no special schema.org markup exists for them, while recommending you keep it for rich-result eligibility. Microsoft says schema helps LLMs understand content for Copilot. A Search/Atlas study from December 2024 found no correlation between schema coverage and citation rates.
  • Neither OpenAI nor Perplexity has disclosed whether their retrieval pipelines preserve schema, which means confident claims in either direction are running ahead of the evidence.
  • Google withdrew FAQPage rich results on 7 May 2026. The markup remains valid and machine-readable, so removing it treats a search feature decision as a machine-readability decision.
  • Tier one markup – Organisation, Product, Article, Breadcrumb – earns its place on entity clarity alone, independent of any AI citation effect.
  • A seven-month server log study across around 900 domains found no llms.txt requests from GPTBot, ClaudeBot, PerplexityBot or Google-Extended (digitalapplied, 2026). It should not be a paid deliverable.

The CXO takeaway

The useful conclusion here is uncomfortable for both sides of the argument. Schema is almost certainly worth deploying, and almost certainly will not produce the citation gains it is frequently sold as producing.

That combination has a practical consequence for how proposals should be read. A supplier presenting schema deployment as the central mechanism of an AI visibility programme has mistaken infrastructure for strategy. A supplier who has removed FAQ markup because rich results were withdrawn has mistaken a feature deprecation for a technical instruction. Neither is a small error, because both redirect budget away from the things that do reliably move citation – entity accuracy, independent third-party credibility, and content that answers a question completely and originally.

When the evidence on an intervention is genuinely contested, the right response is to fund it at the level its cost justifies rather than at the level its marketing suggests, and to say plainly which parts of the programme rest on judgement rather than proof. Very little content in this category does that, which is itself a reason to be sceptical of most of it.

The open question worth watching is whether OpenAI or Perplexity ever disclose how their ingestion handles structured data. Google has now told us plainly that markup is not required for its AI features. Until the other two say anything at all, anyone claiming certainty about schema and AI citation is telling you more about their sales process than about the systems.

Frequently Asked Questions

Does schema markup help you get cited by ChatGPT?

The honest answer is that nobody outside OpenAI knows. OpenAI has never disclosed whether its retrieval pipeline preserves schema. Microsoft has said structured data helps its AI systems understand content, and Google says markup is not required for its own AI features though it remains worth keeping for rich results. Neither statement covers ChatGPT. Deploy schema for entity clarity and accept that its effect on ChatGPT citation is unproven.

Should I remove FAQ schema now that Google dropped FAQ rich results?

No. Google withdrew the visual rich result on 7 May 2026 and retired the supporting documentation the following month, but the markup remains valid and continues to describe a page's question-and-answer structure to any system that reads it. Removing it treats a search feature decision as a machine-readability decision.

Which schema types matter most for AI search?

Organisation and Product carry the most reliable value because they establish entity identity, which every retrieval system depends on. Article markup declaring a named author and a genuine last-modified date supports expertise and freshness signals. FAQPage and HowTo are worth deploying where the content genuinely has that structure, on a cost-versus-uncertainty judgement rather than on proof.

Do I need an llms.txt file?

There is no evidence that it works. A seven-month server log study across around 900 domains recorded no requests for llms.txt from GPTBot, ClaudeBot, PerplexityBot or Google-Extended; John Mueller of Google has said Search ignores it; and Google's AI features documentation states that no AI text files are needed. Deploy it if you wish, since it costs nothing, but it should not be a paid deliverable.

Can schema markup alone improve AI visibility?