Structured Data Markup for AI Answer Eligibility
Schema markup now determines AI citation eligibility, not search visibility.

Structured data's job changed sometime around March 2026, and most sites still haven't caught up. The markup that used to earn a star rating or an FAQ dropdown on Google now does something different: it tells an AI system who you are, whether your claims hold up, and whether it should cite you. That's a harder job, and most of the schema built for the old purpose flatly fails at the new one.
For years, structured data's whole pitch was visual. Mark up your review content, get stars next to your listing. Mark up an FAQ, get a dropdown that eats more space on the results page than the listing sitting underneath it. The incentive was rich results, plain and simple, and site owners chased them accordingly.
That incentive is mostly gone now. FAQPage rich results launched in May 2019, got restricted to government and health sites in August 2023, and were pulled entirely on May 7, 2026. HowTo rich results followed a similar arc, with Google phasing out desktop and then mobile display before retiring the format entirely. Two of the most commonly implemented schema types, gone from the visual layer of search.
Then came the March 2026 core update, which finished rolling out on April 8. FAQ rich result impressions across tracked sites dropped 47% in its wake, and the update's real target was a habit, not a schema type: tagging content that wasn't the actual point of the page. Sites had been marking up sidebar text, boilerplate disclaimers, anything with a question mark, just to angle for the dropdown. Google shut that down. What survives has to earn its keep somewhere other than the visual layer of search, and that somewhere is AI answer generation.
AI Answer Engines' Use of Structured Data, and Where They Diverge
Strip the mechanism down to its parts. Schema takes content an AI system would otherwise have to interpret from raw text, full of ambiguity and inference, and turns it into labeled entities the system can pull out with confidence. A paragraph reading "founded in 2014 by a former Stripe engineer" needs parsing. A schema property stating foundingDate: 2014 needs none. Less interpretation means fewer errors, and fewer errors means an AI system trusts the claim enough to cite it.
Not every AI system reads schema the same way, though, and that gap affects how confidently AI systems attribute citations, a risk most implementation guides don't address. JSON-LD placed in the document head is still Google's preferred format after March 2026, and neither Microdata nor RDFa has shown any real edge in AI citation performance by comparison. Google's own AI surfaces, AI Overviews and AI Mode, actively use schema when generating answers, leveraging it to verify claims and resolve which entity is being referenced.
Third-party LLMs don't necessarily bother. ChatGPT, Claude, and Perplexity do not reliably parse JSON-LD semantically during their own retrieval passes. Testing shows some of these systems pull the plain text sitting inside a JSON-LD block and treat it like ordinary page copy rather than structured data, but that's a fallback, not something to build a strategy around. A page can be flawlessly marked up for Google's crawler and still get read as plain unstructured text by Perplexity. That's a platform difference, not a bug in anyone's implementation, and it means no one schema strategy can assume one system's behavior speaks for all of them.
The schema types that move the needle for AI citation in 2026
Judged strictly on AI visibility, apart from any rich-result ambition, only a handful of schema types actually carry weight. Most others are noise.
Organization schema anchors brand identity: name, official URL, logo, contact details, and how the entity connects to others in its space. If it is skipped, an AI system has to guess who's actually behind the content, and every guess lowers citation confidence. The sameAs property carries most of the load here, linking out to LinkedIn, Wikipedia, Crunchbase, X, Facebook, a Google Business Profile. Each link gives an AI engine another place to cross-check identity, and consistency across those sources builds trust the way a paper trail builds a fraud case. knowsAbout adds a second layer, flagging subject-matter expertise directly so an AI system can identify topical authority when resolving a source. sameAs identifiers pointing to Wikidata, Crunchbase, or GRID sharpen Knowledge Graph recognition enough that sites with clean entity schema get cited more often, simply because the AI can resolve who the source is without guesswork.
Person and Author schema does the same job at the individual level, reinforcing the expertise and trust signals that fall under the E-E-A-T framework. The properties that matter: jobTitle, worksFor, sameAs (pointing to LinkedIn and published work), alumniOf, hasCredential, and knowsAbout. Google's AI Overviews and Perplexity both weigh author expertise heavily on YMYL topics, the "your money or your life" categories where a wrong answer carries real cost. The bigger gain comes from wiring these pieces together: an Article tied to an Author tied to an Organization, using @graph and @id inside one JSON-LD block, builds something close to an internal knowledge graph. Three schema blocks sitting on a page with no reference to each other don't add up to that. They're three disconnected labels, full stop.
Article, NewsArticle, and BlogPosting schema round this out by telling an AI system what a piece of content actually is: who wrote it, when it went live, when it last changed. dateModified gets skipped constantly, but it matters, since AI Overviews and AI Mode both favor content that reads as recent and clearly attributed. Tie the author field to a Person schema, and tie that to an author bio page, and the credibility layer compounds instead of sitting flat.
The implementation mistakes that make technically valid schema strategically useless
A page can pass Google's Rich Results Test with zero errors and still do almost nothing for AI visibility. Validity and strategic value are separate questions, and most schema audits only check the first one.
The most damaging mistake, by a wide margin, is markup that doesn't match the page. AI systems check for consistency between what the schema claims and what the visible content actually says. When schema describes something that isn't there, or tags something that isn't the page's real subject, that mismatch reads as a trust problem. This pattern drew scrutiny in the March 2026 core update period. Content-schema parity isn't a best practice anymore, it's an enforcement surface, and sites that treated schema as a tagging exercise disconnected from what the page actually delivers got hit for it.
Disconnected schema blocks are the second failure mode. Article, Author, and Organization schema sitting on the same page without ever referencing one another is far less legible to an AI system than a properly linked graph. The fix is mechanical: use @graph and @id in JSON-LD so entities point directly at each other. An entity linked by @id is traversable. Three separate blocks stay three separate blocks no matter how well each one validates on its own.
Then there's staleness. Prices, dates, and version numbers left untouched for months read as neglect. Retrieval-based AI systems favor fresh data, and stale markup signals a lower-reliability source. Perplexity carries the strongest recency bias of any AI search platform tracked, and a FirstMotion analysis cited by writer.com found that for fast-moving queries, content older than 90 days starts entering a decay window, losing retrieval priority to newer pages.
Limits of Schema Alone for AI Visibility Problems Rooted in Citation Sourcing
AI-generated answers cite third-party content 91% of the time, not brand-owned websites, and that single number should reorder most of what a company spends its schema budget on. An AirOps analysis cited by writer.com found that 85% of brand mentions inside AI search responses come from third-party pages. A brand is far more likely to get cited through someone else's domain than through its own. Schema on a company's own site cannot touch that gap, because the gap is a sourcing problem. It's a sourcing problem, full stop, and no amount of JSON-LD fixes where the citation actually comes from.
That's why schema deployed only on brand-owned pages addresses such a narrow slice of where AI citations originate. An Ahrefs study covering 75,000 brands found brand web mentions correlate with AI Overview visibility at 0.664, against just 0.218 for backlinks, a gap wide enough to show that off-site presence, not on-page markup, is the dominant predictor here.
Platform behavior fragments the picture even further. Each AI system draws from a different pool of preferred sources, meaning a citation that appears in one platform's answers may not surface in another's at all. The citation landscape across AI systems is highly fragmented, with no small set of domains dominating across platforms. That's about as scattered as a citation landscape gets.
Building the citation surface schema alone can't reach: third-party content and entity presence
Distribution, not content quality on its own, is what actually moves the citation number. Publishing the same article across third-party publisher networks, rather than keeping it locked inside a brand's own domain, produced a 239% median increase in AI citations in one analysis. A separate pilot saw citation rates climb from 7.7% to 34%, a substantial lift, with distribution as the only variable that changed.
Grounding pages are one concrete way to act on that. A grounding page is a dedicated fact page built so a specific claim can be found and cited the instant a model runs a retrieval step, and it's a different animal from a general content page written to rank or to inform a reader. The distinction is visible in the On-Model versus Off-Model framing: On-Model SEO covers what a model already knows without live search, while Off-Model SEO covers what's externally checkable the moment the model needs to verify something. Grounding pages live entirely in that Off-Model layer, and schema on those pages raises the model's confidence that it's attributing the right claim to the right source.
Entity presence across third-party platforms fills in the rest of the gap. sameAs links to LinkedIn, Wikipedia, Crunchbase, Wikidata, G2, and Capterra each work as a potential citation source for a different engine: Microsoft Copilot leans hard on LinkedIn for B2B queries, while Perplexity tends to favor third-party and authoritative sources over brand-published content, consistent with its strong recency and sourcing biases. Platforms that measure AI visibility, tracking whether ChatGPT, Claude, Gemini, and Perplexity actually name and cite a source, consistently report that clean Organization and Person schema, particularly sameAs links to verified identities, correlate with higher citation rates. Letterstory is one option among the services working in this space, and the logic is straightforward: markup like this works because it lets an AI system resolve identity with confidence instead of guessing. Getting the right entity onto the right platform, for the right engine, is the part schema was never built to handle by itself.


