Why Data Integrity Is The New Technical SEO: From Crawling To Trust

Why Data Integrity Is The New Technical SEO: From Crawling To Trust

Previously two years, Google has dropped assist for 9 ItemTypes from its wealthy end result search gallery. This has occurred not too lengthy after ChatGPT’s launch and when mass adoption started:

Rich results supported by Google search over time
Picture from writer, July 2026

The query stays whether or not this decline will proceed, however the most recent removal – FAQ/FAQPage – has since brought about some debate over the position of schema.org inside the way forward for Search.

Schema Is Useless, Proper?

Whereas some carry out tests and experiments to grasp whether or not schema actually makes a optimistic influence on being cited inside platform responses, Gianluca Fiorelli notably observed that we could also be performing these checks on restricted datasets. With that in thoughts, let’s remind ourselves of the wording of the deprecation message for FAQ wealthy outcomes:

“…We will probably be dropping the FAQ search look, wealthy end result report, and assist within the Wealthy outcomes check in June 2026.”

Discover right here what they didn’t point out – which is that the usage of FAQ schema is now not required. It is because the deprecation is that of wealthy outcomes solely – a show characteristic. Schema itself is a comprehension layer – figuring out entities and the relationships between them. Is schema useless? In my view, it’s removed from it. Whereas some properties deprecate, others, reminiscent of Product, are extended.

That being stated, I’m additionally conscious that including schema isn’t a magic bullet that contributes in direction of quotation development. Nonetheless, that development goes past the metrics we’ve been accustomed to depend on reminiscent of citations, impressions, and so on. Suganthan Mohanadasan wrote a great piece about how schema has three “lives”:

  1. Google’s index pipeline.
  2. LLM pretraining (oblique).
  3. LLM runtime retrieval.

SEOs have been traditionally centered on No. 1 as one thing that may positively contribute in direction of success metrics. However schema goes past what we’re used to, or precisely, report on. Schema isn’t dying; one show characteristic it benefited from is diminishing as a substitute.

An search engine optimization’s Largest Menace: Ambiguity

Schema is an ontology that, as an internet commonplace, can contribute in direction of information integrity. The danger to information integrity is ambiguity. Ambiguity results in hallucinations. Hallucinations snowball. Ultimately the end result compounds, which might result in inaccurate outcomes and even incorrect LLM pre-training which might have longer-lasting results.

If an agent can misinterpret you, in some unspecified time in the future it can. LLMs can then threat touring into “semantic drift” detracted from the details and in favor of narrative. This was explored inside a chunk titled “Sangue e Grafi: Teaching a Small Model to Read the Bloodline” by Andrea Volpini and Chiara Carrozza the place frontier fashions tended to fall for narrative over details, whereas a small mannequin given knowledge-graph instruments drew degree with them.

Sangue e Grafi, by WordLift
Picture from writer, July 2026

→ Additional studying: Information Retrieval Part 1: Disambiguation

5 Layers Of Information Integrity

All this corroborates my perception that an search engine optimization’s position is to maximise information integrity, of which schema performs a task. Under, I illustrate 5 layers of what information integrity can embody:

5 layers of data integrity
Picture from writer, July 2026
  1. Entities: What exists, and what that factor is. Factor, Group, and Particular person, stabilized with @ids and tied out to Wikidata, GS1, ISNI, or ORCID so an agent is aware of your “Apple” from the fruit.
  2. Relationships: How these issues join. @id and sameAs, RDF. Yoast SEO’s schema aggregation feature and EntityMap by Dixon Jones.
  3. Format: How construction is serialized and served. JSON-LD, RDFa, and Microdata. Markdown, too (encompassing LLMs.txt, brokers.md, OKF) and endpoints (content material negotiation, ARD, MCP).
  4. Actions: What may be completed, declared to brokers. Schema.org Actions reminiscent of BuyAction, plus the newer WebMCP, ACP, and UCP.
  5. Notion: Grounding, third-party notion, sentiment, and so on.

Aggregation, Steering, And Consumption

In a post I wrote in October final yr, I stated, “SEOs should think about each side of the online and the best way to serve each.” The rising protocols (all of which have been launched previously two years) present this to be true, the place a brand new “agentic grounding stack” typically adopts one among three targets:

Protocol Aim What It Does 
sitemap.xml AggregationEach canonical URL right into a single XML index.
llms.txt AggregationAbstract of a website’s content material with vital data and hyperlinks to additional studying.
Yoast Schema Aggregation AggregationWeb page-level JSON-LD into one linked site-wide graph.
EntityMap AggregationA website’s entity declarations into one specific map.
Data Catalog AggregationStructured, unstructured, and SaaS information right into a ruled context engine.
OKF AggregationWeb site information right into a markdown bundle at /okf/.
ARD · ai-catalog.json AggregationA site’s instruments and brokers right into a catalog; registries federate above it.
OpenKB AggregationSupply paperwork compiled right into a markdown wiki.
Schema.org SteeringThe shared vocabulary that tells machines what issues imply.
brokers.md SteeringHow brokers ought to symbolize and work together with you.
Markdown for Brokers ConsumptionSimilar URL served as clear markdown by way of content material negotiation.
Markdown alternate output ConsumptionA separate .md model linked with rel=alternate.
/crawl endpoint ConsumptionRenders a web page, or whole website, as clear markdown on demand.
WebMCP ConsumptionExposes a website’s actions as instruments an agent can invoke.
NLWeb ConsumptionIngests schema, feeds , and sitemaps to reply natural-language queries.
ACP ConsumptionAgent checkout inside ChatGPT in opposition to service provider product information.
UCP ConsumptionA typical language for agent commerce actions throughout surfaces.

These three targets assist scale back the variety of requests whereas growing token effectivity. A few of the above protocols have been lined in additional element inside Search Engine Journal, together with my very own on ACP and UCP and Emina Demiri-Watson’s thorough article on OKF, ARD, and others earlier this month.

However there’s one thing none of those protocols have…

There Is No Consensus Or Agreed Customary

Schema.org was born out of consensus between Google, Microsoft/Bing, and Yahoo! (Yandex becoming a member of later) who launched it beneath joint governance. The identical occurred 5 years earlier with the XML sitemap. When the major search engines wanted a normal, they merely sat down and created one – collectively.

Nothing like that is occurring now, and it comes at the detriment of SEOs who genuinely need readability on what to implement and what to not implement for websites they work on. Even fundamental details about consumption are contested, the place the debate over markdown is a great example of this.

Whereas these debates proceed, there’s no room the place platforms are convening and agreeing to 1 common commonplace. The ecosystem has modified so dramatically that these firms will not be within the enterprise of Search and the nice of the online, however should now navigate how their companies have an effect on jobs, economies, livelihoods, and the way forward for humanity as a complete. As such, I simply don’t consider questions posed by SEOs are on the prime of their priorities.

What Can You Do About It Now?

Wanting again on the 5 layers of knowledge integrity, the 4 you may have management over may be illustrated beneath when what an agentic grounding stack can appear like:

Agentic grounding stack options
Picture from writer, July 2026

There’s rather a lot to think about, and all have completely different targets and technical debt. Resolve that are most relevant to you, in addition to adopting something that ought to not require an excessive amount of technical debt.

If I needed to decide an order, it will be this.

  • Stabilize your @ids and add sameAs links out to Wikidata and the opposite authorities first, as a result of every thing else stands on it.
  • Then check how you’re truly interpreted, with NLWeb, quite than assuming the graph reads the best way you supposed.
  • If you’re in ecommerce, audit the product feed earlier than touching something shiny, and take a look at BuyAction whereas you’re there: solely ReadAction and SearchAction are deployed at any actual scale right now, so the sphere is genuinely open. Look into the current information about what has been added to the Product schema.
  • Attempt to implement WebMCP. It may be completed on any web site, and doesn’t must be ecommerce.
  • Markdown serving and content material negotiation can wait till you may have engineering capability to spare (until you should utilize Cloudflare’s Markdown for Agents).
  • Look into OKF and ARD. When Google launches new protocols, I all the time take discover – particularly in terms of how an agent or LLM understands a website as a complete.

None of this can be a guess on a particular protocol. Implementing any of those reduces the chance that an LLM has to go “the great distance spherical” to type the reply it desires to reply with. By hedging bets to welcome any agent from any platform, this can even make it easier to assume extra about precisely how your website could also be interpreted by them, and the way this improves when these protocols are adopted.

Even if you wish to make small experiments away from bigger websites, it’s price exploring not solely to see if there are optimistic outcomes from it, but additionally to grasp how all of them work in apply. That is precisely what I’ve completed not too long ago, the place I’ve rebuilt my private web site, which has a number of “agentic prepared” protocols working, together with content material negotiation, markdown alternate, llms.txt and WebMCP.

Don’t Chase The Protocol, Personal The Layers Beneath

Proper now, it appears there isn’t any single “winner” that can progress from a proposal or protocol to an internet commonplace. There’s no consortium to repeat what was completed with the XML sitemap and Schema.

The stack is now huge, however there’s one factor all of them share. Whether or not it aggregates, guides, or consumes, each is fed by the identical substrate: Correct entities, specific relationships, and content material a machine can learn with out guessing or researching additional.

Rankings have been the success metrics of the outdated net. Belief, integrity, accuracy, and validity. Incomes that is nonetheless the position of an search engine optimization.

Extra Assets:


Featured Picture: hmorena/Shutterstock


#Information #Integrity #Technical #search engine optimization #Crawling #Belief

Leave a Reply

Your email address will not be published. Required fields are marked *