Episode 6: How to Help Google/Amazon Make Sense of a Chaotic, Unstructured Web
|

#SEOisAEO: Episode 6 - How to Help Google/Amazon Make Sense of a Chaotic, Unstructured Web - Jason Barnard with Arnout Hellemans, Martha Van Berkel & Aaron Bradley

Published by: Semrush. Host: Jason Barnard. Guests: Arnout Hellemans, Martha Van Berkel, Aaron Bradley. October 09, 2018


Google’s mission has always been to organise the world’s information. Amazon’s mission is to be the most customer-centric company on earth. In 2018, both were moving toward increasingly sophisticated answer-driven experiences - but they were approaching the problem from fundamentally different directions.

Google answers questions. Amazon completes actions.

Episode 6 of the #SEOisAEO series examined what that difference means for brands trying to be found, understood, and recommended across both ecosystems.

Arnout Hellemans, Martha Van Berkel, and Aaron Bradley joined Jason Barnard to tackle the coordination problem sitting underneath both platforms: the web is a mess.

It is contradictory, inconsistent, and largely unstructured. The same brand might be described differently on its own website, its Wikipedia entry, its Google Business Profile, its Wikidata record, and dozens of other sources. Each inconsistency creates another problem for a machine trying to determine what is true.

The practical framing developed during the episode was not about gaming the machine, but about helping it.

A brand that explicitly defines its entity - who it is, what it does, where it operates, who founded it, and what it is related to - in structured, machine-readable language is making the machine’s job easier. Instead of forcing an algorithm to extract meaning from chaotic prose and reconcile contradictory information, the business can provide clearer signals from the outset.

The original Semrush webinar listing for Episode 6 of the #SEOisAEO series, featuring Jason Barnard, Arnout Hellemans, Martha Van Berkel, and Aaron Bradley.
The original Semrush webinar listing for Episode 6 of the #SEOisAEO series, featuring Jason Barnard, Arnout Hellemans, Martha Van Berkel, and Aaron Bradley.

The Web Is a Chaotic, Unstructured Information Environment

The fundamental problem is simple: machines have to make sense of information that humans created without machines in mind.

People can move between a company’s website, a Wikipedia page, a marketplace listing, a social profile, and a third-party directory and intuitively recognise that the information refers to the same business.

An algorithm has to work that out.

It needs to determine whether two names refer to the same entity.

It needs to reconcile conflicting descriptions.

It needs to establish which facts belong to which entity.

And it needs to decide which sources deserve greater confidence.

This is the challenge that makes structured information so valuable.

Help the Machine Instead of Making It Guess

The episode’s central practical principle is straightforward:

The easier you make your information for the machine to understand, the less the machine has to guess.

That means clearly defining the important entities associated with a business.

Who is the organisation?

What does it do?

Where does it operate?

Who founded it?

Which products or services does it offer?

Which people, organisations, places, and concepts is it related to?

These are not merely branding questions.

They are questions an answer engine needs to resolve before it can confidently use a business as a source of information or recommend it as an answer.

The presentation frames this challenge through Relevancy, Understanding, and Credibility. The machine first needs to identify potentially relevant information, then understand that information, and finally determine which sources it can trust.

Schema.org: A Shared Vocabulary for the Machine

Martha Van Berkel brought Schema.org into the discussion as a way of providing a common vocabulary for describing information.

Schema helps businesses communicate information in a form that machines can interpret consistently.

Instead of leaving the algorithm to infer whether a piece of content describes a person, organisation, product, event, or another entity, structured data allows that meaning to be stated more explicitly.

The important point is that Schema is not useful simply because it is a technical SEO requirement.

Its deeper value is that it creates shared language between businesses and machines.

That makes structured data relevant not only to Google, but potentially to the wider ecosystem of platforms and services that need to interpret information about businesses and entities.

From Structured Data to Linked Data

Aaron Bradley expanded the discussion beyond Schema itself into the broader ecosystem of structured and Linked Data.

The important idea is connection.

An entity does not exist in isolation.

A person belongs to organisations.

An organisation operates in places.

A product belongs to a company.

A business offers services.

A publication may be written by a person and published by an organisation.

These relationships form a network.

Linked Data provides a way of thinking about how those relationships can be connected across the web rather than leaving each piece of information isolated inside one webpage.

That is where the idea of a Knowledge Graph becomes relevant.

Knowledge Graphs Need Connected Information

A Knowledge Graph is more useful when information is not simply a collection of disconnected facts.

The relationships between the facts matter.

This is why consistent entity definitions are so important.

If a company identifies itself one way on one platform and another way somewhere else, the machine has to resolve the discrepancy.

If the same entity is consistently identified and described across sources, the machine has a better opportunity to connect those references.

The objective is therefore not just to create more data.

It is to create coherent data.

The Role of HTML and Page Structure

Arnout Hellemans brought the conversation back to the practical reality of websites.

Structured data does not exist in isolation from the page itself.

The HTML structure, page architecture, content blocks, tables, lists, and other elements all affect how information can be extracted and interpreted.

This connects directly with the concerns raised in Episode 5.

A machine may encounter information through explicit Schema, but it may also need to interpret the underlying HTML and natural-language content.

The cleaner and more meaningful that structure is, the easier the extraction problem becomes.

Tables and Lists Are More Complicated Than They Look

The presentation spends particular attention on tables and lists.

For humans, these formats seem straightforward.

A table has columns and rows.

A list has items.

But a machine needs to understand what the table represents, what each column means, which values belong together, and what relationships exist between the items.

The same applies to lists.

Is the list merely a collection of items?

Is it ordered?

Does the order have meaning?

Does one item belong to another?

What does the heading above the list represent?

These apparently simple structures can contain a surprising amount of implicit meaning.

That makes clear HTML and explicit structure valuable to machines trying to extract information from the page. (slideshare.net)

The DOM and Machine Extraction

The presentation also looks at the DOM - the Document Object Model that represents the structure of a webpage.

For machines, the DOM can provide a layer of semi-structured information even when a website does not use comprehensive Schema Markup.

This reinforces an important point:

Machine-readable information is not limited to JSON-LD.

The machine can learn from the structure of the page itself.

That makes good HTML, clear page architecture, meaningful headings, and well-structured content part of the wider machine-understanding strategy.

Free Text Is Still the Biggest Challenge

Structured data can explicitly identify information.

But the vast majority of the web is still written in natural language.

Machines therefore also need to understand free text.

This is where the presentation moves into Natural Language Processing, relatedness and co-occurrence, topical silos, and semantic triples. (slideshare.net)

Humans understand relationships naturally.

If a sentence says that a company founded a product division in a particular city, a human reader can normally understand who founded what, when, and where.

The machine has to extract those relationships from the words.

That is why semantics matters.

The challenge is not simply recognising vocabulary.

It is understanding meaning and relationships in context.

Content Creation for Machine Understanding

This naturally leads to one of the most important questions for marketers:

How should we create content when machines are interpreting the free text we write?

The answer is not to write for machines instead of humans.

It is to write clear content whose meaning is easier for machines to interpret.

That means being explicit about entities and relationships.

It means avoiding unnecessary ambiguity.

It means structuring information logically.

And it means ensuring that important claims are supported consistently rather than buried inside vague or contradictory language.

The objective is to make the content easier for both audiences to understand.

Google and Amazon: Different Jobs, Different Needs

The episode also examines the differences between Google and Amazon.

Google’s core role is information discovery and answering questions.

Amazon’s environment is much more closely connected to products, commerce, and actions.

That distinction matters because the machine’s information requirements depend partly on what it needs to help the user accomplish.

A search for information and a request to purchase a product are not identical problems.

One may require an answer.

The other may require an action.

For brands, this means the same entity information may need to work across different ecosystems and different user journeys.

One Entity, Many Platforms

This is where the coordination problem becomes particularly important.

A business may have information spread across:

  • Its own website
  • Google
  • Amazon
  • Wikipedia
  • Wikidata
  • Local business platforms
  • Social networks
  • Industry directories
  • Third-party publications

The machine has to determine whether all those references describe the same entity and whether the information is consistent.

That means brands should not think of each platform as an isolated optimisation exercise.

They should think about the entity behind those platforms.

The objective is to create a consistent digital representation that machines can reconcile across sources.

Schema Strategy: Short-Term Tactic or Long-Term Infrastructure?

The presentation also asks an important strategic question: How should businesses approach Schema in the short and long term? (slideshare.net)

The short-term temptation is to ask:

“Which Schema types can I add to this page?”

The longer-term question is more powerful:

“What does the machine need to understand about this business?”

That shift changes Schema from a checklist into an information strategy.

Instead of marking up isolated pages, businesses can think about their broader ecosystem of entities, properties, and relationships.

That perspective leads naturally toward Knowledge Graph thinking.

The Position 0 Profile

The presentation brings the discussion together through a Position 0 profile.

Among the elements it highlights are:

  • Knowledge Graph first
  • The best of the Top 10
  • An answer algorithm on top of the search algorithm
  • The Ultimate Source of Truth
  • User signals
  • Schema as a foundation, but not enough on its own
  • Semantic HTML5
  • Clear content blocks (slideshare.net)

Together, these ideas demonstrate that AEO is not a single optimisation tactic.

It is an information-understanding problem.

The machine needs structured signals.

It needs well-organised content.

It needs meaningful relationships.

It needs consistent information.

And it needs evidence that users find the resulting answers useful.

The Ultimate Source of Truth

One of the strongest ideas in the presentation is the concept of becoming an Ultimate Source of Truth.

This does not simply mean publishing the most content.

It means creating information that is sufficiently clear, consistent, and authoritative for the machine to rely on it when constructing an answer.

That requires brands to take responsibility for their own information.

If the machine has to discover the truth about the brand by piecing together contradictory descriptions from dozens of sources, there is more uncertainty.

If the brand clearly communicates who it is and how its information fits together, there is less ambiguity.

That is the practical advantage of helping the machine rather than leaving it to guess.

Make the Algorithm’s Job Easier

This is ultimately the through-line of Episode 6.

Do the machine’s work for it where you can.

Define the entities.

Structure the information.

Use a shared vocabulary.

Connect related entities.

Make the HTML meaningful.

Write clear natural-language content.

Keep important information consistent across platforms.

Provide structured evidence rather than relying entirely on the machine to infer meaning from unstructured prose.

The goal is not to manipulate the algorithm.

It is to reduce uncertainty.

A Brand That Makes Itself Easy to Understand Gets Understood

Episode 6 takes the #SEOisAEO discussion deeper into the technical infrastructure behind machine understanding.

Google, Amazon, and other platforms may approach search and actions differently, but they all face a similar underlying problem:

the web contains an enormous amount of information that was not created in a consistent, machine-readable way.

Businesses can either add to that uncertainty or reduce it.

Schema.org, Linked Data, Knowledge Graphs, semantic HTML, clear page structures, NLP-friendly content, and consistent entity information all provide ways to reduce the ambiguity.

The goal is not to make the algorithm work harder.

It is to make the information easier to understand.

And that leaves Episode 6 with a simple principle:

A brand that makes itself easy to understand gets understood. A brand that leaves the interpretation to the algorithm gets interpreted - and the algorithm will use whatever it finds.


Published by: Semrush
Series: #SEOisAEO
Episode: 6
Host: Jason Barnard
Guests: Arnout Hellemans, Martha Van Berkel, Aaron Bradley.
Original webinar date: October 9, 2018
Topic: Structured Data, Schema.org, Knowledge Graphs, Linked Data, Semantic HTML, Natural Language Processing, Google, Amazon, and Answer Engine Optimization

Presentation deck: How to Help Google/Amazon Make Sense of a Chaotic, Unstructured Web

Similar Posts