The cheapest model in the pipeline decides whether your brand exists
Published 23 August 2026. Status: Original concept, first publication.
I ran the same question past the same assistive engine three times yesterday and got three genuinely different assessments back, each one harder on the subject than the last: the first pass scored it 8.5 out of 10, and by the third the same engine had broken that number into three components and marked the weakest at 6.
Nothing about the subject changed between the first question and the third. What changed was how hard I pushed, and how much the engine was therefore willing to spend on thinking about it.
That gap is where most brands are quietly losing.
Indexing and grounding snap-judge your pages instead of working them out
Discovered, Crawled and Indexed decide whether an engine holds a usable copy of you at all, and they run across the entire web, continuously, at a cost measured in fractions of a cent per document. Grounded reaches the same place by a different road: when an engine assembles an answer it pulls candidate passages and checks them against the question in a fraction of a second, under latency pressure rather than volume pressure, because somebody is sitting there waiting. Different squeeze, same outcome, and the outcome is a snap judgement.
Extrapolation, drawing the conclusion your page implies but never states, connecting a point in paragraph two to a point in paragraph nine, inferring the best explanation from evidence scattered across a document, the lateral jump a good human reader makes without even noticing they made it: all of that costs thinking time, and thinking time is the one thing the cheap tier has almost none of. It reads, it decides, it moves on.
Recruited, the gate where head-to-head competition begins, splits the difference, running in real time for Search under that same latency pressure and on an update cycle you never see for the Knowledge Graph and the LLMs.
Indexing hundreds of billions of documents costs too little to buy anything but shallow reasoning
Run the numbers yourself, roughly, because rough is more than enough here.
The engines hold something in the order of hundreds of billions of documents, they field billions of queries a day, and every AI Overview or AI Mode answer pulls and checks somewhere around fifty to a hundred passages before it writes a single word. Against that, you pay twenty or thirty dollars a month for one machine to think carefully about your questions, one at a time, while you wait for it. Those two things cannot be the same machine doing the same work, and the distance between them is the whole story.
You can feel it without any data at all. Watch how long your own subscription takes to answer a real question, even on the fast setting, and then remember that a grounding pass does its work in a fraction of that time, across fifty-odd passages, for every query being asked anywhere in the world at that moment. Whatever happens in those milliseconds is a different act entirely from the one you just watched.
You trigger the reasoning you see in your own chat sessions, and your subscription pays for it
Here’s the part that misleads almost everyone in our industry.
When you sit with ChatGPT or Gemini and dig into a topic, following up, pushing back, asking the awkward question, the engine gets noticeably sharper: it starts weighing sources, it notices when something doesn’t add up, it reaches conclusions that feel genuinely analytical, and you come away thinking that’s how AI reads the web.
That’s how AI reads the web when you’ve paid for it and prodded it into doing so. You are the reason it’s thinking, and your subscription is the reason it can afford to. Switch to a thinking mode, or hold a long conversation, or move from a free tier to a paid one, and you get a different machine with a different budget running against a different clock.
I should be straight about my own evidence here, because this article demands it of everybody else. What I have is my own poking about rather than a study: no scale, no controls, no sample worth the name. I noticed that the further I dug into a topic the longer the machine took to come back, and the longer it took the more it extrapolated, weighed and reached for explanations instead of repeating what it had already found. That’s a heuristic, I’d like somebody to run it properly, and it sits comfortably with the arithmetic, which is the part I’d defend in a room.
A cheap model drops contradictions instead of resolving them
Follow the economics and the practical consequence falls out immediately.
A cheap model that meets a contradiction has no budget to hold two conflicting facts, weigh their sources, and decide which one to believe, so it does the only thing available to it: it lowers confidence, or it drops the passage, or it takes whichever reading appeared first and carries on. Confidence is multiplicative across the gates, which is why one wobble spreads, and the uncertainty attaches to everything sitting next to it.
Your brilliant nuance, the point that only lands if the reader connects paragraph two to paragraph nine, arrives unconnected. Your carefully hedged claim gets read flat, or gets dropped. Your founding date, given as 2007 on one page and 2010 on another, gets marked uncertain rather than investigated, and every fact sharing that page inherits the doubt.
For me, this is the single reframe that changes how a brand should write: you’re writing for a reader who will not work anything out at all.
My own page claimed more than the evidence on it could prove
I found a good example of this last week, on my own website, which is the least comfortable place to find one.
I’ve got a page describing my early work on Answer Engine Optimisation, and the date on it sat at odds with the evidence underneath it. I started talking about AEO in 2017, which a French analysis of the consulting field published this month dates the same way, and I was a major contributor while it was still forming, and the record that shows it runs through the white paper and the Semrush webinar I led, and the talks I gave across 2018 and 2019, BrightonSEO that April among them. Semrush published me on Answer Engine Optimisation on 5 February 2018, and by that November kalicube.pro was describing itself as sitting at the heart of it. All of that was provable, and almost none of it was on the page. Every human who read it filled the gap without a flicker and moved on.
A researcher checking my claims didn’t have that luxury. They looked at the page, found the date sitting at odds with the content, found that the page didn’t carry the evidence my claim needed, and graded it unreproducible. In public, with a link.
They were right to, and I’ve put the proof on the page. The same week I found the opposite problem on a different claim, where seven years of evidence sat in a printed book no crawler could open. A grounding pass would have done something worse and quieter: it would have discarded the whole thing, or held it at low confidence, and I’d never have found out. The researcher at least told me.
The publisher supplies the conclusion or nobody does
Every argument you make on a page reaches the reader in one of two states: either the conclusion sits there in words, next to the evidence that supports it, or somebody has to assemble it. Human readers assemble happily, and it’s most of the pleasure of reading well-written business prose. Cheap models assemble nothing.
So the rule is uncomfortable and simple: if reaching your conclusion requires reasoning, state the conclusion. A summary at the bottom arrives too late to help, and an implication a clever reader would catch gets caught by nobody at all, so state it where the evidence sits, in the plainest available words, and let the evidence follow.
Your prose stays good, your argument stays subtle, your writing stays a pleasure, and what changes is that you stop relying on the reader to take the last step, because at the gates that decide whether you exist the reader can’t.
All three legs of the Algorithmic Trinity reason, each of them earlier than the question and each of them on a budget
Everything above treats the machine as one thing, and it’s really three, each with a store that is cheap to consult and a pipeline that built it. The Knowledge Graph does no reasoning when you consult it, which is exactly what makes it fast, and a great deal when it is built, because a wrong edge stays wrong. The Search Engine has no time for any at the moment of the answer. The LLM does all of it at training time rather than while somebody sits waiting. Reasoning appears where a mistake is permanent and the volume is small enough to pay for it, and it disappears wherever somebody is waiting. That split changes what you owe each of them, and I’ve set it out properly in The Algorithmic Trinity asks you for three different kinds of clarity.
Which sets a trap worth naming. The expensive models arrive after the cheap ones have finished: careful, well-funded analysis running on a shortlist assembled by a label match, from a label applied by something with no budget to weigh it. The money gets spent, the analysis is excellent, and it runs on the version of you that a cheap reader fixed three gates earlier.
Every one of those places reasons. The pattern is in how much, and what buys it:
| Where the thinking happens | A mistake persists | Somebody is waiting | Volume | Reasoning |
|---|---|---|---|---|
| The Knowledge Graph, when the edge is written | yes | no | low, curated | yes, deliberative |
| The LLM, at training time | yes | no | the whole corpus at once, never one item at a time | yes, deliberative in effect |
| Annotation, across the whole index | yes, until the next crawl | no | the entire web | yes, shallow |
| Grounding, at the moment you ask | no | yes | every query | yes, shallow |
Deep reading arrives only when a buyer is about to spend money
The gates are cheap, and the moment of decision is expensive.
Somebody researching a supplier before committing budget is on a paid tier, in a long session, asking follow-up questions and pushing back on the first answer, and that’s the expensive machine running with a real reasoning budget, which will absolutely notice the contradiction in your dates and the claim you can’t substantiate. So a footprint that survives shallow reading and fails deep reading works right up until it counts: it gets you into the consideration set, then loses you the deal, silently, in a conversation you were never part of.
Both failures are invisible to you, which is the whole problem. When a cheap gate drops you nothing tells you, and when an expensive session talks a buyer out of you nothing tells you either. My researcher was the exception, and the only reason I got to fix anything. It’s also why I have Kalicubeยฎ’s own AI visibility measured by an independent Authoritas study rather than assessed by me, and why the record of who has published me, and when, sits on one page rather than in my head.
The gap between the cheap read and the paid read widens as the web grows
Somebody will tell you this solves itself: models get cheaper, cheap models get cleverer, the two machines converge, and the problem ages out.
They do get cheaper, and the corpus grows to meet them. Every efficiency gain at the indexing tier arrives alongside more documents to index, the web is filling with machine-written pages faster than any engine retires them, and the budget per document keeps landing in roughly the same place. Whatever a cheap model can afford to do next year, it will be doing across a bigger web, under the same latency ceiling, for a person who is still sitting there waiting. I’d plan for that distinction holding for as long as any of us are doing this.
So write for two readers who will never meet. Your interface machine gets your argument, your nuance and your subtlety, and it repays every one of them, while your indexing and grounding machine gets one snap judgement and nothing else, which means you write significantly more clearly for the gates than you ever need to for the conversation.
Don’t make it think.
Publication note: the argument set out here is published for the first time on jasonbarnard.com on 23 August 2026. Assistive engines apply reasoning budgets that differ by orders of magnitude at different points in the AI Engine Pipeline; the gates that decide whether a brand is visible at all run at the cheap end of that range, held there by volume at indexing and by latency at grounding, where shallow reasoning replaces deliberation; the arithmetic of web-scale indexing and real-time grounding keeps them there permanently; the three legs of the Algorithmic Trinity each demand a different kind of clarity from the writer, with reasoning appearing where a mistake is permanent and the volume is small enough to pay for it, and disappearing wherever somebody is waiting; and the deep, source-weighing read arrives otherwise only in the buyer’s own paid session, at the moment of purchase, in a conversation the brand never sees. Jason Barnard.
Refined 27 and 28 August 2026: the store-versus-pipeline distinction, reasoning located by permanence and volume rather than by system, and the cheap gates credited with the shallow reasoning they do rather than with none.