Discovery in the DSCRI Pipeline: The Bot Will Never Find You If You Wait to Be Found
By Jason Barnard
The pipeline does not start when the bot arrives at your page. It starts when the system decides whether you are worth visiting at all.
Most brands treat Discovery as a passive condition: the site is live, the pages exist, and the assumption is that the bots will find them eventually. That assumption is wrong, and it is expensive. Discovery is an active signal, not a passive hope, and the difference between a brand the system finds quickly and a brand the system ignores lies less in the quality of the content than in whether the brand gave the system a reason to look.
Fabrice Canel, Principal Program Manager at Microsoft responsible for Bing’s crawling infrastructure, put it directly: “You want to be in control of your SEO. You want to be in control of a crawler, and IndexNow, with sitemaps, enable this control.” Control, not hope. The push layer changes the economics of Discovery entirely, and every brand still waiting to be found is ceding that control to chance.
The system asks two questions at Discovery, and most brands only answer one
The first question is whether the URL exists. That is the question most brands think about: is the page published, is it in the sitemap, does it resolve correctly. Most brands answer yes.
The second question is whether the URL belongs to an entity the system already trusts. That is the question most brands ignore. Content without entity association arrives as an orphan: the URL exists, but the system has no prior relationship with the entity that published it, no confidence to anchor against, and orphans wait at the back of the queue regardless of how well-written the content is.
Entity Home is the term I use for the canonical web property you control, the primary discovery anchor that establishes who you are before the system reads anything you publish. Get the Entity Home right and every page published from it inherits a trust context. Get it wrong or leave it ambiguous, and every page publishes into a void.
Three mechanisms feed Discovery, and together they are the full picture
XML sitemaps are the census: a structured declaration of what exists, updated when new content is published, giving the system a complete map of the territory. They are the baseline, and every brand should have them configured correctly.
Internal linking is the road network: the system discovers pages by following links, and the architecture of those links determines which pages get found, how quickly, and with what context attached to each journey. A page with strong internal links from high-authority pages within the site arrives at Selection with a different signal than a page that exists in the sitemap and nowhere else.
IndexNow is the telegraph: a direct push notification to search engines the moment content is published or updated. Instead of waiting for the next crawl cycle, the brand signals immediately that something new exists and is ready to be evaluated. The economics change completely: Discovery moves from a function of crawl scheduling to a function of publishing cadence.
For me, the Common Crawl test is the most honest diagnostic available
Common Crawl is an independent, open-access web crawl covering approximately three billion pages. It is not a search engine and not a commercial product: it is an autonomous bot that indexes what it finds, without editorial intervention, without relationships with publishers, without the cooperative signals that Google and Bing use to prioritise crawling.
Perplexity and OpenAI use Common Crawl data precisely because it is free and independently operated. A brand in the Common Crawl dataset has cleared an objective credibility threshold: an independent third-party bot found the site worth crawling without any of the push mechanisms that accelerate commercial crawlers. Being in Common Crawl does not guarantee strong Discovery performance with Google or Bing, but being absent from it is a signal worth taking seriously.
I use it as a diagnostic for clients: if Common Crawl has not found your pages, you have a Discovery problem that sitemaps and IndexNow will help but that ultimately points to an entity trust deficit. The bot found you unremarkable, and that verdict is worth more than any technical audit.
The Nested Audience Model starts here
This series walks five infrastructure gates, each with the same primary audience: the bot. Everything from Discovery through Indexing is optimisation for machine readability before a single algorithm makes a competitive decision. The content can be brilliant. If the bot cannot find it, fetch it, and file it, the algorithm never sees it.
I call this the Nested Audience Model: the audiences are nested in sequence, not running in parallel. Bot first, then algorithm, then person. The ARGDW competitive phase that follows DSCRI is entirely an algorithm audience. The eventual human recommendation is entirely downstream of both. Fix Discovery and you open the door to everything that follows, leave it broken and nothing upstream matters.
The push layer changes what game you are playing at this gate, and the gap between brands using IndexNow, structured feeds, and MCP connections and brands still waiting for crawl schedules is widening with every cycle. The bot will never find you if you wait to be found. Start pushing.
The Complete Ten-Gate AI Engine Pipeline
- Discovery in the DSCRI Pipeline: The Bot Will Never Find You If You Wait to Be Found
- Selection in the DSCRI Pipeline: The Bot Decided Your Page Wasn’t Worth Its Time
- Crawling in the DSCRI Pipeline: The Bot Arrived at Your Page and Brought a Briefing Document
- Rendering in the DSCRI Pipeline: The Bot Sees a Different Page Than Your Customers Do
- Indexing in the DSCRI Pipeline: Stored Is Not the Same as Understood
- Annotation in the ARGDW Pipeline: The Bots Stored Your Page but the Algorithms Don’t Understand It
- Recruitment in the ARGDW Pipeline: The Trick Is to Charm the Algorithmic Trinity
- Grounding in the ARGDW Pipeline: The Truth-Check That Decides Whether the AI Uses Your Brand or Your Competitor’s at the Moment of Display in Assistive Engines
- Display in the ARGDW Pipeline: Your AI Salesforce Is Recommending Your Competitor, Not You
- Won in the ARGDW Pipeline: 95% of Your Market Is Not Buying Right Now. Who Does the Assistive Engine Choose When They Are?
This is the first in a five-part series on the DSCRI infrastructure gates in Jason Barnard’s ten-gate AI Engine Pipeline (part of the 15-gate Kalicubeยฎ Framework). The next piece covers Selection: where the system uses entity confidence to decide how many of your pages are worth fetching, and why crawl budget was always the symptom of a deeper problem.