|

The substrate decides now. Not the artifact. Week 33 Newsletter Summary

30-Second Take

Week 33 produced two hard findings that look unrelated.

Inside the enterprise: 95% of organisations stalled or cancelled an AI project in the past year. The named blocker isn’t model quality. It’s that nobody can state, in machine-readable form, what is actually true about the business. The organisations that have governed their context catch roughly twice as many confidently-wrong agent answers as the ones that haven’t.

Outside the enterprise: a study of 60 ChatGPT conversations found brands sitting in the model’s prior (what it already believed before searching) got cited 33 times more often than brands it discovered through retrieval. Retrieved pages earned citations 3.1% of the time. In 21 of 27 conversations, ChatGPT named brands before fetching any search results at all. The search was confirmation, not discovery.

Those aren’t two stories. They’re one. In both cases, the entity doesn’t hold an authoritative machine-readable account of what is true about itself. In both cases, a third party holds it instead, and the entity is discovering this when it checks why its AI projects keep failing or why a competitor keeps appearing in answers where it should.

The same week: every governance product launched sits between the agent and the resource rather than underneath it, measurement consolidated around Nielsen absorbing DoubleVerify for $2.15 billion, and every major model got faster and cheaper without moving much on capability. The pattern underneath all of it: value moved from the artifact to the substrate that decides whether the artifact is trusted.

Key Developments

AEO: the model decides which brands to name before it searches

Executive read: your owned surface is the least persuasive evidence about you

The study that rewrites the AEO category (caveat: 60 conversations, one analyst, not peer reviewed, treat as directional): ChatGPT named brands before fetching any search results in 21 of 27 conversations. Brands present in its initial search queries were cited 68.9% of the time. Brands discovered only through retrieval were cited 2.1% of the time. That’s a 33x gap. Retrieved pages overall earned citations 3.1% of the time, meaning 97 in every 100 pages actually read were discarded without citation.

The corroboration landed the same week. AirOps analysed 3.5 billion citations and found creator and social citation share grew 140% from August 2025 to June 2026, while brand.com citation share fell 10%. YouTube drove most of that growth, up 158%. Reddit, Wikipedia and LinkedIn account for 99% of UGC citations in SaaS ChatGPT answers. More than half of US shoppers verify AI recommendations on Reddit before buying, trusting that community over professional critics and influencers as the deciding check. Google is testing an AI-first homepage for signed-out users, removing the classic Search button.

The consolidated read: your owned surface is becoming the least persuasive evidence about you. The entity record that decides your citations is assembled elsewhere, by third parties, about you, without you. The model’s prior behaves like an attitude in the psychological sense, an evaluative disposition formed largely from other people’s testimony rather than from the subject’s own self-description. You don’t change an attitude by talking louder about yourself. You change it by changing who else vouches, and in which forums.

Practitioner read: instrument the prior, not the crawl

If the model names brands before it searches, crawl-based visibility tooling is instrumenting the wrong moment. The early observable: capture what queries the model writes when asked a category question. Those query compositions are cheap to observe and show you what the model already believes before it touches your content.

The citation data repoints the budget. If 99% of UGC citations in your category come from Reddit, Wikipedia and LinkedIn, and those platforms respond to named humans rather than brand accounts, the relevant work is getting credible people to say true things in those three places. That’s closer to analyst relations and comms than to SEO, which means the budget may not currently sit where the agency is looking.

One counter-signal worth noting: ShieldFont, a font-based anti-scraping technique that swaps 24.5% of a page’s words into subtly wrong substitutes visible only in raw HTML. Tested against six scraper pipelines, more than 90% of ShieldFont pages were rejected outright. Deliberate corpus poisoning as a defensive posture raises a governance question for any brand considering it.

Enterprise AI: the governance blocker finally has numbers attached

Executive read: 95% stalled, and governance is the named cause

Two numbers that belong together. Cloudera research (11 August, vendor-funded, check methodology before citing): 95% of surveyed enterprises delayed or cancelled AI projects over the past year, governance and compliance among the biggest blockers. Nearly three-quarters said AI made data governance more complex, not less. This is the first time the “AI makes your data problem worse before it makes it better” claim has arrived with a survey rather than as consultant folklore.

The sharper finding came separately: a survey of 101 enterprises found organisations running governed semantic layers substantially better at catching confidently-wrong agent answers, roughly twice as many caught. Treat the exact multiple as directional until the funder is confirmed, but the direction is the point. This is an outcome number attached to the least glamorous layer in the stack.

Gartner added the cost context (10 August): AI-optimised infrastructure spending projected to grow 96% in 2026 to $42 billion, with inference overtaking training as the largest demand source. Two-thirds of surveyed enterprises run AI workloads in production with no visibility into infrastructure cost or utilisation. Spend is doubling, cost is moving from a project line to an operating line, and most organisations can’t see what it costs.

Practitioner read: the boring layer now has a business case

The semantic layer finding matters commercially because it’s the first time the boring layer has an outcome attached rather than a virtue. “Governance is important” is unsellable to a finance audience. “The organisations that governed their context catch twice as many wrong answers” is a business case.

Supporting data worth filing:

  • Capital One built its multi-agent platform around open-weight models, citing control over deployment and infrastructure.
  • Agentic memory is named as the next infrastructure problem after longer context windows. Agents running across hours or days need persistent memory that decides what to retain, retrieve and discard, a context-architecture problem wearing an infrastructure hat.
  • Enterprise 30TB SSD prices are up roughly 6.5x year over year. Storage is becoming a first-order AI cost.
  • IBM is launching a dedicated OpenAI consulting practice and plans to certify tens of thousands of consultants. Read that as supply-side pricing pressure on generalist AI advisory, arriving fast.

Governance products sit above the problem, not underneath it

Executive read: a gateway can audit access, but it cannot audit truth

Look at everything shipped this week: A10 AI Gateway, AWS AgentCore Observability, JumpCloud, Brex, Citrix Platform Flex, Tines. Every one sits between the agent and the resource.

That position is commercially excellent and architecturally shallow. Excellent because it’s installable without anyone having to agree on anything, produces a dashboard in a week, and can be sold jointly to a CISO and a CFO (this is now the third consecutive week that joint pitch has held). Shallow because a gateway can tell you an agent read the revenue table nine hundred times. It cannot tell you the revenue table was wrong.

The thing none of this week’s products build is the assertion layer: a permissioned, corroborated, machine-readable layer that adjudicates between competing definitions and carries provenance per claim. That can’t be shipped without organisational agreement. Someone has to decide what a “customer” is, and someone with authority has to say that finance’s definition wins over marketing’s and then live with the consequences. No vendor can do that work. No procurement process is shaped to buy it. So the market routes around it, and the 95% keeps stalling.

Practitioner read: MCP sprawl is invisible to shadow IT tooling

Shadow IT detection watches network traffic. MCP servers frequently run locally on a machine over stdio (two processes talking through standard input and output on the same computer), so no packets cross a monitored network boundary. An employee can wire an AI agent into a production database via a local MCP server and the security tooling registers zero events. The tools that caught shadow IT can’t see this.

Also worth noting: MCP’s new session model hands session handles to the model, meaning applications now own correlation, authorisation and session management. More scalable, more application responsibility.

The judgment gap: output is outrunning our ability to verify it

Executive read: the failure mode is plausible work, not broken work

The most commercially useful cluster of the week for anyone operating inside or advising an organisation that has already bought the tools.

AI is expanding our ability to produce work faster than our ability to judge it. The failure mode isn’t an obviously broken draft, it’s a plausible one, arriving organised and confident enough that reviewing it feels like editing rather than investigating. One study found consultant accuracy on a harder business problem fell by up to 24 percentage points among consultants using AI, the same tool that produced large gains on easier tasks. Same tool, harder problem, worse outcome. Verification is now the constraint on adoption, not capability.

Named hard problems after “hiring the agent”: tacit standards, company-specific evals, feedback ownership, permissions, liability, self-improving learning loops. Anthropic’s own research named multi-agent failure at scale as an emerging concern, individually benign agent behaviours compounding into systemic failure faster than institutions can oversee them.

People moves that signal where organisations think this lands:

  • Target hired its first Chief AI Officer (Chandhu Nair, ex-Lowe’s).
  • Ford named a chief data, AI and analytics officer (Mano Mannoochahr, ex-Verizon).
  • OpenAI lost both its COO and CRO within days, ahead of an anticipated IPO.

Practitioner read: an AI writing policy is becoming a buyer question

Clay’s AI writing policy is the first well-circulated attempt at a written organisational standard for AI-assisted work. The core principle: employees must fully own AI-assisted writing. Generating long documents from short prompts is discouraged. Writing is treated as a way to develop understanding, not bypass it. The operating principle is ownership, not disclosure. Expect “do you have an AI writing policy?” to become a standard buyer question within two quarters.

Spec-driven development applies directly: as code generation gets cheap, the bottleneck shifts to intent, coordination and review. The living, versioned specification becomes the source of truth. Same logic applies to any AI-accelerated workflow.

The Big Question: what is missing on both sides of the firewall?

Two of this week’s hardest findings are the same failure viewed from opposite sides of the firewall: 95% of enterprises have stalled an AI project on governance, and brands are cited 33 times more often from what the model already believes than from anything it retrieves. So, what is the one thing missing in both cases, and why is every vendor this week selling something that sits above it rather than building it?

The answer: an assertion layer. Not a data warehouse (holds values, not authority). Not a data dictionary (catalogues disagreement, carries a low price in the buyer’s mind, routes to the wrong room). A permissioned, corroborated, machine-readable layer that adjudicates competing claims, carries provenance per assertion (who stated it, on whose authority, corroborated by what, expires when), and is enforced at query time rather than consulted by the human who remembers it exists.

The assertion layer can’t be installed without organisational agreement. It’s not primarily a software purchase, it’s a political act. Someone has to end a disagreement that’s been live for years and put their name on the outcome. No vendor can do that, and the market routes around it accordingly.

What makes this week different from previous “governance is the blocker” weeks: the semantic layer finding attaches an outcome number. “Governed context measurably reduces confidently-wrong output” is the first version of that argument that survives contact with a CFO. The exact multiple needs verification, but the shape of the claim is now fundable in a way it wasn’t before.

Monday Move

For executives: run the diff in both directions

Ask two models the ten questions your buyers actually ask. Record what they assert about you before they search. Then take your five most-used business metrics and ask which system is authoritative for each, and who decided. Most organisations can’t answer either question. Both are cheap to run and immediately alarming.

For practitioners: stop measuring retrieval, start measuring the prior

If the model names brands before it fetches, crawl-based visibility tooling is instrumenting the wrong moment. Measure what it says unprompted, per model, and track the divergence between models over time. Divergence is where the entity record is thinnest and where intervention has the most leverage. Per-model divergence, not a single-model score, is the measurement that shows where you’re weak.

Quotable Take

Your owned surface is becoming the least persuasive evidence about you. The entity record that decides your citations is assembled elsewhere, by third parties, about you, without you. You don’t change that by publishing more or structuring your content better. You change it by changing who else vouches and in which forums. That’s an identity and corroboration problem, not a content problem.

Frequently Asked Questions

  • Why does ChatGPT cite some brands 33 times more often than others?

    Because the model usually decides which brands to name before it searches. In a study of 60 ChatGPT conversations, brands present in the model’s initial search queries were cited 68.9% of the time, while brands discovered only through retrieval were cited 2.1% of the time. In 21 of 27 conversations the model named brands before fetching any search results at all, which makes the search step confirmation rather than discovery.

  • What is an assertion layer?

    An assertion layer is a permissioned, corroborated, machine-readable layer that adjudicates between competing definitions inside an organisation and carries provenance for each claim: who stated it, on whose authority, corroborated by what, and when it expires. It is enforced at query time rather than consulted by whoever remembers it exists. It is not a data warehouse, which holds values but not authority, and not a data dictionary, which catalogues disagreement without resolving it.

  • Why are 95% of enterprise AI projects stalling or being cancelled?

    Governance and compliance, not model quality. Cloudera research published on 11 August found 95% of surveyed enterprises delayed or cancelled AI projects over the past year, with governance among the biggest blockers, and nearly three-quarters saying AI made data governance more complex rather than simpler. The underlying problem is that no one can state, in machine-readable form, what is actually true about the business.

  • Where do AI answer engines get their citations?

    Increasingly from third-party and user-generated sources rather than owned brand sites. AirOps analysed 3.5 billion citations and found creator and social citation share grew 140% between August 2025 and June 2026 while brand.com citation share fell 10%. YouTube grew 158%. Reddit, Wikipedia and LinkedIn account for 99% of UGC citations in SaaS ChatGPT answers.

  • How should a brand measure AI search visibility?

    Measure the model’s prior rather than crawl coverage. Ask each model the questions your buyers ask, record what it asserts about you before it searches, and track divergence between models over time. Per-model divergence, not a single blended visibility score, shows where the entity record about you is thinnest and where intervention has the most leverage.

Similar Posts

Leave a Reply