Executive summary
Four figures anchor this paper, and they tell a consistent story. Agents are already in production at scale, 60% of companies, per the research this paper draws on, are running them for real buyer-facing tasks, and Salesforce reports a fifth of its own sales volume as agent-originated already. Yet the same agents fail a core buying task, pricing, roughly one time in five, even on well-resourced B2B sites, and more than half the crawl traffic hitting the open web today isn't even trying to answer a task, it's training data collection. Enterprises are underwriting agent traffic they haven't measured, on infrastructure economics that shifted twice this year alone. This paper's four-dimension framework, task completion, machine-readability under load, fallback exposure, and access economics, is built to give that measurement a shape.
This gap between adoption and readiness is the paper's central concern. Three out of four companies already running agents in production, per the research cited above, are actively investing further, which means the volume of agent-mediated buyer interactions is set to grow from here, not plateau. Enterprises that wait for a mature, industry-standard measurement framework before acting are choosing to scale an unmeasured surface for another year or more. The four dimensions below are built to be run now, with tools and access most enterprise marketing and RevOps teams already have, rather than waiting on a vendor category that hasn't fully formed yet.
Why agent readiness is a different problem
Our existing GEO readiness framework measures accessibility, structure, authority, and coverage, the inputs that decide whether an engine cites you in a generated answer. That framework still holds, and every enterprise should still be running it. But it answers a narrower question than the one enterprises now actually face: can an agent, acting on a buyer's behalf and reporting back a specific answer, complete a defined task using only your site, or does it give up and hand the answer to a directory, a reseller, or a competitor.
Kevin Indig's July 2026 study is the sharpest data point available on this distinction. He ran 100 B2B products through three buyer-relevant tasks, pricing and features, integrations, and security and compliance, five times each, logging whether the agent answered from the vendor's own site or fell back to a third party. That's a research design closer to the technical audits we run than to a typical citation-tracking study: it measures whether the site functions for the task, not whether it theoretically could. The framework below builds directly on what that study, and the crawler-economics research that followed it this quarter, implies for how enterprises should be scoring their own sites.
It's worth naming the underlying shift plainly, because it changes what "the buyer's journey" even means as a phrase. For most of the last two decades, that journey assumed a human clicking through a sequence of pages, forming an impression along the way, brand tone, design quality, social proof, all doing real persuasive work. An agent completing a pricing lookup on a buyer's behalf skips almost all of that. It doesn't form an impression from your homepage hero section. It wants a number, an integration list, or a compliance answer, and it wants it in a form it can extract without executing a sales pitch. Sites built primarily to persuade a human through that journey are, by construction, not built for this second kind of visitor, and most enterprise sites were built for exactly one of those two visitors, not both.
Dimension 1: Task completion
Task completion asks the most direct question in this framework: given a specific buyer task, can an agent finish it, end to end, using only your domain. Indig's study is built around exactly this question, and the results split sharply by task type. Integrations and security and compliance questions got answered correctly from first-party sources 92 to 93% of the time. Pricing questions dropped to 79%.
| TASK TYPE | FIRST-PARTY ANSWER RATE | SHARE OF THIRD-PARTY CITATIONS |
|---|---|---|
| Integrations | 92-93% | Low |
| Security and compliance | 92-93% | Low |
| Pricing and features | 79% | 77% of all third-party citations |
That last column is the one enterprises underweight. Pricing questions alone generated 77% of all the third-party citations Indig logged across the entire study. In other words: when an agent gives up on a first-party answer, it is overwhelmingly a pricing question that caused the failure, and it is overwhelmingly a third party, a review site, a reseller, a comparison page, that gets the fallback citation instead. That's the exact dynamic we wrote about when the study first landed: sites built as showrooms, full of narrative and nurture sequences, do not serve an agent that just wants the fact.
It's worth asking why pricing specifically is the weak point, when integrations and security pages score so much higher. The honest answer is that most enterprise sales teams have spent years deliberately obscuring pricing, gating it behind a demo request, a sales call, or a vague "contact us for enterprise pricing" page, because that friction has historically been good for the sales motion: it puts a human rep between the buyer and the number, and lets that rep control the framing. That's a defensible tactic when the buyer is a person willing to fill out a form. It's a losing tactic when the buyer is an agent that treats the absence of a static answer as a failed task and moves on to whichever competitor made the number available. The gating strategy that protected margin in a human-mediated sale is the same strategy actively costing visibility in an agent-mediated one.
Integrations and security pages score higher for a related, structural reason: they're usually built as reference documentation from the start, a list of supported integrations, a compliance matrix, a security whitepaper with concrete certifications named. That format was never designed to persuade through narrative, so it happens to already be close to what an agent needs. Pricing pages, built by a different team with a different goal, persuade first and inform second, which is exactly backwards from what an agent-facing page needs to do.
Scoring task completion for your own site means running the same design at a smaller scale: pick your three to five highest-stakes buyer tasks, run each through a general-purpose agent five separate times, and log the first-party completion rate for each. A site scoring above 90% on every task is rare and defensible. A site scoring in the high 70s on pricing specifically is, per Indig's benchmark, the enterprise norm, not an outlier, which is itself the finding worth acting on. Extend the same test beyond pricing to whatever task actually decides your deals, an ROI calculation, a specific compliance certification, a data-residency guarantee, and you'll typically find the same pattern: whichever page was built to persuade rather than inform is the one an agent fails on first.
Dimension 2: Machine-readability under load
A site can pass a task-completion test once and still fail the deeper test that matters for a repeat buyer or a repeat agent visit: does the answer hold up under load, tested more than once, the way a real agent traffic pattern actually behaves. Indig's July 13, 2026 piece on Growth Memo, "Where AI agents get stuck on your site," is the sharper data point here, and it documents exactly the kind of variance enterprises don't expect: the same page, tested in the same hour, swinging from a 73% to a 92.5% correct-answer rate on repeat runs with no change to the underlying content.
That swing is the whole dimension in one number. A page that answers correctly nine times out of ten and incorrectly the tenth time looks fine in a single spot-check and fails silently at scale, exactly when agent traffic is compounding the way Salesforce's own agent-attributed sales figures suggest it now is. The causes we see most often in client audits cluster into a small, fixable set: JavaScript-rendered pricing calculators or configurators that don't resolve to a static, crawlable value; chat-widget gating, where the real answer only surfaces after a conversational flow an agent won't complete the way a human would; PDF-only spec sheets that some agent pipelines parse cleanly and others don't; and cookie-consent or login walls that intermittently block an automated fetch depending on how the request is structured.
Each of these four failure modes shares a common trait worth naming: none of them show up in a standard SEO or content audit, because none of them are visible-content problems. A human reviewer looking at the rendered page sees the price, sees the answer, sees nothing wrong. The failure only appears when the exact retrieval path an agent uses, often a stripped-down fetch without a full browser session, without cookies persisted the way a returning human's would be, hits a wall the human reviewer never encounters. That's precisely why this dimension requires its own testing discipline rather than folding into existing QA, the failure is invisible to every test built around a human user.
There's also a scale dimension to this variance that the same-page swing understates. Indig's test ran on a single page in a single hour. An enterprise site with hundreds or thousands of product, pricing, and spec pages is running this same dice roll across every one of them, continuously, as agent traffic grows. A 73-to-92.5% swing on one page is a data point. The same swing distributed across a full site's worth of buyer-relevant pages is a structural reliability problem, and it's one most enterprises have no monitoring in place to even detect, let alone fix.
Scoring this dimension requires repeat testing, not a single pass. Run your task-completion test from Dimension 1 at least five times, spaced across a day, and treat any answer that isn't stable across all five runs as a readiness failure, even if most of the individual runs succeeded. A single clean pass tells you almost nothing, given the variance Indig found on the same page in the same hour. For enterprises with the engineering capacity, the more rigorous version of this test runs weekly on an automated schedule against your highest-value pages, the same way uptime monitoring runs continuously rather than as an occasional manual check, because a page that was reliable in June has no guarantee of staying reliable in September as the underlying agent tooling itself changes.
Dimension 3: Fallback exposure
When a first-party answer fails, either through the task-completion gap in Dimension 1 or the reliability gap in Dimension 2, the agent doesn't simply return nothing. It falls back to a third party, and Dimension 3 asks a question enterprises rarely track: which third party, and how much of that traffic and trust transfer is recoverable.
The data across our own research this year points to the same handful of fallback categories, repeatedly. Muck Rack's research puts earned media, journalism and press coverage a brand doesn't own, at 82% of all AI citations tracked. Separately, our own analysis of third-party citation composition in B2B categories shows roughly 90% of recommendations pointing somewhere other than the brand's own domain. When we ran the numbers on typical PR outreach lists against the journalists AI engines actually cite, the overlap was as low as 2%, meaning most enterprises are not even positioned to influence the fallback destination their own failed agent tasks are routing to.
Where fallback traffic lands when a first-party agent task fails
Fallback exposure scoring has two parts: how often does your site force a fallback at all, which is the inverse of your Dimension 1 completion rate, and where does that fallback traffic land when it happens. A site with a 79% first-party completion rate on pricing, matching Indig's benchmark, is handing roughly one in five pricing evaluations to a third party by default. The readiness question is whether that third party is one you have any relationship with, or one you've never engaged, which is the far more common answer given the 2% overlap figure above.
It's worth distinguishing fallback exposure from simple competitive loss, because the two are often conflated. Losing a deal to a competitor with a genuinely better product is a normal, healthy market outcome. Losing visibility to a reseller, an outdated aggregator, or a stale review site, none of which are actually competing on product merit, is a pure measurement and infrastructure failure, one with no upside for anyone except the accidental beneficiary of your Dimension 1 gap. The framework's job is separating those two outcomes, since only one of them is fixable by the enterprise itself.
There's a second-order risk inside fallback exposure worth calling out explicitly: accuracy. A third party answering on your behalf, a reseller, an outdated review, a comparison page written by a competitor, has no obligation to get your pricing or positioning right, and often doesn't. We've seen fallback citations quote pricing tiers that were retired two product cycles ago, or describe an integration that was deprecated a year prior. An enterprise scoring poorly on Dimension 1 isn't just losing the chance to make its own case; it's ceding the accuracy of its own story to a source with no incentive to keep that story current.
The practical response isn't to try to eliminate fallback entirely, that's unrealistic even for a site that scores well on Dimension 1, since some share of agent queries will always exceed what any first-party site anticipates. The realistic goal is making sure the fallback destinations that do get cited are ones your team already has a relationship with, through the same earned-media and third-party citation work covered above, so that when the agent does look elsewhere, it's looking somewhere your story is told accurately.
Dimension 4: Access economics
The final dimension is the one most enterprises treat as a pure infrastructure cost rather than a readiness variable, and 2026's crawler-economics data argues that's a mistake. By June 2026, AI training crawls made up 50.6% of all traffic on Cloudflare's network, against just 10.7% for legitimate search bots, the traffic that actually sends value back. Separate Cloudflare research found more than half of AI bot crawls fetch a page that hasn't changed since the previous visit, meaning a large share of that access cost is pure overhead with no new information changing hands.
The economics shifted twice in 2026 alone. Cloudflare's original Pay Per Crawl model let publishers meter and bill crawler access directly, and adoption moved fast: publishers logged well over a billion metered 402 responses a day at points during the year. But as we covered this quarter, Cloudflare has since pivoted toward pay-per-citation, compensating publishers only when content is actually used in a generated answer, alongside a crackdown on "mixed-use" crawlers that misrepresent themselves as ordinary search bots while feeding AI training pipelines. Both changes point the same direction: the cost of staying reachable to agents is no longer a flat infrastructure line item, it's a variable tied to whether that reachability converts into an actual cited or completed task.
It's also the dimension most likely to keep changing underneath enterprises that score well on it today. Task completion and machine-readability are largely under a site owner's own control; access economics depends on decisions made by Cloudflare, by competing CDNs, and eventually by regulators, none of which any single enterprise controls directly. That's not a reason to skip scoring it, it's a reason to re-score it more often than the other three dimensions, since the ground underneath it moved twice in a single year already.
This has a direct, practical implication for enterprises that manage their own infrastructure rather than sitting behind a CDN with a built-in metering layer. Blocking AI crawlers wholesale, still a common default in enterprise robots.txt configurations written before agent traffic existed, doesn't just forfeit citation upside. Given the task-completion data in Dimension 1, it guarantees a fallback to a third party for every buyer task an agent attempts, since a blocked crawler can't complete any task on your domain at all, first-party answer rate included. A blanket block converts every one of your agent-mediated buyer interactions, the 79% pricing tasks that would have succeeded and the 21% that wouldn't have, into the same 100% fallback rate. That's a strictly worse outcome than even the weakest first-party completion rate in Indig's dataset.
Scoring access economics means answering three questions concretely: is your infrastructure explicitly allowing the agent user agents relevant to your buyers, rather than blocking them by default or by country; do you have visibility into what share of your crawl traffic is redundant re-fetching versus genuinely new agent activity; and are you positioned to capture value, through a citation-based or pay-per-use model, from the access you're already granting, rather than treating every crawl as pure cost. Enterprises running their own infrastructure without a CDN-level metering option should still track the first two questions manually; server log analysis, filtered for known AI user agents, is a low-cost way to at least establish a baseline before any pricing model is available to layer on top.
The scoring rubric
The model below is our proposed scoring rubric, illustrative in its weighting rather than a benchmarked industry standard, since agent-readiness data is still early relative to the citation-readiness research behind our original GEO framework. Each dimension is scored 0 to 25, for a possible 100 points total.
| SCORE BAND | LABEL | WHAT IT MEANS |
|---|---|---|
| 80-100 | Agent-ready | Agents complete core buyer tasks reliably and fallback exposure is limited and understood. |
| 60-79 | Partially exposed | Some tasks complete reliably; others, often pricing, route to third parties you don't influence. |
| Below 60 | Agent-blind | Most buyer tasks fail first-party, fallback destinations are unmanaged, and access economics are untracked. |
Applying Indig's benchmark figures to this rubric as an illustrative example: a site scoring 92% on integrations and security tasks, but only 79% on pricing, with no visibility into fallback destinations and no active management of crawler economics, would likely land in the 60-79 "partially exposed" band overall, dragged down not by its strongest dimension but by the combination of a weaker Dimension 1 score on pricing and a near-total gap on Dimensions 3 and 4. That combination, strong on paper, exposed in practice, is the profile this framework is built to catch.
It's also worth stating what this rubric deliberately doesn't do: it doesn't attempt to convert a score into a projected revenue impact. Enterprises will be tempted to ask for that conversion immediately, since it's the number that gets budget approved. We're withholding it on purpose, for the same data-honesty reason we withhold invented statistics anywhere else in our work: no dataset yet exists tying agent-readiness scores to closed revenue at the scale needed to make that claim responsibly. What the four figures in this paper's executive summary do support is a directional argument, agent traffic is large and growing, task failure is common, and the infrastructure economics underneath it all are shifting fast, which is enough to justify running the audit even without a revenue model attached to the output yet.
A brief note on how we weight the four dimensions equally in this first version, rather than weighting task completion more heavily as the seemingly more important variable. The reason is that a strong Dimension 1 score with a weak Dimension 4 score is a fragile position, not a strong one: a site with excellent task completion today, but blocked or throttled crawler access tomorrow, loses the benefit of Dimension 1 entirely the moment access changes, which given how fast the economics moved in 2026 alone is not a hypothetical risk. Equal weighting reflects that all four dimensions are load-bearing, and a genuinely high score requires clearing a real bar on each, not compensating for a weak one with strength elsewhere.
A worked example
To make the rubric concrete, consider a hypothetical mid-market B2B software vendor, illustrative rather than a specific named account, running this framework for the first time. Its integrations and security pages, built as structured reference documentation years ago, score close to Indig's 92-93% benchmark on Dimension 1 for those task types: call it 22 out of 25. Its pricing page, gated behind a "request a quote" form with no static number anywhere in the rendered HTML, fails every agent task attempt outright: a 4 out of 25 on that same dimension, dragging the blended Dimension 1 score down substantially once pricing is weighted in as one of the three to five core tasks the framework recommends testing.
Its Dimension 2 score depends on whether that same gated pricing page at least fails consistently, which it does, since a hard gate always fails the same way, versus the JavaScript-configurator problem, which fails inconsistently and would score worse under repeat testing. Assume a consistent, if total, failure: 18 out of 25, penalized for the pricing gate but not for added same-page variance. Its Dimension 3 score is weak, 8 out of 25, since it has no active earned-media program and no visibility into where its pricing fallback traffic lands. Its Dimension 4 score is middling, 15 out of 25: crawlers aren't blocked outright, but nobody on the team has reviewed the robots.txt configuration since before agent traffic existed, and no pay-per-citation or equivalent model has been evaluated.
Total: 22 + 18 + 8 + 15 = 63, squarely in the "partially exposed" band. The instructive part of this hypothetical isn't the specific numbers, it's the diagnosis they produce: a single fix, un-gating pricing behind a static number, would move Dimension 1 and Dimension 2 both, likely pushing the total score into the 80s without touching Dimension 3 or 4 at all. That's the value of scoring the dimensions separately rather than reporting one blended number: it tells you which single fix has the highest leverage, instead of leaving you with an aggregate score and no sense of where to start.
Where this diverges from citation readiness
It's worth being explicit about why a brand can score well on our original GEO readiness framework, which measures accessibility, structure, authority, and coverage, and still score poorly here. Citation readiness is largely a content and structure problem: is the page reachable, is the fact extractable, does the source carry trust. Agent readiness adds an operational layer on top: can the specific task actually be completed, does the answer hold up on repeat testing, and what happens, concretely, when it doesn't. A brand can have excellent content and a broken pricing configurator at the same time, and the citation framework won't catch that failure, because it isn't testing task completion, it's testing whether the page could theoretically be cited if an agent got that far.
This is also why the two frameworks should run together, not as a replacement for one another. A site with strong Dimension 1-4 scores here, but weak accessibility or authority scores on the original framework, is agent-ready for the buyers who already trust it, but invisible to the ones who haven't found it yet. The reverse, strong citation readiness with weak agent readiness, is arguably the more dangerous blind spot in 2026, because it's the one enterprises are least likely to be measuring at all right now. A marketing team celebrating a rising mention rate has no way of knowing, from that number alone, whether the agents generating those mentions are also completing the tasks that actually close deals.
What this means by role
The four dimensions land differently depending on who in the organization owns the fix, and naming that ownership up front is usually the difference between a framework that gets run once and one that gets run quarterly. Marketing and content teams own most of Dimension 1 and a meaningful share of Dimension 2, since the decision to gate pricing behind a form, or to build a spec sheet as a PDF instead of a page, is usually a content and positioning call rather than an engineering one. RevOps and sales leadership own the harder conversation underneath Dimension 1: whether gating pricing behind a human rep is still worth the friction it now costs in agent-mediated deals, a tradeoff that used to have an obvious answer and no longer does.
It's worth adding a note on budget ownership specifically, since that's usually the harder organizational question than task ownership. None of the four dimensions require a large new line item to begin measuring; the audit itself is a few days of cross-functional testing time, not a new tooling purchase. The fixes that follow, un-gating a pricing page, rebuilding a PDF spec sheet as a static page, do carry real engineering cost, but they're the same class of cost as any other conversion-rate-optimization fix, evaluated against the same kind of payback logic, not a speculative AI investment with no clear return.
Engineering and platform teams own most of Dimension 2 directly, the JavaScript-rendering and consent-wall issues that show up as same-page variance, along with the technical half of Dimension 4, robots.txt configuration, crawler-traffic monitoring, and evaluating pay-per-citation or equivalent infrastructure options as they mature. Communications and PR teams own Dimension 3, since fixing fallback exposure runs through the same earned-media and third-party relationship work that's always been their mandate, just newly measured against a different outcome. No single role owns all four dimensions, which is precisely why this framework works best run as a cross-functional quarterly review rather than a single team's audit.
What to do this quarter
Run the four-dimension test on your three highest-stakes buyer tasks this quarter, not next year's roadmap. Pricing first, since it's where Indig's data shows both the lowest first-party completion rate and the highest share of third-party fallback citations. Test each task five times, spaced across a day, to catch the same-page variance Dimension 2 is built to surface. Log where the fallback lands when a task fails, and check that destination against your own PR and citation-building relationships, the same audit discipline we open every engagement with. And treat your crawler access settings as a live economic decision, not a set-and-forget robots.txt line, given how fast the underlying terms shifted in 2026 alone. Track your progress against the same topical-authority data that shows most categories are still winnable, and against our reporting model built to track outcomes like this one alongside pipeline, not in isolation.
We're running this exact framework as the opening audit on new GEO engagements starting this quarter, the same way our original readiness framework became the opening audit for citation work two years ago. Two current engagements illustrate both ends of the spectrum: a B2B agentic security platform that scored strongly on task completion and machine-readability from the start, because its documentation was already built API-first and pricing was never gated behind a calculator, and a DevSecOps platform that had the opposite starting profile, strong content, weak task completion, and needed the pricing and integrations pages rebuilt as static, agent-parseable pages before its citation work could compound. Both are proof the same framework, applied honestly, tells you which kind of problem you actually have before you spend a quarter solving the wrong one. Citation readiness got you into the conversation. Agent readiness decides whether you get the sale.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.