The Measurement Void: Agentic AI's $1.5 Billion Story and the 95% That Never Shipped
Hook: Two Numbers That Should Never Have Shared a Headline
Seven billion. That is the Agentic Work Unit figure Salesforce put in front of investors to prove its agent platform had momentum, bundled alongside a $1.5 billion ARR claim. Then there is the other number, the one that rarely makes the deck: ninety-five percent of enterprise GenAI pilots have failed to reach expected returns, according to MIT's NANDA research. And a third, quieter still: eighty percent of enterprises have embedded AI somewhere, but only thirty-one percent run agents in production.
These figures do not describe a market that is growing. They describe a market that is being measured into existence.
I have spent twelve years reading on-chain accounting, and I recognize this pattern instantly. It is the same arithmetic that let a DeFi protocol claim billions in total value locked while most of that capital was recursive, self-referential, and one bad oracle call away from evaporation. The number was real. The meaning behind it was not. What agentic AI has done in the last eighteen months is invent a new unit of account that no competitor can audit, no buyer can compare, and no regulator has touched β and then pointed at the growth of that unit as evidence of value. That is not a metric. That is a narrative wearing a metric's clothes. The question every allocator should be asking, and almost none are, is what happens when the clothes come off.
Context: How the Agent Economy Ended Up Without a Balance Sheet
To understand why the measurement problem is structural rather than cosmetic, you have to understand what an agent actually is beneath the marketing.
A large language model answering a question is a single forward pass, more or less. You give it context, it produces tokens, you pay for those tokens, and the transaction closes. Pricing this is trivial: input tokens plus output tokens, multiplied by a rate card. The unit economics are legible because the computation is legible. One request, one bill.
An agent is not that. An agent is a loop. You give it a goal, it plans, it calls tools, it observes the results, it critiques its own output, it retries, it calls more tools, it reconciles conflicting information, and eventually β if you are lucky β it terminates with something that resembles an answer. Sometimes it terminates because it succeeded. Sometimes it terminates because it hit a step limit. Sometimes it terminates because it hallucinated a completion and confidently stopped. The bill, however, is always the same shape: you pay for every step of that loop, including the ones that were wasted.
This is where the McKinsey finding that sixty percent of agentic AI spend goes into iterative response optimization β the check, correct, improve cycle β stops being a budget line and starts being an architectural confession. The cost center is not generation. The cost center is doubt. The agent's expense is proportional to how uncertain it is about its own work, multiplied by how many times you let it second-guess itself before you accept the output.
The enterprise software industry has spent two decades learning to price deterministic work. A CRM seat costs what it costs because the work it enables is bounded and countable. An API call costs what it costs because the computation is fixed. Agentic AI broke both assumptions at once. The work is unbounded, the computation is stochastic, and the outcome quality varies wildly across inputs of apparently identical difficulty. You cannot price that with a seat, and you cannot price it honestly with a token, because the honest token price of a task depends on how badly the model performed on that specific task on that specific day.
So the industry did what industries always do when they cannot price something honestly. It invented a proxy, optimized the proxy, and started selling the proxy. Salesforce called its proxy the Agentic Work Unit. It is not alone in spirit, only in branding.
There is a timing tell in all of this. The discourse that produced these numbers sits around mid-2026, with Salesforce's September Dreamforce as the next scheduled moment where the vendor narrative gets refreshed. I flag this deliberately, because it means the specific product claims and price points circulating in this conversation sit outside the window I can independently verify. What I can verify, and what matters more than any single figure, is the shape of the incentive structure producing those claims. The measurement void is not a bug in the agent economy. It is the load-bearing wall of the agent economy's valuation.
Core: The Correction Loop Is the Product, and That Is the Problem
Let me start where the money actually goes.
The reflex most people have when they hear 'AI agent' is to imagine a smarter chatbot. That mental model is wrong in a way that distorts every downstream financial calculation. A chatbot's cost is dominated by the generation step. An agent's cost is dominated by the verification step. When you read that sixty percent of spend flows into iterative response optimization, you are reading a description of a system whose economics live in reflection, self-critique, and tool-mediated retry.
In the technical literature this maps onto a family of patterns β ReAct-style reasoning and acting, Reflexion-style self-correction, and the broader generator-verifier split where one model produces and another adjudicates. What these patterns share is a brutal property: each additional correction cycle costs the same as the first pass but delivers diminishing accuracy gains. The first retry often buys a meaningful jump in reliability. The fifth retry frequently buys noise. The tenth retry buys the illusion of diligence and a bill to match.
I have a personal frame for this. In 2020, during the DeFi Summer liquidity crunch, I watched Compound's governance forums while a price spike threatened the cToken collateral factors. What struck me was not the spike itself but how the protocol's safety depended on a liquidation cascade behaving the way the model said it would β and how little margin there was between 'the model is right' and 'the model is catastrophically wrong.' I published a rapid technical breakdown within hours, citing on-chain metrics from Etherscan and predicting cascade failure if minting was not paused. The lesson I took from that week was not about Compound specifically. It was that systems which depend on iterative error-correction to stay safe are only as strong as their weakest correction step, and you almost never see that step in the dashboard.
An agent architecture is that lesson industrialized. Reliability is not a property the agent has. It is a property the agent's loop produces, and only for the classes of input the loop has been tuned against. Take it out of the controlled pilot and drop it into an unpredictable production environment and you get distribution shift β the same phenomenon that turns a backtested trading strategy into a losing one the moment conditions move. Tool calls fail. Context windows get polluted by irrelevant retrievals. Multi-step errors accumulate because each step's output becomes the next step's premise. The MIT NANDA finding that ninety-five percent of pilots miss expectations is not a story about lazy enterprises. It is a story about systems whose error-correction machinery does not transfer cleanly from the lab to the wild.
The corollary is uncomfortable for everyone selling agents. If the cost center is verification, then the most competitively valuable architectural work is the work that reduces the number of verification cycles required β better single-pass reasoning quality, stronger dedicated verifiers, more aggressive caching and reuse of intermediate results. This is where the technical edge will actually be decided, and it is almost entirely invisible from the outside. A vendor that has halved its average loop depth has a structurally better product than one that has not, but you cannot tell from either UI, and neither vendor will volunteer the number, because the loop depth is precisely what their cost model is built to monetize.
Where the money the buyer pays actually lands
Here is the part that should worry a finance team more than any model benchmark.
When a vendor charges by tokens consumed, they are charging for the agent's uncertainty. A clean task that the model nails on the first pass costs a little. A messy task that sends the model into a nine-round spiral of doubt, tool calls, and half-corrections costs many multiples of that β and the buyer pays the full tab for the spiral. In effect, token-based pricing transfers the agent's inefficiency directly onto the customer's P&L. The customer is not paying for outcomes. The customer is paying for the vendor's loop count, and the vendor has no incentive to reduce it.
This is why the more sophisticated procurement conversations have started to migrate toward cost-per-outcome. Buyers have realized that a deterministic KPI like 'records processed' or 'tokens consumed' is a measurement of activity, not value. And here is the structural trap: cost-per-outcome sounds obviously correct, but it is far harder than it looks. It requires someone to define what counts as an outcome, to verify that the outcome was actually achieved, to adjudicate disputes when the agent claims success and the buyer disagrees, and to price the tail risk of an outcome that looked good for a quarter and turned out to be wrong. That is not a pricing model. That is an audit function, and it does not exist yet at scale.
Core: AWU Is a Moated Vanity Metric, and the Moat Is the Point
Now the unit itself.
An Agentic Work Unit is a proprietary measure of activity. Salesforce reports seven billion of them. What is one AWU? The honest answer, as far as any external observer can tell, is that it is whatever Salesforce says it is. It is not derived from a public formula. It is not comparable to any competitor's metric. It is not auditable by a third party. And that is not a flaw in the design. That is the design.
I have seen this exact move before, and I have been on both sides of it. When a market has no shared unit of truth, the first large player to establish a proprietary unit gets to define the conversation, and everyone else is forced to argue on that player's terms. During the Axie Infinity cycle in 2021, I audited token emission schedules and found a seventy-two-hour window where staking rewards outpaced inflation. I quantified it, published a strategy to a private channel, and the model returned twenty-two percent in four days. The thing that made that window exploitable was not that the numbers were hidden β it was that most participants were looking at the headline metric, not the flow behind it. The headline metric told a story about growth. The emission schedule told a story about who was actually paying for it.
AWU is a headline metric. It is engineered to be unfalsifiable in the near term and impressive in the medium term. Seven billion work units is a number with no denominator that would embarrass it β no failure rate attached, no cost-per-successful-outcome, no error distribution. It measures the volume of activity passing through the platform. The volume of activity passing through a platform is exactly the metric an intermediary wants you to watch, because it is the metric that grows when the intermediary becomes more central β regardless of whether the customer gets richer.
This is the venue-versus-user distinction that anyone who has actually traded understands at a bone level. An exchange can print record volume while its average customer loses money. A chain can print record transactions while most are bots and arbitrageurs extracting from retail. Activity is the venue's revenue. Value is the user's outcome. When a vendor conflates the two, they are not lying about the number. They are lying about what the number means, and they are relying on you not to notice the difference.
There is a second layer to the moat. A proprietary metric is not just marketing. It is a switching cost. Once an enterprise has built internal reporting around AWU β dashboards, board decks, procurement benchmarks, analyst briefings β it cannot easily migrate to a competitor whose agents are measured in a different unit, because the enterprise would have to rebuild its entire evidence base from scratch. The metric becomes sticky in the same way a data format becomes sticky. Nobody wants to rip out a reporting layer and start over. So the buyer stays, not because the product is superior, but because the ruler is theirs.
Core: The Budget Misallocation Nobody Audits
Here is the finding that the vendor narrative quietly buries, and it is the most actionable thing in this entire conversation.
More than half of enterprise GenAI budgets flow into sales and marketing. And the scenario with the most consistently measurable ROI is back-office automation. The front office gets the money. The back office gets the returns. Nobody audits the gap.
Think about why this happens. Sales and marketing have always been the loudest buyers of anything that promises more pipeline, because their performance is measured in growth and growth buys organizational status. An AI agent that promises better lead scoring, faster outreach, more personalized campaigns, is an easy sell to a revenue organization that is under quarterly pressure. The pitch lands. The pilot launches. And then the measurement problem bites, because sales and marketing outcomes are among the hardest things in any business to attribute to a single tool. Did the agent close the deal, or did the rep, or did the timing, or did the market? The metric dissolves the moment you look at it directly.
Back-office automation is the opposite. A document-processing agent that extracts fields from invoices is measured against a ground truth: either the field is right or it is not. A reconciliation agent is measured against a ledger: either the books balance or they do not. A compliance agent is measured against a rule set: either the control was satisfied or it was not. These outcomes are countable, falsifiable, and attributable β precisely the properties the front office lacks and precisely the properties finance departments need to justify a line item.
So why does the money not go there first? Because the back office does not have the political power to demand an AI budget, and the front office does. The most measurable ROI sits in the least empowered department. This is a governance failure masquerading as a technology strategy, and it is quietly burning well over half the capital that enterprises are deploying into agentic AI.
I recognize this failure mode from the Terra-Luna collapse in 2022. I treated it as a data-rich post-mortem rather than a tragedy, and within forty-eight hours I published a dissection of the UST de-pegging mechanism, citing specific smart-contract vulnerabilities in Anchor Protocol. The lesson that survived wasn't about algorithmic stablecoins. It was that capital had been flowing toward the highest-yield narrative, and the highest-yield narrative was exactly where the risk was concentrated and the measurement was weakest. Back-office automation is the highest-yield narrative in agentic AI. It is also the one nobody is funding, because the people who would benefit from funding it do not control the budget. The misallocation is not irrational in the sense of being random. It is rational for the individual decision-maker and destructive for the enterprise β which is the most dangerous kind of misallocation there is.
Core: The Token Math Nobody Publishes
Now zoom beneath the business layer into the physics that constrains it.
If sixty percent of agentic spend is iterative optimization, then the compute profile of an agent is not 'conversation-grade.' It is 'task-grade.' A single agent doing a real multi-step job β retrieving, planning, calling tools, checking, retrying β can consume one to two orders of magnitude more inference than a single chat interaction. Multiply that across a production fleet and you get an inference load that the conversational-AI era never priced in.
This is the hidden ceiling on agentic AI. It is not model capability. It is the cost curve of inference compute, and it bites hardest exactly where agents are most valuable: long, complex, tool-heavy tasks with lots of correction loops. A short task with a clean first pass is cheap to run and cheap to sell. A long task that spirals through five correction cycles is expensive to run and β under token pricing β equally expensive to sell, which means the vendor is happy and the buyer is bleeding.
The demand side already senses this. When buyers push for cost-per-outcome, part of what they are buying is protection against the loop count. They are saying, in effect, that they refuse to pay for the agent's uncertainty. That is a demand for inference efficiency whether or not anyone utters the phrase. The agent economy is, at bottom, a bet on a falling inference-cost curve, and very few of the valuations being placed today are explicitly underwritten by that belief.
There is a quiet upside here that the negativity of the discourse obscures. If you can lower the average loop depth β better single-pass quality, cheaper verifier models, more aggressive KV-cache and prefix-cache reuse β then you lower cost-per-outcome without touching the underlying model's intelligence. That is a pure margin extraction, and it is arguably worth more commercially than a marginal benchmark improvement. Architecture efficiency is the new competitive axis, and it is invisible in the marketing.
We do not have the average token-per-task figure for a production agent fleet. We do not have the distribution. We do not have the sensitivity of cost to loop depth. This is the single most important missing measurement in the entire industry, because it determines which agent products can ever cross from pilot to production economics, and it is the number every vendor has an incentive to suppress.
Core: Crypto Already Ran This Experiment
The reason I keep reaching for crypto parallels is not rhetorical. It is that the digital-asset industry spent a decade learning β painfully β what happens when an industry adopts a metric that is easy to inflate and hard to falsify.
Total value locked. It was the defining metric of the last cycle. It appeared in every pitch deck, every dashboard, every ranking. It was also enormously gameable, because capital that is deposited and redeposited across protocols can be counted more than once, and because TVL says nothing about whether the capital is productive or merely parked. A protocol with a billion in TVL could be a billion in genuine, sticky, revenue-generating liquidity, or a billion in recursive leverage that unwinds the moment incentives stop. The number was identical in both cases. The value was not.
The industry eventually learned to ask harder questions. Not just how much is locked, but what it earns. Not just deposits, but net flows. Not just activity, but retained value. That shift from vanity metrics to real-yield metrics took years and cost a lot of people a lot of money to learn. Agentic AI is running the identical experiment in real time, with AWU playing the role TVL played, and with the same likely outcome: a reckoning in which the metric that carried the narrative gets repriced against the metric that carries the cash flow.
There is a further parallel that should make enterprise buyers sit up. In crypto, the entities that survived the metric reckoning were the ones whose revenue was real and whose costs were legible. The entities that did not were the ones whose valuation rested on a number nobody could decompose. Agentic AI has both populations right now. There are agent deployments with genuine, falsifiable, attributable outcomes β mostly in the back office, ironically unglamorous, ironically unloved. And there are agent deployments whose entire justification is an activity metric that grows because the platform is being used more, not because the customer is getting more.
Arbitrage isn't about seeing a different number. It is about seeing the same number and understanding what produced it. When I looked at BlackRock's S-1 filings ahead of the 2024 spot Bitcoin ETFs, the edge was not a secret document. It was the discipline of reading the submission timeline and the legal comments and asking what they implied about probability, rather than what they implied about enthusiasm. I put a 94 percent approval probability on the table by May of that year, anchored to specific legal precedents, and organized three analysts to track the filings. When approval came, that earlier work was cited in major financial press β not because I had access nobody else had, but because I had interpreted the same public information more rigorously. The same discipline applies here. Every number in the agent economy is public. The edge is in decomposing it.
Contrarian: The Angle You Will Not Read Anywhere Else
The consensus critique of agentic AI is that the technology is overhyped and the returns will disappoint. That critique is directionally right and analytically lazy. It misses the two things that actually matter.
The first contrarian claim: the measurement void is not a temporary phase of immaturity, and it will not be resolved by the market naturally discovering a good standard. The reason is incentive, not complexity. A shared, auditable, cross-comparable outcome metric would be wonderful for buyers and catastrophic for the largest sellers, because it would let buyers compare agents on the only axis that matters β cost per real, verified outcome β and that comparison would expose exactly how much of the current pricing is being charged for loop counts the customer never wanted. The incumbent with the proprietary metric has every reason to defend it and no reason to converge. The third parties calling for standardization β the analyst houses, the consultancies β have their own subtle incentive: a chaotic measurement landscape generates enormous demand for their services as the 'meaning-making layer.' Gartner predicts cancellations. McKinsey diagnoses overspend. The Futurum Group tracks the shift toward financial ROI. All three benefit from being the ones who explain the numbers to a confused market. Their calls for standardization are not cynical, but they are not disinterested either, and the fact that nobody in this conversation discloses that tension tells you how young the discourse is.
The second contrarian claim is the one that should genuinely unsettle a buyer: measuring activity instead of outcomes is not merely inefficient. It is a form of liability accumulation. Every token spent on a correction loop that ran without a verifiable outcome is a cost recognized as a benefit. Every AWU counted without an accompanying success rate is a data point that will eventually be repriced. The standard framing is 'activity metrics are non-ideal but tolerable.' The harder framing is 'an activity metric that inflates a vendor's ARR is a deferred write-down sitting on the customer's books.' When Gartner predicts that forty percent of agentic AI projects will be cancelled, the market reads that as a forecast about vendors. It is more accurately a forecast about the buyers who are currently reporting benefits they cannot measure and will be forced to stop reporting them. The cancellation is downstream. The overstatement already happened.
And here is the piece that almost no analyst has connected: the overstatement is concentrated in the exact scenario where measurement is hardest, which is the front office. Half the budget is flowing into sales and marketing agents, where outcomes are least attributable and where 'activity' is most easily mistaken for 'value.' The most inflated claims live where the ground truth is weakest, which means the coming repricing will be sharpest exactly where the budget is heaviest. The back office, starved of funding, is where the claims are conservative because they have to be β a reconciliation tool either balances the ledger or it does not.
The final contrarian observation is about who actually benefits from the void. It is not the model providers, who sell tokens and get paid for loop depth regardless of outcome. It is the vendors who own the customer relationship and can therefore define the unit of account. Defining the unit is worth more than producing the intelligence. In the agent economy, the ruler matters more than the thing being measured, and the industry is currently fighting over who holds the ruler, not over who builds the best agent. That fight is invisible if you only watch benchmarks. It is the whole game if you watch ARR.
Takeaway: The Two Variables That Decide Everything
Forget the benchmarks. Two variables will determine whether agentic AI is a durable industry or a repriced narrative, and both are measurable within the next twelve to eighteen months.
The first is who owns the cost-per-outcome standard. If a credible, third-party-audited, cross-vendor outcome metric emerges and gets adopted by major enterprise buyers, the pricing power shifts decisively to the buy side, loop-count billing dies, and the vendors with genuine architectural efficiency win. If no such standard emerges, the proprietary metric survives, pricing stays opaque, and the buyer keeps paying for uncertainty indefinitely. Watch for whether any standardization effort gains real enterprise adoption, or whether it remains an analyst talking point with no procurement teeth.
The second is whether the inference-cost curve descending is actually captured by the loop-count curve descending. If correction-loop depth falls β through better verifiers, cheaper validation models, aggressive cache reuse β then cost-per-outcome improves without any change in model intelligence, and agentic AI climbs down the cost curve toward viability. If loop depth stays flat while model quality plateaus, the economics get stuck, and the addressable market is quietly capped at the tasks wealthy enough to afford the spiral.
If neither variable moves within a year and a half, the warning embedded in all of this stops being a warning and becomes an accounting fact: at that point, the volume a platform processes is no longer evidence of value delivered to the customer. It is a record of cost transferred onto the customer. Activity becomes a liability, counted as an asset, for as long as the ruler stays proprietary. We don't get to choose whether the repricing happens. We only get to choose whether we have already done the math before it does.