<?xml version="1.0" encoding="UTF-8" ?><!-- generator=Zoho Sites --><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><atom:link href="https://www.aheadcrm.co.nz/blogs/tag/Pricing/feed" rel="self" type="application/rss+xml"/><title>aheadCRM - Blog #Pricing</title><description>aheadCRM - Blog #Pricing</description><link>https://www.aheadcrm.co.nz/blogs/tag/Pricing</link><lastBuildDate>Wed, 23 Sep 2026 07:52:34 -0700</lastBuildDate><generator>http://zoho.com/sites/</generator><item><title><![CDATA[Usage-Based Pricing for Copilot Is Good for Microsoft's Investors. Read That Sentence Again.]]></title><link>https://www.aheadcrm.co.nz/blogs/post/usage-based-pricing-for-copilot-is-good-for-microsofts-investors-read-that-sentence-again</link><description><![CDATA[TheStreet ran a piece this week arguing that, of Microsoft's two Copilot announcements, the shift to usage-based pricing matters more to investors tha ]]></description><content:encoded><![CDATA[<div class="zpcontent-container blogpost-container "><div data-element-id="elm_I4c1lqnDTGyUwhvs3m8HUg" data-element-type="section" class="zpsection "><style type="text/css"></style><div class="zpcontainer-fluid zpcontainer"><div data-element-id="elm_3_BOC-oNSD6e_4nuIxYi5A" data-element-type="row" class="zprow zprow-container zpalign-items- zpjustify-content- " data-equal-column=""><style type="text/css"></style><div data-element-id="elm_hReP0NBfQJSCzo2rxmtQzA" data-element-type="column" class="zpelem-col zpcol-12 zpcol-md-12 zpcol-sm-12 zpalign-self- "><style type="text/css"></style><div data-element-id="elm_67tNEolpSCKsWp5BOxfUcw" data-element-type="text" class="zpelement zpelem-text "><style></style><div class="zptext zptext-align-center " data-editor="true"><div><p>TheStreet <a href="https://www.thestreet.com/technology/microsoft-copilot-power-user-pricing">ran a piece</a> this week arguing that, of Microsoft's two Copilot announcements, the shift to usage-based pricing matters more to investors than the DeepSeek flirtation. That read is correct. It is also the tell.</p><p>Here is what Microsoft actually did. Copilot Cowork, the agent that reaches across Microsoft 365 to run multi-step work on your data, is coming off the flat per-seat add-on and moving onto consumption billing the company calls &quot;Copilot Credits.&quot; Charles Lamanna, who runs Copilot, told Axios the product could not be offered on an unlimited-use basis. The users he pointed to are the ones doing hundreds of tasks a week. He called them &quot;way productive.&quot; And then he said the part vendors normally keep off the slide: their costs go very high.</p><p>So the most productive users are the expensive ones. Hold that thought, because the whole argument lives there.</p><h1 class="wp-block-heading">What &quot;good for investors&quot; is really saying</h1><p>A pricing model earns the label &quot;good for investors&quot; when three things are true. Revenue starts to track cost-to-serve. Revenue scales with consumption instead of sitting flat per seat. And the vendor stops eating the margin on its heaviest users. All three are true here. None of them is a statement about whether a customer got value.</p><p>That is the gap I want to sit in for a minute.</p><p>Usage-based pricing meters an input. Tokens, compute, credits, whatever the unit. The customer does not buy tokens because they want tokens. They want a finished report, a resolved ticket, a reconciled spreadsheet. The token count is the cost of producing the outcome, not the outcome. And the relationship between the two is loose at best.</p><p>TSIA <a href="https://www.tsia.com/blog/ai-pricing-models-usage-based-outcome-based-hybrid">put it plainly</a> in its May analysis of AI pricing: usage does not equal value, and consumption models often fail to reflect actual business value. That is not a critic talking. That is a research firm whose audience is the vendors building these models.</p><h1 class="wp-block-heading">Agents make the coupling worse, not better</h1><p>A chatbot answers and stops. An agent keeps going. It reads files, calls tools, checks its own work, hits a wall, tries again. Each of those steps burns compute, and the steps are what get metered. Every retry, every verbose detour, every loop the agent runs to second-guess itself adds to the bill. The customer pays for all of it.</p><p>Now ask yourself the hard question. Does a workflow that took the agent three retries and a long chain of self-checks deliver more value than the same workflow done cleanly in one pass? Of course not. It delivers the same outcome and costs more. Under seat pricing, that inefficiency was the vendor's problem. Under usage pricing, it is line-itemed onto the buyer's invoice.</p><p>The billing platform Flexprice, which sells the plumbing for this, says it out loud to its own customers: <a href="https://flexprice.io/blog/how-to-price-ai-agent-usage-based-pricing">retries, loops, and background jobs are friction</a>, not value, and usage-based pricing only works when customers gain something real as the meter climbs. Their warning to vendors is the buyer's whole case.</p><h1 class="wp-block-heading">The people who get punished are the people who bought in</h1><p>We do not have to guess how this lands, because GitHub Copilot already ran the experiment. On June 1 it moved to token billing. The median user barely noticed. The pain landed on the top five to ten percent, and it landed hard: community projections of bills jumping ten to fifty times, one developer modeling a move from roughly $29 a month to nearly $750, another claiming <a href="https://www.reddit.com/r/GithubCopilot/comments/1tqca76/comment/oofol56/?screen_view_count=25&amp;rdt=65117">$50 to $3,000</a>. TechCrunch called it <a href="https://techcrunch.com/2026/05/30/what-a-joke-github-copilots-new-token-based-billing-spurs-consternation-among-devs/">the end of Copilot's golden age</a>.</p><p>Look at who those heavy users are. They are not abusers. They are the people who took the vendor's three-year advice to use the tool for everything, built agentic workflows around it, and made it part of how they work. The pricing change penalizes exactly the depth of adoption every vendor claims to want. And the old safety net, where running out of premium budget dropped you to a cheaper model so you could keep working, is gone. What is gone, too, is the cost ceiling.</p><p>Lamanna's &quot;<em>way productive</em>&quot; power user and GitHub's top-decile developer are the same person. The model charges most to the customer who is succeeding most. Reward and penalty have swapped places. The reward now goes to the vendor.</p><h1 class="wp-block-heading">The structural problem, which is bigger than the bill</h1><p>Now the second half, and this part is more important than any individual invoice.</p><p>Think about where accountability for value sits in each pricing model. With outcome-based pricing, the vendor gets paid when a result lands and not before. Fin's (formerly known as Intercom) Fin charges 99 cents per resolution, billed only when the customer confirms the AI actually solved the problem. Under that model, every failed attempt costs the vendor. So the vendor has a direct, financial reason to make the agent efficient, accurate, and sparing with compute. Their margin depends on it.</p><p>Usage pricing inverts that incentive. The vendor is paid for activity regardless of whether the activity worked. An agent that burns more tokens, retries more often, and reasons more verbosely produces more revenue, not less. I am not claiming Microsoft will deliberately bloat Cowork to pump credits. I am saying the financial pressure that used to push toward lean, effective agents has been switched off; and switched-off incentives have a tendency of showing up in the product eventually.</p><p>The demand side pulls in the same direction. There is a name for it now: tokenmaxxing, the workplace habit of treating AI usage as a proxy for productivity, where people get judged on how many token they burn rather than on what they shipped. Built In's <a href="https://builtin.com/articles/ai-tokenmaxxing">writeup</a> is blunt about it: the habit rewards visible activity, not results. So stack the three forces. Buyers under pressure to run up consumption as a status signal, a vendor that meters by consumption, and an agent that inflates consumption on its own. Everything drives the meter up. Nothing points it at the outcome.</p><p>That is the real cost of the model. It moves the vendor one step further from owning the question of whether you got value, and it hands that entire question to you. The vendor essentially plays <a href="https://en.wikipedia.org/wiki/Pontius_Pilate">Pontius Pilate</a>. The buyer now runs FinOps for AI. You set the budget caps. You write the spending policies. You read the consumption dashboard. You type /cost to see what a task burned. Microsoft, to its credit, is shipping all of those controls, and they are better than the ones GitHub fumbled out the door. But notice what they are. They are tools for the customer to govern value. They are not the vendor guaranteeing it.</p><h1 class="wp-block-heading">The honest counterargument</h1><p>I would be doing the same vendor-spin thing I just criticized if I left out the other side.</p><p>Flat pricing for agentic tools is inherently unsustainable. The economics are upside down: the model subsidizes the heaviest five percent and overcharges the lightest fifty. Metered billing is the rational fix for that, and for a low-volume or experimental buyer it is a better deal than paying a fat seat fee to barely use the thing. Aligning price with cost-to-serve is a good thing. It is just a vendor virtue, not a customer one, and the trick to watch is anyone presenting the first as if it were the second.</p><p>Cheaper models do not solve the coupling problem. A fine-tuned DeepSeek on Azure lowers the unit price of the metered thing. It does not make the metered thing track value. You are paying less per token for a number that still has a loose relationship to your outcome. And who knows how many additional tokens a potentially inferior model burns.</p><h1 class="wp-block-heading">Where this lands</h1><p>On my orchestration battleground, Cowork is the M365 layer that coordinates work across your apps and your Graph. Pricing that orchestration by consumption reframes it from a capability you own into a utility you rent by the drink. That is a substantial shift in who carries the risk when an orchestrated workflow goes long, and it is the buyer.</p><p>So, on the null hypothesis. Is usage-based pricing good for customers? Largely no, and for the reasons the question assumed. The metered unit is loosely coupled to value, agents widen that gap rather than closing it, the model bills the most engaged users the most, and it relocates the entire burden of value accountability from the vendor onto the buyer. Good for investors and good for customers are not in alignment here. On this one, they partly trade off.</p><h1 class="wp-block-heading">Three things to do if you are buying.</h1><p>Model your power users, not your average. The average user will not break your budget. The fifteen people who actually adopted the thing will, and they may very well be the ones delivering your return.</p><p>Make the vendor define the unit before you sign. If a task can cost anywhere from a few credits to a few hundred depending on how many times the agent talks to itself, that is not a price, it is a range. And you'll end up at the upper end, trust me. Ask vendors to commit to a per-outcome cost and watch how fast the conversation gets vague.</p><p>Push for outcome terms on anything that has a definable outcome. Resolution, completion, ticket closed. If the vendor will only price the effort and not the result, they are telling you something about how confident they are in the result.</p><p>The interesting question is not whether usage pricing is here. It is. The question is whether buyers will accept a model where the vendor is paid the same whether the agent nails it on the first try or flails through ten, or whether the market pushes back toward paying for outcomes the way Fin does. Or Zendesk. Or Hubspot. Or others. I do not know which way that goes. But the vendor whose margin improves when its agent works harder is not, structurally, the vendor most motivated to make the agent work better.</p></div></div>
</div></div></div></div></div></div> ]]></content:encoded><pubDate>Sat, 20 Jun 2026 14:52:55 -0400</pubDate></item><item><title><![CDATA[Pega's fix for runaway AI costs: stop the agents from thinking at runtime]]></title><link>https://www.aheadcrm.co.nz/blogs/post/pegas-fix-for-runaway-ai-costs-stop-the-agents-from-thinking-at-runtime</link><description><![CDATA[The news At its PegaWorld conference in Las Vegas on June 8, 2026, Pegasystems announced Pega Infinity 26, which it says will be available in Q3 2026. ]]></description><content:encoded><![CDATA[<div class="zpcontent-container blogpost-container "><div data-element-id="elm_HrRQc_alQ96g_QqJuzyRnQ" data-element-type="section" class="zpsection "><style type="text/css"></style><div class="zpcontainer-fluid zpcontainer"><div data-element-id="elm_3kdzRqWNTbe_GPrWYK9FqQ" data-element-type="row" class="zprow zprow-container zpalign-items- zpjustify-content- " data-equal-column=""><style type="text/css"></style><div data-element-id="elm_nL5sYD0pQ3C-jmvJDvjsKA" data-element-type="column" class="zpelem-col zpcol-12 zpcol-md-12 zpcol-sm-12 zpalign-self- "><style type="text/css"></style><div data-element-id="elm_9PZYyHcUScSjjUL0AChrbw" data-element-type="text" class="zpelement zpelem-text "><style></style><div class="zptext zptext-align-center " data-editor="true"><div><h1 class="wp-block-heading">The news</h1><p>At its <a href="https://www.pega.com/events/pegaworld">PegaWorld</a> conference in Las Vegas on June 8, 2026, Pegasystems announced Pega Infinity 26, which it says will be available in Q3 2026. The principal change is commercial: <a href="https://www.pega.com/about/news/press-releases/pega-eliminates-ai-token-tax-more-efficient-way-build-and-run-agentic">Pega is moving away from per-token pricing</a> for its AI agents toward a flat charge per completed &quot;case,&quot; which it defines as a task carried out from start to finish, such as a customer changing an order, a loan approval, or a claim. Pega frames the move as removing what it calls the &quot;<em>AI token tax</em>&quot;.</p><p>The pricing change rests on an architecture Pega calls Predictable AI. Reasoning-heavy AI work is concentrated at design time, when workflows are authored in Pega Blueprint and the new Infinity Studio. At runtime, a lighter-weight model identifies the user's intent, selects a pre-approved workflow, and executes it step by step; where an individual step requires a language model, for example to parse a document or summarize a prior interaction, that step is given bounded instructions rather than open-ended latitude. Pega gives two reasons: more consistent outcomes, because agents follow approved workflows rather than re-reasoning each request, and more predictable cost, because the heavier processing happens only once during design rather than on every transaction.</p><p>The architecture is not new to this release. Pega introduced <a href="https://www.pega.com/about/news/press-releases/new-pega-predictable-ai-agents-combine-power-reasoning-predictability">Predictable AI Agents</a> in May 2025 and <a href="https://www.pega.com/insights/articles/introducing-pega-infinity-25-agentic-platform-enterprise-transformation">integrated them into Pega Infinity '25</a>, which reached general availability in December 2025. Infinity 26 primarily adds the outcomes-based pricing model, alongside a companion announcement that <a href="https://www.businesswire.com/news/home/20260608601073/en/Pega-Powers-AI-Agents-to-Reliably-Drive-Mission-Critical-Work">exposes Pega processes as Model Context Protocol (MCP) servers</a>, allowing third-party agents from Anthropic, OpenAI, Google, and AWS to call them under Pega's governance controls. The release cites no named customer, quotes analyst <a href="https://www.linkedin.com/in/lizkmiller/">Liz Miller of Constellation Research</a>. The &quot;more than 20x&quot; savings figure comes from Pega's AI Token Cost Calculator and is qualified as applying &quot;<em>depending on workflow complexity and scale</em>&quot;.</p><h1 class="wp-block-heading">The bigger picture</h1><p>Two industry currents explain the timing of this announcement.</p><p>The first is pricing. The customer-service software market has spent the past year and a half moving away from per-seat and per-token models toward charging for outcomes. Intercom Fin charges $0.99 per resolution. HubSpot cut its customer agent to $0.50 per resolved conversation in April. Zendesk runs around $1.50 per automated resolution on committed volume and has been selling outcome-based pricing since 2024. Salesforce launched Agentforce at $2.00 per conversation, a unit so loose that only roughly 8,000 of its 150,000-plus customers adopted it, which forced a pivot to per-action Flex Credits and Agentic Work Units. Sierra, Decagon, and Ada <a href="https://www.saastr.com/hubspot-switching-ai-pricing-from-per-use-to-per-resolution-but-does-it-really-matter/">all sell per-outcome</a> on custom enterprise contracts. Gartner, <a href="https://www.gartner.com/en/newsroom/press-releases/2026-03-25-gartner-predicts-that-by-2030-performing-inference-on-an-llm-with-1-trillion-parameters-will-cost-genai-providers-over-90-percent-less-than-in-2025">in a March 2026 forecast</a>, projects that the cost of running inference on a trillion-parameter model will fall more than 90% by 2030, while noting that those provider-side savings will not fully reach customers and that agentic models consume between 5 and 30 times more tokens per task than a standard chatbot. Not all of it will reach the buyers, though. The unit price of thinking is falling while the number of units per task climbs, which is the squeeze every vendor in this market is now pricing against. Pega's per-&quot;case&quot; charge belongs to this trend, with its unit defined differently from a customer-service &quot;resolution&quot;: a case spans a back-office task such as a loan approval or an insurance claim run end to end, rather than a single support interaction.</p><p>The second current is a deep disagreement across the industry about how much freedom an AI agent should have at runtime. One camp ships prompt-based tooling and lets agents reason and plan at each step, treating flexibility as the key point. Another constrains agents to pre-approved workflows and treats unbounded runtime reasoning as a liability, especially in regulated processes. Pega sits firmly in the second camp, <a href="https://diginomica.com/pegas-agentic-approach-puts-workflows-first-prompts-second-heres-why-matters-enterprise-ai-adoption">and its CEO has said publicly that competitors asking users to write prompts are setting themselves up for trouble</a>. The context underneath the argument is not trivial. A widely cited 2025 <a href="http://blog.aheadcrm.co.nz/2025/10/the-great-genai-divide-debunking-myth.html">MIT study from its NANDA initiative</a> found that roughly 95% of enterprise generative AI pilots produced no measurable return on the profit line, which the authors attributed less to model quality than to a &quot;learning gap&quot; in how organizations integrated the tools. This is the line the market is arguing about right now, and the vendors have started to pick sides.</p><h1 class="wp-block-heading">My point of view and analysis</h1><p>Start with the part Pega frames as leadership. On price, Pega is not leading, it is catching up, and the per-&quot;case&quot; charge is the same outcome-based move the customer-service vendors made first, just dressed for a different room. Credit where it is due, however, because the chosen unit is better than most: a completed back-office case is harder to game than a support &quot;resolution&quot; and maps to work a CFO already values. That is a real distinction. It is also a modest one, and it is not a first.</p><p>On the architecture, Pega's CEO is not entirely wrong about the risk he is arguing against. Letting a model improvise its way through a regulated claims process is asking for trouble, and the graveyard of failed genAI pilots is full of companies that could not audit what their agents did. The trouble is that the cure and the original promise of agentic AI pull in opposite directions.</p><p>Here is the question I cannot get my head around. There is real value in customer interactions that follow a rote path, and a great deal of work is exactly that; so Pega serving the rote case cheaply and consistently is a good thing, period. But the value of an agentic system was supposed to be the other case: the request that does not fit the workflow as designed, the genuinely novel situation. Pega's architecture is built to do the opposite of reasoning through those at runtime. So how does the system know it can safely run the rote workflow if it never reasons through the case at the outset? Pega's answer is the lightweight intent query that does the routing, which means the only runtime intelligence in the loop is intent classification, and classification is itself probabilistic and perfectly able to misroute. A request that matches no workflow then has three exits: forced onto the nearest approved path, escalated to a human, or handed to Blueprint to generate a workflow on the fly. However, that third option is the one Pega spends the whole pitch warning against, because runtime generation in a regulated process is precisely what it calls dangerous. You cannot headline determinism and keep on-the-fly generation as the safety valve without owning the contradiction.</p><p>There is a distinction underneath all of this. Deterministic guardrails wrapped around a probabilistic system set the boundaries of acceptable action without collapsing the space inside them. The agent still reasons; it simply cannot climb the fence. Pega is doing something else. At runtime, the approved space is the entire space. There is no reasoning inside the fence, because the fence is the answer. That is not an agent operating within guardrails. It is a workflow engine with a probabilistic front desk. For loan approvals and claims that may well be the right trade, and it should simply be named as one. The industry spent two years insisting agents would handle the unscripted long tail, and Pega's bet is that the long tail is where you get hurt, so it designed the long tail out. They may be right about the risk while conceding the promise without saying so. This is BPM, Pega's home turf since 1983, with an AI intake layer on the front. Calling it agentic is generous.</p><p>So here is what I would do before believing the deck. Ask Pega for one named production customer, on the record, who has run this at scale and watched the cost curve flatten, because a calculator output is not a reference you can phone. Then get the definition of a billable &quot;case&quot; in writing, including what happens when the workflow misroutes, fails, or escalates to a human, because &quot;resolution&quot; was always a vendor-defined word and &quot;case&quot; is no different, and that ambiguity surfaces on the invoice rather than in the contract. Finally, ask the uncomfortable one: what share of your real request volume does not map cleanly to a pre-approved workflow today, and what does Pega do with that slice? If the answer is &quot;a human takes it&quot; or &quot;Blueprint writes a new one live,&quot; you are buying a very capable workflow engine, which may be exactly what you need, as long as you buy it with your eyes open.</p><p>The token critique landed because it is true, and the architecture is sensible for the work Pega is aiming at. I am just not convinced the market asked for agents that are forbidden from thinking the moment a request gets interesting, and I would like to know whether buyers are actually asking for this or whether the industry has decided the long tail was a bad idea all along.</p></div></div>
</div></div></div></div></div></div> ]]></content:encoded><pubDate>Sat, 13 Jun 2026 11:55:10 -0400</pubDate></item><item><title><![CDATA[The Illusion of Value: Why Salesforce’s Agentic Work Unit is the New &quot;Bad Query&quot; of the AI Era]]></title><link>https://www.aheadcrm.co.nz/blogs/post/the-illusion-of-value-why-salesforces-agentic-work-unit-is-the-new-bad-query-of-the-ai-era</link><description><![CDATA[The News On February. 25, 2026, Salesforce announced a pricing and metrics update . During the company’s Q4 FY2026 earnings call, CEO Marc Benio ff, toge ]]></description><content:encoded><![CDATA[<div class="zpcontent-container blogpost-container "><div data-element-id="elm_hvMRJzNwQUSSARB5L90iuQ" data-element-type="section" class="zpsection "><style type="text/css"></style><div class="zpcontainer-fluid zpcontainer"><div data-element-id="elm_oG6IdhkHRval9EYuUjZDAw" data-element-type="row" class="zprow zprow-container zpalign-items- zpjustify-content- " data-equal-column=""><style type="text/css"></style><div data-element-id="elm_QjSd8e74QwSgDT50CsOPQg" data-element-type="column" class="zpelem-col zpcol-12 zpcol-md-12 zpcol-sm-12 zpalign-self- "><style type="text/css"></style><div data-element-id="elm_dqnQmre_QWiwssBRMKxp3g" data-element-type="text" class="zpelement zpelem-text "><style></style><div class="zptext zptext-align-center " data-editor="true"><div><h1 class="wp-block-heading">The News</h1><p>On February. 25, 2026, Salesforce announced <a href="https://www.salesforce.com/news/stories/agentic-work-units/">a pricing and metrics update</a>. During the company’s Q4 FY2026 earnings call, CEO <a href="https://www.linkedin.com/in/marcbenioff/">Marc Benio</a>ff, together with CMO <a href="https://www.linkedin.com/in/patricks/">Patrick Stokes</a>, unveiled the <em>Agentic Work Unit</em> (AWU). Positioned as a metric to quantify the labor performed by autonomous digital systems, Salesforce defines an AWU as one discrete task accomplished by an AI agent.</p><p>According to Salesforce, this discrete task represents the exact moment &quot;<em>raw intelligence is converted into real work</em>&quot;. It is not a fixed unit but measured as a processed prompt, a completed reasoning chain, or an invoked tool. Salesforce explicitly designed the AWU to move the industry conversation away from the raw consumption of Large Language Model (LLM) tokens. As Benioff noted, tokens only measure &quot;how much an AI talks,&quot; whereas the AWU is intended to measure actual business execution.</p><p>The scale of this rollout is massive. Salesforce reported that its platform has already processed over 19 trillion AI tokens, translating them into 2.4 billion Agentic Work Units, with 771 million AWUs delivered in the fourth quarter alone. This new metric serves as the underlying foundation for Salesforce's evolving Agentforce monetization strategy.</p><h1 class="wp-block-heading">The bigger picture</h1><p>Following a nearly 18-month period of <a href="https://www.saastr.com/salesforce-now-has-3-pricing-models-for-agentforce-and-maybe-right-now-thats-the-way-to-do-it/">pricing triangulation</a>, which included a $2.00 per conversation model and a $0.10 per action &quot;Flex Credit&quot; model, Salesforce is leveraging the AWU to track system utilization, even as it wraps enterprise purchasing in familiar, unmetered per-user license agreements starting at $125 per user per month.&nbsp;&nbsp;</p><p>To understand the significance of the Agentic Work Unit, one must view it through the lens of a broader industry crisis: the so-called &quot;SaaSpocalypse&quot; and the looming threat of the seat cannibalization trap. For two decades, the Software-as-a-Service business model has been dominated by seat-based licensing. However, as agentic AI systems mature and promise to be capable of executing multi-step workflows autonomously, they inherently reduce the need for human software operators. If an AI agent resolves 84% of tier-one support tickets without human intervention, the enterprise requires fewer human support seats. I have repeatedly written about pricing models, e.g., <a href="https://customerthink.com/which-ai-pricing-models-work-best-for-customers/">here, as part of my CustomerThink column</a>.</p><p>This dynamic has forced the software industry into a frantic transition toward usage-based and attempts at outcome-based pricing models. Usage-based pricing, popularized by cloud infrastructure providers like AWS and data platforms like Snowflake, charges customers based on system consumption, e.g., compute seconds or data processed, or, these days, tokens consumed. While this protects the vendor's margins and aligns with their variable cloud GPU costs, it shifts the financial risk of system inefficiency entirely onto the buyer. The vendors essentially play the role of <a href="https://en.wikipedia.org/wiki/Pontius_Pilate">Pontius Pilate</a> and wash their hands in innocence.</p><p>Conversely, agile AI disruptors and customer service incumbents are aggressively pioneering true outcome-based pricing, where the billable event is delayed until a verified business success is achieved. For instance, <a href="https://www.intercom.com/help/en/articles/9061614-intercom-plans-explained">Intercom's Fin AI agent</a> charges a strict $0.99 per successful resolution, while <a href="https://support.zendesk.com/hc/en-us/articles/6931689272090-Moving-to-automated-resolutions-from-existing-bot-pricing-plans#topic_e4x_z1s_y1c">Zendesk recently started a $1.50 per automated resolution model</a>. In these models, if the AI fails to resolve the customer's issue, the customer pays nothing. Correspondingly, the vendor has skin in the game and needs to be interested in its software actually delivering value.</p><p>Salesforce’s introduction of the AWU represents a kind of a middle ground. Industry analysts like Constellation Research’s <a href="https://www.linkedin.com/in/lizkmiller/">Liz Miller</a> observe that the AWU acts as a <a href="https://www.cio.com/article/4138622/awu-by-salesforce-a-shiny-new-metric-that-tells-cios-little-of-value.html">placeholder for the agentic era</a>, much like clicks and likes functioned in the early days of online and social media. It is still a usage-based consumption metric masquerading as an outcome metric. The industry is currently witnessing a tug-of-war: legacy giants are deploying metrics like the AWU to track utilization and justify high enterprise license costs, while pure-play AI vendors intend to leverage outcome-based pricing as a competitive weapon to steal market share by guaranteeing and demonstrating return on investment.</p><h1 class="wp-block-heading">My point of view and analysis</h1><p>As someone who has spent years helping organizations unlock their potential through digital transformation initiatives, I look at the Agentic Work Unit highly skeptical. When evaluating generative and agentic AI investments, the critical measure is the ability to deliver measurable business results, not just technological activity. In this context, the AWU represents a fundamental conflation: it equates doing work with achieving outcomes, which simply is not true.</p><p>By defining an AWU as a discrete task, such as invoking an API or triggering a workflow, Salesforce has created a metric that measures machine exertion rather than enterprise value. In the realm of autonomous systems, an AI agent can execute thousands of discrete tasks, burn through immense computational resources, and work incredibly hard while achieving absolutely nothing of commercial consequence.</p><p>Working hard on the wrong thing still doesn’t deliver results.</p><p>To fully grasp why measuring discrete AI tasks is a poor proxy for value, consider the analogy with an unoptimized database query vs. an optimized one in cloud data warehouses. A highly optimized SQL query returns a vital dataset in seconds for pennies. Conversely, a poorly written query forces the database engine into massive data scans. It might be running for hours and consuming plenty of CPU and memory resources. From the vendor's billing perspective, the system successfully performed the discrete scanning tasks it was instructed to execute. However, it results in a massive consumption bill for the customer. The business gains little value, as much of it is harvested by the vendor; even worse, if the result is wrong. Yet the financial penalty is severe.&nbsp;&nbsp;</p><p>The Agentic Work Unit operates exactly on this flawed economic principle. Autonomous AI agents are still highly susceptible to unique failure modes, e.g., <a href="https://arxiv.org/html/2502.19918v2">the infinite reasoning loop</a>. If an agent encounters an ambiguous prompt or lacks solid memory tracking, it may repeatedly call the same tool or query the same database in an endless cycle due to perfection bias. While engineers desperately build so-called Meta-Reasoners to halt this wasted computation, the AWU metric actively monetizes it. If a confused agent loops fifty times before timing out, it has successfully generated fifty AWUs delivering zero result. The customer is actively billed for the machine's confusion.</p><p>Furthermore, agentic workflows suffer from compounding hallucinations, or <a href="https://failingfast.io/autocomplete-was-never-the-point/#autocomplete---intellisense-on-crack">epistemic debt</a>. If an agent hallucinates a false premise in step one, it will still confidently execute subsequent tools based on that fabrication. By the time the workflow concludes, the agent may have triggered dozens of AWUs across multiple enterprise systems, corrupting data and requiring costly human remediation.</p><p>Ultimately, a dashboard celebrating 2.4 billion AWUs gives the illusion of massive productivity, but it is a vanity metric. If those tasks were merely redundant internal data reshuffling or failed reasoning loops, the actual profit multiplier of the organization remains unchanged. An Agentic Work Unit quantifies motion, but motion is not progress. Until AI pricing models mature to align the cost of digital labor with the verified delivery of business outcomes, enterprises must continue to treat effort-based metrics like the AWU with extreme skepticism.</p><p>Just my $.02. What do you think?</p></div></div>
</div></div></div></div></div></div> ]]></content:encoded><pubDate>Sun, 01 Mar 2026 16:16:16 -0500</pubDate></item></channel></rss>