/ Insights / View Recording: From Pilot to Portfolio: What Scaling AI Actually Costs Insights View Recording: From Pilot to Portfolio: What Scaling AI Actually Costs September 3, 2026 From Pilot to Portfolio: What Scaling AI Actually Costs As organizations move beyond isolated AI pilots and into managing portfolios of agents, the conversation inevitably shifts to economics. What does it cost to scale, govern, and sustain AI at enterprise levels? In this session, we’ll examine how leading organizations are approaching the transition from licenses to consumption-based models, including budgeting strategies, cost-per-outcome metrics, and the trade-offs executives are making along the way. You’ll leave with practical benchmarks and planning insights to help shape your own AI investment strategy. Most organizations have proven that AI can work. A pilot ran, results looked promising, and leadership approved the next step. Then the invoice arrives from production and nobody can explain the gap between what the prototype suggested and what the system actually costs to operate. The problem is not that AI got more expensive. The problem is that pilot economics are protected economics, and treating a prototype as a forecast produces a number that looks precise and is fiction. The gap compounds quietly. Pilots run on cleaner inputs than the production mix, with cooperative users who know what they are testing and low or irregular volume that never lets the expensive paths compound. A delivery team quietly rescues failures by hand, and that expert labor is never booked as cost. There are no service level agreements, no support rotation, and no incident process. When the workload moves into production, every one of those protections disappears at once, and the commercial bill fragments at the same time: seat licenses, included entitlements, pooled credits, pay-as-you-go meters, product-specific reservations, and broader pre-purchase commitments now all appear on the same workload. Teams double count by adding consumption on top of entitlements they already own, or undercount by assuming the included grants will never run out. Brian Haydin, Solutions Architect at Concurrency, works through what actually changes as an AI workload moves from a hand-built pilot to a production agent to a governed portfolio. The session uses a factory analogy throughout: proving one unit can be made is a different discipline from operating a system that produces acceptable goods repeatedly. Haydin covers why the commercial bill fragments across Microsoft 365 Copilot, Copilot Studio, and Microsoft Foundry, how Copilot credits became a common currency across those capabilities, how to build a fully loaded cost model for an accepted business outcome, the two-part workload scorecard, the difference between a completed run and an accepted one, and how funding posture should shift as demand moves from uncertain to durable. Rather than comparing model price sheets or handing out a per-agent price range, this session demonstrates how to build the cost denominator itself. Attendees learn why a vendor meter is not a cost to serve, how to assemble the five cost layers of build, run, govern, adopt, and recover, how to allocate shared platform costs fairly instead of charging them in full or treating them as free, why a named business owner is a precondition for scaling, how to set an allowable cost ceiling before work begins, and why the median and the 95th percentile must be read together. The key takeaway is clear: The token bill is the visible lever, not the dominant one. In a worked example, halving AI and tool runtime saved $1,500 a month, while reducing the human review rate from 25 percent to 15 percent saved $6,250 while preserving the same 8,500 accepted outcomes at the same quality and risk threshold. Organizations that measure cost per accepted outcome rather than cost per token can defend a scale decision with evidence, spot when a team is optimizing easy cases while exceptions grow more expensive, and decide at each gate whether a workload should scale, tune, reroute, constrain, consolidate, or retire. WHAT YOU’LL LEARN Why Pilot Economics Mislead The five protections a pilot enjoys: cleaner inputs, cooperative users, low volume, manual rescue by the delivery team, and no service obligation Why expert labor invested during the pilot becomes an invisible cost later Why a pilot is allowed to end after producing learning, without moving to production The Three Stages and Their Economic Questions Pilot asks whether it can work, and acceptable spend is the cost of reducing uncertainty Production asks whether it can work reliably, tying cost to throughput, quality, service levels, exceptions, and support Portfolio asks which workloads earned the right to scale, introducing allocation, reuse, commitments, and opportunity cost How the Commercial Bill Fragments Combining seat licenses, included entitlements, pooled credits, pay-as-you-go meters, reservations, and pre-purchase commitments on a single workload How Copilot credits became a common currency across Microsoft 365 Copilot, Copilot Studio, and Microsoft Foundry The double-count and undercount errors teams make when entitlements and consumption are read as one picture Matching Commercial Posture to the Demand Curve Keep structure flexible during exploration, when the spend is buying information Attach meters and attribute consumption once demand becomes observable, to learn the baseline and the spikes Three common mistakes: committing before the workload is proven, staying fully pay-as-you-go after a stable base is obvious, and buying for the average while ignoring peaks Vendor Meter Versus Cost to Serve The vendor bill describes how the supplier charges, in tokens, credits, provisioned throughput, and search operations The business question is what it cost to produce the result, including integration, human review, exceptions, and evaluation Microsoft Foundry cost guidance states explicitly that its charges are a subset of total solution cost Building the Fully Loaded Model Five cost layers: build, run, govern, change management and adoption, and recovery from rejected outcomes and incidents Fair allocation of shared platform capabilities such as identity, networking, data access, retrieval, telemetry, and security Shared does not mean generic, and centralized cost does not mean free Completed Is Not Accepted An agent can finish every step and still miss the quality threshold, violate policy, or fail to complete automatically Attempts measure demand placed on the system, accepted outcomes measure throughput that cleared the business gate, and the difference is failures, rework, and escalation The business owns the outcome definition and the cost ceiling, engineering owns observability, tracing, escalation, and latency guardrails The Scorecard and the Decision Part one names the work: business owner, accepted outcome, attempts, accepted outcomes, and an allowable cost ceiling set before launch Part two explains the cost: acceptance rate, straight-through rate, median cost per accepted outcome, and the 95th percentile tail Every gate ends in a decision to scale, tune, reroute, constrain, consolidate, or retire FREQUENTLY ASKED QUESTIONS Is this session technical or business-focused? Both. The session is built around the handoff between finance and engineering, and it argues that the metrics break apart when either side defines acceptance alone. Business leaders get a funding and scorecard framework. Engineering leaders get the telemetry and testing discipline required to produce the numbers. Who should attend? Anyone accountable for the cost or the outcome of an AI workload: technology and finance leaders, FinOps practitioners, architects, product owners, and the named business owners who must decide whether a given agent scales or retires. Do I need experience with AI agents or FinOps? No. The session uses a factory analogy to explain the difference between proving one unit can be made and operating a system that produces acceptable results repeatedly, so the concepts stand on their own. What is an accepted outcome? An accepted outcome is a completed run that also cleared the business quality gate: the case closed, the shipping discrepancy was resolved, the payment was applied. A run that reaches the end of its steps but misses the threshold, breaks policy, or requires manual rescue is completed, not accepted. Why does AI cost control matter right now? Because the way AI is purchased changed. Licenses did not disappear, they were joined by entitlements, pooled credits, consumption meters, reservations, and pre-purchase commitments, so a single workload can now draw on all of them at once. What is the most important takeaway from this session? That the human review and exception path, not the token meter, is usually the larger economic lever. In the worked example, reducing the review rate was four times more valuable than halving AI and tool runtime, and it preserved the same accepted outcomes at the same quality. ABOUT THE SPEAKER Brian Haydin, Solutions Architect, helps organizations move AI workloads from prototype to governed production at a cost they can defend. His focus areas include agent operations, workload scorecards, fully loaded unit economics, and the funding decisions that determine which agents scale and which are retired. TRANSCIPT Transcription Collapsed Transcription Expanded Brian Haydin Hey everybody, good morning. Brian Hayden, for those of you who haven’t seen one of my webinars before or have met me. I’m a solution architect here at Concurrency, and I know it looks like it’s night time. Sometimes I do these webinars from Europe. I’m actually in Milwaukee. 0:0:34.12 –> 0:0:56.892 Brian Haydin And hopefully the thunderstorm behind us doesn’t distract me too much today. Today’s topic, you know, I would start by saying that the cost of, you know, AI in this day, it’s changing dramatically, right? But it’s never really just a finance question. And it’s never really just a technical question. It’s kind of a handoff in between both of those. 0:0:57.372 –> 0:1:18.572 Brian Haydin And so the title here today is From Pilot to Portfolio and talking about what scaling AI actually costs. And the visual spine is going to be pretty simple. It’s a pilot and it’s a, you know, hand-built prototype is what we’re going to be talking about. And then it moves into a production agent and we’re going to use that as kind of like a repeatable 0:1:19.12 –> 0:1:40.972 Brian Haydin sort of factory. You know, think of a portfolio as a plant with several, like a factory with several lines, you know, that are sharing infrastructure, they’re sharing people, controls, you know, and capital as well. And we’re going to use that idea, you know, throughout the seminar, because, you know, not because every company here is a manufacturer, but because everybody understands the difference between 0:1:41.372 –> 0:1:51.452 Brian Haydin proving that one unit can be made and then operating a system that can make goods, you know, final goods, like repeatedly over and over again. 0:1:52.412 –> 0:2:15.452 Brian Haydin So before we talk, just a little bit about CONSULTANCY. We are a 35 year old, 35 plus year old consulting company, solution integration company. And, you know, we do a lot of these talks. Obviously, you’re here today. Search for us on LinkedIn, engage with us. You’ll find out about more of our LinkedIn 0:2:15.852 –> 0:2:34.502 Brian Haydin activities, you know, and there’s a QR code or just search for Concurrency Inc. on LinkedIn. But I want to talk a little bit about three things that we’re going to do today. First, I want to talk about why the bill changes when an AI workload moves from a 0:2:34.572 –> 0:2:54.172 Brian Haydin protected Pilot into production, and then it moves into a portfolio, right? And then second, we’re going to give you a reusable way to calculate the fully loaded cost of an accepted business outcome. Not a token, just a request, but a completed run. And then third, I want to talk about how the funding posture should change. 0:2:55.292 –> 0:3:14.12 Brian Haydin you know, as the demand becomes a little bit more certain. So we’re not going to spend, you know, this time comparing model price sheets. I don’t think that’s, you know, that’s going to be helpful. Those prices, they move. And the contractual terms, you know, are different for everybody. And then, you know, the other, you know, the other idea is that a cheap meter 0:3:14.292 –> 0:3:34.252 Brian Haydin if you’re, you know, if you look at it that way, can still produce expensive work. I’ll unpack that a little bit later. I’m probably going to be talking quite a bit about Microsoft, you know, as the examples, because that’s where concurrency lives. But I would just acknowledge that the things that we’re talking about today, they’re not Microsoft specific, you know, per se, they’re portable. 0:3:34.892 –> 0:3:54.652 Brian Haydin And you’ll be able to use these ideas, whether you’re building in the Microsoft ecosystem or not. So, you know, let’s see, what does an AI agent cost? And let’s start behind, let’s start with that question. And I, you know, if you came here thinking that I was going to give you a range and a number. 0:3:55.52 –> 0:4:14.732 Brian Haydin of what a lightweight agent costs and what an enterprise agent costs and what a multi-agent system costs. I’m sorry to disappoint you, you’re not going to get that number, because the real answer is that there is no responsible universal number. A service, like a service assistant, like a coding workflow, 0:4:15.692 –> 0:4:34.412 Brian Haydin and agents that are authorized to execute financial transactions, they do not share the same volume. They don’t share the same quality bar or the risks, human oversight. And so they don’t have the same commercial model. And if you put one price on all three of those things, the number becomes like 0:4:34.652 –> 0:4:53.412 Brian Haydin It might look precise, but it’s fiction. So a practical test would be, you know, whether whenever someone gives you like a per agent price, like ask them back, like what is what, ask what workflow, what volume. 0:4:53.492 –> 0:5:18.132 Brian Haydin What is the acceptance bar that that assumes? What’s the human oversight that, you know, is included in this? And then, you know, finally, what the operating boundary is that that number assumes. If any of those answers are missing, then I would say that that range is, you know, it’s kind of irrelevant what they said. But that doesn’t mean that we can’t make these economics understandable or unknowable. It means that we have to change 0:5:18.212 –> 0:5:22.172 Brian Haydin The frame of reference of what that unit means, O. 0:5:24.572 –> 0:5:43.452 Brian Haydin Before we build the method, I want to show you a little bit about what the results might look like. These 2 changes are applied to the same hypothetical synthetic workflow. 10,000 monthly attempts, it’s got the same number of accepted outcomes, and let’s just assume that there’s the same quality thresholds. 0:5:43.812 –> 0:6:3.92 Brian Haydin I make one change and it saves $1,500 a month. And the other change could potentially save $6,250. And so the question might be like, which lever, you know, do you think is doing which? And don’t overthink anything. I’m not defining what those levers are yet. 0:6:3.172 –> 0:6:21.532 Brian Haydin But, you know, I’m just saying like, could you imagine like, you know, what, which one would be which? And, you know, most cost conversations start with like the visible, you know, meters, like the dollar amounts. And most people initially choose the lever that seems to make that meter go down the most. 0:6:22.92 –> 0:6:41.212 Brian Haydin you know, and you know, what it reveals is that the workflow carries some sort of economic mass that we have to consider. And so a lot of times we withhold the labels, you know, that, you know, for the, for what these levers are doing, you know, and we don’t really discuss what those outcomes are, you know, 0:6:41.732 –> 0:6:58.972 Brian Haydin And so what I want to do is I want to I want these numbers to actually earn their standing in this. And so when we get to the end, this is going to make a little bit more sense, you know, but I just wanted to start with this is where we’re going today is to look at what these levers are going to do and cause, you know, and we’ll talk about the changes at the end of it. 0:7:1.132 –> 0:7:19.612 Brian Haydin So a pilot is a hand-built prototype. Its job, you know, is to answer basically one question when you’re doing a pilot. Can we make this thing work well enough to justify learning more? I think that’s an important way to look at what, you know, what a prototype is or what a pilot is. 0:7:19.972 –> 0:7:41.212 Brian Haydin because learning really is the goal. And we’re going to tolerate conditions that would be irrational if we were to put this into production and try to scale it out. We may use the strongest model at the very beginning for every one of the requests. We might curate the test cases. The project team might do some manual rescue failure, you know, rescues for the failures that happen. 0:7:42.172 –> 0:8:0.332 Brian Haydin You get the domain expert involved and that’s going to answer questions in real time. You know, all that kind of stuff, right? And that’s not a bad practice. It’s the rational way that we reduce the uncertainty and the unpredictability of building new systems. Now on the production side, when we want to get there, that’s a different system. 0:8:0.732 –> 0:8:24.12 Brian Haydin We need repeatable throughput, we need inspections, we need exception handling, we need agent ops, security, support, and most importantly, we need to have a defined owner. The portfolio, that adds like another layer on top of that. So when we get like past that production, we start building a portfolio of multi-agent systems. 0:8:25.772 –> 0:8:48.172 Brian Haydin We have several lines competing, you know, for shared people, shared capital. They’re reusing some of the common infrastructure, so they’re sharing some of those components as well. And then we need decisions about which line, you know, is going to expand, which ones change, you know, what combines with another, or, you know, I’ve talked about this in my other talks, like what agents are we going to shut down? 0:8:48.812 –> 0:9:12.852 Brian Haydin So factories, they also carry, they also care about things like yield and capacity and changeover and inspection, downtime, scraps. Those concepts, I think they, you know, for a lot of us that are on this call, probably working in manufacturing, but they map pretty cleanly to like, you know, the same things that agents do and that we look at when we’re building agents, acceptable outcomes. What is the 0:9:12.932 –> 0:9:32.172 Brian Haydin throughput of the work. What happens when things like models and workflow changes? How much of it goes through like a human review? What do we do with the incidents and how do we manage reworking some of the processes? So, you know, I think that the mistake really isn’t, you know, it’s not building a cheap. 0:9:32.212 –> 0:9:52.252 Brian Haydin pilot, the mistake that people make sometimes is treating the prototype as like a forecast of what we’re going to be doing. So pilot economics are protected economics. And 5 protections show up pretty repeatedly in the conversations that I have. 0:9:52.732 –> 0:10:11.852 Brian Haydin The inputs are cleaner than, you know, than the production mix. The users are pretty cooperative and they know, you know, they know what and how and you know what they’re testing. Volume is usually pretty low. Sometimes it’s irregular. And so the expensive paths that an agent might go through 0:10:12.572 –> 0:10:30.812 Brian Haydin They’re not compounding on each other. You’re not seeing a lot of effects on that. A delivery team provides manual rescue. And we never really book that as like some of the cost that goes into this. And there’s no real service obligation, you know, that’s been set up. No SLAs, no support rotation, no incident process. 0:10:31.172 –> 0:10:50.652 Brian Haydin things along that nature. So I think that, you know, there’s still like kind of this promise that the result is still going to be reliable, you know, after the Pilot, you know, when we just roll it in, the model, the data, the policies, none of that changes, right? So, you know, I would also say like, 0:10:50.692 –> 0:11:9.52 Brian Haydin The, like, the Pilot does also consume infrastructure. It still has, like, you know, some of the identity stuff, the search storage, those things, you know, those capabilities, you know, get built into it as well, you know, and we need to be, you know, just sort of mindful of what those resources are costing as well. 0:11:10.252 –> 0:11:28.572 Brian Haydin I would also say like, there’s one kind of like warning sign that tends to creep up in these conversations is that like the expert label that labor that exists in practice, you know, like when we’re building it, it disappears like later, right? So the people that are investing their time 0:11:28.892 –> 0:11:48.92 Brian Haydin to make sure that the Pilot is successful, they become kind of like a transparent sort of cost or an invisible cost, I would say, if you want to think of it that way. And so those are some of the protections, you know, that I, you know, wanted to call out. Now, 0:11:49.932 –> 0:12:10.492 Brian Haydin As you know, as we mature a little bit, each stage that we’re going to go through is going to ask a different economic question. At the beginning, when we’re doing the pilot, we started off with, can it work? The acceptable spend, you know, is the cost of, you know, the spend that we’re doing, what’s acceptable here is the cost of reducing 0:12:10.852 –> 0:12:31.292 Brian Haydin that uncertainty to making sure that it works. But in production, the shift here is can it work, is can it work reliably? Not just can it work, but can it work reliably? So the cost becomes tied to throughput, you know, becomes tied to quality, service levels, you know, the exceptions and the support. And, you know, 0:12:32.172 –> 0:12:52.332 Brian Haydin just measuring the technically like a completed run, it’s not enough. You need to look at, you know, the reliability of it as well. And then on the Portfolio side, that question shifts to which workloads earned the right to actually scale. And that’s where things like allocation and reuse and commitments, opportunity cost, 0:12:53.52 –> 0:13:12.12 Brian Haydin they start to enter the conversation. So every dollar given to one workload is a dollar that’s not given to another at this point. So looking at this, you know, from left to right, the progression is capability, reliability, and, you know, and economic control. And so 0:13:12.92 –> 0:13:32.92 Brian Haydin That’s more or less what the storyline is for today. First, I want to show how the commercial bill fragments. And then I want to move us from technical consumption to the cost of the accepted work. And that’s an important distinction, not the cost of, you know, running it, but the cost of the accepted work, what actually completed. 0:13:32.612 –> 0:13:50.892 Brian Haydin And then we’re going to return to the Portfolio. And I wanted to decide, you know, help you decide how uncertainty, proven demand, and ongoing change should be funded as well. And you know, when you go through this, there’s stops, right? There’s gaps, you know, control points. So those conditions change too. 0:13:51.452 –> 0:14:9.612 Brian Haydin You know, and what I mean by that is that a Pilot can end, you know, can end after it creates some learning, and that’s it. It doesn’t have to move into production, and a production workload could pause when quality or controls start to fall outside of the boundaries that you set at the beginning of it. 0:14:10.212 –> 0:14:30.812 Brian Haydin And then a portfolio workload should lose funding, you know, or could lose funding when another approach creates more accepted work, more, you know, more work that actually gets the job done for the same amount of money or the same constrained dollar. So let’s take a look at the 1st place budgets, you know, become confused. And that’s the way that 0:14:31.52 –> 0:14:51.92 Brian Haydin AI is getting purchased today. So we’ve had some webinars, you know, talking about this in pretty good detail, but, you know, I described a transition from licenses to consumption-based models when I wrote up what we would talk about today. And directionally, that’s still valid. 0:14:51.852 –> 0:15:13.212 Brian Haydin But the normalized enterprise reality is a little bit messier today. And, you know, it’s also a little bit more important. The licenses didn’t disappear. People are still paying that. But the bill, it’s become a little bit more fragmented. You now have to combine things like seat licenses with the included entitlements. How many credits are you getting allocated? 0:15:13.532 –> 0:15:36.412 Brian Haydin as you know, the pooled credits. And then there’s also the pay as you go meteor, you know, that gets added to that. Product specific reservations might come into play. And then there’s like this idea of doing broader pre-purchase kind of commitments. So the Portfolio that that you’re building might contain some of them, or it might contain all of them. 0:15:36.892 –> 0:15:55.612 Brian Haydin And Microsoft gave us, like, a really good example this summer: the copilot Studio, you know, story around copilot credits, and that kind of became this common currency that we’re using across different kinds of capabilities, right? copilot Cowork, copilot Studio, you know, and then… 0:15:55.972 –> 0:16:12.972 Brian Haydin That also incorporates this whole idea of like the pay as you go and the pre-purchase, you know, options. So you can, you know, with cowork and 365 cowork, just pay as you go, run the meter, or, you know, that can affect like copilot credits as well, right? So. 0:16:14.812 –> 0:16:33.292 Brian Haydin And then, so that’s Microsoft 365 copilot, copilot Studio, and then Microsoft Foundry has its own model, you know, for, you know, consumptive based. And then you get this like add-on here where you can do the predictable demand and pre-purchasing of, you know, the credits as well. 0:16:33.732 –> 0:16:54.892 Brian Haydin and that can reduce some of the cost. So try to stitch that all together and that fragmentation of the bill becomes pretty apparent, right? And teams can double count things by treating included entitlement, you know, and the consumptive cost that goes on top of that, you know, as like the full picture. 0:16:55.52 –> 0:17:13.932 Brian Haydin Like are they they don’t count the free stuff and then they count the, you know, the the consumption on it on top of that as well. Or maybe on the on the flip side of it, they undercount, you know, some of the credits or some of the the operating expense of this by assuming that we’re never going to run out of like the. 0:17:14.892 –> 0:17:32.972 Brian Haydin the included entitlement grants, because we just haven’t done that before. So they can, you know, kind of miss the boat on that as well. So I just want to call out these different, you know, these different ways that you can, you know, that cost can surface, you know, because, and then also that it’s going to change, right? 0:17:33.452 –> 0:17:40.92 Brian Haydin because like it just hasn’t been super consistent up until this point. But so. 0:17:41.172 –> 0:18:0.492 Brian Haydin With all that kind of change, I think it’s important to think about that the commercial posture, you know, should follow this demand curve as well. So going back to like the Pilot, during that exploration phase, you want to keep that commercial structure flexible. 0:18:1.52 –> 0:18:22.892 Brian Haydin you’re really paying for information, right? So I like to think of it like when I, when I’m doing like really deep research, you know, in ChatGPT that, you know, I’m sending it out in the wild, do as much damage as you can to my consumption because I really need to find this information. So it’s going to be a little bit more expensive in that ideation phase. 0:18:23.692 –> 0:18:47.972 Brian Haydin But that, you know, the caution here is like, don’t use that as a long-term commitment, you know, because it’s, it isn’t like, it isn’t the right projection. As the demand, as you start to run these pilots and the demand becomes a little bit more observable, that you can like attach a meter to how much is happening and attribute it to the workloads that are actually requested. 0:18:48.92 –> 0:19:7.132 Brian Haydin testing it. That helps you to learn the base, you know, the baseline, and then observe some of the spikes that might come in, you know, a little bit later. And then also like helping you measure what the expensive paths of making some of those decisions, automations, or… 0:19:7.612 –> 0:19:27.612 Brian Haydin you know, activities would be. So that’s pretty helpful. And then, you know, only after you have a workload, you know, sort of built out, you know, and that diversified Portfolio, you’re going to be able to show durable demand, and then you can evaluate, you know, a longer term commitment. 0:19:28.92 –> 0:19:48.412 Brian Haydin So, I would, you know, commit against a stable base, you know, once you have that, and not some sort of optimistic, you know, adoption curve that you think, you know, think might exist when you’re just in the middle of doing the Pilot. You know, what’s going to wind up happening there is that you’re going to have a lot of… 0:19:49.52 –> 0:19:53.852 Brian Haydin discounted commitment that, you know, just winds up getting burned, you know, and flushed on the toilet. 0:19:56.92 –> 0:20:15.292 Brian Haydin So 3 mistakes I think are, you know, are pretty common. Committing before the workload is proven, staying entirely pay as you go after a stable base is, you know, has become obvious, and then buying only for the average while ignoring like the actual peaks, you know, that are happening. 0:20:15.772 –> 0:20:35.52 Brian Haydin So the right answer is probably going to combine commitment for the base with variable capacity and then baking in some, you know, some of the uncertainty. So here’s a distinction I think that prevents most bad estimates. 0:20:35.452 –> 0:20:59.132 Brian Haydin So A vendor meter is not your cost, you know, to serve. What the vendor bill tells you how, like the vendor bill is telling you how the supplier is charging you. It shows like the input and output tokens, copilot credits, things like provision throughput, search operations, that type of stuff, right? 0:20:59.572 –> 0:21:20.972 Brian Haydin And those numbers, they they matter, but they are, they’re the cash cash costs, but like the engineering team is going to need to understand that so they can optimize some of those workloads, you know, to control those kind of costs, but the business is probably going to be asking a different question. 0:21:21.412 –> 0:21:41.532 Brian Haydin what did it cost us to produce that actual result? They don’t really care about the token cost or like how many credits were required for each round trip. What they want to know is like when I got done with this unit of work, how much did that actually cost me? And so that includes everything, you know, that the platform has, the model. 0:21:41.932 –> 0:22:2.372 Brian Haydin But it also includes things like integration with other systems. It includes the time for human review and, you know, exceptions, evaluations, you know, all that kind of stuff, right? So that’s like, you know, that’s the all-in cost that you need to look at. And Microsoft’s Foundry 0:22:2.452 –> 0:22:21.972 Brian Haydin Cost guidance has some, you know, pretty explicit, you know, foundry charges that you should think about. And they’re only like, they’re only that subset of the overall application of the solution cost. I mean, they call that out pretty directly. And I think that most of us need to start thinking about like, 0:22:22.52 –> 0:22:42.332 Brian Haydin the, you know, the other costs. And then when we, like, when we get done with the Pilot, we should, you know, think about what those additional costs are going to be. So that, that’s going to help us, like, project what something is going to cost to operate in the long term. 0:22:42.892 –> 0:23:3.612 Brian Haydin And so, you know, don’t confuse just those like supplier meters, you know, with the overall spend. I’m trying to think of like this dashboard that one of my customers has, and it does this like, does this like, you know, respond to an e-mail inbox on a question. 0:23:4.132 –> 0:23:24.972 Brian Haydin And we had calculated that every time an e-mail comes into the inbox, it cost a average of $6, let’s just call it $6 to respond to that cost. Now that was just basically a person, their labor time to like answer the e-mail, go into the ERP system, make the change, and be done with it. 0:23:26.12 –> 0:23:44.252 Brian Haydin And then we measured that against like the all-in cost of each one of the requests and it was like a dollar, you know, so definitely like a definitely savings, but that was across an entire platform. You know, we included like things like how many requests per month versus what the licensing cost, what the token cost was or the. 0:23:44.492 –> 0:24:3.692 Brian Haydin the copilot credits. And then you can kind of compare and contrast, like, you know, how much does it cost to do it for a person? How much does it cost to do it for an agent? And I think that’s like, you know, getting, you know, one cycle that cost attributed to that one cycle is really a great way to look at it. So. 0:24:6.252 –> 0:24:30.52 Brian Haydin The model, you know, I think the model is 1, the point that I’m trying to make is that the model is just one line in a larger system. And above that, like the cost, you know, there’s other things that people can see, model inference and the direct tools that are involved during the run. And then you’ve got the, you know, capabilities that make the workload usable in an enterprise. So identity, networking, data access, 0:24:30.132 –> 0:24:52.732 Brian Haydin retrieval, you’ve got integration, evaluation, telemetry, security. Some of those are variable costs. Some of these are fixed, and some of them are shared across a lot of different workloads. I would say that some, you know, probably existed before the agent and need, you know, need us to do a fair distribution of the allocation rather than like the full charge. 0:24:53.852 –> 0:25:16.172 Brian Haydin You know, but the point isn’t to burden every idea with an accounting, with, you know, with the accounting exercises. The point is to prevent some of these invisible ideas or the invisible work as being treated as like they were just free. So, you know, I think someone still has to operate, you know, what you assembled. So you have to kind of include, you know, those costs as well. 0:25:16.812 –> 0:25:34.252 Brian Haydin And then make sure that the scorecard that you build preserves that end-to-end result, why the components that you might shift around underneath it can be, you know, moved around and shuffled around. Like you might change models, you might change platforms, you might convert from a licensing model to was pay as you go. 0:25:35.372 –> 0:25:54.492 Brian Haydin So make sure that you have a lot of those things like accounted for. And then the useful, like a useful fully loaded view follows like the entire workflow through its cycle. So 0:25:54.892 –> 0:26:15.132 Brian Haydin First, you build it. And so you’ve got discovery, process redesign, data preparation. Those activities, you know, need to be incorporated into what it’s going to take to build something. Then you’ve got the run. So that’s your model usage, your tool usage, you know, Azure search and storage, orchestration. 0:26:16.12 –> 0:26:39.52 Brian Haydin You’ve got then the third one, you’ve got your governance costs, permissions, traces, evaluations, policy enforcement. These are things that get baked into these agents as well. Then you’ve got your change management and adoption. You’re going to be investing in training. You’re going to be investing in model and prompt updates as you kind of like, you know, 0:26:39.452 –> 0:26:59.52 Brian Haydin do your regression testing, and then, you know, the ongoing ownership really, right? And then finally, you’ve got like the recovery aspect of it, the fact that agents don’t always work on, you know, the way that you expect them to, so rejected outcomes, incidents, bake that into there. 0:26:59.852 –> 0:27:18.572 Brian Haydin Those are all separate, fixed, and, you know, separate costs. And, you know, some of these are fixed and some of these are variable. So I would say that, like, you know, this is a pretty good way of like breaking down each one of the elements that you want to put into your cost breakdown when you’re planning this out. 0:27:20.12 –> 0:27:39.772 Brian Haydin At A portfolio scale, now you have like a little bit different metrics. You have, and you have a choice. So for example, 10 agents, you know, you could look at the 10 agents that are deployed in your portfolio as 10 independent products, you know, each recreating all of those things that we just talked about. 0:27:40.52 –> 0:27:50.172 Brian Haydin On the previous two slides, or you know, you could, you know, look at it as like one governed kind of plant as a portfolio as well, so… 0:27:52.892 –> 0:27:53.452 Brian Haydin Arm. 0:27:55.852 –> 0:28:16.52 Brian Haydin But shared, I would say like the caution I want to make, you know, here is that shared doesn’t really mean generic. And centralized, you know, cost doesn’t mean that it’s free. So you got to make sure that you kind of bake some of the cost distribution in there as well. I would say like reusable components, you know, 0:28:17.212 –> 0:28:33.772 Brian Haydin They have real, they have real costs that are associated with them, you know, and you need to make sure that you allocate those across it. So maybe you have a platform team, you know, that is supporting it, and somehow you have to break out the shared, you know, the shared workload on that as well. 0:28:35.612 –> 0:28:39.132 Brian Haydin So, let’s see, like… 0:28:40.292 –> 0:28:51.132 Brian Haydin Yeah, I think that I think that kind of makes sense to just like make sure that you’re planning out like some of those shared resources across the across the overall cost of it. 0:28:53.972 –> 0:29:12.732 Brian Haydin Another important topic to talk about a little bit is tokens. And I wanted to think about tokens a little bit differently and not just this specific cost. You can use tokens almost as a little bit of telemetry. They can tell you, I would say they can tell something. 0:29:12.772 –> 0:29:32.412 Brian Haydin Pretty useful about the behavior of a system, and you know, like a token. What I mean by that is that a token rolls into the request, and the request that you that you made might turn into several model calls, it might turn into a bunch of different retrieval steps, it might call some tools. 0:29:32.972 –> 0:29:51.692 Brian Haydin And then those calls can become another agent run, and then, and then only after all that, you have a run that becomes a complete, you know, complete the workflow. So then you can use, you know, all those tokens kind of aggregate up into, you know, how much work is actually going into it. 0:29:52.492 –> 0:30:10.892 Brian Haydin And once you get all that work done, you now have to decide was that result actually what you wanted to get done. So, so that’s like, you know, all this different costs rolling up into an acceptable outcome. So each one of those units that I talked about, though, has a different job. You’ve got a cost 0:30:10.932 –> 0:30:30.652 Brian Haydin per token, you know, and you’ve got a cost per completed workflow and you need to separate those. And the engineers can help you compare like different trade-offs. So when I do one model versus another, if I increase the thinking. 0:30:31.452 –> 0:30:53.292 Brian Haydin you know, depth, you know, I can, you know, I can measure whether I’m getting the same level of throughput on the acceptable outcomes. And then, and then you measure those trade-offs back and forth, right? So, you know, so you got to be able to like do AB testing in a couple of different cases, because you need to have the ability to compare 0:30:53.852 –> 0:31:16.412 Brian Haydin the resolved cases that happened in one model versus the resolved cases in the other, and which one actually got the results at a lower cost when you compute all the token aspect of it, you know, as part of it. That’s just one way of doing it. And then to do that, you need to have a little bit of discipline as well around FinOps. 0:31:17.292 –> 0:31:37.212 Brian Haydin building in some of the actual telemetry around token utilization and measuring when it got to a complete, you know, acceptable, acceptable outcome. And so the token economics is something that we’re going to use. 0:31:37.532 –> 0:31:47.532 Brian Haydin Because it’s one of the more measurable, it’s one of the more measurable things that that come out of out of out of an acceptable an acceptable outcome. 0:31:50.412 –> 0:31:57.292 Brian Haydin But when I say acceptable outcome, I also want to say that completed doesn’t mean accepted. 0:31:59.452 –> 0:32:18.292 Brian Haydin So one of the other concepts that I want to bring up is that an agent can reach the end of its steps and still produce a result that basically failed, that it didn’t get the answer, that met the quality threshold that you previously 0:32:18.372 –> 0:32:38.412 Brian Haydin defined or the action that it recommended might, you know, violate some sort of policy or it didn’t complete things in an automated fashion. So those are all things that can happen. And like, but the job got done, it just didn’t result in a acceptable action. 0:32:38.572 –> 0:32:42.972 Brian Haydin And so most often, what’s going to happen is that people are going to retry. 0:32:44.412 –> 0:33:4.212 Brian Haydin So there’s going to be either manual intervention that we have to account for, or they’re going to retry with different parameters or provide some more information, et cetera. So only when we get to that right side of the screen where it’s actually accepted, where the case was closed, where the 0:33:7.212 –> 0:33:24.412 Brian Haydin shipping discrepancy was figured out or the payment was applied to an invoice. That’s the acceptable standard that we’re going to get to and we have to measure everything that happened in between there. So I think I’ve hammered that home pretty well. 0:33:25.132 –> 0:33:45.372 Brian Haydin at this point, you know, but just to bring it up, like, as a, you know, as a who owns what part of the acceptance, I would say that the business owns, you know, things like what outcome were we buying, what makes it acceptable, 0:33:45.812 –> 0:34:5.12 Brian Haydin and what is the maximum cost or risk that we’re willing to tolerate, you know, or to accomplish that task. And then we need this partnership on the other end with engineering to answer the questions. Can we observe the attempts? Can we measure the acceptance result? How do we 0:34:5.132 –> 0:34:26.92 Brian Haydin you know, elevate to human intervention? Can we trace the tool path to make sure that it was doing the things that we expected to do? And then are we putting guardrails around the latency? So if you aren’t defining that in the business, and if you’re not finding the engineering, having a good partnership with 0:34:26.132 –> 0:34:49.532 Brian Haydin this, then all these metrics that we’re trying to, that we’re trying to compute, they’re going to break apart. So one of the things that I said before was that a purely technical definition, you know, is separate from like the business definition as well. And so a purely technical definition rewards completions, you know, 0:34:49.852 –> 0:35:10.652 Brian Haydin if the business doesn’t have alignment with it and rejects it. So it’s like this, this, you know, this kind of mess if they’re not working together. So, and one of the things that I said before is that it’s important to have somebody on the business that’s going to own the agents. And that’s why, you know, I think that’s why this is important. If you don’t have a known 0:35:11.452 –> 0:35:32.972 Brian Haydin being named business owner, then it’s, you know, not ready for you to take it to that next level, to that scale level, because you need the sponsor to build those kind of acceptable outcomes and definitions so that the engineering team can actually deliver against those results. So, and then 0:35:33.612 –> 0:35:51.932 Brian Haydin Why do I say a named owner? And the named ownership comes because you can’t have a, like when you have a committee making decisions, then like, then there’s no person, there’s no place to go when there’s a discrepancy or somebody has to make a decisions. And several groups may 0:35:52.812 –> 0:36:13.772 Brian Haydin you know, collectively start to approve the definition, but there’s no accountable business owner to decide, you know, what the, you know, what’s acceptable and what’s not acceptable. So I always say like, find a specific named person that’s going to be the decision maker, the product owner, if you will, if you want to follow. 0:36:14.52 –> 0:36:18.812 Brian Haydin a little bit more of the agile methodology, a named product owner would be, you know, one thing to do. 0:36:22.12 –> 0:36:46.12 Brian Haydin All right, so we named, we the first half of the scorecard names the work, you know, and we start with the business owner and then we move in the accepted outcome. And those might sound a little bit more like administrative, but they keep the metric from drifting into whatever was the easiest thing to count. 0:36:46.732 –> 0:37:8.652 Brian Haydin But then we capture 2 volumes. We capture like the attempts and the accepted outcomes. Attempts tell you the demand placed on the system and the accepted outcomes are going to tell you things like the throughput that cleared the business gate that we defined. The difference is going to be the failures, the rework, the escalations, things like that. 0:37:9.452 –> 0:37:31.372 Brian Haydin and maybe, you know, in like the cases that got unresolved. And then finally, we want to set an allowable cost ceiling. And that ceiling comes from prior process, like a service margin, avoided loss, available capacity. It doesn’t have to be a perfect number for that. It just, you know, on day one, but it has to be, 0:37:31.412 –> 0:37:54.812 Brian Haydin it has to be something that’s called out, you know, before you get started. Because without a ceiling, lower cost is always celebrated when, you know, even when the workflow is like not acceptable, when that acceptable work didn’t get met. And then, you know, the ceiling doesn’t have to be equal to the cost of like, I would say that the ceiling doesn’t have to be like the cost of the manual process. 0:37:56.92 –> 0:38:13.372 Brian Haydin But that might be a good starting point, you know, but set some sort of allowable cost ceiling on each one of these turns. I would say important because one of the conversations I have with a lot of my customers is talking about the ROI, you know, and 0:38:13.972 –> 0:38:33.772 Brian Haydin while a pilot answers questions and provides learning, our agents are intended to make our work more efficient, right? So we definitely want to, you know, reduce that burden of cost. So the second scorecard that we want to build 0:38:34.252 –> 0:38:53.292 Brian Haydin explains the cost, right? So the acceptance rate tells us how much demand becomes usable throughput, and straight through rate tells us how much clears without human intervention. So we got this idea that median cost per accepted outcome 0:38:53.532 –> 0:39:5.452 Brian Haydin Like, we, if we can calculate the median cost per accepted outcome, um, that’s going to describe what, like, the normal path is across the board, uh, and uh, so the… 0:39:6.652 –> 0:39:25.812 Brian Haydin then like if you’ve got the 95th percentile exposing like just the expensive one, you know, then the loops and the usual tool paths, you know, that can kind of get hidden when you’re just looking at like 1 edge of the spectrum. 0:39:26.412 –> 0:39:28.652 Brian Haydin So, um… 0:39:29.772 –> 0:39:36.652 Brian Haydin The cost per accepted outcome is, you know, something that you measure against like the ceiling. 0:39:38.652 –> 0:39:57.852 Brian Haydin And the other measures tend to explain why those moved and what the team can change. So the tail metric answers the practical question, how bad are the expensive past we experienced? You know, you know, how bad are the expensive past we experience, you know, when they’re often 0:39:57.932 –> 0:40:17.772 Brian Haydin you know, when they happen often enough for us to actually care about. So if the median improves while the 95 percent deteriorates, then, you know, the team might actually be optimizing the easy cases and allowing the exceptions to become way more costly for us. And that becomes kind of dangerous for us to 0:40:17.892 –> 0:40:36.892 Brian Haydin to build an operational plan against. And the last field that is up here is the decision. So like at the end of the day, at the end of the day, what we want to drive towards is this decision that we’re going to either scale this, we’re going to tune this. 0:40:37.212 –> 0:40:57.932 Brian Haydin we’re going to reroute this, constrain it, consolidate it, retire, almost like we would do with like an app rationalization, you know, traditionally. So the scorecards that we’re building out here, scorecard one, part one and part two, help us, you know, make that decision, you know, to scale, tune, route. 0:40:58.12 –> 0:40:59.292 Brian Haydin Or retire. 0:41:1.772 –> 0:41:2.252 Brian Haydin On. 0:41:3.212 –> 0:41:7.52 Brian Haydin I’ve got a little bit of a formula here that, you know, that can help. 0:41:10.572 –> 0:41:29.92 Brian Haydin Add the fair share of, add the fair share of the shared platform costs. Remember if we talked about like this idea that there’s shared resources, so take that fair share of the shared platform cost, you know, add that with the direct runtime and tools, add that. 0:41:29.212 –> 0:41:49.212 Brian Haydin to the ongoing operations and the human review and the exception handling and the cost of failures and rework. Take all of that and divide that total by the accepted business, the number of accepted business outcomes. And so that’s the formula. The allocation does not need false. 0:41:49.252 –> 0:42:8.12 Brian Haydin precision, you know, like don’t make any of this up. Shared costs can begin with, you know, showing a practical project environment, attributable consumption cost, you know, those types of things. And then when… 0:42:8.492 –> 0:42:9.52 Brian Haydin Um… 0:42:10.132 –> 0:42:35.12 Brian Haydin when the recording becomes a little bit more sophisticated, then you can offer better precisions, you know, precision a little bit later. I would say for high consequence types of workflows, things where the risk of wrong is up, keep the expected loss visible as, you know, a separated like risk layer adjustment or risk adjustment. 0:42:35.132 –> 0:42:56.412 Brian Haydin Layer, so a workflow that saves like a way to think of this is like a workflow that saves 50 cents per case, while increasing the probability of a costly error might look more efficient, but it, but the direct cost view, you know, becomes a becomes a little bit more visible. 0:42:56.732 –> 0:43:15.932 Brian Haydin when you drive up those costly errors. So this would be like the typical formula that I would run and then run the formula as like a recurring series. So don’t do this just one time at the end of the pilot or when you’re going back to your leadership to 0:43:16.172 –> 0:43:36.892 Brian Haydin to, you know, justify reinvestment, run them on a regular basis, weekly, monthly, quarterly, you know, whatever the cadence is that you want to show. And why? Because a trend shows whether the scale is starting to amortize, you know, the fixed cost. It can show things like the quality of 0:43:37.932 –> 0:43:58.92 Brian Haydin of things are improving the economics of it, you know, after it gets deployed, after it gets used, after people start learning how to use the tool that you’re building. And then if you look at that trend, now you can actually start to do some of the AB testing and see if it’s improving the results economically over time. 0:43:58.652 –> 0:43:59.132 Brian Haydin Um… 0:44:0.652 –> 0:44:20.412 Brian Haydin So, what does this look like? I would say, like, so if we take a if we take a step back and think about like a synthetic workflow that that we build, and it has 10,000 attempts in a month, like just sort of hypothetically thinking of a… 0:44:20.852 –> 0:44:38.892 Brian Haydin of a scenario. 7,500, you know, take the straight pass through path, right? They just go straight through. And 7,000 of those clear the acceptance gate. And that means that 500 are rejected or require some sort of rework. So it’s the math that we’re trying to get through. 0:44:40.492 –> 0:45:0.892 Brian Haydin The remaining 2500 need 5 minutes of human review. Let’s like hypothetically say that. So review is productive in this model. It rescues 1500 cases and then it becomes accepted, then it becomes, you know, the accepted outcomes. 1000 reviewed cases still fail the gate. 0:45:1.12 –> 0:45:22.412 Brian Haydin or require, you know, some sort of additional rework. And the result is that 8,500 accepted outcomes and 1500 that don’t clear the acceptance during this cycle. So that distinction, I think, is, you know, is pretty important for us to call out. 2500 review cases are not the 1500 failed cases. 0:45:22.452 –> 0:45:41.212 Brian Haydin Review isn’t like, don’t treat review as pure waste. It’s part of the operating design that makes, you know, the outcomes actually viable. And over time, you should see that population distribution should move. You’ll get better straight through behavior, and that may reduce some of the review. 0:45:41.852 –> 0:46:0.852 Brian Haydin You might have new policies, and they may temporarily increase it. They may, you know, decrease it, increase it. But the useful question is not whether the percentages are starting to, you know, look impressive. It’s whether the movement preserves the acceptance, the risk, and the service while improving the fully loaded unit. 0:46:1.132 –> 0:46:23.132 Brian Haydin I hope that makes sense to you. The target, like the optimization target, shouldn’t be like that we’re going to delete the reviewer, that we’re going to remove the human in a loop. It’s, you know, it’s going to be to improve routing, you know, the routing distribution, the context, the checks, the deterministic checks, the workflow quality. 0:46:23.852 –> 0:46:38.92 Brian Haydin and the authority design so that fewer of these cases actually need to have somebody inspect it and you still get the same 8,500 outcomes at the same quality, you know, and that’s the preservation of that. 0:46:41.92 –> 0:47:1.292 Brian Haydin And then we price the system. So we’ve got AI and direct tool run runtime that in this hypothetical scenario cost $3,000, you know, for the month. And then we calculated human review cost 15,612, you know, 15,625 dollars. 0:47:2.92 –> 0:47:24.52 Brian Haydin That is 2500 reviews, 5 minutes a piece at a arbitrary loaded labor rate of $75 an hour. So we allocate $8,000 for shared platform and the operating support, failure and rework, that line item is going to cost, you know, add another 4000. 0:47:24.492 –> 0:47:44.92 Brian Haydin The total is $30,625. Now, divide that by the 800, the 8500 acceptable, you know, accepted outcomes, and then the fully loaded cost is $3.60 per accepted outcome. There’s a lot of math on the screen. 0:47:44.972 –> 0:48:7.292 Brian Haydin So I hope this is how we’ve been, like that other calculation that I talked about, this is essentially how we went around and built that $5 versus a dollar type of thing. So I would say like, now that we did the math, the specific number, that’s not like, that doesn’t have to be the benchmark. If you do some things like change the labor rate, 0:48:7.852 –> 0:48:28.692 Brian Haydin the review duration, the failure cost, then the number is going to change, right? We have a lot of like variables that are in here. And so, but the benchmark here that we’re setting is the method, and then, and then also the transparency of how we’re actually allocating the cost of scaling these things. And so one of the thing is like, 0:48:28.972 –> 0:48:53.612 Brian Haydin that I would say is intentionally that we didn’t that we didn’t put up on the screen is the value of the accepted outcome. So, $3.60 could be could be a great number, if you know, or it could be a terrible number depending on what the result is worth. So, if it cost me $3.60 to do an accepted 0:48:53.732 –> 0:49:13.532 Brian Haydin outcome, and you know it only provided $1.00 of value, I’m in the hole, right? So, but today, what we’re talking about is building the cost denominator and the value comparison belongs in the ROI layer. 0:49:14.492 –> 0:49:20.12 Brian Haydin So, you know, that isn’t necessarily the point of making the calculation, right? 0:49:21.292 –> 0:49:40.652 Brian Haydin So now that we have some of those laid out, remember when we talked about the two labels, the two levers that we had before, now I can add those labels back to you and you’ll understand this a little bit more. If we somehow 0:49:40.692 –> 0:50:1.292 Brian Haydin cut AI and tool runtime in half, then the monthly savings is $1,500. You know, that’s useful, but it’s bounded, I think, by the size of the line that we had. If we improve the workflow so that the review rate falls from 25% to 15%, 0:50:1.772 –> 0:50:20.92 Brian Haydin we have 1000 fewer cases requiring 5 minutes to review. And at that same loaded rate, that saves $6,250 while preserving the 8, you know, the 8,500 accepted outcomes, you know, with the same quality and the same risk threshold. 0:50:20.652 –> 0:50:39.772 Brian Haydin So that second lever is actually four times larger than the first lever, right? And I hope this is starting to make a little bit more sense why we started out with those numbers. So improving the workflow and, you know, cutting down the amount of human intervention is actually more important. 0:50:39.812 –> 0:50:58.332 Brian Haydin important to us than cutting down on the economics of what the act, what’s going into it, like the model cost, et cetera, like that. So you wanted, I know it was a long-winded way to get to the punch line here, you know, on this. 0:50:59.132 –> 0:51:16.492 Brian Haydin But we should be focusing, I think, our efforts on, you know, improving lever B rather than, you know, constantly thinking about token costs, you know, you know, to improve the cost of these, you know, these agents that we’re building. So, 0:51:18.12 –> 0:51:37.52 Brian Haydin So this, like, this was like a fake, like made-up kind of like use case. And let’s, you know, I guess like taking a look at this, like outside of the lens of what I put together for this webinar, I just wanted to talk about some of the studies and very recent 0:51:37.212 –> 0:51:58.612 Brian Haydin August, McKinsey had a illustration for a banking customer service agent, and the token cost represented roughly 20 to 25 percent of the variable run cost, while the human oversight in that use case represented roughly 70 to 75 percent. And in banking onboarding, the same analysis 0:51:58.732 –> 0:52:23.372 Brian Haydin expected risk and functional experts to review roughly 10 to 20% of agentic runs that they were doing. So those are those are workflow specific figures. They’re not like, I wouldn’t use that as permission to tell every client that people will be doing 70, will be 70% of their cost. But what the evidence, you know, supports in this study 0:52:23.532 –> 0:52:43.92 Brian Haydin is that is the shape of the problem, that the fully loaded workflow, it still includes humans. It includes humans and deterministic systems, infrastructure, orchestration, and the ongoing operations. The visible token bill, you know, that still exists, but it 0:52:43.212 –> 0:53:3.612 Brian Haydin it’s probably not the dominant lever that you have to change. And so the recommendation that came out of there is that is pretty much the same as what we just covered, you know, today is, you know, don’t stop at the model selection. Redesign things like the exception paths and simplify the review processes that go into it. 0:53:4.172 –> 0:53:24.12 Brian Haydin that are going to require the more costly human oversight, the more costly human intervention. And so the same article emphasized that scale can amortize, you know, fixed costs and that reusable capabilities can improve the economics across, you know, multiple workflows, so across all those workflows. 0:53:24.652 –> 0:53:43.452 Brian Haydin So that I think is like, you know, the better punchline out of all this. So then the obvious question is, you know, to like, the obvious question out of that then is, how do we drive to near 0? 0:53:44.492 –> 0:54:3.332 Brian Haydin you know, review. If people become the expense, you know, that expensive line, then how do we automate the people out? And I think like that can lower the visible operating costs and, you know, increase expected loss. But the review may be the control that makes the whole workflow viable. 0:54:3.852 –> 0:54:22.92 Brian Haydin So in a high value transaction, you know, where you’ve got maybe ambiguous matches, policy exceptions, or like, you know, you know that there’s going to be low confidence recommendations that deserve human judgment, you have to 0:54:22.212 –> 0:54:40.812 Brian Haydin you have to have that as a consideration. You know, the target is the right is the right review. So the target should be route routine cases all the way through and send ambiguous or consequential cases, you know, to the right person efficiently. 0:54:41.532 –> 0:54:45.612 Brian Haydin Get it to the right person at the first, you know, the first time. 0:54:46.652 –> 0:55:0.652 Brian Haydin And so it’s, you know, it that that’s really that’s really what’s going to drive the cost down. So I got a couple more slides, but we’re running out of time here, so… 0:55:5.212 –> 0:55:26.172 Brian Haydin Going back to, going back to something that I talked about is that you want to forecast. So agentic activities are kind of like this distribution curve, right? So you have a lot of things that happen sort of in like the happy path. 0:55:27.132 –> 0:55:45.212 Brian Haydin and things that can be automated. And then you have these P95s, these things at the end. But those are typically the exceptions that require the most amount of human intervention. And so, like, what I really wanted to talk about on this slide is that 0:55:45.532 –> 0:56:4.572 Brian Haydin make sure that you’re accounting for that, because that’s where a lot of your cost is going to happen. That’s where the most human intervention is. If you’re trying to plan out your cost for the median, where it’s not likely to have a lot of human intervention, you’re going to miss out on some of the most expensive stuff that the agents 0:56:5.292 –> 0:56:30.12 Brian Haydin that’s going to cost you to scale this agent out. And then lastly, I would say that you want to follow the evidence and make sure that if you have the telemetry baked into it and you’re monitoring the budgets, not just of the tokens, but of the people, that you can… 0:56:30.332 –> 0:56:48.412 Brian Haydin You can make a much better case, you know, to reinvest or to scale to scale the agent. So, at each one of these, at each one of these gates, whether you’re exploring and beginning, you know, to prove that the capability worth, you know, worth investing in and proving, 0:56:48.972 –> 0:57:9.52 Brian Haydin each one of those like stages, you should be able to quantify what the effect was, what the cost was, and then make a decision whether you should move past that into the next, into the next gate, if that makes sense. Just to let everybody out of here on time, I ran a little bit short with the last couple of slides. 0:57:9.932 –> 0:57:26.812 Brian Haydin But I think I hit the punchline in most of what I wanted to talk about on this webinar today. So nonetheless, if you would give us a little bit of feedback, I invite you to reach out to us and follow us on LinkedIn again. And thanks for spending time with me today.
Events Best of FabCon Europe: The Fabric Highlights That Matter FabCon Europe is one of the biggest events in the Microsoft Fabric community, generating a flood of announcements, roadmap updates, and expert insights. In this session, Suneer cuts through the noise to bring you the developments that matter most—and what they could mean for your data strategy. We’ll cover the features, trends, and discussions that… October 8, 2026
Events Show Me the Money: Measuring Real ROI on AI Agents Anyone can build an impressive agent demo. Proving it delivers business value is what gets executive buy-in. In this session, we’ll break down how to measure AI agent success in terms your CFO and leadership team actually care about. Learn which metrics matter, how to establish meaningful baselines before deployment, and how to separate real… September 24, 2026
Events Copilot 201: From Adoption to Enablement Getting Copilot into users’ hands is easy. Getting it into their daily workflow is where the real work begins. This 200-level session is built for champions, IT leaders, and enablement teams tasked with driving adoption beyond the pilot phase. You’ll learn practical strategies to increase usage, uncover high-impact use cases by role, measure what adoption… September 17, 2026