Two weeks of work in four hours, for seventy-five dollars
Somebody at Jefferson Health needed to analyze a big pile of files. He asked for access to one of the agentic AI tools, got it that morning, and by the end of the day he had finished work that would have taken him about two weeks.
The token bill for that day was $75.
Luis Taveras, Jefferson's chief digital and information officer, told Becker's exactly what I would have said in his chair. A highly compensated person turning two weeks into four hours for seventy-five bucks is a fantastic trade. Then he said the second part, which is the part I want to talk about.
"But if he does that every day, the number is going to get pretty high."
That sentence is the whole 2026 AI problem in eighteen words, and I have not been able to stop thinking about it since I read it.
First, what you're actually buying
If nobody has explained tokens to you in plain English yet, let me do it, because I’m an old-school nerd.
A token is a chunk of text, roughly four characters. When somebody types a question into an AI tool, the system counts the tokens going in and the tokens coming back out, and bills you for both. Output usually costs more than input.
That's it. That's the unit. You are buying words, by the syllable, wholesale.
Two things follow from that, and both of them are uncomfortable for your expense account.
The first is that nobody using the tool can see the meter. Your coder, your denials analyst, your CDI specialist, none of them have any idea that the sixteen follow-up questions they just asked cost anything at all. There's no gauge. There's no little dollar figure in the corner. It feels like the free version, because it looks exactly like the free version.
The second is subtler, and I think meaner. Taveras said it bluntly: these models are chatty by design, and their chattiness is billable. You ask a simple question, you get an answer, and then the thing offers to do four more things for you. Since output costs more than input, every one of those helpful little suggestions is revenue for somebody. The tool is not neutral about how long the conversation goes.
I'm not saying that's sinister. I'm saying it's a business model, and you should recognize it, because you have spent your entire career recognizing exactly this pattern in payer behavior.
No one really has a plan or forecast
Philip Payne is the chief health AI officer at BJC. When Becker's asked him about forecasting token spend, he did not hedge:
He said he doesn't believe anybody has solved the problem of accurate, reproducible forecasting of token utilization.
Not "it's hard." Not "we're refining our model." Nobody has solved it. Tokenization varies by platform, usage varies by use case, and the scaling effects vary by volume, so the three variables you'd need to build a forecast all move independently.
Meanwhile the exposure is not hypothetical. Uber blew through its entire annual agentic AI budget in three months and had to cap employee use. Google reported seven times the token consumption year over year. Epic says more than 85% of its customers are now live on its generative AI tools, and those tools run on consumption underneath, whether or not your contract is written that way.
Adam Landman at Brown University Health put his finger on the structural problem: Token pricing is hard for any organization because cost scales with adoption, and adoption is the goal.
Read that again if you're the person who has to sign the contract. Every other technology purchase you have ever made got cheaper per unit as more people used it. This one gets more expensive, forever, and success looks identical to overspending until the invoice shows up.
Putting some discipline into the mix
The good news is that the CIOs who are ahead of this are not being precious about it. Their answers are unglamorous, brilliant, and completely stealable.
Jefferson built a ladder. Routine question, use Google. Need more, use the basic Copilot license that's already bundled into the Microsoft agreement at no additional cost. Need more than that, Copilot Plus at a fixed monthly fee. Only the people who genuinely need frontier capability get the metered tools, and they don't get them until they've been through an education process on what the things cost. Jefferson also allocates tokens by user type rather than handing everyone the same bucket, and sends people reports on their own consumption.
Emory picks the lowest-intensity model that can do the job, and makes vendors open up their financial operations data before signing, so the system can forecast consumption. Advocate Health reframed the governance question entirely, from "which model is best" to which is the cheapest one that reliably gets the job done. Eric Kirkendall there says they're moving from cost per token to cost per outcome delivered, which is the right North Star and about eighteen months ahead of where most organizations are.
Brown keeps most of its scaled capabilities on flat per-user pricing and only pilots the consumption-priced tools in small groups with hard usage caps, specifically to generate real utilization data before forecasting. Underneath all of it, Landman says, is FinOps: real-time visibility, so that if spend moves unexpectedly they know within hours, while the number is still small.
And then there's the University of Utah Health. They put several million dollars into on-premises NVIDIA infrastructure, so the token economics live inside their own four walls. They call it AI sovereignty. It's a capital answer to an operating problem, and I expect to see a lot more of it by 2028.
Now the second number nobody has
Here's where I got genuinely uneasy this week.
Deloitte surveyed 64 healthcare CFOs at medium and large organizations, half of them health systems over a billion in revenue. Fierce Healthcare covered the findings, and one data point stopped me cold.
They sorted respondents into "AI scalers," the 44% who have deployed the technology most broadly, and everyone else. Three-quarters of the scalers plan to increase their gen AI and agentic AI investment over the next twelve months.
Only 18% of those scalers report mature financial attribution, meaning they can consistently measure what AI is actually doing to their financials.
Among the organizations earlier in their AI journey, that number is 31%.
The further along you get, the less able you are to tell what it's returning. Measurement doesn't improve with scale; it degrades. And the group that can see the least is the group spending the most and planning to spend more.
Deloitte's framing of the broader readiness problem was that it may fundamentally be a design gap, meaning expectations for CFOs expanded faster than the systems built to support them. I think that's generous but accurate. Nearly three in five of these CFOs are targeting two points or more of margin improvement in the next two years, and 47% don't believe their organization is prepared for what's coming at it.
That’s our reality. And it’s not a comfortable one.
What proof looks like when somebody bothers
I don't want to leave you thinking the returns aren't real, because they are, and the organizations getting clean numbers are getting good ones.
BJC and WashU put just under 500 primary care providers on ambient documentation and deliberately declined to measure ROI up front. Fix the experience, they figured, and the money follows. It did, on three different clocks. Note-writing time dropped about 7% almost immediately, then roughly doubled to 15% by month five, which is more than an hour a week back per physician. After-hours charting didn't move at all for a month, then declined steadily. Revenue came last and slowest: a 3% weekly lift in work RVUs, close to $2 million a year across 200-plus doctors.
Payne's summary is the most honest sentence I've read about ambient AI economics. At a minimum, he says, the technology is self-sustaining financially.
At Onvida Medical Group in Yuma, CFO Mike Lewis went the other way and was prescriptive about quantifiable outcomes from day one. He reported a 14% increase in patient-facing time, a 27% drop in after-hours documentation, 35% average documentation time savings, and about $24,000 in revenue yield improvement driven mostly by throughput, since clinicians averaged one additional patient per day. He also saw revenue capture go up and charge lag come down, which is the part that should interest anyone reading this from a revenue cycle seat.
Mercy reported a 22% cut in nurse flowsheet documentation time per shift, and at one location, incremental nurse overtime fell somewhere between 29% and 56% in the first twelve weeks.
Those are real numbers from organizations that decided in advance what they were going to count. That is the only variable that separates them from everybody else.
What It Means
In RCM 2030 I gave you a cost-to-collect target of three percent or better for large systems, calculated as total revenue cycle expense divided by cash collected. I've been using some version of that benchmark for most of my career.
Here's my question, and I'd like you to go find out rather than guess: Is your AI token budget inside that denominator?
Because if your ambient documentation tool, your denial appeal drafting tool, your coding assistant, and your patient messaging tool are all metered, and that consumption is landing in an IT budget line somewhere instead of in revenue cycle expense, then your cost to collect is understated, and you don't know by how much. You are reporting a number to your board that is quietly wrong, and it will get more wrong every month, because cost scales with adoption and adoption is the goal.
It's fixable in a quarter if somebody decides to own it.
What I'd do between now and January
Find the meter. Ask your CIO for a consumption report by tool and by user type for the last six months. If nobody can produce one, that is your answer and your first project. You cannot manage what you cannot see, and Landman's standard of knowing within hours to days is the right one.
Put token spend in cost to collect. Every AI tool touching a revenue cycle workflow gets allocated to revenue cycle expense. Restate the last four quarters so you have a real trend line rather than a cliff.
Pick your denominator before you deploy. Decide what you're counting before the tool goes live, the way Onvida did, or decide deliberately that you're measuring experience first and revenue later, the way BJC did. Both work. What doesn't work is deciding afterward, which is how you end up in that 18%.
Build the ladder. Not everyone needs a frontier model. Most people asking most questions need the tool you're already paying a flat fee for. Tiering access is the single cheapest cost control available to you and it costs nothing but a policy.
Get the pricing structure into the contract conversation. Flat per-user pricing where you can get it. Hard caps and small pilots where you can't. Vendor financial operations data before signature, the way Emory demands it. And ask every vendor, out loud, what happens to your bill if usage triples, because it will.
Where this goes
By 2030, I think the health systems still standing will be the ones that treated 2026 and 2027 as a margin recovery sprint rather than a policy waiting game. I've been saying versions of that for a while and nothing this week changed my mind.
What did change is that I now think AI cost governance is going to be one of the things that sorts the survivors, and I did not expect that a year ago. Hospital operating margins are running near one percent. The bulk of the Medicaid reductions land across late 2026 and 2027. Into that, we're layering a technology whose cost is unbounded, unforecastable, and rises with exactly the adoption we're all working to drive.
The organizations that win this are not going to be the ones that negotiated the cheapest token rate. They're going to be the ones who knew what they were spending and what they were getting, in the same report, on the same page, in the same quarter.
If you want a structured read on where your organization actually stands, the RCM AI Readiness Scorecard covers five domains including governance and cash forecasting. You check what's running today, not what's planned.

