// Episode 012

The Meter Room: Read the Meter Before the Invoice Does

3 September 2026 · 19 min listen · 3,110 words

Software used to be priced at the signature. Now it accrues at the meter. The first stop on the Grid tour: the unit, the rate, the commitment, and the true-up.

// Listen
ConsumptionAI SpendFinOps
In This Episode

What You'll Take Away

A note on evidence

Ramp's benchmarks are drawn from its own customers' payment data rather than a survey, and Ramp sells the token spend product that produces them. The Coinbase figures come from the company's own published playbook. A widely circulated token price-decline statistic was deliberately excluded because it traces to a platform vendor measuring its own customer base. The commitment-sizing guidance is flagged in the episode as opinion.

Key Terms

Defined Plainly

Effective rate
Dollars on the invoice line divided by units consumed. The only rate that is real, because the same unit can bill at full, cached, or batch pricing on the same meter in the same month.
Unit, rate, commitment, true-up
The four parts of every consumption agreement. What makes the meter click, what each click costs, the floor you promised to burn, and what happens when actual crosses committed.
Prompt caching
Engineering requests to reuse context that has already been paid for. The mechanism behind a fifth-of-list effective rate, and the term an AngelList controller had never heard until a spend briefing surfaced ten thousand dollars a month of it.
Model tier migration
Teams upgrading to premium models for quality, usually invisibly. Ramp's data names it the single biggest driver of surprise cost: premium models went from 5.7% of business token cost in June 2025 to 55.9% by April 2026.
Estimated read
The utility practice of billing a guess from history when nobody reads the meter, reconciled later. When you do not read your own meter, you are living inside the vendor's estimate of you.
Demand charge
The commercial power premium pegged to your single worst fifteen-minute window of the month. Consumption billing is demand-charge economics, and an agent workflow with no per-run ceiling is a peak nobody is watching.
Full Transcript

Read the Episode

The complete transcript — 3,110 words. Search it, quote it, or send the relevant section to a colleague who needs to see it.

Show / hide transcript

Cold Open: Twenty Point Seven Times

Hey everybody, and welcome back to the Operational ITAM Podcast. I'm Bill Van Nort, and today we are reading a meter. On July sixteenth, a New York fintech company called Ramp put out a press release with a number in it that I have been thinking about ever since.

Ramp processes corporate card and bill payments for more than seventy thousand American businesses, which means they are not surveying anybody. They are watching the money move. And what they reported is this.

Since June of twenty twenty-five, AI token spend across their customers has increased twenty point seven times. Not twenty percent. Twenty point seven times, in about thirteen months.

That is not a line item growing. That is a new category of spend arriving at full speed, and it does not behave like anything your procurement process was built to approve. Nobody signed a purchase order for that number.

It accrued. One request at a time, on a meter almost nobody in the organization has ever read.

Good morning, good afternoon, or good evening, wherever you're listening from. This is the show where we take the unglamorous machinery of enterprise technology and make it make sense. Grab your coffee.

Last episode we did something we had never done before. We closed the library. Eleven episodes in those rooms, and we walked out with full honors and stepped into the Grid, because the estate we manage stopped behaving like a collection and started behaving like a utility.

Today is the first stop on the tour I promised you, the meter room. In episode eleven we laid out the five forms AI takes in your estate, and one of those forms was Metered, the form where spend scales with usage. Today we walk all the way into that one.

And honestly, the warning shot came even earlier. Back in episode nine, Zylo's SaaS Management Index told us seventy-eight percent of IT leaders got hit with unexpected consumption or AI charges. At the time, that was a renewals problem.

It is not a renewals problem anymore. It is the shape of the whole market.

Reading the New Bill

So let me give you the numbers, and I'll name the sources, because on this show, we do that. Ramp publishes benchmarks from their token spend product, real payments, not survey answers, and the April numbers tell you everything about how this category behaves. The median business paid about two thousand two hundred forty-six dollars a month for AI.

The average paid about one hundred forty thousand eight hundred forty-two dollars. Sit with that gap for a second. When the average is sixty times the median, that is not a typo.

That is a distribution with a long, expensive tail, and the distance between typical and tail is usually one team, or one automated workflow, running without a ceiling. Same dataset. The effective rate across everything averaged seventy-two cents per million tokens, but a thousand dollars buys you roughly fourteen and a half billion tokens on the cheapest mainstream model and about seven hundred four million on a flagship.

Same money. Twenty times difference, decided by a default someone set in an afternoon. And the FinOps Foundation, in their sixth annual State of FinOps survey, nearly twelve hundred practitioners, found that ninety-eight percent now manage AI spend.

Two years ago that was thirty-one percent. The single most requested capability in the entire survey was granular AI spend monitoring. Tokens.

Requests. GPU utilization. The people whose whole job is reading bills are telling us the bill they cannot read yet.

So here is the thesis. The price of software used to be set at the signature. Now it accrues at the meter.

A seat is a price you approved once. A meter is a price you discover every month. And the discipline this demands is simple to say and rare to find.

Read the meter before the invoice does.

What a Meter Is: Unit, Rate, Commitment, True-Up

So let's open the panel and look at what a meter actually is, because every consumption agreement in your estate, AI or otherwise, is built from the same four parts, and you need to be able to name all four for every meter you own. Part one is the unit. What makes the meter click forward by one?

For a language model it is the token, and here is the trap inside the trap. Tokens in and tokens out are different units at different prices, and output typically runs three to five times the cost of input. For infrastructure it might be the GPU hour.

For a platform agent it might be the message, or the run, or a credit, which is a currency the vendor invented and controls. If you cannot say what the unit is, stop. Everything else about the meter is unknowable until you can.

Part two is the rate. And this is where consumption pricing gets genuinely strange, because the same unit does not have one price. The same million tokens can bill at the full rate, at a cached rate, or at a batch rate, on the same meter, in the same month.

Ramp's April data makes this concrete. Businesses running one popular mid-tier model paid an effective sixty-two cents per million tokens against a list price of three dollars. That entire gap is caching.

Same model. Same work. Fifth of the price, because somebody engineered the requests to reuse what was already paid for.

When your effective rate and the list price diverge that much, the list price stops being information. Your effective rate, your dollars divided by your units, is the only rate that is real. Part three is the commitment.

Most enterprise consumption deals have a floor. You promised to burn a number, and the meter runs against that promise. Which means a meter can hurt you in both directions.

Overage above the commitment, or a commitment you never consume, which is shelfware with a clock on it. Reading the meter mid-cycle means comparing the accrual against the commitment while there is still time to steer. And I will hand you some humility on this one, mine.

Years ago, on a purchase order I signed, I sized a consumption commitment straight off the vendor's projection of our growth, because the discount at that tier was beautiful and the projection came with a very confident chart. We never read the meter mid-cycle. Not once.

The true-up arrived at the end of the year like a weather event, and I remember thinking we had been surprised. We had not been surprised. We had been unread.

The meter knew the whole time. So let me give you an opinion here, clearly labeled as an opinion. Size a consumption commitment to what your own history says you will actually burn, your typical month, not the vendor's growth curve for you.

The counter argument is real, and I'll give it to you fairly. Bigger commitments buy better rates, and if your usage genuinely is doubling, committing low means paying overage on growth you should have seen coming. Fine.

Then let your own meter readings make that case, three months of them, and commit on your evidence instead of their enthusiasm. And part four is the true-up. What happens when actual crosses committed?

What rate applies above the line? When does the excess bill, and can it carry over? Those four answers live in the contract, and I will tell you from experience, almost nobody who watches the dashboard has read the contract, and almost nobody who read the contract watches the dashboard.

Real-World Savings: Coinbase and AngelList

Now let me show you what this looks like when real organizations do it well, because two examples crossed the wire this summer and they teach opposite ends of the same lesson. First one. In late June, Coinbase's chief executive, Brian Armstrong, posted their internal playbook, and the trade press picked it apart in detail.

Coinbase cut its AI spend nearly in half while its token usage kept growing. Read that again. Half the spend.

More usage. And the part I want you to hear is how. Not caps.

Armstrong said ninety-one percent of employees never even hit the old usage caps, so the caps were friction pretending to be control. His words on how you actually do it, quote, not with friction and spend alerts, with better defaults, routing, and caching, end quote. Their gateway defaults route routine work to inexpensive models, harder work escalates when it earns it, and their caching went from hitting five percent of requests to sixty percent.

Every one of those decisions happened before the meter turned. Second example, from the other side of the house. In Ramp's launch announcement, the controller at AngelList describes getting a weekly spend briefing that surfaced a term he had never heard.

Prompt caching. Not a controller's phrase. He routed it to engineering, they found they had been losing about ten thousand dollars a month, and the fix went in the same day.

Think about what actually happened there. The control lived in engineering. The visibility lived in finance.

The money got saved on the day those two met in the middle. That meeting place is the meter room, and in most organizations, that room is empty.

How to Read Usage: The Four Steps

Let's take a quick break. If you're getting value from this show, subscribe wherever you're listening. And if you know somebody who just opened an AI invoice they could not explain, send them this episode, because the explanation is in part two.

Every episode with full transcripts is at operationalitam.com. Okay, part two, how you actually read a meter before the invoice does.

Step one, find every meter. A meter you have not located is a meter the vendor reads alone, and your meters are hiding in all five forms from episode eleven. If you did that episode's homework, the shadow AI census, you already hold the beginning of your meter list.

Expense data, single sign-on, asset records. Every AI vendor on that census either bills a seat or bills a meter, and your first pass is just sorting them into those two piles. Add the cloud AI services riding inside your existing cloud bill, because those meters do not even arrive as separate invoices.

They arrive as line items inside a bill you already pay, under your cloud provider's name, which is exactly why the census keeps missing them. Step two, reconcile every invoice line to a meter you found. And I mean arithmetic, not vibes.

Take the dollars on the line, divide by the units consumed, and write down the effective rate. Then do last month. A line you cannot tie to a meter is not a cost you approved.

It is a cost you absorbed. And an effective rate that moved is telling you one of exactly three stories. The mix changed, meaning somebody moved work to a pricier model.

The caching changed, meaning the engineering got better or worse. Or the price changed, meaning the vendor moved the rate. Three stories, three different owners, and the invoice total tells you none of them.

The effective rate tells you all of them. Step three, watch the drivers, not the total. Ramp's data says the single biggest driver of surprise cost is model tier migration, teams upgrading to premium models for quality, usually invisibly.

Premium models went from five point seven percent of business token cost in June of twenty twenty-five to fifty-five point nine percent by April. Ten months. And the difference per workflow can run ten to one hundred times.

The other driver is agents, and I want to say this one carefully, because it is the future of this problem. An agent takes as many steps as it decides the task needs. Which means you approved the task, and the agent is deciding the bill.

Any automated workflow without a per-run ceiling is a meter with no shutoff valve. Step four, set the defaults, because this is the actual governance layer. The Coinbase lesson is not that they saved money.

It is where the saving lived. In the gateway default, in the routing rule, in the cache policy. Decisions made once, upstream, that shape every request afterward.

A spend alert fires after the money is gone. A default decides whether it ever leaves. And one thing I am deliberately not giving you today.

There is a token price decline statistic making the rounds, a big satisfying percentage. I chased it to its source, the way we did with refresh numbers back in episode seven, and it traces to a platform vendor measuring its own customer base. So you will not hear it here.

What I can give you is Ramp's paired numbers across the same window. Token usage among businesses up one thousand one percent, January twenty twenty-five to April twenty twenty-six. Spend up four hundred ninety-seven percent.

Falling unit prices are real, and they are not saving anybody money, because usage is growing faster than price is falling, and it is not close.

Budgeting the Unknown: Elena's Question

Which brings us to a listener question, and this one landed in my inbox from Elena in Cincinnati. Elena writes, our finance team wants next year's AI budget as one fixed number. With consumption pricing that feels impossible.

What do I even give them? Elena, your instinct is right, and here is the move. Do not give them one number, because one number is a guess dressed up as a commitment, and finance people can smell that.

Give them three numbers, reported separately, never blended. Number one, the committed floor. That is contractual.

It is the most honest number you own. Number two, the expected run rate, built from this year's effective rates and named drivers, and put a real buffer on it, because Ramp's data shows month over month swings above forty percent are common even with stable headcount. When finance asks why the buffer is that big, do not defend it.

Show them the swing data and let the data do the arguing. Number three, the uncapped exposure, and next to every meter that creates it, the specific control that caps it. The per-run ceiling, the gateway default, the alert threshold.

Here is the reframe, Elena. You cannot promise finance what a meter will read, and you should stop pretending you can. What you can promise is which drivers are governed and who owns each one.

Budget the drivers. Report the meter. That conversation is the whole job.

Utility Lessons: The Meter Room

Alright. The Grid room. We are standing in the meter room this week, and if you have ever actually looked at the utility meters in a commercial building, you know three things about them that map onto everything we just covered.

First, the meter spins whether anyone reads it. It does not care about your budget cycle, your approval workflow, or your fiscal year. Consumption is recorded the moment it happens, and the only question is who reads it first, you or the biller.

Second, utilities have a thing called an estimated read. When nobody physically reads the meter, the utility guesses your usage from history and bills the guess, and they reconcile it later. And I want you to notice whose model produces the guess.

When you do not read your own meter, you are living inside the vendor's estimate of you. Third, and this is my favorite, commercial power bills carry something called a demand charge. You do not just pay for the energy you used.

You pay a premium pegged to your single worst fifteen-minute window of the month. Your peak sets the price. Now think about an agent workflow with no ceiling, spiking on a Friday night.

Consumption billing is demand-charge economics, and your estate is full of peaks nobody is watching. The utilities figured all this out a century ago. They read on a schedule.

They audit the estimates. They engineer the peaks down. None of it is new.

It is just new to us.

Measurement Discipline

Which brings us to today's principle. And let's take the whole run from the top. Hardware Asset Management is a custody discipline.

Software Asset Management is an evidence discipline. Audit defense is a process discipline. Settlement is a commercial discipline.

Shadow IT is a service discipline. Refresh is an economics discipline. Disposal is a liability discipline.

Renewal is a leverage discipline. Tooling is a judgment discipline. The AI estate is a governance discipline.

And consumption? Consumption is a measurement discipline. The seat asked you to count people.

The meter asks you to measure flow. And a flow you do not measure is a price you do not know, on a bill you have already agreed to pay.

Homework and Next Steps

Class dismissed. Here is your homework, and it needs about an hour. Pull your largest AI or consumption invoice, last full month.

Go line by line and do two things for every line. First, name the meter behind it. The unit, and the rate.

Second, divide the dollars by the units and write down the effective rate, then do the same for the month before. When you are done you will hold a one-page list of two kinds of findings. Lines with no meter you can name, and rate movements with no story you can tell.

Date the page. Sign it. That page is your first meter reading, and whoever you show it to will understand this episode without hearing a word of it.

Next episode, we walk from the meter room to the standards body, because here is what happened while nobody in ITAM was watching. The FinOps Foundation built a common language for exactly the bills we struggled with today. It is called FOCUS, the FinOps Open Cost and Usage Specification, and this June their steering committee ratified version one point four, which adds invoice reconciliation, the exact discipline we just did by hand, straight into the standard, with native AI support already on the roadmap for the next version.

These are the people who renamed their own flagship conference around the token. The meter readers built themselves a language. Next week, we learn to speak it, and I will show you what it means for how ITAM and FinOps stop working the same estate in different rooms.

And listen, the case files are open. One situation, one page. The constraint you were under, what you did, what happened.

No company names, I will sanitize everything. Send me one worth working at the website, and I will build an episode around it.

I'm Bill Van Nort, this is the Operational ITAM Podcast. Find every meter. Compute your own effective rate.

Set the default before the invoice sets it for you. I'll talk to you next week. Take care.

Sources & Further Reading

Everything Referenced

Listener Case Files

Got a Situation Like This?

Send it over — anonymized, sanitized, no company names. Real constraints, real politics, real budgets. Situations get worked on air.

Submit a Case File
From Listening to Doing

Apply This to Your Own Environment

The podcast covers the principles. An Executive Briefing applies them to your renewal calendar, your portfolio, and your actual numbers. 45–60 minutes, no charge.