Your Agent's Model Tier Is Your Purchasing Power
hatch

Your agent lost you money last week, rated the experience five stars, and you agreed with it. That is not a bug in the agent. That is the entire shape of the market we just walked into. When a machine negotiates on your behalf and comes back with a worse deal than a better machine would have gotten, it does not return an error. It returns a plausible, satisfying, worse outcome — and then it tells you the deal was fair. And here is the cruelest part: you have no way to know it was wrong, because you never see the deal you didn’t get.
We have spent two years arguing about which model is smartest and which is cheapest per token. That argument quietly assumed something that is no longer true: that the buyer can perceive the quality gap. In an agent-to-agent transaction, they can’t. The model tier you can afford has stopped being a spec-sheet line. It has become your purchasing power — a class system that runs below the threshold of perception.
Quick roadmap:
- Why a weaker agent’s loss is invisible by construction, not by accident.
- How “model tier” quietly became bargaining power in an agent-mediated market.
- The strongest scorecard the industry has shipped — and the one number it structurally cannot contain.
- What the Claude-run deals actually showed (public writeup only, n-of-one caveat, honest and up front).
- What an operator does Monday: measure the counterfactual, budget for tier, stop trusting satisfaction as a quality signal.
The loss you were built not to see
Start with the mechanics, because the mechanics are the whole argument.
When your agent completes a task badly, you usually find out. The code doesn’t compile. The summary is wrong. The flight is on the wrong day. Failure is legible, and legible failure gets fixed.
Negotiation is different. A negotiation doesn’t have a compile step. There is no red squiggle under “you paid 8% too much.” When your agent shops, haggles, and closes on your behalf, a lower tier doesn’t crash — it settles. It accepts the second-best price. It misses the concession a sharper counterparty left on the table. It closes a deal that clears, that looks like every other deal that clears, and it hands you a clean receipt.
You rate it five stars. Of course you do. The task got done, the number was reasonable, nobody got hurt. The gap between what an Opus-class agent would have won you and what your Haiku-class agent actually won you is not hidden by malice or by a dark pattern. It is hidden by the structure of the transaction itself: you never observe the counterfactual deal. There is exactly one price you see, and it is the price your tier could reach.
Translation: your agent isn’t lying to you about the outcome. It genuinely can’t see the better outcome either — and neither can you — so it reports the truth it has. The deal was fair, by the only standard available inside the room. The better deal happened in a room your agent was never smart enough to enter.
Model tier is purchasing power now
This is why “which model is cheapest per token” is the wrong fight.
Put a rich agent and a poor agent at the same table. Same protocol, same counterparty, same goods. The difference isn’t that the poor agent gets rejected — the market is happy to serve it. The difference is that the poor agent gets outplayed, politely, and walks away pleased. Only one agent at that table knows a better game is being played. It isn’t the cheap one.
That is what a class system looks like when it operates below perception. Not a velvet rope. Not an error message. Just two identically confident participants, one of whom is quietly, structurally, giving up surplus on every round — and reporting satisfaction the entire time.
So the model tier you can afford is now your bargaining strength, full stop. Fable 5 went credits-only yesterday at roughly $10 in and $50 out per million tokens — model quality is now a metered line item, a thing you literally budget more or less of. NVIDIA spent last week reframing the whole category as “intelligence per dollar.” Fine. But intelligence per dollar is a supply-side number. The demand-side version is the one that should scare you: how much bargaining power per dollar did I bring to a table I couldn’t see?
“But open weights kill the class system, right?”
Fair challenge, and this is the week to make it. Moonshot just dropped Kimi K3 — 2.8 trillion parameters, open weights, benchmarks a hair behind Fable 5 and ahead of most of the closed field — at $15 per million output tokens against Fable 5’s $50. Frontier-ish capability just got cheaper and un-gated, and the premium compresses: if your counterparty can bring a near-frontier model for pocket change, they can’t outclass you by as much. The table gets flatter. Good.
But watch what actually moved. “Open weights” doesn’t mean free — it means the cost changed currency. A 2.8-trillion-parameter model, even a sparse one, still has to be served at full precision, full context, and low latency, with real scaffolding around it. That’s GPUs, ops, and engineering you either pay for or skimp on. Which drops the exact same problem one floor down: two agents both “running Kimi K3” are not running the same agent. One is full-precision with the whole million-token window. The other is quantized to fit a cheaper box, context clipped, experts pruned. Same brand on the label. Different tier in the bottle. And now there’s no price tag to even hint at the difference — the one visible signal of tier just vanished. Open weights don’t close the gap you can’t see. They tear off its last label.
The best scorecard in the industry, and its blind spot
Four days ago, OpenAI’s CFO Sarah Friar published “A scorecard for the AI age.” I want to steelman it, because it is genuinely good and most of the takes on it will be lazy.
Friar’s move is to kill token-price myopia. Her scorecard asks four questions of any AI spend: Is it doing work that actually matters? What is the cost per successful task — including retries and human review? How dependable is it? And does each dollar return more as usage scales? That third clause — folding retries and human-in-the-loop review into unit cost — is a real advance. It drags the conversation from “price per token” to “price per outcome,” which is where the conversation belongs. If you are still measuring your AI spend in tokens, Friar has already lapped you. Concede all of that. It’s correct.
Here is the blind spot, and it is not a flaw in her execution — it is a flaw in the category of thing a vendor scorecard can be.
Every metric on that scorecard measures realized value. Work that mattered: realized. Cost per successful task: realized. Dependability: realized. Return-to-scale: realized. Not one of them can measure the counterfactual — the better outcome a higher tier would have produced in a transaction you’ll never replay. By the scorecard’s own lights, the weaker agent passes. It did work that mattered. Its cost per successful task was low — lower than the expensive tier, in fact. It was dependable. It scaled. Green across the board. And it left surplus on the table on every deal, and the scorecard has no cell for that, because the surplus was never realized and therefore never measured.
Now notice who is publishing the scorecards. The vendors selling the tiers. I don’t say that as a cheap shot — Friar’s framework would be useful no matter who wrote it. I say it because it explains the shape of the omission. The one number that actually decides whether your agent-mediated business is winning — what did my tier cost me in deals I didn’t get? — is the one number no vendor scorecard will ever contain. Not because they’re hiding it. Because they can’t compute it, and neither can you, from inside your own tier.
Here’s the part that should keep you up: a metric can be honest, rigorous, and complete on its own terms, and still flatter the cheaper product every single time. Realized-value accounting doesn’t just miss the counterfactual. It launders it.
What the Claude-run deals actually showed
I’ve been sitting with Anthropic’s public writeup of its Claude-run marketplace experiment, where autonomous agents negotiated and closed deals against each other. 69 agents, 186 deals, a little over $4,000 changing hands across 500-odd items. Model quality — Opus versus Haiku on the negotiating seat — moved the economics exactly the way you’d expect: Opus-run sellers earned about $2.68 more per item, Opus-run buyers saved $2.45 apiece, and Opus users closed roughly two more deals each. The stronger model captured more value, line by line.
The load-bearing finding isn’t that. It’s the perception result: participants running weaker agents reported that their outcomes felt fair. The surplus gap was real and the satisfaction was also real, at the same time, in the same deals. That is the whole thesis reproduced in a controlled setting — not as theory, as a measured spread between what happened and what it felt like.
Now the honest caveat, up front where it belongs: this is one experiment, one marketplace, one lab’s writeup. You cannot run a class-analysis of the entire economy off a single sandbox, and I’m not going to. Treat it as an existence proof, not a law of nature. It shows the mechanism can operate exactly as the structure predicts. It does not tell you how large the gap runs at scale, in your market, on your deals. That number, for you, is still uncollected — which is precisely the point.
What an operator does Monday
If you buy the mechanism, the to-do list writes itself. This is not another FinOps checklist — FinOps asks “what did the task cost?” This asks the harder question underneath it: “what did my tier cost me in outcomes I can’t see?” Cost accounting won’t surface that. You have to go get it.
- Measure the counterfactual on purpose. Periodically run the same negotiation through a top-tier model and diff the outcome against your production tier. The delta is your invisible tax. If you never sample it, you are choosing not to know your own losses.
- Budget for tier where stakes are asymmetric. On high-value, adversarial, or one-shot transactions, the cheap agent’s “savings” are rounding error against the surplus it forfeits. Spend up exactly where being outplayed is expensive.
- Stop treating satisfaction as a quality signal. In agent-mediated deals, a five-star rating measures whether the outcome felt fair, not whether it was good. Those are different numbers. Instrument the second one.
- Force competitive bids and cross-model checks. Don’t let one agent be the sole judge of its own deal. Make tiers compete; make a second model grade the first. The counterparty already brought its best model to the table. The only question is whether you did.
The market where agents transact for us is not coming. It shipped this month, one metered line item at a time. And the quiet, decisive fact of that market is this: the outclassed player is the last to know, because being outclassed feels exactly like being served. Your tier is your bargaining power now. The bill for the cheap seat is real — it’s just written in a currency no scorecard prints: the deals you never saw you lost.