Cost Optimization Just Got a Second Author

Cost Optimization Just Got a Second Author
A model rewrote its own GPU serving code, cut OpenAI’s inference cost, and got itself a price cut in the field — while nobody prompted it to think about money.
That’s not a caching trick. That’s not a router sending easy questions to a cheaper tier. That’s the thing traditionally done to a model — the entire FinOps-for-AI playbook — being done by one, autonomously, at the infrastructure layer. And it landed in the same week I was already forwarding myself an open-source tool that does the mirror-image job on AWS: scan the account, find the waste, hand you a ranked list with a dollar figure attached, no ticket, no Slack thread, no waiting on next quarter’s audit. Cost optimization didn’t get better tooling. It got a second author, on both sides of the meter, in the same seven days.