I Approved US$100,000 for AI Hardware.
Then I Didn’t Spend It.
AMD says its new data center chip delivers up to 18× lower cost per token than its own previous generation. Everyone read that as a story about NVIDIA. It isn’t. It is a story about every organisation that owns compute — including the one I nearly built for ourselves.
Two numbers came out of AMD’s keynote. Almost everyone quoted the wrong one.
On 23 July, AMD launched the Instinct MI400 series and the Helios rack built around it. The number that travelled was up to 30% more tokens per dollar than NVIDIA’s Vera Rubin NVL72 — AMD’s own estimate, on a workload AMD chose, using projected hardware pricing. It is a competitive claim by a vendor against a rival who gets to answer next year. Interesting; not decisive.
The number that matters was the other one: up to 18× lower cost per token than the MI355X, with up to 34× the throughput. The MI355X is not a competitor’s product. It is AMD’s own previous generation.
Read that comparison again. The threat to anyone holding AI compute is not the company across the street. It is last year’s version of the same company.
A question I asked in July, answered harder than I expected
Earlier this month I wrote that the building pays back but the silicon might not — that compute price-performance improves roughly 20–30% a year, so a two-year-old facility sells compute costing nearly twice what a new competitor pays, on the same power bill.
An 18× generational claim is not that curve. It is a different shape of problem. Twenty per cent a year is something an operator can outrun with utilisation, contracts and discipline. An order of magnitude in one generation cannot be outrun by management at all, because the gap is not operational — it is silicon.
Which forced me to re-examine a decision of my own, at a very different scale.
The US$100,000 I approved and never spent
I budgeted US$100,000 for an internal inference cluster — Mac Studios with 512GB of unified memory, roughly ten machines at list price. Not for a product. For our own team’s daily work.
The case was good, and I still think it was good. Fixed cost instead of per-token billing. Data that never leaves the office. No dependency on a provider’s pricing decisions, no surprise rate card, no vendor quietly deciding our margins for us.
I could not justify the throughput. The money is still unspent.
The first comparison is the obvious one. That same US$100,000, spent on rented inference, buys somewhere between 10 and 200 billion tokens depending on which tier of model you use. Even the pessimistic end of that range is far more than a company our size consumes in a year — and it arrives with no procurement, no rack, no cooling, and nobody on staff keeping it alive.
But that is not what killed it.
The number that actually killed it: idle time
An office is bursty. We work roughly 2,100 hours a year out of 8,760. Even if every machine were saturated during every working hour — which never happens, because demand arrives in spikes around a few people’s sessions — the cluster sits idle about three-quarters of the year.
Owned hardware bills you for idle time. Rented tokens don’t. A tokens-per-second figure assumes a machine running flat out. Divide by a realistic duty cycle and the effective cost per delivered token is several times worse than the number that justified the purchase.
This is the part no spec sheet shows you, and it is why so much in-house AI hardware disappoints. For a service running 24/7 at scale, ownership can absolutely win — utilisation is the whole game, and at scale you have it. For internal company use, the machine spends most of its life as an expensive space heater.
Then add the 18×. Even if I had made the utilisation work, I would be holding an asset whose successor makes it an order of magnitude less efficient — while the rented alternative gets cheaper every time any vendor ships. That is not two risks stacked; it is the same risk arriving twice.
Eighteen months of payback on a machine that might be beaten in twelve is not an investment. It is a subscription with worse terms: you pay everything upfront, and you cannot cancel.
And “beaten in twelve” isn’t hypothetical for the exact machine I was pricing. As I write this, Bloomberg’s Mark Gurman reports that Apple is set to skip the M5 entirely for its next Mac mini and iMac and jump straight to the M6 — and that it’s deprioritising the high-end M6 chips to rush the M7 generation. The high end is the class of silicon a Mac Studio runs on. So the cluster I nearly bought wouldn’t just have faced a normal yearly step-down; its whole generation was about to be leapfrogged by the vendor itself. When the manufacturer is compressing its own upgrade cycle, “wait” stops being caution and starts being arithmetic.
The rule I use now
“Never own hardware” is the opposite mistake. It hands your cost structure to someone else, and anyone who has watched an API price change knows that is its own exposure.
Buy compute you can amortise in months. Rent compute that is improving faster than you can depreciate it.
In practice that means we own the small, boring, always-on infrastructure — the things whose requirements do not change and whose payback is measured in months. And we rent frontier inference, where an entire industry is competing to push my cost line down. I would rather be the beneficiary of that competition than a stranded owner on the wrong side of it.
This isn't just my RM100k — the credit market just made the same call at $167 billion
While I was declining to spend a hundred thousand dollars, Oracle was doing the opposite at a scale that’s hard to picture — and on 9 July 2026, it got graded for it. S&P Global cut Oracle’s credit rating from BBB to BBB−, one notch above junk. Not for a bad quarter. For its AI infrastructure buildout.
The numbers behind the downgrade are the same argument as this article, three orders of magnitude larger: capital expenditure projected at $90–95 billion for fiscal 2027 (up from a $60 billion forecast), $167 billion in total debt, free cash flow heading toward negative $42 billion, all funding a roughly $250 billion data-centre buildout — with about half of its committed future revenue tied to a single customer, OpenAI.
Strip away the zeros and it’s my decision inverted. Oracle is buying the depreciating layer — the same silicon that gets an order of magnitude cheaper per generation, from the top of this piece — and borrowing heavily to own it. When you owe $167 billion against assets whose efficiency the next chip cuts by multiples, the lender notices. The equity market shrugged; the credit market, which cares about getting paid back rather than about the story, did not.
I’m not calling Oracle’s bet wrong — at hyperscale, with utilisation they can actually achieve, owning may pay off, and they’re making that wager with open eyes. The point is narrower: the same posture that felt risky to me at RM100,000 is rated risky at $167 billion, for the same reason. And here is the asset Oracle no longer has that I did — the ability to not commit. I could leave the money unspent. A company mid-buildout, funded by debt, cannot. Optionality is the cheapest thing to own in an accelerating market, and it never appears on anyone’s balance sheet.
Two honest caveats
18× is a vendor’s headline. It is an “up to” figure, on a workload the vendor selected, measured against the vendor’s own older product. Real-world gains will be smaller and workload-dependent. I trust the direction, not the magnitude — and the direction is enough to change a decision.
Cost per token is not the only reason to own a machine. Privacy, latency, sovereignty, and the ability to keep working when someone else’s service is down are all real, and none of them appear on a tokens-per-dollar chart. If those are why you are buying, this argument does not apply to you. Just say so out loud — so you are not accidentally defending a privacy decision with economics that will not hold.
If you never buy a GPU, this still applies to you
Most Malaysian businesses will never make a decision about a data center chip. Nearly all of them are about to make the same decision in smaller form: buy AI capacity, or rent it. A server under the stairs. A licence bought for three years. A platform paid annually because the discount looked good.
The pattern holds at every scale. If the thing you are buying improves by an order of magnitude per generation, ownership is a liability dressed as an asset. If it is stable and boring, ownership is fine — and usually cheaper.
The trap is not buying too early. It is buying early and then defending the purchase — running last year’s economics because you already paid, while a competitor who committed to nothing quietly runs this year’s.
The bottom line
I do not regret budgeting the US$100,000. Doing the work is what produced the answer. What I would regret is having spent it to justify the budget — which is how most of these purchases actually happen.
Published by IMA AI — July 2026. We run our own infrastructure decisions on the same math we publish — and this one ended with the budget unspent.