All articles
B.01 Unit economics

Your GPU bill is part of your product economics

Inference compute isn’t overhead — it’s cost of goods sold. Once you see it that way, everything about how you buy GPUs changes.

SEP 30, 2026/4 MIN READ

Here’s a question most AI infrastructure buyers never ask directly: is your GPU bill a cost center, or is it cost of goods sold? Read that again, because the answer determines what you’re allowed to spend on it — and how you should negotiate it.

If GPUs are overhead — like office internet or developer laptops — you minimize them and move on. But if GPUs manufacture your product — if every API response, every generated frame, every inference call consumes GPU time — then compute sits inside your gross margin. It is one of the primary things that decides what you can charge, and what you get to keep.

01 / THE MARGIN EQUATION

Classic SaaS companies run gross margins around 70–80% precisely because the marginal cost of serving one more customer is near zero. AI companies don’t get that deal: every additional inference call has a real, measurable GPU cost attached.

That means unit compute economics compound exactly like gross margin. Consider an API startup running 100 GPUs for production inference, continuously:

A $4.38M/year swing on the same fleet, same utilization, same product — the only variable is the hourly rate. That’s the kind of number that changes a pricing meeting. It’s also, incidentally, a line item management can actually defend.

And it scales both directions. At 1,000 GPUs, the illustrative difference crosses $40M a year. At that point the compute bill isn’t an infrastructure detail — it’s the single biggest lever on your company’s valuation.

02 / PRICING POWER STARTS AT THE UNIT LEVEL

When you price your product, you are implicitly pricing your compute. A per-token or per-generation price carries a GPU-cost floor; if that floor drops by roughly half, you can lower prices to win deals, hold prices to expand margin, or split the difference. Companies that can’t move their compute cost negotiate their product price with one hand tied.

03 / WHY THIS CHANGES HOW YOU BUY COMPUTE

Treating GPUs as COGS reframes every infrastructure decision. Variable hourly compute becomes a planning problem: finance can’t model gross margin on a line item that swings with market availability. Committed, predictable compute becomes finance-friendly infrastructure — the number your CFO puts in the model is the number you pay.

It also flips how you read a utilization dashboard. Under on-demand pricing, high utilization is budget pressure. Under a committed model, high utilization is the machine working exactly as designed — the infrastructure cost is known, and every increment of usage improves unit economics.

The companies that will win the inference market won’t necessarily have the best models. They’ll be the ones whose compute economics let them scale without flinching at the bill. Your GPU spend isn’t a utility invoice to minimize — it’s line one of your product economics. Buy it accordingly.

Next article
Baseload vs burst: the 80/20 of GPU infrastructure

* Figures are illustrative examples based on 24/7 utilization and a $10/hr vs. $5/hr rate comparison. Actual pricing depends on GPU type, node configuration, commitment length, region, networking and deployment. $5/hr is a target rate for qualifying long-term commitments, not a universal price.