All articles
B.02 Infrastructure strategy

Baseload vs burst: the 80/20 of GPU infrastructure

Most of your GPU fleet is doing something predictable. Price it accordingly — and keep the cloud for everything else.

SEP 23, 2026/4 MIN READ

Try this with your own numbers: chart your team’s GPU-hour consumption week by week for the last quarter, from quietest week to busiest. For most production AI companies, the shape is unmistakable — a tall, steady floor with spikes stacked on top.

That floor is your baseload: the compute you’d need even if the quarter had gone perfectly average — production inference, always-on API capacity, the training runs baked into your weekly rhythm. The spikes are everything else: launches, experiments, capacity you’ve never seen before.

01 / THE POWER GRID FIGURED THIS OUT DECADES AGO

Electricity grids have run on exactly this logic for a century. Baseload plants — nuclear, hydro — generate cheap power constantly, 24 hours a day. Peaker plants sit idle until demand surges, then command premium rates precisely because they exist for flexibility. Nobody runs an industrial plant on peaker pricing and calls it a strategy.

GPU infrastructure is the same shape with worse labeling. On-demand cloud is peak pricing, charged around the clock, to workloads that were never going to shut off. If your quietest week still burns thousands of GPU-hours, you’re buying baseload capacity at peaker prices every single day.

02 / THE 80/20 RULE OF FLEETS

A typical production fleet lands near an 80/20 split: roughly four-fifths of compute is predictable baseload, one-fifth is variable burst. Lock in the predictable 80 at a committed rate and your entire cost structure shifts — while the flexible 20 stays on the cloud, exactly where elasticity belongs.

This is deliberately not all-or-nothing. The companies that hesitate about long-term infrastructure usually imagine committing 100% of future compute. The correct move is narrower: commit the portion you’d pay for anyway, and stop paying the flexibility premium on capacity that was never going to be flexible.

03 / HOW TO FIND YOUR BASELOAD

Look at the last 90 days of GPU-hours per week. Sort ascending. The average of the lowest quartile — or more conservatively, the 10th-percentile week — is a defensible estimate of your baseload. That’s the compute that is, functionally, permanent. Everything above it is your burst budget.

Two sanity checks: if your baseload estimate keeps rising month over month, that’s demand growth you can actually plan around. And if no week in the last quarter dropped below the estimate, you can trust the number the way finance trusts it — as a line item, not a guess.

Compute has the same profile as electricity: mostly steady, sometimes spiky. Price the steady part like it’s steady. Keep the cloud for the storms. The 80/20 split isn’t a theory — it’s what your utilization chart already says, if you read it honestly.