GPU capacity for AI

Scale AI without runaway GPU costs.

For AI teams with sustained GPU demand: NVIDIA B200 / B300 capacity, 12–24-month terms, and a ~$5/GPU-hour target for qualifying deployments. Price is indicative; availability and final terms are deployment-specific.

Reserve a predictable baseline; keep cloud for bursts.

A stylized ice fortress built from layered GPU server modules with ice crystals and cyan circuit glow
~$5/GPU-hour* target rate
12–24 month commitments
41%
Illustrative spend reduction: $10 vs. $5/GPU-hour at 85% on-demand utilization*
B200 / B300
NVIDIA B200 / B300 capacity
24/7
Dedicated capacity for production workloads
Buyer fit

For AI teams with a steady GPU baseline.

Best fit: 24/7-class workloads, typically $50K+ in monthly GPU spend, and a baseline you can forecast for 12–24 months.

LLM Inference

Reserve capacity for production inference and forecast spend with greater confidence.

Video Generation

Manage the compute cost of high-volume video generation.

Audio & Multimodal AI

Plan dedicated capacity for sustained audio and multimodal workloads.

Fine-Tuning & Distillation

Support recurring fine-tuning, RL, distillation and evaluation workloads.

Model Training

Plan training capacity without relying solely on on-demand availability.

Available capacity

NVIDIA B200 / B300 capacity, confirmed per deployment.

B200 / B300 capacity is offered for qualifying deployments. Availability is subject to confirmation; this page does not publish live inventory by region or standard provisioning lead times.

A frosted ice-blue GPU server module sliding out of a rack in a dark data hall with cyan circuit glow
NVIDIA

B200

Production AI workloads
  • LLM inference
  • Training
  • Fine-tuning
  • Distillation
  • High-throughput workloads
NVIDIA

B300

Next-generation AI workloads
  • Large-scale inference
  • Training
  • Reasoning workloads
  • Multimodal workloads
  • High-density deployments

Confirm these in the written offer

GPU model and reserved count
Node, memory and interconnect
Datacenter region and network scope
Provisioning date and availability window
Uptime SLA and support coverage
Included services, fees and payment schedule
The source page does not publish region-level inventory, delivery dates, or standard SLA/payment terms. Treat these as unconfirmed until documented for your deployment.
Economics

Model the cost of a committed GPU baseline.

At 24/7 usage, a $10 vs. $5/GPU-hour comparison halves the hourly-rate cost. At lower on-demand utilization, the committed fleet may still be billed for every reserved hour—model both assumptions below.

Typical on-demand rate
~$10
per GPU-hour; variable
ICE Castle target rate*
~$5
per GPU-hour on a qualifying commitment
One GPU at $10/hour, 24/7
$87,600
per year
One GPU at $5/hour, 24/7
$43,800
per year
Illustrative annual difference
~$43,800
per GPU / year
41%
Illustrative annual spend reduction at 85% on-demand utilization, with reserved hours billed in full.
Estimate depends on utilization and contract billing.

At 32 GPUs running 24/7, a $5 vs. $10/GPU-hour scenario implies an illustrative ~$1.4M annual difference in GPU spend—before other infrastructure costs.

*Illustrative only: $10 vs. $5/GPU-hour. Actual pricing depends on GPU, configuration, quantity, region and term. The ~$5 target applies only to qualifying deployments; it is not a quote.

Fleet cost estimator

Estimate annual spend and savings.

Compare current on-demand spend with reserved capacity billed for every reserved hour. This is an estimate, not a quote.

32
$10.00
85%
$5.00
On-demand spend at utilization$2,382,720
Committed cost at 100% reserved hours$1,401,600
Difference vs. on-demand$981,120

Model: on-demand cost = GPUs × 8,760 × utilization × current rate. Commitment cost = GPUs × 8,760 × committed rate (all reserved hours billed). The public page does not confirm actual utilization billing; confirm invoicing and included services in a written quote.

FAQ

Resolve the key procurement questions.

01

What if our GPU demand changes during the term?

The offer describes a 12–24-month commitment. The public page does not state reduction, cancellation or ramp-down rights; have those terms written into the agreement before signing.

02

How do we verify capacity and delivery?

Request a dated, region-specific quote naming GPU model and count, configuration, and provisioning date. The page does not publish live inventory or a standard lead time.

03

What is included, and when are we billed?

Payment cadence and inclusions are not published. Ask the quote to itemize compute, networking, storage, egress, support, setup fees, invoice timing and any minimums.

04

What SLA and support will we receive?

These are deployment-specific in the available copy. Require the uptime target, response model, support coverage and remedies to be stated in the contract.

05

Can we keep our existing cloud for burst demand?

Yes. The intended approach is to commit only the predictable baseline and retain cloud for spikes, experiments and workloads you cannot forecast.

Next step

Check capacity and pricing.

Share the basics. We’ll confirm availability and follow up with a deployment-specific quote.

Use your work email. Detailed spend information can be shared later.

We’ll confirm availability and follow up by email.