For AI teams with sustained GPU demand: NVIDIA B200 / B300 capacity, 12–24-month terms, and a ~$5/GPU-hour target for qualifying deployments. Price is indicative; availability and final terms are deployment-specific.
Reserve a predictable baseline; keep cloud for bursts.

Best fit: 24/7-class workloads, typically $50K+ in monthly GPU spend, and a baseline you can forecast for 12–24 months.
Reserve capacity for production inference and forecast spend with greater confidence.
Manage the compute cost of high-volume video generation.
Plan dedicated capacity for sustained audio and multimodal workloads.
Support recurring fine-tuning, RL, distillation and evaluation workloads.
Plan training capacity without relying solely on on-demand availability.
B200 / B300 capacity is offered for qualifying deployments. Availability is subject to confirmation; this page does not publish live inventory by region or standard provisioning lead times.

At 24/7 usage, a $10 vs. $5/GPU-hour comparison halves the hourly-rate cost. At lower on-demand utilization, the committed fleet may still be billed for every reserved hour—model both assumptions below.
At 32 GPUs running 24/7, a $5 vs. $10/GPU-hour scenario implies an illustrative ~$1.4M annual difference in GPU spend—before other infrastructure costs.
*Illustrative only: $10 vs. $5/GPU-hour. Actual pricing depends on GPU, configuration, quantity, region and term. The ~$5 target applies only to qualifying deployments; it is not a quote.
Compare current on-demand spend with reserved capacity billed for every reserved hour. This is an estimate, not a quote.
Model: on-demand cost = GPUs × 8,760 × utilization × current rate. Commitment cost = GPUs × 8,760 × committed rate (all reserved hours billed). The public page does not confirm actual utilization billing; confirm invoicing and included services in a written quote.
The offer describes a 12–24-month commitment. The public page does not state reduction, cancellation or ramp-down rights; have those terms written into the agreement before signing.
Request a dated, region-specific quote naming GPU model and count, configuration, and provisioning date. The page does not publish live inventory or a standard lead time.
Payment cadence and inclusions are not published. Ask the quote to itemize compute, networking, storage, egress, support, setup fees, invoice timing and any minimums.
These are deployment-specific in the available copy. Require the uptime target, response model, support coverage and remedies to be stated in the contract.
Yes. The intended approach is to commit only the predictable baseline and retain cloud for spikes, experiments and workloads you cannot forecast.
Share the basics. We’ll confirm availability and follow up with a deployment-specific quote.