Skip to content
AI Tools · GPU Cloud and AI Compute

Top 7 GPU Cloud and AI Compute Providers 2026: Price Per GPU-Hour by Chip Class, Verified From Each Provider's Own Pricing Page

CoreWeave, Lambda, Together, RunPod, Vast.ai, Modal and Crusoe compared on published price per GPU-hour for H100, H200, B200 and Blackwell Ultra, plus commitment terms, storage, egress and availability. Every figure read from the provider's own page on 2026-09-18.

By ·Sep 18, 2026·16 min·7 tools compared
GPU CloudAI ComputeCoreWeaveLambdaRunPodModalCrusoeVast.aiH100B200AI Tools

The answer, before the table

Price per GPU-hour by chip class is the whole decision, so start there. Every figure below was read from the provider's own pricing page on 18 September 2026.

  • Cheapest H100 you can actually buy: Together preemptible at $1.99 per GPU-hour, or RunPod Community PCIe at $1.99. Both trade away guaranteed availability or vetted hosting.
  • Cheapest dedicated-provider on-demand H100: Crusoe at $3.90, then Modal at an effective $3.95, then Lambda and Together at $3.99.
  • Cheapest published B200: Lambda at $6.69 per GPU-hour on 8x instances.
  • Blackwell Ultra self-serve: RunPod B300 at $6.94 Community and $7.89 Secure, and Modal B300 at an effective $7.10. Almost everyone else says contact sales.
  • Scale and rack-scale systems: CoreWeave, which publishes GB200 NVL72 at $42.00 per hour for 4 GPUs and charges nothing for egress.
  • Lowest floor of all, with the least certainty: Vast.ai, a marketplace that publishes no prices at all.

A three-times spread for the same silicon is normal here. It is explained by availability guarantees, interconnect, hosting model and egress policy, not by the chip.

What changed with current-generation accelerators

Three shifts matter for a 2026 purchase.

Blackwell set the new price ceiling, and H100 became the commodity floor. B200 runs $5.98 to $8.60 per GPU-hour across these providers against $1.99 to $6.16 for H100. Blackwell Ultra (B300, GB300) is $6.94 to $7.89 on RunPod and an effective $7.10 on Modal, while Fireworks publishes $15.00 for B300 and $20.00 for GB300 on its inference-oriented GPU tier. If your workload runs fine on H100, this is the best time in three years to buy it.

The unit of sale changed at the top end. CoreWeave sells GB200 NVL72 at $42.00 per hour for 4 GPUs, because a rack-scale NVL system is the product rather than an individual accelerator. HGX lines are sold by the 8-GPU node. Per-GPU figures for CoreWeave in this comparison are derived by division and are honest for comparison, not for purchase.

The newest parts retreated behind sales. HGX B300 and GB300 NVL72 are contact sales on CoreWeave. B200, GB200 NVL72 and MI355X are contact sales on Crusoe, as is every Crusoe spot rate. GB200 NVL72 and B300 are contact sales on Together. A buyer comparing published prices in 2026 is largely comparing last-generation economics.

Where the hourly rate stops being the real cost

This is the part that decides the invoice, and none of it is in the headline number.

Cost Who charges nothing Who charges, and how much
Egress CoreWeave (free transfer), Crusoe ("does not charge for network ingress or egress"), Lambda ("no egress fees") Together, Modal and Vast.ai publish no policy. Hyperscalers meter it
Idle storage CoreWeave cold object storage at $0.015/GB/month RunPod volume disk at $0.20/GB/month while idle, double the $0.10 running rate
Minimum purchase Lambda (1 GPU), RunPod (1 GPU), Modal (per second), Vast.ai (per second) CoreWeave's smallest HGX H100 purchase is a $49.24/hour 8-GPU node
Idle GPU time Modal, which scales to zero Every hourly rental, where a four-hour-a-day endpoint pays for 24

On a pipeline moving 50 TB a month, egress alone is a five-figure annual difference between two providers whose hourly rates look identical.

How the hyperscalers price against them

AWS publishes per-accelerator-hour rates through EC2 Capacity Blocks for ML: $5.191 for P5 (H100), $5.97 for P5e (H200), $10.582 for P6e (B200), $14.04 for P6-B300 and $0.596 for Trn1 (Trainium). The reservation fee is charged up front when you schedule, with an operating-system fee on top for non-Linux images. AWS's own P5 instance page publishes no on-demand hourly price at all, which tells you something about how hyperscaler GPU pricing is meant to be consumed.

Those rates sit at or above the top of the specialist range, before metered egress. What the premium buys is everything around the GPU: IAM, VPC, the compliance attestations your auditor already accepts, and capacity you can add without a new vendor review. For a regulated buyer that is frequently worth it. For a team whose requirement is simply GPUs, it is not.

Quick Comparison

ProviderH100 (per GPU-hour)H200 (per GPU-hour)B200 (per GPU-hour)Blackwell Ultra (B300 / GB300)EgressBilling Granularity and Commitment
CoreWeave$6.16 (HGX H100 8-GPU node at $49.24/hr)$6.31 (HGX H200 8-GPU node at $50.44/hr)$8.60 (HGX B200 8-GPU node at $68.80/hr)HGX B300 and GB300 NVL72 are contact sales. GB200 NVL72 is $42.00/hr for 4 GPUs, $10.50 per GPU-hourFree, for internet egress and internal transfersSold by node, not by GPU. No minimum commitment stated in the on-demand pricing tables
Lambda$3.99 on 8x SXM, $4.29 on 1x SXM, $3.29 on 1x PCIeNot listed in the on-demand table$6.69 on 8x, $6.99 on 1xNot listed"Transparent pricing with no egress fees"Per-GPU on-demand from a single card. 1-Click Clusters run 2 weeks to 1 year
Together AI$3.99 on-demand, $3.19 to $3.69 reserved, $1.99 preemptible$5.99 on-demand, $3.99 to $4.99 reserved, $2.99 preemptible$8.19 on-demand, $6.79 to $7.99 reserved, $4.09 preemptibleGB200 NVL72 and B300 are contact salesNot published on the pricing pageReserved terms run 7 to 180+ days. Preemptible for fault-tolerant work
RunPod$2.69 Community, $3.49 Secure (SXM). PCIe $1.99 Community, $2.89 Secure$3.59 Community, $4.59 Secure$5.98 Community, $6.79 SecureB300 $6.94 Community, $7.89 SecureNo egress charge listed in the pricing documentationPer-GPU, no minimum. Two tiers: Community (third-party hosts) and Secure (vetted data centres)
Vast.aiMarketplace rate, not publishedMarketplace rate, not publishedMarketplace rate, not publishedMarketplace rate, not publishedNot published on the pricing pagePer-second billing with no minimum or rounding. On-demand, interruptible ("50%+ cheaper") and reserved for 1, 3 or 6 months ("up to 50% off")
Modal$3.95 (H100 SXM5 at $0.001097/second)$4.54 ($0.001261/second)$6.25 ($0.001736/second)B300 $7.10 ($0.001972/second)Not published on the pricing pagePer-second billing with scale to zero. Starter $0 + $30/month included compute, Team $250/month + $100/month included compute
Crusoe$3.90 (HGX, 80GB)$4.29 (HGX, 141GB)Contact salesGB200 NVL72 and MI355X are contact sales"Crusoe Cloud does not charge for network ingress or egress"Per-GPU on-demand published. Spot pricing and reserved terms are contact sales

CoreWeave

H100 (per GPU-hour)
$6.16 (HGX H100 8-GPU node at $49.24/hr)
H200 (per GPU-hour)
$6.31 (HGX H200 8-GPU node at $50.44/hr)
B200 (per GPU-hour)
$8.60 (HGX B200 8-GPU node at $68.80/hr)
Blackwell Ultra (B300 / GB300)
HGX B300 and GB300 NVL72 are contact sales. GB200 NVL72 is $42.00/hr for 4 GPUs, $10.50 per GPU-hour
Egress
Free, for internet egress and internal transfers
Billing Granularity and Commitment
Sold by node, not by GPU. No minimum commitment stated in the on-demand pricing tables

Lambda

H100 (per GPU-hour)
$3.99 on 8x SXM, $4.29 on 1x SXM, $3.29 on 1x PCIe
H200 (per GPU-hour)
Not listed in the on-demand table
B200 (per GPU-hour)
$6.69 on 8x, $6.99 on 1x
Blackwell Ultra (B300 / GB300)
Not listed
Egress
"Transparent pricing with no egress fees"
Billing Granularity and Commitment
Per-GPU on-demand from a single card. 1-Click Clusters run 2 weeks to 1 year

Together AI

H100 (per GPU-hour)
$3.99 on-demand, $3.19 to $3.69 reserved, $1.99 preemptible
H200 (per GPU-hour)
$5.99 on-demand, $3.99 to $4.99 reserved, $2.99 preemptible
B200 (per GPU-hour)
$8.19 on-demand, $6.79 to $7.99 reserved, $4.09 preemptible
Blackwell Ultra (B300 / GB300)
GB200 NVL72 and B300 are contact sales
Egress
Not published on the pricing page
Billing Granularity and Commitment
Reserved terms run 7 to 180+ days. Preemptible for fault-tolerant work

RunPod

H100 (per GPU-hour)
$2.69 Community, $3.49 Secure (SXM). PCIe $1.99 Community, $2.89 Secure
H200 (per GPU-hour)
$3.59 Community, $4.59 Secure
B200 (per GPU-hour)
$5.98 Community, $6.79 Secure
Blackwell Ultra (B300 / GB300)
B300 $6.94 Community, $7.89 Secure
Egress
No egress charge listed in the pricing documentation
Billing Granularity and Commitment
Per-GPU, no minimum. Two tiers: Community (third-party hosts) and Secure (vetted data centres)

Vast.ai

H100 (per GPU-hour)
Marketplace rate, not published
H200 (per GPU-hour)
Marketplace rate, not published
B200 (per GPU-hour)
Marketplace rate, not published
Blackwell Ultra (B300 / GB300)
Marketplace rate, not published
Egress
Not published on the pricing page
Billing Granularity and Commitment
Per-second billing with no minimum or rounding. On-demand, interruptible ("50%+ cheaper") and reserved for 1, 3 or 6 months ("up to 50% off")

Modal

H100 (per GPU-hour)
$3.95 (H100 SXM5 at $0.001097/second)
H200 (per GPU-hour)
$4.54 ($0.001261/second)
B200 (per GPU-hour)
$6.25 ($0.001736/second)
Blackwell Ultra (B300 / GB300)
B300 $7.10 ($0.001972/second)
Egress
Not published on the pricing page
Billing Granularity and Commitment
Per-second billing with scale to zero. Starter $0 + $30/month included compute, Team $250/month + $100/month included compute

Crusoe

H100 (per GPU-hour)
$3.90 (HGX, 80GB)
H200 (per GPU-hour)
$4.29 (HGX, 141GB)
B200 (per GPU-hour)
Contact sales
Blackwell Ultra (B300 / GB300)
GB200 NVL72 and MI355X are contact sales
Egress
"Crusoe Cloud does not charge for network ingress or egress"
Billing Granularity and Commitment
Per-GPU on-demand published. Spot pricing and reserved terms are contact sales
1

CoreWeave

Best for Enterprise

Best for: Multi-node training at scale, rack-scale Blackwell systems, and teams that move terabytes in and out

CoreWeave is the largest of the specialist GPU clouds and the one that publishes the most complete price list, including rack-scale GB200 NVL72 capacity that most competitors will only quote. It sells by the node rather than the GPU, which rules it out for single-card experimentation and makes it the right shape for real multi-node training. Free egress and $0.015 to $0.07 per GB per month storage make the total bill predictable in a way hyperscaler pricing is not.

Pros

  • Published node pricing across the full range: HGX H100 $49.24/hr, HGX H200 $50.44/hr, HGX B200 $68.80/hr, A100 $21.60/hr, L40S $18.00/hr, RTX PRO 6000 Blackwell high-memory $20.00/hr, GH200 $6.50/hr for a single GPU
  • GB200 NVL72 is published at $42.00 per hour for 4 GPUs, which is $10.50 per GPU-hour, rather than hidden behind sales
  • Internet egress and internal transfers are free, which removes the single largest hidden cost of moving a training pipeline to a cloud
  • Storage is cheap and tiered: object storage at $0.06 hot, $0.03 warm and $0.015 cold per GB per month, distributed file storage at $0.070 per GB per month
  • No minimum commitment appears in the on-demand pricing tables

Cons

  • GPUs are sold in 8-GPU nodes for the HGX lines, so there is no cheap way to run one card
  • HGX B300 and GB300 NVL72 are contact sales, so the current top of the range is not self-serve
  • Public IP addresses are billed separately at $4.00 per IP per month
  • Per-GPU H100 pricing of $6.16 is higher than Crusoe at $3.90, Lambda at $3.99 and RunPod Community at $2.69
  • Now also owns Weights & Biases, so a team buying both compute and ML tooling may be concentrating vendor risk
Honest Weakness: On raw H100 price per GPU-hour, CoreWeave is among the most expensive providers here, roughly 58% above Crusoe's published $3.90 and more than double RunPod's Community tier. Node-only sales make that worse for anyone whose job does not need eight GPUs, because the minimum purchase is $49.24 an hour. CoreWeave earns its price on fabric, capacity at scale and free egress, none of which matter to a team running single-GPU fine-tunes. If that is your workload, you are paying for infrastructure you will not use.

Why node pricing changes the comparison

Every per-GPU figure for CoreWeave in the table above is derived by dividing the published node price by the GPU count, because that is how the capacity is sold. It is an honest comparison number and a misleading purchase number. You cannot buy $6.16 of CoreWeave H100. The smallest unit is $49.24 an hour, which is roughly $35,500 a month if left running.

Free egress is the underrated line item

Training pipelines move data constantly: datasets in, checkpoints out, evaluation artefacts back and forth. Hyperscalers meter that. CoreWeave, Crusoe and Lambda do not. On a pipeline shifting 50 TB a month, hyperscaler egress at typical published rates is a five-figure annual line that simply does not exist on these providers.

Ownership and consolidation

CoreWeave completed its acquisition of Weights & Biases on 5 May 2025. If you are also evaluating W&B Weave for evaluation and tracing, note that you would be buying compute and tooling from the same company.

North America on-demand, per node: GB200 NVL72 (4 GPUs) $42.00/hr, HGX B200 (8 GPUs) $68.80/hr, HGX H200 (8 GPUs) $50.44/hr, HGX H100 (8 GPUs) $49.24/hr, RTX PRO 6000 Blackwell high-memory (8 GPUs) $20.00/hr, A100 (8 GPUs) $21.60/hr, L40S (8 GPUs) $18.00/hr, L40 (8 GPUs) $10.00/hr, GH200 (1 GPU) $6.50/hr. GB300 NVL72 and HGX B300 are contact sales. Storage: object hot $0.06, warm $0.03, cold $0.015 per GB per month; distributed file storage $0.070 per GB per month. Data transfer free. Public IP $4.00 per IP per month. Checked on coreweave.com/pricing, 2026-09-18.

Visit CoreWeave
2

Lambda

Best Overall

Best for: Clear per-GPU on-demand pricing from a single card upward, with no egress fees

Lambda publishes the cleanest per-GPU price list in this comparison and lets you rent a single card. H100 SXM is $3.99 per GPU-hour on an 8x instance and $4.29 on a 1x, and B200 is $6.69 on 8x and $6.99 on 1x. The narrow gap between one GPU and eight is unusual and valuable: you can prototype on one card at almost the same unit rate you will pay in production. Lambda's own page states there are no egress fees.

Pros

  • Per-GPU on-demand pricing from a single card, with a small premium for smaller configurations
  • B200 at $6.69 per GPU-hour on 8x instances is the cheapest published on-demand Blackwell rate among the dedicated providers here
  • H100 SXM at $3.99 per GPU-hour on 8x undercuts CoreWeave's derived $6.16 substantially
  • Wide older-generation catalogue for cheap experimentation: A100 40GB $1.99, A10 $1.29, A6000 $1.09, V100 $0.79, Quadro RTX 6000 $0.69
  • States transparent pricing with no egress fees
  • 1-Click Clusters give reserved multi-node capacity on terms from 2 weeks to 1 year without an enterprise contract

Cons

  • 1-Click Cluster rates are higher than on-demand: H100 runs $5.54 to $6.16 per GPU-hour and B200 runs $8.87 to $9.86, above the $3.99 and $6.69 on-demand rates
  • H200 does not appear in the published on-demand instance table
  • No Blackwell Ultra (B300 or GB300) capacity is listed
  • The pricing page does not publish a persistent storage rate, stating only that rates appear when you create a filesystem
  • On-demand availability for current-generation parts is not guaranteed, and the page does not publish capacity commitments
Honest Weakness: Lambda's reserved pricing is higher than its on-demand pricing, which reads like an error and is not one. A 1-Click Cluster is a different product from an on-demand instance: dedicated interconnect, guaranteed capacity for the term, multi-node fabric. Buyers who assume committing saves money will commit to a worse rate. Decide first whether you need guaranteed multi-node capacity with a fast fabric. If you do, the premium is the price of certainty. If you do not, stay on demand.

The one-to-eight GPU gap

On most providers, renting one GPU carries a large premium over renting a full node. Lambda charges $4.29 for one H100 against $3.99 per GPU on eight, a 7.5% difference. That makes it the most sensible place to develop a training script before scaling it, because your cost model does not change shape when you scale.

Older generations are still the right answer often

A100 40GB at $1.99 and A10 at $1.29 remain the correct hardware for LoRA fine-tuning small models, for embedding workloads and for anything memory-light. Pairing those rates with an efficiency library that cuts VRAM sharply is how a fine-tuning programme stays inexpensive.

On-demand per GPU-hour: B200 SXM6 $6.69 (8x), $6.79 (4x), $6.89 (2x), $6.99 (1x). H100 SXM $3.99 (8x), $4.09 (4x), $4.19 (2x), $4.29 (1x). H100 PCIe $3.29 (1x). GH200 $2.29 (1x). A100 SXM 80GB $2.79 (8x). A100 SXM 40GB $1.99 (1x). A100 PCIe 40GB $1.99 (4x and 2x). A10 $1.29. A6000 $1.09. V100 $0.79 (8x). Quadro RTX 6000 $0.69. 1-Click Clusters, 2 weeks to 1 year: B200 $9.86 (16 GPUs), $9.36 (64), $8.87 (256+); H100 $6.16 (16), $5.85 (64), $5.54 (256). Commitments beyond 1 year are custom. The page states no egress fees and does not publish a storage rate. Checked on lambda.ai/pricing and lambda.ai/service/gpu-cloud, 2026-09-18.

Visit Lambda
3

Together AI GPU Clusters

Runner Up

Best for: Fault-tolerant training that can take preemption, and teams that also want inference on the same account

Together publishes three prices per chip: on-demand, reserved and preemptible. That third column is the reason to look. H100 preemptible at $1.99 per GPU-hour is roughly half its own on-demand rate and among the lowest published figures anywhere for that chip. If your training loop checkpoints properly, you are paying an availability premium you do not need. The same account also covers inference and fine-tuning, which removes a vendor from the stack.

Pros

  • Three published price tiers per chip, so the preemption trade-off is explicit rather than a sales conversation
  • Preemptible rates are aggressive: H100 $1.99, H200 $2.99, B200 $4.09 per GPU-hour
  • Reserved terms start at 7 days, far shorter than the multi-month commitments most providers require: H100 $3.19 to $3.69, H200 $3.99 to $4.99, B200 $6.79 to $7.99
  • On-demand H100 at $3.99 matches Lambda and undercuts CoreWeave's derived $6.16
  • The same platform covers dedicated inference endpoints and fine-tuning, so a training-to-serving pipeline stays with one vendor

Cons

  • GB200 NVL72 and B300 are contact sales, so the current top of the range is not self-serve
  • The pricing page does not publish storage rates or an egress policy, both of which are decisive at scale
  • Preemptible capacity is only usable if your training code checkpoints and resumes cleanly, which is a real engineering prerequisite
  • Dedicated inference on HGX H100 is advertised at $3.99 against a $5.49 list price under a promotion dated to 09/30/26, so the durable rate is the higher one
Honest Weakness: Together does not publish storage or egress terms on its pricing page, and those are exactly the costs this comparison exists to expose. CoreWeave and Crusoe state free egress in writing; Lambda states no egress fees; Together says nothing either way. For a workload moving large checkpoints and datasets, an unstated egress policy is a material unknown. Get it in writing before you compare its headline $1.99 preemptible rate against a provider that has already told you transfer is free.

Preemptible capacity is the real product

The gap between $3.99 and $1.99 on H100 is the whole argument. Training is the canonical preemptible workload: it is long-running, it checkpoints naturally, and a restart costs minutes rather than a customer. If your training script cannot resume from a checkpoint, fixing that is a day of work that halves your compute bill permanently.

Seven-day reserved terms

Most reserved GPU capacity is sold in months or years. A 7-day reservation is short enough to cover a single training campaign, which is a genuinely different commercial shape and worth knowing about when a project needs guaranteed capacity for a fortnight and nothing after.

GPU Clusters per GPU-hour. On-demand: HGX H100 $3.99, HGX H200 $5.99, HGX B200 $8.19. Reserved, 7 to 180+ days: H100 $3.19 to $3.69, H200 $3.99 to $4.99, B200 $6.79 to $7.99, with longer commitments yielding better rates. Preemptible: H100 $1.99, H200 $2.99, B200 $4.09. GB200 NVL72 and B300 are contact sales. Dedicated inference on-demand: HGX H100 $3.99 promotional through 09/30/26 against a $5.49 list price, HGX B200 $8.99; H200 and B300 are contact sales. Storage and egress are not published on the pricing page. Checked on together.ai/pricing, 2026-09-18.

Visit Together AI GPU Clusters
4

RunPod

Best Value

Best for: The cheapest per-GPU rates available self-serve, when you can accept third-party hosting for part of the work

RunPod splits its catalogue into Community Cloud, where capacity comes from third-party hosts, and Secure Cloud, in vetted data centres, and publishes both prices side by side. That transparency is the product. H100 SXM is $2.69 on Community and $3.49 on Secure, and even the Secure tier undercuts most dedicated providers. The consumer-card catalogue (RTX 5090 at $0.69, RTX 4090 at $0.34, RTX A5000 at $0.16) makes it the cheapest place to iterate.

Pros

  • Lowest published self-serve rates for current chips: H100 PCIe at $1.99 Community, H100 SXM at $2.69, H200 at $3.59, B200 at $5.98, B300 at $6.94
  • Secure Cloud remains competitive at $3.49 for H100 SXM while running in vetted data centres
  • Deep consumer and workstation catalogue for cheap iteration: RTX 5090 $0.69, RTX 4090 $0.34, RTX A5000 $0.16, A40 $0.35, RTX A6000 $0.33 on Community
  • B300 is available self-serve at $6.94 Community and $7.89 Secure, while several dedicated providers only quote Blackwell Ultra
  • Storage rates are published in full: container disk $0.10/GB/month, network storage $0.07/GB/month under 1 TB and $0.05 over, high-performance network storage $0.14/GB/month
  • No egress charge appears in the pricing documentation

Cons

  • Idle volume storage costs $0.20/GB/month, double the $0.10 charged while running, which penalises exactly the pattern of stopping a pod between sessions
  • Community Cloud capacity is hosted by third parties, which is a real compliance and data-handling question, not a formality
  • Availability of any specific GPU in any specific region varies, and the pricing page is not an availability guarantee
  • Multi-node training with a high-performance interconnect is not the platform's strength compared with CoreWeave or Lambda clusters
Honest Weakness: The idle volume charge inverts the incentive the platform otherwise creates. Volume disk is $0.10 per GB per month while a pod runs and $0.20 per GB per month while it is stopped, so a 500 GB dataset left attached between sessions costs $100 a month to do nothing. Teams adopt RunPod to avoid paying for idle GPUs and then quietly pay for idle disk instead. Move datasets to network storage at $0.07 or $0.05 per GB per month, or off the platform entirely, between campaigns.

Community versus Secure is a compliance decision

Community Cloud capacity is supplied by third-party hosts. For public datasets, open-weight models and experimentation, the roughly 25% saving is free money. For customer data, regulated workloads or anything under a data-processing agreement, Secure Cloud is the only defensible choice, and at $3.49 per H100-hour it is still cheaper than most dedicated providers.

Where consumer cards actually win

An RTX 4090 at $0.34 an hour has 24 GB of memory. Paired with a memory-efficient training library, that is enough for LoRA fine-tuning on small and mid-size models, for embedding generation and for most inference experimentation. A large share of work that gets run on an H100 does not need one.

Per hour, Community / Secure. B300 $6.94 / $7.89. B200 $5.98 / $6.79. H200 $3.59 / $4.59. H100 NVL $2.59 / $3.19. H100 SXM $2.69 / $3.49. H100 PCIe $1.99 / $2.89. RTX PRO 6000 $1.69 / $2.09. A100 SXM $1.39 / $1.59. A100 PCIe $1.19 / $1.59. L40S $0.79 / $1.09. RTX 6000 Ada $0.74 / $0.84. L40 $0.69 / $0.82. RTX 5090 $0.69 / $0.99. L4 $0.44 / $0.49. A40 $0.35 / $0.49. RTX 4090 $0.34 / $0.74. RTX A6000 $0.33 / $0.53. RTX 3090 $0.22 / $0.50. RTX A5000 $0.16 / $0.27. Storage: container disk $0.10/GB/month, volume disk $0.10/GB/month running and $0.20/GB/month idle, network storage $0.07/GB/month under 1 TB and $0.05/GB/month over, high-performance network storage $0.14/GB/month. No egress charge listed. Checked on runpod.io/pricing, 2026-09-18.

Visit RunPod
5

Modal

Fastest

Best for: Bursty and intermittent workloads that should cost nothing when idle

Modal is a serverless compute platform rather than a GPU rental service, and it prices accordingly: per second, with scale to zero. H100 SXM5 is $0.001097 per second, which is $3.95 per GPU-hour, and B300 is $0.001972 per second, which is $7.10. The reason to choose it is not the hourly rate, it is that an endpoint handling traffic for four hours a day costs a sixth of what an always-on instance does, with no orchestration work from you.

Pros

  • Per-second billing with scale to zero, so idle capacity costs nothing
  • Competitive effective hourly rates: H100 SXM5 $3.95, H200 $4.54, B200 $6.25, B300 $7.10, A100 80GB $2.50, A100 40GB $2.10, L40S $1.95, A10 $1.10, L4 $0.80, T4 $0.59
  • CPU and memory are billed separately and transparently at $0.0000131 per physical core-second and $0.00000222 per GiB-second, so mixed workloads are costable
  • Storage is $0.09 per GiB per month with the first 1 TiB free
  • Starter plan is $0 with $30 a month of included compute, which is enough to evaluate the platform properly
  • Blackwell Ultra (B300) is available self-serve, which several dedicated providers do not offer

Cons

  • Team plan carries a $250 a month platform fee against $100 a month of included compute, so the fixed cost is real
  • Serverless abstraction means less control over placement, interconnect and node topology than a raw GPU rental
  • Not the right shape for a long multi-node training run with a high-performance fabric
  • The pricing page does not publish an egress policy
  • A minimum of 0.125 physical CPU cores is billed alongside every workload
Honest Weakness: Modal is priced for intermittency and punishes the opposite. An H100 running continuously costs $3.95 an hour, roughly $2,880 a month, which is comparable to dedicated providers but with a serverless platform's constraints on topology and placement. Teams that adopt Modal for a bursty inference endpoint and then move a continuous training job onto it get the worst of both: platform limitations without the price advantage. Match the workload shape to the billing shape, or the abstraction costs you money.

Where per-second billing actually pays

Take an inference endpoint receiving real traffic for four hours a day. On a dedicated H100 instance at $3.99 an hour you pay for 24, which is roughly $2,900 a month. On Modal at $3.95 an hour of actual use you pay for 4, which is roughly $480. The saving is not in the rate, it is in the 20 hours you do not buy. Invert the workload to a continuous training run and the arithmetic disappears entirely.

Cold starts

Scale to zero means the first request after an idle period pays a start-up cost. That is the trade for not paying while idle and it is the thing to test with your actual model size before committing. Our inference and model hosting comparison covers cold-start behaviour across serving platforms in more detail.

Per second: B300 $0.001972, B200 $0.001736, H200 SXM $0.001261, H100 SXM5 $0.001097, RTX PRO 6000 $0.000842, A100 80GB $0.000694, A100 40GB $0.000583, L40S $0.000542, A10 $0.000306, L4 $0.000222, T4 $0.000164. Those are $7.10, $6.25, $4.54, $3.95, $3.03, $2.50, $2.10, $1.95, $1.10, $0.80 and $0.59 per hour respectively. CPU $0.0000131 per physical core-second with a 0.125 core minimum; memory $0.00000222 per GiB-second. Storage $0.09 per GiB per month with 1 TiB free. Plans: Starter $0 base with $30 a month included compute; Team $250 a month base with $100 a month included compute; Enterprise custom with volume discounts. Checked on modal.com/pricing, 2026-09-18.

Visit Modal
6

Crusoe

Honorable Mention

Best for: The lowest published on-demand H100 and H200 rates from a dedicated provider, with free ingress and egress

Crusoe publishes $3.90 per GPU-hour for HGX H100 and $4.29 for HGX H200, the lowest on-demand rates from a dedicated provider in this comparison, and states plainly that it does not charge for network ingress or egress. Storage is cheap at $0.07 to $0.08 per GiB per month. It is also the only provider here publishing an AMD option, MI300X at $3.45. The catch is that everything current-generation is behind a sales conversation.

Pros

  • Lowest published dedicated-provider on-demand rates: H100 $3.90 and H200 $4.29 per GPU-hour
  • States in writing that it does not charge for network ingress or egress
  • Cheap storage: persistent disks $0.08 per GiB per month, shared disks $0.07, object storage $0.06
  • A100 80GB SXM at $2.30 and PCIe at $2.00, plus L40S at $1.50, cover cost-sensitive workloads
  • AMD MI300X published at $3.45 per GPU-hour, a genuine alternative for teams wanting to reduce NVIDIA dependence

Cons

  • B200, GB200 NVL72 and MI355X are all contact sales, so no current-generation NVIDIA capacity is self-serve priced
  • Spot pricing requires contacting sales, with no public rates at all
  • Reserved and commitment discounts are described only as tailored agreements, with no figures published
  • Smaller catalogue and smaller footprint than CoreWeave, which matters for regional availability
  • The published prices cover the previous generation, so a buyer planning a Blackwell migration cannot price the path
Honest Weakness: Crusoe publishes the best prices for the chips you may be moving away from and no prices at all for the chips you are moving to. H100 at $3.90 and H200 at $4.29 are genuinely leading rates, but B200, GB200 NVL72 and MI355X are all contact sales, as is every spot rate. A buyer choosing Crusoe on published pricing is choosing on last-generation economics, then negotiating blind for the next generation. Get Blackwell pricing in writing before you build a roadmap on the H100 number.

The AMD option

MI300X at $3.45 per GPU-hour with 192 GB of memory is the only published AMD price in this comparison. For inference on large models, memory capacity per GPU often matters more than raw throughput, and 192 GB against an H100's 80 GB changes how many cards a model needs. The cost is software maturity: your stack has to work on ROCm, and that is a real porting question rather than a flag change.

On-demand per GPU-hour: NVIDIA H200 141GB HGX $4.29, H100 80GB HGX $3.90, A100 80GB SXM $2.30, A100 80GB PCIe $2.00, L40S 48GB $1.50, AMD MI300X 192GB $3.45. NVIDIA GB200 NVL72, NVIDIA B200 180GB HGX and AMD MI355X 288GB are contact sales. All spot pricing is contact sales. Reserved terms are described as tailored agreements with no published figures. Storage: persistent disks $0.08 per GiB per month, shared disks $0.07, object storage $0.06. The page states that Crusoe Cloud does not charge for network ingress or egress. Checked on crusoe.ai/cloud/pricing, 2026-09-18.

Visit Crusoe
7

Vast.ai

Honorable Mention

Best for: The cheapest possible floor price, when you can tolerate a marketplace and interruption

Vast.ai is a marketplace rather than a cloud. Prices are set by supply and demand across more than 40 data centres and over 68 GPU types, billed per second with no minimum and no rounding. Three purchase modes exist: on-demand with guaranteed uptime, interruptible described as 50% or more cheaper, and reserved for 1, 3 or 6 months at up to 50% off. It routinely produces the lowest rate available for a given chip, and it publishes no fixed price at all.

Pros

  • Marketplace competition across 40+ data centres typically produces the lowest available rate for any given chip
  • Per-second billing with no minimum hours and no rounding up
  • Three modes with clear trade-offs: on-demand with guaranteed uptime, interruptible at 50%+ cheaper, and reserved at up to 50% off for 1, 3 or 6 months
  • 68+ GPU types, including both current data-centre parts and older consumer hardware
  • Suits fault-tolerant training and batch inference exceptionally well

Cons

  • No exact prices are published anywhere on the pricing page, so you cannot budget without opening the console
  • No storage or bandwidth pricing is published either
  • Hosts are independent operators, which makes this the weakest fit in this comparison for regulated data
  • Quality, network and reliability vary by host, and the platform's rating system is the only signal
  • Prices move with demand, so a cost model built this month may not hold next month
Honest Weakness: You cannot plan a budget from Vast.ai's own pricing page, because it contains no prices. That is a defensible consequence of being a marketplace and it is still a real problem: every other provider here lets you build a spreadsheet before you create an account. Combined with independent hosts and demand-driven rates, that makes Vast.ai excellent for opportunistic cheap compute and poor for anything a finance team has to forecast or a compliance team has to approve.

How to use a marketplace safely

Treat it as spot capacity regardless of which mode you buy. Checkpoint aggressively, keep the dataset somewhere you control rather than on the instance, assume any given host can disappear, and never put regulated or customer data on it. Under those constraints it is the cheapest compute in this comparison by a comfortable margin.

No fixed prices published. Rates are set by supply and demand across 40+ data centres and 68+ GPU types, and are visible only in the live console. Billing is per second with no minimum hours and no rounding up. Purchase modes: On-Demand with guaranteed uptime and no interruptions; Interruptible, described as 50%+ cheaper for preemptible workloads; Reserved at up to 50% off with 1, 3 or 6-month commitments. No storage or bandwidth charges are published on the pricing page. Checked on vast.ai/pricing, 2026-09-18.

Visit Vast.ai

Which One Should You Pick?

Use CaseOur Recommendation
Multi-node pretraining or large-scale fine-tuning with a fast interconnectCoreWeave or Lambda 1-Click Clusters. CoreWeave publishes rack-scale GB200 NVL72 at $42.00 per hour for 4 GPUs and charges nothing for egress; Lambda clusters run from 2 weeks to 1 year at $5.54 to $6.16 per H100-hour. Both are more expensive per GPU than the marketplaces and both are the right answer here.
A training run that checkpoints cleanly and can survive interruptionTogether preemptible at $1.99 per H100-hour, or Vast.ai interruptible. You are paying roughly half for the same silicon. If your training code cannot resume from a checkpoint, fixing that is a day of work that permanently halves this bill.
An inference endpoint with real traffic for a few hours a dayModal, for per-second billing and scale to zero. Four hours of daily use on an H100 is roughly $480 a month against roughly $2,900 for an always-on instance. Test cold-start latency with your actual model size before committing.
Prototyping a training script before scaling itLambda for a single H100 at $4.29, only 7.5% above its 8-GPU rate, so your cost model does not change shape when you scale. Or RunPod consumer cards at $0.16 to $0.69 an hour if the model fits in 24 to 32 GB.
Customer data or a regulated workloadCoreWeave, Crusoe, Lambda or RunPod Secure Cloud. Avoid RunPod Community and Vast.ai entirely: both source capacity from third-party hosts, which is a data-handling question your security review will not wave through.
A pipeline that moves tens of terabytes in and out every monthCoreWeave or Crusoe, both of which state in writing that data transfer is free, or Lambda, which states no egress fees. Together and Modal do not publish an egress policy, so get it in writing before comparing their rates.
You only need to run inference on a stock open-weight modelDo not buy GPU capacity at all. A serverless inference API charges per token and costs nothing when idle. Renting a GPU to serve a model someone else already hosts is the most common and most expensive mistake in this category.
You want to reduce NVIDIA dependenceCrusoe publishes AMD MI300X at $3.45 per GPU-hour with 192 GB of memory, the only published AMD rate here. Budget engineering time for the ROCm port; the discount is not free.

How we evaluated

Last verified: 18 September 2026. Every price on this page was read from the provider's own pricing page on that date, and each entry names the page. No figure is carried from a comparison site, a pricing aggregator or a vendor's marketing blog. That includes the hyperscaler section. The AWS rates quoted are the per-accelerator figures AWS publishes on its EC2 Capacity Blocks for ML pricing page, because AWS's own P5 instance page publishes no on-demand hourly rate at all.

Where a provider publishes nothing, the page says so. Vast.ai publishes no fixed prices anywhere, only a marketplace model and three purchase modes. Together publishes no storage rate and no egress policy. Modal publishes no egress policy. Lambda publishes no storage rate on its pricing page. Crusoe publishes no spot rates and no reserved discounts. Those gaps are reported as findings, not filled in.

Where a provider sells by node, the per-GPU figure in the comparison table is derived by dividing the published node price by its GPU count, and the derivation is shown. That is an honest comparison number and a misleading purchase number, because you cannot buy a fraction of a CoreWeave node.

The criteria, in the order they decide real outcomes:

  • Published price per GPU-hour, by chip class. H100, H200, B200 and Blackwell Ultra, because that is how the decision is actually framed and because mixing generations is how comparisons become dishonest.
  • Purchase modes and commitment terms. On-demand, reserved, preemptible and spot, with the term lengths each provider requires. A price without its availability guarantee is not a price.
  • Availability of current-generation parts. Specifically whether Blackwell and Blackwell Ultra capacity is self-serve priced or contact sales, because that distinction now separates the providers more than their rates do.
  • Storage. Published per-GB-per-month rates, including the difference between running and idle volumes, which is where several providers recover the discount they gave on compute.
  • Egress. Stated in writing, or not stated. This is the largest single hidden cost in a training pipeline and the one most often left out of a comparison.
  • Minimum purchase unit and billing granularity. Per second, per GPU, or per 8-GPU node. This decides whether a small job is cheap or absurd.
  • Hosting model. Whether capacity comes from vetted data centres or third-party hosts, which is a compliance decision rather than a price one.

Vendor selection

Seven providers were chosen to cover the distinct shapes this market has: hyperscale specialist (CoreWeave), transparent per-GPU rental (Lambda, Crusoe), rental plus a serving and tuning platform (Together), two-tier marketplace (RunPod), pure marketplace (Vast.ai), and serverless per-second compute (Modal). The hyperscalers are covered as a pricing reference point rather than as entries, because their GPU offerings are bought for reasons that have little to do with GPU price.

One ownership change is worth naming: CoreWeave completed its acquisition of Weights & Biases on 5 May 2025. A team buying both GPU capacity and ML tooling may be consolidating onto one vendor without intending to.

What we did not do

No provider on this page was benchmarked. No instance was launched, no throughput was measured, and there are no hands-on testing claims anywhere in this comparison. Performance on identical silicon is dominated by interconnect, storage throughput and your own code, none of which generalise from a vendor test. What this page offers instead is an accurate, dated price comparison built from primary sources, an explicit statement of each provider's weakest point, and the non-obvious costs that change the total.

Prices in this category move faster than in any other we track. Re-check the provider's own page before you commit, and treat any rate on any comparison site, including this one, as a snapshot rather than a quote.

Note

Editorial independence: this is a vendor-neutral comparison with no paid placements, sponsorships or affiliate links. Rankings reflect fit for the stated use cases, not commercial relationships.

Sources

Frequently Asked Questions

What does an H100 actually cost per hour in 2026?
Between $1.99 and $6.16 per GPU-hour across the seven providers here, depending entirely on what you are willing to give up, and every figure here came from the provider's own page on 18 September 2026. The floor is Together preemptible at $1.99 and RunPod Community PCIe at $1.99, both of which trade away guaranteed availability or vetted hosting. The middle is Crusoe at $3.90, Modal at $3.95 effective, and Lambda and Together on-demand at $3.99. The top of the dedicated range is CoreWeave at a derived $6.16 per GPU, sold only as a $49.24 8-GPU node. AWS publishes $5.191 per accelerator-hour for P5 H100 capacity through EC2 Capacity Blocks for ML. A three-times spread for identical silicon is normal in this market, and it is entirely explained by availability guarantees, interconnect, hosting model and egress policy.
What changed with current-generation accelerators?
Three things. Blackwell became the price-setting tier and carries a clear premium. B200 runs $5.98 to $8.60 per GPU-hour across these providers against $1.99 to $6.16 for H100. Blackwell Ultra (B300, GB300) runs $6.94 to $7.89 on RunPod and $7.10 on Modal, while Fireworks publishes $15.00 for B300 and $20.00 for GB300 on its inference-oriented GPU tier. Second, rack-scale systems changed the unit of sale: CoreWeave's GB200 NVL72 is sold at $42.00 per hour for 4 GPUs, not per GPU, because the rack is the product. Third, the top of the range retreated behind sales. HGX B300 and GB300 NVL72 are contact sales on CoreWeave, B200 and GB200 are contact sales on Crusoe, and GB200 NVL72 and B300 are contact sales on Together. Meanwhile H100 has become the commodity floor, which is why it is the chip most worth shopping hard on.
Where does the quoted hourly rate stop being the real cost?
Four places, in rough order of how much money they move. Egress: CoreWeave and Crusoe state transfer is free and Lambda states no egress fees, while Together, Modal and Vast.ai publish no policy at all and hyperscalers meter it. On a pipeline shifting 50 TB a month, that is a five-figure annual difference. Storage: RunPod charges $0.20 per GB per month for idle volumes against $0.10 while running, so a 500 GB dataset left attached costs $100 a month to do nothing, while CoreWeave cold object storage is $0.015 per GB per month. Minimum purchase unit: CoreWeave's smallest HGX H100 purchase is a $49.24-per-hour node, so a single-GPU job pays for eight. Utilisation: an always-on instance you use four hours a day costs six times what per-second billing would. Model all four before you compare headline rates.
How do the hyperscalers price against these providers?
AWS publishes per-accelerator-hour rates through EC2 Capacity Blocks for ML: $5.191 for P5 (H100), $5.97 for P5e (H200), $10.582 for P6e (B200), $14.04 for P6-B300 and $0.596 for Trn1 (Trainium). Those sit at or above the top of the specialist range, and the reservation fee is charged up front at the time you schedule, with an operating-system fee on top for non-Linux images. AWS's own P5 instance page publishes no on-demand hourly price at all, which is itself informative: hyperscaler GPU pricing is not designed to be read off a page and compared. Add metered egress, which the specialists largely do not charge, and the practical gap is wider than the hourly rates suggest. What you buy for the premium is everything around the GPU: IAM, VPC, compliance attestations your auditor already accepts, and the ability to add capacity without a new vendor review. For many regulated buyers that is worth it. For a team whose only requirement is GPUs, it is not.
What is the difference between on-demand, reserved, preemptible and spot?
On-demand means you pay a published rate and keep the instance until you stop it. Reserved means you commit to a term in exchange for a lower rate: Together's reserved terms start at 7 days, Vast.ai offers 1, 3 and 6 months, and Lambda's 1-Click Clusters run 2 weeks to a year. Preemptible and spot mean the provider can reclaim your capacity, typically at roughly half price: Together publishes H100 preemptible at $1.99 against $3.99 on-demand, and Vast.ai describes interruptible as 50% or more cheaper. One counterintuitive case is worth knowing: Lambda's reserved cluster rates are higher than its on-demand rates, because a 1-Click Cluster is a different product with a dedicated fabric and guaranteed capacity, not a discount for committing. Read what the commitment buys before assuming it saves money.
Who should not buy GPU cloud capacity at all?
Three groups, and together they are most of the people who ask. Anyone who only needs inference on a stock open-weight model. A serverless inference API charges per token and costs nothing when idle. Renting a GPU to serve a model someone else already hosts is the most expensive mistake in this category. Anyone whose workload runs a few hours a week: per-second serverless compute or a managed API will be cheaper than any hourly rental, including the marketplaces. And anyone without someone on the team who can babysit a training run, because a failed job on a $49-an-hour node fails silently and keeps billing. Buy raw GPU capacity when you are training or fine-tuning your own models, when you need a specific chip or interconnect, or when the volume makes per-token pricing more expensive than per-hour.
Is a marketplace like Vast.ai or RunPod Community safe to use?
It depends entirely on the data. Both source capacity from third-party hosts rather than vetted data centres, which is exactly why they are cheap. For public datasets, open-weight models, benchmarking and experimentation, the 25% to 50% saving carries no meaningful risk. For customer data, anything under a data-processing agreement, or anything your security team has attested about, they are not defensible, and no rating system changes that. The practical pattern is to split the work: iterate on marketplace capacity with synthetic or public data, then run the production job on RunPod Secure at $3.49 per H100-hour, Crusoe at $3.90 or a dedicated provider.
How stable are these prices?
Less stable than any other category we track, which is why every figure on this page carries the date it was checked, 18 September 2026. Two patterns recur. Promotions get quoted as list prices: Together's dedicated HGX H100 inference rate is advertised at $3.99 against a $5.49 list, with the promotion dated to 09/30/26, so the durable number is the higher one. And generational transitions reprice everything beneath them: as Blackwell Ultra capacity broadens, H100 rates fall further, which is good news if you are buying and a reason not to sign a long commitment at today's H100 price. Re-check the provider's own page before you commit, and never take a rate from a comparison site, including this one, as current.

About the author

is the founder and creator of LoginRadius, a customer identity platform he built and scaled to over a billion users. He is now the founder of GrackerAI, a GEO platform for B2B SaaS and cybersecurity teams, and has spent more than 15 years building identity and security products.

Related Comparisons