Top 7 GPU Cloud and AI Compute Providers 2026: Price Per GPU-Hour by Chip Class, Verified From Each Provider's Own Pricing Page
CoreWeave, Lambda, Together, RunPod, Vast.ai, Modal and Crusoe compared on published price per GPU-hour for H100, H200, B200 and Blackwell Ultra, plus commitment terms, storage, egress and availability. Every figure read from the provider's own page on 2026-09-18.
The answer, before the table
Price per GPU-hour by chip class is the whole decision, so start there. Every figure below was read from the provider's own pricing page on 18 September 2026.
- Cheapest H100 you can actually buy: Together preemptible at $1.99 per GPU-hour, or RunPod Community PCIe at $1.99. Both trade away guaranteed availability or vetted hosting.
- Cheapest dedicated-provider on-demand H100: Crusoe at $3.90, then Modal at an effective $3.95, then Lambda and Together at $3.99.
- Cheapest published B200: Lambda at $6.69 per GPU-hour on 8x instances.
- Blackwell Ultra self-serve: RunPod B300 at $6.94 Community and $7.89 Secure, and Modal B300 at an effective $7.10. Almost everyone else says contact sales.
- Scale and rack-scale systems: CoreWeave, which publishes GB200 NVL72 at $42.00 per hour for 4 GPUs and charges nothing for egress.
- Lowest floor of all, with the least certainty: Vast.ai, a marketplace that publishes no prices at all.
A three-times spread for the same silicon is normal here. It is explained by availability guarantees, interconnect, hosting model and egress policy, not by the chip.
What changed with current-generation accelerators
Three shifts matter for a 2026 purchase.
Blackwell set the new price ceiling, and H100 became the commodity floor. B200 runs $5.98 to $8.60 per GPU-hour across these providers against $1.99 to $6.16 for H100. Blackwell Ultra (B300, GB300) is $6.94 to $7.89 on RunPod and an effective $7.10 on Modal, while Fireworks publishes $15.00 for B300 and $20.00 for GB300 on its inference-oriented GPU tier. If your workload runs fine on H100, this is the best time in three years to buy it.
The unit of sale changed at the top end. CoreWeave sells GB200 NVL72 at $42.00 per hour for 4 GPUs, because a rack-scale NVL system is the product rather than an individual accelerator. HGX lines are sold by the 8-GPU node. Per-GPU figures for CoreWeave in this comparison are derived by division and are honest for comparison, not for purchase.
The newest parts retreated behind sales. HGX B300 and GB300 NVL72 are contact sales on CoreWeave. B200, GB200 NVL72 and MI355X are contact sales on Crusoe, as is every Crusoe spot rate. GB200 NVL72 and B300 are contact sales on Together. A buyer comparing published prices in 2026 is largely comparing last-generation economics.
Where the hourly rate stops being the real cost
This is the part that decides the invoice, and none of it is in the headline number.
| Cost | Who charges nothing | Who charges, and how much |
|---|---|---|
| Egress | CoreWeave (free transfer), Crusoe ("does not charge for network ingress or egress"), Lambda ("no egress fees") | Together, Modal and Vast.ai publish no policy. Hyperscalers meter it |
| Idle storage | CoreWeave cold object storage at $0.015/GB/month | RunPod volume disk at $0.20/GB/month while idle, double the $0.10 running rate |
| Minimum purchase | Lambda (1 GPU), RunPod (1 GPU), Modal (per second), Vast.ai (per second) | CoreWeave's smallest HGX H100 purchase is a $49.24/hour 8-GPU node |
| Idle GPU time | Modal, which scales to zero | Every hourly rental, where a four-hour-a-day endpoint pays for 24 |
On a pipeline moving 50 TB a month, egress alone is a five-figure annual difference between two providers whose hourly rates look identical.
How the hyperscalers price against them
AWS publishes per-accelerator-hour rates through EC2 Capacity Blocks for ML: $5.191 for P5 (H100), $5.97 for P5e (H200), $10.582 for P6e (B200), $14.04 for P6-B300 and $0.596 for Trn1 (Trainium). The reservation fee is charged up front when you schedule, with an operating-system fee on top for non-Linux images. AWS's own P5 instance page publishes no on-demand hourly price at all, which tells you something about how hyperscaler GPU pricing is meant to be consumed.
Those rates sit at or above the top of the specialist range, before metered egress. What the premium buys is everything around the GPU: IAM, VPC, the compliance attestations your auditor already accepts, and capacity you can add without a new vendor review. For a regulated buyer that is frequently worth it. For a team whose requirement is simply GPUs, it is not.
Related comparisons
- Top 5 AI Inference and Model Hosting Platforms 2026 if you only need to serve a model, in which case you should probably not buy GPU capacity at all.
- Top 8 Fine-Tuning and Model Customization Platforms 2026 for what you would actually run on this hardware.
- Top 10 MLOps Platforms 2026 for the orchestration layer above it.
- AI models directory for the memory footprint of the models you are sizing hardware for.
Quick Comparison
| Provider | H100 (per GPU-hour) | H200 (per GPU-hour) | B200 (per GPU-hour) | Blackwell Ultra (B300 / GB300) | Egress | Billing Granularity and Commitment |
|---|---|---|---|---|---|---|
| CoreWeave | $6.16 (HGX H100 8-GPU node at $49.24/hr) | $6.31 (HGX H200 8-GPU node at $50.44/hr) | $8.60 (HGX B200 8-GPU node at $68.80/hr) | HGX B300 and GB300 NVL72 are contact sales. GB200 NVL72 is $42.00/hr for 4 GPUs, $10.50 per GPU-hour | Free, for internet egress and internal transfers | Sold by node, not by GPU. No minimum commitment stated in the on-demand pricing tables |
| Lambda | $3.99 on 8x SXM, $4.29 on 1x SXM, $3.29 on 1x PCIe | Not listed in the on-demand table | $6.69 on 8x, $6.99 on 1x | Not listed | "Transparent pricing with no egress fees" | Per-GPU on-demand from a single card. 1-Click Clusters run 2 weeks to 1 year |
| Together AI | $3.99 on-demand, $3.19 to $3.69 reserved, $1.99 preemptible | $5.99 on-demand, $3.99 to $4.99 reserved, $2.99 preemptible | $8.19 on-demand, $6.79 to $7.99 reserved, $4.09 preemptible | GB200 NVL72 and B300 are contact sales | Not published on the pricing page | Reserved terms run 7 to 180+ days. Preemptible for fault-tolerant work |
| RunPod | $2.69 Community, $3.49 Secure (SXM). PCIe $1.99 Community, $2.89 Secure | $3.59 Community, $4.59 Secure | $5.98 Community, $6.79 Secure | B300 $6.94 Community, $7.89 Secure | No egress charge listed in the pricing documentation | Per-GPU, no minimum. Two tiers: Community (third-party hosts) and Secure (vetted data centres) |
| Vast.ai | Marketplace rate, not published | Marketplace rate, not published | Marketplace rate, not published | Marketplace rate, not published | Not published on the pricing page | Per-second billing with no minimum or rounding. On-demand, interruptible ("50%+ cheaper") and reserved for 1, 3 or 6 months ("up to 50% off") |
| Modal | $3.95 (H100 SXM5 at $0.001097/second) | $4.54 ($0.001261/second) | $6.25 ($0.001736/second) | B300 $7.10 ($0.001972/second) | Not published on the pricing page | Per-second billing with scale to zero. Starter $0 + $30/month included compute, Team $250/month + $100/month included compute |
| Crusoe | $3.90 (HGX, 80GB) | $4.29 (HGX, 141GB) | Contact sales | GB200 NVL72 and MI355X are contact sales | "Crusoe Cloud does not charge for network ingress or egress" | Per-GPU on-demand published. Spot pricing and reserved terms are contact sales |
CoreWeave
- H100 (per GPU-hour)
- $6.16 (HGX H100 8-GPU node at $49.24/hr)
- H200 (per GPU-hour)
- $6.31 (HGX H200 8-GPU node at $50.44/hr)
- B200 (per GPU-hour)
- $8.60 (HGX B200 8-GPU node at $68.80/hr)
- Blackwell Ultra (B300 / GB300)
- HGX B300 and GB300 NVL72 are contact sales. GB200 NVL72 is $42.00/hr for 4 GPUs, $10.50 per GPU-hour
- Egress
- Free, for internet egress and internal transfers
- Billing Granularity and Commitment
- Sold by node, not by GPU. No minimum commitment stated in the on-demand pricing tables
Lambda
- H100 (per GPU-hour)
- $3.99 on 8x SXM, $4.29 on 1x SXM, $3.29 on 1x PCIe
- H200 (per GPU-hour)
- Not listed in the on-demand table
- B200 (per GPU-hour)
- $6.69 on 8x, $6.99 on 1x
- Blackwell Ultra (B300 / GB300)
- Not listed
- Egress
- "Transparent pricing with no egress fees"
- Billing Granularity and Commitment
- Per-GPU on-demand from a single card. 1-Click Clusters run 2 weeks to 1 year
Together AI
- H100 (per GPU-hour)
- $3.99 on-demand, $3.19 to $3.69 reserved, $1.99 preemptible
- H200 (per GPU-hour)
- $5.99 on-demand, $3.99 to $4.99 reserved, $2.99 preemptible
- B200 (per GPU-hour)
- $8.19 on-demand, $6.79 to $7.99 reserved, $4.09 preemptible
- Blackwell Ultra (B300 / GB300)
- GB200 NVL72 and B300 are contact sales
- Egress
- Not published on the pricing page
- Billing Granularity and Commitment
- Reserved terms run 7 to 180+ days. Preemptible for fault-tolerant work
RunPod
- H100 (per GPU-hour)
- $2.69 Community, $3.49 Secure (SXM). PCIe $1.99 Community, $2.89 Secure
- H200 (per GPU-hour)
- $3.59 Community, $4.59 Secure
- B200 (per GPU-hour)
- $5.98 Community, $6.79 Secure
- Blackwell Ultra (B300 / GB300)
- B300 $6.94 Community, $7.89 Secure
- Egress
- No egress charge listed in the pricing documentation
- Billing Granularity and Commitment
- Per-GPU, no minimum. Two tiers: Community (third-party hosts) and Secure (vetted data centres)
Vast.ai
- H100 (per GPU-hour)
- Marketplace rate, not published
- H200 (per GPU-hour)
- Marketplace rate, not published
- B200 (per GPU-hour)
- Marketplace rate, not published
- Blackwell Ultra (B300 / GB300)
- Marketplace rate, not published
- Egress
- Not published on the pricing page
- Billing Granularity and Commitment
- Per-second billing with no minimum or rounding. On-demand, interruptible ("50%+ cheaper") and reserved for 1, 3 or 6 months ("up to 50% off")
Modal
- H100 (per GPU-hour)
- $3.95 (H100 SXM5 at $0.001097/second)
- H200 (per GPU-hour)
- $4.54 ($0.001261/second)
- B200 (per GPU-hour)
- $6.25 ($0.001736/second)
- Blackwell Ultra (B300 / GB300)
- B300 $7.10 ($0.001972/second)
- Egress
- Not published on the pricing page
- Billing Granularity and Commitment
- Per-second billing with scale to zero. Starter $0 + $30/month included compute, Team $250/month + $100/month included compute
Crusoe
- H100 (per GPU-hour)
- $3.90 (HGX, 80GB)
- H200 (per GPU-hour)
- $4.29 (HGX, 141GB)
- B200 (per GPU-hour)
- Contact sales
- Blackwell Ultra (B300 / GB300)
- GB200 NVL72 and MI355X are contact sales
- Egress
- "Crusoe Cloud does not charge for network ingress or egress"
- Billing Granularity and Commitment
- Per-GPU on-demand published. Spot pricing and reserved terms are contact sales
CoreWeave
Best for EnterpriseBest for: Multi-node training at scale, rack-scale Blackwell systems, and teams that move terabytes in and out
“CoreWeave is the largest of the specialist GPU clouds and the one that publishes the most complete price list, including rack-scale GB200 NVL72 capacity that most competitors will only quote. It sells by the node rather than the GPU, which rules it out for single-card experimentation and makes it the right shape for real multi-node training. Free egress and $0.015 to $0.07 per GB per month storage make the total bill predictable in a way hyperscaler pricing is not.”
Pros
- Published node pricing across the full range: HGX H100 $49.24/hr, HGX H200 $50.44/hr, HGX B200 $68.80/hr, A100 $21.60/hr, L40S $18.00/hr, RTX PRO 6000 Blackwell high-memory $20.00/hr, GH200 $6.50/hr for a single GPU
- GB200 NVL72 is published at $42.00 per hour for 4 GPUs, which is $10.50 per GPU-hour, rather than hidden behind sales
- Internet egress and internal transfers are free, which removes the single largest hidden cost of moving a training pipeline to a cloud
- Storage is cheap and tiered: object storage at $0.06 hot, $0.03 warm and $0.015 cold per GB per month, distributed file storage at $0.070 per GB per month
- No minimum commitment appears in the on-demand pricing tables
Cons
- GPUs are sold in 8-GPU nodes for the HGX lines, so there is no cheap way to run one card
- HGX B300 and GB300 NVL72 are contact sales, so the current top of the range is not self-serve
- Public IP addresses are billed separately at $4.00 per IP per month
- Per-GPU H100 pricing of $6.16 is higher than Crusoe at $3.90, Lambda at $3.99 and RunPod Community at $2.69
- Now also owns Weights & Biases, so a team buying both compute and ML tooling may be concentrating vendor risk
Why node pricing changes the comparison
Every per-GPU figure for CoreWeave in the table above is derived by dividing the published node price by the GPU count, because that is how the capacity is sold. It is an honest comparison number and a misleading purchase number. You cannot buy $6.16 of CoreWeave H100. The smallest unit is $49.24 an hour, which is roughly $35,500 a month if left running.
Free egress is the underrated line item
Training pipelines move data constantly: datasets in, checkpoints out, evaluation artefacts back and forth. Hyperscalers meter that. CoreWeave, Crusoe and Lambda do not. On a pipeline shifting 50 TB a month, hyperscaler egress at typical published rates is a five-figure annual line that simply does not exist on these providers.
Ownership and consolidation
CoreWeave completed its acquisition of Weights & Biases on 5 May 2025. If you are also evaluating W&B Weave for evaluation and tracing, note that you would be buying compute and tooling from the same company.
North America on-demand, per node: GB200 NVL72 (4 GPUs) $42.00/hr, HGX B200 (8 GPUs) $68.80/hr, HGX H200 (8 GPUs) $50.44/hr, HGX H100 (8 GPUs) $49.24/hr, RTX PRO 6000 Blackwell high-memory (8 GPUs) $20.00/hr, A100 (8 GPUs) $21.60/hr, L40S (8 GPUs) $18.00/hr, L40 (8 GPUs) $10.00/hr, GH200 (1 GPU) $6.50/hr. GB300 NVL72 and HGX B300 are contact sales. Storage: object hot $0.06, warm $0.03, cold $0.015 per GB per month; distributed file storage $0.070 per GB per month. Data transfer free. Public IP $4.00 per IP per month. Checked on coreweave.com/pricing, 2026-09-18.
Lambda
Best OverallBest for: Clear per-GPU on-demand pricing from a single card upward, with no egress fees
“Lambda publishes the cleanest per-GPU price list in this comparison and lets you rent a single card. H100 SXM is $3.99 per GPU-hour on an 8x instance and $4.29 on a 1x, and B200 is $6.69 on 8x and $6.99 on 1x. The narrow gap between one GPU and eight is unusual and valuable: you can prototype on one card at almost the same unit rate you will pay in production. Lambda's own page states there are no egress fees.”
Pros
- Per-GPU on-demand pricing from a single card, with a small premium for smaller configurations
- B200 at $6.69 per GPU-hour on 8x instances is the cheapest published on-demand Blackwell rate among the dedicated providers here
- H100 SXM at $3.99 per GPU-hour on 8x undercuts CoreWeave's derived $6.16 substantially
- Wide older-generation catalogue for cheap experimentation: A100 40GB $1.99, A10 $1.29, A6000 $1.09, V100 $0.79, Quadro RTX 6000 $0.69
- States transparent pricing with no egress fees
- 1-Click Clusters give reserved multi-node capacity on terms from 2 weeks to 1 year without an enterprise contract
Cons
- 1-Click Cluster rates are higher than on-demand: H100 runs $5.54 to $6.16 per GPU-hour and B200 runs $8.87 to $9.86, above the $3.99 and $6.69 on-demand rates
- H200 does not appear in the published on-demand instance table
- No Blackwell Ultra (B300 or GB300) capacity is listed
- The pricing page does not publish a persistent storage rate, stating only that rates appear when you create a filesystem
- On-demand availability for current-generation parts is not guaranteed, and the page does not publish capacity commitments
The one-to-eight GPU gap
On most providers, renting one GPU carries a large premium over renting a full node. Lambda charges $4.29 for one H100 against $3.99 per GPU on eight, a 7.5% difference. That makes it the most sensible place to develop a training script before scaling it, because your cost model does not change shape when you scale.
Older generations are still the right answer often
A100 40GB at $1.99 and A10 at $1.29 remain the correct hardware for LoRA fine-tuning small models, for embedding workloads and for anything memory-light. Pairing those rates with an efficiency library that cuts VRAM sharply is how a fine-tuning programme stays inexpensive.
On-demand per GPU-hour: B200 SXM6 $6.69 (8x), $6.79 (4x), $6.89 (2x), $6.99 (1x). H100 SXM $3.99 (8x), $4.09 (4x), $4.19 (2x), $4.29 (1x). H100 PCIe $3.29 (1x). GH200 $2.29 (1x). A100 SXM 80GB $2.79 (8x). A100 SXM 40GB $1.99 (1x). A100 PCIe 40GB $1.99 (4x and 2x). A10 $1.29. A6000 $1.09. V100 $0.79 (8x). Quadro RTX 6000 $0.69. 1-Click Clusters, 2 weeks to 1 year: B200 $9.86 (16 GPUs), $9.36 (64), $8.87 (256+); H100 $6.16 (16), $5.85 (64), $5.54 (256). Commitments beyond 1 year are custom. The page states no egress fees and does not publish a storage rate. Checked on lambda.ai/pricing and lambda.ai/service/gpu-cloud, 2026-09-18.
Together AI GPU Clusters
Runner UpBest for: Fault-tolerant training that can take preemption, and teams that also want inference on the same account
“Together publishes three prices per chip: on-demand, reserved and preemptible. That third column is the reason to look. H100 preemptible at $1.99 per GPU-hour is roughly half its own on-demand rate and among the lowest published figures anywhere for that chip. If your training loop checkpoints properly, you are paying an availability premium you do not need. The same account also covers inference and fine-tuning, which removes a vendor from the stack.”
Pros
- Three published price tiers per chip, so the preemption trade-off is explicit rather than a sales conversation
- Preemptible rates are aggressive: H100 $1.99, H200 $2.99, B200 $4.09 per GPU-hour
- Reserved terms start at 7 days, far shorter than the multi-month commitments most providers require: H100 $3.19 to $3.69, H200 $3.99 to $4.99, B200 $6.79 to $7.99
- On-demand H100 at $3.99 matches Lambda and undercuts CoreWeave's derived $6.16
- The same platform covers dedicated inference endpoints and fine-tuning, so a training-to-serving pipeline stays with one vendor
Cons
- GB200 NVL72 and B300 are contact sales, so the current top of the range is not self-serve
- The pricing page does not publish storage rates or an egress policy, both of which are decisive at scale
- Preemptible capacity is only usable if your training code checkpoints and resumes cleanly, which is a real engineering prerequisite
- Dedicated inference on HGX H100 is advertised at $3.99 against a $5.49 list price under a promotion dated to 09/30/26, so the durable rate is the higher one
Preemptible capacity is the real product
The gap between $3.99 and $1.99 on H100 is the whole argument. Training is the canonical preemptible workload: it is long-running, it checkpoints naturally, and a restart costs minutes rather than a customer. If your training script cannot resume from a checkpoint, fixing that is a day of work that halves your compute bill permanently.
Seven-day reserved terms
Most reserved GPU capacity is sold in months or years. A 7-day reservation is short enough to cover a single training campaign, which is a genuinely different commercial shape and worth knowing about when a project needs guaranteed capacity for a fortnight and nothing after.
GPU Clusters per GPU-hour. On-demand: HGX H100 $3.99, HGX H200 $5.99, HGX B200 $8.19. Reserved, 7 to 180+ days: H100 $3.19 to $3.69, H200 $3.99 to $4.99, B200 $6.79 to $7.99, with longer commitments yielding better rates. Preemptible: H100 $1.99, H200 $2.99, B200 $4.09. GB200 NVL72 and B300 are contact sales. Dedicated inference on-demand: HGX H100 $3.99 promotional through 09/30/26 against a $5.49 list price, HGX B200 $8.99; H200 and B300 are contact sales. Storage and egress are not published on the pricing page. Checked on together.ai/pricing, 2026-09-18.
RunPod
Best ValueBest for: The cheapest per-GPU rates available self-serve, when you can accept third-party hosting for part of the work
“RunPod splits its catalogue into Community Cloud, where capacity comes from third-party hosts, and Secure Cloud, in vetted data centres, and publishes both prices side by side. That transparency is the product. H100 SXM is $2.69 on Community and $3.49 on Secure, and even the Secure tier undercuts most dedicated providers. The consumer-card catalogue (RTX 5090 at $0.69, RTX 4090 at $0.34, RTX A5000 at $0.16) makes it the cheapest place to iterate.”
Pros
- Lowest published self-serve rates for current chips: H100 PCIe at $1.99 Community, H100 SXM at $2.69, H200 at $3.59, B200 at $5.98, B300 at $6.94
- Secure Cloud remains competitive at $3.49 for H100 SXM while running in vetted data centres
- Deep consumer and workstation catalogue for cheap iteration: RTX 5090 $0.69, RTX 4090 $0.34, RTX A5000 $0.16, A40 $0.35, RTX A6000 $0.33 on Community
- B300 is available self-serve at $6.94 Community and $7.89 Secure, while several dedicated providers only quote Blackwell Ultra
- Storage rates are published in full: container disk $0.10/GB/month, network storage $0.07/GB/month under 1 TB and $0.05 over, high-performance network storage $0.14/GB/month
- No egress charge appears in the pricing documentation
Cons
- Idle volume storage costs $0.20/GB/month, double the $0.10 charged while running, which penalises exactly the pattern of stopping a pod between sessions
- Community Cloud capacity is hosted by third parties, which is a real compliance and data-handling question, not a formality
- Availability of any specific GPU in any specific region varies, and the pricing page is not an availability guarantee
- Multi-node training with a high-performance interconnect is not the platform's strength compared with CoreWeave or Lambda clusters
Community versus Secure is a compliance decision
Community Cloud capacity is supplied by third-party hosts. For public datasets, open-weight models and experimentation, the roughly 25% saving is free money. For customer data, regulated workloads or anything under a data-processing agreement, Secure Cloud is the only defensible choice, and at $3.49 per H100-hour it is still cheaper than most dedicated providers.
Where consumer cards actually win
An RTX 4090 at $0.34 an hour has 24 GB of memory. Paired with a memory-efficient training library, that is enough for LoRA fine-tuning on small and mid-size models, for embedding generation and for most inference experimentation. A large share of work that gets run on an H100 does not need one.
Per hour, Community / Secure. B300 $6.94 / $7.89. B200 $5.98 / $6.79. H200 $3.59 / $4.59. H100 NVL $2.59 / $3.19. H100 SXM $2.69 / $3.49. H100 PCIe $1.99 / $2.89. RTX PRO 6000 $1.69 / $2.09. A100 SXM $1.39 / $1.59. A100 PCIe $1.19 / $1.59. L40S $0.79 / $1.09. RTX 6000 Ada $0.74 / $0.84. L40 $0.69 / $0.82. RTX 5090 $0.69 / $0.99. L4 $0.44 / $0.49. A40 $0.35 / $0.49. RTX 4090 $0.34 / $0.74. RTX A6000 $0.33 / $0.53. RTX 3090 $0.22 / $0.50. RTX A5000 $0.16 / $0.27. Storage: container disk $0.10/GB/month, volume disk $0.10/GB/month running and $0.20/GB/month idle, network storage $0.07/GB/month under 1 TB and $0.05/GB/month over, high-performance network storage $0.14/GB/month. No egress charge listed. Checked on runpod.io/pricing, 2026-09-18.
Modal
FastestBest for: Bursty and intermittent workloads that should cost nothing when idle
“Modal is a serverless compute platform rather than a GPU rental service, and it prices accordingly: per second, with scale to zero. H100 SXM5 is $0.001097 per second, which is $3.95 per GPU-hour, and B300 is $0.001972 per second, which is $7.10. The reason to choose it is not the hourly rate, it is that an endpoint handling traffic for four hours a day costs a sixth of what an always-on instance does, with no orchestration work from you.”
Pros
- Per-second billing with scale to zero, so idle capacity costs nothing
- Competitive effective hourly rates: H100 SXM5 $3.95, H200 $4.54, B200 $6.25, B300 $7.10, A100 80GB $2.50, A100 40GB $2.10, L40S $1.95, A10 $1.10, L4 $0.80, T4 $0.59
- CPU and memory are billed separately and transparently at $0.0000131 per physical core-second and $0.00000222 per GiB-second, so mixed workloads are costable
- Storage is $0.09 per GiB per month with the first 1 TiB free
- Starter plan is $0 with $30 a month of included compute, which is enough to evaluate the platform properly
- Blackwell Ultra (B300) is available self-serve, which several dedicated providers do not offer
Cons
- Team plan carries a $250 a month platform fee against $100 a month of included compute, so the fixed cost is real
- Serverless abstraction means less control over placement, interconnect and node topology than a raw GPU rental
- Not the right shape for a long multi-node training run with a high-performance fabric
- The pricing page does not publish an egress policy
- A minimum of 0.125 physical CPU cores is billed alongside every workload
Where per-second billing actually pays
Take an inference endpoint receiving real traffic for four hours a day. On a dedicated H100 instance at $3.99 an hour you pay for 24, which is roughly $2,900 a month. On Modal at $3.95 an hour of actual use you pay for 4, which is roughly $480. The saving is not in the rate, it is in the 20 hours you do not buy. Invert the workload to a continuous training run and the arithmetic disappears entirely.
Cold starts
Scale to zero means the first request after an idle period pays a start-up cost. That is the trade for not paying while idle and it is the thing to test with your actual model size before committing. Our inference and model hosting comparison covers cold-start behaviour across serving platforms in more detail.
Per second: B300 $0.001972, B200 $0.001736, H200 SXM $0.001261, H100 SXM5 $0.001097, RTX PRO 6000 $0.000842, A100 80GB $0.000694, A100 40GB $0.000583, L40S $0.000542, A10 $0.000306, L4 $0.000222, T4 $0.000164. Those are $7.10, $6.25, $4.54, $3.95, $3.03, $2.50, $2.10, $1.95, $1.10, $0.80 and $0.59 per hour respectively. CPU $0.0000131 per physical core-second with a 0.125 core minimum; memory $0.00000222 per GiB-second. Storage $0.09 per GiB per month with 1 TiB free. Plans: Starter $0 base with $30 a month included compute; Team $250 a month base with $100 a month included compute; Enterprise custom with volume discounts. Checked on modal.com/pricing, 2026-09-18.
Crusoe
Honorable MentionBest for: The lowest published on-demand H100 and H200 rates from a dedicated provider, with free ingress and egress
“Crusoe publishes $3.90 per GPU-hour for HGX H100 and $4.29 for HGX H200, the lowest on-demand rates from a dedicated provider in this comparison, and states plainly that it does not charge for network ingress or egress. Storage is cheap at $0.07 to $0.08 per GiB per month. It is also the only provider here publishing an AMD option, MI300X at $3.45. The catch is that everything current-generation is behind a sales conversation.”
Pros
- Lowest published dedicated-provider on-demand rates: H100 $3.90 and H200 $4.29 per GPU-hour
- States in writing that it does not charge for network ingress or egress
- Cheap storage: persistent disks $0.08 per GiB per month, shared disks $0.07, object storage $0.06
- A100 80GB SXM at $2.30 and PCIe at $2.00, plus L40S at $1.50, cover cost-sensitive workloads
- AMD MI300X published at $3.45 per GPU-hour, a genuine alternative for teams wanting to reduce NVIDIA dependence
Cons
- B200, GB200 NVL72 and MI355X are all contact sales, so no current-generation NVIDIA capacity is self-serve priced
- Spot pricing requires contacting sales, with no public rates at all
- Reserved and commitment discounts are described only as tailored agreements, with no figures published
- Smaller catalogue and smaller footprint than CoreWeave, which matters for regional availability
- The published prices cover the previous generation, so a buyer planning a Blackwell migration cannot price the path
The AMD option
MI300X at $3.45 per GPU-hour with 192 GB of memory is the only published AMD price in this comparison. For inference on large models, memory capacity per GPU often matters more than raw throughput, and 192 GB against an H100's 80 GB changes how many cards a model needs. The cost is software maturity: your stack has to work on ROCm, and that is a real porting question rather than a flag change.
On-demand per GPU-hour: NVIDIA H200 141GB HGX $4.29, H100 80GB HGX $3.90, A100 80GB SXM $2.30, A100 80GB PCIe $2.00, L40S 48GB $1.50, AMD MI300X 192GB $3.45. NVIDIA GB200 NVL72, NVIDIA B200 180GB HGX and AMD MI355X 288GB are contact sales. All spot pricing is contact sales. Reserved terms are described as tailored agreements with no published figures. Storage: persistent disks $0.08 per GiB per month, shared disks $0.07, object storage $0.06. The page states that Crusoe Cloud does not charge for network ingress or egress. Checked on crusoe.ai/cloud/pricing, 2026-09-18.
Vast.ai
Honorable MentionBest for: The cheapest possible floor price, when you can tolerate a marketplace and interruption
“Vast.ai is a marketplace rather than a cloud. Prices are set by supply and demand across more than 40 data centres and over 68 GPU types, billed per second with no minimum and no rounding. Three purchase modes exist: on-demand with guaranteed uptime, interruptible described as 50% or more cheaper, and reserved for 1, 3 or 6 months at up to 50% off. It routinely produces the lowest rate available for a given chip, and it publishes no fixed price at all.”
Pros
- Marketplace competition across 40+ data centres typically produces the lowest available rate for any given chip
- Per-second billing with no minimum hours and no rounding up
- Three modes with clear trade-offs: on-demand with guaranteed uptime, interruptible at 50%+ cheaper, and reserved at up to 50% off for 1, 3 or 6 months
- 68+ GPU types, including both current data-centre parts and older consumer hardware
- Suits fault-tolerant training and batch inference exceptionally well
Cons
- No exact prices are published anywhere on the pricing page, so you cannot budget without opening the console
- No storage or bandwidth pricing is published either
- Hosts are independent operators, which makes this the weakest fit in this comparison for regulated data
- Quality, network and reliability vary by host, and the platform's rating system is the only signal
- Prices move with demand, so a cost model built this month may not hold next month
How to use a marketplace safely
Treat it as spot capacity regardless of which mode you buy. Checkpoint aggressively, keep the dataset somewhere you control rather than on the instance, assume any given host can disappear, and never put regulated or customer data on it. Under those constraints it is the cheapest compute in this comparison by a comfortable margin.
No fixed prices published. Rates are set by supply and demand across 40+ data centres and 68+ GPU types, and are visible only in the live console. Billing is per second with no minimum hours and no rounding up. Purchase modes: On-Demand with guaranteed uptime and no interruptions; Interruptible, described as 50%+ cheaper for preemptible workloads; Reserved at up to 50% off with 1, 3 or 6-month commitments. No storage or bandwidth charges are published on the pricing page. Checked on vast.ai/pricing, 2026-09-18.
Which One Should You Pick?
| Use Case | Our Recommendation |
|---|---|
| Multi-node pretraining or large-scale fine-tuning with a fast interconnect | CoreWeave or Lambda 1-Click Clusters. CoreWeave publishes rack-scale GB200 NVL72 at $42.00 per hour for 4 GPUs and charges nothing for egress; Lambda clusters run from 2 weeks to 1 year at $5.54 to $6.16 per H100-hour. Both are more expensive per GPU than the marketplaces and both are the right answer here. |
| A training run that checkpoints cleanly and can survive interruption | Together preemptible at $1.99 per H100-hour, or Vast.ai interruptible. You are paying roughly half for the same silicon. If your training code cannot resume from a checkpoint, fixing that is a day of work that permanently halves this bill. |
| An inference endpoint with real traffic for a few hours a day | Modal, for per-second billing and scale to zero. Four hours of daily use on an H100 is roughly $480 a month against roughly $2,900 for an always-on instance. Test cold-start latency with your actual model size before committing. |
| Prototyping a training script before scaling it | Lambda for a single H100 at $4.29, only 7.5% above its 8-GPU rate, so your cost model does not change shape when you scale. Or RunPod consumer cards at $0.16 to $0.69 an hour if the model fits in 24 to 32 GB. |
| Customer data or a regulated workload | CoreWeave, Crusoe, Lambda or RunPod Secure Cloud. Avoid RunPod Community and Vast.ai entirely: both source capacity from third-party hosts, which is a data-handling question your security review will not wave through. |
| A pipeline that moves tens of terabytes in and out every month | CoreWeave or Crusoe, both of which state in writing that data transfer is free, or Lambda, which states no egress fees. Together and Modal do not publish an egress policy, so get it in writing before comparing their rates. |
| You only need to run inference on a stock open-weight model | Do not buy GPU capacity at all. A serverless inference API charges per token and costs nothing when idle. Renting a GPU to serve a model someone else already hosts is the most common and most expensive mistake in this category. |
| You want to reduce NVIDIA dependence | Crusoe publishes AMD MI300X at $3.45 per GPU-hour with 192 GB of memory, the only published AMD rate here. Budget engineering time for the ROCm port; the discount is not free. |
How we evaluated
Last verified: 18 September 2026. Every price on this page was read from the provider's own pricing page on that date, and each entry names the page. No figure is carried from a comparison site, a pricing aggregator or a vendor's marketing blog. That includes the hyperscaler section. The AWS rates quoted are the per-accelerator figures AWS publishes on its EC2 Capacity Blocks for ML pricing page, because AWS's own P5 instance page publishes no on-demand hourly rate at all.
Where a provider publishes nothing, the page says so. Vast.ai publishes no fixed prices anywhere, only a marketplace model and three purchase modes. Together publishes no storage rate and no egress policy. Modal publishes no egress policy. Lambda publishes no storage rate on its pricing page. Crusoe publishes no spot rates and no reserved discounts. Those gaps are reported as findings, not filled in.
Where a provider sells by node, the per-GPU figure in the comparison table is derived by dividing the published node price by its GPU count, and the derivation is shown. That is an honest comparison number and a misleading purchase number, because you cannot buy a fraction of a CoreWeave node.
The criteria, in the order they decide real outcomes:
- Published price per GPU-hour, by chip class. H100, H200, B200 and Blackwell Ultra, because that is how the decision is actually framed and because mixing generations is how comparisons become dishonest.
- Purchase modes and commitment terms. On-demand, reserved, preemptible and spot, with the term lengths each provider requires. A price without its availability guarantee is not a price.
- Availability of current-generation parts. Specifically whether Blackwell and Blackwell Ultra capacity is self-serve priced or contact sales, because that distinction now separates the providers more than their rates do.
- Storage. Published per-GB-per-month rates, including the difference between running and idle volumes, which is where several providers recover the discount they gave on compute.
- Egress. Stated in writing, or not stated. This is the largest single hidden cost in a training pipeline and the one most often left out of a comparison.
- Minimum purchase unit and billing granularity. Per second, per GPU, or per 8-GPU node. This decides whether a small job is cheap or absurd.
- Hosting model. Whether capacity comes from vetted data centres or third-party hosts, which is a compliance decision rather than a price one.
Vendor selection
Seven providers were chosen to cover the distinct shapes this market has: hyperscale specialist (CoreWeave), transparent per-GPU rental (Lambda, Crusoe), rental plus a serving and tuning platform (Together), two-tier marketplace (RunPod), pure marketplace (Vast.ai), and serverless per-second compute (Modal). The hyperscalers are covered as a pricing reference point rather than as entries, because their GPU offerings are bought for reasons that have little to do with GPU price.
One ownership change is worth naming: CoreWeave completed its acquisition of Weights & Biases on 5 May 2025. A team buying both GPU capacity and ML tooling may be consolidating onto one vendor without intending to.
What we did not do
No provider on this page was benchmarked. No instance was launched, no throughput was measured, and there are no hands-on testing claims anywhere in this comparison. Performance on identical silicon is dominated by interconnect, storage throughput and your own code, none of which generalise from a vendor test. What this page offers instead is an accurate, dated price comparison built from primary sources, an explicit statement of each provider's weakest point, and the non-obvious costs that change the total.
Prices in this category move faster than in any other we track. Re-check the provider's own page before you commit, and treat any rate on any comparison site, including this one, as a snapshot rather than a quote.
Editorial independence: this is a vendor-neutral comparison with no paid placements, sponsorships or affiliate links. Rankings reflect fit for the stated use cases, not commercial relationships.
Sources
- CoreWeave pricing
- Lambda pricing and Lambda GPU Cloud
- Together AI pricing
- RunPod pricing
- Vast.ai pricing
- Modal pricing
- Crusoe Cloud pricing
- AWS EC2 Capacity Blocks for ML pricing
- AWS EC2 P5 instances
- Fireworks AI pricing (B300 and GB300 reference rates)
- CoreWeave completes acquisition of Weights & Biases
Frequently Asked Questions
What does an H100 actually cost per hour in 2026?
What changed with current-generation accelerators?
Where does the quoted hourly rate stop being the real cost?
How do the hyperscalers price against these providers?
What is the difference between on-demand, reserved, preemptible and spot?
Who should not buy GPU cloud capacity at all?
Is a marketplace like Vast.ai or RunPod Community safe to use?
How stable are these prices?
Related Comparisons
LLM Evaluation and Prompt Management
Top 7 LLM Evaluation and Prompt Management Platforms 2026: How to Stop Shipping Prompt Changes Blind
7 tools compared
Fine-Tuning and Model Customization
Top 8 Fine-Tuning and Model Customization Platforms 2026: What It Costs, When It Wins, and Why OpenAI Is Shutting Its Own Down
8 tools compared
AI Legal / Contract
Top 5 AI Legal and Contract Tools 2026: Harvey vs Spellbook vs Ironclad vs LegalOn vs Luminance
5 tools compared
AI Sales / SDR
Top 5 AI Sales / SDR Tools in 2026
5 tools compared