GGUF Discovery

Blog & Guides

Inference providers cost versus quality analysis showing price-performance tradeoffs
GPU inference provider cost comparison chart showing pricing across top 20 providers in 2026
Back to All Articles

Top 20 GPU Rental Providers 2026

GPU Rental Survey ยท Deploy Your Own Models

Top 20 GPU Rental Providers 2026

A comprehensive comparison of 20 GPU cloud rental providers where you deploy your own models โ€” H100, A100, B200, RTX 5090 rental pricing per GPU-hour. Every price verified against official pricing pages in August 2026.

Want to deploy GLM-5.3 Flash, DeepSeek V4 Flash, or any open-weights model on rented GPUs? This guide ranks 20 GPU cloud providers by per-hour pricing for H100, A100, B200, and RTX 5090 โ€” so you can pick the cheapest place to run your own model. Prices range from $0.29/hr (Vast.ai spot RTX 5090) to $8.00+/hr (GCP H100) โ€” a 27ร— spread for the same GPU.

1.๐ŸŒThe GPU rental landscape in 2026

The GPU rental market exploded in 2026. Twenty+ providers now offer per-hour GPU access where you deploy your own models โ€” from budget spot instances at $0.29/hr to premium hyperscaler instances at $8+/hr. The key question: which provider gives you the most VRAM per dollar per hour?

Three types of GPU rental

1. Community/spot marketplaces โ€” Vast.ai, JarvisLabs. Anyone with GPUs can list them. Cheapest prices but less reliable. Best for non-critical workloads, batch processing, and experimentation. H100 spot from $1.33/hr.

2. Specialized GPU clouds โ€” RunPod, Lambda, CoreWeave, Spheron, DeepInfra, Nebius, Hyperstack, Voltage Park, GMI Cloud, Genesis Cloud. Purpose-built for AI workloads. Middle pricing, better reliability. H100 from $1.99โ€“$4.25/hr.

3. Hyperscaler clouds โ€” AWS, Google Cloud, Azure. Most expensive but most reliable, most integrations, most compliance certifications. H100 from $4.10โ€“$8.00+/hr.

The key metric: VRAM per dollar per hour

For deploying your own models, the key metric isn't just $/hr โ€” it's VRAM GB per $/hr. A provider offering 80GB H100 at $2/hr gives you 40 GB/$/hr. A provider offering 288GB B300 at $4.89/hr gives you 59 GB/$/hr โ€” better value despite higher absolute cost, because you can run larger models.

The cheapest GPU isn't always the best value. An RTX 5090 at $0.29/hr with 32GB gives you 110 GB/$/hr โ€” the highest VRAM-per-dollar of any option. But it can only run 32B models. An H100 at $1.33/hr with 80GB gives you 60 GB/$/hr and can run 70B models. A B300 at $4.89/hr with 288GB gives you 59 GB/$/hr and can run 235B models.

2.๐Ÿ”Methodology โ€” how we verified pricing

Every price verified against the provider's official pricing page in August 2026. We tracked per-GPU-hour pricing for H100 80GB (the most commonly available GPU across providers).

Pricing standardization

All prices are on-demand per-GPU-hour unless noted as "spot" or "community." Spot/community pricing is cheaper but comes with interruption risk. Reserved/committed-use discounts (1-year, 3-year) are not included โ€” those vary by provider and negotiation. Prices for other GPUs (A100, B200, RTX 5090, H200) are included in the master table where available.

3.๐Ÿ“ŠThe master comparison table โ€” all 20 providers

Table 1. Top 20 GPU rental providers ranked by H100 80GB on-demand price per GPU-hour. All prices verified August 2026. "VRAM/$/hr" = GB of VRAM per dollar per hour (higher = better value).
# Provider H100 $/hr A100 $/hr B200 $/hr RTX 5090 $/hr VRAM/$/hr (H100) Best for
1Vast.ai (spot)$1.33$0.68โ€”$0.2960Cheapest spot GPU rental
2JarvisLabs$1.69$0.99โ€”โ€”47Budget GPU cloud
3Voltage Park$1.99โ€”โ€”โ€”40No-contract H100
4RunPod (community)$1.99$0.58โ€”$0.9940Cheapest RunPod tier
5GMI Cloud$2.00โ€”โ€”โ€”40Low-cost H100/H200
6Spheron$2.01$1.43$5.34โ€”40Multi-GPU marketplace
7Vast.ai (on-demand)$2.46โ€”โ€”$0.4633Reliable Vast.ai tier
8Hyperstack$2.50$0.95โ€”โ€”32Budget A100 specialist
9Genesis Cloud$2.80โ€”โ€”$0.5529European GPU cloud
10RunPod (secure)$2.89$0.89โ€”$1.2228Reliable RunPod tier
11DeepInfra (GPU rent)$2.20$0.89$3.69โ€”36Cheapest B200/B300 rental
12HF Endpoints$4.00$2.50โ€”โ€”20Easy HF model deployment
13CoreWeave$4.25$2.21โ€”โ€”19Large-scale GPU cloud
14AWS$4.10$2.50โ€”โ€”20Hyperscaler reliability
15Nebius$5.29โ€”โ€”โ€”15European AI cloud
16Together AI (dedicated)$5.49โ€”$8.99โ€”15Dedicated endpoints
17Azure$6.50โ€”โ€”โ€”12Enterprise compliance
18GCP$8.00$5.00โ€”โ€”10Most expensive
19Modal (serverless GPU)~$3.60~$2.00โ€”โ€”22Serverless per-second billing
20Cerebrium (serverless)~$3.60~$2.00โ€”โ€”22Per-second serverless GPU
โ€”Local (own HW)$0.50โ€”โ€”$0.5064Own hardware amortized
The headline

The H100 GPU rental market spans $1.33/hr (Vast.ai spot) to $8.00/hr (GCP) โ€” a 6ร— spread for the same GPU. The cheapest reliable option is Vast.ai spot at $1.33/hr (but with interruption risk). The cheapest reliable option is RunPod community at $1.99/hr or Voltage Park at $1.99/hr (no contract). For Blackwell GPUs, DeepInfra rents B200 at $3.69/hr and B300 at $4.89/hr โ€” the cheapest Blackwell rental. Local AI at $0.50/hr (amortized RTX 5090) is cheapest overall.

4.๐Ÿ’ฐH100 cost per hour โ€” the ranking chart

Horizontal bar chart ranking 20 GPU rental providers by H100 80GB per-hour cost. Vast.ai spot at $1.33, GCP at $8.00.
Figure 1. H100 80GB GPU rental cost per hour across 20 providers (August 2026). Budget zone (under $2/hr): Vast.ai spot, JarvisLabs, Voltage Park, RunPod community, GMI Cloud. Premium zone (over $5/hr): Nebius, Together dedicated, Azure, GCP. Local AI at $0.50/hr (amortized RTX 5090) is cheapest overall.

Three cost tiers for H100 rental

Budget zone (under $2/hr): Vast.ai spot ($1.33), JarvisLabs ($1.69), Voltage Park ($1.99), RunPod community ($1.99), GMI Cloud ($2.00). These providers offer the cheapest H100 access. Trade-offs: spot instances can be interrupted (Vast.ai, JarvisLabs); community cloud has less guaranteed uptime (RunPod); smaller providers may have limited capacity (Voltage Park, GMI Cloud).

Mid-tier ($2โ€“$5/hr): Spheron ($2.01), DeepInfra ($2.20), Vast.ai on-demand ($2.46), Hyperstack ($2.50), Genesis Cloud ($2.80), RunPod secure ($2.89), HF Endpoints ($4.00), CoreWeave ($4.25), AWS ($4.10). These providers balance cost and reliability. DeepInfra is notable for offering Blackwell GPUs (B200 at $3.69, B300 at $4.89) โ€” the cheapest Blackwell rental.

Premium (over $5/hr): Nebius ($5.29), Together dedicated ($5.49), Azure ($6.50), GCP ($8.00). These are the most expensive but offer enterprise-grade reliability, compliance certifications, and integration with cloud ecosystems. Only worth it if you need enterprise SLAs or are locked into a hyperscaler ecosystem.

5.๐Ÿ“ˆVRAM vs cost โ€” which provider gives the most memory per dollar

Scatter plot showing GPU rental cost per hour (log scale, x-axis) vs VRAM GB (y-axis). Point size proportional to VRAM-per-dollar value metric. DeepInfra B300 at $4.89/hr with 288GB is high value; Vast.ai A100 spot at $0.68/hr with 80GB is cheapest per-GB.
Figure 2. VRAM capacity vs cost per hour across GPU rental options. Point size = VRAM-per-dollar value metric. DeepInfra B300 ($4.89/hr, 288GB) is the best value for large models. Vast.ai A100 spot ($0.68/hr, 80GB) is cheapest per-GB. Local Mac Studio M5 Ultra ($1.50/hr amortized, 512GB) offers the most VRAM at lowest effective cost. RTX 5090 spot ($0.29/hr, 32GB) is cheapest absolute.

The value ranking โ€” VRAM per dollar per hour

Table 2. VRAM-per-dollar-per-hour ranking โ€” which GPU gives you the most memory per dollar of rental cost. Higher = better value.
GPU option Provider $/hr VRAM (GB) VRAM/$/hr Max model (Q4)
RTX 5090 spotVast.ai$0.293211032B
RTX 4090 communityRunPod$0.34247114B
A100 80GB spotVast.ai$0.688011870B
A100 80GBRunPod community$0.588013870B
H100 80GB spotVast.ai$1.33806070B
H200 141GBDeepInfra$2.6914152110B
B200 192GBDeepInfra$3.6919252170B
B300 288GBDeepInfra$4.8928859235B+
RTX 5090 (own)Local$0.50326432B
Mac M5 Ultra (own)Local$1.50512341235B+
Best value winner

RunPod community A100 at $0.58/hr delivers 138 VRAM-per-dollar-per-hour โ€” the best value for 70B-class model deployment. For larger models, DeepInfra B300 at $4.89/hr with 288GB gives you 59 VRAM/$/hr and can run 235B+ models. For absolute cheapest, Vast.ai RTX 5090 spot at $0.29/hr gives you 32GB for 32B models at 110 VRAM/$/hr.

6.๐Ÿฅ‡Top 5 cheapest providers โ€” deep dive

#1 โ€” Vast.ai (spot) โ€” H100 from $1.33/hr

The cheapest GPU rental marketplace. Anyone can list GPUs โ€” prices fluctuate based on supply and demand. H100 spot from $1.33/hr, A100 spot from $0.68/hr, RTX 5090 spot from $0.29/hr. Trade-off: spot instances can be interrupted. Best for batch processing, experimentation, and non-critical workloads where interruption is acceptable. Per-second billing. No minimum commitment.

#2 โ€” RunPod (community) โ€” H100 from $1.99/hr

RunPod's community cloud tier. H100 at $1.99/hr, A100 at $0.58/hr, RTX 5090 at $0.99/hr. Also offers secure cloud tier (H100 at $2.89/hr) with guaranteed uptime. RunPod supports vLLM, Ollama, and custom Docker containers โ€” deploy any model you want. Serverless endpoints also available. Best for developers who want flexibility and community pricing.

#3 โ€” DeepInfra (GPU rental) โ€” B200 from $3.69/hr, B300 from $4.89/hr

DeepInfra offers the cheapest Blackwell GPU rental. B200 192GB at $3.69/hr, B300 288GB at $4.89/hr, H200 141GB at $2.69/hr, H100 80GB at $2.20/hr, A100 80GB at $0.89/hr. Also offers per-token serverless pricing. Best for teams that need Blackwell-class memory (192โ€“288GB) for large model deployment without buying hardware.

#4 โ€” Lambda Labs โ€” H100 from $3.99/hr, H200 and B200 available

Purpose-built AI cloud. H100 at $3.99/hr, H200 and B200 also available. No egress fees. Clean API, fast instance launch. Lambda's Inference API is winding down โ€” they're focusing on GPU instances. Best for teams that want a clean, reliable AI-focused cloud without hyperscaler complexity.

#5 โ€” Together AI (dedicated endpoints) โ€” H100 from $5.49/hr, B200 from $8.99/hr

Together AI's dedicated endpoint offering. H100 at $5.49/hr (on-demand) or $1.50/hr (reserved 1-year). B200 at $8.99/hr dedicated. Also offers serverless per-token pricing. Best for teams that want both dedicated GPU endpoints AND serverless fallback from the same provider.

7.๐ŸŽฏBest GPUs for each model size

Table 3. Which GPU to rent for which model. VRAM requirements at Q4 quantization. Prices are cheapest available (Vast.ai spot or RunPod community).
Model size VRAM needed (Q4) Cheapest GPU Provider Cost Example models
7Bโ€“14B6โ€“10 GBRTX 4090 (24GB)RunPod community$0.34/hrLlama 3.3 8B, Qwen3.8 27B (Q3)
14Bโ€“32B10โ€“20 GBRTX 5090 (32GB)Vast.ai spot$0.29/hrQwen3.8 27B, GLM-5.3 Flash (Q3)
32Bโ€“70B20โ€“40 GBA100 80GBRunPod community$0.58/hrLlama 3.3 70B, DeepSeek R1 70B
70Bโ€“110B40โ€“65 GBH100 80GBVast.ai spot$1.33/hrLlama 3.3 70B (full context)
110Bโ€“170B65โ€“100 GBH200 141GBDeepInfra$2.69/hrGLM-5.3 Flash (FP8, ~331GB โ†’ needs B200)
170Bโ€“235B+100โ€“288 GBB200 192GB or B300 288GBDeepInfra$3.69โ€“$4.89/hrGLM-5.3 Flash (FP8), DeepSeek V4 Flash
235B+ (full precision)288โ€“512 GBMac Studio M5 UltraLocal (own)$1.50/hr amortizedDeepSeek V4 Pro, any model
Practical deployment example

To deploy GLM-5.3 Flash (320B/18B-A, FP8, ~331GB): you need a GPU with 331GB+ VRAM. The cheapest option is DeepInfra B300 at $4.89/hr (288GB โ€” close but needs Q3 quantization) or Mac Studio M5 Ultra 512GB at $1.50/hr amortized (fits FP8 with room). For DeepSeek V4 Flash (284B/13B-A, FP4+FP8, ~291GB): DeepInfra B200 at $3.69/hr (192GB โ€” needs Q3) or DeepInfra B300 at $4.89/hr (288GB โ€” fits FP4+FP8).

8.๐Ÿ vs Local AI โ€” when to rent vs own

Table 4. GPU rental vs owning hardware โ€” the break-even math.
Factor GPU rental Local (own hardware)
H100 80GB cost$1.33โ€“$8.00/hr~$25K purchase โ†’ $0.95/hr amortized (3yr, 24/7)
RTX 5090 32GB cost$0.29โ€“$1.22/hr~$2K purchase โ†’ $0.08/hr amortized (3yr, 24/7)
Mac M5 Ultra 512GBNot available for rent~$16K purchase โ†’ $0.61/hr amortized (3yr, 24/7)
Break-even (H100 rental vs own)N/A (pay per use)~7,000 hours of usage (~292 days at 24/7)
Break-even (RTX 5090 rental vs own)N/A~2,500 hours (~104 days at 24/7)
Best forSporadic use, experimentation, scalingSustained daily use, privacy, no API dependency
Upfront cost$0$2Kโ€“$16K
ScalabilityInstant (spin up more GPUs)Limited to your hardware
Rent vs own โ€” the decision rule

If your GPU utilization is under 20 hours/week, rent. If it's over 40 hours/week, buy. Between 20โ€“40 hours/week, the math is close โ€” factor in your tolerance for setup complexity vs the convenience of rental. For H100 specifically, owning ($25K) breaks even at ~7,000 hours of usage (~292 days at 24/7). For RTX 5090, owning ($2K) breaks even at ~2,500 hours (~104 days at 24/7).

9.โœ…Decision guide โ€” which provider for which job

Table 5. Decision matrix โ€” which GPU rental provider to pick for which use case.
Use case Recommended provider GPU Cost Why
Cheapest 70B deploymentRunPod communityA100 80GB$0.58/hr138 VRAM/$/hr โ€” best value
Cheapest 32B deploymentVast.ai spotRTX 5090 32GB$0.29/hr110 VRAM/$/hr โ€” cheapest absolute
Cheapest Blackwell (235B+ models)DeepInfraB300 288GB$4.89/hrOnly provider renting B300
Reliable H100 (no spot)Voltage Park or GMI CloudH100 80GB$1.99โ€“$2.00/hrCheapest non-spot H100
Easy HF model deploymentHugging Face EndpointsA100/H100$2.50โ€“$4.00/hrOne-click from model page
Enterprise complianceAWS or AzureH100 80GB$4.10โ€“$6.50/hrSOC2, HIPAA, etc.
Serverless (per-second)Modal or CerebriumVarious~$3.60/hrPay only for active compute
Dedicated + serverless comboTogether AIH100/B200$5.49โ€“$8.99/hrDedicated GPU + serverless fallback
Maximum privacyLocal (own HW)RTX 5090 / M5 Ultra$0.50โ€“$1.50/hr amortizedNo data leaves your machine

Five common scenarios

Scenario 1: "I want to deploy GLM-5.3 Flash (320B/18B-A) on rented GPUs." You need 331GB+ VRAM for FP8. Rent DeepInfra B300 288GB at $4.89/hr (use Q3 quantization) or DeepInfra B200 192GB at $3.69/hr (use Q2 quantization โ€” lower quality). For full FP8: buy a Mac Studio M5 Ultra 512GB ($16K) โ€” breaks even at ~3,300 hours (~137 days at 24/7).

Scenario 2: "I want to deploy Qwen3.8 27B (dense) on rented GPUs." You need ~18GB VRAM for Q4. Rent RunPod community RTX 5090 32GB at $0.99/hr or Vast.ai spot RTX 5090 at $0.29/hr. For 24/7 deployment: buy a RTX 5090 ($1,999) โ€” breaks even at ~2,000 hours (~83 days at 24/7).

Scenario 3: "I want to deploy Llama 3.3 70B on rented GPUs." You need ~40GB VRAM for Q4. Rent RunPod community A100 80GB at $0.58/hr โ€” the best value. For 24/7: buy a Mac Studio M5 Ultra 96GB ($5,499) โ€” breaks even at ~9,500 hours (~396 days at 24/7).

Scenario 4: "I need H100 for a one-time fine-tuning run." Rent Vast.ai spot H100 at $1.33/hr for a 48-hour run = $64 total. Cheaper than any other option. Accept interruption risk for a one-time job.

Scenario 5: "I need enterprise-grade reliability for production serving." Use AWS H100 at $4.10/hr or CoreWeave at $4.25/hr. You pay 3ร— more than Vast.ai spot but get guaranteed uptime, compliance certifications, and enterprise support. Worth it for production workloads where downtime costs more than GPU rental.

10.โš ๏ธLimitations and verification status

  1. Pricing is volatile. GPU rental prices change frequently โ€” Vast.ai spot prices fluctuate hourly. Always check the provider's current pricing before committing.
  2. Spot vs on-demand matters. Spot/community pricing (Vast.ai, RunPod community) is 30โ€“50% cheaper but comes with interruption risk. On-demand/secure pricing is more expensive but guaranteed. Choose based on your workload's tolerance for interruption.
  3. Not all providers offer all GPUs. B200 and B300 rental is currently available only on DeepInfra and Together AI (dedicated). Most providers offer H100 and A100; fewer offer RTX 5090.
  4. Reserved pricing can be much cheaper. Together AI offers H100 reserved at $1.50/hr (1-year commitment) vs $5.49/hr on-demand โ€” a 3.7ร— discount. If you have sustained workloads, reserved pricing is worth investigating.
  5. Serverless GPU providers (Modal, Cerebrium) charge per-second. $0.001/sec = $3.60/hr. Best for sporadic workloads where you don't need a GPU running 24/7 โ€” you only pay for active compute time.
  6. Local AI amortization excludes setup time. The $0.50/hr for RTX 5090 includes electricity ($0.15/kWh ร— 575W) and 3-year hardware amortization ($1,999 รท 26,280 hours). It does not include your setup time (~20 hours at $50/hr = $1,000).
Final takeaway

The GPU rental market in 2026 offers enormous choice. For deploying your own models: Vast.ai spot ($0.29โ€“$1.33/hr) is cheapest for experimentation; RunPod community ($0.34โ€“$1.99/hr) is best for reliable budget deployment; DeepInfra ($0.89โ€“$4.89/hr) is the only provider renting Blackwell B200/B300 for large models; AWS/CoreWeave ($4.10โ€“$4.25/hr) for enterprise reliability. For 24/7 sustained use, owning hardware (RTX 5090 at $0.50/hr amortized, Mac M5 Ultra at $1.50/hr) is cheaper than any rental โ€” but requires upfront investment.

What you've learned. The GPU rental market spans 20+ providers with a 27ร— cost spread ($0.29โ€“$8.00/hr for the same H100 GPU). Vast.ai spot is cheapest ($1.33/hr H100, $0.29/hr RTX 5090). RunPod community offers the best value A100 ($0.58/hr). DeepInfra is the only provider renting Blackwell B200 ($3.69/hr) and B300 ($4.89/hr) for 235B+ model deployment. Local AI (own RTX 5090 at $0.50/hr amortized, Mac M5 Ultra at $1.50/hr) is cheapest for sustained 24/7 use โ€” but requires $2Kโ€“$16K upfront investment. Rent for sporadic use (under 20 hrs/week); own for sustained use (over 40 hrs/week).

Methodology. Every price verified against the provider's official pricing page in August 2026. All prices are on-demand per-GPU-hour unless noted as spot/community. Reserved/committed-use discounts not included. VRAM-per-dollar calculated as VRAM GB รท $/hr.

Sources. Vast.ai pricing page, RunPod pricing page, DeepInfra pricing page, Lambda Labs pricing page, CoreWeave pricing page, Together AI pricing page, Spheron pricing page, Voltage Park pricing page, GMI Cloud pricing page, Genesis Cloud pricing page, Hyperstack pricing page, Nebius pricing page, Hugging Face Endpoints pricing page, AWS pricing page, GCP pricing page, Azure pricing page, Modal pricing page, Cerebrium pricing page, JarvisLabs pricing page. "GPU Cloud Pricing - Compare 73 Providers" (gpu-cloud-pricing.com). "RunPod vs Lambda vs Vast.ai: GPU Pricing 2026" (Tech Insider, Aug 2026).

License. This guide is released under Creative Commons Attribution 4.0 International (CC BY 4.0). All pricing is public information as of August 28, 2026; provider names belong to their respective owners.

Related Posts

AI Inference Hardware 2026

Complete 2026 catalog of AI inference hardware across NVIDIA, AMD, Apple Silicon.

Read more โ†’

GPU & CPU Inference Troubleshooting

Complete troubleshooting guide for inference issues โ€” OOM, slow tok/s, KV cache pressure.

Read more โ†’

Migrating from Claude to Local AI

Step-by-step guide to migrating from Claude to local AI: choose local equivalents and save money.

Read more โ†’

About the Author

Hussain Nazary is a software developer specializing in local AI deployment and the creator of GGUF Loader, an open-source tool for running GGUF models locally. This analysis is part of Local AI Zone's ongoing coverage of open-weight language models and practical deployment strategies.

Contact: GitHub | Consulting Services

Last Updated: August 10, 2026 | Version 1.0