Top 20 GPU Rental Providers 2026
A comprehensive comparison of 20 GPU cloud rental providers where you deploy your own models โ H100, A100, B200, RTX 5090 rental pricing per GPU-hour. Every price verified against official pricing pages in August 2026.
1.๐The GPU rental landscape in 2026
The GPU rental market exploded in 2026. Twenty+ providers now offer per-hour GPU access where you deploy your own models โ from budget spot instances at $0.29/hr to premium hyperscaler instances at $8+/hr. The key question: which provider gives you the most VRAM per dollar per hour?
Three types of GPU rental
1. Community/spot marketplaces โ Vast.ai, JarvisLabs. Anyone with GPUs can list them. Cheapest prices but less reliable. Best for non-critical workloads, batch processing, and experimentation. H100 spot from $1.33/hr.
2. Specialized GPU clouds โ RunPod, Lambda, CoreWeave, Spheron, DeepInfra, Nebius, Hyperstack, Voltage Park, GMI Cloud, Genesis Cloud. Purpose-built for AI workloads. Middle pricing, better reliability. H100 from $1.99โ$4.25/hr.
3. Hyperscaler clouds โ AWS, Google Cloud, Azure. Most expensive but most reliable, most integrations, most compliance certifications. H100 from $4.10โ$8.00+/hr.
The key metric: VRAM per dollar per hour
For deploying your own models, the key metric isn't just $/hr โ it's VRAM GB per $/hr. A provider offering 80GB H100 at $2/hr gives you 40 GB/$/hr. A provider offering 288GB B300 at $4.89/hr gives you 59 GB/$/hr โ better value despite higher absolute cost, because you can run larger models.
The cheapest GPU isn't always the best value. An RTX 5090 at $0.29/hr with 32GB gives you 110 GB/$/hr โ the highest VRAM-per-dollar of any option. But it can only run 32B models. An H100 at $1.33/hr with 80GB gives you 60 GB/$/hr and can run 70B models. A B300 at $4.89/hr with 288GB gives you 59 GB/$/hr and can run 235B models.
2.๐Methodology โ how we verified pricing
Every price verified against the provider's official pricing page in August 2026. We tracked per-GPU-hour pricing for H100 80GB (the most commonly available GPU across providers).
All prices are on-demand per-GPU-hour unless noted as "spot" or "community." Spot/community pricing is cheaper but comes with interruption risk. Reserved/committed-use discounts (1-year, 3-year) are not included โ those vary by provider and negotiation. Prices for other GPUs (A100, B200, RTX 5090, H200) are included in the master table where available.
3.๐The master comparison table โ all 20 providers
| # | Provider | H100 $/hr | A100 $/hr | B200 $/hr | RTX 5090 $/hr | VRAM/$/hr (H100) | Best for |
|---|---|---|---|---|---|---|---|
| 1 | Vast.ai (spot) | $1.33 | $0.68 | โ | $0.29 | 60 | Cheapest spot GPU rental |
| 2 | JarvisLabs | $1.69 | $0.99 | โ | โ | 47 | Budget GPU cloud |
| 3 | Voltage Park | $1.99 | โ | โ | โ | 40 | No-contract H100 |
| 4 | RunPod (community) | $1.99 | $0.58 | โ | $0.99 | 40 | Cheapest RunPod tier |
| 5 | GMI Cloud | $2.00 | โ | โ | โ | 40 | Low-cost H100/H200 |
| 6 | Spheron | $2.01 | $1.43 | $5.34 | โ | 40 | Multi-GPU marketplace |
| 7 | Vast.ai (on-demand) | $2.46 | โ | โ | $0.46 | 33 | Reliable Vast.ai tier |
| 8 | Hyperstack | $2.50 | $0.95 | โ | โ | 32 | Budget A100 specialist |
| 9 | Genesis Cloud | $2.80 | โ | โ | $0.55 | 29 | European GPU cloud |
| 10 | RunPod (secure) | $2.89 | $0.89 | โ | $1.22 | 28 | Reliable RunPod tier |
| 11 | DeepInfra (GPU rent) | $2.20 | $0.89 | $3.69 | โ | 36 | Cheapest B200/B300 rental |
| 12 | HF Endpoints | $4.00 | $2.50 | โ | โ | 20 | Easy HF model deployment |
| 13 | CoreWeave | $4.25 | $2.21 | โ | โ | 19 | Large-scale GPU cloud |
| 14 | AWS | $4.10 | $2.50 | โ | โ | 20 | Hyperscaler reliability |
| 15 | Nebius | $5.29 | โ | โ | โ | 15 | European AI cloud |
| 16 | Together AI (dedicated) | $5.49 | โ | $8.99 | โ | 15 | Dedicated endpoints |
| 17 | Azure | $6.50 | โ | โ | โ | 12 | Enterprise compliance |
| 18 | GCP | $8.00 | $5.00 | โ | โ | 10 | Most expensive |
| 19 | Modal (serverless GPU) | ~$3.60 | ~$2.00 | โ | โ | 22 | Serverless per-second billing |
| 20 | Cerebrium (serverless) | ~$3.60 | ~$2.00 | โ | โ | 22 | Per-second serverless GPU |
| โ | Local (own HW) | $0.50 | โ | โ | $0.50 | 64 | Own hardware amortized |
The H100 GPU rental market spans $1.33/hr (Vast.ai spot) to $8.00/hr (GCP) โ a 6ร spread for the same GPU. The cheapest reliable option is Vast.ai spot at $1.33/hr (but with interruption risk). The cheapest reliable option is RunPod community at $1.99/hr or Voltage Park at $1.99/hr (no contract). For Blackwell GPUs, DeepInfra rents B200 at $3.69/hr and B300 at $4.89/hr โ the cheapest Blackwell rental. Local AI at $0.50/hr (amortized RTX 5090) is cheapest overall.
4.๐ฐH100 cost per hour โ the ranking chart
Three cost tiers for H100 rental
Budget zone (under $2/hr): Vast.ai spot ($1.33), JarvisLabs ($1.69), Voltage Park ($1.99), RunPod community ($1.99), GMI Cloud ($2.00). These providers offer the cheapest H100 access. Trade-offs: spot instances can be interrupted (Vast.ai, JarvisLabs); community cloud has less guaranteed uptime (RunPod); smaller providers may have limited capacity (Voltage Park, GMI Cloud).
Mid-tier ($2โ$5/hr): Spheron ($2.01), DeepInfra ($2.20), Vast.ai on-demand ($2.46), Hyperstack ($2.50), Genesis Cloud ($2.80), RunPod secure ($2.89), HF Endpoints ($4.00), CoreWeave ($4.25), AWS ($4.10). These providers balance cost and reliability. DeepInfra is notable for offering Blackwell GPUs (B200 at $3.69, B300 at $4.89) โ the cheapest Blackwell rental.
Premium (over $5/hr): Nebius ($5.29), Together dedicated ($5.49), Azure ($6.50), GCP ($8.00). These are the most expensive but offer enterprise-grade reliability, compliance certifications, and integration with cloud ecosystems. Only worth it if you need enterprise SLAs or are locked into a hyperscaler ecosystem.
5.๐VRAM vs cost โ which provider gives the most memory per dollar
The value ranking โ VRAM per dollar per hour
| GPU option | Provider | $/hr | VRAM (GB) | VRAM/$/hr | Max model (Q4) |
|---|---|---|---|---|---|
| RTX 5090 spot | Vast.ai | $0.29 | 32 | 110 | 32B |
| RTX 4090 community | RunPod | $0.34 | 24 | 71 | 14B |
| A100 80GB spot | Vast.ai | $0.68 | 80 | 118 | 70B |
| A100 80GB | RunPod community | $0.58 | 80 | 138 | 70B |
| H100 80GB spot | Vast.ai | $1.33 | 80 | 60 | 70B |
| H200 141GB | DeepInfra | $2.69 | 141 | 52 | 110B |
| B200 192GB | DeepInfra | $3.69 | 192 | 52 | 170B |
| B300 288GB | DeepInfra | $4.89 | 288 | 59 | 235B+ |
| RTX 5090 (own) | Local | $0.50 | 32 | 64 | 32B |
| Mac M5 Ultra (own) | Local | $1.50 | 512 | 341 | 235B+ |
RunPod community A100 at $0.58/hr delivers 138 VRAM-per-dollar-per-hour โ the best value for 70B-class model deployment. For larger models, DeepInfra B300 at $4.89/hr with 288GB gives you 59 VRAM/$/hr and can run 235B+ models. For absolute cheapest, Vast.ai RTX 5090 spot at $0.29/hr gives you 32GB for 32B models at 110 VRAM/$/hr.
6.๐ฅTop 5 cheapest providers โ deep dive
#1 โ Vast.ai (spot) โ H100 from $1.33/hr
The cheapest GPU rental marketplace. Anyone can list GPUs โ prices fluctuate based on supply and demand. H100 spot from $1.33/hr, A100 spot from $0.68/hr, RTX 5090 spot from $0.29/hr. Trade-off: spot instances can be interrupted. Best for batch processing, experimentation, and non-critical workloads where interruption is acceptable. Per-second billing. No minimum commitment.
#2 โ RunPod (community) โ H100 from $1.99/hr
RunPod's community cloud tier. H100 at $1.99/hr, A100 at $0.58/hr, RTX 5090 at $0.99/hr. Also offers secure cloud tier (H100 at $2.89/hr) with guaranteed uptime. RunPod supports vLLM, Ollama, and custom Docker containers โ deploy any model you want. Serverless endpoints also available. Best for developers who want flexibility and community pricing.
#3 โ DeepInfra (GPU rental) โ B200 from $3.69/hr, B300 from $4.89/hr
DeepInfra offers the cheapest Blackwell GPU rental. B200 192GB at $3.69/hr, B300 288GB at $4.89/hr, H200 141GB at $2.69/hr, H100 80GB at $2.20/hr, A100 80GB at $0.89/hr. Also offers per-token serverless pricing. Best for teams that need Blackwell-class memory (192โ288GB) for large model deployment without buying hardware.
#4 โ Lambda Labs โ H100 from $3.99/hr, H200 and B200 available
Purpose-built AI cloud. H100 at $3.99/hr, H200 and B200 also available. No egress fees. Clean API, fast instance launch. Lambda's Inference API is winding down โ they're focusing on GPU instances. Best for teams that want a clean, reliable AI-focused cloud without hyperscaler complexity.
#5 โ Together AI (dedicated endpoints) โ H100 from $5.49/hr, B200 from $8.99/hr
Together AI's dedicated endpoint offering. H100 at $5.49/hr (on-demand) or $1.50/hr (reserved 1-year). B200 at $8.99/hr dedicated. Also offers serverless per-token pricing. Best for teams that want both dedicated GPU endpoints AND serverless fallback from the same provider.
7.๐ฏBest GPUs for each model size
| Model size | VRAM needed (Q4) | Cheapest GPU | Provider | Cost | Example models |
|---|---|---|---|---|---|
| 7Bโ14B | 6โ10 GB | RTX 4090 (24GB) | RunPod community | $0.34/hr | Llama 3.3 8B, Qwen3.8 27B (Q3) |
| 14Bโ32B | 10โ20 GB | RTX 5090 (32GB) | Vast.ai spot | $0.29/hr | Qwen3.8 27B, GLM-5.3 Flash (Q3) |
| 32Bโ70B | 20โ40 GB | A100 80GB | RunPod community | $0.58/hr | Llama 3.3 70B, DeepSeek R1 70B |
| 70Bโ110B | 40โ65 GB | H100 80GB | Vast.ai spot | $1.33/hr | Llama 3.3 70B (full context) |
| 110Bโ170B | 65โ100 GB | H200 141GB | DeepInfra | $2.69/hr | GLM-5.3 Flash (FP8, ~331GB โ needs B200) |
| 170Bโ235B+ | 100โ288 GB | B200 192GB or B300 288GB | DeepInfra | $3.69โ$4.89/hr | GLM-5.3 Flash (FP8), DeepSeek V4 Flash |
| 235B+ (full precision) | 288โ512 GB | Mac Studio M5 Ultra | Local (own) | $1.50/hr amortized | DeepSeek V4 Pro, any model |
To deploy GLM-5.3 Flash (320B/18B-A, FP8, ~331GB): you need a GPU with 331GB+ VRAM. The cheapest option is DeepInfra B300 at $4.89/hr (288GB โ close but needs Q3 quantization) or Mac Studio M5 Ultra 512GB at $1.50/hr amortized (fits FP8 with room). For DeepSeek V4 Flash (284B/13B-A, FP4+FP8, ~291GB): DeepInfra B200 at $3.69/hr (192GB โ needs Q3) or DeepInfra B300 at $4.89/hr (288GB โ fits FP4+FP8).
8.๐ vs Local AI โ when to rent vs own
| Factor | GPU rental | Local (own hardware) |
|---|---|---|
| H100 80GB cost | $1.33โ$8.00/hr | ~$25K purchase โ $0.95/hr amortized (3yr, 24/7) |
| RTX 5090 32GB cost | $0.29โ$1.22/hr | ~$2K purchase โ $0.08/hr amortized (3yr, 24/7) |
| Mac M5 Ultra 512GB | Not available for rent | ~$16K purchase โ $0.61/hr amortized (3yr, 24/7) |
| Break-even (H100 rental vs own) | N/A (pay per use) | ~7,000 hours of usage (~292 days at 24/7) |
| Break-even (RTX 5090 rental vs own) | N/A | ~2,500 hours (~104 days at 24/7) |
| Best for | Sporadic use, experimentation, scaling | Sustained daily use, privacy, no API dependency |
| Upfront cost | $0 | $2Kโ$16K |
| Scalability | Instant (spin up more GPUs) | Limited to your hardware |
If your GPU utilization is under 20 hours/week, rent. If it's over 40 hours/week, buy. Between 20โ40 hours/week, the math is close โ factor in your tolerance for setup complexity vs the convenience of rental. For H100 specifically, owning ($25K) breaks even at ~7,000 hours of usage (~292 days at 24/7). For RTX 5090, owning ($2K) breaks even at ~2,500 hours (~104 days at 24/7).
9.โ Decision guide โ which provider for which job
| Use case | Recommended provider | GPU | Cost | Why |
|---|---|---|---|---|
| Cheapest 70B deployment | RunPod community | A100 80GB | $0.58/hr | 138 VRAM/$/hr โ best value |
| Cheapest 32B deployment | Vast.ai spot | RTX 5090 32GB | $0.29/hr | 110 VRAM/$/hr โ cheapest absolute |
| Cheapest Blackwell (235B+ models) | DeepInfra | B300 288GB | $4.89/hr | Only provider renting B300 |
| Reliable H100 (no spot) | Voltage Park or GMI Cloud | H100 80GB | $1.99โ$2.00/hr | Cheapest non-spot H100 |
| Easy HF model deployment | Hugging Face Endpoints | A100/H100 | $2.50โ$4.00/hr | One-click from model page |
| Enterprise compliance | AWS or Azure | H100 80GB | $4.10โ$6.50/hr | SOC2, HIPAA, etc. |
| Serverless (per-second) | Modal or Cerebrium | Various | ~$3.60/hr | Pay only for active compute |
| Dedicated + serverless combo | Together AI | H100/B200 | $5.49โ$8.99/hr | Dedicated GPU + serverless fallback |
| Maximum privacy | Local (own HW) | RTX 5090 / M5 Ultra | $0.50โ$1.50/hr amortized | No data leaves your machine |
Five common scenarios
Scenario 1: "I want to deploy GLM-5.3 Flash (320B/18B-A) on rented GPUs." You need 331GB+ VRAM for FP8. Rent DeepInfra B300 288GB at $4.89/hr (use Q3 quantization) or DeepInfra B200 192GB at $3.69/hr (use Q2 quantization โ lower quality). For full FP8: buy a Mac Studio M5 Ultra 512GB ($16K) โ breaks even at ~3,300 hours (~137 days at 24/7).
Scenario 2: "I want to deploy Qwen3.8 27B (dense) on rented GPUs." You need ~18GB VRAM for Q4. Rent RunPod community RTX 5090 32GB at $0.99/hr or Vast.ai spot RTX 5090 at $0.29/hr. For 24/7 deployment: buy a RTX 5090 ($1,999) โ breaks even at ~2,000 hours (~83 days at 24/7).
Scenario 3: "I want to deploy Llama 3.3 70B on rented GPUs." You need ~40GB VRAM for Q4. Rent RunPod community A100 80GB at $0.58/hr โ the best value. For 24/7: buy a Mac Studio M5 Ultra 96GB ($5,499) โ breaks even at ~9,500 hours (~396 days at 24/7).
Scenario 4: "I need H100 for a one-time fine-tuning run." Rent Vast.ai spot H100 at $1.33/hr for a 48-hour run = $64 total. Cheaper than any other option. Accept interruption risk for a one-time job.
Scenario 5: "I need enterprise-grade reliability for production serving." Use AWS H100 at $4.10/hr or CoreWeave at $4.25/hr. You pay 3ร more than Vast.ai spot but get guaranteed uptime, compliance certifications, and enterprise support. Worth it for production workloads where downtime costs more than GPU rental.
10.โ ๏ธLimitations and verification status
- Pricing is volatile. GPU rental prices change frequently โ Vast.ai spot prices fluctuate hourly. Always check the provider's current pricing before committing.
- Spot vs on-demand matters. Spot/community pricing (Vast.ai, RunPod community) is 30โ50% cheaper but comes with interruption risk. On-demand/secure pricing is more expensive but guaranteed. Choose based on your workload's tolerance for interruption.
- Not all providers offer all GPUs. B200 and B300 rental is currently available only on DeepInfra and Together AI (dedicated). Most providers offer H100 and A100; fewer offer RTX 5090.
- Reserved pricing can be much cheaper. Together AI offers H100 reserved at $1.50/hr (1-year commitment) vs $5.49/hr on-demand โ a 3.7ร discount. If you have sustained workloads, reserved pricing is worth investigating.
- Serverless GPU providers (Modal, Cerebrium) charge per-second. $0.001/sec = $3.60/hr. Best for sporadic workloads where you don't need a GPU running 24/7 โ you only pay for active compute time.
- Local AI amortization excludes setup time. The $0.50/hr for RTX 5090 includes electricity ($0.15/kWh ร 575W) and 3-year hardware amortization ($1,999 รท 26,280 hours). It does not include your setup time (~20 hours at $50/hr = $1,000).
The GPU rental market in 2026 offers enormous choice. For deploying your own models: Vast.ai spot ($0.29โ$1.33/hr) is cheapest for experimentation; RunPod community ($0.34โ$1.99/hr) is best for reliable budget deployment; DeepInfra ($0.89โ$4.89/hr) is the only provider renting Blackwell B200/B300 for large models; AWS/CoreWeave ($4.10โ$4.25/hr) for enterprise reliability. For 24/7 sustained use, owning hardware (RTX 5090 at $0.50/hr amortized, Mac M5 Ultra at $1.50/hr) is cheaper than any rental โ but requires upfront investment.
Related Posts
AI Inference Hardware 2026
Complete 2026 catalog of AI inference hardware across NVIDIA, AMD, Apple Silicon.
Read more โGPU & CPU Inference Troubleshooting
Complete troubleshooting guide for inference issues โ OOM, slow tok/s, KV cache pressure.
Read more โMigrating from Claude to Local AI
Step-by-step guide to migrating from Claude to local AI: choose local equivalents and save money.
Read more โAbout the Author
Hussain Nazary is a software developer specializing in local AI deployment and the creator of GGUF Loader, an open-source tool for running GGUF models locally. This analysis is part of Local AI Zone's ongoing coverage of open-weight language models and practical deployment strategies.
Contact: GitHub | Consulting Services
Last Updated: August 10, 2026 | Version 1.0