Cloud GPU Pricing: The Complete Picture
The advertised hourly rate of a cloud GPU is not what you actually pay. Real cloud GPU costs include hidden fees, billing granularity penalties, and infrastructure overhead that can inflate your bill by 20–40% beyond the headline price.
This guide focuses on what other pricing comparisons miss: the total cost of running a GPU workload, not just the sticker price. We compare three types of providers, expose hidden costs, and provide five concrete strategies to reduce your GPU spend.
For GPU spec comparisons and workload recommendations, see our best GPU for AI guide. This article focuses on providers and pricing — where you rent, not what you rent.
Three Types of Cloud GPU Providers
The cloud GPU market has three distinct tiers, each with different pricing structures, trade-offs, and target customers.
Hyperscalers (AWS, Google Cloud, Azure)
The major cloud providers offer GPU instances as part of their broad infrastructure platform. You get enterprise SLAs, global regions, deep integration with managed services (storage, ML platforms, monitoring), and guaranteed availability for reserved instances.
The premium: $3.00–$6.98/hr for an H100 on-demand. Reserved instances (1–3 year commitments) reduce this by 30–50%, but require long-term lock-in.
Specialized GPU Clouds (Lambda, CoreWeave)
Focused exclusively on GPU infrastructure. These providers offer competitive pricing ($2.00–$3.00/hr for H100), configurations optimized for AI workloads, and AI-focused support — but fewer managed services than hyperscalers.
GPU Marketplaces (GPUnex, Vast.ai, RunPod)
Aggregate GPU capacity from hundreds of distributed data centers. By pooling supply from many providers, marketplaces offer the most competitive pricing ($1.49–$2.50/hr for H100) with flexible billing — often per-second rather than per-hour. Trade-offs include variable reliability and less integrated tooling.
Pricing Comparison by GPU Model
Here is what each GPU actually costs across provider types in early 2026:
| GPU Model | Hyperscaler On-Demand | Hyperscaler Reserved | Specialized Cloud | GPU Marketplace |
|---|---|---|---|---|
| H100 80GB | $3.00–$6.98/hr | $2.00–$4.50/hr | $2.00–$3.00/hr | $1.49–$2.50/hr |
| A100 80GB | $1.50–$3.50/hr | $1.00–$2.50/hr | $0.80–$1.50/hr | $0.72–$1.50/hr |
| L40S 48GB | $1.20–$2.50/hr | $0.80–$1.80/hr | $0.80–$1.50/hr | $0.60–$1.20/hr |
| L4 24GB | $0.50–$1.00/hr | $0.30–$0.70/hr | $0.30–$0.60/hr | $0.20–$0.50/hr |
| RTX 4090 24GB | N/A (not offered) | N/A | $0.40–$0.80/hr | $0.40–$0.80/hr |
Key observation: The price gap between hyperscaler on-demand and marketplace pricing is 2–4.7x for the same GPU hardware. For an H100, the difference between $6.98/hr (AWS on-demand) and $1.49/hr (marketplace spot) is $5.49/hr — which compounds to $4,008/month for a single GPU running 24/7.
Hidden Costs That Inflate Your GPU Bill
The hourly GPU rate is only part of the story. Five categories of hidden costs regularly surprise teams:
1. Data Egress Fees
Moving data out of a cloud provider’s network costs money — typically $0.05–$0.12 per GB. Training a model that downloads 1 TB of training data and exports checkpoints regularly can add $50–$120 in egress fees per training run. Multi-cloud workflows (training on one provider, inference on another) multiply these costs.
2. Storage Costs
GPU instances need disk space for datasets, model checkpoints, and logs. Cloud storage costs $0.02–$0.08/GB/month. A 5 TB dataset stored for 3 months costs $300–$1,200 — invisible in the GPU hourly rate but very visible on the bill.
3. Networking and Load Balancing
Production inference requires load balancers, API gateways, and inter-node networking. These services add $50–$200+/month depending on traffic volume. Multi-GPU training clusters incur additional costs for high-speed inter-node networking.
4. Minimum Commitments and Rounding
Some providers bill in hourly increments — a 35-minute job costs the same as a 60-minute job. Per-second billing (offered by some marketplaces) eliminates this waste. For bursty workloads with many short jobs, per-hour billing can waste 15–30% of your spend on unused time.
5. API and Management Service Fees
Managed ML services (SageMaker, Vertex AI) charge additional fees on top of GPU compute. These can add 10–30% to the base GPU cost for the convenience of managed training and deployment pipelines.
Key insight: A workload that costs $3.15/hr in GPU compute may cost $4.00–$4.50/hr in total when you add egress, storage, networking, and management fees. Always calculate the total cost of your workload, not just the GPU line item.
How to Cut Cloud GPU Costs by 40%
Five proven strategies that, combined, can reduce your total GPU spend by 30–40%:
1. Use Spot/Preemptible Instances (Save 15–20%)
Spot instances offer 40–60% discounts over on-demand pricing. The catch: they can be interrupted with short notice (typically 2 minutes to 2 hours). For fault-tolerant workloads — training with checkpointing, batch data processing, non-urgent inference — spot instances are the single largest cost lever.
Works best for: Model training with regular checkpoints, data preprocessing, batch inference, experimentation.
Does not work for: Production inference serving live users, time-critical jobs, workloads that cannot tolerate interruption.
2. Right-Size Your GPU (Save 8–12%)
Many teams default to the most powerful GPU “to be safe.” A workload that runs fine on an A100 ($0.72/hr) does not need an H100 ($3.15/hr). Our best GPU for AI guide provides a workload decision framework — matching GPU model to task can reduce costs by 50%+ for over-provisioned workloads.
Common mistakes: Using H100 for inference when L40S suffices. Using A100 for fine-tuning small models when RTX 4090 works. Running batch inference on on-demand instances when spot is available.
3. Schedule and Auto-Scale (Save 5–8%)
GPU instances running 24/7 when your team works 8 hours/day waste two-thirds of spend. Auto-scaling policies that spin up GPUs for training jobs and terminate them when idle can reclaim significant waste. Even simple scheduling — “GPUs on at 8 AM, off at 8 PM” — saves 50% on development instances.
4. Migrate Workloads to Marketplaces (Save 10–15%)
For workloads that do not require hyperscaler-specific managed services, moving to a GPU marketplace can save 3–6x on compute. The key requirement: your workload must be containerized (Docker) and not dependent on provider-specific APIs (SageMaker, Vertex AI).
5. Use Reserved/Committed Capacity (Save 5–10% additional)
For predictable, sustained workloads (production inference, ongoing training), reserved capacity commitments offer 30–50% discounts over on-demand pricing. The trade-off is commitment — typically 1–3 years. Combine reserved capacity for your baseline with spot/marketplace for variable load.
Provider-by-Provider Breakdown
AWS (Amazon Web Services)
GPU instances: P5 (H100), P4 (A100), G5 (A10G), G6 (L4). Pricing: H100 on-demand: ~$4.50–$6.98/hr. Reserved (1-year): ~$3.00–$4.50/hr. Spot: ~$1.50–$2.50/hr. Strengths: Broadest managed service ecosystem (SageMaker, EKS, S3), global availability, enterprise support. Watch out for: Egress fees ($0.09/GB), complex pricing tiers, instance availability fluctuations for on-demand.
Google Cloud Platform
GPU instances: A3 (H100), A2 (A100), G2 (L4). Pricing: H100 on-demand: ~$3.50–$5.50/hr. Committed (1-year): ~$2.50–$3.50/hr. Spot: ~$1.50–$2.00/hr. Strengths: Strong TPU alternative, excellent JAX/TensorFlow integration, Vertex AI managed platform, competitive spot pricing. Watch out for: Smaller GPU fleet than AWS, less availability in some regions.
Microsoft Azure
GPU instances: ND H100 (H100), NC A100 (A100), NC T4 (T4). Pricing: H100 on-demand: ~$3.00–$5.00/hr. Reserved (1-year): ~$2.00–$3.50/hr. Strengths: OpenAI integration (Azure OpenAI Service), strong enterprise identity/security, hybrid cloud capabilities. Watch out for: GPU availability can be constrained in popular regions.
GPU Marketplaces
Available GPUs: H100, A100, L40S, L4, RTX 4090, and more. Pricing: H100: $1.49–$2.50/hr. A100: $0.72–$1.50/hr. RTX 4090: $0.40–$0.80/hr. Strengths: Lowest prices, flexible billing (per-second on some platforms), no commitments, wide GPU selection including consumer cards. GPUnex, for example, offers GPUs starting at $0.39/hr with per-second billing and pre-installed AI frameworks. Watch out for: Variable reliability depending on provider, less managed tooling, may require more self-service configuration.
When to Use Which Provider Type
| Scenario | Best Provider Type | Why |
|---|---|---|
| Enterprise production inference | Hyperscaler | SLAs, global availability, managed services |
| Large-scale training (1000+ GPUs) | Hyperscaler or Specialized Cloud | Guaranteed capacity, NVLink networking |
| Cost-optimized training | GPU Marketplace (spot) | Lowest prices, checkpointing handles interruptions |
| Development and experimentation | GPU Marketplace | Per-second billing, no commitments, cheap iteration |
| Small team / startup | GPU Marketplace | Lowest barrier, flexible scaling, no minimums |
| Batch inference | GPU Marketplace or Spot | Tolerates interruption, price-sensitive |
| Regulated workloads (finance, health) | Hyperscaler | Compliance certifications, audit trails |
For the rent-vs-buy question — whether cloud rental makes more sense than owning hardware — see our detailed cost comparison. For startups budgeting their first GPU spend, see our AI startup compute costs guide.
Frequently Asked Questions
What is the cheapest way to rent an H100?
GPU marketplaces offer the lowest H100 rates at $1.49–$2.50/hr. Hyperscaler spot instances come next at $1.50–$2.50/hr but can be interrupted. The cheapest reliable option is a marketplace with a provider that has high uptime scores. For non-time-critical workloads, spot pricing on any platform offers the deepest discounts.
Are hidden costs really that significant?
Yes. For a typical training workload, hidden costs (egress, storage, networking) add 20–40% on top of the GPU compute bill. A $3.15/hr H100 instance running a training job with 2 TB of data transfer and 500 GB of checkpoint storage can effectively cost $4.00–$4.50/hr in total. Always estimate total workload cost, not just GPU per-hour rate.
Should I use one provider or multiple?
Multiple. Use GPU marketplaces for cost-sensitive workloads (development, batch processing, inference). Use hyperscalers for production inference and regulated workloads. Use spot pricing across all providers for fault-tolerant jobs. Diversifying prevents lock-in and lets you optimize cost at each tier.
What is per-second billing and does it matter?
Per-second billing charges you only for the exact time your GPU is active. Per-hour billing rounds up — a 35-minute job costs the same as 60 minutes. For workloads with many short jobs (inference, experimentation, CI/CD), per-second billing saves 15–30%. For long-running training jobs, the difference is negligible.
How have GPU prices changed over the past year?
H100 prices dropped 44% between June 2024 and June 2025, driven by expanding supply and marketplace competition. A100 prices fell even more as the GPU ages. Prices are expected to continue declining gradually as NVIDIA and AMD increase production, though new-generation GPUs (B200, Rubin) will command premium pricing at launch. For broader market context, see our OpenAI-NVIDIA market analysis.