GPUs
H100 vs H200
What actually changed between NVIDIA's two Hopper-generation data-center accelerators.
Updated
The NVIDIA H100 and H200 are two versions of the same Hopper generation. The headline change is memory: the H200 keeps Hopper compute but moves from 80 GB of HBM3 to 141 GB of faster HBM3e. That makes this a workload question, not a simple “newer is better” question. This guide compares the common SXM configurations (H100 SXM 80GB and H200 SXM 141GB), explains where the extra memory matters, and links to live pricing so you can compare cost for your own job.
H100 vs H200 specifications
| Spec | H100 SXM 80GB | H200 SXM 141GB |
|---|---|---|
| Architecture | NVIDIA Hopper | NVIDIA Hopper (same compute architecture) |
| Memory capacity | 80 GB | 141 GB |
| Memory type | HBM3 | HBM3e |
| Memory bandwidth | Up to 3.35 TB/s | Up to 4.8 TB/s |
| Max thermal design power | Up to 700 W (configurable) | Up to 700 W (configurable) |
| Interconnect | NVLink 900 GB/s, PCIe Gen5 128 GB/s | NVLink 900 GB/s, PCIe Gen5 128 GB/s |
| Generation context | Announced 2022; first Hopper data-center GPU | Announced 2023; memory upgrade of Hopper |
Source: NVIDIA H100 and H200 product specifications for SXM form factors. PCIe and NVL versions have different memory, bandwidth and power figures. Full entries: NVIDIA H100 specs and NVIDIA H200 specs.
What the difference means in practice
LLM inference, KV cache and longer context
During inference a GPU must hold the model weights plus a key-value (KV) cache that grows with context length and the number of concurrent requests. Generating each token also means reading weights from memory, so many inference workloads are limited by memory bandwidth rather than raw compute. The H200’s 141 GB leaves more room for KV cache after weights are loaded, which can allow longer contexts or more parallel sequences on one GPU, and its higher bandwidth can help token generation. A model that only just fits on an H100, or needs two H100s, may fit on a single H200. Whether that shows up as higher throughput depends on the serving engine, precision, batching strategy and model.
Training: when memory is or is not the bottleneck
For training jobs where the model, optimizer states and activations already fit comfortably in 80 GB, the H100 and H200 share the same compute architecture, so the gain per GPU may be small. Where memory is the constraint, the H200 can reduce the need for activation checkpointing, aggressive sharding or very small micro-batches, which can simplify the setup and improve how well the GPU is used. Measure on your own training code before assuming either outcome.
Batch size and concurrency
More memory per GPU usually means room for larger batches or more simultaneous users before running out of memory. For serving, that can lower cost per token even if the hourly price is higher, because each GPU handles more work. It is not automatic: latency targets, scheduler behaviour and the model’s attention pattern all affect how much extra concurrency is usable.
Multi-GPU deployment
Both GPUs use fourth-generation NVLink at 900 GB/s and are commonly deployed in eight-GPU HGX servers. Memory is not automatically pooled across GPUs; large models are split with tensor, pipeline or expert parallelism, which adds communication. Fitting a model on fewer H200s can reduce that communication, while very large models still need many GPUs of either type. Because the platforms are similar, existing H100 cluster designs and software generally carry over.
Choose H100 when…
- Your model, context length and batch size fit comfortably in 80 GB per GPU.
- The workload is mainly compute-bound, such as many training jobs on mid-sized models.
- Lower hourly rates or wider availability on your preferred provider matter more than memory headroom.
- You already run H100 clusters and want consistent hardware across jobs.
Choose H200 when…
- Your model plus KV cache exceeds, or sits close to, 80 GB per GPU.
- You serve long contexts or many concurrent users and memory limits your batch size.
- Fitting the model on fewer GPUs would cut parallelism overhead or total GPU count.
- Your own tests show memory bandwidth limits token generation speed.
Not sure which applies? Estimate memory needs with the LLM VRAM calculator, then test a short run on each GPU.
H100 and H200 cloud pricing
Prices below are read from the site’s GPU pricing database at page load, showing only rows with a recorded hourly rate. Rates change often and depend on region, commitment and configuration, so confirm with the provider. Compare more providers on GPU cloud prices and the GPU cloud provider directory.
| GPU | Provider | Per GPU-hour | Updated |
|---|---|---|---|
| NVIDIA H100 | DigitalOcean | 4.41 USD | 29 September 2026 |
| NVIDIA H100 PCIe | Lambda | 3.29 USD | 7 October 2026 |
| NVIDIA H100 PCIe | Runpod | 2.89 USD | 29 September 2026 |
| NVIDIA H100 PCIe | Scaleway | 2.73 EUR | 29 September 2026 |
| NVIDIA H100 SXM | Lambda | 4.29 USD | 7 October 2026 |
| NVIDIA H100 SXM | Nebius | 4.50 USD | 7 October 2026 |
| NVIDIA H100 SXM | Runpod | 3.49 USD | 29 September 2026 |
| NVIDIA H100 SXM | Scaleway | 6.61 EUR | 29 September 2026 |
| NVIDIA H100 SXM | Verda | 3.25 USD | 29 September 2026 |
| NVIDIA H200 | DigitalOcean | 4.47 USD | 29 September 2026 |
| NVIDIA H200 | Nebius | 5.40 USD | 7 October 2026 |
| NVIDIA H200 | Runpod | 4.59 USD | 29 September 2026 |
| NVIDIA H200 | Verda | 4.00 USD | 29 September 2026 |
A higher hourly rate can still mean a lower cost per job if the H200 needs fewer GPUs or serves more requests. Model this with the GPU cloud cost calculator.
Frequently asked questions
Is the H200 faster than the H100?
Both use the same Hopper compute architecture, so peak compute per GPU is broadly similar. The H200 has more memory (141 GB vs 80 GB) and higher memory bandwidth (about 4.8 TB/s vs 3.35 TB/s). Workloads limited by memory capacity or bandwidth, such as large-model inference, can benefit; compute-bound work may see little difference. Real results depend on the model, software and system.
Can software written for the H100 run on the H200?
Generally yes. Both are Hopper GPUs and use the same CUDA software stack, so code and frameworks that target H100 normally run on H200 without changes. Check your framework and driver versions with your provider.
Does the H200 use more power than the H100?
NVIDIA lists a maximum thermal design power of up to 700 W for both the H100 SXM and H200 SXM. Actual draw depends on the workload and how the system is configured.
How do I know if my model fits on one H100 or H200?
Estimate the memory for model weights, runtime overhead and KV cache with the LLM VRAM calculator, then compare it with 80 GB or 141 GB. Treat the result as a planning estimate and test on the real hardware.
Is the H200 more expensive to rent?
On the same provider the H200 is usually listed above the H100, but rates change often and vary by region, commitment and provider. Check the current rows on the GPU prices page rather than relying on a fixed number.
GPU Data Hub Daily
Stay Ahead of the AI Infrastructure Economy
The most important GPU, AI, data-center, semiconductor and cloud developments delivered directly to your inbox.
Cite this page
“H100 vs H200.” GPU Data Hub. https://gpudatahub.com/guides/h100-vs-h200
You are welcome to reference and link to this page. Please link to the URL above.