H200 vs MI300X: what you're actually paying for
I spent a chunk of last month on a question that sounds simple. Fixed budget, a 70B model to serve. Which GPU do I rent?
Every comparison I found opened with a TFLOPS table, which is close to useless for what I was doing. Token generation isn't compute-bound. It's memory-bound. Your GPU spends most of the decode phase waiting on memory rather than doing math, so a chart of peak FLOPS is describing a bottleneck you don't have.
Here's what I found once I started looking at the parts that matter.
The H200 is an H100 with better memory
That's not a criticism, it's the design. Same GH100 die, same Tensor Cores, same peak FLOPS. NVIDIA swapped the memory subsystem: six stacks of 24 GB HBM3e instead of five of 16 GB HBM3. You get 141 GB at 4.8 TB/s where the H100 gave you 80 GB at 3.35 TB/s.
The consequence is more interesting than the spec. H100 and H200 perform almost identically on prefill, since chewing through your prompt is compute-bound and the compute didn't change. The gap only opens during decode. Serving chat means you live in decode, so you feel the whole 43% bandwidth increase. Doing bulk document processing with short outputs, you might barely notice.
The MI300X got to its numbers a different way. AMD went chiplet, eight compute dies and four I/O dies stitched together with Infinity Fabric, and that packaging is what allowed 192 GB of HBM3 at 5.3 TB/s. More capacity and more bandwidth than the H200.
| Spec | H200 SXM | MI300X |
|---|---|---|
| Memory | 141 GB HBM3e | 192 GB HBM3 |
| Bandwidth | 4.8 TB/s | 5.3 TB/s |
| Architecture | Hopper, monolithic | CDNA 3, chiplet |
| Power | 700 W | 750 W |
| GPU to GPU | NVLink, 900 GB/s | Infinity Fabric, about 896 GB/s |
| Software | CUDA | ROCm |
On memory, AMD wins. That part isn't really debatable. Whether it matters is a different question.
What the extra 51 GB is actually for
The useful question isn't which number is bigger. It's how many GPUs you need to hold your model.
Weights run about 2 bytes per parameter at FP16, roughly 1 at FP8. Then there's KV cache, which grows with batch size and context length, and which in my experience is what actually breaks things. People size for weights, forget cache, then wonder why throughput falls apart at batch 32.
A 70B model at FP16 is around 140 GB of weights. That technically fits in an H200's 141 GB and leaves you nothing for cache, so realistically you're running two. The same model on an MI300X leaves roughly 50 GB of headroom. One GPU.
That's the real pitch, more than the bandwidth number. When a model fits on one device you delete a whole category of problems: no tensor parallelism, no per-layer synchronization, one thing to fail instead of eight. I'd take that simplification over a few percent of bandwidth most days.
Try other sizes in the will-it-fit calculator on the home page.
PCIe or NVLink, and why it decides everything above one GPU
PCIe Gen5 x16 moves about 128 GB/s bidirectional. NVLink on an H200 SXM moves 900 GB/s per GPU. Roughly seven times the bandwidth, and it's the entire reason SXM nodes exist as a product.
Why that gap matters so much: tensor parallelism splits each layer across GPUs, so every layer needs an all-reduce, for every token. Eighty layers and a few hundred output tokens means tens of thousands of collectives per request. Over PCIe that collective becomes the bottleneck, and you end up with expensive silicon sitting idle waiting on its neighbours.
The heuristic I've settled on: PCIe is fine when each GPU runs its own copy of the model. You want NVLink when GPUs have to cooperate on one model. Four independent replicas of a 30B? PCIe cards, save the money. One 405B split eight ways? The interconnect is the thing you're buying, not the GPU.
There's also an H200 NVL variant, a PCIe card that bridges up to four GPUs over NVLink. It exists for standard air-cooled racks that can't host an 8-GPU SXM baseboard. Useful, but slower than SXM once you're doing serious tensor-parallel work.
One correction while I'm here, because it caught me out. Several sites list the H200 as having NVLink 5.0 at 1.8 TB/s. It doesn't. H200 is Hopper, so it's NVLink 4.0 at 900 GB/s. The 1.8 TB/s figure belongs to Blackwell. I'd built half a spreadsheet on the wrong number before I noticed, so if you're comparing tables, check that row.
Going from 1 to 8
Nobody gets linear scaling. Here's what each step buys:
- One GPU is simplest and often fastest per token, assuming it fits. Zero collectives.
- Two is the cheapest way to double memory, with a direct peer link and modest overhead.
- Four tends to be the sweet spot for serving 70B-class models at volume.
- Eight is the standard node, an HGX H200 or an eight-way MI300X box. On the NVIDIA side you get a full NVSwitch fabric, so any GPU talks to any other at full speed.
This is where the picture flips. In published 8-GPU training comparisons, NVLink nodes land around 87% scaling efficiency against roughly 78% for MI300X, and the gap grows with node count. Per-chip memory doesn't help you when the constraint is how well eight chips pretend to be one.
So AMD wins when you're trying to use fewer GPUs, and NVIDIA wins when you need many working together. Those are different problems, and people conflate them constantly.
What it costs, and who the neoclouds are
Neocloud is the label for GPU specialists like CoreWeave, Lambda, Nebius, Crusoe, RunPod, and GMI, as opposed to the big three. They undercut hyperscalers because selling compute is the whole business rather than a line item carrying a lot of platform overhead.
The spread is genuinely wide. Published H200 rates this year have run from around $2.60 an hour at the cheap end to $6.31 at CoreWeave. That's about $2,600 a month on a single GPU, or somewhere near $21,000 a month across eight, for identical silicon. MI300X has come down further, as low as roughly $1.71 an hour at Crusoe, with spot going under a dollar.
Run the memory-per-dollar math and it's stark. An H200 at $2.60 is about $0.018 per GB-hour. An MI300X at $1.71 is about $0.009. Half the price for the same gigabyte.
Before acting on that, though, the hourly rate isn't the bill:
- Minimum node size. CoreWeave won't sell you fewer than eight GPUs. You pay for the node whether you use it or not.
- Egress. Moving weights and datasets out can quietly eat the savings. Some providers charge nothing, some very much do.
- Preemption. Cheap often means spot. Fine for training with checkpoints, bad for anything user-facing.
- Utilization. A GPU you rent and don't saturate is the most expensive one you'll ever buy.
That last one is the one I keep relearning.
How I'd choose
I'd reach for the MI300X when the model fits in 192 GB but not 141 GB, when I'm serving long contexts where KV cache dominates, when the stack is already vLLM or SGLang on ROCm, and when cost per gigabyte matters more than shaving the last few milliseconds.
I'd reach for the H200 when I'm doing tensor parallelism at four or eight GPUs, when I depend on CUDA-specific kernels, when latency consistency is a product requirement rather than a nice-to-have, or when I'm training rather than serving.
Reduced to one line: AMD sells memory, NVIDIA sells interconnect and a software stack that already works. Which is the better buy depends on whether your problem is fitting the model or making a pile of GPUs behave like one.
I'd rather measure this than argue about it, so the next post is a real batch-size sweep on a single GPU with the configs published. If a number here is wrong or has gone stale, tell me and I'll fix it.
Figures checked August 2026. GPU pricing moves fast, so check current rates before deciding anything on them.