7B/8B inference
Recommended
NVIDIA · Hopper
Best for high-memory LLM inference, quantized 70B+ workloads and batch serving with generous VRAM headroom.
Monthly rental
~ USD 3.21/hour based on a 730-hour month
Availability: Limited availability
Compare GPUs| Vendor | NVIDIA |
|---|---|
| Model | NVIDIA H200 SXM 141GB |
| SKU | H200 SXM |
| Architecture | Hopper |
| Hardware class | datacenter gpu |
| VRAM | 141 GB |
|---|---|
| Memory type | HBM3e |
| Memory bandwidth | 4800 GB/s |
| Precision | Dense | With sparsity |
|---|---|---|
| FP64 | 34 TFLOPS | — |
| FP64 Tensor | 67 TFLOPS | — |
| FP32 | 67 TFLOPS | — |
| TF32 Tensor | 494.5 TFLOPS | 989 TFLOPS |
| BF16 Tensor | 989.5 TFLOPS | 1979 TFLOPS |
| FP16 Tensor | 989.5 TFLOPS | 1979 TFLOPS |
| FP8 Tensor | 1979 TFLOPS | 3958 TFLOPS |
| INT8 Tensor | 1979 TOPS | 3958 TOPS |
| Form factor / interface | SXM module |
|---|---|
| Interface | PCIe Gen5 x16 host interface |
| TDP | 700 W |
| NVLink | 900 GB/s NVLink |
| NVSwitch | Supported |
| MIG | Up to 7 MIGs @ 18 GB |
| Max GPUs per node | 8 |
| CUDA | Supported |
|---|---|
| Verification | OFFICIAL_VENDOR |
| Source | www.nvidia.com |
Recommended
Recommended with quantisation or FP16/BF16 headroom review
Recommended with quantisation for production headroom
Recommended for quantised inference; review context and concurrency
Suitable for LoRA/QLoRA; full fine-tuning depends on model size and topology
Use multi-GPU nodes with high-bandwidth topology
80 GB HBM3 · Available
USD 1,860/month
80 GB HBM2e · Limited availability
USD 1,579/month
80 GB HBM2e · Available
USD 915/month
NVIDIA H200 SXM 141GB is rated Excellent for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.
The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.
Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.