7B/8B inference
Recommended
NVIDIA · Hopper
Best for production LLM serving, 70B-class quantized inference and high-bandwidth multi-GPU training.
Monthly rental
~ USD 2.55/hour based on a 730-hour month
Availability: Available
Compare GPUs| Vendor | NVIDIA |
|---|---|
| Model | NVIDIA H100 SXM 80GB |
| SKU | H100 SXM |
| Architecture | Hopper |
| Hardware class | datacenter gpu |
| VRAM | 80 GB |
|---|---|
| Memory type | HBM3 |
| Memory bandwidth | 3350 GB/s |
| Precision | Dense | With sparsity |
|---|---|---|
| FP64 | 34 TFLOPS | — |
| FP64 Tensor | 67 TFLOPS | — |
| FP32 | 67 TFLOPS | — |
| TF32 Tensor | 494.5 TFLOPS | 989 TFLOPS |
| BF16 Tensor | 989.5 TFLOPS | 1979 TFLOPS |
| FP16 Tensor | 989.5 TFLOPS | 1979 TFLOPS |
| FP8 Tensor | 1979 TFLOPS | 3958 TFLOPS |
| INT8 Tensor | 1979 TOPS | 3958 TOPS |
| Form factor / interface | SXM module |
|---|---|
| Interface | PCIe Gen5 x16 host interface |
| TDP | 700 W |
| NVLink | 900 GB/s NVLink |
| NVSwitch | Supported |
| MIG | Up to 7 MIGs @ 10 GB |
| Max GPUs per node | 8 |
| CUDA | Supported |
|---|---|
| Verification | OFFICIAL_VENDOR |
| Source | www.nvidia.com |
Recommended
Recommended with quantisation or FP16/BF16 headroom review
Recommended with quantisation for production headroom
Recommended for quantised inference; review context and concurrency
Suitable for LoRA/QLoRA; full fine-tuning depends on model size and topology
Use multi-GPU nodes with high-bandwidth topology
80 GB HBM2e · Limited availability
USD 1,579/month
80 GB HBM2e · Available
USD 915/month
48 GB GDDR6 ECC · Available
USD 506/month
NVIDIA H100 SXM 80GB is rated Very Good for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.
The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.
Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.