NVIDIA H100 PCIe 80GB GPU accelerator

NVIDIA · Hopper

NVIDIA H100 PCIe 80GB

Best for powerful single-node inference, smaller fine-tuning jobs and cost-efficient Hopper access.

Monthly rental

USD 1,579/month

~ USD 2.16/hour based on a 730-hour month

Availability: Limited availability

Compare GPUs

Technical specifications

Identity

Vendor NVIDIA
Model NVIDIA H100 PCIe 80GB
SKU H100 PCIe
Architecture Hopper
Hardware class datacenter gpu

Memory

VRAM 80 GB
Memory type HBM2e
Memory bandwidth 2000 GB/s

Compute

Precision Dense With sparsity
FP64 26 TFLOPS
FP64 Tensor 51 TFLOPS
FP32 51 TFLOPS
TF32 Tensor 378 TFLOPS 756 TFLOPS
BF16 Tensor 756.5 TFLOPS 1513 TFLOPS
FP16 Tensor 756.5 TFLOPS 1513 TFLOPS
FP8 Tensor 1513 TFLOPS 3026 TFLOPS
INT8 Tensor 1513 TOPS 3026 TOPS

Platform

Form factor / interface PCIe dual-slot accelerator card
Interface PCIe Gen5 x16
TDP 350 W
NVLink 600 GB/s NVLink Bridge
NVSwitch Not supported
MIG Up to 7 MIGs @ 10 GB
Max GPUs per node 8

Software

CUDA Supported
Verification OFFICIAL_VENDOR
Source www.nvidia.com

Workload fit

  • LLM inferenceVery Good
  • LLM fine-tuningVery Good
  • Full model trainingVery Good
  • Image / video generationVery Good
  • Embeddings / rerankingExcellent

Frameworks

  • PyTorch
  • TensorFlow
  • JAX
  • Hugging Face
  • vLLM
  • TensorRT
  • ONNX Runtime
  • CUDA

Model compatibility guidance

7B/8B inference

Recommended

13B/14B inference

Recommended with quantisation or FP16/BF16 headroom review

30B/32B inference

Recommended with quantisation for production headroom

65B/70B inference

Recommended for quantised inference; review context and concurrency

Fine-tuning

Suitable for LoRA/QLoRA; full fine-tuning depends on model size and topology

Large training

Use multi-GPU nodes with high-bandwidth topology

Related GPUs

AI GPU hosting FAQ

Is NVIDIA H100 PCIe 80GB suitable for LLM inference?

NVIDIA H100 PCIe 80GB is rated Very Good for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.

How is monthly GPU pricing calculated?

The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.

Can I rent more than one GPU?

Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.