7B/8B inference
Recommended
NVIDIA · Ampere
Best for proven CUDA workloads, LoRA fine-tuning, research and stable production inference.
Monthly rental
~ USD 1.25/hour based on a 730-hour month
Availability: Available
Compare GPUs| Vendor | NVIDIA |
|---|---|
| Model | NVIDIA A100 SXM 80GB |
| SKU | A100 SXM 80GB |
| Architecture | Ampere |
| Hardware class | datacenter gpu |
| VRAM | 80 GB |
|---|---|
| Memory type | HBM2e |
| Memory bandwidth | 2039 GB/s |
| Precision | Dense | With sparsity |
|---|---|---|
| FP64 | 9.7 TFLOPS | — |
| FP64 Tensor | 19.5 TFLOPS | — |
| FP32 | 19.5 TFLOPS | — |
| TF32 Tensor | 156 TFLOPS | 312 TFLOPS |
| BF16 Tensor | 312 TFLOPS | 624 TFLOPS |
| FP16 Tensor | 312 TFLOPS | 624 TFLOPS |
| INT8 Tensor | 624 TOPS | 1248 TOPS |
| Form factor / interface | SXM module |
|---|---|
| Interface | PCIe Gen4 x16 host interface |
| TDP | 400 W |
| NVLink | 600 GB/s NVLink |
| NVSwitch | Supported |
| MIG | Up to 7 MIGs @ 10 GB |
| Max GPUs per node | 8 |
| CUDA | Supported |
|---|---|
| Verification | OFFICIAL_VENDOR |
| Source | www.nvidia.com |
Recommended
Recommended with quantisation or FP16/BF16 headroom review
Recommended with quantisation for production headroom
Recommended for quantised inference; review context and concurrency
Suitable for LoRA/QLoRA; full fine-tuning depends on model size and topology
Use multi-GPU nodes with high-bandwidth topology
80 GB HBM3 · Available
USD 1,860/month
80 GB HBM2e · Limited availability
USD 1,579/month
48 GB GDDR6 ECC · Available
USD 506/month
NVIDIA A100 SXM 80GB is rated Very Good for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.
The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.
Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.