NVIDIA · Platform · Superchip
NVIDIA GH200 Grace Hopper Superchip 96GB
Exact GH200 96GB reference specifications and AI deployment guidance, separated from FusionPointAI commercial inventory.
Verified technical specifications
| Manufacturer | NVIDIA |
|---|---|
| Exact SKU | GH200 96GB |
| Product family | NVIDIA GH200 |
| Architecture | Hopper |
| Generation | Grace Hopper |
| Entity type | Superchip |
| Lifecycle | Recent generation |
| Launch date | The manufacturer has not published this information. |
| VRAM | 96 GB |
| Memory technology | HBM3 |
| Memory bandwidth | 4,000 GB/s |
| Power | 1000 W |
| Interface | NVLink-C2C |
| Form factor | CPU-GPU · Superchip |
| Interconnect | NVLink-C2C |
| MIG / partitioning | Supported |
| AI precision support | FP64, FP32, TF32, BF16, FP16, FP8, INT8 |
| Software stack | cuda, vllm, sglang, tensorrt-llm |
AI workload suitability
LLM inference
High memory capacity supports large checkpoints and multi-GPU inference when runtime and topology are certified.
Fine-tuning and training
Suitable for limited fine-tuning and development when software support, cooling and memory capacity are verified.
Image, video and embeddings
FP16 acceleration supports common generative and vector workloads within the listed VRAM envelope.
Power and cooling
1000 W published power requires chassis, airflow and power delivery for this exact form factor.
Multi-GPU use
NVLink-C2C is the recorded interconnect path; topology must still be validated at node level.
Software ecosystem
cuda, vllm, sglang, tensorrt-llm
Product timeline
Media gallery
Official manufacturer source
https://www.nvidia.com/en-us/data-center/grace-hopper-superchip/
Last verified: 2026-08-24