7B/8B inference
Recommended
NVIDIA · Ada Lovelace
Best for small-model inference, embeddings, reranking and lightweight production workloads.
Monthly rental
~ USD 0.34/hour based on a 730-hour month
Availability: Available
Compare GPUs| Vendor | NVIDIA |
|---|---|
| Model | NVIDIA L4 24GB |
| SKU | L4 |
| Architecture | Ada Lovelace |
| Hardware class | datacenter gpu |
| VRAM | 24 GB |
|---|---|
| Memory type | GDDR6 |
| Memory bandwidth | 300 GB/s |
| Precision | Dense | With sparsity |
|---|---|---|
| FP32 | 30.3 TFLOPS | — |
| TF32 Tensor | 60 TFLOPS | 120 TFLOPS |
| BF16 Tensor | 121 TFLOPS | 242 TFLOPS |
| FP16 Tensor | 121 TFLOPS | 242 TFLOPS |
| FP8 Tensor | 242.5 TFLOPS | 485 TFLOPS |
| INT8 Tensor | 242.5 TOPS | 485 TOPS |
| Form factor / interface | PCIe low-profile accelerator card |
|---|---|
| Interface | PCIe Gen4 x16 |
| TDP | 72 W |
| NVLink | Not supported |
| NVSwitch | Not supported |
| MIG | Not supported |
| Max GPUs per node | 8 |
| CUDA | Supported |
|---|---|
| Verification | OFFICIAL_VENDOR |
| Source | www.nvidia.com |
Recommended
Recommended with quantisation or FP16/BF16 headroom review
Limited
Usually requires multiple GPUs or stronger quantisation
Limited to small/adapted workloads
Not the primary choice for large distributed training
48 GB GDDR6 ECC · Available
USD 506/month
80 GB HBM3 · Available
USD 1,860/month
80 GB HBM2e · Limited availability
USD 1,579/month
NVIDIA L4 24GB is rated Suitable for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.
The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.
Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.