7B/8B inference
Recommended
NVIDIA · Ada Lovelace
Best for image generation, computer vision, embeddings and mid-size LLM inference.
Monthly rental
~ USD 0.69/hour based on a 730-hour month
Availability: Available
Compare GPUs| Vendor | NVIDIA |
|---|---|
| Model | NVIDIA L40S 48GB |
| SKU | L40S |
| Architecture | Ada Lovelace |
| Hardware class | datacenter gpu |
| VRAM | 48 GB |
|---|---|
| Memory type | GDDR6 ECC |
| Memory bandwidth | 864 GB/s |
| Precision | Dense | With sparsity |
|---|---|---|
| FP32 | 91.6 TFLOPS | — |
| TF32 Tensor | 183 TFLOPS | 366 TFLOPS |
| BF16 Tensor | 362.05 TFLOPS | 733 TFLOPS |
| FP16 Tensor | 362.05 TFLOPS | 733 TFLOPS |
| FP8 Tensor | 733 TFLOPS | 1466 TFLOPS |
| INT8 Tensor | 733 TOPS | 1466 TOPS |
| Form factor / interface | PCIe dual-slot accelerator card |
|---|---|
| Interface | PCIe Gen4 x16 |
| TDP | 350 W |
| NVLink | Not supported |
| NVSwitch | Not supported |
| MIG | Not supported |
| Max GPUs per node | 8 |
| CUDA | Supported |
|---|---|
| Verification | OFFICIAL_VENDOR |
| Source | www.nvidia.com |
Recommended
Recommended with quantisation or FP16/BF16 headroom review
Recommended with quantisation for production headroom
Usually requires multiple GPUs or stronger quantisation
Suitable for LoRA/QLoRA; full fine-tuning depends on model size and topology
Not the primary choice for large distributed training
24 GB GDDR6 · Available
USD 250/month
80 GB HBM3 · Available
USD 1,860/month
80 GB HBM2e · Limited availability
USD 1,579/month
NVIDIA L40S 48GB is rated Very Good for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.
The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.
Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.