NVIDIA L40S 48GB GPU accelerator

NVIDIA · Ada Lovelace

NVIDIA L40S 48GB

Best for image generation, computer vision, embeddings and mid-size LLM inference.

Monthly rental

USD 506/month

~ USD 0.69/hour based on a 730-hour month

Availability: Available

Compare GPUs

Technical specifications

Identity

Vendor NVIDIA
Model NVIDIA L40S 48GB
SKU L40S
Architecture Ada Lovelace
Hardware class datacenter gpu

Memory

VRAM 48 GB
Memory type GDDR6 ECC
Memory bandwidth 864 GB/s

Compute

Precision Dense With sparsity
FP32 91.6 TFLOPS
TF32 Tensor 183 TFLOPS 366 TFLOPS
BF16 Tensor 362.05 TFLOPS 733 TFLOPS
FP16 Tensor 362.05 TFLOPS 733 TFLOPS
FP8 Tensor 733 TFLOPS 1466 TFLOPS
INT8 Tensor 733 TOPS 1466 TOPS

Platform

Form factor / interface PCIe dual-slot accelerator card
Interface PCIe Gen4 x16
TDP 350 W
NVLink Not supported
NVSwitch Not supported
MIG Not supported
Max GPUs per node 8

Software

CUDA Supported
Verification OFFICIAL_VENDOR
Source www.nvidia.com

Workload fit

  • LLM inferenceVery Good
  • LLM fine-tuningSuitable
  • Full model trainingSuitable
  • Image / video generationVery Good
  • Embeddings / rerankingExcellent

Frameworks

  • PyTorch
  • TensorFlow
  • JAX
  • Hugging Face
  • vLLM
  • TensorRT
  • ONNX Runtime
  • CUDA

Model compatibility guidance

7B/8B inference

Recommended

13B/14B inference

Recommended with quantisation or FP16/BF16 headroom review

30B/32B inference

Recommended with quantisation for production headroom

65B/70B inference

Usually requires multiple GPUs or stronger quantisation

Fine-tuning

Suitable for LoRA/QLoRA; full fine-tuning depends on model size and topology

Large training

Not the primary choice for large distributed training

Related GPUs

AI GPU hosting FAQ

Is NVIDIA L40S 48GB suitable for LLM inference?

NVIDIA L40S 48GB is rated Very Good for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.

How is monthly GPU pricing calculated?

The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.

Can I rent more than one GPU?

Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.