NVIDIA L4 24GB GPU accelerator

NVIDIA · Ada Lovelace

NVIDIA L4 24GB

Best for small-model inference, embeddings, reranking and lightweight production workloads.

Monthly rental

USD 250/month

~ USD 0.34/hour based on a 730-hour month

Availability: Available

Compare GPUs

Technical specifications

Identity

Vendor NVIDIA
Model NVIDIA L4 24GB
SKU L4
Architecture Ada Lovelace
Hardware class datacenter gpu

Memory

VRAM 24 GB
Memory type GDDR6
Memory bandwidth 300 GB/s

Compute

Precision Dense With sparsity
FP32 30.3 TFLOPS
TF32 Tensor 60 TFLOPS 120 TFLOPS
BF16 Tensor 121 TFLOPS 242 TFLOPS
FP16 Tensor 121 TFLOPS 242 TFLOPS
FP8 Tensor 242.5 TFLOPS 485 TFLOPS
INT8 Tensor 242.5 TOPS 485 TOPS

Platform

Form factor / interface PCIe low-profile accelerator card
Interface PCIe Gen4 x16
TDP 72 W
NVLink Not supported
NVSwitch Not supported
MIG Not supported
Max GPUs per node 8

Software

CUDA Supported
Verification OFFICIAL_VENDOR
Source www.nvidia.com

Workload fit

  • LLM inferenceSuitable
  • LLM fine-tuningLimited
  • Full model trainingLimited
  • Image / video generationSuitable
  • Embeddings / rerankingExcellent

Frameworks

  • PyTorch
  • TensorFlow
  • JAX
  • Hugging Face
  • vLLM
  • TensorRT
  • ONNX Runtime
  • CUDA

Model compatibility guidance

7B/8B inference

Recommended

13B/14B inference

Recommended with quantisation or FP16/BF16 headroom review

30B/32B inference

Limited

65B/70B inference

Usually requires multiple GPUs or stronger quantisation

Fine-tuning

Limited to small/adapted workloads

Large training

Not the primary choice for large distributed training

Related GPUs

AI GPU hosting FAQ

Is NVIDIA L4 24GB suitable for LLM inference?

NVIDIA L4 24GB is rated Suitable for inference in the FusionPointAI catalogue. Final sizing still depends on model size, precision, context length and concurrency.

How is monthly GPU pricing calculated?

The public monthly price is the authoritative price version. The displayed hourly equivalent is derived from that monthly price using a 730-hour month.

Can I rent more than one GPU?

Multi-GPU configurations are supported when inventory and node topology allow it. This SKU currently lists up to 8 GPUs per node.