Skip to main content

Meta Llama · General · Llama 3.1

Llama 3.1 405B: Specs, GPU requirements, hosting & deployment

Llama 3.1 405B is a recent General release in the Meta Llama family, certified from its immutable official source revision.

Technical deployment reference illustration for Llama 3.1 405B
Reference deployment flow; exact architecture facts are listed separately.
Total parameters405.0B
Active parametersDense model
Context window131,072 tokens
Release statusRecent · Recent
ArchitectureDense Transformer
ModalitiesText
LicenseLlama 3.1 Community License
Deployment stateDeployable with certified configuration
Commercial stateCommercial Use Allowed With Conditions
Release date2024-07-16

Architecture, strengths and use cases

Llama 3.1 405B is an official self-hostable Meta release in the Meta Llama General track. It uses Dense Transformer, carries 405B, supports a verified context window of 131,072 tokens, and exposes only source-certified checkpoints and runtime combinations. Its release, license and deployment facts are tied to the official repository revision rather than inferred from the family name.

Llama 3.1 405B uses Dense Transformer with 405B. The stored architecture profile records 64 layers, hidden width 8192, 64 attention heads and a 131,072 token context ceiling where the official configuration publishes those fields.

Verified Dense Transformer architecture
131,072 token certified context
Official BF16 checkpoint coverage
Private LLM inference
Agentic workflows
Long-context document processing

Certified checkpoints and quantizations

Only exact, source-certified variants attached to this release are shown.

Llama-3.1-405B-Instruct

meta-llama/Llama-3.1-405B-Instruct

Official · BF16 · safetensors

Runtime certified: Transformers

Precision / quantizationFormatBitsPublisher trustDeployment state
BF16safetensors16OfficialRuntime certified

Recommended hardware

Use only the listed transformers runtime configuration with the exact checkpoint precision. The recommendation engine checks weight memory, KV cache, runtime overhead, safety headroom, GPU architecture and topology before ranking rentable and reference hardware.

Estimated weight memory777 GB
KV cache2 GB
Runtime overhead196.86 GB
Total with safety margin1210.07 GB
Technically compatible reference sizing

No current FusionPointAI inventory passes every runtime, architecture, topology and memory gate for this exact checkpoint. Use the estimator for global compatible hardware.

License and self-hosting considerations

The full 405B resident weight set must fit across the selected GPUs. Context and concurrency add KV-cache memory; MoE active parameters do not replace resident-weight memory.

Llama 3.1 405B is published under Llama 3.1 Community License. The release-scoped official license link and verification date are retained with the catalogue record.

Commercial Use Allowed With Conditions. Commercial use is conditional. Review the linked official terms before deployment.

Alternatives and sibling variants

Llama 4 Scout 17B-16E, Llama 4 Maverick 17B-128E

Release timeline

meta/llama/llama-3-1-405b @ ec5a654d6f9ae0f8c3e0ad5cf9d102d9937c302c51a497a7ae316cdb15dfed83

Official source

Frequently asked questions

How much VRAM does Llama 3.1 405B require?

The exact estimate depends on checkpoint precision, context and concurrency; the page recommendation uses the certified memory engine and compatible runtime configurations.

Which runtimes support Llama 3.1 405B?

Certified upstream runtime evidence is available for transformers.

Can Llama 3.1 405B be used commercially?

The stored commercial classification is Commercial Use Allowed With Conditions. Consult the official Llama 3.1 Community License terms for the final obligations.

Official references

Last verified: 2026-08-26