Skip to main content

Meta Llama · General · Llama 4

Llama 4 Maverick 17B-128E: Specs, GPU requirements, hosting & deployment

Llama 4 Maverick 17B-128E is a current General release in the Meta Llama family, certified from its immutable official source revision.

Technical deployment reference illustration for Llama 4 Maverick 17B-128E
Reference deployment flow; exact architecture facts are listed separately.
Total parameters400.0B
Active parameters17.0B
Context window1,000,000 tokens
Release statusCurrent · Current
ArchitectureMoE Transformer
ModalitiesText
LicenseLlama 4 Community License
Deployment stateDeployable with certified configuration
Commercial stateCommercial Use Allowed With Conditions
Release date2025-04-01

Architecture, strengths and use cases

Llama 4 Maverick 17B-128E is an official self-hostable Meta release in the Meta Llama General track. It uses MoE Transformer, carries 400B / 17B MoE, supports a verified context window of 1,000,000 tokens, and exposes only source-certified checkpoints and runtime combinations. Its release, license and deployment facts are tied to the official repository revision rather than inferred from the family name.

Llama 4 Maverick 17B-128E uses MoE Transformer with 400B / 17B MoE. The stored architecture profile records 64 layers, hidden width 8192, 64 attention heads and a 1,000,000 token context ceiling where the official configuration publishes those fields.

Verified MoE Transformer architecture
1,000,000 token certified context
Official FP8 checkpoint coverage
Private LLM inference
Agentic workflows
Long-context document processing

Certified checkpoints and quantizations

Only exact, source-certified variants attached to this release are shown.

Llama-4-Maverick-17B-128E-Instruct-FP8

meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8

Official · FP8 · safetensors

Runtime certified: Transformers

Precision / quantizationFormatBitsPublisher trustDeployment state
FP8safetensors8OfficialRuntime certified

Recommended hardware

Use only the listed transformers runtime configuration with the exact checkpoint precision. The recommendation engine checks weight memory, KV cache, runtime overhead, safety headroom, GPU architecture and topology before ranking rentable and reference hardware.

Estimated weight memory383.7 GB
KV cache1 GB
Runtime overhead125.89 GB
Total with safety margin633.13 GB
Technically compatible reference sizing

No current FusionPointAI inventory passes every runtime, architecture, topology and memory gate for this exact checkpoint. Use the estimator for global compatible hardware.

License and self-hosting considerations

The full 400B resident weight set must fit across the selected GPUs. Context and concurrency add KV-cache memory; MoE active parameters do not replace resident-weight memory.

Llama 4 Maverick 17B-128E is published under Llama 4 Community License. The release-scoped official license link and verification date are retained with the catalogue record.

Commercial Use Allowed With Conditions. Commercial use is conditional. Review the linked official terms before deployment.

Alternatives and sibling variants

Llama 4 Scout 17B-16E, Llama 3.1 405B

Curated comparisons

Release timeline

meta/llama/llama-4-maverick-17b-128e @ 0b389dbe619938c2b9ee945f693f2816c90ed1c6f2754c113b652873e163735f

Official source

Frequently asked questions

How much VRAM does Llama 4 Maverick 17B-128E require?

The exact estimate depends on checkpoint precision, context and concurrency; the page recommendation uses the certified memory engine and compatible runtime configurations.

Which runtimes support Llama 4 Maverick 17B-128E?

Certified upstream runtime evidence is available for transformers.

Can Llama 4 Maverick 17B-128E be used commercially?

The stored commercial classification is Commercial Use Allowed With Conditions. Consult the official Llama 4 Community License terms for the final obligations.

Official references

Last verified: 2026-08-26