Skip to main content

Phi · General · Phi-4

Phi-4 Mini: Specs, GPU requirements, hosting & deployment

Phi-4 Mini is a current General release in the Phi family, certified from its immutable official source revision.

Technical deployment reference illustration for Phi-4 Mini
Reference deployment flow; exact architecture facts are listed separately.
Total parameters3.8B
Active parametersDense model
Context window131,072 tokens
Release statusCurrent · Current
ArchitectureDense Transformer
ModalitiesText
Licensemit
Deployment stateDeployable with certified configuration
Commercial stateCommercial Use Allowed
Release date2025-02-19

Architecture, strengths and use cases

Phi-4 Mini is an official self-hostable Microsoft release in the Phi General track. It uses Dense Transformer, carries 3.8B, supports a verified context window of 131,072 tokens, and exposes only source-certified checkpoints and runtime combinations. Its release, license and deployment facts are tied to the official repository revision rather than inferred from the family name.

Phi-4 Mini uses Dense Transformer with 3.8B. The stored architecture profile records 64 layers, hidden width 8192, 64 attention heads and a 131,072 token context ceiling where the official configuration publishes those fields.

Verified Dense Transformer architecture
131,072 token certified context
Official BF16, Unknown checkpoint coverage
Private LLM inference
Agentic workflows
Long-context document processing

Certified checkpoints and quantizations

Only exact, source-certified variants attached to this release are shown.

Phi-4-mini-instruct

microsoft/Phi-4-mini-instruct

Official · BF16 · safetensors

Runtime certified: Transformers

Precision / quantizationFormatBitsPublisher trustDeployment state
BF16safetensors16OfficialRuntime certified

Recommended hardware

Use only the listed transformers runtime configuration with the exact checkpoint precision. The recommendation engine checks weight memory, KV cache, runtime overhead, safety headroom, GPU architecture and topology before ranking rentable and reference hardware.

Estimated weight memory7.29 GB
KV cache2 GB
Runtime overhead6.39 GB
Total with safety margin19.45 GB

Recommended best value

1× NVIDIA A100 SXM 80GB

80 GB total VRAM · nvlink

60.55 GB headroom · Available

$915.00 / month

Lowest-cost viable

1× NVIDIA L4 24GB

24 GB total VRAM · pcie

4.55 GB headroom · Available

$250.00 / month

Highest performance

1× NVIDIA H200 SXM 141GB

141 GB total VRAM · nvswitch

121.55 GB headroom · Limited capacity

$2,345.00 / month

License and self-hosting considerations

The full 3.8B resident weight set must fit across the selected GPUs. Context and concurrency add KV-cache memory; MoE active parameters do not replace resident-weight memory.

Phi-4 Mini is published under mit. The release-scoped official license link and verification date are retained with the catalogue record.

Commercial Use Allowed. The official license permits commercial use; its obligations still apply.

Release timeline

microsoft/phi/phi-4-mini @ f79166923a03646efde5aa2c4b86837297c6ab419752f980c51f2e0325778f1c

Official source

Frequently asked questions

How much VRAM does Phi-4 Mini require?

The exact estimate depends on checkpoint precision, context and concurrency; the page recommendation uses the certified memory engine and compatible runtime configurations.

Which runtimes support Phi-4 Mini?

Certified upstream runtime evidence is available for transformers.

Can Phi-4 Mini be used commercially?

The stored commercial classification is Commercial Use Allowed. Consult the official mit terms for the final obligations.

Official references

Last verified: 2026-09-03