Skip to main content

GPT-OSS · General · GPT-OSS

GPT-OSS 120B: Specs, GPU requirements, hosting & deployment

GPT-OSS 120B is a current General release in the GPT-OSS family, certified from its immutable official source revision.

Technical deployment reference illustration for GPT-OSS 120B
Reference deployment flow; exact architecture facts are listed separately.
Total parameters117.0B
Active parameters5.1B
Context window131,072 tokens
Release statusCurrent · Current
ArchitectureGptOssForCausalLM
ModalitiesText
Licenseapache-2.0
Deployment stateDeployable with certified configuration
Commercial stateCommercial Use Allowed
Release date2025-08-04

Architecture, strengths and use cases

GPT-OSS 120B is an official self-hostable OpenAI release in the GPT-OSS General track. It uses GptOssForCausalLM, carries 117B / 5.1B MoE, supports a verified context window of 131,072 tokens, and exposes only source-certified checkpoints and runtime combinations. Its release, license and deployment facts are tied to the official repository revision rather than inferred from the family name.

GPT-OSS 120B uses GptOssForCausalLM with 117B / 5.1B MoE. The stored architecture profile records 36 layers, hidden width 2880, 64 attention heads and a 131,072 token context ceiling where the official configuration publishes those fields.

Verified GptOssForCausalLM architecture
131,072 token certified context
Official MXFP4, Unknown checkpoint coverage
Private LLM inference
Agentic workflows
Long-context document processing

Certified checkpoints and quantizations

Only exact, source-certified variants attached to this release are shown.

gpt-oss-120b

openai/gpt-oss-120b

Official · MXFP4 · safetensors

Runtime certified: Transformers

Precision / quantizationFormatBitsPublisher trustDeployment state
MXFP4safetensors4OfficialRuntime certified

Recommended hardware

Use only the listed transformers runtime configuration with the exact checkpoint precision. The recommendation engine checks weight memory, KV cache, runtime overhead, safety headroom, GPU architecture and topology before ranking rentable and reference hardware.

Estimated weight memory56.12 GB
KV cache0.14 GB
Runtime overhead26.69 GB
Total with safety margin102.86 GB
Technically compatible reference sizing

No current FusionPointAI inventory passes every runtime, architecture, topology and memory gate for this exact checkpoint. Use the estimator for global compatible hardware.

License and self-hosting considerations

The full 117B resident weight set must fit across the selected GPUs. Context and concurrency add KV-cache memory; MoE active parameters do not replace resident-weight memory.

GPT-OSS 120B is published under apache-2.0. The release-scoped official license link and verification date are retained with the catalogue record.

Commercial Use Allowed. The official license permits commercial use; its obligations still apply.

Alternatives and sibling variants

GPT-OSS 20B, gpt-oss-120b-Eagle3-v3, granitelib-rag-gpt-oss-r1.0

Release timeline

openai/gpt-oss/gpt-oss-120b @ 776a162c0c69e2eae8df64b37ac803df6f2dfe928aad74f49583f0532c30e946

Official source

Frequently asked questions

How much VRAM does GPT-OSS 120B require?

The exact estimate depends on checkpoint precision, context and concurrency; the page recommendation uses the certified memory engine and compatible runtime configurations.

Which runtimes support GPT-OSS 120B?

Certified upstream runtime evidence is available for transformers.

Can GPT-OSS 120B be used commercially?

The stored commercial classification is Commercial Use Allowed. Consult the official apache-2.0 terms for the final obligations.

Official references

Last verified: 2026-09-03