Skip to main content

GLM · General · GLM-4.5

GLM-4.5: Specs, GPU requirements, hosting & deployment

GLM-4.5 is a legacy General release in the GLM family, certified from its immutable official source revision.

Technical deployment reference illustration for GLM-4.5
Reference deployment flow; exact architecture facts are listed separately.
Total parameters355.0B
Active parameters32.0B
Context window131,072 tokens
Release statusLegacy · Legacy
ArchitectureGlm4MoeForCausalLM
ModalitiesText
Licensemit
Deployment stateDeployable with certified configuration
Commercial stateCommercial Use Allowed
Release date2025-07-20

Architecture, strengths and use cases

GLM-4.5 is an official self-hostable Z.ai release in the GLM General track. It uses Glm4MoeForCausalLM, carries 355B / 32B MoE, supports a verified context window of 131,072 tokens, and exposes only source-certified checkpoints and runtime combinations. Its release, license and deployment facts are tied to the official repository revision rather than inferred from the family name.

GLM-4.5 uses Glm4MoeForCausalLM with 355B / 32B MoE. The stored architecture profile records 92 layers, hidden width 5120, 96 attention heads and a 131,072 token context ceiling where the official configuration publishes those fields.

Verified Glm4MoeForCausalLM architecture
131,072 token certified context
Official BF16 checkpoint coverage
Private LLM inference
Agentic workflows
Long-context document processing

Certified checkpoints and quantizations

Only exact, source-certified variants attached to this release are shown.

GLM-4.5

zai-org/GLM-4.5

Official · BF16 · safetensors

Runtime certified: SGLang Transformers

Precision / quantizationFormatBitsPublisher trustDeployment state
BF16safetensors16OfficialRuntime certified

Recommended hardware

Use only the listed sglang, transformers runtime configuration with the exact checkpoint precision. The recommendation engine checks weight memory, KV cache, runtime overhead, safety headroom, GPU architecture and topology before ranking rentable and reference hardware.

Estimated weight memory681.08 GB
KV cache2.88 GB
Runtime overhead173.38 GB
Total with safety margin1063.09 GB
Technically compatible reference sizing

No current FusionPointAI inventory passes every runtime, architecture, topology and memory gate for this exact checkpoint. Use the estimator for global compatible hardware.

License and self-hosting considerations

The full 355B resident weight set must fit across the selected GPUs. Context and concurrency add KV-cache memory; MoE active parameters do not replace resident-weight memory.

GLM-4.5 is published under mit. The release-scoped official license link and verification date are retained with the catalogue record.

Commercial Use Allowed. The official license permits commercial use; its obligations still apply.

Alternatives and sibling variants

GLM-5.3, GLM-5.2, GLM-5.1

Release timeline

zai/glm/glm-4-5 @ c6d66eec6dc6a38e2df9beb8f766aeef48d3d4d0b2fefe3f53c2d2820f56fdb3

Official source

Frequently asked questions

How much VRAM does GLM-4.5 require?

The exact estimate depends on checkpoint precision, context and concurrency; the page recommendation uses the certified memory engine and compatible runtime configurations.

Which runtimes support GLM-4.5?

Certified upstream runtime evidence is available for sglang, transformers.

Can GLM-4.5 be used commercially?

The stored commercial classification is Commercial Use Allowed. Consult the official mit terms for the final obligations.

Official references

Last verified: 2026-09-03