GLM-5.1
Official · BF16 · safetensors
Runtime certified: Transformers
| Precision / quantization | Format | Bits | Publisher trust | Deployment state |
|---|---|---|---|---|
| BF16 | safetensors | 16 | Official | Runtime certified |
GLM · General · GLM-5.1
GLM-5.1 is a previous General release in the GLM family, certified from its immutable official source revision.
GLM-5.1 is an official self-hostable Z.ai release in the GLM General track. It uses GlmMoeDsaForCausalLM, carries 744B / 40B MoE, supports a verified context window of 202,752 tokens, and exposes only source-certified checkpoints and runtime combinations. Its release, license and deployment facts are tied to the official repository revision rather than inferred from the family name.
GLM-5.1 uses GlmMoeDsaForCausalLM with 744B / 40B MoE. The stored architecture profile records 78 layers, hidden width 6144, 64 attention heads and a 202,752 token context ceiling where the official configuration publishes those fields.
Only exact, source-certified variants attached to this release are shown.
Official · BF16 · safetensors
Runtime certified: Transformers
| Precision / quantization | Format | Bits | Publisher trust | Deployment state |
|---|---|---|---|---|
| BF16 | safetensors | 16 | Official | Runtime certified |
Official · FP8 · safetensors
Runtime certified: Transformers
| Precision / quantization | Format | Bits | Publisher trust | Deployment state |
|---|---|---|---|---|
| FP8 | safetensors | 8 | Official | Runtime certified |
Use only the listed transformers runtime configuration with the exact checkpoint precision. The recommendation engine checks weight memory, KV cache, runtime overhead, safety headroom, GPU architecture and topology before ranking rentable and reference hardware.
No current FusionPointAI inventory passes every runtime, architecture, topology and memory gate for this exact checkpoint. Use the estimator for global compatible hardware.
The full 744B resident weight set must fit across the selected GPUs. Context and concurrency add KV-cache memory; MoE active parameters do not replace resident-weight memory.
GLM-5.1 is published under mit. The release-scoped official license link and verification date are retained with the catalogue record.
Commercial Use Allowed. The official license permits commercial use; its obligations still apply.
GLM-5.3, GLM-5.2, GLM-5
The exact estimate depends on checkpoint precision, context and concurrency; the page recommendation uses the certified memory engine and compatible runtime configurations.
Certified upstream runtime evidence is available for transformers.
The stored commercial classification is Commercial Use Allowed. Consult the official mit terms for the final obligations.
Last verified: 2026-09-03