GLM / GLM-5.3-Flash
GLM-5.3-Flash FP8 quantization
Estimated until measured during engineering
Before → After
Base releaseVerified checkpoint
320.000B
320.000B
quantizationFP8
Quality retention target
Quality retention target
Production artifactArtifact
Lineage and benchmark report
Lineage and benchmark report
GPU targeting
Weights and KV cacheEstimated until measured during engineering
Estimated VRAMMeasured and certified results are delivered with the project.
Runtime packagingtransformers, vllm, sglang, tokenspeed, ktransformers