Z.ai published the GLM-5.3 text model card. The card lists options for local deployment.

GLM-5.3 uses the same base model as GLM-5.2. Z.ai obtained the model improvements through post-training.

Z.ai says GLM-5.3 improved its result by 50% over GLM-5.2 on the internal Z.ai Code Bench. Z.ai also claims public leadership on Terminal Bench 3.0 and Agents’ Last Exam.

RadixArk published GLM-5.3-NVFP4 as a quantized version of zai-org/GLM-5.3-BF16. This version quantizes the routed experts in 75 MoE layers, which account for 96.2% of the parameters.

Claim check:

  • Z.ai published the GLM-5.3 text model card. Among other things, it gives readers options for local deployment. (confirmed by the publication itself: evidence; «GLM-5.3 supports deployment with the following frameworks. Feel free to try them out:»)
  • Z.ai says GLM-5.3 is based on the same base model as GLM-5.2 and that its improvements came from post-training. (confirmed by the publication itself: evidence; «GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training.»)
  • The GLM-5.3 card claims a 50% improvement over GLM-5.2 on the internal Z.ai Code Bench and public SOTA on Terminal Bench 3.0 and Agents’ Last Exam. (confirmed by the publication itself: evidence; «Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents’ Last Exam.»)
  • Z.ai also claims that GLM-5.3 leads on CyberGym for vulnerability discovery and exceeds GLM-5.2 by more than twofold on exploitation benchmarks. (confirmed by the publication itself: evidence; «Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.»)
  • The GLM-5.3 card describes the reasoning_effort parameter with low, high, and max levels. When the parameter is absent, it uses max. (confirmed by the publication itself: evidence; «GLM-5.3 supports controlling the thinking budget through the reasoning_effort parameter, which accepts three levels: low , high , and max . It defaults to max if not passed (or if set to any other value).»)
  • DavidsYoung published a 3.25-bpw GLM-5.3 EXL3/TR3 quantization with mixed K3/K4 levels and no calibration data. (confirmed by the publication itself: evidence; «GLM-5.3 — EXL3/TR3 3.25 bpw (mixed K3/K4 trellis, data-free)»)
  • In this EXL3/TR3 build, the routed experts in layers 3–78 are quantized as 192 K3 experts and 64 K4 experts per layer, averaging 3.25 bpw. (confirmed by the publication itself: evidence; «routed experts (layers 3–78, incl. MTP-78) EXL3 trellis, per layer 192 experts K3 + 64 experts K4 (avg 3.25 bpw), mcg codebook»)
  • This EXL3/TR3 build specifies the exllamav3-b12x/sparkinfer-lineage stack with a mixed K-level patch. The standard exllamav3 loader does not support it. (confirmed by the publication itself: evidence; «TR3 mixed-bit layout (carrier BF16 shards + per-layer trellis payloads) — serve with an exllamav3-b12x/sparkinfer-lineage stack, TP4 (+ DCP4/MTP3), FP8 KV. The mixed-K projection-tiers patch is REQUIRED; a stock loader that assumes a uniform K per layer will produce fluent garbage. Not loadable by vanilla exllamav3 model loading.»)
  • DavidsYoung reports an average KLD of 0.026103 with FP8 KV and 0.026776 in a CN3 reproduction on held-out confirmation windows for the 3.25-bpw build. (confirmed by the publication itself: evidence; «Measured and independently reproduced 2026-08-29: mean KLD 0.026103 (fp8 KV; CN3 reproduction 0.026776) on held-out confirmation windows.»)
  • RadixArk published GLM-5.3-NVFP4 as a quantized version of zai-org/GLM-5.3-BF16, created with NVIDIA Model Optimizer using the expert-only NVFP4 W4A4 scheme. (confirmed by the publication itself: evidence; «The RadixArk GLM-5.3-NVFP4 model is the quantized version of zai-org/GLM-5.3-BF16 . The quantization was produced at RadixArk using NVIDIA Model Optimizer , following an expert-only NVFP4 W4A4 recipe.»)
  • In GLM-5.3-NVFP4, the routed experts in 75 MoE layers are quantized. They account for 96.2% of the parameters, and the claimed checkpoint size drops from 1,507 to 465 GB. (confirmed by the publication itself: evidence; «The routed experts of the 75 MoE layers use NVFP4 W4A4 quantization with group size 16 — 57,600 linear entries, or 96.2% of parameters — with FP8-E4M3 block scales and static per-tensor activation scales. Sparse attention including the IndexShare indexer, shared experts, routers, the three dense MLP layers, all norms, embeddings, lm_head , and all MTP tensors retain the source BF16 precision. Checkpoint size is reduced from 1,507 GB to 465 GB.»)
  • In a published RadixArk test of GLM-5.3-NVFP4 on eight NVIDIA B300 GPUs, the model scored 97.42% on GSM8K and 94.17% on AIME 2026. (confirmed by the publication itself: evidence; «The benchmark results below were produced with this NVFP4 checkpoint on 8x NVIDIA B300 GPUs using a TP8 SGLang deployment. Benchmark Evaluation protocol Score GSM8K Full 1,319-example split, single-shot, sgl-eval 97.42% (1,285/1,319) AIME 2026 30 problems x 16 rollouts, pass@1, sgl-eval 94.17% (majority@16 100% )»)

Publications:

Primary sources:

score 42.3 · kind release · revision 4 · stories st-1171e08