Z.ai опублікувала картку текстової моделі GLM-5.3. У картці наведено варіанти її локального розгортання.
GLM-5.3 використовує ту саму базову модель, що й GLM-5.2. Z.ai отримала поліпшення моделі післянавчанням.
За заявою Z.ai, GLM-5.3 поліпшила результат на 50% порівняно з GLM-5.2 у внутрішньому Z.ai Code Bench. Z.ai також заявляє про відкрите лідерство в Terminal Bench 3.0 та Agents’ Last Exam.
RadixArk опублікувала GLM-5.3-NVFP4 як квантизовану версію zai-org/GLM-5.3-BF16. У цій версії квантуються маршрутизовані експерти 75 MoE-шарів, що становить 96,2% параметрів.
Перевірка тверджень:
- Z.ai опублікувала картку текстової моделі GLM-5.3. Вона, зокрема, надає читачеві варіанти її локального розгортання. (підтверджено самою публікацією: доказ; «GLM-5.3 supports deployment with the following frameworks. Feel free to try them out:»)
- Z.ai заявляє, що GLM-5.3 заснована на тій самій базовій моделі, що й GLM-5.2, а поліпшення отримано післянавчанням. (підтверджено самою публікацією: доказ; «GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training.»)
- У картці GLM-5.3 заявлено поліпшення на 50% порівняно з GLM-5.2 у внутрішньому Z.ai Code Bench та відкритий SOTA у Terminal Bench 3.0 і Agents’ Last Exam. (підтверджено самою публікацією: доказ; «Stronger Coding: GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents’ Last Exam.»)
- Z.ai також заявляє для GLM-5.3 лідерство на CyberGym у пошуку вразливостей і більш ніж дворазову перевагу над GLM-5.2 на бенчмарках експлуатації. (підтверджено самою публікацією: доказ; «Emergent Cyber Capability: As we scaled post-training, cyber capability developed faster than we expected. GLM-5.3 is state of the art on CyberGym for vulnerability discovery, and its gains are largest further up the exploitation chain, where it more than doubles GLM-5.2 on exploitation benchmarks.»)
- Картка GLM-5.3 описує параметр reasoning_effort з рівнями low, high і max. За відсутності параметра використовується max. (підтверджено самою публікацією: доказ; «GLM-5.3 supports controlling the thinking budget through the reasoning_effort parameter, which accepts three levels: low , high , and max . It defaults to max if not passed (or if set to any other value).»)
- DavidsYoung опублікував 3,25-bpw квантизацію GLM-5.3 EXL3/TR3 зі змішаними рівнями K3/K4 без калібрувальних даних. (підтверджено самою публікацією: доказ; «GLM-5.3 — EXL3/TR3 3.25 bpw (mixed K3/K4 trellis, data-free)»)
- У цій EXL3/TR3-збірці маршрутизовані експерти шарів 3–78 квантизовано як 192 експерти K3 і 64 експерти K4 на шар, у середньому до 3,25 bpw. (підтверджено самою публікацією: доказ; «routed experts (layers 3–78, incl. MTP-78) EXL3 trellis, per layer 192 experts K3 + 64 experts K4 (avg 3.25 bpw), mcg codebook»)
- Для цієї EXL3/TR3-збірки вказано стек exllamav3-b12x/sparkinfer-lineage з патчем змішаних рівнів K. Стандартний завантажувач exllamav3 її не підтримує. (підтверджено самою публікацією: доказ; «TR3 mixed-bit layout (carrier BF16 shards + per-layer trellis payloads) — serve with an exllamav3-b12x/sparkinfer-lineage stack, TP4 (+ DCP4/MTP3), FP8 KV. The mixed-K projection-tiers patch is REQUIRED; a stock loader that assumes a uniform K per layer will produce fluent garbage. Not loadable by vanilla exllamav3 model loading.»)
- DavidsYoung повідомляє для 3,25-bpw збірки середній KLD 0,026103 з FP8 KV і 0,026776 у відтворенні CN3 на відкладених вікнах підтвердження. (підтверджено самою публікацією: доказ; «Measured and independently reproduced 2026-08-29: mean KLD 0.026103 (fp8 KV; CN3 reproduction 0.026776) on held-out confirmation windows.»)
- RadixArk опублікувала GLM-5.3-NVFP4 як квантизовану версію zai-org/GLM-5.3-BF16, створену NVIDIA Model Optimizer за схемою expert-only NVFP4 W4A4. (підтверджено самою публікацією: доказ; «The RadixArk GLM-5.3-NVFP4 model is the quantized version of zai-org/GLM-5.3-BF16 . The quantization was produced at RadixArk using NVIDIA Model Optimizer , following an expert-only NVFP4 W4A4 recipe.»)
- У GLM-5.3-NVFP4 квантуються маршрутизовані експерти 75 MoE-шарів: це 96,2% параметрів, а розмір чекпойнта заявлено зменшеним з 1 507 до 465 ГБ. (підтверджено самою публікацією: доказ; «The routed experts of the 75 MoE layers use NVFP4 W4A4 quantization with group size 16 — 57,600 linear entries, or 96.2% of parameters — with FP8-E4M3 block scales and static per-tensor activation scales. Sparse attention including the IndexShare indexer, shared experts, routers, the three dense MLP layers, all norms, embeddings, lm_head , and all MTP tensors retain the source BF16 precision. Checkpoint size is reduced from 1,507 GB to 465 GB.»)
- В опублікованому RadixArk тесті GLM-5.3-NVFP4 на восьми NVIDIA B300 модель отримала 97,42% на GSM8K і 94,17% на AIME 2026. (підтверджено самою публікацією: доказ; «The benchmark results below were produced with this NVFP4 checkpoint on 8x NVIDIA B300 GPUs using a TP8 SGLang deployment. Benchmark Evaluation protocol Score GSM8K Full 1,319-example split, single-shot, sgl-eval 97.42% (1,285/1,319) AIME 2026 30 problems x 16 rollouts, pass@1, sgl-eval 94.17% (majority@16 100% )»)
Публікації:
- https://huggingface.co/davidsyoung/GLM-5.3-EXL3-TR3-3.25bpw
- https://huggingface.co/RadixArk/GLM-5.3-NVFP4
Першоджерела:
- https://huggingface.co/zai-org/GLM-5.3
- https://huggingface.co/zai-org/GLM-5.3-BF16
- https://huggingface.co/davidsyoung/GLM-5.3-EXL3-TR3-3.42bpw
оцінка 42.3 · тип release · ревізія 4 · історії st-1171e08