В руководстве Z.ai по GLM-5.3-Flash сообщается, что модель полностью доступна в GLM Coding Plan. Это выпуск модели с нативной мультимодальностью и втрое большей квотой.

У модели 320 млрд параметров, из которых активны 18 млрд. Гибридная архитектура sparse и linear attention, по данным Z.ai, сокращает вычисления attention в 3,01 раза и KV-кэш в 4,44 раза относительно GLM-5.3.

Модель принимает видео, изображения, текст и файлы, а выдаёт текст. Она поддерживает Function Calling, Context Caching и структурированный вывод JSON.

Проверка утверждений:

  • Z.ai в руководстве по GLM-5.3-Flash сообщила, что модель полностью доступна в GLM Coding Plan; читателю это даёт доступ к нативной мультимодальности и втрое большей квоте. (подтверждено самой публикацией: доказательство; «GLM-5.3-Flash is now fully available on the GLM Coding Plan . With native multimodal capabilities and 3× the quota, it delivers a smoother and more cost-effective coding experience.»)
  • У GLM-5.3-Flash 320 млрд параметров, из которых активны 18 млрд; гибридная архитектура sparse и linear attention, по данным Z.ai, сокращает вычисления attention в 3,01 раза и KV-кэш в 4,44 раза относительно GLM-5.3. (подтверждено самой публикацией: доказательство; «GLM-5.3-Flash has 320B total parameters with 18B activated. As the first open-source frontier model to combine sparse and linear attention, it significantly cuts computation and serving costs while preserving long-context quality — reducing attention computation and KV cache by 3.01× and 4.44× versus GLM-5.3.»)
  • Модель принимает видео, изображения, текст и файлы, выдаёт текст, поддерживает контекст до 1 млн токенов и максимальный вывод 128 тыс. токенов. (подтверждено самой публикацией: доказательство; «Input Modality Video / Image / Text / File Output Modality Text Context Length 1M Maximum Output Tokens 128K»)
  • В API изображение передаётся через блок image_url как URL или Base64 Data URL; в одном запросе можно добавить несколько изображений. (подтверждено самой публикацией: доказательство; «Image Parameters :Add a content block with type: image_url to messages[].content[] , and pass the image URL (recommended) or a Base64 Data URL through image_url.url . Multiple images can be added by including multiple image_url content blocks.»)
  • В GLM Coding Plan GLM-5.3-Flash даёт втрое больше квоты, чем GLM-5.3. (подтверждено самой публикацией: доказательство; «Now fully available, GLM-5.3-Flash can be used with your preferred tools, with 3× the available quota compared with GLM-5.3.»)
  • Вызовы в GLM Coding Plan в непиковые часы, включая все выходные, расходуют половину обычных баллов. (подтверждено самой публикацией: доказательство; «The new GLM Coding Plan adopts a points-based quota system with transparent usage limits. Model calls made during off-peak hours, including all day on weekends, consume only 50% of the standard points.»)
  • Документация Z.ai указывает, что GLM-5.3-Flash поддерживает Function Calling, Context Caching, структурированный вывод JSON и нативный мультимодальный ввод изображений, видео и файлов. (подтверждено самой публикацией: доказательство; «Function Calling : Provides powerful tool-calling capabilities and supports integration with a wide range of external tools. Context Caching : Uses an intelligent caching mechanism to optimize long-conversation performance. Structured Output : Supports structured output formats such as JSON for seamless system integration. Visual Understanding : Native multimodal input supporting images, videos, and files.»)
  • Карточка модели Z.ai на Hugging Face сообщает, что GLM-5.3-Flash можно развёртывать локально через SGLang, vLLM, TokenSpeed, Transformers, KTransformers и Unsloth. (подтверждено самой публикацией: доказательство; «GLM-5.3-Flash supports deployment with the following frameworks. Feel free to try them out: - SGLang — see cookbook - vLLM — see recipes - TokenSpeed — see here - Transformers — see transformers docs - KTransformers — see tutorial - Unsloth — see guide»)
  • Параметр reasoning_effort у GLM-5.3-Flash принимает уровни low, high и max; по умолчанию используется max. (подтверждено самой публикацией: доказательство; «GLM-5.3-Flash supports controlling the thinking budget through the reasoning_effort parameter, which accepts three levels: low, high, and max. It defaults to max if not passed (or if set to any other value).»)
  • Unsloth сообщает, что GLM-5.3-Flash можно запускать в интерфейсе Unsloth Desktop. (подтверждено самой публикацией: доказательство; «You can now run GLM-5.3-Flash in our Unsloth Desktop UI.»)
  • В карточке cyankiwi размер модели GLM-5.3-Flash-AWQ-INT4 указан как 212,72 GB. (подтверждено самой публикацией: доказательство; «Model Size 212.72 GB»)
  • The New Stack опубликовал материал-сравнение GLM-5.3-Flash и GLM-5.3, сфокусированный на времени и стоимости, а не только на спецификациях. (подтверждено самой публикацией: доказательство; «GLM-5.3-Flash vs. GLM-5.3: Time and money, not the spec sheet»)

Первоисточники:

оценка 35.9 · тип release · ревизия 4 · истории st-3lteq7