The Qwen3.8-Flash-Next-FP8 repository contains model weights quantized to FP8 and configuration files for the post-trained model in the Hugging Face Transformers format. FP8 is an 8-bit floating-point format.
The artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, and TokenSpeed. The quantization uses blocks of 128, and its performance metrics are nearly identical to those of the original model.
Claim check:
- The Qwen3.8-Flash-Next-FP8 repository contains model weights quantized to FP8 and configuration files for the post-trained model in the Hugging Face Transformers format. (confirmed by the publication itself: evidence; «This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format.»)
- The artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, and TokenSpeed. (confirmed by the publication itself: evidence; «These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.»)
- The quantization uses blocks of 128, and its performance metrics are nearly identical to those of the original model. (confirmed by the publication itself: evidence; «The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.»)
Publications:
Primary sources:
- https://huggingface.co/Qwen/Qwen3.8-Flash-Next-FP8
- https://huggingface.co/spiritfather/Qwen3.8-Flash-Next-heretic-2-GGUF
score 27.8 out of 100 · kind: release · update 3