GPTProto

Exploring Quantization Backends in Diffusers

Hugging Face:Blog(RSS)··Tutorials

Hugging Face 发布 Diffusers 量化后端指南,对比 bitsandbytes、torchao、Quanto、GGUF 和 FP8 Layerwise Casting 在 FLUX.1-dev 上把 BF16 加载内存约 31.447 GB 压缩到约 10.6 至 23.7 GB 的表现。