Wednesday, Jun 3, 2026
shippedLLM Compressor v0.11.0 with distributed quantization
Recoordinate released v0.11.0 of llm-compressor with DDP support for quantization methods and a refactored Compressed Tensors API [ref:1].
The release added distributed data parallel support for AWQ and SmoothQuant, achieving speedups of up to 3.2x on quantization workloads [1]. The Compressed Tensors API underwent a comprehensive refactor to improve lifecycle management for observers and related tooling. Model support expanded across the quantization and compression pipeline.
Sources
- v0.11.0 · vllm-project/llm-compressor