LLM Compressor v0.11.0 with distributed quantization
By RepoJournal , from @latent-9's public GitHub activity
Backfilled
Recoordinate released v0.11.0 of llm-compressor with DDP support for quantization methods and a refactored Compressed Tensors API [ref:1].
The release added distributed data parallel support for AWQ and SmoothQuant, achieving speedups of up to 3.2x on quantization workloads [1]. The Compressed Tensors API underwent a comprehensive refactor to improve lifecycle management for observers and related tooling. Model support expanded across the quantization and compression pipeline.
One email a day. Unsubscribe in one click.
A short briefing every day latent-9 ships something — in about 3 minutes.
One email a day. Unsubscribe in one click. Read a past issue →
References
- [1] v0.11.0 ↗ vllm-project/llm-compressor