$ the-wire · showcase
Multi-eos completions in TRL, Kandinsky 6 lands in diffusers
By RepoJournal · Filed · About Hugging Face · Composed from the cited sources · methodology
TRL's trainers now end completions on every eos id a model declares rather than only the tokenizer's, Kandinsky 6 arrives in diffusers with new pipelines and schedulers, and xet-core ships 1.7.0 with client-side telemetry and RAM-aware download buffers.
End completions on every eos id the model declares in GRPO, RLOO and Distillation huggingface/trl
GRPOTrainer, RLOOTrainer and DistillationTrainer only watched the tokenizer's eos token, so models like Phi-3.5 and Gemma 3/4, which close a turn with ids their generation config declares (<|end|>, <end_of_turn>, <turn|>), kept generating past the intended end. They now stop on any eos id in the generation config; the change is part of a larger effort tracked in the same repo.
Kandinsky6 (#14949) huggingface/diffusers
The Kandinsky 6 implementation adds TI2VA and SR pipelines plus PiflowScheduler and MMAudioVAE, alongside docs and cleanup of redundant bigvgan code and manual checkpoint loading. einops and pydantic are dropped from the hard dependency set; torchvision, av and librosa become lazy imports.
[hf-xet v1.7.0]: Telemetry, Dynamic Download Buffers, and fixes huggingface/xet-core
by github-actions[bot]
hf-xet 1.7.0 adds client-side telemetry reports on every upload and download, and sizes download buffers dynamically from the RAM available to the machine or container. Release notes also list GIL release while waiting among the fixes.
Add PEFT support to AsyncDistillationTrainer huggingface/trl
AsyncDistillationTrainer gains peft_config, ported from the AsyncGRPO LoRA support: if the student's vLLM server was started with --enable-lora only the adapter is pushed on each sync, otherwise the adapter is merged and full weights go over NCCL as before. Teacher servers are unchanged, and compute_loss now avoids running the full causal forward baked into model.base_model on a PeftModel.
Refuse nn.DataParallel and drop the multi-GPU slow test job huggingface/trl
TRL now refuses nn.DataParallel, which Trainer applies when several GPUs are visible to one process (plain python, a notebook, accelerate launch --num_processes 1) and whose replicas crash at the first step on the fused LM head, chunked cross-entropy and Liger forwards. The multi-GPU slow test job is dropped: single-process multi-GPU runs need a real distributed launch instead.
feat: add machine readability to kernels variants huggingface/kernels
The long tail: a nix-builder change shares nixpkgs between build sets with equal backend versions, kernels variants gained a --json --only-compatible machine-readable output, v6 picked up rust cpu support, diffusers documented training on Jobs and kept explicit timesteps in FlowMatchEulerDiscreteScheduler.set_timesteps, and the doc-builder workflow pin moved.