$ the-wire · showcase
TRL moves every trainer onto the fused LM head
By RepoJournal · Filed · About Hugging Face · Composed from the cited sources · methodology
TRL's fused LM head now scores tokens for SFT, DPO and KTO, collapsing four separate loss paths into one and speeding up training across the board.
Score SFT tokens with the fused LM head huggingface/trl
SFT now scores tokens through the fused LM head the way GRPO, RLOO, DPO and KTO already do, with the loss computed from per-token log-probs. Its chunked cross-entropy and Liger loss path are gone, and nll and dft both run on the fused path; measured against main across 23 configs on H100 at 1 and 8 GPUs, step time improved 1.18x at the median (range 1.04x to 1.36x) with peak VRAM unchanged.
Score DPO and KTO tokens with the fused LM head huggingface/trl
DPO and KTO follow the same conversion: policy and reference models get a fused head via add_fused_lm_head, every scoring call passes fused_lm_head=True, and the [batch, seq, vocab] logits are never built. The full-logits path and the use_liger_kernel chunked path are removed; the kernel gains three optional per-token outputs so existing logging is unaffected.
Skip softmax over single-label CrossEncoder scores (#4123) huggingface/sentence-transformers
predict(apply_softmax=True) is documented to softmax only when model.num_labels > 1, but single-label scores are still [batch, 1] at the check, so everything collapsed to 1.0. Because rank() accepts only single-label models and passes apply_softmax through, rank(..., apply_softmax=True) was returning every score as 1.0 in input order.
Apply model2vec token mapping and weights when loading StaticEmbedding huggingface/sentence-transformers
StaticEmbedding.load, from_model2vec and from_distillation now fold model2vec's optional mapping and per-token weights into the matrix as embedding[id] = embeddings[mapping[id]] * weights[id]. Models shipping the extra tensors, such as minishlab/potion-code-16M, were producing different embeddings than model2vec, with cosine similarities of 0.64 and 0.93 in the reported case.
Fix redundant DDP synchronization in cached losses huggingface/sentence-transformers
Cached losses replay a forward and backward per mini-batch per input column, and under DDP each replay synchronized gradients; one example in issue 4027 hit eight synchronization rounds. This removes the per-mini-batch gradient sync and adds distributed regression tests for gradient equivalence, accumulation, and frozen inputs.
[docs] Community methods huggingface/diffusers
The Community methods docs drop Token merging, DeepCache and TGATE because they largely target older UNet models and Diffusers' built-in caching now covers similar ground, and remove the xFormers page in favor of the Attention backends doc, with inbound links redirected there. The OpenEnv release train and an active-adapter fix for text-encoder-only LoRAs round out a quiet day elsewhere.