$ the-wire · showcase
QLoRA keeps DoRA in float32, flash-attention kwargs stop leaking into vision encoders
By RepoJournal · Filed · About Hugging Face · Composed from the cited sources · methodology
The day's changelog is dominated by silent-correctness fixes in the training stack, where a dtype that rounds away optimizer updates and kwargs meant for text that reach multimodal encoders are the kind of bug you only notice in your eval numbers.
Keep the DoRA magnitude vector in float32 under QLoRA huggingface/trl
The bf16 downcast QLoRA applies to every trainable parameter was also catching the DoRA magnitude vector, whose optimizer updates can be smaller than bfloat16 can represent, so the vector silently froze. The magnitude vector now stays in float32; under FSDP2 the QDoRA path now fails at wrap time with "FSDP expects uniform original parameter dtype" rather than training on a frozen vector.
Don't forward text-level flash-attention kwargs to multimodal encoders (#49227) huggingface/transformers
Text-level flash-attention kwargs were being passed through into multimodal encoders, where they don't belong. The commit removes that forwarding, which matters if you run a vision-language model with flash-attention settings set at the text level.
refactor(dtype): unify the dtype configuration of policies huggingface/lerobot
LeRobot policies configured precision in incompatible ways: seven policies stored `dtype` as a string with their own private converter, and fastwam, vla_jepa, evo1 and groot used separate keys (`torch_dtype`, `vlm_dtype`, `model_params_fp32`). The refactor introduces `PreTrainedConfig.dtype: torch.dtype | None` as the single dtype key for every policy. This is a breaking config change: the old ...
Honor logits_scaling and lm_head_multiplier in the fused LM head huggingface/trl
The fused LM head and the chunked log-prob and loss paths project hidden states themselves instead of calling the model's `forward`, so they have to reproduce any logit scaling the model applies. They already covered `logit_scale` (Cohere) and `output_multiplier` (Muse Glimmer); this adds `logits_scaling`, which Granite (including Granite 4.0 and `granite4_vision`) and MiniCPM3 divide by.
Add Trainer.loss_is_scaled_for_ga to declare whether compute_loss already scales for gradient accumulation huggingface/transformers
A single `Trainer.loss_is_scaled_for_ga` flag lets `compute_loss` declare that it already scales for gradient accumulation, so the Trainer stops double-scaling it. Custom losses that return a pre-scaled value can now say so instead of dividing again.
fix(annotations): store VQA bbox and keypoint coordinates as [0, 1] image fractions (#4819) huggingface/lerobot
In the long tail: VQA bounding-box and keypoint coordinates are now stored as [0, 1] image fractions rather than pixel values, and `tasks-v0.21.57` shipped from huggingface.js.
Action items