130 wires and counting

$ follow Hugging Face

Keep up with Hugging Face in about 3 minutes: what actually shipped — the commits, pull requests, releases, and security advisories that matter.

or

fair warning: these emails are deeply technical. diffs, version numbers, CVEs, benchmark deltas. if that's not your idea of a good read, this isn't your newsletter.

Folds into your digest — weekly by default, monthly if you prefer. Unsubscribe in one click.

$ status

wire 2026-10-03
stories 48

© 2026 RepoJournal Home Showcase How it works Privacy

$ the-wire · showcase

QLoRA keeps DoRA in float32, flash-attention kwargs stop leaking into vision encoders

By RepoJournal · Filed · About Hugging Face · Composed from the cited sources · methodology

The day's changelog is dominated by silent-correctness fixes in the training stack, where a dtype that rounds away optimizer updates and kwargs meant for text that reach multimodal encoders are the kind of bug you only notice in your eval numbers.

Keep the DoRA magnitude vector in float32 under QLoRA huggingface/trl

by qgallouedec

The bf16 downcast QLoRA applies to every trainable parameter was also catching the DoRA magnitude vector, whose optimizer updates can be smaller than bfloat16 can represent, so the vector silently froze. The magnitude vector now stays in float32; under FSDP2 the QDoRA path now fails at wrap time with "FSDP expects uniform original parameter dtype" rather than training on a frozen vector.

Don't forward text-level flash-attention kwargs to multimodal encoders (#49227) huggingface/transformers

by BADAOUI Abdennacer

Text-level flash-attention kwargs were being passed through into multimodal encoders, where they don't belong. The commit removes that forwarding, which matters if you run a vision-language model with flash-attention settings set at the text level.

refactor(dtype): unify the dtype configuration of policies huggingface/lerobot

by HaomingSong

LeRobot policies configured precision in incompatible ways: seven policies stored `dtype` as a string with their own private converter, and fastwam, vla_jepa, evo1 and groot used separate keys (`torch_dtype`, `vlm_dtype`, `model_params_fp32`). The refactor introduces `PreTrainedConfig.dtype: torch.dtype | None` as the single dtype key for every policy. This is a breaking config change: the old ...

Honor logits_scaling and lm_head_multiplier in the fused LM head huggingface/trl

by YaseenBashaT

The fused LM head and the chunked log-prob and loss paths project hidden states themselves instead of calling the model's `forward`, so they have to reproduce any logit scaling the model applies. They already covered `logit_scale` (Cohere) and `output_multiplier` (Muse Glimmer); this adds `logits_scaling`, which Granite (including Granite 4.0 and `granite4_vision`) and MiniCPM3 divide by.

Add Trainer.loss_is_scaled_for_ga to declare whether compute_loss already scales for gradient accumulation huggingface/transformers

by qgallouedec

A single `Trainer.loss_is_scaled_for_ga` flag lets `compute_loss` declare that it already scales for gradient accumulation, so the Trainer stops double-scaling it. Custom losses that return a pre-scaled value can now say so instead of dividing again.

fix(annotations): store VQA bbox and keypoint coordinates as [0, 1] image fractions (#4819) huggingface/lerobot

by Maxime Ellerbach

In the long tail: VQA bounding-box and keypoint coordinates are now stored as [0, 1] image fractions rather than pixel values, and `tasks-v0.21.57` shipped from huggingface.js.

Quick answers

What shipped in Hugging Face on October 3, 2026?
The day's changelog is dominated by silent-correctness fixes in the training stack, where a dtype that rounds away optimizer updates and kwargs meant for text that reach multimodal encoders are the kind of bug you only notice in your eval numbers. In total, 24 commits, 22 pull requests, and 2 releases landed.
Who contributed to Hugging Face on October 3, 2026?
14 developers shipped this update, including qgallouedec, Behrooz Azarkhalili, YaseenBashaT, aazizyan, albertvillanova, BADAOUI Abdennacer, Yih-Dar, and HaomingSong, and 6 more.
What were the notable Hugging Face updates?
Keep the DoRA magnitude vector in float32 under QLoRA, Don't forward text-level flash-attention kwargs to multimodal encoders (#49227), and refactor(dtype): unify the dtype configuration of policies.