$ the-wire · showcase
lerobot fp16 training lands on every topology
By RepoJournal · Filed · About Hugging Face · Composed from the cited sources · methodology
lerobot finished fp16 mixed-precision training across single-GPU, DDP and FSDP2/HSDP, and transformers fixed a mamba2 chunk scan that was materializing a 4 GiB intermediate at Bamba-9B shapes.
feat(train): add support for fp16 mixed precision (#4652) huggingface/lerobot
The loss scaler machinery was never wired up, so `--accelerator.mixed_precision=fp16` was accepted in non-sharded training while doing nothing, and it was rejected outright under sharded training. Now `--accelerator.mixed_precision=fp16` works under single-GPU, DDP and FSDP2/HSDP: accelerate builds a plain `torch.amp.GradScaler` for FSDP2 and torch reduces the scaler's found-inf flag across the...
Contract the mamba2 chunk scan with einsum instead of broadcast-then-sum (#48978) huggingface/transformers
The fallback `mamba2_chunk_scan` built each contraction as a broadcast product followed by a sum, materializing the un-summed tensor first: G's is (batch, chunks, chunk, chunk, heads, state) in float32, 4 GiB per sequence at bamba-9B's shapes. Reading the same data with einsum removes that intermediate, and BambaModelIntegrationTest no longer OOMs asking for 8 GiB on a 22.3 GiB runner for a 10-...
refactor(processors): one path for building processors from a checkpoint huggingface/lerobot
`make_pre_post_processors` used to resolve the pretrained path through a dispatcher plus three hardcoded `isinstance` branches for groot, molmoact2 and evo1; it now always resolves `make_{type}_pre_post_processors_from_pretrained` by naming convention, and evo1 loads through a from_pretrained hook. If you maintain a policy with custom checkpoint loading, you no longer need to edit `policies/fac...
Keep image processor backends in sync on keys and dtypes (#48739) huggingface/transformers
The torchvision backends of Idefics2, Idefics3 and SmolVLM allocated the batch `pixel_attention_mask` buffer without a dtype, so int64 masks built by `pad()` were stored as float32 while the PIL backend returned int64. The Fuyu PIL backend computed original image sizes and then dropped them, so its output lacked the `image_sizes` key the torchvision backend emits. Both backends now agree on key...
fix(train): only record gpu_mem_gb when the tracker declares it (#4714) huggingface/lerobot
Elsewhere: `gpu_mem_gb` is now recorded only when the tracker declares it, fixing a red CI run, and the Fuyu/Idefics work and the Hungarian matcher fix round out a quiet transformers day.