$ the-wire · showcase
Transformers 5.18 ships Nemotron 3 diarization, TRL drops vLLM 0.20.0
By RepoJournal · Filed · About Hugging Face · Composed from the cited sources · methodology
Transformers 5.18.0 lands a streaming speaker-diarization model while TRL removes a vLLM version, adds selective activation checkpointing to SFT, and corrects how MFU is computed on non-H100 hardware.
Release 5.18.0 huggingface/transformers
Nemotron 3 Diarization is an open-weight speaker-diarization model that determines "who spoke when" in real-world audio, with both streaming and offline inference, up to eight speakers, and speaker outputs ordered by first arrival. A follow-up fixes the last STFT frame being dropped in streaming mode.
[MFU] Use device peak flops instead of hardcoded H100's huggingface/trl
Async trainers computed MFU without passing the device's peak flops, so the formula silently fell back to the H100 default. A new _PEAK_FLOPS_BY_DEVICE map (sourced from torchtitan and validated against NVIDIA specs, precision-aware) makes the number meaningful on non-H100 hardware.
Add selective activation checkpointing (SAC) to SFT huggingface/trl
Turn on with gradient_checkpointing_kwargs={"selective_activation_checkpointing": True} in SFTConfig: attention output is saved during forward instead of recomputed in backward, recovering most of the checkpointing slowdown at long context for one extra hidden-state-sized tensor per layer. Eager, like torchtitan under FSDP2. No torch.compile required.
Drop vLLM 0.20.0 support (#7435) huggingface/trl
vLLM 0.20.0 is no longer supported by TRL. If your training stack pins it, move off before upgrading.
Deprecate opencode_env and pi_env in favour of harbor_env huggingface/OpenEnv
OpenCode and Pi are validated Harbor harnesses (harness="opencode", harness="pi") with the same token-level capture for training, so opencode_env and pi_env are a redundant second path. Importing either now emits a FutureWarning naming the HarborSessionFactory call to use instead, and both are removed in OpenEnv 0.8.0.
Recommend upload_folder / hf upload for large uploads (#2832) huggingface/hub-docs
upload_large_folder is deprecated in the docs; upload_folder and hf upload now split large uploads into several commits and resume if interrupted. The transform/CI test page also picked up a scoped device memory panel, run-resolving links, and a WDYT? popup that streams an LLM's read on a failing test plus its real billed cost.
Action items
- → Replace opencode_env and pi_env imports with the harbor_env / HarborSessionFactory path before OpenEnv 0.8.0 lands huggingface/OpenEnv [plan]
- → Migrate off vLLM 0.20.0 in TRL training stacks huggingface/trl [plan]
- → Replace upload_large_folder calls with upload_folder or hf upload huggingface/hub-docs [plan]