$ the-wire · showcase
AnyLanguageModel catches up to Foundation Models 26 and 27
By RepoJournal · Filed · About Hugging Face · Composed from the cited sources · methodology
Three stacked changes in huggingface/AnyLanguageModel land dynamic instructions, per-request context resolution, and reasoning transcript entries, while trl fixes a trainer-side truncation bug that silently produced empty completions.
Add sessions with dynamic instructions huggingface/AnyLanguageModel
Adds LanguageModelSession.init(model:dynamicInstructions:history:), which evaluates dynamic instructions before every request, including the continuation after tool calls, and sends the resolved instructions and tools in that request's context. The resolved instructions never enter the session transcript, so a restored conversation does not replay stale or re-evaluated instructions.
Resolve a request context before each model request huggingface/AnyLanguageModel
Introduces LanguageModelSession.RequestContext, carrying the transcript, instructions, and tools for a single request, plus resolvedRequestContext(). Every built-in adapter now reads those values from a context resolved before each request, and when a request produces tool calls the adapter runs them with the tools from that same context rather than reusing the ones it started with.
Represent reasoning as transcript entries huggingface/AnyLanguageModel
Reasoning becomes a first-class transcript entry, Transcript.Entry.reasoning(Transcript.Reasoning), with a stable ID, display segments, an opaque signature, and metadata. The built-in Anthropic provider fills these entries for thinking and redacted thinking in both streaming and nonstreaming responses, which is what lets it replay its own reasoning after a transcript is restored.
Discover environment tools on the class, not the instance huggingface/trl
GRPOTrainer and the async rollout worker used inspect.getmembers(env, predicate=inspect.ismethod) to find an environment's tools; because inspect calls getattr on every listed name before applying the predicate, every property on the environment ran once at trainer init and again for each rollout before every step. Discovery now happens on the class instead of the instance, so cached and side-e...
fix(cpo, orpo): truncate responses independently to prevent empty completions huggingface/trl
CPOTrainer and ORPOTrainer truncated responses against the length of the longer response, which could leave a slice with nothing in it; they now compute an independent budget from len(answer_tokens["prompt_input_ids"]) and bound it with max(0, ...) so a negative slice cannot silently drop tokens. It is the first half of the fix tracked for the CPO/SimPO truncation bug, so expect the remaining h...