-
Notifications
You must be signed in to change notification settings - Fork 2.9k
Pull requests: huggingface/trl
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Bump the actions group with 3 updates
dependencies
Pull requests that update a dependency file
github_actions
Pull requests that update GitHub Actions code
#6540
opened Jul 25, 2026 by
dependabot
Bot
Loading…
Fix SFTTrainer chunked-CE patch crash when forward is not a bound method
#6539
opened Jul 25, 2026 by
verma8076
Loading…
3 of 8 tasks
Update quantization tests for bitsandbytes 0.50.0
#6538
opened Jul 25, 2026 by
DaoyuanLi2816
Contributor
•
Draft
4 of 8 tasks
[DistillationTrainer refactor] Heavy-clean the Liger path: share extraction with the chunked path
#6537
opened Jul 24, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Wire compute_loss to the chunked JSD path; delete the full-logit loss
#6530
opened Jul 24, 2026 by
qgallouedec
Member
Loading…
Point the opencode example at huggingface/OpenEnv
#6529
opened Jul 24, 2026 by
sergiopaniego
Member
Loading…
3 of 6 tasks
Point the opencode example at openenv.core.sandbox
#6528
opened Jul 24, 2026 by
sergiopaniego
Member
•
Draft
2 of 6 tasks
Fix FSDP2 reference log-prob precomputation
#6527
opened Jul 24, 2026 by
DaoyuanLi2816
Contributor
Loading…
4 of 8 tasks
2
[DistillationTrainer refactor] Add the chunked JSD loss (unwired)
#6526
opened Jul 24, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Add
_get_last_hidden_state (unwired)
#6525
opened Jul 24, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Align the vLLM ctor block with GRPO
#6524
opened Jul 24, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Align the generation_kwargs dict with GRPO
#6523
opened Jul 23, 2026 by
qgallouedec
Member
Loading…
[DistillationTrainer refactor] Switch generation to GRPO's stack; delete the buffer
#6522
opened Jul 23, 2026 by
qgallouedec
Member
Loading…
Async GRPO: vision-language model (image) support
#6515
opened Jul 23, 2026 by
adithya-s-k
Collaborator
Loading…
Add bounded MoE expert usage metrics to SFTTrainer
#6514
opened Jul 23, 2026 by
kaining-never-stop
•
Draft
5 of 8 tasks
Respect
TQDM_DISABLE in DPO/KTO/BCO reference log-prob loops
#6507
opened Jul 22, 2026 by
qgallouedec
Member
Loading…
Default
use_bias_correction_kl=True in GRPOConfig
#6503
opened Jul 22, 2026 by
gowtham-sai-yadav
Contributor
Loading…
4 of 8 tasks
fix(dpo): apply f_divergence_type consistently in apo_down loss
#6475
opened Jul 20, 2026 by
Ewertonslv
•
Draft
4 of 8 tasks
Raise a clear error for QLoRA +
vllm_mode="server" instead of a cryptic vLLM crash
#6440
opened Jul 18, 2026 by
DaoyuanLi2816
Contributor
Loading…
1 task done
[GOLD] Support iterable datasets and collate each generation batch once per accumulation window
#6438
opened Jul 17, 2026 by
Strongich
Contributor
Loading…
Add privileged-context distillation to GOLD
#6437
opened Jul 17, 2026 by
eshwanthkartitr
Loading…
6 of 8 tasks
[OpenReward] Return None for unrewarded rollouts to prevent default-0.0 advantage poisoning
#6430
opened Jul 17, 2026 by
AUTHENSOR
Loading…
4 of 8 tasks
[SDPO] Port unscorable-mask reward defense from GRPO
#6429
opened Jul 17, 2026 by
AUTHENSOR
Loading…
5 of 8 tasks
Previous Next
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.