-
Notifications
You must be signed in to change notification settings - Fork 1.7k
Pull requests: modelscope/ms-swift
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(trainer): support selective logits with sequence parallel
#10039
opened Sep 4, 2026 by
zjn20030811
Loading…
Support mixed text and raw-token rollout content in Swift template
#10029
opened Sep 2, 2026 by
kerbeans
Loading…
refactor(sequence_parallel): centralize SP wiring behind SPStrategy facade
#10028
opened Sep 2, 2026 by
werwrewe
Loading…
1 of 5 tasks
feat(gkd): add agentic monitoring metrics
#10026
opened Sep 2, 2026 by
Casuallkk
Loading…
1 of 4 tasks
perf(rlhf): vectorize padding-free log-probability restore
#10021
opened Sep 1, 2026 by
Ruihan11
Contributor
Loading…
1 of 5 tasks
perf(npu): vectorize ring attention LSE extractio
#10002
opened Aug 28, 2026 by
Ruihan11
Contributor
Loading…
1 of 5 tasks
docs: document OrcaRouter as an OpenAI-compatible sampling provider
#9995
opened Aug 27, 2026 by
nissrin2020ali-ux
Loading…
1 of 4 tasks
feat(fsdp2): load model on meta device for non-rank0 ranks (0 CPU RAM per worker)
#9982
opened Aug 24, 2026 by
cben484
Contributor
Loading…
fix: align channel loss with sequence parallel labels
#9977
opened Aug 24, 2026 by
taking-lying-flat
Contributor
Loading…
1 of 4 tasks
feat(grpo): add M2PO for stale-rollout training
#9965
opened Aug 21, 2026 by
primorLee
Contributor
Loading…
2 of 4 tasks
Add --dataloader_multiprocessing_context to work around Python 3.14 incompatibility
#9941
opened Aug 18, 2026 by
sliedes
Contributor
Loading…
1 of 4 tasks
[Megatron] Preserve RNG state across checkpoint resume
#9935
opened Aug 17, 2026 by
taking-lying-flat
Contributor
Loading…
Fix reward model margin broadcasting and alignment
#9927
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
Fix DPO IPO log-prob normalization
#9925
opened Aug 16, 2026 by
taking-lying-flat
Contributor
Loading…
1 of 4 tasks
Fix BatchSamplerShard tail sampling
#9907
opened Aug 14, 2026 by
taking-lying-flat
Contributor
Loading…
[Train] Reduce padding-free embedding output memory
#9893
opened Aug 11, 2026 by
taking-lying-flat
Contributor
Loading…
fix(deploy): handle reasoning-only stream chunks
#9881
opened Aug 11, 2026 by
taking-lying-flat
Contributor
Loading…
1 of 4 tasks
Fix IterablePackingDataset workers to use spawn
#9876
opened Aug 9, 2026 by
zupengwang
Loading…
1 of 4 tasks
[Feature] Add --streaming_shard: split the streaming dataset across data-parallel ranks
#9860
opened Aug 5, 2026 by
Rapisurazurite
Loading…
1 of 4 tasks
Fix Qwen3.5 multimodal packing kwargs compatibility
#9838
opened Aug 3, 2026 by
taking-lying-flat
Contributor
Loading…
Support model initialization without loading pretrained weights
#9821
opened Jul 30, 2026 by
LiXinYuECNU
Loading…
feat(sft):Add optional fused linear ce for qwen2/qwen3 on npu
#9784
opened Jul 22, 2026 by
yangguang-zhang
Loading…
1 of 4 tasks
feat: support epoch-based checkpoint intervals
#9782
opened Jul 22, 2026 by
tongchen126
Contributor
Loading…
2 of 4 tasks
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.