Repository navigation
【No.3】unify role inference infrastructure - #347
ldemon2333 wants to merge 14 commits into
Conversation
Nyanpasu 审查看板审查状态: ✅ 已通过 审查版本: 43e1a21 目标分支: main 本轮复核(head 43e1a21):上轮两项遗留 P2(Teacher client 按样本重建;defer 预检拒绝现有示例)已修复并经代码与本地实验证实;17 项既往发现保持已解决或已取代,无新增。defer 示例完成预检级验证,多机执行由维护者 nightly 覆盖。结论 APPROVE。
审查发现待处理
已解决或已取代
提交范围 · 接收 98 · 建议移出 0 · 待确认 0接收 98 个文件 · 建议移出 0 个文件 · 待确认 0 个文件。移出与待确认部分暂停深审,不代表审查通过。
精简审查与验证依据
生产代码的必要性与替代方案
测试的必要性与替代方案
Powered by Nyanpasu with glm-5.3-flash max, please check the suggestions carefully.
|
||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
6a05181 to
9f82f32
Compare
9f82f32 to
44ba49a
Compare
44ba49a to
eaa767a
Compare
|
@SigureMo 任务已完成,请求 review。可以 approve workflows 吗 |
eaa767a to
d38bbec
Compare
d38bbec to
13c0cb7
Compare
rai-studio-bot
left a comment
There was a problem hiding this comment.
已复核 13c0cb7。批量缩容能保留成功结果,但仍有一项多节点部分清理后无法重试的 P2,详见新增行级评论;此前三项问题保持已修复。
本地 50 项 scale-in 测试因缺少 Ray/SGLang 依赖全部跳过。依赖隔离验证覆盖正常部分失败重试,并复现了上述问题;未执行真实多节点或 GPU 验收。
| with self._engine_lifecycle_lock: | ||
| for i, _ in live_actors: | ||
| for i in released: | ||
| group.all_engines[i] = None |
There was a problem hiding this comment.
请保留多节点副本在部分 shutdown 成功后的可重试清理入口。若 node0 成功、follower 暂时失败,这里会留下 all_engines=[None, follower],group 保持 DRAINING;但 _select_engines_for_removal() 只遍历 group.engines(node0 切片),后续请求无法再次选中它。Eviction 同样拒绝 node0 为空的副本,recover 则跳过 DRAINING 弹性组,导致剩余 worker/PG 持续占用,并维持权重更新 fence。
用实际方法的依赖隔离复现确认:即使 follower 随后恢复可正常 shutdown,新缩容请求仍返回 “No engines selected for removal”。建议从副本中任一残留 worker 识别待清理目标,或显式保存待重试的副本状态,同时让重试跳过已清理的 head;补充“head 成功、follower 失败后再次清理”的多节点契约测试。
There was a problem hiding this comment.
981a275 已修复数量缩容和 eviction 的残留 follower 清理,F4 部分解决;按原 engine_urls 重试仍失败。
_resolve_scale_in_url_candidates() 现在会查询残留 follower,但真实 SGLangEngine.get_url()(relax/backends/sglang/sglang_engine.py:909–914)对非 node0 返回 None,因此无法匹配原 URL。用实际管理器方法并按该契约设置 follower 后,隔离验证仍得到 “No engines selected for removal”;改用数量缩容则可完成清理。
建议保留副本原 URL 与逻辑索引的关联,让 URL 重试也能选中残留 worker,并补充 follower URL 为 None、重复原 URL 请求的回归。当前新增测试只以数量重试,未覆盖此路径。
There was a problem hiding this comment.
7813d61 已修复原始同 URL 重试,但恢复后地址变化的场景仍有缺口。replica_urls 只在首次注册和 URL 选择时更新;RolloutServer.recover() 重建 head、重新分配端口后没有刷新它。
若 URL 从 A 变为 B,随后按数量缩容或 eviction 清理时 head 成功、follower 失败,按 B 重试会回退到缓存 A,仍报 “No engines selected for removal”。当前实际方法的隔离验证确认:普通同 URL 重试通过,恢复后旧缓存场景失败。
建议在恢复完成后刷新地址,或在所有移除路径清理 head 前保存当前地址,并补充 recovery → 部分清理 → 当前 URL 重试回归。F4 保持部分解决。
There was a problem hiding this comment.
@rai-studio-bot 目前通过 ray-native 扩容申请了多节点 replica 实现是不是有问题啊,具体分析一下,(relax/distributed/ray/rollout.py : _scale_out_ray_native),以每台 8 卡、每副本 16 卡为例)
1797 行算出 actors_per_replica = 2,并为每个副本创建含 16 个 GPU bundle 的 PG。
1888 行按 2 对齐预留 rank:head 拿到偶数 rank,follower 的 rank 也预留了。
2046 行创建 EngineGroup 时写死了 all_engines=[None]。
514 行按 slot 数创建 actor,只创建出 i=0 这一个,也只用到第一台机器的 bundle(523 行)。
这个 actor 的 server 参数是 nnodes=2、node_rank=0、tp_size=16(sglang_engine.py:1740),它会一直等待一个永远不会创建的 node_rank 1。
There was a problem hiding this comment.
是,这条链路有确定的 slot 数量错误,你指出的位置是关键。我直接检查了当前代码,并用实际 EngineGroup.start_engines() 和地址分配函数做了隔离验证。
按每节点 8 卡、每副本 16 卡、sglang_pp_size=1,假设预留的起始 rank 为 2:
| 环节 | 当前实际行为 |
|---|---|
| PG | 申请 16 个单 GPU bundle;资源申请本身不会创建 SGLang worker |
| rank | 预留 rank 2、3,但只是在分配编号 |
_bring_up_single_replica() |
all_engines=[None] 只有一个 worker slot |
start_engines() |
只执行 i=0,创建 rank 2,使用排序后的 bundle 0、base GPU 0 |
| server 参数 | nnodes=2、node_rank=2 % 2=0、tp_size=16 // 1=16 |
| 缺失部分 | 应有的 rank 3 / node_rank 1 / bundle 8 从未创建或调用 init |
根因是把“一个逻辑副本”误当成“一个 Ray actor”。all_engines 存每个节点的 worker,engines 才通过步长 nodes_per_engine 选出逻辑副本的 head。地址分配函数虽然会为 follower rank 写入同一个 dist_init_addr,但它不会创建 follower。现有 GPU 索引校验也只检查不越界,8 个使用槽位小于等于 16,因而拦不住缺 worker。
最小修复方向是在创建 group 时使用与 rank 预留相同的 actor 数:
actors_per_replica = max(1, gpus_per_engine // self.args.num_gpus_per_node)
new_group = EngineGroup(
...,
all_engines=[None] * actors_per_replica,
...,
)在理想的两节点各 8 bundle 布局下,隔离验证结果是:当前单 slot 只创建 (rank=2, bundle=0);两个 slot 则创建 (2, 0) 和 (3, 8),两个 init 使用同一个 rendezvous 地址。仍应让 finalizer 接收 new_group.engines(一个 head),不要把两个 worker 计成两个副本。
“永远等待”需稍作限定:head 缺少 peer,不能正常完成分布式初始化;但 _bring_up_single_replica() 外层 wait_for 有初始化超时,会进入 PROVISION_TIMEOUT 和 rollback,也可能先收到后端初始化异常。是否连 rollback 都卡住是另一项问题,不能仅凭缺 follower 就断言。
建议补一个不替换 EngineGroup 的 8×2-node / 16-GPU scale-out 测试,断言两个 actor、两个 init、不同 node_rank、相同 dist_init_addr,且只发布一个逻辑副本;再在真实两节点上验证初始化、权重同步与失败回滚。当前隔离验证只能证明创建/参数链,不能替代 GPU 验收。
另外,PR merge-base 2a8d2ed 的同一路径也已有 all_engines=[None](当时第 2018 行),所以这是当前代码仍存在的基线缺陷,不是本轮重构新引入。我此前对缩容重试的验证不能覆盖这条实际扩容创建路径。
There was a problem hiding this comment.
已在 2101a0f0 复核修复:recover() 在 init 全部成功后立即调用 refresh_replica_urls() 记录重建 head 的最新地址——后续清理部分失败时缓存持有的已是恢复后的 URL B;_resolve_scale_in_url_candidates 优先向存活 head 实时查询 get_url 并回写缓存,仅在 head 不可达时回退缓存;重建前先弹出旧 URL;迟到的旧 head 探测被 all_engines[...] is not head 校验挡住,无法覆盖新地址;URL 获取失败的回滚会丢弃旧地址而不是保留 A。
tests/inference/test_rollout_recovery_urls.py 以真实 RolloutManager 方法覆盖了本条原始场景(head 消失后按恢复后 URL 重试部分清理)及回滚、迟到探测、scaled-out 仅刷新重建 head 等变体。同时 all_engines=[None] * actors_per_replica 修复了本线程中另行确认的多节点扩容缺 follower 基线缺陷,tests/inference/test_multinode_scale_out.py 含拒绝单 slot 旧行为的回归测试。本地 120 项相关测试通过,该 head 8 项 CI 全绿。此项已解决。
da83371 to
7506d13
Compare
981a275 to
27e29cd
Compare
rai-studio-bot
left a comment
There was a problem hiding this comment.
已完成 a6e84d8 复核,并独立检查新增静态审查记录。新增两项 P2:终止 actor 的局部清理/恢复无法推进,以及文档误判 Serve 异步析构支持,详见行级评论。原 URL 清理重试问题保持部分解决。
文档中的 P1 结论未全部得到验证:deferred 消费有 commit gate,失败会终止运行;external/custom producer 保留兼容路径。普通请求超时验证仅发送一次 POST,不能据此认定重放。
本轮 client/gateway 测试 52 passed、1 skipped;依赖隔离验证复现 dead actor 的重复恢复失败。未执行真实 Ray actor 故障注入、多节点或 GPU 验收,当前版本暂无 CI checks。
a6e84d8 to
7813d61
Compare
rai-studio-bot
left a comment
There was a problem hiding this comment.
已完成 7813d61 跨模块复核。原 URL 重试问题及终止 actor 清理问题均有进展,但恢复后地址缓存、子进程释放确认仍需补齐;已在原线程补充证据,维持 P2。错误审查文档已删除,对应问题已解决。未发现其他新增问题。
本地定向测试 186 项通过,1 项因缺少 Ray 无法执行;另有管理器/lifecycle 隔离测试 44 项、deferred 子集 11 项通过,缩容隔离验证复现剩余边界。隔离验证不代表真实 Ray、完整 OPD 或多节点/GPU 验收,当前版本暂无 CI checks。
|
你这模型不行啊,出现幻觉了吧
链接打不开啊 以及报错分析能不能从实际报错开始分析?怎么突然跳到了依赖问题呢?麻烦输出点自己的理解哈,不要全依赖 agent 的判断,这样我们会怀疑对代码的理解程度 @rai-studio-bot 你来分析下报错呢?给个建议 |
# ♻️ Refactor ## Complete the Phase 6 ownership migration - Move Rollout engine-group ownership, memory lifecycle, discovery observation, recovery, and cleanup into the shared InferenceManager core - Keep generation, evaluation, data sources, and elastic workload behavior in the Rollout facade and workload layer - Preserve legacy engine-lock and weight-update interfaces --- # 🐛 Bug Fix ## Isolate failed elastic groups from production - Exclude non-ACTIVE scaled groups from memory operations, weight synchronization, discovery markers, and Rollout engine helpers - Keep DRAINING groups attached for cleanup and placement-group ownership retries - Let the shared manager reopen admission only after full memory restoration and backend continuation --- # ✅ Tests ## Verify Rollout ownership and topology filtering - Cover ACTIVE topology selection, legacy five-tuples, lifecycle rollback, elastic ownership, and workload contracts
# ♻️ Refactor ## Complete the Phase 7 migration - Connect remaining Rollout, Agentic, Teacher, Autoscaler, and OPD callers to unified discovery and lifecycle paths - Preserve Native and Agentic Rollout HTTP capacity from the configured per-engine concurrency and replica count - Remove the MultiEngineManager compatibility module and the GenRM-specific SGLangEngine subclass ## Remove legacy Python import paths - Require callers and fresh Ray workers to import InferenceManager and SGLangEngine directly - Preserve runtime role APIs and schema compatibility while dropping obsolete implementation aliases --- # 🐛 Bug Fix ## Retry partial logical replica cleanup - Identify logical replicas through any surviving worker after the head has stopped - Keep partial DRAINING replicas selectable by target count, URL, and eviction - Retry only residual actors while preserving confirmed cleanup results --- # ✅ Tests ## Verify final migration and cleanup boundaries - Cover managed Teacher clients, loop-local transport cache identity, shutdown ownership, old interface adapters, and placement rejection - Cover multi-node follower cleanup retries through scale-in and graceful eviction --- # ✅ Tests ## Keep CPU CI collectible without SGLang - Skip real Teacher and SGLang imports before collection enters the optional backend - Limit runtime skips to dependent tests and fixtures while retaining pure CPU coverage
189538c to
ec350f6
Compare
rai-studio-bot
left a comment
There was a problem hiding this comment.
ec350f6 本轮增量未发现新增问题:已核对权重版本查询、发布失败的状态更新及异常传播、路由异常的 Ray 序列化。/model_info 提供所需字段,与固定 SGLang 版本兼容(v0.5.17、v0.5.15)。
三个定向测试文件 85 passed,覆盖查询失败、未发布版本、版本 0、失败后重试和异常序列化;后端测试使用替身,未执行真实 SGLang/GPU 或多节点验收。当前提交尚无 CI 结果。
整体仍有 F4:恢复后 URL 重试 与 F6:actor 死亡后的子进程清理确认 部分解决,维持补充意见,不作通过结论。F7 保持已解决。
rai-studio-bot
left a comment
There was a problem hiding this comment.
复核 2785a46c:8 项精简建议全部落地,新增枚举重构无问题
3a563718(Add Role State Enum Class)与 2785a46c(Simplify code)已逐项复核:
- F10–F17 全部按建议实现并验证(死链/死模型/无引用符号删除、共享
genrm_ray_gpu_fractionhelper 并新增直接执行真实 manager 方法的新测试、max_samples透传与容量测试、phase字段删除、共享测试替身模块),已在各原始线程逐项回复确认。改动文件编译通过、全树无残留引用;本地可收集的 CPU 测试 187 + 14 项通过(失败/报错均为基线既有的本地缺依赖问题:torch/numpy/sglang_router;11 项新 placement 用例在 CI 依赖集下可执行)。 RoleState枚举重构为纯机械替换:状态集合不变,get_state导出与错误信息经.value保持纯字符串契约(新测试显式断言导出类型为str),lifecycle coordinator 测试全部通过。- F4/F6 相关代码(
inference_manager.py、distributed/ray/rollout.py)本轮未改动,维持部分解决;F9 待更新(描述提交范围与 13 个文件链接仍指向1054831/7813d61,描述快照示例仍含已删除的phase/phase_epoch字段)。 - 当前 head 推送后 CI 尚无结果(排队中),以上为本地验证结论。
当前待处理仅剩:F4/F6(P2,部分解决)与 F9(P3,描述更新)。审查看板已同步。
|
感谢贡献,我近期会开始review |
rai-studio-bot
left a comment
There was a problem hiding this comment.
复核 2101a0f0(fix bug):F4/F6 已解决,全部代码类问题清零,该 head CI 8/8 绿
- F4(恢复后 URL 重试)已解决:
recover()成功后立即refresh_replica_urls()记录重建地址;scale-in 选路优先实时查询存活 head 并回写缓存;重建前弹出旧 URL,迟到探测与回滚均不会保留旧地址。本线程中另行确认的多节点扩容缺 follower 基线缺陷也已修复(all_engines=[None] * actors_per_replica)并附拒绝旧行为的回归测试。 - F6(终止 actor 后的资源确认)已解决:新增
EngineShutdownGuard,终端 actor 死亡不再被计入 shutdown 确认,而是围栏全部生命周期入口并抛出InferenceRecoveryRequired;Controller 重启前逐一确认 rollout/teacher/genrm 后端清理结果,失败即阻止自动重启并给出面向操作者的说明(明确 Ray shutdown 与固定等待不构成资源释放确认);health monitor 复用组清理路径并登记 guard。测试含真实 Ray 子进程 crash 场景与真实 owned PG 的保留断言。 - 附带收益:新增 gateway→admission 双真实组件集成测试(补上此前审计记录的 CPU 级集成缺口);
dispose拒绝在途扩容请求。 - 验证:本地新增/更新测试 120 项通过(1 项依赖 psutil 的子进程用例在 CI 依赖集下通过);该 head 全部 8 项 GitHub CI 成功。
- F9(P3,描述)部分解决:过期提交范围行与已删除的
phase字段示例已清理,但 13 个 Key implementation files 链接仍指向7813d61(rebase 前且已落后三个 head,如genrm.py/inference_manager.py其后均有实质改动),建议改指当前 head2101a0f0。
至此代码层面的审查问题全部解决;唯一待处理项为 F9 的描述链接更新。审查看板已同步。
rai-studio-bot
left a comment
There was a problem hiding this comment.
审查结论:通过 —— 17 项审查发现全部解决或取代
首轮审查中的 F9(描述过期提交引用)已随本次描述更新全部落实:13 个 Key implementation files 链接已指向当前 head 2101a0f0,过期提交范围行与已删除的 phase 字段示例均已清理,代码量统计与当前 diff 完全一致(91 文件、+14067/−1743)。
至此本轮审查的全部 17 项发现(F1–F17,含早期 P1 并发回归、两项曾部分解决的 P2)均在 2101a0f0 上修复并逐项验证;F4 线程中另行确认的多节点扩容缺 follower 基线缺陷也已修复。生产与测试必要性审计完成且决策均已落地;该 head 全部 8 项 GitHub CI 通过。剩余边界仅作者显式声明 SKIP 的多节点真实硬件验收(单节点 4×B200 实测已在描述中记录),供维护者评审时权衡。新增的真实路径回归测试(真实 GenRM manager 的 preflight 一致性、真实 Ray 子进程 crash 场景、gateway→admission 双真实组件集成)显著提升了验证质量。
完整发现索引与各阶段结论见审查看板。
|
辛苦解决下冲突~ |
# 🔩 Chore ## Merge redai-studio/main into feat/task3 - Merge main at 815de3a - Preserve deferred inference commit gates alongside offline DPO and RM actor behavior - Regenerate OpenAPI specifications from the merged service routes --- # 🐛 Bug Fix ## Reject deferred scoring for offline workloads - Apply the offline workload guard consistently to SFT, DPO, and RM - Prevent offline training from waiting for a rollout commit coordinator --- # ✅ Tests ## Cover the merged deferred workload validation - Verify SFT, DPO, and RM are rejected by deferred scoring preflight - Run focused actor and deferred inference tests - Run all pre-commit hooks including gitleaks
|
Thanks for contributing to Relax, @ldemon2333! 感谢你为 Relax 做出贡献! 📚 Contribution guide / 贡献指南Describe the problem, your changes, and how you validated them. Keep each PR focused and run 请说明问题、改动和验证方式,保持 PR 聚焦,并在提交前运行 🛠️ CI commands / CI 指令
Put one command on the first line of a new PR comment. Rerun/cancel require PR authorship or repository write access. 在新 PR 评论的首行写一条指令。PR 作者或有仓库写权限的贡献者可以重跑、取消 CI。 🔎 Merge requirements / 合入条件GitHub:⏳ 合入条件未满足
|
|
Yangruipis
left a comment
There was a problem hiding this comment.
整体问题不大,在 架构统一、RFC 覆盖、运行正确性、抽象上做的都比较好,但是改动量和侵入性相对较大。
一些小问题 comment 了,后期如果通过 review,建议按模块拆分提交和验证,我们内部也会在多机大规模下进行nightly任务验证
| raise ValueError("Deferred Teacher requires SGLang OPD") | ||
| if "genrm" in roles: | ||
| models = getattr(args, "_genrm_instances_resolved", {}) | ||
| if len(models) != 1 or getattr(args, "rm_type", None) != "dapo-genrm" or getattr(args, "custom_rm_path", None): |
There was a problem hiding this comment.
新增 commit:b36a 新增 supports_deferred 相关契约,解耦 genrm role 与特定 reward adapter 计算逻辑
| serve.run(deployment, name="metrics", route_prefix="/metrics") | ||
| logger.info("MetricsService deployed at /metrics") | ||
|
|
||
| def _deploy_teacher_gateway(self) -> None: |
There was a problem hiding this comment.
当前 rollout 和 genrm 都还是各自的 deployment,只有 teacher 用的 InferenceGateway ,这个是短期形态吗,后面会统一嘛?
There was a problem hiding this comment.
目前这种设计是兼容旧接口的一种过渡形态。Teacher 直接使用通用 InferenceGateway,主要是为了在保留现有 OPD Teacher 启动方式、opd_teacher_url(s) 直连兼容路径以及 Controller/Actor 对 TeacherManager 生命周期控制的前提下,最小成本接入统一的 schema v2 discovery 和 routing contract。并且目前的 TeacherManager 已覆盖单模型的 engine 生命周期;多模型编排目前由 OPD helper、Controller 和 Gateway 分散承担,目前实现中并没有像 genrm serve deployment 一样,使用一个专门的 teacher serve deployment 来做多 teacher model 编排。
我觉得目前 genrm 与 teacher role 本身都是作为被动的 serve deployment,两者架构高度相似,可以进行统一
一个 inference role
│
├── 一个 Serve Component
│ │
│ └── managers: dict[model_id, Manager]
│ │
│ ├── model_id=A → Manager A → replicas A
│ ├── model_id=B → Manager B → replicas B
│ └── model_id=C → Manager C → replicas C
│
└── 一个 role 级 InferenceGateway
│
├── 聚合所有 Manager 的 schema v2 snapshot
├── route_key → model_id
├── 从目标模型选择 READY replica
└── 转发推理请求
GenRM 、Teacher 、 rollout 最终的结构:
Rollout role
├── Rollout Service
│ ├── 主动 rollout workload
│ ├── DataSource / TransferQueue
│ ├── weight sync / evaluation / scaling
│ └── 一个多模型 RolloutManager
└── InferenceGateway(role="rollout")
GenRM role
├── GenRM Service
│ ├── tokenizer / chat-template renderer
│ └── dict[model_id, GenRMManager]
│ └── 每个 Manager 管理一个模型的多个 replicas
└── InferenceGateway(role="genrm")
Teacher role
├── Teacher Service
│ └── dict[model_id, TeacherManager]
│ └── 每个 Manager 管理一个模型的多个 replicas
└── InferenceGateway(role="teacher")
| finally: | ||
| self._inference_clients.pop(session, None) | ||
| if client is not None: | ||
| await client.aclose() |
There was a problem hiding this comment.
[P2] 按样本重建 Teacher client 使 discovery 缓存失效
native Rollout 按样本调用 prefill,而这里每次创建并关闭 InferenceClient,因此 Teacher discovery 的 TTL 缓存无法跨样本复用。发现请求又进入串行锁和副本健康 RPC,会增加高并发 OPD 的控制面开销。建议按事件循环复用 client,在对应生命周期结束时关闭,并增加连续两次 prefill 的 discovery 请求次数测试。已确认调用路径,尚未实测吞吐下降幅度。
| raise ValueError( | ||
| "Use --inference-defer-roles with a framework scoring adapter; legacy hooks cannot own this plan" | ||
| ) | ||
| validate_deferred_workload(self.config) |
There was a problem hiding this comment.
[P2] 新 defer 校验使仓库现有示例无法启动
新的 preflight 要求 inference_defer_roles,但 examples/generate_reward_model/run-qwen35-35B-A3B-16xgpu-genrm-397B-defer.sh 仍使用旧 defer flag、rm-type dummy 和 custom post-process hook。该示例会被这里拒绝;只补新 flag 又会被 validate_deferred_workload 的 hook 和 reward adapter 限制拒绝。请同步迁移示例及 README 到框架评分适配器,并验证整套参数通过预检,保留现有安全限制。
# ♻️ Refactor ## Declare deferred reward capability - Add an explicit deferred capability to registered reward specifications - Resolve deferred GenRM scoring through the reward registry - Remove the inference scheduler dependency on the DAPO reward implementation --- # ✅ Tests ## Cover capability-based deferred routing - Accept async reward adapters that declare deferred support regardless of name - Reject async adapters without the deferred capability - Exercise deferred scoring through the registered handler
rai-studio-bot
left a comment
There was a problem hiding this comment.
复核结论:b36aac4b 增量审查完成,无新增问题 —— 有条件等待两项未回应意见
对上一已审 head 2101a0f0 之后的增量完成复核:合并 bb2aa11(同步 main 815de3aa)经 git diff-tree --cc 确认的全部 11 个冲突文件逐一核对,PR 侧的 defer 门控、discovery 与 placement 内容和 main 侧的 offline 模式、DPO/RM、设备模块化变更均完整保留,defer.py 离线拒绝切换为 is_offline_mode 后覆盖 sft/dpo/rm 并有对应参数化测试;b36aac4b 按审查意见把 deferred GenRM 评分从 dapo_genrm 硬编码导入改为 reward registry 的 supports_deferred 能力契约,preflight 会拒绝未声明该能力的 async 适配器(如 remote_rm),评分入口仅在 preflight 通过的 defer 模式可达。本地 tests/inference/test_defer_wiring.py 38 项、rewards/legacy 161 项通过,该 head 全部 CI 通过。本轮无新增发现,既往 17 项发现保持已解决或已取代。
非行级事项(待作者处理,非本审查新发现):维护者审查中的两个 P2 意见截至当前 head 仍无代码回应,我在当前代码上复核后仍然成立——按样本重建 Teacher client 使 discovery 缓存失效(on_policy_distillation.py 未改动),以及新 defer 校验拒绝仓库现有示例(示例仍使用旧 defer flag、--rm-type dummy 与 custom post-process hook,placement.py 仍将其映射为 defer 计划)。审查线程 1 的解耦修复已验证,线程 2 的部署形态问题作者已给出设计与后续统一方案。
完整发现索引、提交范围(94 个文件全部接收)与各阶段结论见审查看板。
# 🐛 Bug Fix ## Use framework-managed deferred scoring - Replace the legacy custom post-process lifecycle with the registered dapo-genrm adapter - Remove unsupported dynamic batching and in-run evaluation from the deferred recipe - Keep the shared 16-GPU Rollout and GenRM placement --- # 📝 Documentation ## Update GenRM defer guidance - Document the explicit deferred role and adapter configuration in English and Chinese - Describe the closed-batch commit flow and current evaluation limitation --- # ✅ Tests ## Cover the complete recipe preflight - Verify the shared placement resolves to defer mode on 16 GPUs - Verify the migrated arguments pass deferred workload validation
# ⚡ Performance ## Reuse event-loop discovery clients - Route managed Teacher requests through the existing loop-local inference client cache - Preserve model routing and group affinity while allowing discovery TTL reuse - Keep the legacy external Teacher transport unchanged --- # ✅ Tests ## Cover discovery reuse and cleanup - Verify concurrent prefill calls share one inference client and one discovery request - Verify both Teacher scoring requests still execute - Close cached clients explicitly at the event-loop lifecycle boundary
rai-studio-bot
left a comment
There was a problem hiding this comment.
结论:通过(APPROVE)
上轮遗留的两项 P2 已在 head 43e1a21b 修复并核实:
- Teacher client 按样本重建使 discovery 缓存失效:
43e1a21b改用generate_with_discovery的事件循环级缓存复用;route_key经 payload 传递、affinity 经X-SMG-Routing-Key头逐请求转发,请求语义不变;客户端在shutdown_async_loop关闭,并有并发 prefill 复用同一 discovery client 的测试覆盖。 - defer 预检拒绝现有示例:
4bf91848将 defer 示例与 README/中英文档迁移到框架评分适配器(--rm-type dapo-genrm+--inference-defer-roles genrm),移除了预检拒绝的 dummy rm、legacy hook、--eval-interval与--use-dynamic-batch-size。已在本地以脚本真实参数复跑参数解析、校验、plan_inference_placement(mode=defer,16 GPU)与validate_deferred_workload,整套参数通过预检;46 项相关测试通过。
既往 17 项发现保持已解决或已取代,本轮无新增发现,CI 在该 head 全绿。边界说明:defer 示例完成的是预检级(解析/校验/placement 计划)验证,多机真实执行仍属作者声明的 SKIP,由内部 nightly 覆盖;维护者线程的解除由维护者确认。
详细范围、阶段结论与发现索引见审查看板。



统一 Rollout、GenRM 与 Teacher 推理基础设施
Relates to #71
feat/task32d88481至2101a0f;Phase 1–7 主干及后续 3 个修复/代码简化提交,共 10 个提交origin/main/2a8d2edrelax/tests/docs/统计口径:以
2a8d2ed..2101a0f为范围,按 GitHub PR 文件差异统计物理增删行;包含 OpenAPI 文档,文件重命名计一个变更条目。What
本 PR 将 Rollout、GenRM 和 managed SGLang Teacher 收敛到共同的 discovery schema、模型路由、HTTP handler、Engine 生命周期核心和 GPU placement preflight,同时保留三个角色各自的业务流程与旧接口。
改造后,请求入口先读取 schema 2 的角色快照,再通过同一个
RoutingSpec和RouteResolver选择模型以及 Router 或逻辑副本。GenRM 和 Teacher 的静态模型使用统一SGLangEngine,明确拒绝 DCS、Actor 权重更新和动态 LoRA。Rollout 保留 Router、动态权重、弹性扩缩和生成业务,但把 Engine group ownership、显存生命周期和 discovery 状态委托给公共InferenceManager核心。Colocated 模式增加两阶段 placement 校验。配置层先生成全角色切片计划,Ray placement group ready 后再核验实际 node、GPU 和 bundle 映射。
split使用不重叠切片;受约束的defer允许 Rollout、GenRM、Teacher 在不同阶段复用同一切片。LifecycleCoordinator只在前一角色 offload 得到确定回执后授予下一角色 lease,并在评分字段写回、TransferQueue 发布成功后才提交训练批次。RFC Decision 对照
InferenceGatewayclass, one instance per roleInferenceGatewayHandler;managed Teacher 使用独立 CPU-onlyInferenceGatewaydeployment;Rollout/GenRM 仍在原角色 deployment 内组合 handlerInferenceManagerclass for all rolesInferenceManager;RolloutManager 组合InferenceManager.for_rollout();旧兼容 Manager 模块已删除SGLangEngine; remove the GenRM subclassSGLangEngine;GenRM 通过InferenceEngineSpec("genrm")注入 STATIC 约束,专用 Engine subclass 已删除RoutingSpec,Gateway 和 directInferenceClient共用RouteResolversplitanddefersplit规划和单节点绑定通过;受约束 defer 控制链已实现,已修复实测发现的空响应 blocker,但完整 GPU 训练闭环尚未重跑RolloutManager、GenRMManager、TeacherManager类名和 Rollout facade,MultiEngineManager兼容模块已删除改造前的问题
三个角色都调用 SGLang,但控制面和数据面彼此分散:
GenRMClient -> /genrm/generate -> GenRM Serve三个路径对“哪个模型”“哪个副本可用”“SLEEPING 是否可路由”“失败后何时刷新地址”的回答不同。资源检查也分散:各角色只看自己的 GPU 数,无法在任何 Engine 启动前验证 Rollout、GenRM、Teacher 和训练角色的组合布局。
当前架构
点击展开 Mermaid 图
flowchart TB subgraph RoleIngress[角色 HTTP 入口] R[Rollout Serve /rollout] G[GenRM Serve /genrm] T[CPU-only InferenceGateway /teacher] end R --> RH[InferenceGatewayHandler] G --> GH[InferenceGatewayHandler] T --> TH[InferenceGatewayHandler] RH --> Client[InferenceClient + RouteResolver] GH --> Client TH --> Client Client -.-> Snapshot[Schema 2 role snapshot] Client --> Router[Rollout SGLang Router] Client --> Direct[GenRM or Teacher logical head] Planner[Placement preflight and physical validation] -.-> Manager[InferenceManager core] Lifecycle[LifecycleCoordinator] -.-> Manager Manager --> Engine[SGLangEngine] Engine --> GPU[SGLang GPU workers]职责边界如下:
InferenceGatewayHandlerRoutingSpec/RouteResolverInferenceManagerLifecycleCoordinatorDeferredBatch从启动到训练消费的端到端关系如下。虚线表示控制或状态发布,实线表示请求、样本或资源流。
点击展开 Mermaid 图
flowchart TB subgraph Startup[启动和资源面] Config[Role configs and model specs] --> Plan[Placement plan] Plan --> PG[Ray placement groups] PG --> Bind[Physical node and GPU validation] Bind --> Managers[Rollout GenRM Teacher Managers] Managers --> Engines[SGLangEngine logical replicas] end subgraph Routing[请求和发现面] Managers -. publish .-> Registry[Schema 2 role snapshots] Registry -. refresh .-> Client[InferenceClient and RouteResolver] Gateway[Role Gateway handler] --> Client Direct[Framework direct callers] --> Client Client --> Router[Rollout Router] Client --> Heads[GenRM or Teacher heads] Router --> Engines Heads --> Engines end subgraph Batch[样本和训练面] Engines --> Generated[Generated groups] Generated --> Staging[Deferred CPU staging] Staging --> Scoring[GenRM Teacher and optional Student stages] Scoring --> Writeback[Validate and write back fields] Writeback --> TQ[TransferQueue train partition] TQ --> Training[Actor Critic ActorFwd Advantages] end Coordinator[LifecycleCoordinator] -. activation lease .-> Managers Coordinator -. commit gate .-> Training Scoring -. role released .-> Coordinator TQ -. put acknowledged .-> CoordinatorGateway 与请求数据流
下图把角色 HTTP 入口和框架内部 direct client 放在同一张图中。两类调用最终都读取 schema 2,并使用相同的
RouteResolver;区别只是请求是否经过角色 Gateway。点击展开 Mermaid 图
flowchart LR subgraph Callers[调用方] Native[Native or Agentic Rollout] Reward[Reward function and GenRMClient] OPD[OpdManager] HTTP[External HTTP caller] end subgraph RoleIngress[角色入口] RolloutAPI[Rollout Serve] GenRMAPI[GenRM Serve] TeacherAPI[Teacher InferenceGateway] end subgraph SharedRouting[共同发现和路由] Handler[InferenceGatewayHandler] Client[InferenceClient] Snapshot[Schema 2 snapshot] Resolver[RouteResolver] end Native --> Client Reward --> GenRMAPI OPD --> Client HTTP --> RolloutAPI HTTP --> GenRMAPI HTTP --> TeacherAPI RolloutAPI --> Handler GenRMAPI --> Handler TeacherAPI --> Handler Handler --> Client Client -. refresh .-> Snapshot Client --> Resolver Snapshot --> Resolver Resolver -->|SGLANG_ROUTER| Router[Rollout model Router] Resolver -->|DIRECT| Head[GenRM or Teacher logical head] Router --> RolloutEngine[Rollout SGLang workers] Head --> StaticEngine[Static SGLang workers]Rollout
Rollout 的“direct client”只绕过角色 Gateway,仍然访问模型 Router,不会随机直连 PD worker。Router、动态权重、scale-out/in、生成和评估业务仍由 Rollout 体系拥有。
GenRM
Raw
input_ids/text请求可绕过 messages adapter,直接保留 SGLang payload。最终 judge 文本如何转换为 reward 仍由 reward adapter 决定。Managed Teacher
External Teacher URL 仍保留原外管路径;只有 managed Teacher 进入统一 discovery、placement 和 defer 生命周期。
HTTP 和 SSE
Handler 移除 hop-by-hop headers,保留 request ID 和
X-SMG-Routing-Key。Raw/generate在选路后去掉逻辑model和route_key;Chat 将逻辑模型名改写为 backendserved_model_name。SSE 字节流不解析重组。正常结束、读取异常、客户端断开和首 chunk 前 ASGI 失败都会关闭 upstream response;首事件发出后不会切副本重放。关闭 HTTP 流不等于 SGLang GPU request 已确认 abort,这仍是严格 drain 的边界。
Discovery schema 2
Schema 2 是规范路由快照,不是旧 worker 诊断列表。核心结构为:
{ "schema_version": 2, "role": "teacher", "registry_epoch": "manager-incarnation", "topology_revision": 3, "routing": { "default_model": "math", "route_key_map": {"math-data": "math"}, "aliases": {"teacher-checkpoint": "math"}, "policy": "round_robin", "policy_revision": 0 }, "models": { "math": { "state": "READY", "route_mode": "DIRECT", "router_url": null, "weight_source": "STATIC", "weight_version": null, "served_model_name": "Qwen/Qwen3-0.6B", "engines": [ { "engine_id": "math/replica-0", "base_url": "http://head:port", "state": "READY", "direct_eligible": true, "generation": 1 } ] } } }状态不是简单的进程存活。一个副本只有 weights、KV cache、CUDA graph 全部 resident,权重有效且 admission 打开时才是 READY。worker 或 endpoint 换代会增加 generation;Manager 状态变化会推进 topology revision。
多节点 logical replica 的所有 follower 都参与健康验证,但只发布 head URL。external Rollout engine(
--rollout-external或 external scale-out group)和未实现health_process的自定义 Rollout engine 没有可探测的本地进程,只按 head 的health_generate和 endpoint 判定;本地 engine 报告进程状态未知时仍会被 fence。PD prefill/decode workers 不进入公共 replica 列表,PD 请求只能走 Router。旧 Rollout
/engines?schema=1仍供 Autoscaler 和诊断兼容使用,只保留 regular head 的 rank、status、URL,过滤 PID、node metadata 和 P/D worker。Schema 1 不参与统一选路。下面是 Gateway 和 direct client 共用的实际选路决策。
点击展开 Mermaid 图
flowchart TD S[Read and validate schema 2 snapshot] --> M{Explicit model present} M -->|yes| MA[Match model ID or alias] M -->|no| K{Route key present} K -->|yes| KM[Resolve route_key_map] K -->|no| D[Use default_model] MA --> Found{Model resolved} KM --> Found D --> Found Found -->|no| Bad[HTTP 400] Found -->|yes| Ready{Model state is READY} Ready -->|no| Sleep[HTTP 503 and refresh boundary] Ready -->|yes| Mode{route_mode} Mode -->|SGLANG_ROUTER| RU[Use model router_url] Mode -->|DIRECT| Candidates[Filter READY and direct_eligible heads] Candidates --> Any{Candidates available} Any -->|no| Unavailable[HTTP 503] Any -->|yes| Affinity{Affinity key present} Affinity -->|yes| Hash[Rendezvous hash by stable engine_id] Affinity -->|no| RR[Process-local round robin] Hash --> Target[RouteTarget] RR --> Target RU --> TargetInferenceManager 与 SGLangEngine
公共 Manager 核心
三角色共享的是同一个资源和生命周期核心,并不意味着三个角色已经具有完全相同的 Ray Actor class。GenRM/Teacher 使用继承,Rollout 使用组合。
点击展开 Mermaid 图
flowchart TB Core[InferenceManager common core] GenRM[GenRMManager Ray actor] -->|inherits| Core Teacher[TeacherManager Ray actor] -->|inherits| Core Rollout[RolloutManager Ray actor] -->|composition| RolloutCore[InferenceManager.for_rollout] RolloutCore --> Core Core --> Observation[InferenceObservation and schema 2 registry] Core --> Slots[Engine slots and logical replicas] Core --> Memory[onload offload and resident tags] Core --> Recovery[retire recover and cleanup] Core --> Ownership[owned or borrowed placement groups] Slots --> Head[Logical head] Slots --> Followers[Internal TP or PP followers] Head --> Engine[SGLangEngine] Followers --> Engine Observation --> Public[Publish head only]InferenceManager统一管理:init.remote();静态角色直接使用公共基类路径。Rollout 因保留大量业务方法,采用 composition:
Engine 能力隔离
三角色实际创建
SGLangEngine。InferenceEngineSpec区分:GenRM/Teacher 的 STATIC snapshot 不发布 Actor weight version,并拒绝 DCS、tensor weight update、seed sync 和动态 LoRA。
GenRM、Teacher 和 Rollout 现在都直接实例化
SGLangEngine。角色差异由构造时传入的InferenceEngineSpec和 server overrides 表达,不再通过 GenRM 专用 Engine subclass 表达。生命周期失败边界
Manager 对每个 logical replica 单独记录成功与失败。Offload 部分成功时,成功项进入 SLEEPING,未确认项进入 FAILED;FAILED 不代表显存已释放,不能把共享 GPU lease 交给下一角色。
Shutdown 必须先等待所有 worker
shutdown.remote(),再 kill actor,最后删除 owned PG。借用 Controller/Service PG 的 Manager 不删除 PG。清理任一步未确认时保留 pending 记录供后续重试。Engine actor 已确认死亡(
RayActorError,不含ActorUnavailableError)不能证明 SGLang 子进程已退出。此时保留 slot、pending 与 PG 的资源 fence,停止重复 shutdown RPC,并抛出InferenceRecoveryRequired阻止自动重建。Controller 在删除 owner/Serve 和重新初始化前严格确认各角色清理;未确认时终止自动恢复,要求独立确认节点/container 清理后启动新作业。ActorUnavailableError或普通超时仍按未确认状态处理,可重试确认式清理。health_check()覆盖每个 logical replica 的全部 worker:head 执行health_generate,每个 worker 报告本地进程状态。发现故障 worker 时尝试整组清理;只有 backend shutdown 已确认才允许释放并重建,terminal actor death 保留未确认 worker 的资源 fence。GenRM/health(schema 1)会调用该检查,因此健康探测可能带来退役副作用。点击展开 Mermaid 图
stateDiagram-v2 [*] --> STARTING STARTING --> READY: full tags and valid weights and admission open READY --> DRAINING: close admission and start offload DRAINING --> SLEEPING: backend release confirmed SLEEPING --> ONLOADING: activation lease granted ONLOADING --> READY: missing tags restored and admission reopened STARTING --> FAILED: startup or health failure DRAINING --> FAILED: release not confirmed ONLOADING --> FAILED: resume not confirmed READY --> STOPPING: shutdown SLEEPING --> STOPPING: shutdown FAILED --> STOPPING: controlled cleanupFAILED只表示该副本不可路由,不表示 GPU 已释放。共享资源只有在 Manager 返回确定的 SLEEPING/offload 结果后才能交给下一角色。Placement:先规划,再绑定,再启动
当前没有名为
PlacementPlanner的 class。设计由函数式接口实现,见relax/inference/placement.py:plan_inference_placement(args):无副作用逻辑 preflight;InferencePlacementPlan:不可变全局计划;validate_bound_placement():PG ready 后物理拓扑校验;model_placement():Manager 查询自身模型切片。点击展开 Mermaid 图
flowchart LR Args[All role resources and parsed configs] --> Plan[plan_inference_placement] Plan --> Logical{Logical layout valid} Logical -->|no| Reject1[Reject before Teacher PG Router or Engine] Logical -->|yes| Saved[Serializable InferencePlacementPlan] Saved --> PG[Create or borrow Ray placement group] PG --> Probe[InfoActor probes node and GPU identity] Probe --> Bind[validate_bound_placement] Bind --> Physical{Physical topology valid} Physical -->|no owned PG| Cleanup[Remove newly owned PG] Physical -->|no borrowed PG| Stop[Stop deployment and preserve external owner] Physical -->|yes| Launch[Launch Managers and complete logical replicas]逻辑 preflight
Planner 汇总 Actor、Critic、Reference、ActorFwd、Rollout、GenRM、Teacher,检查:
TP x PP = GPUs per engine;这一步发生在 Teacher Manager、data source、PG、Router 和 Engine 创建前。失败不会留下 GPU actor。
三种模式
decoupledsplitdefer点击展开 Mermaid 图
flowchart LR Mode{Placement mode} Mode --> Decoupled[Decoupled] Mode --> Split[Split] Mode --> Defer[Defer] Decoupled --> RP[Rollout pool] Decoupled --> GP[GenRM pool] Decoupled --> TP[Teacher pool] Split --> AP[One Actor pool] AP --> RS[Rollout slice] AP --> GS[GenRM slice] AP --> TS[Teacher slice] Defer --> Shared[Same shared bundle set] Shared --> P1[Generation phase: Rollout] P1 --> P2[Score phase: GenRM] P2 --> P3[Score phase: Teacher] P3 --> P4[Optional phase: Student]例如 8-GPU Actor pool 的 split:
8-GPU defer:
物理绑定
PG ready 后,临时 InfoActor 探测每个 bundle 的
bundle_index、node_id、node_ip、gpu_id。绑定验证拒绝:验证通过后 Manager 才能启动 Engine。物理校验失败只删除当前 owner 拥有的 PG,借用 PG 不会被错误删除。
Colocation 与 deferred scoring
Split
Split 的布局和单节点真实 Ray bundle 绑定已经验证。它允许推理角色使用同一 Actor PG 的不同切片;进入 Actor 训练前仍需现有 barrier/offload 机制确保推理显存已释放。
Defer
当前 defer 是受约束的同步 colocate closed-batch 实现。固定阶段为:
生成侧原本准备写入 TQ 的完整 groups 会被
DeferredTransferCollector截获到 CPU staging。DeferredBatch对 Sample 做副本,评分先写 staged fields;只有所有阶段成功并通过校验,才把 reward、Teacher logprobs、top-k 和多模态字段写回原 Sample。LifecycleCoordinator是 CPU-only Ray Actor,维护:Coordinator 内部状态
构造函数维护以下控制面状态:
_operation_timeout_conditionasyncio.Condition_managers_rolesrollout: SLEEPING_batches_Batch;begin_batch时驱逐已提交批次,不保留完整历史_current_failure_operations_committed_retired_operations只有当前批次保存完整状态。
_committed让迟到的wait_committed,以及重复的begin_batch、commit_batch,在批次被驱逐后仍看到 COMMITTED;_retired_operations在窗口内继续拒绝跨批次复用 operation ID。等待早于 1024 个提交窗口且已被驱逐的批次时,wait_committed只能等到调用方超时。每个
_Batch只保存四个控制字段:Coordinator 不存样本、reward、Teacher logits 或 TQ partition。样本和评分结果属于
DeferredBatch;Coordinator 只判断批次能否继续、共享 GPU lease 是否已经安全释放,以及训练消费者能否开始。Batch 与角色 lease 两套状态机
Coordinator 同时维护 batch 状态和角色 lease 状态。Batch 状态回答“这批数据能否训练”,lease 状态回答“当前哪个角色仍可能占用共享 GPU”,两者不能混用。
Batch 状态机:
点击展开 Mermaid 图
stateDiagram-v2 [*] --> PENDING: begin_batch PENDING --> COMMITTED: commit_batch and all roles sleeping PENDING --> FAILED: fail_batch or lifecycle RPC failure COMMITTED --> [*] FAILED --> [*]Coordinator 的 batch state 只有
PENDING、COMMITTED和FAILED。GENERATED、SCORING、COMPLETE、PUBLISHING等更细的数据处理阶段属于DeferredBatch。角色 lease 状态机:
点击展开 Mermaid 图
stateDiagram-v2 [*] --> SLEEPING SLEEPING --> ACTIVE: activate or begin_batch for rollout ACTIVE --> DRAINING: deactivate begins DRAINING --> SLEEPING: all manager offloads succeed SLEEPING --> ONLOADING: activate begins ONLOADING --> ACTIVE: all manager onloads succeed ONLOADING --> BLOCKED: failure or timeout DRAINING --> BLOCKED: failure or timeoutBLOCKED是安全状态,而不是普通失败标签。它表示远端 onload/offload 的结果不确定,宁可停止当前训练,也不能把同一块 GPU 分配给下一个角色,BLOCKED作为共享GPU 的 fail-stop 安全栅栏。Coordinator 不会因为调用超时或远端 RPC 稍后完成,就自动将角色改为SLEEPING并把同一组 GPU 交给下一个角色。Manager RPC 超时后 lease 保持 BLOCKED,即使远端稍后完成也不自动放行。Rollout 为避免 Actor self-RPC deadlock,执行本地 offload 后向 Coordinator 发送 acknowledgment;Teacher/GenRM 由 Coordinator 直接调用远程 Manager。
TQ
async_put(train_N, is_last=true)成功后,先执行本地 finalizer(debug 保存和指标),再调用commit_batch(N);finalizer 失败时先清理train_N分区再报错,不会 commit。Actor、Critic、ActorFwd 和 Advantages 在消费前调用wait_committed(N)。评分或发布失败会唤醒 waiters 并抛出 batch error,不产生训练许可。生成阶段在
begin_batch之后失败时,Rollout 在本地执行 offload;只有确认释放后才发送 acknowledgment,否则保留 lease,并在 batch 错误中标注 release unconfirmed。Student restore 失败同样会在finally中释放 Rollout lease。当前 defer 明确拒绝 partial rollout、dynamic batch、自定义 reward/convert/filter、SFT、非 Megatron backend、带评估数据的 defer、多 GenRM 模型和多组 Rollout/PD 等尚无阶段 adapter 的组合。
完整数据流如下。训练消费者可以先等待 batch ID,但只有评分、字段回写和 TQ 写入全部完成后才会收到 commit。
点击展开 Mermaid 图
sequenceDiagram participant W as RolloutWorkload participant C as LifecycleCoordinator participant B as DeferredBatch participant R as RolloutManager local lifecycle participant G as GenRM Manager participant T as Teacher Manager participant Q as TransferQueue participant A as Training consumers A->>C: wait_committed(batch N) W->>C: begin_batch(N) W->>W: generate and capture groups in CPU staging W->>B: complete(groups) B->>B: clone samples and prepare Teacher inputs B->>R: local Rollout offload R-->>B: offload confirmed B->>C: acknowledge_rollout_offloaded(N) opt GenRM deferred B->>C: activate(genrm, operation ID) C->>G: onload.remote() G-->>C: confirmed B->>B: compute reward and write staged samples B->>C: deactivate(genrm, operation ID) C->>G: offload.remote() G-->>C: confirmed end opt Teacher deferred B->>C: activate(teacher, operation ID) C->>T: onload.remote() T-->>C: confirmed B->>B: Teacher prefill and field writeback B->>C: deactivate(teacher, operation ID) C->>T: offload.remote() T-->>C: confirmed end opt Student scoring required B->>C: begin_student_scoring(N) B->>R: local Rollout onload and version check B->>B: student prefill B->>R: local Rollout offload R-->>B: confirmed B->>C: acknowledge_rollout_offloaded(N) end B->>B: validate and copy fields to original samples B->>Q: async_put(train_N, is_last=true) Q-->>B: write returned successfully B->>C: commit_batch(N) C-->>A: wake all waiters A->>Q: consume committed training partitionShared co-resident
第一版拒绝同卡同阶段常驻。
_genrm_colocate_with_rollout如果没有显式 defer 会在 preflight 报错。减小mem_fraction_static不能绕过,因为 Ray reservation、峰值显存和生命周期冲突仍然存在。HTTP API 变化
新增公开角色路由
/rollout/generate/rollout/chat/completions/rollout/health/genrm/chat/completions/genrm/v1/chat/completions/genrm/v1/models/genrm/engines/teacher/generate/teacher/chat/completions/teacher/v1/chat/completions/teacher/v1/models/teacher/engines/teacher/health合计新增 13 个角色级公开路由。
已有路由的行为变化
/rollout/v1/chat/completions/rollout/v1/models/rollout/engines/genrm/generate/genrm/health?schema=2返回统一 health没有新增公开
/activate、/offload或/commit_batch。这些是内部 Ray RPC。Engine admission 内部入口
严格 defer Engine head 前增加 CPU admission server,允许的路径包括:
/generate、/v1/chat/completions、/chat/completions、/v1/completions;/abort_request、/close_session;/v1/models、/server_info、/get_server_info、/model_info、/get_model_info;/health、/health_generate。Weight update、pause、onload/offload 不在公共 admission allowlist。Gateway 和 Router 访问 guard URL,底层 SGLang 端口不作为框架 direct client 的旁路。
Key implementation files
relax/inference/gateway.pyrelax/inference/client.pyrelax/inference/routing.pyrelax/inference/specs.pyrelax/inference/compat.pyrelax/components/inference_gateway.pyrelax/distributed/ray/inference_manager.pyrelax/backends/sglang/sglang_engine.pyrelax/inference/placement.pyrelax/distributed/ray/inference_lifecycle.pyrelax/inference/defer.pyrelax/engine/rollout/deferred_scoring.pyrelax/distributed/ray/rollout.py兼容性
本 PR 保留以下兼容面:
get_rollout_engines_and_lock()五元组;{response: string}返回形状;GenRMClient、旧 health/metrics;get_urls()和旧 args URL 适配;RolloutManager、GenRMManager、TeacherManager命名;SGLangEngine的显式 role spec 构造;/engines?schema=1;兼容不包括危险布局。旧配置如果表达同卡同阶段共同驻留、无法确定 owner 或不满足完整 replica,会提前报错。
旧兼容 Manager 模块及其导入路径已经删除。GenRM 和 Teacher 直接继承
InferenceManager。Failure semantics
InferenceRecoveryRequired,禁止自动重建,需确认节点/container 清理后启动新作业health_check()失败并尝试整组清理;成功 worker 可清理,未确认的 terminal worker 保持 fence,不能释放共享 PG 或重建begin_batchfinally中 offload 并 ack Rollout leasetrain_N分区,不 commit点击展开 Mermaid 图
flowchart TD Failure[Lifecycle or scoring failure] --> Fence[Close admission and remove route eligibility] Fence --> Known{GPU release confirmed} Known -->|yes| Sleeping[Role becomes SLEEPING] Known -->|no or timeout| Blocked[Role remains FAILED or BLOCKED] Blocked --> StopNext[Do not activate conflicting role] Blocked --> FailBatch[Fail current batch and wake waiters] Sleeping --> Data{Scoring and TQ publish complete} Data -->|yes| Commit[Commit training batch] Data -->|no| FailBatchHTTP POST 发出后,公共 client 不自动重放,因为响应丢失时 backend 可能已经执行请求。旧 GenRMClient 仍有自己的有限重试策略,因此系统没有端到端 exactly-once 保证。
Validation
CPU 回归和单节点 4×B200 验证均通过,本次 未执行多节点测试
Acceptance checks
状态定义:
PASS表示当前代码和对应回归已满足;PARTIAL表示实现与 CPU/单节点证据成立,但最终 GPU 闭环仍缺失;SKIP表示当前环境不具备所需硬件,未把静态分析或单节点结果冒充为通过。InferenceGatewayHandler;Rollout/GenRM 组合 handler,Teacher 使用独立 CPU-onlyInferenceGatewaydeployment。test_gateway.py与test_teacher_gateway_wiring.py覆盖三角色协议和部署资源。InferenceManager,Rollout 通过InferenceManager.for_rollout()组合相同 ownership core;旧MultiEngineManager已删除。test_inference_manager.py、test_rollout_core.py和 Manager 继承断言通过。InferenceClient和RouteResolver;test_gateway_and_direct_resolver_choose_same_affinity_target、gateway transport 与 resolver 测试通过。plan_inference_placement()和validate_deferred_workload()在 Manager、Router、Engine 创建前拒绝切片重叠、资源超限和无阶段 adapter 的 defer 组合;placement/defer preflight 测试通过。test_deferred_ppo_rejects_persistent_ray_actor_reservation_conflict等测试通过。STATICEngine spec,拒绝 DCS、tensor/distributed update、LoRA 和 seed sync;能力测试通过,GenRM 的 DCS/tensor update 拒绝也有单节点真实 SGLang 证据。DeferredBatch先校验并写回 Teacher sampled/top-k 字段,再async_put(train_N),最后commit_batch(N);CPU workflow 测试通过。尚未运行真实 Teacher → TQ → Actor 训练闭环。DRAINING,PG 仍由InferenceManager.remove_group()按 ownership 释放。单节点实测也确认 Teacher-owned PG 被移除、GenRM borrowed PG 保留。Type of change