Skip to content

【No.3】unify role inference infrastructure - #347

Open
ldemon2333 wants to merge 14 commits into
redai-studio:mainfrom
ldemon2333:feat/task3
Open

ldemon2333 wants to merge 14 commits into
redai-studio:mainfrom
ldemon2333:feat/task3

Conversation

@ldemon2333

@ldemon2333 ldemon2333 commented Sep 21, 2026 •

Copy link
Copy Markdown

统一 Rollout、GenRM 与 Teacher 推理基础设施

Relates to #71

项目 内容
分支 feat/task3
代码范围 2d88481 至 2101a0f;Phase 1–7 主干及后续 3 个修复/代码简化提交,共 10 个提交
比较基线 origin/main / 2a8d2ed
测试结果 架构主干已经建立;个人资源有限,只做了单 node 多卡测验,多 node 测验未做
类别 变更文件 新增 删除 净增 改动量 新增占比 改动量占比
核心代码 relax/ 45 5,775 1,120 +4,655 6,895 41.05% 43.61%
pytest tests/ 45 8,285 469 +7,816 8,754 58.90% 55.37%
文档 docs/ 1 7 154 -147 161 0.05% 1.02%
合计 91 14,067 1,743 +12,324 15,810 100.00% 100.00%

统计口径:以 2a8d2ed..2101a0f 为范围,按 GitHub PR 文件差异统计物理增删行;包含 OpenAPI 文档,文件重命名计一个变更条目。

What

本 PR 将 Rollout、GenRM 和 managed SGLang Teacher 收敛到共同的 discovery schema、模型路由、HTTP handler、Engine 生命周期核心和 GPU placement preflight,同时保留三个角色各自的业务流程与旧接口。

改造后,请求入口先读取 schema 2 的角色快照,再通过同一个 RoutingSpec 和 RouteResolver 选择模型以及 Router 或逻辑副本。GenRM 和 Teacher 的静态模型使用统一 SGLangEngine,明确拒绝 DCS、Actor 权重更新和动态 LoRA。Rollout 保留 Router、动态权重、弹性扩缩和生成业务,但把 Engine group ownership、显存生命周期和 discovery 状态委托给公共 InferenceManager 核心。

Colocated 模式增加两阶段 placement 校验。配置层先生成全角色切片计划,Ray placement group ready 后再核验实际 node、GPU 和 bundle 映射。split 使用不重叠切片;受约束的 defer 允许 Rollout、GenRM、Teacher 在不同阶段复用同一切片。LifecycleCoordinator 只在前一角色 offload 得到确定回执后授予下一角色 lease,并在评分字段写回、TransferQueue 发布成功后才提交训练批次。

RFC Decision 对照

Topic RFC Proposal 当前实现
API One CPU-only InferenceGateway class, one instance per role 三角色共用 InferenceGatewayHandler;managed Teacher 使用独立 CPU-only InferenceGateway deployment;Rollout/GenRM 仍在原角色 deployment 内组合 handler
Engine lifecycle One InferenceManager class for all roles GenRM/Teacher 直接继承 InferenceManager;RolloutManager 组合 InferenceManager.for_rollout();旧兼容 Manager 模块已删除
SGLang One SGLangEngine; remove the GenRM subclass 三角色均实例化 SGLangEngine;GenRM 通过 InferenceEngineSpec("genrm") 注入 STATIC 约束,专用 Engine subclass 已删除
Routing One model-routing spec shared by gateway and direct clients schema 2 内统一 RoutingSpec,Gateway 和 direct InferenceClient 共用 RouteResolver
Placement Resolve and validate all GPU bundles before engine startup 配置 preflight 加 PG ready 后物理绑定验证均已接线;多节点真实硬件仍未验证
Colocation Support split and defer split 规划和单节点绑定通过;受约束 defer 控制链已实现,已修复实测发现的空响应 blocker,但完整 GPU 训练闭环尚未重跑
Shared co-resident Reject in the first version 同卡同阶段 Rollout/GenRM 配置若没有显式 defer 会在 preflight 拒绝
Migration Start with discovery and client; keep current managers P1 先引入 schema/client,随后迁移 Gateway、Manager、Placement、defer 和 Rollout ownership;保留 RolloutManager、GenRMManager、TeacherManager 类名和 Rollout facade,MultiEngineManager 兼容模块已删除

改造前的问题

三个角色都调用 SGLang,但控制面和数据面彼此分散:

角色 原请求路径 原生命周期/资源路径 主要问题
Rollout 生成函数直接使用 Router 地址,Chat 另有代理逻辑 RolloutManager、RolloutServer、EngineGroup 自己管理 discovery、业务和资源 ownership 混在大型 Manager 中
GenRM GenRMClient -> /genrm/generate -> GenRM Serve 旧 GenRM Manager 与专用 Engine 链路 专用 Engine、专用 cache、offset 与恢复逻辑
Teacher OPD 从 args 读取启动时写入的 URL TeacherManager 另行创建/借用 PG 地址可能过期,Teacher 结果与训练提交没有统一门槛

三个路径对“哪个模型”“哪个副本可用”“SLEEPING 是否可路由”“失败后何时刷新地址”的回答不同。资源检查也分散:各角色只看自己的 GPU 数,无法在任何 Engine 启动前验证 Rollout、GenRM、Teacher 和训练角色的组合布局。

当前架构

点击展开 Mermaid 图
flowchart TB
    subgraph RoleIngress[角色 HTTP 入口]
        R[Rollout Serve /rollout]
        G[GenRM Serve /genrm]
        T[CPU-only InferenceGateway /teacher]
    end

    R --> RH[InferenceGatewayHandler]
    G --> GH[InferenceGatewayHandler]
    T --> TH[InferenceGatewayHandler]

    RH --> Client[InferenceClient + RouteResolver]
    GH --> Client
    TH --> Client

    Client -.-> Snapshot[Schema 2 role snapshot]
    Client --> Router[Rollout SGLang Router]
    Client --> Direct[GenRM or Teacher logical head]

    Planner[Placement preflight and physical validation] -.-> Manager[InferenceManager core]
    Lifecycle[LifecycleCoordinator] -.-> Manager
    Manager --> Engine[SGLangEngine]
    Engine --> GPU[SGLang GPU workers]
Loading

职责边界如下:

组件 负责 不负责
InferenceGatewayHandler 请求校验、模型选择、HTTP/SSE 转发、错误映射 GPU 分配、自动唤醒、reward 解析
RoutingSpec/RouteResolver 模型、Router/direct replica 和 affinity 规则 Engine 生命周期
InferenceManager Engine slot、logical replica、onload/offload、recovery、PG ownership、snapshot 训练业务和跨角色调度
Placement planner 全局逻辑切片和实际 bundle 绑定校验 动态按流量重排、显存容量证明
LifecycleCoordinator 共享 GPU lease、operation 去重、batch commit gate 样本内容、TQ 存储、自动 reconcile
DeferredBatch staging、评分顺序、结果校验和字段回写 PG、Engine 和 discovery

从启动到训练消费的端到端关系如下。虚线表示控制或状态发布,实线表示请求、样本或资源流。

点击展开 Mermaid 图
flowchart TB
    subgraph Startup[启动和资源面]
        Config[Role configs and model specs] --> Plan[Placement plan]
        Plan --> PG[Ray placement groups]
        PG --> Bind[Physical node and GPU validation]
        Bind --> Managers[Rollout GenRM Teacher Managers]
        Managers --> Engines[SGLangEngine logical replicas]
    end

    subgraph Routing[请求和发现面]
        Managers -. publish .-> Registry[Schema 2 role snapshots]
        Registry -. refresh .-> Client[InferenceClient and RouteResolver]
        Gateway[Role Gateway handler] --> Client
        Direct[Framework direct callers] --> Client
        Client --> Router[Rollout Router]
        Client --> Heads[GenRM or Teacher heads]
        Router --> Engines
        Heads --> Engines
    end

    subgraph Batch[样本和训练面]
        Engines --> Generated[Generated groups]
        Generated --> Staging[Deferred CPU staging]
        Staging --> Scoring[GenRM Teacher and optional Student stages]
        Scoring --> Writeback[Validate and write back fields]
        Writeback --> TQ[TransferQueue train partition]
        TQ --> Training[Actor Critic ActorFwd Advantages]
    end

    Coordinator[LifecycleCoordinator] -. activation lease .-> Managers
    Coordinator -. commit gate .-> Training
    Scoring -. role released .-> Coordinator
    TQ -. put acknowledged .-> Coordinator
Loading

Gateway 与请求数据流

下图把角色 HTTP 入口和框架内部 direct client 放在同一张图中。两类调用最终都读取 schema 2,并使用相同的 RouteResolver;区别只是请求是否经过角色 Gateway。

点击展开 Mermaid 图
flowchart LR
    subgraph Callers[调用方]
        Native[Native or Agentic Rollout]
        Reward[Reward function and GenRMClient]
        OPD[OpdManager]
        HTTP[External HTTP caller]
    end

    subgraph RoleIngress[角色入口]
        RolloutAPI[Rollout Serve]
        GenRMAPI[GenRM Serve]
        TeacherAPI[Teacher InferenceGateway]
    end

    subgraph SharedRouting[共同发现和路由]
        Handler[InferenceGatewayHandler]
        Client[InferenceClient]
        Snapshot[Schema 2 snapshot]
        Resolver[RouteResolver]
    end

    Native --> Client
    Reward --> GenRMAPI
    OPD --> Client
    HTTP --> RolloutAPI
    HTTP --> GenRMAPI
    HTTP --> TeacherAPI

    RolloutAPI --> Handler
    GenRMAPI --> Handler
    TeacherAPI --> Handler
    Handler --> Client
    Client -. refresh .-> Snapshot
    Client --> Resolver
    Snapshot --> Resolver

    Resolver -->|SGLANG_ROUTER| Router[Rollout model Router]
    Resolver -->|DIRECT| Head[GenRM or Teacher logical head]
    Router --> RolloutEngine[Rollout SGLang workers]
    Head --> StaticEngine[Static SGLang workers]
Loading

Rollout

Native or Agentic Rollout
  -> generate_with_discovery()
  -> GET /rollout/engines?schema=2
  -> RouteResolver selects logical model
  -> model route_mode=SGLANG_ROUTER
  -> POST Router /generate

Rollout 的“direct client”只绕过角色 Gateway,仍然访问模型 Router,不会随机直连 PD worker。Router、动态权重、scale-out/in、生成和评估业务仍由 Rollout 体系拥有。

GenRM

Reward function
  -> legacy GenRMClient
  -> POST /genrm/generate with messages and route_key
  -> InferenceGatewayHandler
  -> select model through RoutingSpec
  -> render with selected model tokenizer/template
  -> refresh snapshot and recheck readiness
  -> POST selected DIRECT head /generate
  -> return {response: text}

Raw input_ids/text 请求可绕过 messages adapter,直接保留 SGLang payload。最终 judge 文本如何转换为 reward 仍由 reward adapter 决定。

Managed Teacher

OpdManager
  -> InferenceClient using /teacher discovery
  -> route_key selects Teacher model
  -> group affinity selects stable logical replica
  -> POST direct logical head /generate
  -> write sampled or top-k OPD fields

External Teacher URL 仍保留原外管路径;只有 managed Teacher 进入统一 discovery、placement 和 defer 生命周期。

HTTP 和 SSE

Handler 移除 hop-by-hop headers,保留 request ID 和 X-SMG-Routing-Key。Raw /generate 在选路后去掉逻辑 model 和 route_key;Chat 将逻辑模型名改写为 backend served_model_name。

SSE 字节流不解析重组。正常结束、读取异常、客户端断开和首 chunk 前 ASGI 失败都会关闭 upstream response;首事件发出后不会切副本重放。关闭 HTTP 流不等于 SGLang GPU request 已确认 abort,这仍是严格 drain 的边界。

Discovery schema 2

Schema 2 是规范路由快照,不是旧 worker 诊断列表。核心结构为:

{
  "schema_version": 2,
  "role": "teacher",
  "registry_epoch": "manager-incarnation",
  "topology_revision": 3,
  "routing": {
    "default_model": "math",
    "route_key_map": {"math-data": "math"},
    "aliases": {"teacher-checkpoint": "math"},
    "policy": "round_robin",
    "policy_revision": 0
  },
  "models": {
    "math": {
      "state": "READY",
      "route_mode": "DIRECT",
      "router_url": null,
      "weight_source": "STATIC",
      "weight_version": null,
      "served_model_name": "Qwen/Qwen3-0.6B",
      "engines": [
        {
          "engine_id": "math/replica-0",
          "base_url": "http://head:port",
          "state": "READY",
          "direct_eligible": true,
          "generation": 1
        }
      ]
    }
  }
}

状态不是简单的进程存活。一个副本只有 weights、KV cache、CUDA graph 全部 resident,权重有效且 admission 打开时才是 READY。worker 或 endpoint 换代会增加 generation;Manager 状态变化会推进 topology revision。

多节点 logical replica 的所有 follower 都参与健康验证,但只发布 head URL。external Rollout engine(--rollout-external 或 external scale-out group)和未实现 health_process 的自定义 Rollout engine 没有可探测的本地进程,只按 head 的 health_generate 和 endpoint 判定;本地 engine 报告进程状态未知时仍会被 fence。PD prefill/decode workers 不进入公共 replica 列表,PD 请求只能走 Router。

旧 Rollout /engines?schema=1 仍供 Autoscaler 和诊断兼容使用,只保留 regular head 的 rank、status、URL,过滤 PID、node metadata 和 P/D worker。Schema 1 不参与统一选路。

下面是 Gateway 和 direct client 共用的实际选路决策。

点击展开 Mermaid 图
flowchart TD
    S[Read and validate schema 2 snapshot] --> M{Explicit model present}
    M -->|yes| MA[Match model ID or alias]
    M -->|no| K{Route key present}
    K -->|yes| KM[Resolve route_key_map]
    K -->|no| D[Use default_model]
    MA --> Found{Model resolved}
    KM --> Found
    D --> Found
    Found -->|no| Bad[HTTP 400]
    Found -->|yes| Ready{Model state is READY}
    Ready -->|no| Sleep[HTTP 503 and refresh boundary]
    Ready -->|yes| Mode{route_mode}
    Mode -->|SGLANG_ROUTER| RU[Use model router_url]
    Mode -->|DIRECT| Candidates[Filter READY and direct_eligible heads]
    Candidates --> Any{Candidates available}
    Any -->|no| Unavailable[HTTP 503]
    Any -->|yes| Affinity{Affinity key present}
    Affinity -->|yes| Hash[Rendezvous hash by stable engine_id]
    Affinity -->|no| RR[Process-local round robin]
    Hash --> Target[RouteTarget]
    RR --> Target
    RU --> Target
Loading

InferenceManager 与 SGLangEngine

公共 Manager 核心

三角色共享的是同一个资源和生命周期核心,并不意味着三个角色已经具有完全相同的 Ray Actor class。GenRM/Teacher 使用继承,Rollout 使用组合。

点击展开 Mermaid 图
flowchart TB
    Core[InferenceManager common core]
    GenRM[GenRMManager Ray actor] -->|inherits| Core
    Teacher[TeacherManager Ray actor] -->|inherits| Core
    Rollout[RolloutManager Ray actor] -->|composition| RolloutCore[InferenceManager.for_rollout]
    RolloutCore --> Core

    Core --> Observation[InferenceObservation and schema 2 registry]
    Core --> Slots[Engine slots and logical replicas]
    Core --> Memory[onload offload and resident tags]
    Core --> Recovery[retire recover and cleanup]
    Core --> Ownership[owned or borrowed placement groups]

    Slots --> Head[Logical head]
    Slots --> Followers[Internal TP or PP followers]
    Head --> Engine[SGLangEngine]
    Followers --> Engine
    Observation --> Public[Publish head only]
Loading

InferenceManager 统一管理:

  • 扁平 Engine slots 与跨节点 logical replica;
  • 并行 actor 创建和并行 init.remote();
  • weights、KV cache、CUDA graph resident tags;
  • READY、DRAINING、SLEEPING、ONLOADING、FAILED、STOPPING 状态;
  • 重复 onload/offload 的 no-op;
  • 部分 onload 失败的补偿:Rollout 对已恢复的 replica 补偿 offload;静态角色仅在 backend shutdown 确认后清理并重建,terminal actor death 保留资源 fence;
  • 故障 replica 的完整清理与恢复;terminal actor death 的资源释放未确认时停止自动重建;
  • owned/borrowed placement group;
  • schema 2 discovery 和 endpoint generation。

静态角色直接使用公共基类路径。Rollout 因保留大量业务方法,采用 composition:

RolloutManager facade
  -> RolloutWorkload for generate/evaluate/data source
  -> InferenceManager.for_rollout() for resource ownership

Engine 能力隔离

三角色实际创建 SGLangEngine。InferenceEngineSpec 区分:

  • role;
  • STATIC、ACTOR 或 DCS weight source;
  • 允许的 generate/chat/lifecycle 操作;
  • strict drain;
  • checkpoint、CPU backup 和静态配置。

GenRM/Teacher 的 STATIC snapshot 不发布 Actor weight version,并拒绝 DCS、tensor weight update、seed sync 和动态 LoRA。

GenRM、Teacher 和 Rollout 现在都直接实例化 SGLangEngine。角色差异由构造时传入的 InferenceEngineSpec 和 server overrides 表达,不再通过 GenRM 专用 Engine subclass 表达。

生命周期失败边界

Manager 对每个 logical replica 单独记录成功与失败。Offload 部分成功时,成功项进入 SLEEPING,未确认项进入 FAILED;FAILED 不代表显存已释放,不能把共享 GPU lease 交给下一角色。

Shutdown 必须先等待所有 worker shutdown.remote(),再 kill actor,最后删除 owned PG。借用 Controller/Service PG 的 Manager 不删除 PG。清理任一步未确认时保留 pending 记录供后续重试。

Engine actor 已确认死亡(RayActorError,不含 ActorUnavailableError)不能证明 SGLang 子进程已退出。此时保留 slot、pending 与 PG 的资源 fence,停止重复 shutdown RPC,并抛出 InferenceRecoveryRequired 阻止自动重建。Controller 在删除 owner/Serve 和重新初始化前严格确认各角色清理;未确认时终止自动恢复,要求独立确认节点/container 清理后启动新作业。ActorUnavailableError 或普通超时仍按未确认状态处理,可重试确认式清理。

health_check() 覆盖每个 logical replica 的全部 worker:head 执行 health_generate,每个 worker 报告本地进程状态。发现故障 worker 时尝试整组清理;只有 backend shutdown 已确认才允许释放并重建,terminal actor death 保留未确认 worker 的资源 fence。GenRM /health(schema 1)会调用该检查,因此健康探测可能带来退役副作用。

点击展开 Mermaid 图
stateDiagram-v2
    [*] --> STARTING
    STARTING --> READY: full tags and valid weights and admission open
    READY --> DRAINING: close admission and start offload
    DRAINING --> SLEEPING: backend release confirmed
    SLEEPING --> ONLOADING: activation lease granted
    ONLOADING --> READY: missing tags restored and admission reopened
    STARTING --> FAILED: startup or health failure
    DRAINING --> FAILED: release not confirmed
    ONLOADING --> FAILED: resume not confirmed
    READY --> STOPPING: shutdown
    SLEEPING --> STOPPING: shutdown
    FAILED --> STOPPING: controlled cleanup
Loading

FAILED 只表示该副本不可路由,不表示 GPU 已释放。共享资源只有在 Manager 返回确定的 SLEEPING/offload 结果后才能交给下一角色。

Placement:先规划,再绑定,再启动

当前没有名为 PlacementPlanner 的 class。设计由函数式接口实现,见 relax/inference/placement.py:

  • plan_inference_placement(args):无副作用逻辑 preflight;
  • InferencePlacementPlan:不可变全局计划;
  • validate_bound_placement():PG ready 后物理拓扑校验;
  • model_placement():Manager 查询自身模型切片。
点击展开 Mermaid 图
flowchart LR
    Args[All role resources and parsed configs] --> Plan[plan_inference_placement]
    Plan --> Logical{Logical layout valid}
    Logical -->|no| Reject1[Reject before Teacher PG Router or Engine]
    Logical -->|yes| Saved[Serializable InferencePlacementPlan]
    Saved --> PG[Create or borrow Ray placement group]
    PG --> Probe[InfoActor probes node and GPU identity]
    Probe --> Bind[validate_bound_placement]
    Bind --> Physical{Physical topology valid}
    Physical -->|no owned PG| Cleanup[Remove newly owned PG]
    Physical -->|no borrowed PG| Stop[Stop deployment and preserve external owner]
    Physical -->|yes| Launch[Launch Managers and complete logical replicas]
Loading

逻辑 preflight

Planner 汇总 Actor、Critic、Reference、ActorFwd、Rollout、GenRM、Teacher,检查:

  • resource 结构和 GPU 预算;
  • 多模型预算之和;
  • GPU budget 能否整除每 engine GPU 数;
  • TP x PP = GPUs per engine;
  • DP attention 下 TP 可被 DP 整除;
  • 跨节点 engine 使用完整节点;
  • split 切片不超出 Actor pool;
  • defer role 只允许 GenRM/Teacher;
  • shared placement 必须支持 train/inference offload;
  • Ray fractional GPU reservation 在每个 bundle 上不超过 1;
  • 同卡同阶段共同驻留没有显式 defer 时拒绝。

这一步发生在 Teacher Manager、data source、PG、Router 和 Engine 创建前。失败不会留下 GPU actor。

三种模式

模式 当前行为
decoupled 每个推理角色使用独立 pool/PG,offset 可各自从 0 开始
split Rollout、GenRM、Teacher 在 Actor pool 内使用连续且不重叠的切片
defer deferred roles 可以在不同 phase 使用相同 bundle offset,由 Coordinator 串行授予 lease
点击展开 Mermaid 图
flowchart LR
    Mode{Placement mode}
    Mode --> Decoupled[Decoupled]
    Mode --> Split[Split]
    Mode --> Defer[Defer]

    Decoupled --> RP[Rollout pool]
    Decoupled --> GP[GenRM pool]
    Decoupled --> TP[Teacher pool]

    Split --> AP[One Actor pool]
    AP --> RS[Rollout slice]
    AP --> GS[GenRM slice]
    AP --> TS[Teacher slice]

    Defer --> Shared[Same shared bundle set]
    Shared --> P1[Generation phase: Rollout]
    P1 --> P2[Score phase: GenRM]
    P2 --> P3[Score phase: Teacher]
    P3 --> P4[Optional phase: Student]
Loading

例如 8-GPU Actor pool 的 split:

bundle:   0 1 2 3 4 5 6 7
Rollout:  R R R R . . . .
GenRM:    . . . . G G . .
Teacher:  . . . . . . T T

8-GPU defer:

generation:    R R R R R R R R
GenRM phase:   G G G G G G G G
Teacher phase: T T T T T T T T

物理绑定

PG ready 后,临时 InfoActor 探测每个 bundle 的 bundle_index、node_id、node_ip、gpu_id。绑定验证拒绝:

  • 重复或未知 bundle index;
  • 计划切片越界;
  • 两个 bundle 指向相同物理 GPU;
  • 一个 worker 的 GPU 跨节点混放;
  • 本地 GPU ID 不连续递增;
  • 跨节点 replica 重复使用同一节点。

验证通过后 Manager 才能启动 Engine。物理校验失败只删除当前 owner 拥有的 PG,借用 PG 不会被错误删除。

Colocation 与 deferred scoring

Split

Split 的布局和单节点真实 Ray bundle 绑定已经验证。它允许推理角色使用同一 Actor PG 的不同切片;进入 Actor 训练前仍需现有 barrier/offload 机制确保推理显存已释放。

Defer

当前 defer 是受约束的同步 colocate closed-batch 实现。固定阶段为:

Generate
  -> Rollout offload and acknowledgment
  -> optional GenRM onload, score, offload
  -> optional Teacher onload, prefill, offload
  -> optional Student policy restore, version check, prefill, offload
  -> validate and write back fields
  -> convert to training batch
  -> TransferQueue async_put
  -> commit batch
  -> training consumers proceed

生成侧原本准备写入 TQ 的完整 groups 会被 DeferredTransferCollector 截获到 CPU staging。DeferredBatch 对 Sample 做副本,评分先写 staged fields;只有所有阶段成功并通过校验,才把 reward、Teacher logprobs、top-k 和多模态字段写回原 Sample。

LifecycleCoordinator 是 CPU-only Ray Actor,维护:

  • 一个 pending closed batch;
  • 每个 role 的 SLEEPING/ACTIVE/ONLOADING/DRAINING/BLOCKED lease;
  • operation ID 去重;
  • run 级失败;
  • commit waiters。

Coordinator 内部状态

构造函数维护以下控制面状态:

字段 含义
_operation_timeout 一次 Manager lifecycle RPC 的截止时间,默认 180 秒
_condition 同时保护共享状态并唤醒 commit waiters 的 asyncio.Condition
_managers role 到 Manager handles tuple 的映射
_roles role 当前 lease 状态;初始含 rollout: SLEEPING
_batches 当前批次的 _Batch;begin_batch 时驱逐已提交批次,不保留完整历史
_current 当前批次 ID
_failure run 级首个永久失败信息
_operations 当前批次及未完成 operation 的 ID 到 fingerprint 和共享 Task 的映射
_committed 最近 1024 个已提交批次 ID 的有界历史
_retired_operations 已驱逐 operation 的 fingerprint,最多保留 8192 条

只有当前批次保存完整状态。_committed 让迟到的 wait_committed,以及重复的 begin_batch、commit_batch,在批次被驱逐后仍看到 COMMITTED;_retired_operations 在窗口内继续拒绝跨批次复用 operation ID。等待早于 1024 个提交窗口且已被驱逐的批次时,wait_committed 只能等到调用方超时。

每个 _Batch 只保存四个控制字段:

state = "PENDING"
error = None
rollout_released = False
student_scoring_started = False

Coordinator 不存样本、reward、Teacher logits 或 TQ partition。样本和评分结果属于 DeferredBatch;Coordinator 只判断批次能否继续、共享 GPU lease 是否已经安全释放,以及训练消费者能否开始。

Batch 与角色 lease 两套状态机

Coordinator 同时维护 batch 状态和角色 lease 状态。Batch 状态回答“这批数据能否训练”,lease 状态回答“当前哪个角色仍可能占用共享 GPU”,两者不能混用。

Batch 状态机:

点击展开 Mermaid 图
stateDiagram-v2
    [*] --> PENDING: begin_batch
    PENDING --> COMMITTED: commit_batch and all roles sleeping
    PENDING --> FAILED: fail_batch or lifecycle RPC failure
    COMMITTED --> [*]
    FAILED --> [*]
Loading

Coordinator 的 batch state 只有 PENDING、COMMITTED 和 FAILED。GENERATED、SCORING、COMPLETE、PUBLISHING 等更细的数据处理阶段属于 DeferredBatch。

角色 lease 状态机:

点击展开 Mermaid 图
stateDiagram-v2
    [*] --> SLEEPING
    SLEEPING --> ACTIVE: activate or begin_batch for rollout
    ACTIVE --> DRAINING: deactivate begins
    DRAINING --> SLEEPING: all manager offloads succeed
    SLEEPING --> ONLOADING: activate begins
    ONLOADING --> ACTIVE: all manager onloads succeed
    ONLOADING --> BLOCKED: failure or timeout
    DRAINING --> BLOCKED: failure or timeout
Loading

BLOCKED 是安全状态,而不是普通失败标签。它表示远端 onload/offload 的结果不确定,宁可停止当前训练,也不能把同一块 GPU 分配给下一个角色,BLOCKED 作为共享GPU 的 fail-stop 安全栅栏。Coordinator 不会因为调用超时或远端 RPC 稍后完成,就自动将角色改为 SLEEPING 并把同一组 GPU 交给下一个角色。

Manager RPC 超时后 lease 保持 BLOCKED,即使远端稍后完成也不自动放行。Rollout 为避免 Actor self-RPC deadlock,执行本地 offload 后向 Coordinator 发送 acknowledgment;Teacher/GenRM 由 Coordinator 直接调用远程 Manager。

TQ async_put(train_N, is_last=true) 成功后,先执行本地 finalizer(debug 保存和指标),再调用 commit_batch(N);finalizer 失败时先清理 train_N 分区再报错,不会 commit。Actor、Critic、ActorFwd 和 Advantages 在消费前调用 wait_committed(N)。评分或发布失败会唤醒 waiters 并抛出 batch error,不产生训练许可。

生成阶段在 begin_batch 之后失败时,Rollout 在本地执行 offload;只有确认释放后才发送 acknowledgment,否则保留 lease,并在 batch 错误中标注 release unconfirmed。Student restore 失败同样会在 finally 中释放 Rollout lease。

当前 defer 明确拒绝 partial rollout、dynamic batch、自定义 reward/convert/filter、SFT、非 Megatron backend、带评估数据的 defer、多 GenRM 模型和多组 Rollout/PD 等尚无阶段 adapter 的组合。

完整数据流如下。训练消费者可以先等待 batch ID,但只有评分、字段回写和 TQ 写入全部完成后才会收到 commit。

点击展开 Mermaid 图
sequenceDiagram
    participant W as RolloutWorkload
    participant C as LifecycleCoordinator
    participant B as DeferredBatch
    participant R as RolloutManager local lifecycle
    participant G as GenRM Manager
    participant T as Teacher Manager
    participant Q as TransferQueue
    participant A as Training consumers

    A->>C: wait_committed(batch N)
    W->>C: begin_batch(N)
    W->>W: generate and capture groups in CPU staging
    W->>B: complete(groups)
    B->>B: clone samples and prepare Teacher inputs
    B->>R: local Rollout offload
    R-->>B: offload confirmed
    B->>C: acknowledge_rollout_offloaded(N)

    opt GenRM deferred
        B->>C: activate(genrm, operation ID)
        C->>G: onload.remote()
        G-->>C: confirmed
        B->>B: compute reward and write staged samples
        B->>C: deactivate(genrm, operation ID)
        C->>G: offload.remote()
        G-->>C: confirmed
    end

    opt Teacher deferred
        B->>C: activate(teacher, operation ID)
        C->>T: onload.remote()
        T-->>C: confirmed
        B->>B: Teacher prefill and field writeback
        B->>C: deactivate(teacher, operation ID)
        C->>T: offload.remote()
        T-->>C: confirmed
    end

    opt Student scoring required
        B->>C: begin_student_scoring(N)
        B->>R: local Rollout onload and version check
        B->>B: student prefill
        B->>R: local Rollout offload
        R-->>B: confirmed
        B->>C: acknowledge_rollout_offloaded(N)
    end

    B->>B: validate and copy fields to original samples
    B->>Q: async_put(train_N, is_last=true)
    Q-->>B: write returned successfully
    B->>C: commit_batch(N)
    C-->>A: wake all waiters
    A->>Q: consume committed training partition
Loading

Shared co-resident

第一版拒绝同卡同阶段常驻。_genrm_colocate_with_rollout 如果没有显式 defer 会在 preflight 报错。减小 mem_fraction_static 不能绕过,因为 Ray reservation、峰值显存和生命周期冲突仍然存在。

HTTP API 变化

新增公开角色路由

Role Method Path Semantics
Rollout POST /rollout/generate Raw input IDs/text generation,支持 raw streaming
Rollout POST /rollout/chat/completions Chat alias
Rollout GET /rollout/health 统一 schema 2 readiness
GenRM POST /genrm/chat/completions Chat alias
GenRM POST /genrm/v1/chat/completions OpenAI Chat proxy
GenRM GET /genrm/v1/models 逻辑模型列表
GenRM GET /genrm/engines schema 2 discovery
Teacher POST /teacher/generate Raw generation/OPD prefill
Teacher POST /teacher/chat/completions Chat alias
Teacher POST /teacher/v1/chat/completions OpenAI Chat proxy
Teacher GET /teacher/v1/models 逻辑模型列表
Teacher GET /teacher/engines schema 2 discovery
Teacher GET /teacher/health schema 2 health

合计新增 13 个角色级公开路由。

已有路由的行为变化

Method Path Change
POST /rollout/v1/chat/completions 改用公共 handler 和 resolver
GET /rollout/v1/models 返回 registry 逻辑模型 IDs
GET /rollout/engines 默认 schema 2;显式 schema 1 为过滤后的兼容诊断
POST /genrm/generate 保留 messages 协议,同时接受 raw input
GET /genrm/health 默认旧 schema 1;?schema=2 返回统一 health

没有新增公开 /activate、/offload 或 /commit_batch。这些是内部 Ray RPC。

Engine admission 内部入口

严格 defer Engine head 前增加 CPU admission server,允许的路径包括:

  • 生成:/generate、/v1/chat/completions、/chat/completions、/v1/completions;
  • 取消:/abort_request、/close_session;
  • 元数据:/v1/models、/server_info、/get_server_info、/model_info、/get_model_info;
  • 健康:/health、/health_generate。

Weight update、pause、onload/offload 不在公共 admission allowlist。Gateway 和 Router 访问 guard URL,底层 SGLang 端口不作为框架 direct client 的旁路。

Key implementation files

Area Implementation
Gateway protocol and SSE relax/inference/gateway.py
Discovery-aware direct client relax/inference/client.py
Shared routing resolver relax/inference/routing.py
Schema 2 snapshots relax/inference/specs.py
Legacy Manager aggregation relax/inference/compat.py
CPU-only Teacher Gateway relax/components/inference_gateway.py
Common Manager lifecycle relax/distributed/ray/inference_manager.py
Unified SGLang Engine relax/backends/sglang/sglang_engine.py
Placement planning and binding relax/inference/placement.py
Shared-GPU lease coordinator relax/distributed/ray/inference_lifecycle.py
Deferred producer wiring relax/inference/defer.py
Deferred scoring and writeback relax/engine/rollout/deferred_scoring.py
Rollout facade and ownership adapter relax/distributed/ray/rollout.py

兼容性

本 PR 保留以下兼容面:

  • Rollout 原有 generate/evaluate/predict、step、scale-out/in、恢复和权重同步接口;
  • get_rollout_engines_and_lock() 五元组;
  • GenRM messages 请求和 {response: string} 返回形状;
  • GenRMClient、旧 health/metrics;
  • Teacher get_urls() 和旧 args URL 适配;
  • RolloutManager、GenRMManager、TeacherManager 命名;
  • GenRM 通过 SGLangEngine 的显式 role spec 构造;
  • Rollout /engines?schema=1;
  • external Teacher/外管生命周期。

兼容不包括危险布局。旧配置如果表达同卡同阶段共同驻留、无法确定 owner 或不满足完整 replica,会提前报错。

旧兼容 Manager 模块及其导入路径已经删除。GenRM 和 Teacher 直接继承 InferenceManager。

Failure semantics

Failure Current behavior
Registry unavailable Gateway alive,但 readiness false;生成返回 503
Model SLEEPING/DRAINING/FAILED 不发送 GPU 请求,返回 503 和 Retry-After
Upstream transport failure Gateway 返回 502
Request deadline 返回 504
POST timeout or cancellation 失效 discovery 缓存,不重放 POST
Invalid payload/model/route key 返回 400
Engine lifecycle partial failure 成功 replica 与失败 replica 分开记录;FAILED 不发布 URL
Engine actor confirmed dead 保留 slot/pending/PG fence,停止旧 shutdown RPC;抛出 InferenceRecoveryRequired,禁止自动重建,需确认节点/container 清理后启动新作业
Multi-node follower death health_check() 失败并尝试整组清理;成功 worker 可清理,未确认的 terminal worker 保持 fence,不能释放共享 PG 或重建
Lifecycle RPC timeout Coordinator 保留 BLOCKED lease,后续角色不激活
Scoring/validation failure Batch FAILED,不转换、不发布训练许可
Producer failure after begin_batch 本地 offload Rollout;确认释放才 ack,否则保留 lease;batch FAILED
Student restore failure finally 中 offload 并 ack Rollout lease
TQ put failure 不 commit,消费者失败退出
Finalizer failure after TQ put 清理 train_N 分区,不 commit
Owned PG cleanup failure 保留 pending cleanup 并向上抛错
Borrowed PG shutdown Manager 停 Engine,但不删除外部 PG
点击展开 Mermaid 图
flowchart TD
    Failure[Lifecycle or scoring failure] --> Fence[Close admission and remove route eligibility]
    Fence --> Known{GPU release confirmed}
    Known -->|yes| Sleeping[Role becomes SLEEPING]
    Known -->|no or timeout| Blocked[Role remains FAILED or BLOCKED]
    Blocked --> StopNext[Do not activate conflicting role]
    Blocked --> FailBatch[Fail current batch and wake waiters]
    Sleeping --> Data{Scoring and TQ publish complete}
    Data -->|yes| Commit[Commit training batch]
    Data -->|no| FailBatch
Loading

HTTP POST 发出后,公共 client 不自动重放,因为响应丢失时 backend 可能已经执行请求。旧 GenRMClient 仍有自己的有限重试策略,因此系统没有端到端 exactly-once 保证。

Validation

CPU 回归和单节点 4×B200 验证均通过,本次 未执行多节点测试

Acceptance checks

状态定义:PASS 表示当前代码和对应回归已满足;PARTIAL 表示实现与 CPU/单节点证据成立,但最终 GPU 闭环仍缺失;SKIP 表示当前环境不具备所需硬件,未把静态分析或单节点结果冒充为通过。

Acceptance check Status Evidence and remaining boundary
Same gateway class serves Rollout, GenRM, and Teacher. PASS 三个入口都构造 InferenceGatewayHandler;Rollout/GenRM 组合 handler,Teacher 使用独立 CPU-only InferenceGateway deployment。test_gateway.py 与 test_teacher_gateway_wiring.py 覆盖三角色协议和部署资源。
Same manager class owns all three engine-pool types. PASS GenRM/Teacher 直接继承 InferenceManager,Rollout 通过 InferenceManager.for_rollout() 组合相同 ownership core;旧 MultiEngineManager 已删除。test_inference_manager.py、test_rollout_core.py 和 Manager 继承断言通过。
Gateway and direct clients share routing rules. PASS 两条路径都使用 InferenceClient 和 RouteResolver;test_gateway_and_direct_resolver_choose_same_affinity_target、gateway transport 与 resolver 测试通过。
Discovery never exposes internal TP/PP workers as replicas. PASS schema 2 仅发布 logical replica head;多节点 follower 只参与健康检查,PD prefill/decode 只经 Router。CPU 契约测试与单节点 TP4 Teacher、双 TP2 GenRM 实测通过;跨节点运行仍见最后一项。
Split overlap and invalid defer schedules fail before startup. PASS plan_inference_placement() 和 validate_deferred_workload() 在 Manager、Router、Engine 创建前拒绝切片重叠、资源超限和无阶段 adapter 的 defer 组合;placement/defer preflight 测试通过。
Shared co-resident layouts fail before startup. PASS 第一版要求显式 defer;同卡同阶段共同驻留和 Ray fractional reservation 冲突会在 preflight 抛错,test_deferred_ppo_rejects_persistent_ray_actor_reservation_conflict 等测试通过。
Static models cannot register for DCS or dynamic weight updates. PASS GenRM/Teacher 使用 STATIC Engine spec,拒绝 DCS、tensor/distributed update、LoRA 和 seed sync;能力测试通过,GenRM 的 DCS/tensor update 拒绝也有单节点真实 SGLang 证据。
Teacher defer writes Teacher outputs back before training. PARTIAL DeferredBatch 先校验并写回 Teacher sampled/top-k 字段,再 async_put(train_N),最后 commit_batch(N);CPU workflow 测试通过。尚未运行真实 Teacher → TQ → Actor 训练闭环。
PG ownership and rollback are covered by tests. PASS 测试覆盖 owned/borrowed PG、启动失败回滚、批量 scale-in 部分成功、shutdown 重试和 binding 失败;清理失败的 group 保持 DRAINING,PG 仍由 InferenceManager.remove_group() 按 ownership 释放。单节点实测也确认 Teacher-owned PG 被移除、GenRM borrowed PG 保留。
Multi-node GPU tests pass, or skip with an explicit hardware reason. SKIP 当前只有单节点 4×B200 环境,不具备至少两个 Ray GPU 节点,无法执行跨节点 TP/PP、节点故障和跨节点 PG 验证;相关 logical topology/CPU 测试通过,但不计作多节点 GPU 通过。

Type of change

  • Feature:统一 discovery、routing、Gateway 和 deferred scoring
  • Refactor:公共 Manager/Engine lifecycle 和 Rollout ownership
  • Reliability:placement preflight、admission fencing、失败回滚、PG ownership
  • Compatibility:保留 Manager 类名、client、schema 1 和接口形状
  • Breaking import cleanup:删除旧 Manager 兼容模块及其 Python import path
  • Tests and validation evidence
  • GenRM compatibility Engine subclass removed
  • Full test suite green
  • Real offload/onload lifecycle passed
  • Multi-node GPU and full training acceptance passed

@rai-studio-bot

rai-studio-bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Nyanpasu 审查看板

审查状态: ✅ 已通过

审查版本: 43e1a21

目标分支: main

本轮复核(head 43e1a21):上轮两项遗留 P2(Teacher client 按样本重建;defer 预检拒绝现有示例)已修复并经代码与本地实验证实;17 项既往发现保持已解决或已取代,无新增。defer 示例完成预检级验证,多机执行由维护者 nightly 覆盖。结论 APPROVE。

审查阶段进度范围与结果
常规审查 ✅ 已完成 本轮增量(b36aac4b→43e1a21b):Teacher client 改为循环级缓存复用(语义不变,shutdown_async_loop 关闭);defer 示例与 README/中英文档迁移到框架评分适配器。46 项相关测试通过;脚本真实参数复跑解析、校验、placement 与 defer 校验通过预检。线程 3、4 修复核实。CI 全绿。
深度审查 ✅ 已完成 必要性审计沿用既往决策,增量复核未发现失效:client 复用删除按 prefill 重建机制,示例迁移复用既有 registry 能力契约。新增 preflight 测试复刻 flag,本轮以真实参数实验补足。多节点硬件为作者声明 SKIP。

审查发现

待处理
编号 严重性 问题状态规则来源
暂无待处理的记录。
已解决或已取代
编号 严重性 问题状态规则来源
F1 High severity 保留 Rollout 配置的 HTTP 并发容量 ✅ 已解决 —
F2 Medium severity 恢复 external/debug 模式聊天入口 ✅ 已解决 —
F3 Medium severity 在 deferred scoring 后补齐统计与样本导出 ✅ 已解决 —
F4 Medium severity 保留多节点部分清理后的重试入口 ✅ 已解决 —
F5 Medium severity 修正 Serve 异步析构的文档误判 ✅ 已解决 —
F6 Medium severity 为终止 actor 提供可完成的清理恢复路径 ✅ 已解决 —
F7 Medium severity Teacher 测试缺少可选依赖跳过保护 ✅ 已解决 —
F8 Medium severity 恢复测试替身留下的父包模块缓存 ➖ 已被后续变更取代 —
F9 Low severity 更新描述中的过期提交引用 ✅ 已解决 —
F10 Medium severity 清理 GenRM 旧 engine 调用链
精简建议(非阻塞)
✅ 已解决 —
F11 Low severity 删除不可达的 GenerateRequest 分支
精简建议(非阻塞)
✅ 已解决 —
F12 Low severity 移除未被调用的 teacher_base_url
精简建议(非阻塞)
✅ 已解决 —
F13 Low severity 顺手移除无引用的 EnginesInfoResponse
精简建议(非阻塞)
✅ 已解决 —
F14 Low severity 共享 GenRM GPU fraction 公式
精简建议(非阻塞)
✅ 已解决 —
F15 Low severity 把 collector 的 max_samples 传入 DeferredBatch
精简建议(非阻塞)
✅ 已解决 —
F16 Low severity 处理从未被发布的 phase/phase_epoch 字段
精简建议(非阻塞)
✅ 已解决 —
F17 Low severity 共享 sglang_engine_module 替身 fixture
精简建议(非阻塞)
✅ 已解决 —
提交范围 · 接收 98 · 建议移出 0 · 待确认 0

接收 98 个文件 · 建议移出 0 个文件 · 待确认 0 个文件。移出与待确认部分暂停深审,不代表审查通过。

文件结论仓库维护必要性依据替代去向或方案
relax/inference/__init__.py
relax/inference/admission.py
relax/inference/client.py
relax/inference/compat.py
relax/inference/defer.py
relax/inference/engine_spec.py
relax/inference/gateway.py
relax/inference/placement.py
relax/inference/registry.py
relax/inference/routing.py
relax/inference/specs.py
接收 统一 discovery schema、模型路由、Gateway handler、placement preflight、admission 与推理客户端的公共基础设施;本轮 Teacher client 切换到 generate_with_discovery 循环级缓存复用,其余为前几轮已审内容。 PR #347(Relates to #71 RFC);调用方 relax/core/controller.py、relax/components/*、relax/engine/rollout/sglang_rollout.py(多轮已核对);Yangruipis 对 defer.py 与业务耦合的审查意见 按角色保留独立 discovery/路由路径无法提供统一选路与 placement 校验契约;保留 dapo_genrm 硬依赖则维持与特定 reward adapter 的耦合(审查线程要求解耦)。
relax/backends/sglang/sglang_engine.py
relax/distributed/ray/genrm.py
relax/distributed/ray/inference_lifecycle.py
relax/distributed/ray/inference_manager.py
relax/distributed/ray/multi_engine_manager.py
relax/distributed/ray/multi_instance_orchestrator.py
relax/distributed/ray/placement_group.py
relax/distributed/ray/rollout.py
relax/distributed/ray/rollout_workload.py
relax/distributed/ray/teacher_manager.py
接收 公共 Manager 生命周期核心、LifecycleCoordinator 共享 GPU lease 门控与统一 SGLangEngine;MultiEngineManager 删除。本轮增量全部来自 main 侧合并(offline 命名、设备模块化),未触及本层职责;merge 分辨已逐文件核对无 PR 内容丢失。 PR 描述 RFC Decision 对照表;调用方 relax/core/service.py、relax/core/controller.py 与 relax/inference/defer.py 按角色保留独立 Manager 会重复生命周期与 PG ownership;删除统一核心会让 defer 的 GPU lease 门控无处承载。
relax/components/actor.py
relax/components/actor_fwd.py
relax/components/advantages.py
relax/components/critic.py
relax/components/genrm.py
relax/components/inference_gateway.py
relax/components/rollout.py
relax/core/controller.py
relax/core/service.py
接收 角色服务接入统一 Gateway/Manager 与训练批次 commit 门控。本轮 merge 分辨要点:actor._wait_for_data 将 main 的 is_offline_mode/分区检查重构与 PR 的 wait_inference_commit 门控合并,语义等价(None 判定等价);controller 仅注释更新。 PR 描述;调用链 relax/entrypoints/train.py → core/service → core/controller 不接入则新基础设施无法使用;去掉 commit 门控会让 defer 训练消费在评分写回前读取 TQ。
relax/agentic/pipeline/runtime.py
relax/agentic/pipeline/transfer.py
relax/agentic/rollout.py
relax/backends/megatron/actor.py
relax/engine/rewards/__init__.py
relax/engine/rewards/registry.py
relax/engine/rollout/deferred_scoring.py
relax/engine/rollout/on_policy_distillation.py
relax/engine/rollout/sglang_rollout.py
relax/utils/arguments.py
relax/utils/async_utils.py
relax/utils/autoscaler/autoscaler_service.py
relax/utils/env.py
relax/utils/health_monitor.py
relax/utils/opd/opd_utils.py
relax/utils/utils.py
接收 deferred scoring 与原生/Agentic/OPD 生成路径调用点;registry 能力位契约不变。本轮 on_policy_distillation 移除按 prefill 的 client 重建(对应审查线程),未托管 aiohttp 直连路径不变。 PR 描述;defer 解耦审查线程(维护者指出与 dapo_genrm 业务逻辑耦合);调用方 sglang_rollout 生成路径、opd_utils 在 defer.py 内联 if rm_type=="dapo-genrm" 维持旧耦合(已被审查否决);或为每个 adapter 写独立 defer 分支(重复注册表已有的 mode/能力机制)。
tests/backends/sglang/engine_module_stub.py
tests/backends/sglang/test_genrm_offload_drain.py
tests/backends/sglang/test_router_registration.py
tests/backends/sglang/test_weight_version.py
tests/components/test_genrm_engine_pick.py
tests/core/test_controller_inference_restart_fence.py
tests/core/test_service_affinity.py
tests/distributed/ray/conftest.py
tests/distributed/ray/test_coordination.py
tests/distributed/ray/test_genrm_recovery.py
tests/distributed/ray/test_inference_manager.py
tests/distributed/ray/test_multi_engine_manager.py
tests/distributed/ray/test_opd_multi_teacher_orchestration.py
tests/distributed/ray/test_scale_in.py
tests/distributed/ray/test_scale_out_registration_order.py
tests/distributed/ray/test_teacher_manager.py
tests/distributed/ray/test_teacher_manager_factory.py
tests/inference/test_admission.py
tests/inference/test_client.py
tests/inference/test_defer_wiring.py
tests/inference/test_deferred_scoring.py
tests/inference/test_engine_capabilities.py
tests/inference/test_gateway.py
tests/inference/test_gateway_admission_integration.py
tests/inference/test_legacy_contracts.py
tests/inference/test_lifecycle_coordinator.py
tests/inference/test_managed_teacher_client.py
tests/inference/test_manager_discovery.py
tests/inference/test_manager_lifecycle.py
tests/inference/test_multinode_scale_out.py
tests/inference/test_placement.py
tests/inference/test_placement_wiring.py
tests/inference/test_rollout_chat_compat.py
tests/inference/test_rollout_contracts.py
tests/inference/test_rollout_core.py
tests/inference/test_rollout_ownership.py
tests/inference/test_rollout_recovery_urls.py
tests/inference/test_rollout_terminal_cleanup.py
tests/inference/test_routing_registry.py
tests/inference/test_shutdown_ownership.py
tests/inference/test_static_placement_wiring.py
tests/inference/test_teacher_gateway_wiring.py
tests/inference/test_terminal_actor_cleanup.py
tests/inference/test_weight_publication.py
tests/test_s3_model_loader.py
tests/utils/test_health_monitor.py
接收 新推理基础设施回归测试;本轮新增 client 复用断言与 defer 示例预检用例(复刻脚本 flag)。 GitHub CI 执行 tests/inference/*、tests/distributed/ray/*、tests/backends/sglang/*;b36aac4b 提交声明的测试意图 仓库外验证无法保护回归;能力门控若无 remote_rm 拒绝用例则退化为按名字白名单。
docs/public/openapi/genrm.json
docs/public/openapi/rollout.json
接收 仓库既有 docs/public/openapi/ 每角色一份的提交式 OpenAPI 规范;本轮因 merge 携入 main 的路由/规范变化并按合并说明再生成(rollout.json 两侧本都有该文件,main 侧更新后重新出现在 PR diff 中)。 docs/public/openapi/ 目录既有模式;bb2aa116 合并说明 "Regenerate OpenAPI specifications from the merged service routes" 改为 docs 构建期生成会偏离仓库既有提交式规范模式;保持同步提交即可。
examples/generate_reward_model/README.md
examples/generate_reward_model/run-qwen35-35B-A3B-16xgpu-genrm-397B-defer.sh
接收 迁移 defer 示例与 README 到框架评分适配器(--rm-type dapo-genrm + --inference-defer-roles genrm),移除预检拒绝的 dummy rm、legacy hook、eval 与 dynamic-batch flag;真实参数复跑解析、校验、placement 与 defer 校验通过预检。 维护者审查线程要求迁移示例并验证通过预检;defer.py 校验器禁用 flag 清单。 旧 flag+hook 会被预检直接拒绝;框架适配器是现有 registry 契约的最小迁移路径。
docs/en/examples/generative-reward-model.md
docs/zh/examples/generative-reward-model.md
接收 与迁移后示例同步的中英文 GenRM 示例页(flag 与三阶段 defer 说明),属示例配套文档。 示例脚本迁移提交;docs/en|zh/examples/ 既有中英对齐模式。 不同步文档会让用户照旧 flag 启动即被预检拒绝;单开 PR 会脱离本次修复证据链。
精简审查与验证依据
审查范围进度结论
生产代码 ✅ 已完成 43 个生产文件审计完成于 ec350f6,决策记录见下表;其中 F10–F16 对应的 remove/merge 决策已在 2785a46 实现,F14 并新增直接执行真实 GenRMManager._ray_resource_kwargs 的 preflight/executor 一致性测试(补上此前缺口)。本地缺 torch/numpy 的部分以静态检索 + 编译 + 可收集测试为证据。 本轮(b36aac4b)按新 head 复核:新增 supports_deferred 能力位为审查线程 1 要求的最小解耦契约,未发现可删除的重复机制;既往 43 文件决策在增量核对后沿用。
测试 ✅ 已完成 测试审计完成于 ec350f6 并随后续提交扩展复核:F17 共享替身已实现;2101a0f0 新增 5 个 F4/F6 回归文件与 gateway→admission 集成测试(补上审计记录的 CPU 级双真实组件集成缺口)。本地缺依赖路径的检测力以静态阅读 + 该 head 8 项 CI 绿色为准;多节点 GPU 验收仍为作者声明的 SKIP。 本轮(b36aac4b)增量测试复核:离线拒绝参数化(sft/dpo/rm)与能力路由接受/拒绝用例针对新契约的独立边界,38 项 defer 测试本地通过。

生产代码的必要性与替代方案

范围必须保留的契约更简单的方案结论依据与限制
relax/components/genrm.py 旧 engine 调用链(~110 行)与 tests/components/test_genrm_engine_pick.py GenRM /generate 与 /chat 由 InferenceGatewayHandler 提供服务;调用方为 GenRMClient 兼容路径与 test_gateway.py 删除 _EngineCacheState/_engine_caches/_pick_engine/_call_engine/_call_legacy_payload、对应测试与无消费者的 Envs.GENRM_ENGINE_RETRY_ATTEMPTS 删除 全树静态检索(head/merge-base/PR 全部提交)确认零生产调用方;本地有界删除后可收集 CPU 套件与基线一致;engine-pick 测试因缺 torch 本地跳过(F10)
relax/components/genrm.py Message/GenerateRequest/GenerateResponse 与 isinstance 分支 FastAPI 端点签名 request: Request,全树无 GenerateRequest 构造 仅保留 handle_generate 路径 删除 静态检索 + 删除实验套件不变(F11)
relax/inference/compat.py teacher_base_url OPD teacher URL 规范化内联于 opd_utils.py:420;该函数无生产调用方 删除函数、urlsplit/urlunsplit 导入及其单测 删除 git grep 于 2d88481/f782419/HEAD~1/HEAD 均无调用方;删除后 test_client 等 148 项通过(F12)
relax/components/rollout.py EnginesInfoResponse(95–99) /engines 端点无 response_model;head 与 merge-base 均无引用 删除该模型 删除 静态检索;rollout AST 测试 17 passed/12 skipped,跳过均为既有(F13)
schema 2 的 phase/phase_epoch 字段 快照公开面;无生产者设置,唯一读者 compat.py:76 恒读 0 删除字段与描述示例,或让 Coordinator 真正发布 phase 删除 publish() 三个调用方均用默认值;描述示例展示的 phase_epoch=8 系统从不产生(F16)
relax/distributed/ray/multi_engine_manager.py 的删除 旧兼容层;本 PR 以 InferenceManager 取代 保留兼容层 删除 删除正确:全树零悬空引用(子任务 B 核实),保留反而需要维护两套权威状态
GenRM per-actor Ray GPU fraction 公式(placement.py:232-238 preflight 与 genrm.py:142-146 executor) preflight 校验份额须与 Ray 实际分配一致;调用方 plan_inference_placement 与 GenRMManager._ray_resource_kwargs 提取共享 helper genrm_persistent_gpu_fraction(args) 合并 本地提取 12 行 helper 后 325 项相关测试通过;无可直接执行 genrm 侧的测试(现有测试用替身),以导入对称性验证(F14)
defer.py 与 deferred_scoring.py 的 65536 sample 上限 DeferredTransferCollector 与 DeferredBatch 代表同一采样上限 构造 DeferredBatch 时透传 collector.max_samples 合并 静态确认 _score_and_publish 未透传;两个 defer 测试模块因缺 numpy 本地无法运行,记录为缺口(F15)
relax/inference/admission.py 与 sglang_engine/gateway 门控分工 engine 侧 sleep/drain 准入门(发布为 engine base_url)与 gateway/Router 访问控制 合并为单一门控 保留 admission 是 engine 直连边界的准入门,gateway/Router 是访问控制;重试语义可观察地不同且有测试固定——合并会失去 engine 边界防护
relax/inference/specs.py/registry.py/routing.py 三模块划分 schema 定义 / 可变发布 / 解析选路 合并为单模块 保留 职责互斥,无重复权威状态(子任务 A 逐项核实)
defer 管线三模块(defer.py / deferred_scoring.py / inference_lifecycle.py) 捕获 / staging 与字段回写 / lease 与 commit 门控 合并为单模块 保留 分属不同权威;operation 去重仅存在于 coordinator,无其他重复(子任务 B 核实);batch 状态与 lease 状态为有意的两个状态机
manager 与 engine 两侧 resident-tag 状态 manager 缓存供 schema 2 发布与部分 onload 差分;engine 侧门控 continue_generation/open_admission 单一权威来源 保留 manager 缓存按 worker 身份键控;manager 旁路恢复路径(RolloutServer.recover 直呼 resume_memory_occupation)依赖 engine 侧状态
controller/service/opd_utils 三处 actor pool 绑定校验 PG 创建后的物理绑定校验;调用方为各自 PG 生命周期 共享单一校验 helper 保留 三处守卫不同 PG 与清理所有权(controller 移除 owned PG、Service 仅非 shared、opd 移除 owned shared PG);模式重复约 8 行/处,合并需参数化清理语义(父任务结论)
父任务负责的 16 个调用点/编排文件(components 四角色、core/controller、core/service、agentic×3、megatron/actor、rewards、utils×5) wait_inference_commit 训练消费门控、discovery 切换保留 external 回退、权重更新 mark_inference_weights_updating/ready 围栏、defer 捕获钩子 维持旧路径不接入新基础设施 保留 父任务逐文件复核:均为最小调用点迁移,无重复逻辑;wait_inference_commit 在非 defer 模式安全 no-op
relax/inference/client.py 双缓存与 gateway 请求路径 事件循环键控 httpx 连接池(async_utils 关闭)与 TTL discovery 快照缓存 cache_ttl=0 时跳过 deepcopy/缓存写入的快速路径 保留 两缓存职责不同;网关路径每次请求 2 次 deepcopy + 3 次快照解析为可优化点,2 行快速路径本地验证通过,属性能微优化而非维护问题,未作为发现发布
RolloutWorkload facade 与 inference_manager 组合 测试固定的 ownership 接缝(test_rollout_ownership.py:12-70) 回退继承或并入 RolloutManager 保留 薄但职责明确;子任务 B 核实无重复
on_policy_distillation.py 双 OPD 路径(legacy 容忍 vs deferred 严格) 两条路径均有真实调用方与不同失败语义 合并为单一路径 保留 子任务 B 追踪调用方后否决合并候选
relax/inference/defer.py 对 GenRM 评分适配器的选择(b36aac4b 的 RewardSpec.supports_deferred 契约) deferred 评分经由 reward registry 解析适配器;preflight 拒绝未声明 deferred 能力的 async 适配器(relax/engine/rewards/registry.py:62、relax/inference/defer.py:64-77) 在 defer.py 内联 `` `rm_type == "dapo-genrm"` `` 判断,不在 registry 增加能力位 保留 审查线程 PRRT_kwDOSBAwr86rA-z6 明确要求解耦 defer 与 dapo-genrm 业务;能力位使未来 adapter 无需修改 defer.py 即可声明 deferred 支持,registry 已有 mode/label_matcher 同类机制,成本一个 bool 字段。运行时核对:dapo-genrm=async+deferred、remote_rm=async+不可 deferred;test_defer_wiring.py 38 项与 rewards/legacy 161 项本地通过。

测试的必要性与替代方案

范围必须保留的契约更简单的方案结论依据与限制
tests/test_s3_model_loader.py 与 tests/backends/sglang/test_router_registration.py 的 sglang_engine_module 替身 fixture 两文件分别保护 S3 模型加载与 router 注册边界(测试本身不合并) 共享 http_utils/logging 替身 fixture 或 helper,仅保留各自 ServerArgs 差异(object vs 完整 dataclass) 合并 本 PR 需在两处加入完全相同的 find_available_port 存根行,同步负担已体现(F17,已发布行级评论)
tests/inference/* 22 个文件(274 项可收集测试) 路由/发现解析、gateway HTTP/SSE 错误映射、admission 围栏、engine 能力拒绝、manager 生命周期与 PG ownership、placement 拒绝、defer 顺序与 commit 门控、legacy 兼容面、rollout ownership、权重发布 合并或删除重复用例 保留 按契约分组后各家族保护不同边界;3 个定向变异(admission 503 无 Retry-After、specs._url 接受任意 URL、inference_manager 杀死未确认 engine)均被预期测试以预期原因捕获;无可删除用例
tests/distributed/ray/*、tests/backends/sglang/*、tests/core/test_service_affinity.py、tests/test_s3_model_loader.py(15 路径,13 家族) manager 基础生命周期、scale-in/重试/协调/扩容注册、genrm/teacher 子类、SGLang engine HTTP 控制面(含新权重版本发布)、服务亲和、S3 加载;test_multi_engine_manager→test_inference_manager 为随生产模块删除的重命名 合并或删除重复用例 保留 3 个定向变异(partial-resume 记为全量、follower 查 /model_info、多节点 shutdown worker slot 泄漏)均检出;两个多节点重试测试 setup 相近但覆盖不同重试入口(数量收养 vs 同 URL 重解析),保留两者
test_shutdown_ownership.py 与 test_teacher_gateway_wiring.py 的 _run_global_restart 脚手架 两测试分别验证 shutdown 所有权与 Teacher gateway 接线的不同资源 共享 controller fixture 保留 子任务标记为低优先级脚手架共享候选;两测试验证不同资源,共享收益低,不值得为此改动(未作为发现发布)
本地无法执行的 11 个测试路径(缺 torch/numpy/sglang/sglang_router) CPU CI 中跳过或仅由 CI 执行的边界 — 保留 检测力以静态阅读 + 该 head 8 项 GitHub CI 绿色为准;scale-in 重试与 terminal-actor 清理边界(对应 F4/F6)确认有真实生产入口与独立终态 oracle;已记录缺口:无 CPU 级 gateway→admission→client 双真实组件集成测试
Powered by Nyanpasu with glm-5.3-flash max, please check the suggestions carefully.

rai-studio-bot

This comment was marked as resolved.

@ldemon2333 ldemon2333 changed the title feat(inference): unify role inference infrastructure 【No.3】: unify role inference infrastructure Sep 21, 2026
@ldemon2333 ldemon2333 changed the title 【No.3】: unify role inference infrastructure 【No.3】unify role inference infrastructure Sep 21, 2026
rai-studio-bot

This comment was marked as resolved.

rai-studio-bot

This comment was marked as resolved.

@ldemon2333

ldemon2333 commented Sep 21, 2026 •

Copy link
Copy Markdown
Author

@SigureMo 任务已完成,请求 review。可以 approve workflows 吗

rai-studio-bot

This comment was marked as resolved.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已复核 13c0cb7。批量缩容能保留成功结果,但仍有一项多节点部分清理后无法重试的 P2,详见新增行级评论;此前三项问题保持已修复。

本地 50 项 scale-in 测试因缺少 Ray/SGLang 依赖全部跳过。依赖隔离验证覆盖正常部分失败重试,并复现了上述问题;未执行真实多节点或 GPU 验收。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

Comment on lines 4303 to 4305
with self._engine_lifecycle_lock:
for i, _ in live_actors:
for i in released:
group.all_engines[i] = None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 优先级:P2

请保留多节点副本在部分 shutdown 成功后的可重试清理入口。若 node0 成功、follower 暂时失败,这里会留下 all_engines=[None, follower],group 保持 DRAINING;但 _select_engines_for_removal() 只遍历 group.engines(node0 切片),后续请求无法再次选中它。Eviction 同样拒绝 node0 为空的副本,recover 则跳过 DRAINING 弹性组,导致剩余 worker/PG 持续占用,并维持权重更新 fence。

用实际方法的依赖隔离复现确认:即使 follower 随后恢复可正常 shutdown,新缩容请求仍返回 “No engines selected for removal”。建议从副本中任一残留 worker 识别待清理目标,或显式保存待重试的副本状态,同时让重试跳过已清理的 head;补充“head 成功、follower 失败后再次清理”的多节点契约测试。

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

981a275 已修复数量缩容和 eviction 的残留 follower 清理,F4 部分解决;按原 engine_urls 重试仍失败。

_resolve_scale_in_url_candidates() 现在会查询残留 follower,但真实 SGLangEngine.get_url()(relax/backends/sglang/sglang_engine.py:909–914)对非 node0 返回 None,因此无法匹配原 URL。用实际管理器方法并按该契约设置 follower 后,隔离验证仍得到 “No engines selected for removal”;改用数量缩容则可完成清理。

建议保留副本原 URL 与逻辑索引的关联,让 URL 重试也能选中残留 worker,并补充 follower URL 为 None、重复原 URL 请求的回归。当前新增测试只以数量重试,未覆盖此路径。

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

7813d61 已修复原始同 URL 重试,但恢复后地址变化的场景仍有缺口。replica_urls 只在首次注册和 URL 选择时更新;RolloutServer.recover() 重建 head、重新分配端口后没有刷新它。

若 URL 从 A 变为 B,随后按数量缩容或 eviction 清理时 head 成功、follower 失败,按 B 重试会回退到缓存 A,仍报 “No engines selected for removal”。当前实际方法的隔离验证确认:普通同 URL 重试通过,恢复后旧缓存场景失败。

建议在恢复完成后刷新地址,或在所有移除路径清理 head 前保存当前地址,并补充 recovery → 部分清理 → 当前 URL 重试回归。F4 保持部分解决。

@ldemon2333 ldemon2333 Sep 23, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@rai-studio-bot 目前通过 ray-native 扩容申请了多节点 replica 实现是不是有问题啊,具体分析一下,(relax/distributed/ray/rollout.py : _scale_out_ray_native),以每台 8 卡、每副本 16 卡为例)

1797 行算出 actors_per_replica = 2,并为每个副本创建含 16 个 GPU bundle 的 PG。
1888 行按 2 对齐预留 rank:head 拿到偶数 rank,follower 的 rank 也预留了。
2046 行创建 EngineGroup 时写死了 all_engines=[None]。
514 行按 slot 数创建 actor,只创建出 i=0 这一个,也只用到第一台机器的 bundle(523 行)。
这个 actor 的 server 参数是 nnodes=2、node_rank=0、tp_size=16(sglang_engine.py:1740),它会一直等待一个永远不会创建的 node_rank 1。

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

是,这条链路有确定的 slot 数量错误,你指出的位置是关键。我直接检查了当前代码,并用实际 EngineGroup.start_engines() 和地址分配函数做了隔离验证。

按每节点 8 卡、每副本 16 卡、sglang_pp_size=1,假设预留的起始 rank 为 2:

环节 当前实际行为
PG 申请 16 个单 GPU bundle;资源申请本身不会创建 SGLang worker
rank 预留 rank 2、3,但只是在分配编号
_bring_up_single_replica() all_engines=[None] 只有一个 worker slot
start_engines() 只执行 i=0,创建 rank 2,使用排序后的 bundle 0、base GPU 0
server 参数 nnodes=2、node_rank=2 % 2=0、tp_size=16 // 1=16
缺失部分 应有的 rank 3 / node_rank 1 / bundle 8 从未创建或调用 init

根因是把“一个逻辑副本”误当成“一个 Ray actor”。all_engines 存每个节点的 worker,engines 才通过步长 nodes_per_engine 选出逻辑副本的 head。地址分配函数虽然会为 follower rank 写入同一个 dist_init_addr,但它不会创建 follower。现有 GPU 索引校验也只检查不越界,8 个使用槽位小于等于 16,因而拦不住缺 worker。

最小修复方向是在创建 group 时使用与 rank 预留相同的 actor 数:

actors_per_replica = max(1, gpus_per_engine // self.args.num_gpus_per_node)
new_group = EngineGroup(
    ...,
    all_engines=[None] * actors_per_replica,
    ...,
)

在理想的两节点各 8 bundle 布局下,隔离验证结果是:当前单 slot 只创建 (rank=2, bundle=0);两个 slot 则创建 (2, 0) 和 (3, 8),两个 init 使用同一个 rendezvous 地址。仍应让 finalizer 接收 new_group.engines(一个 head),不要把两个 worker 计成两个副本。

“永远等待”需稍作限定:head 缺少 peer,不能正常完成分布式初始化;但 _bring_up_single_replica() 外层 wait_for 有初始化超时,会进入 PROVISION_TIMEOUT 和 rollback,也可能先收到后端初始化异常。是否连 rollback 都卡住是另一项问题,不能仅凭缺 follower 就断言。

建议补一个不替换 EngineGroup 的 8×2-node / 16-GPU scale-out 测试,断言两个 actor、两个 init、不同 node_rank、相同 dist_init_addr,且只发布一个逻辑副本;再在真实两节点上验证初始化、权重同步与失败回滚。当前隔离验证只能证明创建/参数链,不能替代 GPU 验收。

另外,PR merge-base 2a8d2ed 的同一路径也已有 all_engines=[None](当时第 2018 行),所以这是当前代码仍存在的基线缺陷,不是本轮重构新引入。我此前对缩容重试的验证不能覆盖这条实际扩容创建路径。

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@SigureMo 发现一个 bug,可以看一下

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题独立的话,可以提一个 issue,我找时间看一下

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 2101a0f0 复核修复:recover() 在 init 全部成功后立即调用 refresh_replica_urls() 记录重建 head 的最新地址——后续清理部分失败时缓存持有的已是恢复后的 URL B;_resolve_scale_in_url_candidates 优先向存活 head 实时查询 get_url 并回写缓存,仅在 head 不可达时回退缓存;重建前先弹出旧 URL;迟到的旧 head 探测被 all_engines[...] is not head 校验挡住,无法覆盖新地址;URL 获取失败的回滚会丢弃旧地址而不是保留 A。

tests/inference/test_rollout_recovery_urls.py 以真实 RolloutManager 方法覆盖了本条原始场景(head 消失后按恢复后 URL 重试部分清理)及回滚、迟到探测、scaled-out 仅刷新重建 head 等变体。同时 all_engines=[None] * actors_per_replica 修复了本线程中另行确认的多节点扩容缺 follower 基线缺陷,tests/inference/test_multinode_scale_out.py 含拒绝单 slot 旧行为的回归测试。本地 120 项相关测试通过,该 head 8 项 CI 全绿。此项已解决。

rai-studio-bot

This comment was marked as resolved.

rai-studio-bot

This comment was marked as resolved.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已完成 a6e84d8 复核,并独立检查新增静态审查记录。新增两项 P2:终止 actor 的局部清理/恢复无法推进,以及文档误判 Serve 异步析构支持,详见行级评论。原 URL 清理重试问题保持部分解决。

文档中的 P1 结论未全部得到验证:deferred 消费有 commit gate,失败会终止运行;external/custom producer 保留兼容路径。普通请求超时验证仅发送一次 POST,不能据此认定重放。

本轮 client/gateway 测试 52 passed、1 skipped;依赖隔离验证复现 dead actor 的重复恢复失败。未执行真实 Ray actor 故障注入、多节点或 GPU 验收,当前版本暂无 CI checks。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

Comment thread docs_auto/task3/cc.md Outdated
Comment thread relax/distributed/ray/inference_manager.py

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已完成 7813d61 跨模块复核。原 URL 重试问题及终止 actor 清理问题均有进展,但恢复后地址缓存、子进程释放确认仍需补齐;已在原线程补充证据,维持 P2。错误审查文档已删除,对应问题已解决。未发现其他新增问题。

本地定向测试 186 项通过,1 项因缺少 Ray 无法执行;另有管理器/lifecycle 隔离测试 44 项、deferred 子集 11 项通过,缩容隔离验证复现剩余边界。隔离验证不代表真实 Ray、完整 OPD 或多节点/GPU 验收,当前版本暂无 CI checks。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

@SigureMo

Copy link
Copy Markdown
Member

你这模型不行啊,出现幻觉了吧

Failing job: 107152747059

链接打不开啊

以及报错分析能不能从实际报错开始分析?怎么突然跳到了依赖问题呢?麻烦输出点自己的理解哈,不要全依赖 agent 的判断,这样我们会怀疑对代码的理解程度

@rai-studio-bot 你来分析下报错呢?给个建议

# ♻️ Refactor

## Complete the Phase 6 ownership migration

- Move Rollout engine-group ownership, memory lifecycle, discovery observation, recovery, and cleanup into the shared InferenceManager core
- Keep generation, evaluation, data sources, and elastic workload behavior in the Rollout facade and workload layer
- Preserve legacy engine-lock and weight-update interfaces

---

# 🐛 Bug Fix

## Isolate failed elastic groups from production

- Exclude non-ACTIVE scaled groups from memory operations, weight synchronization, discovery markers, and Rollout engine helpers
- Keep DRAINING groups attached for cleanup and placement-group ownership retries
- Let the shared manager reopen admission only after full memory restoration and backend continuation

---

# ✅ Tests

## Verify Rollout ownership and topology filtering

- Cover ACTIVE topology selection, legacy five-tuples, lifecycle rollback, elastic ownership, and workload contracts
# ♻️ Refactor

## Complete the Phase 7 migration

- Connect remaining Rollout, Agentic, Teacher, Autoscaler, and OPD callers to unified discovery and lifecycle paths
- Preserve Native and Agentic Rollout HTTP capacity from the configured per-engine concurrency and replica count
- Remove the MultiEngineManager compatibility module and the GenRM-specific SGLangEngine subclass

## Remove legacy Python import paths

- Require callers and fresh Ray workers to import InferenceManager and SGLangEngine directly
- Preserve runtime role APIs and schema compatibility while dropping obsolete implementation aliases

---

# 🐛 Bug Fix

## Retry partial logical replica cleanup

- Identify logical replicas through any surviving worker after the head has stopped
- Keep partial DRAINING replicas selectable by target count, URL, and eviction
- Retry only residual actors while preserving confirmed cleanup results

---

# ✅ Tests

## Verify final migration and cleanup boundaries

- Cover managed Teacher clients, loop-local transport cache identity, shutdown ownership, old interface adapters, and placement rejection
- Cover multi-node follower cleanup retries through scale-in and graceful eviction

---

# ✅ Tests

## Keep CPU CI collectible without SGLang

- Skip real Teacher and SGLang imports before collection enters the optional backend
- Limit runtime skips to dependent tests and fixtures while retaining pure CPU coverage

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ec350f6 本轮增量未发现新增问题:已核对权重版本查询、发布失败的状态更新及异常传播、路由异常的 Ray 序列化。/model_info 提供所需字段,与固定 SGLang 版本兼容(v0.5.17、v0.5.15)。

三个定向测试文件 85 passed,覆盖查询失败、未发布版本、版本 0、失败后重试和异常序列化;后端测试使用替身,未执行真实 SGLang/GPU 或多节点验收。当前提交尚无 CI 结果。

整体仍有 F4:恢复后 URL 重试 与 F6:actor 死亡后的子进程清理确认 部分解决,维持补充意见,不作通过结论。F7 保持已解决。

Powered by Nyanpasu with gpt-6-astra medium, please check the suggestions carefully.

rai-studio-bot

This comment was marked as resolved.

rai-studio-bot

This comment was marked as resolved.

rai-studio-bot

This comment was marked as resolved.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复核 2785a46c:8 项精简建议全部落地,新增枚举重构无问题

3a563718(Add Role State Enum Class)与 2785a46c(Simplify code)已逐项复核:

  • F10–F17 全部按建议实现并验证(死链/死模型/无引用符号删除、共享 genrm_ray_gpu_fraction helper 并新增直接执行真实 manager 方法的新测试、max_samples 透传与容量测试、phase 字段删除、共享测试替身模块),已在各原始线程逐项回复确认。改动文件编译通过、全树无残留引用;本地可收集的 CPU 测试 187 + 14 项通过(失败/报错均为基线既有的本地缺依赖问题:torch/numpy/sglang_router;11 项新 placement 用例在 CI 依赖集下可执行)。
  • RoleState 枚举重构为纯机械替换:状态集合不变,get_state 导出与错误信息经 .value 保持纯字符串契约(新测试显式断言导出类型为 str),lifecycle coordinator 测试全部通过。
  • F4/F6 相关代码(inference_manager.py、distributed/ray/rollout.py)本轮未改动,维持部分解决;F9 待更新(描述提交范围与 13 个文件链接仍指向 1054831/7813d61,描述快照示例仍含已删除的 phase/phase_epoch 字段)。
  • 当前 head 推送后 CI 尚无结果(排队中),以上为本地验证结论。

当前待处理仅剩:F4/F6(P2,部分解决)与 F9(P3,描述更新)。审查看板已同步。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

@Yangruipis

Copy link
Copy Markdown
Collaborator

感谢贡献,我近期会开始review

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复核 2101a0f0(fix bug):F4/F6 已解决,全部代码类问题清零,该 head CI 8/8 绿

  • F4(恢复后 URL 重试)已解决:recover() 成功后立即 refresh_replica_urls() 记录重建地址;scale-in 选路优先实时查询存活 head 并回写缓存;重建前弹出旧 URL,迟到探测与回滚均不会保留旧地址。本线程中另行确认的多节点扩容缺 follower 基线缺陷也已修复(all_engines=[None] * actors_per_replica)并附拒绝旧行为的回归测试。
  • F6(终止 actor 后的资源确认)已解决:新增 EngineShutdownGuard,终端 actor 死亡不再被计入 shutdown 确认,而是围栏全部生命周期入口并抛出 InferenceRecoveryRequired;Controller 重启前逐一确认 rollout/teacher/genrm 后端清理结果,失败即阻止自动重启并给出面向操作者的说明(明确 Ray shutdown 与固定等待不构成资源释放确认);health monitor 复用组清理路径并登记 guard。测试含真实 Ray 子进程 crash 场景与真实 owned PG 的保留断言。
  • 附带收益:新增 gateway→admission 双真实组件集成测试(补上此前审计记录的 CPU 级集成缺口);dispose 拒绝在途扩容请求。
  • 验证:本地新增/更新测试 120 项通过(1 项依赖 psutil 的子进程用例在 CI 依赖集下通过);该 head 全部 8 项 GitHub CI 成功。
  • F9(P3,描述)部分解决:过期提交范围行与已删除的 phase 字段示例已清理,但 13 个 Key implementation files 链接仍指向 7813d61(rebase 前且已落后三个 head,如 genrm.py/inference_manager.py 其后均有实质改动),建议改指当前 head 2101a0f0。

至此代码层面的审查问题全部解决;唯一待处理项为 F9 的描述链接更新。审查看板已同步。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

审查结论:通过 —— 17 项审查发现全部解决或取代

首轮审查中的 F9(描述过期提交引用)已随本次描述更新全部落实:13 个 Key implementation files 链接已指向当前 head 2101a0f0,过期提交范围行与已删除的 phase 字段示例均已清理,代码量统计与当前 diff 完全一致(91 文件、+14067/−1743)。

至此本轮审查的全部 17 项发现(F1–F17,含早期 P1 并发回归、两项曾部分解决的 P2)均在 2101a0f0 上修复并逐项验证;F4 线程中另行确认的多节点扩容缺 follower 基线缺陷也已修复。生产与测试必要性审计完成且决策均已落地;该 head 全部 8 项 GitHub CI 通过。剩余边界仅作者显式声明 SKIP 的多节点真实硬件验收(单节点 4×B200 实测已在描述中记录),供维护者评审时权衡。新增的真实路径回归测试(真实 GenRM manager 的 preflight 一致性、真实 Ray 子进程 crash 场景、gateway→admission 双真实组件集成)显著提升了验证质量。

完整发现索引与各阶段结论见审查看板。

Powered by Nyanpasu with glm-5.3[1m] xhigh, please check the suggestions carefully.

@SigureMo

SigureMo commented Oct 8, 2026

Copy link
Copy Markdown
Member

辛苦解决下冲突~

# 🔩 Chore

## Merge redai-studio/main into feat/task3

- Merge main at 815de3a
- Preserve deferred inference commit gates alongside offline DPO and RM actor behavior
- Regenerate OpenAPI specifications from the merged service routes

---

# 🐛 Bug Fix

## Reject deferred scoring for offline workloads

- Apply the offline workload guard consistently to SFT, DPO, and RM
- Prevent offline training from waiting for a rollout commit coordinator

---

# ✅ Tests

## Cover the merged deferred workload validation

- Verify SFT, DPO, and RM are rejected by deferred scoring preflight
- Run focused actor and deferred inference tests
- Run all pre-commit hooks including gitleaks
@ldemon2333 ldemon2333 closed this Oct 8, 2026
@ldemon2333 ldemon2333 reopened this Oct 8, 2026
@rai-studio-bot

rai-studio-bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Thanks for contributing to Relax, @ldemon2333! 感谢你为 Relax 做出贡献!

📚 Contribution guide / 贡献指南

Describe the problem, your changes, and how you validated them. Keep each PR focused and run pre-commit run --all-files before submitting.

请说明问题、改动和验证方式,保持 PR 聚焦,并在提交前运行 pre-commit run --all-files。

English contribution guide · 中文贡献指南

🛠️ CI commands / CI 指令
Command / 指令 Usage / 用途
/rerun Retry failed CI / 重跑失败的 CI
/rerun <target> Rerun a workflow or check / 重跑指定 workflow 或检查
/cancel <workflow> Cancel an entire workflow / 取消整个 workflow
/help Show commands and targets / 查看指令和 target
/review Request a code review / 请求代码 review

Put one command on the first line of a new PR comment. Rerun/cancel require PR authorship or repository write access.

在新 PR 评论的首行写一条指令。PR 作者或有仓库写权限的贡献者可以重跑、取消 CI。

CI usage and targets · CI 用法与 target

🔎 Merge requirements / 合入条件

GitHub:⏳ 合入条件未满足

Requirement / 条件 Status / 状态
审批人数 ❌ 0/1 · 至少 1 位有仓库 write 或更高权限的人 Approve · 处理修改意见后请重新审批:Yangruipis
团队审批(relax-block-reviewers) ⏳ 0/1 · 可联系:Aurelius84, Yangruipis, yuanlehome, SigureMo, NINGBENZHE
CI ✅ 已通过

@ldemon2333

Copy link
Copy Markdown
Author

辛苦解决下冲突~
已解决

@Yangruipis Yangruipis left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整体问题不大,在 架构统一、RFC 覆盖、运行正确性、抽象上做的都比较好,但是改动量和侵入性相对较大。
一些小问题 comment 了,后期如果通过 review,建议按模块拆分提交和验证,我们内部也会在多机大规模下进行nightly任务验证

Comment thread relax/inference/defer.py Outdated
raise ValueError("Deferred Teacher requires SGLang OPD")
if "genrm" in roles:
models = getattr(args, "_genrm_instances_resolved", {})
if len(models) != 1 or getattr(args, "rm_type", None) != "dapo-genrm" or getattr(args, "custom_rm_path", None):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里和业务逻辑dapo-genrm耦合了

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

新增 commit:b36a 新增 supports_deferred 相关契约,解耦 genrm role 与特定 reward adapter 计算逻辑

Comment thread relax/core/controller.py
serve.run(deployment, name="metrics", route_prefix="/metrics")
logger.info("MetricsService deployed at /metrics")

def _deploy_teacher_gateway(self) -> None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

当前 rollout 和 genrm 都还是各自的 deployment,只有 teacher 用的 InferenceGateway ,这个是短期形态吗,后面会统一嘛?

@ldemon2333 ldemon2333 Oct 10, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

目前这种设计是兼容旧接口的一种过渡形态。Teacher 直接使用通用 InferenceGateway,主要是为了在保留现有 OPD Teacher 启动方式、opd_teacher_url(s) 直连兼容路径以及 Controller/Actor 对 TeacherManager 生命周期控制的前提下,最小成本接入统一的 schema v2 discovery 和 routing contract。并且目前的 TeacherManager 已覆盖单模型的 engine 生命周期;多模型编排目前由 OPD helper、Controller 和 Gateway 分散承担,目前实现中并没有像 genrm serve deployment 一样,使用一个专门的 teacher serve deployment 来做多 teacher model 编排。

我觉得目前 genrm 与 teacher role 本身都是作为被动的 serve deployment,两者架构高度相似,可以进行统一

一个 inference role
│
├── 一个 Serve Component
│   │
│   └── managers: dict[model_id, Manager]
│       │
│       ├── model_id=A → Manager A → replicas A
│       ├── model_id=B → Manager B → replicas B
│       └── model_id=C → Manager C → replicas C
│
└── 一个 role 级 InferenceGateway
    │
    ├── 聚合所有 Manager 的 schema v2 snapshot
    ├── route_key → model_id
    ├── 从目标模型选择 READY replica
    └── 转发推理请求

GenRM 、Teacher 、 rollout 最终的结构:

Rollout role
├── Rollout Service
│   ├── 主动 rollout workload
│   ├── DataSource / TransferQueue
│   ├── weight sync / evaluation / scaling
│   └── 一个多模型 RolloutManager
└── InferenceGateway(role="rollout")


GenRM role
├── GenRM Service
│   ├── tokenizer / chat-template renderer
│   └── dict[model_id, GenRMManager]
│       └── 每个 Manager 管理一个模型的多个 replicas
└── InferenceGateway(role="genrm")


Teacher role
├── Teacher Service
│   └── dict[model_id, TeacherManager]
│       └── 每个 Manager 管理一个模型的多个 replicas
└── InferenceGateway(role="teacher")

finally:
self._inference_clients.pop(session, None)
if client is not None:
await client.aclose()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] 按样本重建 Teacher client 使 discovery 缓存失效

native Rollout 按样本调用 prefill,而这里每次创建并关闭 InferenceClient,因此 Teacher discovery 的 TTL 缓存无法跨样本复用。发现请求又进入串行锁和副本健康 RPC,会增加高并发 OPD 的控制面开销。建议按事件循环复用 client,在对应生命周期结束时关闭,并增加连续两次 prefill 的 discovery 请求次数测试。已确认调用路径,尚未实测吞吐下降幅度。

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

修复,见 43e1 commit

Comment thread relax/core/controller.py
raise ValueError(
"Use --inference-defer-roles with a framework scoring adapter; legacy hooks cannot own this plan"
)
validate_deferred_workload(self.config)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] 新 defer 校验使仓库现有示例无法启动

新的 preflight 要求 inference_defer_roles,但 examples/generate_reward_model/run-qwen35-35B-A3B-16xgpu-genrm-397B-defer.sh 仍使用旧 defer flag、rm-type dummy 和 custom post-process hook。该示例会被这里拒绝;只补新 flag 又会被 validate_deferred_workload 的 hook 和 reward adapter 限制拒绝。请同步迁移示例及 README 到框架评分适配器,并验证整套参数通过预检,保留现有安全限制。

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

修复,见 4bf9 commit

# ♻️ Refactor

## Declare deferred reward capability

- Add an explicit deferred capability to registered reward specifications
- Resolve deferred GenRM scoring through the reward registry
- Remove the inference scheduler dependency on the DAPO reward implementation

---

# ✅ Tests

## Cover capability-based deferred routing

- Accept async reward adapters that declare deferred support regardless of name
- Reject async adapters without the deferred capability
- Exercise deferred scoring through the registered handler
@ldemon2333
ldemon2333 requested a review from a team October 10, 2026 12:42

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

复核结论:b36aac4b 增量审查完成,无新增问题 —— 有条件等待两项未回应意见

对上一已审 head 2101a0f0 之后的增量完成复核:合并 bb2aa11(同步 main 815de3aa)经 git diff-tree --cc 确认的全部 11 个冲突文件逐一核对,PR 侧的 defer 门控、discovery 与 placement 内容和 main 侧的 offline 模式、DPO/RM、设备模块化变更均完整保留,defer.py 离线拒绝切换为 is_offline_mode 后覆盖 sft/dpo/rm 并有对应参数化测试;b36aac4b 按审查意见把 deferred GenRM 评分从 dapo_genrm 硬编码导入改为 reward registry 的 supports_deferred 能力契约,preflight 会拒绝未声明该能力的 async 适配器(如 remote_rm),评分入口仅在 preflight 通过的 defer 模式可达。本地 tests/inference/test_defer_wiring.py 38 项、rewards/legacy 161 项通过,该 head 全部 CI 通过。本轮无新增发现,既往 17 项发现保持已解决或已取代。

非行级事项(待作者处理,非本审查新发现):维护者审查中的两个 P2 意见截至当前 head 仍无代码回应,我在当前代码上复核后仍然成立——按样本重建 Teacher client 使 discovery 缓存失效(on_policy_distillation.py 未改动),以及新 defer 校验拒绝仓库现有示例(示例仍使用旧 defer flag、--rm-type dummy 与 custom post-process hook,placement.py 仍将其映射为 defer 计划)。审查线程 1 的解耦修复已验证,线程 2 的部署形态问题作者已给出设计与后续统一方案。

完整发现索引、提交范围(94 个文件全部接收)与各阶段结论见审查看板。

Powered by Nyanpasu with glm-5.3-flash max, please check the suggestions carefully.

# 🐛 Bug Fix

## Use framework-managed deferred scoring

- Replace the legacy custom post-process lifecycle with the registered dapo-genrm adapter
- Remove unsupported dynamic batching and in-run evaluation from the deferred recipe
- Keep the shared 16-GPU Rollout and GenRM placement

---

# 📝 Documentation

## Update GenRM defer guidance

- Document the explicit deferred role and adapter configuration in English and Chinese
- Describe the closed-batch commit flow and current evaluation limitation

---

# ✅ Tests

## Cover the complete recipe preflight

- Verify the shared placement resolves to defer mode on 16 GPUs
- Verify the migrated arguments pass deferred workload validation
# ⚡ Performance

## Reuse event-loop discovery clients

- Route managed Teacher requests through the existing loop-local inference client cache
- Preserve model routing and group affinity while allowing discovery TTL reuse
- Keep the legacy external Teacher transport unchanged

---

# ✅ Tests

## Cover discovery reuse and cleanup

- Verify concurrent prefill calls share one inference client and one discovery request
- Verify both Teacher scoring requests still execute
- Close cached clients explicitly at the event-loop lifecycle boundary
@ldemon2333
ldemon2333 requested a review from yxyOo as a code owner October 10, 2026 13:18

@rai-studio-bot rai-studio-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

结论:通过(APPROVE)

上轮遗留的两项 P2 已在 head 43e1a21b 修复并核实:

  1. Teacher client 按样本重建使 discovery 缓存失效:43e1a21b 改用 generate_with_discovery 的事件循环级缓存复用;route_key 经 payload 传递、affinity 经 X-SMG-Routing-Key 头逐请求转发,请求语义不变;客户端在 shutdown_async_loop 关闭,并有并发 prefill 复用同一 discovery client 的测试覆盖。
  2. defer 预检拒绝现有示例:4bf91848 将 defer 示例与 README/中英文档迁移到框架评分适配器(--rm-type dapo-genrm + --inference-defer-roles genrm),移除了预检拒绝的 dummy rm、legacy hook、--eval-interval 与 --use-dynamic-batch-size。已在本地以脚本真实参数复跑参数解析、校验、plan_inference_placement(mode=defer,16 GPU)与 validate_deferred_workload,整套参数通过预检;46 项相关测试通过。

既往 17 项发现保持已解决或已取代,本轮无新增发现,CI 在该 head 全绿。边界说明:defer 示例完成的是预检级(解析/校验/placement 计划)验证,多机真实执行仍属作者声明的 SKIP,由内部 nightly 覆盖;维护者线程的解除由维护者确认。

详细范围、阶段结论与发现索引见审查看板。

Powered by Nyanpasu with glm-5.3-flash max, please check the suggestions carefully.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants