Field
Value
Status
Draft
Updated
2026-07-17
Scope
Rollout, GenRM, SGLang Teacher
Implementation
Not started
Everything below is proposed unless marked Current .
Decision
Topic
Proposal
API
One CPU-only InferenceGateway class, one instance per role
Engine lifecycle
One InferenceManager class for all roles
SGLang
One SGLangEngine; remove the GenRM subclass
Routing
One model-routing spec shared by gateway and direct clients
Placement
Resolve and validate all GPU bundles before engine startup
Colocation
Support split and defer
Shared co-resident
Reject in the first version
Migration
Start with discovery and client; keep current managers
Why change
Current
Role
API
Lifecycle
Rollout
relax/components/rollout.py
relax/distributed/ray/rollout.py
GenRM
relax/components/genrm.py
relax/distributed/ray/genrm.py
Teacher
raw /generate URLs
relax/distributed/ray/teacher_manager.py
relax/backends/sglang/sglang_engine.py contains both SGLangEngine and GenRMEngine.
flowchart LR
RC[Rollout callers] --> RS[Rollout Serve] --> RM[RolloutManager]
GC[Reward callers] --> GS[GenRM Serve] --> GM[GenRMManager]
TC[OPD callers] --> TU[Teacher URLs] --> TM[TeacherManager]
N[Three API paths<br/>Three lifecycle paths<br/>Three placement paths]
RM -.-> N
GM -.-> N
TM -.-> N
Loading
Main problems:
Problem
Result
Lifecycle code is copied
Fixes land in one role but not the others
Managers calculate role-specific GPU offsets
Invalid layouts fail late
APIs expose different schemas
Callers need role-specific clients
Global args leak into static models
Teacher/GenRM may inherit Rollout-only behavior
Defer is scripted in user code
Resource safety depends on example code
The current GenRM defer flow is implemented in
examples/generate_reward_model/post_process_genrm_swap.py and depends on
the named actor relax_genrm_manager.
Target architecture
flowchart TB
Caller[Training / Rollout / Reward / OPD]
subgraph Control[Control plane - CPU]
Gateway[InferenceGateway<br/>role = rollout / genrm / teacher]
Manager[InferenceManager]
Planner[PlacementPlanner]
Coordinator[LifecycleCoordinator]
end
subgraph Data[Data plane - GPU]
A[Model A<br/>Engine groups / replicas]
B[Model B<br/>Engine groups / replicas]
end
Caller -->|gateway mode| Gateway
Caller -.->|direct mode| A
Gateway -->|route + state| Manager
Planner -->|resolved placement| Manager
Coordinator -->|activate / drain / deactivate| Manager
Manager --> A
Manager --> B
Gateway -.->|GET /role/engines| Caller
Loading
Component
Owns
InferenceGateway
HTTP API, compatibility adapters, routing, proxying, discovery
InferenceManager
Engine pools, health, endpoints, recovery, onload/offload, shutdown
PlacementPlanner
Bundle allocation, topology checks, PG ownership
LifecycleCoordinator
Phase order, drain, activation lease, barriers
RolloutWorkload
Generate/eval/TQ/data flow; no engine lifecycle
The gateway must not use an inference GPU placement group. It stays available
while engines are sleeping, draining, or restarting.
Internal model (proposed)
This is an internal representation, not a new YAML file or CLI contract.
Role -> Model -> Engine Group -> Replica
@dataclass (frozen = True )
class InferenceRoleSpec :
role : str
workload : WorkloadType
deployment : DeploymentSpec
routing : RoutingSpec
models : list [InferenceModelSpec ]
expose_engine_urls : bool = True
@dataclass (frozen = True )
class InferenceModelSpec :
name : str
model_path : str
engine_groups : list [EngineGroupSpec ]
weight_source : WeightSource
route_mode : RouteMode
sampling_defaults : dict [str , object ]
chat_template_kwargs : dict [str , object ]
env_vars : dict [str , str ]
elastic_enabled : bool = False
fault_tolerance_enabled : bool = False
@dataclass (frozen = True )
class ModelPlacement :
pg_owner : Ownership
bundle_indices : tuple [int , ...]
gpu_ids : tuple [int , ...]
activation_group : str | None
activation_phase : str | None
Derived capabilities:
needs_weight_update = weight_source != WeightSource .STATIC
needs_dcs = weight_source == WeightSource .DCS
needs_router = route_mode == RouteMode .SGLANG_ROUTER
The planner validates capacity, parallel layout, multi-node boundaries, split
overlap, defer mutual exclusion, and PG ownership before engine startup.
API
Method
Endpoint
Purpose
GET
/<role>/engines
Discovery and lifecycle state
GET
/<role>/health
Gateway and model health
GET
/<role>/v1/models
OpenAI-compatible model list
POST
/<role>/generate
Raw SGLang payload
POST
/<role>/chat/completions
Relax alias
POST
/<role>/v1/chat/completions
OpenAI-compatible chat
Example discovery response:
{
"role" : " teacher" ,
"topology_revision" : 7 ,
"phase" : " deferred_score" ,
"models" : {
"math-teacher" : {
"state" : " ready" ,
"router_url" : null ,
"engines" : [
{
"engine_id" : " math-teacher/0" ,
"base_url" : " http://node-1:15000" ,
"state" : " ready" ,
"direct_eligible" : true
}
]
}
}
}
Routing order:
explicit model
-> route_key mapping
-> default model
-> 400
selected model
-> SGLang router
or
-> READY && direct_eligible replica
Multi-node TP/PP workers form one logical replica. Discovery exposes only
the ingress/head URL.
PD prefill/decode workers are diagnostic only. Requests use the router.
topology_revision changes after scale, restart, recovery, or endpoint
changes.
Sleeping and draining engines return direct_eligible: false.
Compatibility
Current behavior
Migration behavior
GenRM /generate accepts messages
Keep adapter and {"response": "..."}
Rollout exposes /v1/chat/completions
Keep endpoint
Rollout exposes /engines
Evolve response toward common schema
OPD accepts Teacher engine URLs
Allow raw engine or gateway URL
Existing CLI flags
Convert to internal spec; no first-phase CLI rewrite
Placement modes
Mode
GPU relationship
Concurrency
Decoupled
Separate PG per role
Training and inference may overlap
Colocated split
Disjoint slices in Actor PG
Inference roles may overlap
Colocated defer
Same bundles, different phases
Roles are serialized
Shared co-resident
Same bundles, same phase
Rejected
flowchart LR
subgraph Decoupled
DA[Actor PG]
DR[Rollout PG]
DG[GenRM PG]
DT[Teacher PG]
end
subgraph Split[Colocated split - Actor PG]
SR[GPU 0..7<br/>Rollout]
SG[GPU 8..11<br/>GenRM]
ST[GPU 12..15<br/>Teacher]
end
subgraph Defer[Colocated defer]
PA[Phase A<br/>Rollout READY]
PB[Phase B<br/>Scorer READY]
PC[Phase C<br/>Actor READY]
PA --> PB --> PC --> PA
end
Loading
PG ownership:
Layout
Owner
Who deletes the PG
Decoupled + managed
Manager
Manager
Decoupled + external
External service
External service
Colocated
Controller
Controller
Lifecycle
stateDiagram-v2
[*] --> STARTING
STARTING --> READY
READY --> DRAINING
DRAINING --> SLEEPING
SLEEPING --> ONLOADING
ONLOADING --> READY
STARTING --> DEAD
READY --> DEAD
DRAINING --> DEAD
SLEEPING --> DEAD
Loading
Manager operations are idempotent:
activate (tags = None )
drain ()
deactivate ()
shutdown ()
Topology publication is atomic:
create candidates
-> wait healthy
-> sync weights if required
-> update router
-> publish registry snapshot
-> topology_revision += 1
Defer transition:
sequenceDiagram
participant Coordinator
participant Rollout
participant Registry
participant Scorer
participant Trainer
Coordinator->>Registry: Rollout = DRAINING
Coordinator->>Rollout: stop new work + drain
Coordinator->>Rollout: flush + offload
Coordinator->>Scorer: onload
Coordinator->>Registry: Scorer = READY
Coordinator->>Scorer: batch score
Coordinator->>Scorer: drain + offload
Coordinator->>Trainer: barrier + wake
Loading
The gateway never wakes a sleeping model because a request arrived. It returns
503 with retry guidance.
Deferred OPD
Teacher defer requires a data-flow change, not just different placement.
sequenceDiagram
participant Rollout
participant Samples
participant Teacher
participant Trainer
Rollout->>Samples: samples + teacher inputs + token selection
Note right of Rollout: no inline Teacher request
Rollout->>Rollout: drain + offload
Teacher->>Teacher: onload
Teacher->>Samples: batch prefill / logprobs
Teacher->>Teacher: drain + offload
Samples->>Trainer: samples with teacher logprobs
Trainer->>Trainer: compute OPD loss
Loading
Teacher defer is incomplete until the write-back path is implemented.
Migration
Phase
Deliverable
Public behavior
0
Characterize args, placement, endpoints, lifecycle, DCS/router behavior
No change
1
Common discovery schema, topology revision, common client
Additive
2
One gateway class, multi-model routing, legacy adapters
Compatible
3
One manager and one SGLang engine path
Facades remain
4
Placement planner, decoupled and split
Reject invalid layouts
5
Lifecycle coordinator, GenRM defer, deferred OPD
New framework pipeline
6
Move Rollout engine ownership to unified manager
Workload stays separate
7
Remove old managers, subclass, and named-actor special cases
Breaking cleanup
Start with Phase 0 and Phase 1. Changes to argument parsing, Controller,
Service, Launcher, or public API deletion require separate approval.
Acceptance checks
Open questions
#
Question
1
Keep adapting existing arguments, or design a unified user-facing config in a separate RFC?
2
Should direct clients fall back to the gateway by default?
3
How is external discovery authenticated?
4
Configured or planner-optimized defer sub-phase order?
5
Elastic scale for Rollout only, or also GenRM/Teacher?
6
How long do compatibility facades remain?
Proposed review outcome
Accept architecture
+ approve Phase 0/1
+ defer unified CLI design
+ defer Controller/Service cleanup
+ reject shared co-resident in v1
Decision
InferenceGatewayclass, one instance per roleInferenceManagerclass for all rolesSGLangEngine; remove the GenRM subclasssplitanddeferWhy change
Current
relax/components/rollout.pyrelax/distributed/ray/rollout.pyrelax/components/genrm.pyrelax/distributed/ray/genrm.py/generateURLsrelax/distributed/ray/teacher_manager.pyrelax/backends/sglang/sglang_engine.pycontains bothSGLangEngineandGenRMEngine.flowchart LR RC[Rollout callers] --> RS[Rollout Serve] --> RM[RolloutManager] GC[Reward callers] --> GS[GenRM Serve] --> GM[GenRMManager] TC[OPD callers] --> TU[Teacher URLs] --> TM[TeacherManager] N[Three API paths<br/>Three lifecycle paths<br/>Three placement paths] RM -.-> N GM -.-> N TM -.-> NMain problems:
The current GenRM defer flow is implemented in
examples/generate_reward_model/post_process_genrm_swap.pyand depends onthe named actor
relax_genrm_manager.Target architecture
flowchart TB Caller[Training / Rollout / Reward / OPD] subgraph Control[Control plane - CPU] Gateway[InferenceGateway<br/>role = rollout / genrm / teacher] Manager[InferenceManager] Planner[PlacementPlanner] Coordinator[LifecycleCoordinator] end subgraph Data[Data plane - GPU] A[Model A<br/>Engine groups / replicas] B[Model B<br/>Engine groups / replicas] end Caller -->|gateway mode| Gateway Caller -.->|direct mode| A Gateway -->|route + state| Manager Planner -->|resolved placement| Manager Coordinator -->|activate / drain / deactivate| Manager Manager --> A Manager --> B Gateway -.->|GET /role/engines| CallerInferenceGatewayInferenceManagerPlacementPlannerLifecycleCoordinatorRolloutWorkloadThe gateway must not use an inference GPU placement group. It stays available
while engines are sleeping, draining, or restarting.
Internal model (proposed)
This is an internal representation, not a new YAML file or CLI contract.
Derived capabilities:
The planner validates capacity, parallel layout, multi-node boundaries, split
overlap, defer mutual exclusion, and PG ownership before engine startup.
API
/<role>/engines/<role>/health/<role>/v1/models/<role>/generate/<role>/chat/completions/<role>/v1/chat/completionsExample discovery response:
{ "role": "teacher", "topology_revision": 7, "phase": "deferred_score", "models": { "math-teacher": { "state": "ready", "router_url": null, "engines": [ { "engine_id": "math-teacher/0", "base_url": "http://node-1:15000", "state": "ready", "direct_eligible": true } ] } } }Routing order:
the ingress/head URL.
topology_revisionchanges after scale, restart, recovery, or endpointchanges.
direct_eligible: false.Compatibility
/generateaccepts messages{"response": "..."}/v1/chat/completions/enginesPlacement modes
flowchart LR subgraph Decoupled DA[Actor PG] DR[Rollout PG] DG[GenRM PG] DT[Teacher PG] end subgraph Split[Colocated split - Actor PG] SR[GPU 0..7<br/>Rollout] SG[GPU 8..11<br/>GenRM] ST[GPU 12..15<br/>Teacher] end subgraph Defer[Colocated defer] PA[Phase A<br/>Rollout READY] PB[Phase B<br/>Scorer READY] PC[Phase C<br/>Actor READY] PA --> PB --> PC --> PA endPG ownership:
Lifecycle
stateDiagram-v2 [*] --> STARTING STARTING --> READY READY --> DRAINING DRAINING --> SLEEPING SLEEPING --> ONLOADING ONLOADING --> READY STARTING --> DEAD READY --> DEAD DRAINING --> DEAD SLEEPING --> DEADManager operations are idempotent:
Topology publication is atomic:
Defer transition:
sequenceDiagram participant Coordinator participant Rollout participant Registry participant Scorer participant Trainer Coordinator->>Registry: Rollout = DRAINING Coordinator->>Rollout: stop new work + drain Coordinator->>Rollout: flush + offload Coordinator->>Scorer: onload Coordinator->>Registry: Scorer = READY Coordinator->>Scorer: batch score Coordinator->>Scorer: drain + offload Coordinator->>Trainer: barrier + wakeThe gateway never wakes a sleeping model because a request arrived. It returns
503with retry guidance.Deferred OPD
Teacher defer requires a data-flow change, not just different placement.
sequenceDiagram participant Rollout participant Samples participant Teacher participant Trainer Rollout->>Samples: samples + teacher inputs + token selection Note right of Rollout: no inline Teacher request Rollout->>Rollout: drain + offload Teacher->>Teacher: onload Teacher->>Samples: batch prefill / logprobs Teacher->>Teacher: drain + offload Samples->>Trainer: samples with teacher logprobs Trainer->>Trainer: compute OPD lossTeacher defer is incomplete until the write-back path is implemented.
Migration
Start with Phase 0 and Phase 1. Changes to argument parsing, Controller,
Service, Launcher, or public API deletion require separate approval.
Acceptance checks
Open questions
Proposed review outcome