Skip to content

[RFC] Unified Inference Service #71

Description

@Yangruipis
Field Value
Status Draft
Updated 2026-07-17
Scope Rollout, GenRM, SGLang Teacher
Implementation Not started

Everything below is proposed unless marked Current.

Decision

Topic Proposal
API One CPU-only InferenceGateway class, one instance per role
Engine lifecycle One InferenceManager class for all roles
SGLang One SGLangEngine; remove the GenRM subclass
Routing One model-routing spec shared by gateway and direct clients
Placement Resolve and validate all GPU bundles before engine startup
Colocation Support split and defer
Shared co-resident Reject in the first version
Migration Start with discovery and client; keep current managers

Why change

Current

Role API Lifecycle
Rollout relax/components/rollout.py relax/distributed/ray/rollout.py
GenRM relax/components/genrm.py relax/distributed/ray/genrm.py
Teacher raw /generate URLs relax/distributed/ray/teacher_manager.py

relax/backends/sglang/sglang_engine.py contains both SGLangEngine and GenRMEngine.

flowchart LR
    RC[Rollout callers] --> RS[Rollout Serve] --> RM[RolloutManager]
    GC[Reward callers] --> GS[GenRM Serve] --> GM[GenRMManager]
    TC[OPD callers] --> TU[Teacher URLs] --> TM[TeacherManager]

    N[Three API paths<br/>Three lifecycle paths<br/>Three placement paths]
    RM -.-> N
    GM -.-> N
    TM -.-> N
Loading

Main problems:

Problem Result
Lifecycle code is copied Fixes land in one role but not the others
Managers calculate role-specific GPU offsets Invalid layouts fail late
APIs expose different schemas Callers need role-specific clients
Global args leak into static models Teacher/GenRM may inherit Rollout-only behavior
Defer is scripted in user code Resource safety depends on example code

The current GenRM defer flow is implemented in
examples/generate_reward_model/post_process_genrm_swap.py and depends on
the named actor relax_genrm_manager.

Target architecture

flowchart TB
    Caller[Training / Rollout / Reward / OPD]

    subgraph Control[Control plane - CPU]
        Gateway[InferenceGateway<br/>role = rollout / genrm / teacher]
        Manager[InferenceManager]
        Planner[PlacementPlanner]
        Coordinator[LifecycleCoordinator]
    end

    subgraph Data[Data plane - GPU]
        A[Model A<br/>Engine groups / replicas]
        B[Model B<br/>Engine groups / replicas]
    end

    Caller -->|gateway mode| Gateway
    Caller -.->|direct mode| A
    Gateway -->|route + state| Manager
    Planner -->|resolved placement| Manager
    Coordinator -->|activate / drain / deactivate| Manager
    Manager --> A
    Manager --> B
    Gateway -.->|GET /role/engines| Caller
Loading
Component Owns
InferenceGateway HTTP API, compatibility adapters, routing, proxying, discovery
InferenceManager Engine pools, health, endpoints, recovery, onload/offload, shutdown
PlacementPlanner Bundle allocation, topology checks, PG ownership
LifecycleCoordinator Phase order, drain, activation lease, barriers
RolloutWorkload Generate/eval/TQ/data flow; no engine lifecycle

The gateway must not use an inference GPU placement group. It stays available
while engines are sleeping, draining, or restarting.

Internal model (proposed)

This is an internal representation, not a new YAML file or CLI contract.

Role -> Model -> Engine Group -> Replica
@dataclass(frozen=True)
class InferenceRoleSpec:
    role: str
    workload: WorkloadType
    deployment: DeploymentSpec
    routing: RoutingSpec
    models: list[InferenceModelSpec]
    expose_engine_urls: bool = True


@dataclass(frozen=True)
class InferenceModelSpec:
    name: str
    model_path: str
    engine_groups: list[EngineGroupSpec]
    weight_source: WeightSource
    route_mode: RouteMode
    sampling_defaults: dict[str, object]
    chat_template_kwargs: dict[str, object]
    env_vars: dict[str, str]
    elastic_enabled: bool = False
    fault_tolerance_enabled: bool = False


@dataclass(frozen=True)
class ModelPlacement:
    pg_owner: Ownership
    bundle_indices: tuple[int, ...]
    gpu_ids: tuple[int, ...]
    activation_group: str | None
    activation_phase: str | None

Derived capabilities:

needs_weight_update = weight_source != WeightSource.STATIC
needs_dcs = weight_source == WeightSource.DCS
needs_router = route_mode == RouteMode.SGLANG_ROUTER

The planner validates capacity, parallel layout, multi-node boundaries, split
overlap, defer mutual exclusion, and PG ownership before engine startup.

API

Method Endpoint Purpose
GET /<role>/engines Discovery and lifecycle state
GET /<role>/health Gateway and model health
GET /<role>/v1/models OpenAI-compatible model list
POST /<role>/generate Raw SGLang payload
POST /<role>/chat/completions Relax alias
POST /<role>/v1/chat/completions OpenAI-compatible chat

Example discovery response:

{
  "role": "teacher",
  "topology_revision": 7,
  "phase": "deferred_score",
  "models": {
    "math-teacher": {
      "state": "ready",
      "router_url": null,
      "engines": [
        {
          "engine_id": "math-teacher/0",
          "base_url": "http://node-1:15000",
          "state": "ready",
          "direct_eligible": true
        }
      ]
    }
  }
}

Routing order:

explicit model
  -> route_key mapping
  -> default model
  -> 400

selected model
  -> SGLang router
  or
  -> READY && direct_eligible replica
  • Multi-node TP/PP workers form one logical replica. Discovery exposes only
    the ingress/head URL.
  • PD prefill/decode workers are diagnostic only. Requests use the router.
  • topology_revision changes after scale, restart, recovery, or endpoint
    changes.
  • Sleeping and draining engines return direct_eligible: false.

Compatibility

Current behavior Migration behavior
GenRM /generate accepts messages Keep adapter and {"response": "..."}
Rollout exposes /v1/chat/completions Keep endpoint
Rollout exposes /engines Evolve response toward common schema
OPD accepts Teacher engine URLs Allow raw engine or gateway URL
Existing CLI flags Convert to internal spec; no first-phase CLI rewrite

Placement modes

Mode GPU relationship Concurrency
Decoupled Separate PG per role Training and inference may overlap
Colocated split Disjoint slices in Actor PG Inference roles may overlap
Colocated defer Same bundles, different phases Roles are serialized
Shared co-resident Same bundles, same phase Rejected
flowchart LR
    subgraph Decoupled
        DA[Actor PG]
        DR[Rollout PG]
        DG[GenRM PG]
        DT[Teacher PG]
    end

    subgraph Split[Colocated split - Actor PG]
        SR[GPU 0..7<br/>Rollout]
        SG[GPU 8..11<br/>GenRM]
        ST[GPU 12..15<br/>Teacher]
    end

    subgraph Defer[Colocated defer]
        PA[Phase A<br/>Rollout READY]
        PB[Phase B<br/>Scorer READY]
        PC[Phase C<br/>Actor READY]
        PA --> PB --> PC --> PA
    end
Loading

PG ownership:

Layout Owner Who deletes the PG
Decoupled + managed Manager Manager
Decoupled + external External service External service
Colocated Controller Controller

Lifecycle

stateDiagram-v2
    [*] --> STARTING
    STARTING --> READY
    READY --> DRAINING
    DRAINING --> SLEEPING
    SLEEPING --> ONLOADING
    ONLOADING --> READY
    STARTING --> DEAD
    READY --> DEAD
    DRAINING --> DEAD
    SLEEPING --> DEAD
Loading

Manager operations are idempotent:

activate(tags=None)
drain()
deactivate()
shutdown()

Topology publication is atomic:

create candidates
  -> wait healthy
  -> sync weights if required
  -> update router
  -> publish registry snapshot
  -> topology_revision += 1

Defer transition:

sequenceDiagram
    participant Coordinator
    participant Rollout
    participant Registry
    participant Scorer
    participant Trainer

    Coordinator->>Registry: Rollout = DRAINING
    Coordinator->>Rollout: stop new work + drain
    Coordinator->>Rollout: flush + offload
    Coordinator->>Scorer: onload
    Coordinator->>Registry: Scorer = READY
    Coordinator->>Scorer: batch score
    Coordinator->>Scorer: drain + offload
    Coordinator->>Trainer: barrier + wake
Loading

The gateway never wakes a sleeping model because a request arrived. It returns
503 with retry guidance.

Deferred OPD

Teacher defer requires a data-flow change, not just different placement.

sequenceDiagram
    participant Rollout
    participant Samples
    participant Teacher
    participant Trainer

    Rollout->>Samples: samples + teacher inputs + token selection
    Note right of Rollout: no inline Teacher request
    Rollout->>Rollout: drain + offload
    Teacher->>Teacher: onload
    Teacher->>Samples: batch prefill / logprobs
    Teacher->>Teacher: drain + offload
    Samples->>Trainer: samples with teacher logprobs
    Trainer->>Trainer: compute OPD loss
Loading

Teacher defer is incomplete until the write-back path is implemented.

Migration

Phase Deliverable Public behavior
0 Characterize args, placement, endpoints, lifecycle, DCS/router behavior No change
1 Common discovery schema, topology revision, common client Additive
2 One gateway class, multi-model routing, legacy adapters Compatible
3 One manager and one SGLang engine path Facades remain
4 Placement planner, decoupled and split Reject invalid layouts
5 Lifecycle coordinator, GenRM defer, deferred OPD New framework pipeline
6 Move Rollout engine ownership to unified manager Workload stays separate
7 Remove old managers, subclass, and named-actor special cases Breaking cleanup

Start with Phase 0 and Phase 1. Changes to argument parsing, Controller,
Service, Launcher, or public API deletion require separate approval.

Acceptance checks

  • Same gateway class serves Rollout, GenRM, and Teacher.
  • Same manager class owns all three engine-pool types.
  • Gateway and direct clients share routing rules.
  • Discovery never exposes internal TP/PP workers as replicas.
  • Split overlap and invalid defer schedules fail before startup.
  • Shared co-resident layouts fail before startup.
  • Static models cannot register for DCS or dynamic weight updates.
  • Teacher defer writes Teacher outputs back before training.
  • PG ownership and rollback are covered by tests.
  • Multi-node GPU tests pass, or skip with an explicit hardware reason.

Open questions

# Question
1 Keep adapting existing arguments, or design a unified user-facing config in a separate RFC?
2 Should direct clients fall back to the gateway by default?
3 How is external discovery authenticated?
4 Configured or planner-optimized defer sub-phase order?
5 Elastic scale for Rollout only, or also GenRM/Teacher?
6 How long do compatibility facades remain?

Proposed review outcome

Accept architecture
  + approve Phase 0/1
  + defer unified CLI design
  + defer Controller/Service cleanup
  + reject shared co-resident in v1
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions