Repository navigation
refactor(runtime): derive Session transcripts from RuntimeEvents #4791
Description
Activity
Proposed implementation
I would implement this around one RuntimeEvent-backed transcript reader, preserving
StoredMessageas the public DTO. The required guarantees are a single execution authority, bounded page reads, stable item identities and supported cursor continuation across terminalization/reconnect/restart, and an explicit one-way import boundary. The storage design should establish those guarantees without assuming that a full materialized transcript or every intermediate streaming revision must be retained.1. Remove independent runtime message writers
Remove ordinary conversation
appendMessage/appendMessagescalls fromAiSdkTurn,ToolRuntime, andAgentRun. Give these components canonical execution command capabilities, not access to mutate transcript rows.Preserve specialized tool-preparation, tool-outcome, and terminal commit operations. They already carry domain validation and atomicity requirements; a generic append API must not bypass them. All such operations write into the same RuntimeEvent authority. Event identity and ordering uniqueness belong in storage constraints, with idempotent retries for identical events and explicit rejection of conflicting payloads.
The terminal RuntimeEvent determines semantic completion. Operational run headers must be updated atomically with it or repairable from it. Transcript reads must not choose their authority based on whether the header says the run has finished.
2. Reuse canonical ordering and define item identity
Use or extend the existing Session RuntimeEvent ordinal mechanism. Per-run sequence numbers and timestamps alone are insufficient for ordering a Session containing multiple runs.
Distinguish three properties:
Property Contract Item identity Derived deterministically from canonical message/step/tool identities, scoped where necessary Display position Fixed by a durable first-appearance anchor; later updates do not move the item Content revision Identifies the committed representation being read Reuse existing identity fields when unambiguous. Do not allocate random IDs during projection or assign a different identity when an active item becomes final. If first appearance currently exists only in mutable partial state, persist the anchor needed to preserve its position.
3. Start with indexed, bounded projection reads
Extend the reader around
sessionId, snapshot watermark, item position,maxItems, andmaxBytes. Start with a rebuildable lookup index linking item identities to their contributing events/content segments. Use keyed dependency lookups rather than loading preceding Session history. The same projection rules should serve active and completed items.The boundedness requirement applies to actual storage reads, decoding, and memory, not only response size. Selecting one item and replaying all its events can still be unbounded. Likewise, reading the full range between its first and last event may pull in many unrelated events. Large content must support bounded reads before DTO assembly; a small response fragment must not require reconstructing a huge JSON object first.
Validate this reader with long Sessions, huge individual messages/tool results, many updates to one item, and heavily interleaved events. If direct projection plus the lookup index satisfies the same correctness and resource limits, keep it. If it cannot, add only the required derived metadata, summaries of projection state, or chunked item payloads to remove the measured unbounded work. A full materialized transcript is an implementation option, not a prerequisite.
Any derived payloads have one dedicated projector as their writer, explicit source references, a projection version, and a progress watermark. They are disposable and rebuildable; they never determine execution recovery. Missing/incompatible projections enter bounded preparation or an explicit error state, rather than triggering an unbounded page-time rebuild or a StoredMessage fallback.
4. Specify cursor and streaming guarantees before choosing retention
A stable item position and an exact byte continuation are different contracts. Preserve stable logical positions across completion, reconnect, and restart. Where the protocol returns fragments of a serialized item, bind continuation to the content revision, length/digest, and byte offset of that representation. Never concatenate bytes from different revisions.
Define the supported snapshot/cursor lifetime explicitly. During that lifetime, referenced bytes must remain readable or reconstructible across restart; expiry or incompatibility must be explicit. Replacing exact continuation with “reload the latest item” would change the contract and cannot be presented as equivalent cursor support.
Separate resumable positions from transient subscription IDs. Signed cursors need a key lifecycle compatible with restart, and each resumed request still validates access. Projection format and data-generation changes must also be detectable.
For streaming, first define which output is acknowledged as durable. Existing mutable partial snapshots can represent current progress, but do not by themselves preserve old active snapshots after finalization removes them. Add immutable segments or checkpoints only where needed to back acknowledged revisions and supported active cursors. Batch updates; do not assume one flush per token or indefinite retention of every displayed intermediate state. Any uncommitted preview is explicitly provisional. These changes must not silently alter provider-history replay or compaction under #4779.
5. Recover live delivery from committed positions
Use snapshot-plus-tail delivery with a durable
subscribeAfter(H)contract: read the snapshot throughH, catch up committed changes afterH, then continue following commits. Notifications wake the reader; the committed positions determine what it must deliver. This closes the gap between capturing a watermark and attaching a listener.A crash after commit but before
SessionEventpublication must be recoverable by catch-up. Clients reconcile repeated updates by item identity/revision.SessionEventremains transport and does not need its own authoritative ledger. Durable/overlay presentation concepts may remain if both use the same RuntimeEvent-backed projection.Keep cheap lookup/index updates atomic with canonical commits where practical. Do not automatically place expensive formatting or serialization inside tool and terminal transactions. If a projection is asynchronous, publish its data and progress consistently, serve only through its completed watermark, and catch up idempotently. Choose synchronous versus asynchronous maintenance using the same correctness tests plus write latency, lock duration, and read freshness; neither should be assumed necessary in advance.
6. Compose independent facts and migrate once
UI notes and WorkHub retain their own authorities. Merge their bounded contributions using stable namespaced identities, explicit ordering anchors, and per-source snapshot positions. Do not imply an atomic snapshot across independent stores without a mechanism that supplies one. User-owned facts such as read acknowledgements also remain authoritative in their own domain; unread status combines those facts with projected transcript positions.
New Sessions select the RuntimeEvent authority contract before execution. Supported imports preserve visible identities, use deterministic migration event IDs, validate the supported facts, and publish the prepared data and authority marker atomically. A bounded import may fit one transaction; larger imports can use hidden staging and resumable batches. Retrying preparation must not duplicate history or expose a partially migrated Session.
Unsupported/ambiguous records fail at the compatibility boundary; unmarked 0.1.x data need not be accepted. Before publication, failed preparation can be retried or discarded. After publication, recovery stays on RuntimeEvents: reverting a projector must not reactivate legacy runtime writers. Remove ordinary-read comparison/backfill fallbacks after cutover, retaining only supported explicit import conversion. Update the Runtime Host compatibility epoch for wire or paging semantic changes.
7. Implementation checks
First establish ordering, identity, acknowledgement, and cursor contracts; then implement the bounded reader and necessary storage support. Before Desktop/CLI integration, use production storage in subprocess crash tests at user/stream/tool/terminal commits, projection publication, and import boundaries. Also cover concurrent readers/writers, retries, commit-before-notification crashes, and snapshot/subscription races.
Two checks should drive the implementation:
- Delete only derived transcript data and rebuild. Supported item identities, ordering, content, turn states, and cursor continuations must agree with the canonical facts and legitimate external authorities, without executing tools or contacting providers.
- Increase Session size, item size, update count, and event interleaving independently. A fixed page budget must bound actual read/decoding work and memory. Run these checks under identical cursor and durability contracts when comparing implementations; weakening a guarantee is not evidence that its supporting mechanism is unnecessary.
These are proposed verification criteria, not benchmark results already obtained.
Platform Intended contract on supported local filesystems Verification Linux SQLite atomicity/order and recovery of acknowledged commits after process termination Crash/reopen and concurrent access tests macOS Same application-level contract Same tests on macOS Windows Same application-level contract Same tests, including database locking/reopen behavior Specify power-loss durability separately against SQLite synchronization/VFS and filesystem guarantees; process-kill tests do not prove it. No platform silently switches to a StoredMessage or JSONL authority.
中文版本
实现建议
我建议围绕一个由 RuntimeEvent 支撑的会话读取器实现,保留
StoredMessage作为公开 DTO。必须满足的保证是:执行事实只有一个权威、分页读取有界、条目身份及受支持游标在终止/重连/重启后保持稳定,以及明确的单向导入边界。存储设计应围绕这些保证展开,不预先假定必须建立完整物化会话记录或保存所有流式中间版本。1. 移除独立的 Runtime 消息写入方
删除
AiSdkTurn、ToolRuntime和AgentRun对普通对话的appendMessage/appendMessages调用。这些组件只获得规范执行命令能力,不获得修改会话行的权限。保留工具准备、工具结果和终止等专用提交操作。它们已经承担领域校验和原子性要求,通用追加接口不能绕过这些边界。所有操作写入同一个 RuntimeEvent 权威。事件身份和顺序唯一性由存储约束保证,相同事件可以幂等重试,冲突载荷必须明确拒绝。
终止 RuntimeEvent 决定语义上的完成状态。运行态 Header 应与其原子更新,或能根据它修复。会话读取不能根据 Header 是否显示完成来选择权威来源。
2. 复用规范顺序,定义条目身份
使用或扩展现有 Session RuntimeEvent ordinal 机制。只有 Run 内序号和时间戳,不足以为包含多个 Run 的 Session 排序。
分清三个属性:
属性 契约 条目身份 根据规范消息、Step 或工具身份确定性派生,必要时增加作用域 展示位置 由持久化的首次出现锚点确定,后续更新不移动条目 内容版本 标识当前读取的已提交表示 无歧义时复用已有身份字段。投影时不分配随机 ID,活跃条目完成时不更换身份。如果首次出现目前仅记录在可变 partial 状态中,应持久化维持位置所需的锚点。
3. 从有索引的有界投影读取开始
读取接口围绕
sessionId、快照高水位、条目位置、maxItems和maxBytes展开。先使用可重建的定位索引,把条目身份关联到生成它的事件或内容段。通过按键查询解决依赖,避免加载此前全部 Session 历史。活跃和已完成条目使用相同投影规则。有界性约束实际存储读取量、解码量和内存,而不只是响应大小。选中一个条目后回放它的全部事件,仍可能无界;读取首尾事件之间的完整范围,也可能带入大量无关事件。大内容必须在 DTO 组装前就支持有界读取,不能为了返回一个小片段先重建巨大 JSON 对象。
使用长 Session、单条超大消息或工具结果、同一条目的大量更新,以及高度交错事件验证读取器。如果直接投影加定位索引能满足同样的正确性和资源限制,就保留这个实现;如果不能,再补充消除已测得无界工作所必需的派生元数据、投影状态摘要或分块条目载荷。完整物化会话记录是实现选项,不是前置条件。
任何派生载荷都只有一个专用投影器拥有写权限,并具备明确的来源引用、投影版本和进度高水位。它们可以丢弃和重建,不能决定执行恢复。投影缺失或不兼容时进入有界准备过程或明确报错,不能在分页请求内执行无界重建,也不能回退到 StoredMessage。
4. 先确定游标和流式保证,再选择保留策略
稳定的条目位置和精确的字节续读是不同契约。逻辑位置应在完成、重连和重启后保持稳定。协议返回序列化条目分片时,续读必须绑定该表示的内容版本、长度/摘要和字节偏移,不能拼接不同版本的字节。
明确受支持快照和游标的有效期。有效期内,被引用字节在重启后必须仍可读取或重建;过期和不兼容必须明确报告。把精确续读改成“重新加载最新条目”会改变契约,不能视为等价的游标支持。
可恢复位置与临时订阅 ID 分开。签名游标的密钥生命周期必须支持重启,每次续读仍需验证访问权限。投影格式和数据代次变化也必须能被识别。
流式输出先明确哪些内容被确认已持久化。现有可变 partial 快照可以表示当前进度,但最终完成时删除快照后,它本身不能保留旧活跃快照。只有为已确认版本和受支持的活跃游标提供依据所必需时,才增加不可变内容段或检查点。合并更新,不预设每个 token 刷盘,也不无限保留所有曾显示过的中间状态。未提交预览明确属于暂定内容。这些改动不能静默改变 #4779 负责的 Provider 历史回放或压缩。
5. 从已提交位置恢复实时交付
使用持久化的
subscribeAfter(H)契约衔接快照与后续更新:读取截至H的快照,补读H之后的已提交变化,再继续跟随后续提交。通知负责唤醒读取,已提交位置决定应交付什么,从而消除获取高水位和安装监听器之间的缺口。提交后、
SessionEvent发布前崩溃,必须能通过补读恢复。客户端按条目身份和版本协调重复更新。SessionEvent继续作为传输,无需拥有独立权威账本。durable/overlay 展示概念可以保留,只要两者使用同一个由 RuntimeEvent 支撑的投影。在可行时,将低成本定位索引更新与规范提交放入同一事务。不要默认把昂贵格式转换或序列化放入工具和终止事务。如果投影异步维护,必须一致发布数据和进度,只服务投影已经完成的高水位,并支持幂等追赶。以相同正确性测试,加上写入延迟、锁持有时间和读取新鲜度,选择同步或异步维护;不预先认定某一种必需。
6. 组合独立事实,完成一次性迁移
UI 备注和 WorkHub 保留各自权威。通过稳定、带命名空间的身份、明确排序锚点和各来源快照位置,有界合并它们的贡献。没有相应机制时,不宣称独立存储之间存在原子快照。已读确认等用户拥有的事实也保留自身领域权威;未读状态由这些事实和投影会话位置共同决定。
新建 Session 在执行前确定 RuntimeEvent 权威契约。受支持导入保留可见身份,使用确定性迁移事件 ID,验证受支持事实,并原子发布准备好的数据及权威标记。有界的小导入可以在单个事务内完成;较大导入可以使用隐藏暂存和可续接批次。重试不能重复导入历史,也不能暴露部分迁移的 Session。
不受支持或有歧义的记录在兼容性边界失败;无需接受未标记的 0.1.x 数据。发布前可以重试或丢弃失败的准备,发布后恢复始终使用 RuntimeEvents,回退投影器版本不能重新开启旧 Runtime 写入方。切换后删除普通读取中的比较和回填回退,只保留受支持的显式导入转换。传输或分页语义变化时更新 Runtime Host compatibility epoch。
7. 实施验证
先确定顺序、身份、持久化确认和游标契约,再实现有界读取器及必要存储支持。接入 Desktop/CLI 前,使用生产存储进行子进程崩溃测试,覆盖用户/流式/工具/终止提交、投影发布和导入边界,并覆盖并发读写、重试、提交后通知前崩溃,以及快照与订阅竞争。
用两项检查驱动实现:
- 仅删除派生会话数据并重建。受支持的条目身份、顺序、内容、Turn 状态和游标续读,必须与规范事实及其他合法权威一致,且不执行工具、不联系 Provider。
- 分别增加 Session 大小、单条内容大小、更新次数和事件交错程度。固定分页预算必须约束实际读取/解码工作量和内存。比较实现时保持相同游标及持久化契约,不能把放宽保证后的通过当作支持机制不必要的证据。
以上是拟实施的验证标准,不是已经取得的性能测试结果。
平台 受支持本地文件系统上的目标契约 验证 Linux SQLite 原子性和顺序,进程终止后恢复已确认提交 崩溃重开及并发访问测试 macOS 相同应用层契约 在 macOS 上执行相同测试 Windows 相同应用层契约 相同测试,并覆盖数据库锁定及重开行为 断电持久性应结合 SQLite 同步配置、VFS 和文件系统保证另行说明;杀进程测试不能证明断电安全。任何平台都不能静默切换到 StoredMessage 或 JSONL 权威。
Freeze legacy prefix instead of backfilling
Existing Sessions must continue to open, display full transcripts, and accept new turns without backfill.
Why backfill is impossible
materializeTranscriptLedgerhardcodesmodelHistory: 'conversation_text'(runtime-ledger-repair.ts:112), causingbackfillRuntimeEventsFromStoredMessagesto skiptool_call,tool_result,token_usage, andpermission_decision(runtime-event-backfill.ts:194,:244,:337). Projecting from this backfill yields zero tool cards.- Faithful backfill is impossible in principle:
tool_resultIDs are generated vianewId()(tool-runtime.ts:1940) with no ref in mapper (session-event-runtime-mapper.ts:380);token_usageandturn_stateshare this issue. Fabricating facts never emitted by execution destroys log authority.
Proposal: per-Session freeze watermark
During schema migration, record
transcriptFreeze: { throughSequence, throughOrdinal }usingMAX(session_messages.sequence)andMAX(runtime_session_event_ordinals.ordinal):- Prefix (
sequence <= throughSequence): Immutable read fromsession_messages; chunk-level byte reads (sqlite-session-metadata-store.ts:2600) unchanged. - Suffix (
ordinal > throughOrdinal): Projected fromRuntimeEvents.
Guarantees zero migration failure, O(1) startup, and intact history.
runtime-event-backfill.tsstays strictly for one-way external session imports, which already use this shape.Delivery (3 PRs)
- PR1 (Read switch, zero data risk): Add freeze fields, implement event-backed durable reader (retaining all 4
SessionTranscriptReadermethod signatures), bumpRUNTIME_HOST_COMPATIBILITY_EPOCHand cursor version. No write changes. Revert is safe over untouchedsession_messages. Reader ignores post-freezesession_messages, eliminating the need for dual-write parity verification. - PR2 (Delete execution double-writes): Remove writes in
ai-sdk-turn.ts:937, :2491,tool-runtime.ts:1090/1304/1949/2100,agent-run.ts:689/848/1089,runtime-kernel.ts:2011/2768/1104, and recoveryappendTurnState(session-manager.ts:4709). (PR1 must not be reverted after PR2). - PR3 (Drop compensation layer): Remove
materializeRuntimeEventTranscriptProjection,repairSteeringMessagesOnce,RuntimeReadModelcaches (projectionCache,mergeInFlightProjectionCache,compareProjectionCache),compareRuntimeReadModelMessages, and drift diagnostics.
Blocking
readMessagesconsumersOf 20 real call sites, 3 affect model/orchestration decisions and must migrate before PR2:
goal-coordinator.ts:159(getRecentContext/ tokens; silently degrades to empty context).session-manager.ts:3315(subagent spawnsummaryfallback; partial-text children regress to empty summary).runtime-host-run-command.ts:313, :405(Agent Graph turn outcomes).
- Unaffected: History tools (
execution-composition.ts:466already usesgetMessages) and Desktop search (loadTranscript()). - Validation/Idempotency: 7 checks (
appendUserMessageOnce,turnHasRetainedOutput, etc.) drop with PR2; covered by event PKs,UNIQUE(invocation_id, event_seq), projectedpartialOutputRetained(runtime-event-read-model.ts:1175), and terminal events. - Recovery risk:
hosted-execution-recovery.ts:65andsession-manager.ts:1473.turn_state: runningis recorded byinvocation_openedbut currently unprojected; close this gap before scoping PR2.
Boundedness: invocation page unit
A materialized transcript table is not a prerequisite:
runtime_events_by_run (session_id, run_id, event_seq)andscanRuntimeEventsbudgets enforce a 16MB/invocation cap.readRangeEdgesalready widens pages to Turn boundaries, throwingRangeErrorpastSESSION_TRANSCRIPT_RANGE_MAX_BYTES.- Projection state is invocation-scoped; inter-invocation paging uses
listSessionInvocationsPage. - An in-memory LRU keyed by
(sessionId, runId, terminalEventId)decodes a 16MB Turn once across 512KB pages. Terminal invocations are immutable; no cache invalidation needed. - Only introduce a derived index table if cold bootstrap regresses >3x
readTranscriptPageSnapshotor decoding exceeds >2x delivered bytes on 1000-turn / 16MBtool_result/ 4000-event benchmarks. - Required index: expression index on
(session_id, json_extract(payload_json, '$.refs.providerEventId'))formarkSessionReadThroughMessage(session-store.ts:1174).
Non-execution facts
- Run-internal notes: Project from events, not stored rows (
session_resumefrom terminal event;context_compactedfrom token usage;context_*from provider errors). Precedent:projectTerminalTurnStatesynthesizesstep_limitfromtool_step_cap_reached(runtime-event-read-model.ts:1199). - Session-level notes: (
session_start,mode_change,model_change, plan abandonment) remain insession_messages. Write withanchorOrdinal(MAX(ordinal)) and merge into suffix by(anchorOrdinal, sequence). In PR2, restrict execution callers fromappendMessagetoappendSessionNote. - WorkHub coordination: Lives in
maka_workhub_coordination(core/session.ts:78); handled cleanly via session notes.
Deletions & convergence
Drop 12 execution write sites;
materializeRuntimeEventTranscriptProjection;repairSteeringMessagesOnce;RuntimeReadModelprojection caches and drift diagnostics; active/durable switch insession-transcript-reader.ts:61(liveness derived frominvocation.terminalEvent);readMessagesForRecoveryexecution branch; and "Phase 4 placeholder" framing insession-event-runtime-mapper.ts(now canonical ingress).session_messagesstays strictly for frozen prefixes and session notes.Corrections to issue & proposal
- Issue corrections: Drop "Unmarked 0.1.x data does not need to be supported" (freeze watermark preserves all legacy sessions). Removing comparison/repair paths after cutover is valid only because backfill is dropped.
- Earlier proposal adjustments: Atomic authority marker already exists (
transcriptLedgerVersion+ensureTranscriptLedger+host-session-availability.ts:66). Drop three redundant designs:- Restart-surviving cursor keys: Cursor secrets are ephemeral
randomBytes(32)(session-transcript-pager.ts:99); epoch bumps handle wire changes. - Durable
subscribeAfter(H): Existing pager already checks watermarks and drift;runtime_partial_snapshotsandmergeActiveAssistantStreamshandle stream state. Only the event-backed durable source is needed. - Crash/power-loss test matrix gate: SQLite atomicity is already verified by existing crash/concurrency tests; do not make it a gate.
- Restart-surviving cursor keys: Cursor secrets are ephemeral
中文版本
冻结旧数据前缀,替代全量回填
核心约束:升级后所有存量 Session 必须能正常打开、完整展示 transcript 并接受新 turn,且无需数据回填。
为什么回填在原理上不可行
materializeTranscriptLedger硬编码了modelHistory: 'conversation_text'(runtime-ledger-repair.ts:112),导致backfillRuntimeEventsFromStoredMessages直接跳过tool_call、tool_result、token_usage和permission_decision(runtime-event-backfill.ts:194、:244、:337)。由此投影出的 transcript 没有任何 tool card。- 理论上无法做到保真回填:
tool_result的 ID 为newId()(tool-runtime.ts:1940)且 mapper 中无对应 ref(session-event-runtime-mapper.ts:380);token_usage和turn_state同理。伪造执行期未曾观测到的事实会直接破坏 log 作为权威源的语义。
方案:按 Session 记录冻结水位线(Freeze Watermark)
在 schema migration 事务中,记录每个 Session 的
transcriptFreeze: { throughSequence, throughOrdinal }(取session_messages.sequence与runtime_session_event_ordinals.ordinal的主键MAX):- Prefix(
sequence <= throughSequence):只读访问session_messages,分块字节读取(sqlite-session-metadata-store.ts:2600)保持原样。 - Suffix(
ordinal > throughOrdinal):由RuntimeEvents投影生成。
无迁移失败风险、O(1) 启动、历史完全保真。
runtime-event-backfill.ts仅保留用于单向外部会话导入(导入数据本就符合该结构)。交付规划(3 个 PR)
- PR1(切换读权威,零数据风险): 引入冻结字段,实现基于事件的持久化读取器(保持 4 个
SessionTranscriptReader方法签名),bumpRUNTIME_HOST_COMPATIBILITY_EPOCH与 cursor 版本。写路径完全不动,revert 可无缝回退到未修改的session_messages。读取器会直接忽略冻结水位后的session_messages,因此无需双写对齐校验期。 - PR2(清理执行期双写): 移除
ai-sdk-turn.ts:937, :2491、tool-runtime.ts:1090/1304/1949/2100、agent-run.ts:689/848/1089、runtime-kernel.ts:2011/2768/1104以及session-manager.ts:4709处的恢复写入appendTurnState。(PR1 合并后方可合并 PR2,且 PR1 不可再单独 revert)。 - PR3(移除补偿层): 删除
materializeRuntimeEventTranscriptProjection、repairSteeringMessagesOnce、RuntimeReadModel缓存(projectionCache、mergeInFlightProjectionCache、compareProjectionCache)、compareRuntimeReadModelMessages及漂移诊断代码。
阻塞 PR2 的
readMessages调用点在全部 20 处调用中,有 3 处影响模型或调度决策,必须在 PR2 前完成改造:
goal-coordinator.ts:159(getRecentContext及 continuation 的 token 统计;目前会静默降级为空上下文)。session-manager.ts:3315(subagent spawn 的summaryfallback;部分文本子 agent 会退化为父级上下文中的空 summary)。runtime-host-run-command.ts:313, :405(基于存储 turn 推导的 Agent Graph 结果)。
- 不受影响: History tools(
execution-composition.ts:466已走getMessages)与桌面端搜索(走loadTranscript())。 - 校验与幂等: 7 处检查(如
appendUserMessageOnce、turnHasRetainedOutput等)随写路径移除;现有事件主键、UNIQUE(invocation_id, event_seq)、投影字段partialOutputRetained(runtime-event-read-model.ts:1175)及终态事件已提供替代保证。 - 恢复链路风险:
hosted-execution-recovery.ts:65与session-manager.ts:1473。目前invocation_opened虽记录事实但投影层未生成对应turn_state: running,需在 PR2 定稿前补齐。
有界性:以 Invocation 为分页单元
无需预先引入物化 transcript 表:
- 现有
runtime_events_by_run (session_id, run_id, event_seq)与scanRuntimeEvents的字节预算已限制单 invocation 上限为 16MB。 readRangeEdges已按 Turn 边界展开分页,超过SESSION_TRANSCRIPT_RANGE_MAX_BYTES会抛出RangeError。“单页 <= 单 Turn <= 16MB”属于既有约束。- 投影状态均为 invocation 作用域,跨 invocation 遍历直接复用分页的
listSessionInvocationsPage。 - 内存中维护以
(sessionId, runId, terminalEventId)为 key 的 LRU 缓存投影行:以 512KB 分页遍历 16MB Turn 只需解码一次。终态 invocation 不可变,无需失效逻辑。 - 仅当压测显示冷启动劣于
readTranscriptPageSnapshot3 倍以上,或 1000-turn / 16MBtool_result/ 4000-event 场景下解码字节超输出字节 2 倍以上时,才考虑引入物化索引表。 - 必需索引:在
session-store.ts:1174(markSessionReadThroughMessage)增加(session_id, json_extract(payload_json, '$.refs.providerEventId'))表达式索引,避免全表扫描。
非执行期事实
- Run 内部 notes: 应基于事件投影而非落表存储(如终态事件推导
session_resume,token usage 事件推导context_compacted,provider 报错推导context_*)。参考先例:projectTerminalTurnState已从tool_step_cap_reached合成step_limit(runtime-event-read-model.ts:1199)。 - Session 级别 notes:(
session_start、mode_change、model_change、放弃计划等)保留在session_messages中充当 notes(而非第二权威)。写入时打上anchorOrdinal(MAX(ordinal)),读取时按(anchorOrdinal, sequence)归并入后缀。在 PR2 中收回执行模块的appendMessage权限,改为appendSessionNote。 - WorkHub 协调: 独立存放在
maka_workhub_coordination(core/session.ts:78),天然走 session notes 逻辑,无需特殊处理。
待清理项与收敛标准
清理 12 处执行写入点;
materializeRuntimeEventTranscriptProjection与repairSteeringMessagesOnce;RuntimeReadModel投影缓存及漂移诊断;session-transcript-reader.ts:61的 active/durable 权限切换(活跃态直接由invocation.terminalEvent判定);readMessagesForRecovery执行分支;移除session-event-runtime-mapper.ts的“Phase 4 placeholder”注记(已转为正式 ingress)。session_messages仅保留用于冻结前缀与 session notes。对 Issue 及前期方案的修正
- Issue 修正: 删去“无需支持未标记的 0.1.x 数据”(冻结水位线原生支持全量存量会话)。“切流后移除过渡比对与修复路径”能够成立,前提正是放弃回填。
- 前期方案调整: 权威标记的原子发布已有现成机制(
transcriptLedgerVersion+ensureTranscriptLedger+host-session-availability.ts:66)。去除三处过度设计:- 重启后 cursor key 生命周期: Cursor secret 本就是单次订阅的临时
randomBytes(32)(session-transcript-pager.ts:99),协议变更直接由 epoch 处理。 - 持久化
subscribeAfter(H): 现有 pager 已具备水位校验与漂移报错,runtime_partial_snapshots与mergeActiveAssistantStreams已处理流式状态,仅需接入事件驱动持久化源。 - 跨平台 crash/掉电测试矩阵门禁: SQLite 原子性已有既有 crash 和并发用例覆盖,不应作为本次重构的阻塞门禁。
- 重启后 cursor key 生命周期: Cursor secret 本就是单次订阅的临时
@M4n5ter you built this seam — the transcript reader in #2134, the steering repair in #2420, and the bounded pager in #2922 — so I'd value your read on the proposal above before anyone starts on it.
Two things specifically:
-
The freeze-watermark migration keeps
session_messagesas the read source for everything before the upgrade point, so the overlay/durable split insession-transcript-reader.ts:61collapses into one projection with a watermark. Does that hold up against what perf(runtime-host): replace transcript snapshots with bounded pages #2922 needed from the split? -
turn_state: runninghas no projected row today —invocation_openedis its fact. Recovery reads those rows, so this decides how much of theappendTurnStateremoval is in scope.
-
I support the goal of making
RuntimeEventthe sole semantic authority, but I think we should make this a deliberate hard cutover instead of carrying a frozen legacy prefix or a permanent dual-read path.Proposed target architecture
RuntimeEventis the only persisted source of truth for ordinary execution history.- Transcript rows, turn state, tool calls/results, and other presentation records are projections derived from events.
- Recovery reads canonical execution events directly; it must not depend on projected
StoredMessageDTOs. StoredMessageremains an API/UI projection type, not a second durable execution ledger.- Session-level notes, if they must remain durable, should have an explicit model rather than being mixed into the execution transcript ledger.
- Transcript pagination/cursors should be redesigned around the event-backed model. Invalidating old cursors is acceptable as part of the breaking change.
- Remove execution double-writes, ledger repair, drift comparison, frozen-prefix handling, and other permanent compatibility machinery once the cutover lands.
Migration through the existing Session Importer
Maka already has session import functionality, which is a cleaner migration boundary than adding compatibility to the runtime read path.
I suggest extending the Session Importer to accept the legacy
session_messagesrepresentation and convert it into canonicalRuntimeEvents:- Detect/version the legacy import payload.
- Convert the available legacy transcript data into canonical events, with explicit legacy-import provenance.
- Assign deterministic event IDs and ordering so imports are idempotent and safely retryable.
- Validate the imported session by projecting its transcript through the new event-backed reader.
- Commit each session atomically; failed imports must not leave partially migrated sessions.
- Report information that cannot be reconstructed faithfully instead of silently inventing execution facts.
The migration script can then stay intentionally thin:
read legacy session → build legacy session-import payload → call Maka's normal session import path → record success/failureThis keeps backward compatibility out of the production architecture. The importer may temporarily understand legacy data, and that legacy branch can be deleted after the migration window.
A hard cutover still needs to be complete: transcript reads, recovery, paging, unread markers, goal context, subagent summaries, and Agent Graph consumers should move together. Otherwise we would replace the current dual-ledger problem with a partially migrated system.
中文
我支持让
RuntimeEvent成为唯一语义权威,但建议把这次改造明确做成 hard cutover,而不是长期保留冻结的旧前缀或双读路径。目标架构
RuntimeEvent是普通执行历史唯一的持久化事实来源。- transcript、turn state、tool call/result 等展示记录全部从事件投影生成。
- recovery 直接读取权威执行事件,不能依赖作为展示 DTO 的
StoredMessage。 StoredMessage只保留为 API/UI 投影类型,不再作为第二套执行账本持久化。- session-level note 如果需要持久化,应使用明确的数据模型,而不是继续混在执行 transcript 账本里。
- transcript 分页和 cursor 围绕 event-backed 模型重新设计;作为 breaking change,旧 cursor 失效是可以接受的。
- cutover 完成后,删除执行双写、ledger repair、drift comparison、冻结前缀及其他永久兼容代码。
复用现有 Session Importer 做迁移
Maka 已经有 session import 功能。相比在 runtime 读取路径中增加兼容逻辑,它更适合作为历史数据迁移边界。
建议扩展 Session Importer,让它能够接收旧的
session_messages表示,并转换成规范的RuntimeEvent:- 对旧版 import payload 进行识别和版本管理。
- 将能够恢复的旧 transcript 数据转换为规范事件,并明确标记 legacy import 来源。
- 使用确定性的 event ID 和顺序,保证导入幂等且可以安全重试。
- 通过新的 event-backed reader 投影 transcript,校验导入结果。
- 每个 session 原子提交;导入失败不能留下半迁移状态。
- 无法无损恢复的信息应明确报告,不能静默编造执行事实。
迁移脚本因此可以非常薄:
读取旧 session → 构造 legacy session-import payload → 调用 Maka 正常的 session import 路径 → 记录成功或失败这样历史兼容不会进入生产运行时架构。Importer 可以暂时理解旧数据,迁移窗口结束后再删除对应的 legacy 分支。
Hard cutover 仍然需要一次迁移完整:transcript 读取、恢复、分页、未读标记、goal context、subagent summary 和 Agent Graph consumer 应一起切换。否则只是把当前双账本问题换成半迁移状态。
Thanks @Astro-Han for the detailed writeup — the freeze-watermark breakdown is what sent me digging through the code. A couple of things I ran into:
Backfill isn't actually blocked
The freeze-prefix case hinges on "faithful backfill is impossible in principle." I don't think that holds:
- The conversion pipeline already exists and already runs for existing (non-imported) Sessions —
ensureTranscriptLedgergets called withsource: 'compatibility'(session-manager.ts:2051), not just for imports, and it drivesmaterializeTranscriptLedgerto derive RuntimeEvents from StoredMessages. - It drops tool cards today because of a mode choice, not missing data:
materializeTranscriptLedgerhardcodesmodelHistory: 'conversation_text'(runtime-ledger-repair.ts:112), and the skips fortool_call/tool_result/token_usage/permission_decisionare gated on that mode (runtime-event-backfill.ts:96). The same backfiller emits all of them infullmode. - The linkage is there:
tool_resultcarriesrefs.toolCallId(session-event-runtime-mapper.ts:380), and run ids are already deterministic (sha256(sessionId + '\0' + turnId),runtime-ledger-repair.ts:176). The only non-deterministic bit is the per-event id (newId()), which we'd want to fix anyway for idempotent retry.
So it's really "convert what's derivable, don't invent the rest" — a small set of facts genuinely can't be rebuilt as events (the
turn_state: runninggap below is the clearest), but that's not a wall. @Astro-Han, if you hit a specific case that broke this, I'd like to see it.I'd go with the hard cutover
No release yet, the issue already says unmarked 0.1.x data doesn't need support, and the acceptance criteria explicitly want transitional paths removed and "no permanent parallel authorities." A frozen prefix leaves a durable-read path in the reader indefinitely, which is the thing we're trying to remove. So for the end state I'd take @likun666661's cutover.
But @Astro-Han's staged sequencing is the safer way to get there, and we can reuse the import/materialize pipeline instead of writing new conversion code:
- P0 — close the
turn_state: runninggap (see below). Has to land before any write removal. - PR1 — switch read authority + convert. Collapse the active/durable split in
session-transcript-reader.ts:60to one RuntimeEvent projection (no watermark under a cutover). ReuseensureTranscriptLedger/materializeTranscriptLedger, but (a) run it infullmode so tool/token/permission events survive, (b) derive event ids deterministically fromrefs. BumpRUNTIME_HOST_COMPATIBILITY_EPOCH(112,protocol/index.ts:104) and the cursor version. Report unreconstructable facts instead of faking them. - PR2 — drop the double-writes (
ai-sdk-turn.ts:937/2491,tool-runtime.ts:1090/1304/1949/2100,agent-run.ts:689/848/1089,runtime-kernel.ts:1104/2011/2768, recoveryappendTurnState). Migrate the three consumers that feed model/orchestration first:goal-coordinator.ts:159(token / recent-context),session-manager.ts:3315(subagent summary fallback),runtime-host-run-command.ts:313/405(Agent Graph outcome). History tools (execution-composition.ts:466) and Desktop search (loadTranscript()) already go through a view, so lower risk. - PR3 — delete the compensation layer (
runtime-read-model.tsprojection/merge/compare caches,compareRuntimeReadModelMessages, drift diagnostics,materializeRuntimeEventTranscriptProjection). Legacy conversion stays only in import/materialize, deleted after the window.
A few things from the proposal I'd skip
Same reductions @Astro-Han called out — against the current code these don't look necessary:
- Restart-surviving cursor keys — cursors are ephemeral
randomBytes(32)+ HMAC per subscription (session-transcript-pager.ts:99); the epoch bump covers wire changes. - A new durable
subscribeAfter(H)— the pager already checks watermarks and does byte-offset continuation, andruntime_partial_snapshots+mergeActiveAssistantStreamsalready hold stream state. Only the event-backed source is new. - A materialized transcript table — boundedness is already enforced (16 MiB / 256-msg caps in
protocol/session-transcript.ts:34, turn-boundary alignment inreadRangeEdges, per-invocation scan budgets overruntime_events_by_run). Add an index only if a benchmark regresses.
One scoping note:
UNIQUE(invocation_id, event_seq)is onruntime_events(sqlite-runtime-schema.ts:58), which we're keeping — so removing the double-writes doesn't drop it; it's what keeps idempotency after cutover. What goes away is the StoredMessage-side idempotency (appendUserMessageOnce,turnHasRetainedOutput).The trade-off, stated plainly
Cutover means a few facts in old Sessions that were never events (or only have random ids) get lost or best-effort synthesized. Fine now, since there's no released data to protect — after a real release with real local Sessions, freezing would look a lot better. Which is the main argument for doing it in the pre-release window.
中文版本
谢谢 @Astro-Han 的梳理,冻结水位那套拆解很有帮助,也是它把我带去翻代码的。我碰到几点:
backfill 其实没被堵死
冻结前缀这套的立足点是“回填原理上不可能保真”,我觉得站不住:
- 转换管线其实已经有了,而且已经在为非导入会话跑——
ensureTranscriptLedger会以source: 'compatibility'被调(session-manager.ts:2051),不只用于导入,它驱动materializeTranscriptLedger从 StoredMessage 派生 RuntimeEvent。 - 今天丢 tool card 是模式选择,不是数据缺失:
materializeTranscriptLedger硬编码了modelHistory: 'conversation_text'(runtime-ledger-repair.ts:112),对tool_call/tool_result/token_usage/permission_decision的跳过只在这个模式下发生(runtime-event-backfill.ts:96)。同一个 backfiller 在full模式下会全都产出。 - 链接也在:
tool_result带refs.toolCallId(session-event-runtime-mapper.ts:380),run id 已经是确定性的(sha256(sessionId + '\0' + turnId),runtime-ledger-repair.ts:176)。唯一不确定的是单事件 id(newId()),而这个为了幂等重试本来就要改。
所以本质上是“能派生的就转,剩下的别编”——确实有一小撮事实没法重建成事件(
turn_state: running缺口最典型),但那不是一堵墙。@Astro-Han,你要是手上有把这条打穿的具体 case,很想看看。我会选 hard cutover
还没 release,issue 自己也写了未标记的 0.1.x 数据不必支持,验收标准又明确要求移除过渡路径、“不得留永久并行权威”。冻结前缀会在 reader 里一直留一条 durable 读路径,而这正是我们想去掉的东西。所以终态我倾向 @likun666661 的 cutover。
但到那里的路,@Astro-Han 的分阶段更稳,而且我们可以复用 import/materialize 管线,不用另写转换代码:
- P0 — 补
turn_state: running缺口(见下)。必须先于任何写移除落地。 - PR1 — 切读权威 + 转换。 把
session-transcript-reader.ts:60的 active/durable 分裂折叠成单一 RuntimeEvent 投影(cutover 下没有水位线要维护)。复用ensureTranscriptLedger/materializeTranscriptLedger,但(a)跑full模式,让 tool/token/permission 事件保住,(b)事件 id 从refs确定性派生。bumpRUNTIME_HOST_COMPATIBILITY_EPOCH(112,protocol/index.ts:104)和 cursor 版本。无法重建的事实如实报告,不要伪造。 - PR2 — 删双写(
ai-sdk-turn.ts:937/2491、tool-runtime.ts:1090/1304/1949/2100、agent-run.ts:689/848/1089、runtime-kernel.ts:1104/2011/2768、恢复路径的appendTurnState)。先迁移三个喂给模型 / 编排的消费者:goal-coordinator.ts:159(token / recent-context)、session-manager.ts:3315(subagent summary fallback)、runtime-host-run-command.ts:313/405(Agent Graph 结果)。History tools(execution-composition.ts:466)和桌面搜索(loadTranscript())已经走 view,风险低些。 - PR3 — 删补偿层(
runtime-read-model.ts的投影 / 合并 / 比对缓存、compareRuntimeReadModelMessages、drift 诊断、materializeRuntimeEventTranscriptProjection)。legacy 转换只留在 import/materialize,窗口后删。
方案里我会跳过的几处
跟 @Astro-Han 点的削减一致——对着现在的代码看,这几处不像是必要的:
- 重启后仍有效的 cursor key——cursor 本来就是每订阅临时的
randomBytes(32)+ HMAC(session-transcript-pager.ts:99),wire 变更 epoch bump 就够了。 - 新增持久化
subscribeAfter(H)——pager 已经校验水位、做字节偏移续传,runtime_partial_snapshots+mergeActiveAssistantStreams也已经承载流式状态。新增的只有事件驱动的那个源。 - 物化 transcript 表——有界性已经强制了(16 MiB / 256 条上限,
protocol/session-transcript.ts:34;readRangeEdges的 Turn 边界对齐;基于runtime_events_by_run的 per-invocation 扫描预算)。压测真回退了再加索引。
一个范围上的提醒:
UNIQUE(invocation_id, event_seq)在runtime_events上(sqlite-runtime-schema.ts:58),这张表我们是保留的——所以删双写不会把它删掉,它正是 cutover 后维持幂等的东西。会没的是 StoredMessage 侧的幂等(appendUserMessageOnce、turnHasRetainedOutput)。把取舍摆明
cutover 意味着旧会话里少数从来不是事件(或只有随机 id)的事实会丢、或只能尽力合成。现在没问题,因为没有已发布的数据要保护——等真 release 之后本地有了有价值的会话,冻结就明显更划算了。这也是我主张趁 pre-release 窗口做的主要理由。
- The conversion pipeline already exists and already runs for existing (non-imported) Sessions —
Correction: withdrawing the freeze watermark, and the plan of record
I proposed the per-Session freeze watermark above. I am withdrawing it. @likun666661's objection holds: a frozen prefix is a second read path that never leaves the production architecture, which is the disease this issue is about, not the cure.
What changed is the constraint. Losing a Session or a line of conversation is unacceptable; losing some structure on pre-cutover Sessions is not. That removes the argument that made backfill impossible, and a hard cutover becomes the smaller change.
Plan of record
One converter, one PR. The Session Importer is extended to accept the legacy
session_messagesrepresentation and is invoked once during the upgrade. No schema migration running beside it, no compatibility branch in the runtime read path, and no state where the repository is half migrated.The conversion is a total function. Every legacy row produces something. A row that cannot be mapped to a structured event becomes a text event carrying its rendered content with legacy provenance. Nothing is skipped and nothing is invented. Two things in this codebase make "report and skip" unsafe:
- Three read-model diagnostics are
hardand make the reader throw, so the Session does not open (runtime-event-read-model.ts:74):unsupported_event,incomplete_event,tool_use_id_mismatch. - Startup recovery deletes a Session outright when
revisionState: 'preparing'and no revision user message is found (session-manager.ts:1481). A dropped row becomes an unattended deletion at next launch.
What converts and what does not. @liuxiaocs7 is right that backfill is not blocked, and I have corrected this section accordingly. In
fullmode the tool facts survive:tool_callemitsfunction_callwithcontent.id = message.id(runtime-event-backfill.ts:194),tool_resultemitsfunction_responsewithcontent.id = message.toolUseIdand the same value inrefs.toolCallId(:243), andsafePriorToolCallmatches them inside the turn, sotool_use_id_mismatchcannot fire on a converted pair.token_usageandpermission_decisionconvert too. My earlier reasoning conflated the RuntimeEvent envelopeid, which isnewId()and has to become derived anyway, with the tool use identity, which comes straight from the durable StoredMessage and was never lost.Three classes genuinely do not survive as structured events:
turn_state— unconditionalbreakin both modes (:356). @liuxiaocs7 called this out, correctly, as the thing that must be settled before the writes go away. Under a single-PR cutover there is no "before": the projection only derivesturn_statefrom terminal events (runtime-event-read-model.ts:1182) and has norunningat all, while recovery finds interrupted turns throughlatest.status === 'running'(session-manager.ts:5236). The fact is not missing from the log, since an invocation with an opening and no terminal event was running, which is howruntimeInvocationOutcomealready reads it. Closing that gap is part of the same change, not a predecessor to it.system_note— dropped in both modes;fullmode only adds askipped_high_risk_messagediagnostic (:359). The eight user-visible kinds (session.ts:792) go with it, so the converter handles these itself rather than inheriting the backfiller's behaviour.- Provider-native tool activity whose opaque
providerOutputwas not retained (skipped_provider_native_replay_gap).
Anything in that third category degrades to a marked text event rather than being skipped.
Synthesized, not fabricated. Each converted invocation needs one opening and exactly one terminal event. Both are
role: 'system', author: 'system', facts about the run rather than anything the run said, and both have existing builders for precisely this case:buildInvocationOpenedEventandbuildSyntheticTerminalRuntimeEvent, whose contract already names "the migration of a header whose run never wrote an event". A legacysystem_noteof kinderrororabortdecides the terminal event's status; it must not become anerrorcontent event, which is a hard diagnostic when non-terminal.Event IDs derive from StoredMessage identity so imports are idempotent and resumable, and converted events keep
refs.storedMessageIdso the admission records that reference message ids by value still join.Hard failure moves to the decision path. After the cutover there is no second view to be served "in place of", so a display-side projection gap becomes a marked placeholder row, the shape
archived_tool_result_placeholderalready uses. Recovery and revision decisions keep failing closed, because a gap there changes a destructive outcome.Scope of the single PR
The converter is one part of one change. The cutover moves every consumer in the same PR:
- Read authority: collapse the active/durable split in
session-transcript-reader.ts:61to one event-backed projection. - Recovery:
readMessagesForRecoveryreads the projection. The verification target stays the admission record, which already holds the content and already materializes a missing user message from it. - Paging and cursors, unread markers,
goal-coordinator.ts:159, the subagent summary fallback atsession-manager.ts:3315, and the Agent Graph consumers atruntime-host-run-command.ts:313/405. - Session notes get their own model instead of living in the execution ledger.
turn_state: runningbecomes derivable from the invocation inventory, so recovery keeps recognizing interrupted turns.
Deleted in the same PR: the 12 execution write sites,
materializeRuntimeEventTranscriptProjection,repairSteeringMessagesOnce, theRuntimeReadModelprojection caches and drift diagnostics, the active/durable authority switch insession-transcript-reader.ts:61, the execution branch ofreadMessagesForRecovery, and the duplicate revision check atsession-manager.ts:1481. That last one is a second implementation of a destructive decision already owned bySessionRevisionCoordinator, whose version asks the admission authority first; I could not construct a reachable divergence between them today, but one destructive decision should not have two implementations.Sessions created before RuntimeEvents existed (before 2026-06-14) and still in use are the only case where converted events have no legal Session ordinal, since ordinals are strictly append-only. Those Sessions get their ordinals shifted by the converted event count.
The one-way migration is announced in #4866 per
CONTRIBUTING.md. Technical discussion stays here.Two costs, recorded as facts
- The migration is not reversible once it runs. Downgrading past the cutover cannot read its own data. The epoch covers the protocol, not the data.
- The upgrade window is proportional to total history. It needs a measured upper bound and must resume after a crash, which the deterministic event IDs give us.
Related finding from #4779
#4779 retired the unmarked 0.1.x text summary contract by requiring a
summaryFormatstamp. Every consumer fails open correctly, butexecution-inspectgrades the retired recordscompaction_checkpoint_invalidat error severity, because the superseded classifier keys onsource.policyVersion, a field 0.1.x never wrote (sourcearrived in #955). Every Session carrying a pre-stamp checkpoint therefore reports durable corruption on inspect, and the records that really are damaged stop standing out. It is the same failure mode as the double write: a contract tightens on the current generation's terms, and records written before the contract existed cannot present the evidence the gate asks for. Whatever we adopt here, a record predating a contract must be classified by what it is, not by a self-description its writer had no vocabulary for.
中文
更正:撤回冻结水位线,以及最终方案
上面那份按 Session 记录冻结水位线的方案是我提的,现在撤回。@likun666661 的反对成立:冻结前缀是一条永远不会离开生产架构的第二读路径,那正是本 issue 要治的病,不是解药。
变化的是约束。丢掉 Session 或对话记录不可接受,cutover 之前的 Session 丢掉一些结构可以接受。这拿掉了"回填在原理上不可行"的论据,hard cutover 反而成了更小的改动。
最终方案
一个转换器,一个 PR。 扩展 Session Importer 使其接受旧的
session_messages表示,在升级时调用一次。不并存 schema migration,runtime 读路径不留兼容分支,仓库不出现半迁移状态。转换是全函数。 每一行旧数据都产出东西。映射不出结构化事件的行,变成一条带 legacy 来源标记、装着它原本渲染内容的文本事件。不跳过,也不编造。这个仓库里有两处让"报告并跳过"变得不安全:
- 读模型有三个
hard诊断会让 reader 抛异常,Session 打不开(runtime-event-read-model.ts:74):unsupported_event、incomplete_event、tool_use_id_mismatch。 - 启动恢复时,
revisionState: 'preparing'且找不到 revision user message 的 Session 会被直接删除(session-manager.ts:1481)。丢一行就等于下次启动时的一次无人值守删除。
什么能转,什么不能。 @liuxiaocs7 指出回填并没有被堵死,这是对的,本节据此更正。
full模式下工具事实是保住的:tool_call产出function_call,content.id = message.id(runtime-event-backfill.ts:194);tool_result产出function_response,content.id = message.toolUseId,refs.toolCallId同值(:243);safePriorToolCall在同一个 turn 内撮合两者,所以转换出来的这一对不可能触发tool_use_id_mismatch。token_usage和permission_decision同样能转。我先前的推理把两个身份混为一谈:RuntimeEvent 信封的id是newId(),不确定且本来就要改成派生式;而工具身份直接来自持久化的 StoredMessage,从来没丢过。真正不能以结构化事件形式存活的是三类:
turn_state——两种模式都无条件break(:356)。@liuxiaocs7 指出这件事必须在写路径消失之前解决,这是对的。但在单个 PR 的 cutover 里不存在"之前":投影只从终态事件派生turn_state(runtime-event-read-model.ts:1182),根本没有running;而 recovery 靠latest.status === 'running'找出被中断的 turn(session-manager.ts:5236)。这个事实并没有从日志里丢失——有 opening、无终态事件的 invocation 就是当时在跑,runtimeInvocationOutcome已经是这么读的。补上这个缺口是同一个改动的一部分,不是它的前置。system_note——两种模式都丢;full模式只是多发一条skipped_high_risk_message诊断(:359)。那 8 种用户可见的 kind(session.ts:792)也一并丢,所以转换器自己处理这一类,不沿用 backfiller 的行为。- 未保留不透明
providerOutput的 provider-native 工具活动(skipped_provider_native_replay_gap)。
落在第三类里的东西降级成带标记的文本事件,而不是跳过。
是合成,不是伪造。 每个转换出来的 invocation 需要一个 opening 和恰好一个终态事件。两者都是
role: 'system', author: 'system',是关于 run 的事实而非 run 说过的话,而且都有现成的构造器:buildInvocationOpenedEvent和buildSyntheticTerminalRuntimeEvent,后者的契约里已经写明用于"迁移那些从未写过事件的 run header"。旧的system_note中error/abort两类用来决定终态事件的 status,不能映射成errorcontent 事件(非终态时是 hard 诊断)。event ID 从 StoredMessage 身份派生,导入因此幂等可续跑;转换出来的事件保留
refs.storedMessageId,让按值引用 message id 的 admission 记录仍然 join 得上。hard failure 移到决策路径。 cutover 之后不存在"顶替"的第二视图,所以展示侧的投影缺口变成一条带标记的占位行,沿用
archived_tool_result_placeholder已有的形状。recovery 和 revision 判定继续 fail closed,因为那里的缺口会改变破坏性结果。单个 PR 的范围
转换器只是同一个改动里的一部分。cutover 在同一个 PR 里迁移全部消费方:
- 读权威:把
session-transcript-reader.ts:61的 active/durable 分叉收成一条事件驱动的投影。 - recovery:
readMessagesForRecovery改读投影。校验靶子仍然是 admission 记录——它本来就持有内容,也本来就会在消息缺失时据此现场生成一条。 - 分页与 cursor、未读标记、
goal-coordinator.ts:159、session-manager.ts:3315的 subagent summary 回退,以及runtime-host-run-command.ts:313/405的 Agent Graph 消费方。 - session 级 notes 获得独立模型,不再寄居在执行账本里。
turn_state: running改为从 invocation 清单派生,recovery 仍然认得出被中断的 turn。
同一个 PR 里删除: 12 处执行写入点、
materializeRuntimeEventTranscriptProjection、repairSteeringMessagesOnce、RuntimeReadModel的投影缓存与漂移诊断、session-transcript-reader.ts:61的 active/durable 权威切换、readMessagesForRecovery的执行分支,以及session-manager.ts:1481那处重复的 revision 判定。最后一处是一个破坏性决定的第二份实现,该决定已由SessionRevisionCoordinator拥有,而它那版会先问 admission 权威;今天我构造不出两者分歧的可达路径,但一个破坏性决定不该有两份实现。RuntimeEvent 出现之前(2026-06-14 之前)创建、且仍在使用的 Session 是唯一一种转换事件拿不到合法 Session ordinal 的情况,因为 ordinal 严格追加分配。这些 Session 的 ordinal 按转换事件数整体平移。
这次单向迁移已按
CONTRIBUTING.md在 #4866 通告。技术讨论仍留在本 issue。两个代价,作为事实记录
- 迁移跑完不可回滚。降级到 cutover 之前的版本读不了自己的数据。epoch 管协议,不管数据。
- 升级窗口与历史总量成正比。需要实测上界,并且崩溃后可续跑——确定性 event ID 提供了这个能力。
来自 #4779 的相关发现
#4779 通过要求
summaryFormat标记,退役了未标记的 0.1.x 文本摘要契约。所有消费方都正确地 fail open,但execution-inspect把这些退役记录按 error 级别判为compaction_checkpoint_invalid,因为 superseded 分类器认的是source.policyVersion——一个 0.1.x 从未写过的字段(source是 #955 才有的)。于是每个带有前标记时代 checkpoint 的 Session 一 inspect 就报持久化损坏,真正损坏的记录反而淹没其中。这和双写是同一种失效模式:契约按当代的说法收紧,而在契约存在之前写下的记录拿不出闸口要的证明。无论这里最终采用什么方案,早于契约的记录都必须按它是什么来分类,而不是按一份它的写入方当时没有词汇去写的自述。- Three read-model diagnostics are
I am retracting my freeze watermark proposal and adopting the hard cutover. You are right that keeping a frozen prefix leaves a second read path in production forever. I accept your points without qualification:
RuntimeEventbecomes the single authority,StoredMessageis demoted to an API and UI projection, session notes move to their own explicit model, existing pagination cursors may be invalidated, and every consumer cuts over together.I also agree that we need exactly one converter rather than a schema migration running alongside an importer. Extending the Session Importer is the right approach, run once during upgrade in a single PR so the repository never carries a half-migrated state.
My only correction is on step 6. Reporting what cannot be reconstructed faithfully cannot mean skipping records. Transcript gaps break sessions in this codebase:
- The read model throws on hard diagnostics (
unsupported_event,incomplete_event,tool_use_id_mismatch), which blocks the session from opening (packages/runtime/src/runtime-event-read-model.ts:74). - During startup recovery, a session with
revisionState: 'preparing'is deleted outright if its messages lack a revision user message (packages/runtime/src/session-manager.ts:1481), turning a dropped row into silent data loss at next launch.
Therefore, the conversion must be a total function where every legacy row yields an event. Anything that cannot map to a structured event becomes a plain text event holding its rendered content, tagged with legacy provenance. Degrading structure to text preserves what was actually observed and does not fabricate execution facts.
On what is actually lost: @liuxiaocs7 has since shown that the tool facts do convert in
fullmode, and I have corrected my earlier claim in the plan comment above. Only three classes genuinely do not survive as structured events:turn_state, which the backfiller never converts;system_note, including the eight user-visible kinds; and provider-native tool activity whose opaque provider output was not retained. Those degrade to marked text events rather than being skipped.turn_statedeserves a note. The projection derives it only from terminal events and has norunningat all, while recovery finds interrupted turns throughlatest.status === 'running'. The fact is still in the log, since an invocation with an opening and no terminal event was running, which is howruntimeInvocationOutcomealready reads it. Under a single-PR cutover that work is part of the same change rather than a predecessor to it.Two operational costs should be recorded on this issue: the migration is irreversible once executed, and the upgrade duration scales with total history, requiring a measured upper bound and crash-resumable execution.
中文
我撤回之前的 freeze watermark 提案,并采纳 hard cutover 方案。你的核心反对意见是正确的:保留冻结前缀意味着在生产架构中永久保留第二条读取路径。我完全接受你的以下观点:以
RuntimeEvent作为唯一的权威数据源,StoredMessage降级为仅用于 API 和 UI 的投影类型,会话级别的 notes 采用独立的显式模型,现有的分页 cursor 可以在破坏性变更中失效,且所有消费方一次性同时切换。我也认同只需要一个转换器,而不是让 schema migration 和 importer 并存。扩展 Session Importer 是最合适的落脚点,在升级过程中单次调用,并在单个 PR 中完成,以确保仓库不会处于只迁移了一半的状态。
我唯一的修正针对步骤 6。报告无法如实重建的信息绝不能等同于跳过记录。跳过会导致 transcript 产生空洞,而在这个代码库中,空洞会导致会话异常:
- 三个 hard 诊断(
unsupported_event、incomplete_event、tool_use_id_mismatch)会直接导致读取器抛出异常,从而无法打开会话(packages/runtime/src/runtime-event-read-model.ts:74)。 - 在启动恢复期间,处于
revisionState: 'preparing'的会话如果其消息中缺失 revision 用户消息,会被直接删除而不仅是隐藏(packages/runtime/src/session-manager.ts:1481)。丢掉一行就等于下次启动时的一次无人值守删除。
因此,转换逻辑必须是一个全函数(total function):每一行旧数据都必须产生对应事件。任何无法映射为结构化事件的数据,都转为包含其渲染后文本的纯文本事件,并标记旧版溯源信息。将结构降级为文本保留了记录中实际观察到的事实,并不等同于凭空捏造执行事实。
关于究竟丢了什么:@liuxiaocs7 随后指出工具事实在
full模式下是能转换的,我已在上面那条方案评论里更正了先前的说法。真正不能以结构化事件形式存活的只有三类:turn_state,backfiller 从不转换它;system_note,包括那 8 种用户可见的 kind;以及未保留不透明 provider output 的 provider-native 工具活动。这三类降级成带标记的文本事件,而不是被跳过。turn_state值得单说一句。投影只从终态事件派生它,根本没有running;而 recovery 靠latest.status === 'running'找出被中断的 turn。这个事实仍然在日志里——有 opening、无终态事件的 invocation 就是当时在跑,runtimeInvocationOutcome已经是这么读的。在单个 PR 的 cutover 下,这块工作属于同一个改动,而不是它的前置。我希望在 issue 中记录两项实际代价:migration 一旦执行便不可逆;升级耗时与历史总量成正比,因此需要经过测算的时间上限,并且必须支持崩溃后断点续跑。
- The read model throws on hard diagnostics (
Thanks for reweighing this, and for the two catches — the totality point and the
preparing-session deletion are exactly the kind of thing a naive converter would trip on.Both corrections make sense — adopting them.
Agreed the conversion has to be total: degrade-to-text with legacy provenance rather than skip. Given the read model throws on hard diagnostics (
runtime-event-read-model.ts:74) and recovery deletes apreparingsession missing a revision user message (session-manager.ts:1481), skipping there isn't "reporting," it's silent loss.One thing that reinforces the single-PR + crash-resumable point: derive the ids for the degraded text events deterministically too (from source-row identity + legacy tag), not just the tool events. Then a migration that dies halfway re-runs idempotently instead of double-writing, which is what makes the crash-resumable requirement cheap. And since it's irreversible, a validate-by-projection pass — project each converted session through the new reader before committing — is cheap insurance against shipping a converter that produces sessions which won't open.
On
turn_state: running— agreed, the fact is already in the log and this belongs in the same PR, not ahead of it. The concrete in-scope piece is just: the projection emits a running row from an invocation opening with no terminal event, and recovery stops readinglatest.status === 'running'off StoredMessages and reads that shape instead — the same reasoningruntimeInvocationOutcomealready uses.+1 on recording both operational costs on the issue.
中文
谢谢你重新权衡这件事,也谢谢这两个提醒——全函数那点、还有
preparing会话被删那条,恰恰是天真的转换器最容易踩的坑。两处修正都成立,采纳。
同意转换必须是全函数:无法结构化的降级成带 legacy 溯源的文本事件,而不是跳过。毕竟读模型遇 hard 诊断会抛异常(
runtime-event-read-model.ts:74),恢复又会把缺 revision 用户消息的preparing会话直接删掉(session-manager.ts:1481);在这里跳过不是“报告”,是静默丢数据。有一点能强化你说的单 PR + 崩溃续跑:那些降级的文本事件,id 也一并做确定性派生(用源行身份 + legacy 标记),不只 tool 事件这么做。这样迁移中途挂了可以幂等重跑、而不会双写——崩溃续跑这个要求也就变便宜了。又因为不可逆,加一道 validate-by-projection(每个转换后的会话先用新 reader 投影一遍再提交)是很便宜的保险,免得发出去的转换器产出打不开的会话。
关于
turn_state: running——同意,这个事实本就在日志里,属于同一个 PR、不是它的前置。真正在范围内的就一点:投影从“有 opening、无终态事件的 invocation”产出 running 行,恢复不再从 StoredMessage 读latest.status === 'running',改读这个形态——跟runtimeInvocationOutcome现在的读法同理。两项操作代价记到 issue 上,+1。
- added 7 commits that reference this issue
on Sep 6, 2026 - added a commit that references this issue
on Sep 7, 2026
The problem in one picture
Imagine recording the same conversation in two notebooks. While the conversation is happening, we read notebook A. As soon as it finishes, we switch to notebook B.
Even if both notebooks are meant to say the same thing, they are written separately. A crash, retry, or reconnect can catch them at different points.
That is the current transcript architecture:
flowchart LR subgraph Today["Today: one conversation, two ledgers"] direction TB Run1["Runtime execution"] --> Events1["RuntimeEvent"] Run1 --> Messages1["StoredMessage"] Events1 --> Active["Active transcript"] Messages1 --> Completed["Completed pages"] end subgraph Target["Target: one ledger, many views"] direction TB Run2["Runtime execution"] --> Events2["RuntimeEvent"] Events2 --> Projection["Bounded transcript projection"] Notes["UI notes"] --> Projection WorkHub["WorkHub facts"] --> Projection Projection --> View["StoredMessage view"] endWhat happens today
Ordinary runtime conversation facts are persisted through two paths:
RuntimeEventrecords the execution timeline.AiSdkTurn,ToolRuntime, andAgentRunalso persist equivalentStoredMessageconversation records.Those writes do not share one commit order. The read path then switches authority at the turn boundary:
RuntimeEvent;StoredMessagerecords.The same conversation can therefore be reconstructed from different durable representations before and after completion. Recovery, reconnect, pagination, and migration all depend on those representations remaining equivalent even when their writes settle in different orders.
The architectural rule
RuntimeEventis the single durable semantic authority for ordinary Session execution facts.Everything else has one clear role:
RuntimeEventtimeline. A page must not require loading the full Session.SessionEventremains live transport for active clients; it is not a second storage authority.StoredMessageremains a presentation and protocol DTO. It is not being deleted; normal runtime execution simply stops persisting it as an equivalent second ledger.RuntimeEvent.What changes
RuntimeEventauthority.StoredMessagerecords is removed fromAiSdkTurn,ToolRuntime, andAgentRun.What does not change
StoredMessageremains the client-facing presentation/protocol shape.SessionEventremains the live transport shape.Compatibility and cutover
Acceptance criteria
RuntimeEventauthority.StoredMessageconversation records is removed.Why this is separate from #4779
#4779 establishes the provider-history boundary: what model history is replayed, compacted, and summarized for the next provider request.
This issue changes durable transcript storage and paging. It has different failure modes, compatibility concerns, and recovery tests. Folding it into #4779 would make two authority migrations share one review surface without making either safer.
Out of scope
StoredMessageas a presentation/protocol type.RuntimeEvent.