Skip to content

fix(storage): bound batched clear purges by generation instead of time - #498

Merged
qxcnm merged 1 commit into
qxcnm:mainfrom
xcosmosbox:fix/requestlog-purge-generation-bound
Oct 9, 2026
Merged

qxcnm merged 1 commit into
qxcnm:mainfrom
xcosmosbox:fix/requestlog-purge-generation-bound

Conversation

@xcosmosbox

Copy link
Copy Markdown
Contributor

概述

#494 合并后的数据安全跟进,对应 #494 下的评论。

#494 的清理标记用秒级 now_ts() 区分清空前后的数据,删除条件是 rowid <= max_rowid AND created_at <= cleared_at。SQLite 在最高 rowid 那一行被删除后会复用这个 rowid,所以清空提交后、同一秒内写入的新行可能同时满足两个条件:它会立刻被详情页隐藏,随后被后台分批删除删掉。原有的 reused_rowid_after_clear_is_not_purged 只覆盖了 cleared_at + 5,没有覆盖同一秒。

本 PR 改为用清理代次严格区分清空前后的数据,不再依赖时间。

改动

  • 迁移 143:给四张会被分批删除的请求内容表(预览、manifest、上游尝试、响应 ID 关联)和 request_log_payload_purges 各加一个可空的 generation 列。只是 ALTER TABLE ADD COLUMN,不重写已有数据。迁移中途中断后重跑时,会补齐还没加上的列。
  • 写入记录代次:四处写入都在同一条语句里取当前清理代次((SELECT generation FROM request_log_payload_state WHERE id = 1)),与另一连接上的清空保持原子;预览覆盖写入时也会更新代次。
  • 清空记录代次:清空在标记里记下本次提升后的代次。标记判断一行属于清空前的条件改为「代次低于标记的代次(或是迁移前写入的行)」,清空之后写入的行代次一定不低于标记,无论 rowid 是否复用、是否在同一秒。删除与读取时的可见性判断用同一个条件。
  • 兼容旧标记:升级前留下、尚未删完的标记没有代次,仍按原来的时间条件判断,但只作用于迁移前写入的行,不会误伤升级后的新行。再次清空时会合并成带代次的标记。
  • 删除完成后按标记的全部字段(含代次)比较再删除标记,期间如有新的清空合并进来,不会误删标记。

改动范围

  • Frontend
  • Desktop / Tauri
  • Service
  • Gateway / Protocol Adapter
  • Docs / Governance
  • Workflow / Release

主要文件

  • crates/core/migrations/143_request_log_payload_purge_generation.sql(新)
  • crates/core/src/storage/request_log_payload_purge.rs、request_log_payload_store.rs、request_logs.rs、mod.rs
  • crates/core/src/storage/request_log_payload_purge_tests.rs

验证

本机 2 vCPU / 3.6 GB RAM,构建与测试以单核、低优先级、串行方式运行,基于当前 main(86aa0108)。

  • cargo fmt --all -- --check,WebSocket 依赖 pin 检查
  • cargo test -p codexmanager-core:lib 487(新增 3 个)、集成测试 48,全部通过
    • same_second_rowid_reuse_after_clear_is_not_purged:四张表都先写入一行,清空,删掉这些行让 rowid 被复用,再在与清空同一秒(created_at == cleared_at)写入新行;断言新行确实复用了清空范围内的 rowid、写入时带上了当前代次、清空后立刻可读、分批删除后仍然保留。把判断条件临时改回旧的时间条件后,这个测试会失败(新行在删除前就已不可见)
    • legacy_clear_marker_only_purges_rows_written_before_the_upgrade:升级前留下的无代次标记只删除迁移前写入的行,同秒、同 rowid 范围内的新行保留
    • generation_migration_finishes_when_columns_already_exist:迁移在列已存在时也能完成并记录
  • cargo check --workspace --all-targets,改动文件无新增 warning
  • cargo test --workspace:未执行。本机链接 service 测试二进制会耗尽内存,CI 的 Test SQLite storage maintenance 步骤会跑 core 全量测试

风险与影响面

备注

  • 不包含敏感 token、cookie、API key

A clear marker doomed rows with rowid <= max_rowid AND created_at <=
cleared_at. created_at has one-second resolution and SQLite reuses the
highest rowid once that row is deleted, so a payload row written in the
same second as the clear could match both conditions, be hidden from the
detail view and then be deleted by the background purge.

Migration 143 adds a nullable generation column to the four purged payload
tables and to request_log_payload_purges. Every write stores the clear
generation current in its own statement, and a clear stores the
generation it bumped to. A clear marker now dooms rows of a lower
generation (plus rows written before the migration), which no row written
after the clear can have. Markers left by an older version keep the
time-based condition, restricted to rows written before the upgrade.
Deleting and reading use the same condition.

Add regressions for rowid reuse with created_at == cleared_at in every
purged table, for legacy markers, and for re-running the migration when
the columns already exist.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants