Repository navigation
perf(storage): replace clear-time VACUUM with batched purge and incremental reclaim - #494
Merged
Merged
Conversation
Transaction::execute no longer issues an extra SELECT last_insert_rowid() after every statement. Connection and Transaction now query the rowid on demand from the exact physical handle that ran the INSERT (the held transaction handle, or the single pooled handle), falling back to the last rowid observed from statement results if the lookup fails. Add shim tests for rowid semantics and run them in CI.
SQLx pools default to a 30 min max lifetime and 10 min idle timeout, which silently replaced the single pooled handle. That lost connection-scoped state (last_insert_rowid, busy_timeout, temp_store, journal_size_limit) and wiped in-memory databases. Disable both so a Connection keeps one handle for its whole life, like a real sqlite3 handle.
…mental reclaim Clearing request logs ran `wal_checkpoint(TRUNCATE); VACUUM` after deleting every payload row inside one transaction. With full payload storage that rewrote a multi-GB file, needed as much free disk, held the write lock long enough for gateway writes to hit the 3 s busy timeout, and with temp_store=MEMORY could hold the transient copy in RAM. - clear/prune now only make old data unreachable in a short transaction (generation bump or retention watermark plus per-table rowid bounds in the new request_log_payload_purges table, migration 142) and delete payload rows in adaptive batches with secure_delete enabled; blobs are swept with a reference check inside each delete transaction - rows waiting for the purge are invisible to every reader and are never chosen as parents of new manifests; work left over after the 10 s budget (or after a restart) is continued by a background maintenance thread, and errors after the first commit no longer fail the clear - the maintenance thread also runs retention pruning off the request path, retries busy WAL checkpoints and returns free pages with small incremental_vacuum steps; new databases use auto_vacuum=INCREMENTAL - every connection sets journal_size_limit; space accounting uses live pages (page_count - freelist_count), never the file size - admin-only storage/spaceUsage and storage/reclaim RPCs and a settings card show used/reclaimable space and offer an explicit, disk-checked one-time rebuild for databases created by older versions
This was referenced Oct 7, 2026
When a path did not exist, normalized_components canonicalized its parent and dropped the last component, so a missing mount point such as /definitely-not-a-mount compared equal to / and could win the longest-prefix match. Resolve the deepest existing ancestor instead and re-append the components below it.
Owner
|
补充一项合并后的数据安全跟进:复核 #494 的批量清理边界时发现,清理标记使用秒级 建议尽快补一个 |
Contributor
Author
👌 |
4 of 10 tasks
Contributor
Author
|
@qxcnm 谢谢指出,同一秒内复用 rowid 被误删的问题确实存在,修复单独提在了 #498:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
合并顺序
这是一组 4 个 PR,请按以下顺序合并:
perf(rusqlite)SQLite 垫片:按需读取 rowid、固定物理连接perf(storage)清空日志去掉 VACUUM,改为分批删除与增量回收feat(core)请求内容组提交与磁盘溢出的 core 基础设施feat(requestlog)请求内容捕获切换到有界后台写队列后续 PR:#495 依赖本 PR。
概述
现在清空请求日志时,会在一个事务里删除全部请求内容,然后执行
PRAGMA wal_checkpoint(TRUNCATE); VACUUM;。开启完整请求内容存储后数据库可能有数 GB:VACUUM 要重写整个文件、需要同等大小的空闲磁盘,持有写锁的时间足以让网关写日志等满 3 秒 busy timeout 而报错;在temp_store=MEMORY下临时副本还可能放在内存里。本 PR 去掉清空时的 VACUUM:清空只在短事务里让旧数据不可见,物理删除改为限时分批进行,空闲页由后台小步归还给系统。
改动
request_log_payload_purges(迁移 142)。之后按 rowid 从小到大分批删除:每批 50–4000 行自适应,目标 60 ms,批间停 15 ms;删除期间临时开启secure_delete,结束后恢复。blob 在每个删除事务内检查引用后回收。storage-maintenance:每 30 秒一轮,负责保留期裁剪(不再在网关写日志的线程上同步执行)、续删(每轮最多 5 秒)、wal_checkpoint遇到 busy 时重试,以及在可回收空间不少于 64 MiB 且占 10% 以上时用incremental_vacuum小步归还空间(每步 16–8192 页,目标 50 ms,步间停 200 ms,每轮最多 5 秒)。auto_vacuum=INCREMENTAL,老库不自动转换。每个连接设置journal_size_limit=64 MiB,限制 checkpoint 后留下的 WAL 文件大小。空间统计按page_count - freelist_count计算,不看文件大小。storage/spaceUsage、storage/reclaimRPC(Web 鉴权为 password 模式时同样拒绝成员),以及对应的 Tauri 命令和 Web 映射。设置 → 网关新增「数据库空间」卡片:显示已用 / 可回收空间、数据库与 WAL 文件大小、自动回收模式和后台删除状态,可「立即回收」;老库提供一次性「整理数据库」(启用增量回收并 VACUUM),执行前检查数据目录和临时目录的空闲空间,分批删除未完成时拒绝执行。远程数据库模式下显示不支持。Test SQLite storage maintenance步骤:core 全量测试和 service 维护模块单测。改动范围
主要文件
crates/core/src/storage/request_log_payload_purge.rs(新)、storage_space.rs(新)、request_logs.rs、request_log_payload_store.rscrates/core/migrations/142_request_log_payload_purges.sqlcrates/service/src/storage/maintenance.rs(新)、rpc_dispatch/storage_space.rs(新)、rpc_dispatch/mod.rs、lifecycle/startup.rs、requestlog/seaorm.rsapps/src/app/settings/components/storage-space-card.tsx(新)、apps/src/lib/api/storage-space.ts(新)、apps/src-tauri/src/commands/storage_space.rs(新)验证
本机 2 vCPU / 3.6 GB RAM,构建与测试以单核、低优先级、串行方式运行,均在叠加了 #493 的本分支上执行。
cargo fmt --all -- --check,WebSocket 依赖 pin 检查cargo test -p rusqlite:15/15 通过cargo test -p codexmanager-core:lib 484、集成测试 48,全部通过。新增回归覆盖:清空不执行 VACUUM 并恢复secure_delete;未脱敏内容清空后不残留在数据库和 WAL 文件字节中;分批删除不误删清空后写入的数据,rowid 复用也不误删;时间预算用尽后续删;保留期裁剪重基存活子请求、不复用过期父请求;跨越保留期的预览在删除前即不可见;同一文件的第二个删除立即返回;分批删除未完成时拒绝整理;新库为 INCREMENTAL、老库不自动转换、增量回收与显式转换、每个连接限制 WAL 大小cargo check --workspace --all-targets(包含 service 维护模块单测的类型检查),改动文件无新增 warningnode --test --test-concurrency=1:前端 runtime 测试 250/250 通过pnpm -C apps run build:desktopcargo test --workspace:未执行,原因同上风险与影响面
secure_delete覆盖)。备注