Store.correct() has a TOCTOU race between find and write #16
Description
Activity
I think this one can also be closed. It was fixed by b49f1e6 ("fix: serialize corrections through shared store write path", 2026-09-17), which like #15 didn't reference the issue, so it stayed open.
Scope note: everything below is code reading on current main plus the commit diff — 仅代码阅读层面的个人判断. The tests themselves I did not run(未验证,我没实测过:Windows 机器,store lock 用 fcntl);them passing is CI's word.
What I checked:
correct() now runs the whole read-modify-write under one store_lock: find, validate, mutate, then a new _write_locked helper that writes without re-acquiring the lock (store.py, the correct() around line 355).
The same commit added tests/unit/test_correction_concurrency.py, which drives the exact race from the report: one process pauses inside its find() while a second process corrects the same record, then asserts both updates survive and the search index stays consistent. There is also a delete-racing-corrections case where retained provenance must equal successful corrections.
delete, merge, feedback and record_many hold the lock across their find+write in the same shape, again per code reading.
On that reading, the three-step window described in the report no longer exists on main — 需要确认的话以 CI 和维护者复验为准。
Store.correct() has a TOCTOU race between find and write
Severity: High
File:
packages/core/src/agent_memory/core/store.py:177correct()reads the current record viaself.find(name)(no lock), then acquiresstore_lockonly for the provenance append, then callsself.write()which acquiresa second, independent lock. Between the initial
find()and the finalwrite(),a concurrent writer (another process or adapter hook) can modify the same record, and
correct()will silently overwrite those changes.Why it matters
Two concurrent agents calling
correcton the same memory lose each other's edits.The
record_manybatch path holds the lock for the entire operation (store.py:96), butcorrectsplits it into three non-atomic steps. A sleep cycle running_normalise_datesor
_settle_weightsconcurrently with a user-initiatedcorrectwill produce the samelost-update pattern.