Skip to content

Transient "database is locked" during an event write fails the run but leaves the Claude process running unsupervised #15542

Description

@areidyOTH

What happened

A Claude thread failed about 80 s into its first turn with "Provider error: Provider turn failed." Claude itself had not failed: the underlying claude process kept working with no one watching it. It ran tools (with permission checks bypassed) for another 14+ minutes, and it set up a background systemd timer that competed with a replacement thread the user had started. The UI showed the thread as failed the whole time, and nothing in T3 could stop the process. In the end it had to be found and paused by hand (SIGSTOP).

Diagnosis

  • At 06:59:27.848 UTC, orchestrationV2.EventSink.write (called from orchestrationV2.providerTurnStart.start) for the thread failed after 458 ms with TurnItemPositionStoreError ← SqlError ← Error: database is locked. The transaction was rolled back. The run was then marked failed at 06:59:28.314 with the item "Provider error / Provider turn failed."
  • busy_timeout = 5000 is set (apps/server/src/persistence/Layers/Sqlite.ts:19), yet the write failed after about 0.46 s. That pattern suggests SQLITE_BUSY coming back immediately when a deferred transaction tries to upgrade to a write lock while another connection holds it (the busy handler isn't used in that case). This is inference; not verified. Contention was high at that moment: seven thread launches with worktree provisioning were running, and an external helper was issuing and revoking sessions through t3 auth session issue/revoke CLI processes, which write to the same database.
  • A brief lock while saving one event failed the whole run, instead of the write being retried.
  • When the run failed, the provider session was not stopped. The provider event log for the same providerSessionId keeps receiving assistant / tool_use / tool_result events until 07:13 UTC (14 min after the failure), and the claude --output-format stream-json … process was still alive afterwards.

Steps to reproduce

  1. Start a Claude thread with a long first turn (many tool calls).
  2. While it runs, put write pressure on the database: launch several threads at once, and/or run t3 auth session issue --ttl 5m --json / t3 auth session revoke in a loop from another process.
  3. Once an event write hits database is locked, the thread changes to "Provider turn failed", while ps still shows the claude process running and the provider events log keeps growing.

Version

0.0.46-nightly.20261004.2644 (desktop AppImage, commit 7379933)

Environment

Linux x64 7.0.0-38-generic, Node v26.8.2, claude 2.1.289 (Claude Agent SDK provider, full-access runtime mode)

Evidence

# server.trace.ndjson, span orchestrationV2.EventSink.write, durationMs 458, thread_id <coordinator>
TurnItemPositionStoreError:
    at orchestrationV2.EventSink.write (binCli-*.mjs:83891:21)
    at orchestrationV2.providerTurnStart.start (binCli-*.mjs:229232:59)
    at ServerRuntimeStartup.startEffectWorkerWithRelay (binCli-*.mjs:239491:79)
  [cause]: effect/sql/SqlError: Failed to execute statement
    [cause]: effect/sql/SqlError/UnknownError: Failed to execute statement
        at classifySqliteError (binCli-*.mjs:57423:9)
      [cause]: Error: database is locked
# sibling span sql.transaction -> event db.transaction.rollback, db.name=statev2.sqlite

# thread snapshot
recentRuns: [{status: "failed", startedAt: 06:58:04.087Z, completedAt: 06:59:28.314Z}]
item: {type: "error", title: "Provider error", text: "Provider turn failed."}

# provider/events.<thread>.log, same providerSessionId, after the failure
[06:59:28.449Z] assistant tool_use Bash ...
[06:59:46.157Z] system thinking_tokens ...
...
[07:13:11.038Z] (last event; 3,539 lines total)
$ ps -o stat,etime,args -p <pid>
Tl 24:54 claude --output-format stream-json --verbose --input-format stream-json ... --model claude-opus-5-5[1m] ...

Related issues

#15447 (a Cursor run keeps going after T3 marks it failed): the same "failed run, provider keeps running" problem, seen here with Claude and triggered by a database error rather than a subagent failure. #5099 (busy_timeout never set) was fixed by setting busy_timeout, but the immediate lock failure under load still happens. #6097 (second backend sharing the database) does not apply here; only one backend process had statev2.sqlite open.

Fix applied or workaround

None in T3. The user's replacement coordinator paused the orphaned claude process with SIGSTOP; it is left in place for inspection.

Filed by

Claude Code (claude-opus-5-5) via t3 triage

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions