Skip to content

A reset in the first moments after startup races the what's-new notification's write #419

Description

@DutchJaFO

Description

At startup the what's-new notification (#81) is written on a detached task (Program.cs, the Task.Run
beside WhatsNewNotification.SeedAsync), using the application-version id captured before the task
starts. A POST /admin/database/reset that lands before that write rebuilds the database, removing that
version row, and the write then fails. The notification is lost for that boot, and the log carries an
exception the operator did not cause. Found running #411's automated-test documents.

Reproduction steps

  1. Start a fresh container of the current build with docker run (not test-env.csx, whose one-second
    health poll gives the write time to finish), and poll /api/v1/health in a tight loop.
  2. The instant health answers 200, POST /api/v1/admin/database/reset?allowNoBackup=true with the key.
  3. Read the log.

Reproduced 5 of 5 with the tight poll; 0 of 5 with the one-second poll (negative control).

Expected behaviour

No exception. Either the what's-new write completes against the version it was computed for, or it is
recomputed after the reset, never a write against a row the reset has removed.

Actual behaviour

SQLite Error 19: 'FOREIGN KEY constraint failed', one exception id logged ten times as it is rethrown,
then [Server] Failed to seed the #81 what's-new notification; non-fatal, startup continues. The
what's-new notification is absent until the next restart.

Knowledgebase

The what's-new notification could not be seeded, after a reset
(docs/knowledgebase/whats-new-notification-fails-to-seed-after-a-reset.md), status QTN-KNOWN.

Update it as part of this issue. The condition exists only in development builds, so when the fix
removes it the entry is deleted rather than retired, per docs/knowledgebase.md's retention rule:
its commits are its history. If the fix changes what an operator sees instead of removing it, rewrite
the Symptom, Cause and Remedy to match and drop the status code.

Failing tests

To be written: a unit test that runs a reset between version capture and the what's-new write and
asserts no exception and a notification afterwards; the live reproduction above as its automated
counterpart, run red first.

Test class Test method Status before fix
(to be written) WhatsNew_ResetBeforeWrite_WritesNothingStaleAndThrowsNothing ❌

Definition of done

  • Failing test(s) listed above are red before the fix is written
  • Fix implemented
  • All listed tests pass (green)
  • No regression in related tests
  • Findings summarised in a closing comment

Activity

  1. added this to the Notification system milestone on Sep 22, 2026
  2. DutchJaFO commented on Oct 3, 2026

    @DutchJaFO
    OwnerAuthor

    Implemented and verified — Waiting for release. T1 passed 2026-10-04 on both schema paths.

    The fix is not the one this issue describes, and is broader than it. This asks for the what's-new producer to stop racing a reset. What was built is a rule about the application: no external write reaches the database until startup has finished its own work. The race here is one consequence of the gate opening too early, and the changelog import sat in exactly the same position unnoticed; fixing the producer would have left it there.

    Two things above are superseded. The mechanism is stale — it names the Task.Run beside WhatsNewNotification.SeedAsync, which #424 had already replaced with StartupBackgroundWork; that made the host wait at shutdown and did nothing for this. The second remedy is declined — "recomputed after the reset" would mean a reset reseeding notifications, which breaches CLAUDE.md's endpoint side-effect policy. Nothing is reseeded. What this issue asked for is still delivered, by construction: the write cannot race a reset because a reset cannot arrive until the write has finished.

    A larger defect was found while measuring it. A write during startup was answered 200 with an HTML wait page: an API caller was told its reset had succeeded, handed a web page, and nothing was reset. That is deterministic across the whole startup window rather than a sub-second race. It is now 503, with the body chosen by Accept. /api/v1/version is gated for the same reason and /api/v1/health is the only exempt path.

    The original race did not reproduce here, 0 of 3, against the reporter's 5 of 5 — the window had narrowed, not closed. That is why the evidence is deterministic tests rather than a live reproduction.

    The Knowledgebase entry is deleted, not retired, per docs/knowledgebase.md: its condition never reached a release.

  3. added 13 commits that reference this issue on Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions