Skip to content

fix(e2e): select a farm-local expense range across month boundaries - #1011

Merged
mforce merged 3 commits into
mainfrom
fix/1009-expense-specs-month-boundary
Oct 1, 2026
Merged

mforce merged 3 commits into
mainfrom
fix/1009-expense-specs-month-boundary

Conversation

@mforce

@mforce mforce commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Expense specs fail at the start of a farm month because the default range excludes the seeded rows. The phone optical-size test has the same cause: an empty month renders EmptyState instead of the expense ledger.

openSeededExpenses opens the page and fills both date filters from 30 days before farm-local today through today. It uses the existing farm clock, translated labels, and the phone filter dialog. It waits for the expense response carrying the requested dates, then for the rendered ledger. Listening before navigation also handles a 31-day month-end when the default range already matches. The window assumes a fresh simulation seed and covers its named expenses, dated 2–7 days before the UTC seed anchor. The preflight checks recent eggs, not expenses; an old database can still pass that check while these expense rows have aged out. The seeder and application are unchanged.

The review identified a possible mid-month race where a caller could use the old ledger before the range reload remounted it. I did not reproduce that flake; the helper now owns the response and rendering waits instead of relying on callers.

I checked every /expenses reference. The other callers measure the desktop summary and category controls, which do not require expense rows.

Closes #1009

Verification

  • Live October 1 in America/Chicago: all four affected cases failed before and passed after. All three CI e2e-smoke shards passed on the initial head ab284d7f on the same real date; CI is the authoritative clean-runner result.
  • Final head 1880d761: all three E2E smoke shards passed on the real October 1 Chicago date, including the response wait and immediate rejection handler. CodeQL and every other non-image check passed, including integration and web coverage/build. The CI run is red only because amd64 found ci: move image vulnerability scanning out of CI to a weekly scan of published images #1006’s existing OpenSSL CVE-2026-84782; fail-fast cancelled arm64 and image publishing was skipped.
  • Deterministic February 1 reproduction: the exact advancing-clock recipe below produced four expected baseline failures, then four passes in 14.3 seconds on 83b153ac, before the final promise-handling line.
  • Fresh full quick suite on 83b153ac (before the final promise-handling line): 134 passed, 1 intentionally skipped, 0 failed (6.3 minutes). No reruns. Host one-minute load fell from about 5.6 to 2.2 during the observed run.
  • Earlier full suite on ab284d7f: 133 passed, 1 skipped, 1 timeout (13.3 minutes), with host load around 50–58 on 12 CPUs. The unrelated desktop day-readout-placement.spec.ts:173 English day-strip visibility assertion timed out. An immediate sequential rerun of that case and its phone counterpart passed in 2.2 and 2.1 seconds; no code or deadline changes. Both also passed in the final fresh suite. The skip is the opt-in real token-lifetime test.
  • Final review follow-up handles the response promise immediately, so an earlier navigation/filter failure cannot leave its rejection unhandled; the later await loaded still propagates response failures. A disposable three-case Playwright probe forced navigation to fail before that await. Before the fix, the broken case reported both the navigation error and page.waitForResponse: Test ended; afterward it reported only the intended navigation error. The following two cases passed in both runs, so this probe reproduced the extra rejection, not the wider file cascade.
  • Harness typecheck and git diff --check pass. The unrelated image-scan failure is tracked in ci: move image vulnerability scanning out of CI to a weekly scan of published images #1006; this PR does not change or suppress it.

Repeat the month-boundary proof on any day

I executed these setup, before/after, and browser-probe commands as written on an Ubuntu 26.04 host. The host-downloaded library loaded successfully in the Ubuntu 24.04 app image, including its seeder. These Linux commands use an isolated cw-1009-fixed compose project on ports 18092 and 18900, in a dedicated checkout of this PR. They create only disposable, untracked files. Do not run them while another user is driving that same project. Prerequisites are Docker, Node/npm, Python, and Debian/Ubuntu's apt and dpkg-deb.

FAKETIME='@2026-02-01 12:00:00' starts an advancing wall clock at February 1, not a stopped clock. The API and every one-shot process, including the seeder, inherit it from the app service. The Playwright runner and its Chromium children get the same setting below. All see February 1 in the farm's America/Chicago zone. FAKETIME_DONT_FAKE_MONOTONIC=1 preserves real monotonic timers; FAKETIME_DONT_RESET=1 keeps the runner's child processes on its advancing timeline.

Both stopped-clock and advancing-clock trials reproduced the original failures. The first recipe let libfaketime enable its automatic monotonic workaround, which caused CPU spinning and dialog timeouts on this host. FAKETIME_FORCE_MONOTONIC_FIX=0 removed that overhead: the same desktop correction case passed in 4.3 seconds with the normal 45-second deadline. The recipe below includes this setting in both the app container and the runner. No spec timeout or assertion was relaxed. No clock interception becomes part of the application, simulation harness, or CI.

# Run from the PR checkout's root.
mkdir -p /tmp/cw-1009
(cd /tmp/cw-1009 && apt download libfaketime && dpkg-deb -x libfaketime_*.deb faketime)
export CW1009_FAKE_LIB_DIR="$(dirname "$(find /tmp/cw-1009/faketime -name libfaketime.so.1 -print -quit)")"
bash tools/simulation/bootstrap.sh
python3 - <<'PY'
from pathlib import Path
import os
s = Path('tools/simulation')
c = (s / 'docker-compose.sim.yml').read_text()
c = c.replace('cluckwork-sim', 'cw-1009-fixed')
c = c.replace('127.0.0.1:8081:8080', '127.0.0.1:18092:8080')
c = c.replace('127.0.0.1:8889:8889', '127.0.0.1:18900:8889')
c = c.replace('./out:/app/sim-cast', './out-fixed:/app/sim-cast')
c = c.replace('      ASPNETCORE_URLS:', '''      LD_PRELOAD: /faketime/libfaketime.so.1
      FAKETIME: "@2026-02-01 12:00:00"
      FAKETIME_DONT_FAKE_MONOTONIC: "1"
      FAKETIME_FORCE_MONOTONIC_FIX: "0"
      TZ: UTC
      ASPNETCORE_URLS:''')
c = c.replace('      - ./out-fixed:/app/sim-cast',
    '      - ./out-fixed:/app/sim-cast\n      - ' + os.environ['CW1009_FAKE_LIB_DIR'] + ':/faketime:ro')
(s / 'docker-compose.1009-fixed.yml').write_text(c)
r = (s / 'reset.sh').read_text().replace('cluckwork-sim', 'cw-1009-fixed')
r = r.replace('APP_PORT=8081', 'APP_PORT=18092')
r = r.replace('docker-compose.sim.yml', 'docker-compose.1009-fixed.yml')
r = r.replace('OUT_DIR="$SCRIPT_DIR/out"', 'OUT_DIR="$SCRIPT_DIR/out-fixed"')
(s / 'reset-1009-fixed.sh').write_text(r)
PY
sg docker -c 'bash tools/simulation/reset-1009-fixed.sh'
cd tools/simulation/ui
npm ci
npx playwright install chromium
at_boundary() {
  env LD_PRELOAD="$CW1009_FAKE_LIB_DIR/libfaketime.so.1" \
    FAKETIME='@2026-02-01 12:00:00' FAKETIME_DONT_FAKE_MONOTONIC=1 \
    FAKETIME_FORCE_MONOTONIC_FIX=0 FAKETIME_DONT_RESET=1 TZ=UTC BASE_URL=http://127.0.0.1:18092 "$@"
}

# Run the unchanged specs from the PR's base against the same stack.
mkdir -p repro-specs
for spec in dialog-actions optical-size phone; do
  git show "05d7e7a7ea7389c32a830aa80e99c73489532a09:tools/simulation/ui/specs/$spec.spec.ts" > "repro-specs/$spec.spec.ts"
done
cat > playwright.repro.config.ts <<'TS'
import config from "./playwright.config";
export default { ...config, testDir: "./repro-specs" };
TS
at_boundary npx playwright test --config playwright.repro.config.ts \
  --grep 'Expenses correction|real body text|/expenses shows ten'
# Expected: 4 failures waiting for the missing seeded row or ledger.

at_boundary npx playwright test specs/dialog-actions.spec.ts specs/optical-size.spec.ts specs/phone.spec.ts \
  --grep 'Expenses correction|real body text|/expenses shows ten'
# Expected: 4 passed.

To verify the browser clock itself, run this in the same shell. page.evaluate executes in Chromium's renderer, not Node. The suite uses Playwright's default launch with --no-sandbox; this does not claim coverage of sandboxed Chromium.

at_boundary node --input-type=module - <<'JS'
import { chromium } from '@playwright/test';
const browser = await chromium.launch();
const page = await browser.newPage();
for (let i = 0; i < 2; i++) {
  console.log({ runner: new Date().toISOString(),
    renderer: await page.evaluate(() => new Date().toISOString()) });
  await new Promise(resolve => setTimeout(resolve, 1000));
}
await browser.close();
JS

The advancing probe returned runner 2026-02-01T12:00:00.454Z and renderer 2026-02-01T12:00:00.459Z, then 12:00:01.462Z and 12:00:01.463Z. Both advanced on the same timeline. In the stopped-clock diagnostic, runner and renderer both returned exactly 2026-02-01T12:00:00Z; the API minted a JWT expiring at 12:15:00Z, the seed manifest recorded generatedAtAnchor=2026-02-01, and the UI range was February 1–28. CDP confirmed the ledger was absent from the accessibility tree as well as the DOM.

After the proof, remove the disposable specs and stop only this isolated project:

rm -r repro-specs playwright.repro.config.ts
cd ../../..
sg docker -c 'docker compose -p cw-1009-fixed --env-file tools/simulation/.env.sim -f tools/simulation/docker-compose.1009-fixed.yml down -v'
rm tools/simulation/reset-1009-fixed.sh tools/simulation/docker-compose.1009-fixed.yml

@mforce

mforce commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Review record — Claude Opus 5.5 (local agent, read-only), 2 rounds

Reviewed at the owner's request. CodeRabbit was not triggered.

Round 1 — ab284d7f: 0 must-fix, 1 should-fix, all nine checks sound

Sound:

  • Phone path. On a phone the helper opens the page's Filters dialog through the category chip, fills fromLabel/toLabel, and closes it. A missing control times out; no path skips setting the range.
  • Readiness. The seq ticket discards responses for intermediate ranges.
  • Farm-clock date arithmetic. Date.UTC runs on calendar parts in the farm zone, so January 1, March 1 and both DST transitions are handled.
  • The specs still test what they claim, with no layout change from the filtered state.
  • The range covers the seeded rows for a fresh seed.
  • Suite rules. tEn for every label, no credentials, the phone fixture on phones.
  • Completeness. No other spec needs expense rows.
  • Body claims, and scope: four files under tools/simulation/ui, no seeder, product or CI change.

Raised:

  • Should-fix: the PR body still listed two results as pending.
  • Observation: the comment and body said the range "matches the preflight", but preflight.ts only checks eggs, not expense rows.
  • Observation: the helper relied on the caller's next locator rather than waiting for the reload itself. That can't produce a false green, but the reviewer inferred a flake on ordinary mid-month days: the old month's list is still mounted, and phone.spec clones rows into it.

Fixes — 83b153ac

  • The helper waits for the response carrying the exact new from/to, then for the rendered ledger. It subscribes before navigating, so the 31st of a 31-day month, where the default already equals the requested range and no new request fires, still resolves.
  • The egg-preflight claim is removed.
  • The body reports every result.

Round 2 — delta ab284d7f..83b153ac: 0 must-fix, 1 should-fix

  • Confirmed sound:
    • The wait can't hang silently. It matches GET /api/v1/expenses on decoded from/to, so parameter order and limit/offset don't matter.
    • The initial load satisfies it only on the 31st, when that is genuinely the right range.
    • It handles the last day of 30-day months and February.
    • The listener detaches when it resolves.
  • Should-fix: the response promise was never marked handled. If goto, the chip click or a fill threw first, it later rejected as an unhandled rejection. specs/i18n.spec.ts:84-91 records that such a rejection fails every test in the file, which would hide the real cause.

Fix — 1880d761

  • loaded.catch(() => undefined) immediately after creating the promise; await loaded is unchanged.
  • Probe with goto forced to reject. Before: the navigation error plus page.waitForResponse: Test ended. After: only the navigation error.
  • The wider cascade across the file was not reproduced, and the PR does not claim it.

Evidence the fix works

  • CI e2e-smoke: all 3 shards passed at ab284d7f, on October 1 in Chicago. That run is the live month boundary this issue is about.
  • Local, documented in the body:
    • The advancing-clock February 1 recipe, run as written: 4 failures before, 4 passes after.
    • A fresh full quick suite at 83b153ac: 134 passed, 1 intentional skip, 0 failures, no reruns.

Known red check, unrelated

Image build + Trivy scan fails on the base image's OpenSSL CVE-2026-84782. That is #1006: it also fails on main, and this PR touches no Dockerfile. main has no required status checks.

The review loop was stopped deliberately after round 2. The issue's goal is met and proven in CI, and further rounds would have polished test plumbing rather than found defects.

@mforce
mforce merged commit 944acb0 into main Oct 1, 2026
17 of 19 checks passed
@mforce
mforce deleted the fix/1009-expense-specs-month-boundary branch October 1, 2026 07:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

E2E expense specs fail at farm month boundaries

1 participant