Skip to content

fix(ci): configure apt mirrors before server Chromium install - #191

Merged
lukemaj merged 1 commit into
fork/v2from
fix/server-ci-apt-mirrors
Oct 8, 2026
Merged

lukemaj merged 1 commit into
fork/v2from
fix/server-ci-apt-mirrors

Conversation

@lukemaj

@lukemaj lukemaj commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

What: Configure the existing APT mirrors before server CI installs Chromium dependencies.
Why: Slow Azure Ubuntu package downloads consumed about 390 extra seconds of the server job budget.
So what: Review this separate CI maintenance PR against fork/v2; the maintainer approves merging.

Add only uses: ./.github/actions/setup-apt-mirrors before test_server runs Playwright's install --with-deps --only-shell chromium. The existing action handles Azure archive sources and 15-second APT transport timeout failover. Job timeouts, steps, tests, skips, action code, caches and containers remain unchanged.

The user approved this narrow fix separately from #188. This PR does not close #173. It supports the Objective by keeping the server proof gate usable for fork features.

Evidence

  • Slow server shard 3, feat(wight): continue v2 threads within timer and quota limits #188 head c5013d924c: installer about 407 seconds, 21:04:58 to 21:11:45 UTC, 2026-10-08. Raw log: Fetched 21.5 MB in 6min 32s (54.8 kB/s), with Ubuntu font packages from azure.archive.ubuntu.com.
  • Comparison shard 3, feat(server): restore stream liveness on v2 #184 head 57a59b4fa1d26abe842a12827e776428dc2f2b9f: installer about 17 seconds. Raw log: Fetched 21.5 MB in 2s (10.5 MB/s).
  • Tests were passing at cutoff. Parent evidence reports Wight admission 632 ms for 9 cases; inspected settings logs show 2136 ms for 63 cases versus comparison 3444 ms. Installer delay does not justify changing tests or timeouts.
  • Inspected parent logs /tmp/173-ci-server3-c5013d.log and /tmp/173-ci-server3-base184.log; durable evidence is the linked CI jobs.

Scope and proof

Base: 36d5c0169fc042930aa36e41248c1292cec781d0, refreshed origin/fork/v2.
Head: 7bb2f51e80fd035bf9060717ef7beee70618baf2.
Workspace: dedicated ci-server-apt worktree, branch fix/server-ci-apt-mirrors.

  • git diff --check: passed.
  • actionlint -shellcheck= -pyflakes= .github/workflows/ci.yml: passed.
  • Static assertions passed: removing the two inserted lines restores the base byte for byte; checkout precedes the action; action precedes Chromium. Existing regex matches Azure archive and retains 15-second APT transport configuration.
  • No local builds, test suites, installs, browsers or servers. PR CI provides live runner proof, pending when opened.
  • Exact-head independent review will be recorded in a comment and review/independent status.
  • No version bump under repository policy. Website/SEO impact: none, CI provisioning only.

Modularity and upstream

Existing owner: fork-maintenance. The workflow is already mapped and allowlisted in docs/fork-features.md and scripts/fork-upstream-edits.txt; no ownership documentation change is needed.
Upstream-owned numstat: .github/workflows/ci.yml +2 / -0. One hook into the existing action; no hunk exceeds 15 changed lines.

Inspected upstream main workflow at 9b0df1358f9f9c3277bcb33777c2435b88ee75c4: build/test use this action, server installer still omits it. Open upstream PR and Issue searches for apt mirrors returned no matches. No upstream code was ported.

Elon record

Requirements and who asked: User requested a separate PR adding only the existing mirror action before server Chromium installation.
Deleted: No new mirror code, containers, caches, job timeout changes, skips, test changes, version bump or unrelated docs.
Bottleneck: Azure font package downloads consumed 392 seconds versus 2 seconds in the comparison.
Checked myself: Read logs, action, workflow, map and allowlist; verified the two-line diff, Azure regex and actionlint result.

Model and harness: GPT-6.1-Sol, medium reasoning effort, Codex harness in T3 Code; independent reviewer selected by Prism.

@lukemaj

lukemaj commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Live CI observation on head 7bb2f51e80fd035bf9060717ef7beee70618baf2: run 37847044567, Test Server 3 mirror action succeeded, then Chromium installation succeeded from 21:29:49 to 21:30:05 UTC (16 seconds). This compares with the 407-second installer in the linked slow run. Full CI remains in progress; this is provisioning proof only.

@lukemaj

lukemaj commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Agent work on this PR

Estimated cost unknown · 9 responses · 50 sessions · 7.8 h wall time

Model Responses Tokens Estimated cost
gpt-6.1-sol (codex, medium) 9 533,003 unknown

Flags: 4 human corrections · 89 large tool outputs · 108 repeated commands · 11 repeated failures · 108 repeated reads · 17 repeated skill loads · 1 session without usage records · 9 unpriced responses

Details: snapshot, prices, coverage, counters
  • Task toolboxmd/chromeria#191: outcome unknown (recorded acceptance only; a finished process never implies it).
  • Proof: fix(ci): configure apt mirrors before server Chromium install #191
  • Snapshot 9a102a06ef33e576030fc5fd02d5ac9a0f5cf0f4db6d78f9e8fc43dd87353397, records up to 2026-10-08 21:29 UTC.
  • 50 sessions on claude, codex; AgentsMD 14.6.0.
  • Not counted: 6,330 responses (at least $138.88) in sessions shared with other PRs that worked in no single PR's checkout.
  • 5,007 responses in these sessions worked on other PRs and are counted there.
  • Totals reconcile with the measured sessions: yes. Evidence complete: no.
  • Prices: list-price estimate from T3 local rate table (path withheld) as of 2026-09-28, schedule 696aae45933d0a684a015d3f08cfb6f1aa2fef02e7d272b4689e3a692ea04fff. Unknown prices stay unknown, never zero.
    • T3 LiteLLM rate table when present; bundled schedule covers the rest.
    • Model price sources: T3 local rate table (path withheld).
    • Unpriced: unknown model (9).
  • Native session usage or worker ownership is unavailable.
  • Harness-reported cost: none reported.
  • Usage totals are not billing. Subscription spending is separate and is never posted as spend.
  • Crashed runs are counted separately: 0.

Token counters by model (native counter semantics; never added across semantics):

  • gpt-6.1-sol (codex, medium), codex counters: input includes cached, output includes reasoning: input 528,223; cache-read 495,744; cache-write 0; output 4,780; reasoning 120; harness total 533,003.

Generated tokens (reasoning and other output are disjoint; other output includes code, tool calls and replies):

Model Reasoning Other output Reasoning share Split known for
gpt-6.1-sol (codex, medium) 120 4,660 2.5% 9 of 9

Usage by work phase (this task's own responses):

Phase Model Responses Input Cache read Cache write Output Reasoning Total Priced part
unclassified gpt-6.1-sol (codex, medium) 9 528,223 495,744 0 4,780 120 533,003 unknown

Recorded activity by phase (observations, not separate token bills):

Phase Tool calls MCP results Reads File changes Failed tool results
unclassified 34 18 0 0 1

Selected rates (USD per million tokens). These rates value the report at the selected schedule date; they do not establish historical prices or subscription spending.

Model Input tier Input Cache read Cache write Other output Reasoning
gpt-6.1-sol all unknown unknown unknown unknown unknown

Other output and reasoning are priced without double counting inclusive native output. Missing rates remain unknown.

Local measurement from native records; usage totals are not billing. Updated in place by agent-observer publish.

@lukemaj

lukemaj commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Independent exact-head review: no blocking findings for 7bb2f51 against 36d5c01. The complete diff adds only uses: ./.github/actions/setup-apt-mirrors after checkout and dependency setup, immediately before Playwright Chromium install --with-deps; the existing action handles Ubuntu archive mirror failover and 15-second APT timeouts. Job timeout, tests, skips, and later step logic are unchanged. .github/workflows/ci.yml already has fork-maintenance ownership and allowlist coverage. The PR reports actionlint passed; exact-head CI was still running when checked. The comparison run fetched 21.5 MB in 2 seconds; the older slow run remains in progress, so GitHub did not expose its raw log.

@github-actions github-actions Bot added the vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. label Oct 8, 2026
@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 5.0 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.8 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.9 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 5.0 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.8 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 21.2 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: 029c4e6 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 108.5 KiB
  • Claude decoded thread snapshot: 108.8 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@lukemaj lukemaj mentioned this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XS vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant