Skip to content

land.yml: the landing driver on Actions, with the token job running no tree code - #226

Merged
thejackshelton merged 9 commits into
masterfrom
land-yml
Oct 8, 2026
Merged

thejackshelton merged 9 commits into
masterfrom
land-yml

Conversation

@thejackshelton

@thejackshelton thejackshelton commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Cloud migration item 9: .github/workflows/land.yml, stacked on #225 (item 8). The design is split so that the job holding the landing token runs no code from any PR or merged tree. This follows the PM's ruling on the Claude review of #225: both High findings and Medium 5. Mac runs behave as before: every new behaviour is behind LAND_TRUSTED=1 or LAND_GLOBAL_LOCK=1, which only land.yml sets.

The job that holds LAND_TOKEN runs no tree code (High 1)

  • Its own checkout is the workflow's commit: ref: ${{ github.sha }}, master at dispatch time, with persist-credentials: false. Only that checkout is installed (pnpm install --frozen-lockfile) and run. The job has no other actions/checkout, no working-directory and no cd.
  • LAND_TRUSTED=1 in scripts/land.ts (it requires LAND_CI=only):
    • Refusal guard: run, and so must and heavy, refuses any command whose cwd isn't the trusted checkout (treeCodeRefusal, a Fatal). The local parity:lanes in judgeDevices is refused the same way.
    • No installs in the driver or position worktrees.
    • Merge drivers: set from trustedGitConfig, so each driver is node '<trusted checkout>/scripts/floor-merge.ts' … (and sorted-merge.ts) by absolute path. Anything else is refused.
    • No pull: the driver never runs git pull --ff-only in its checkout, so code that just merged never runs in this job.
  • The tree's own commands run in .github/workflows/land-checks.yml. It is dispatched on master for the tree's commit, pushed apart to land-checks/c-<sha>, with contents: read, actions: read, no secrets and no dependency cache.
    • It runs the install, device-ci.ts merge of a device run's outcomes, pnpm typecheck, evidence:stamp --compare <prev> and parity:lanes, and records each one's exit code and output tail in result.json. The merged records come back as outputs.patch.
    • The driver checks the results strictly against what it asked for (parseChecksResult) and applies the patch with the same tree check as the regen patch.
    • Exit codes are judged as the local run judges them. A failed install fails the PR at install; a merge exit 3 is a refusal. Results that don't read, or Actions not running the workflow, stop the driver as a CI outage, with no PR blamed.
    • Regen, the full test and the device lanes were already dispatched workflows under LAND_CI=only.
  • The token:
    • LAND_TOKEN appears only in the token-check step (as secrets.LAND_TOKEN != '') and in the Land one batch step's env, never in a job env.
    • The driver reads it once, deletes it from its environment, and passes it per call only to git push (a GIT_CONFIG_* extraheader in that call's env) and to its own gh calls (GH_TOKEN in that call's env). Children such as pr:review and the review lookup read with the job's read-only GITHUB_TOKEN.
    • The job's GITHUB_TOKEN permissions are all read.
  • Fail fast: a non-empty queue fails at once without the token, or if the run wasn't dispatched on master.
  • environment: land on the job (High 2).

Queued runs and the cross-host lock (Medium 5)

  • Concurrency is on the land job, not the run. The redispatch job runs after the land job ends and holds no secret: GITHUB_TOKEN with actions: write may dispatch workflows.
  • dispatchLand (scripts/land-state.ts) serves both the re-dispatch and pnpm land:dispatch. GitHub keeps one pending run per concurrency group and cancels the older, which would silently drop its queue. So dispatchLand:
    • lists the land.yml runs and waits while any land job (other than its own run) hasn't started;
    • then dispatches through REST with return_run_details;
    • then watches its run, and dispatches again if a racing dispatch got it cancelled before it started.
  • How the PM dispatches safely (in land.yml's header): pnpm land:dispatch <branch>:<pr>:<clean-head> .... Never gh workflow run land.yml by hand while a land run is queued. To stop after the current batch, set the land-stop label on issue vars.LAND_STOP_ISSUE.
  • Cross-host lock: LAND_GLOBAL_LOCK=1 takes refs/tags/land-lock. It is an annotated tag whose message names the holder: a land.yml run, or a pid on a host.
    • It is created and deleted with git push --force-with-lease, which is atomic, rather than a label, which has no compare-and-set.
    • A holder that is gone is replaced under the lease: its run completed, or its pid on this same host is dead. A holder on another host, or one that can't be read, is never judged gone; the PM deletes it by hand.
    • It is released only while it is still ours.
    • land.yml sets it. The PM's Mac driver must also run with LAND_GLOBAL_LOCK=1 for the two to exclude each other. I left the Mac default off, per "Mac default unchanged".
    • An empty queue takes no lock.

Second round: the security review of 18ff0b9 (2 High, 3 Medium, and the Low on #225)

  • High 1, symlinks.
    • Records are written into a tree's worktree, read from it and removed from it only through writeInTree, readInTree and removeInTree (land-lib.ts). Every directory on the way must be a real directory, and the file is opened with O_NOFOLLOW. The two record write sites (the carried records, and the merge's records in a prepared tree) first call refuseRecordSymlinks: a mode-120000 entry at or above a record, in the source trees or the worktree's index, fails the PR at records.
    • Under LAND_TRUSTED the repository has core.symlinks=false, so git writes a tree's links as plain files.
    • applyRegenPatch refuses a patch that creates or keeps a symlink. This covers both the regen patch and the land-checks records patch, on any host.
    • Artifact files (result.json, outputs.patch, full-test-results.json) are read without following links.
    • seedRegenCache doesn't run under LAND_TRUSTED.
    • The workflow makes the trusted checkout read-only (chmod -R a-w) before the driver step, and checks that nothing outside .git is still writable. Only .git stays writable: the driver's worktrees, fetches, refs, config and merges live there. Its logs, run directory and locks are in $RUNNER_TEMP and /tmp.
    • Write-path audit of land.ts: every remaining write targets a directory the driver made itself (LAND_LOG_DIR, the run directory, /tmp locks, mkdtemp directories, its own worktrees, .git), or goes through git. Git refuses to apply a patch beyond a symbolic link. The ignored-file cleanup in resetWorktree now uses removeInTree too.
  • High 2, no package manager in the token job.
    • The job has no pnpm/action-setup and no install. The workflow sets the git config itself, with node '<checkout>/scripts/floor-merge.ts' and sorted-merge.ts by absolute path. It runs node scripts/land.ts.
    • Under LAND_TRUSTED, pr:review runs as node <checkout>/scripts/pr-review.ts, and treeCodeRefusal refuses pnpm, npx, npm, yarn and corepack anywhere. On the Mac, pnpm -s pr:review is unchanged.
    • The token job's scripts import only node: built-ins and repository scripts. A test walks the import closure with ts.preProcessFile and fails on anything else.
    • This changes the reviewed scope: **/package.json is removed from .macroscope/ignore.md, so every package.json is now reviewed. pnpm-lock.yaml stays ignored, since the token job installs nothing. macroscope-ignore.test.ts now lists package.json among the reviewed paths.
  • Medium 3, the lock on the Mac. A run without LAND_GLOBAL_LOCK=1 still reads refs/tags/land-lock (checkLockFree). If the tag is held by a land.yml run that hasn't completed, the run refuses to start. Any other holder is logged loudly. With no tag, behaviour is unchanged. Taking the lock on the Mac stays opt-in.
  • Medium 4, the stop channel. The check step fails a non-empty queue unless vars.LAND_STOP_ISSUE is an issue number.
  • Medium 5, land-checks.yml.
    • Each tree command runs under timeout --kill-after=60s. A hang becomes that command's recorded exit 124, and the driver fails the PR at that check. The limits, each several times a usual run:

      Command Limit
      install 30 min
      device merge 15 min
      typecheck 40 min
      evidence stamp 20 min
      parity:lanes 20 min
    • The job limit is now 150 min, and the driver's LAND_CHECKS_WAIT default is 165 min.

    • Results upload with if: always().

    • A failed Land checks step is a verdict (VERDICT_STEP), and so is a completed run with no result.json: the driver blames the tree. Only setup steps and the runner count as an outage.

  • Low (Land: driver state off /tmp (review comment, land-stop label, land/proof status, logs, handoff) #225), land/proof: a status with no creator id is ignored as untrusted instead of throwing.

The new tests in land-trusted.test.ts:

  • Symlinks:
    • a planted symlink at a record, or at a directory above one, is refused for write, read and remove, and the trusted file is unchanged;
    • a symlink patch made on a scratch repository is refused before git apply;
    • symlinkEntries and patchHasSymlink are checked directly.
  • No package manager: no pnpm, npx, npm, yarn or corepack in either job's steps; treeCodeRefusal refuses a package manager; pr:review runs with node when trusted; and the import closure test.
  • Workflow checks: the stop-channel check comes after the empty-queue exit, and the read-only step runs right before the driver.
  • The lock check against a scratch bare repository.
  • land-checks.yml: the timeouts, if: always(), and that the record lines are the only tree commands.

Third round: the re-review of 350af11 (one Medium, two Lows)

  • A tree can't steer its own runs into an outage. Once a dispatched run has completed and passed, every byte of its artifact comes from where the tree's code ran. So these now fail the PR at that step instead of stopping the driver:

    • an artifact that is missing ("no valid artifacts found"), or isn't the workflow's shape;
    • land-checks results that don't parse;
    • a patch the driver refuses (PatchRefused: a symlink, not a regular file, or one git apply can't apply to the tree it was made from).

    This covers the land-checks results and records patch, the regen patch and the device outcomes. Network errors, and the driver's own precondition (its worktree is the tree it sent), are still CI trouble.

  • Per-PR outage count (LAND_OUTAGE_EJECT, default 2).

    • Under LAND_CI=only, each build of a PR's position is recorded as the commit status land/outage on the PR's head: pending while it builds, error when it ended in a CI outage, success when it reached a verdict (built, or failed).
    • A run that is killed mid-build (runner killed, or the job hit its 350-minute limit) leaves its pending status, which counts.
    • At admission, that many non-success marks in a row by the trusted writer eject the PR at ci-outage, with the landing-failed label and a comment naming each run and how it ended.
    • The count lives on the head commit, so a new push starts a new count.
  • Stray processes. land-checks.yml kills every process of the runner user that the tree's commands left behind (all but the step's ancestors and the runner's own) before it writes result.json and makes the records patch.

  • Lows.

    • The symlink check of a patch normalises CRLF first, since git apply accepts a CRLF patch.
    • The Mac's read of the lock tag retries transient git errors.
  • Existing tests changed (in land-devices-ci.test.ts): a passed run with an empty device-outcomes artifact, and a regen artifact that isn't exactly outputs.patch, were CiUnavailable and are now the PR's LandFailure, per this ruling. A failed download that isn't "missing" stays CiUnavailable. That test's title now says "a failed download" instead of "a bad artifact".

Fourth round: the re-review of 1bb56c5 (one Medium, one Low)

  • Outages in the other two places a PR's tree runs now count against PRs.
    • Prepared regen: a prepared position's CI regen is recorded on its PR's head. It is pending while the driver waits for it, and error on a CI outage.
    • Full test: a position's full test is recorded on every PR its tree holds (the chain up to that position, as recorded by verifyChain). It is pending during the proof, success on a pass or a test failure, and error on an outage. A batch can't pin a full-test outage on one PR.
    • Solo rebuilds: a PR with any mark since its last verdict is built and proved alone. The new runBatches solo hook starts a batch of 1 for it, and never puts it into an existing batch. So the next outage is pinned on it, and the streak ejects the culprit.
  • Stray processes, safer form. land-checks.yml runs each tree command in a session of its own (setsid timeout …, with no job control, so the session id is the process id). Once the command ends, it kills that session only (pkill -KILL -s <sid>). There is no name list and no user-wide sweep, so nothing of the runner's is at risk. A process that leaves the session through its own setsid can still write the results. Whatever it writes is the tree's artifact, and a bad one fails the PR, never CI.
  • Clearing a streak (Low). pnpm land:clear-outage <pr> writes a trusted land/outage success at the PR's current head after a real GitHub outage. It counts only from the driver's identity or a LAND_PROOF_WRITERS id, like every status. So the PM's account must be on LAND_PROOF_WRITERS for its clear to count.
  • The first real land-checks run must be checked after this lands. land-checks.yml can only be dispatched from master, so neither the session kill nor the rest of the workflow has run yet. The PM's proof after landing: one real land-checks dispatch whose tree plants a stray background process. Check that the session kill removes it before the results are written, and that the run still uploads its artifact. If it doesn't, the cost is bounded: a run the tree breaks is the PR's failure or an outage at that PR, and the eject rule removes such a PR after LAND_OUTAGE_EJECT outages in a row.

Fifth round: the re-review of 5f86880 (one Medium)

  • land/outage now counts builds, not statuses. Before this, a single outage counted twice: once for the pending status and once for the error after it.

    • Each marking site (a position build, a prepared regen, a full test) takes a build id from newBuild(). On Actions the id is the land.yml run and attempt plus a counter; on a Mac it is the host and pid plus a counter. Every status of that build carries the id as [build <id>] in its description.
    • outageStreak judges each build by its newest status. A pending followed by its outage is one build. A pending with nothing after it (the run was killed) is one. A status with no id counts on its own.
    • A prepared regen that ends without an outage now closes its build with a success. Before, its pending was left open and later counted as a killed build.
  • New tests run the real write sequence through writeOutage into a fake status store:

    • building, then outage: count 1, which is above 0 (the PR is rebuilt alone) and below the eject limit of 2;
    • two real outages in a row: count 2, so the PR is ejected;
    • a killed build: count 1;
    • a full test marked with one id and then passing: count 0.

    They also check that every driver mark passes its build's id.

Sixth round: the re-review of b8fc65a (third streak finding): a design and a check for the whole class

The fourth and fifth rounds' build-level marks are replaced, per AGENTS.md step 6.

  • The model: count outages per driver run, not per phase (outageLedger in land-state.ts). Each run records on a PR's head, as land/outage with [run <id> <kind>] in the description:

    • in: the PR was admitted to the run;
    • outage: a CI outage in a step that ran its tree (its build, with its regen, land-checks and devices; a prepared regen; a full test or bisect prefix of a tree holding it);
    • verdict: a verdict about the PR (it failed, though not an ejection for outages; a full test passed on a tree holding it; it merged);
    • left: the run ended normally with neither, as for a PR requeued behind a culprit.

    How a run counts toward the streak:

    • A run counts 1 when it has an outage, or only in (it was killed, maybe by the tree), and no verdict.
    • A run with a verdict ends the streak.
    • A left run is neutral: it neither counts nor resets.
    • Phases that pass write nothing, so they never reset the streak.
    • The run id is the land.yml run and attempt, or on a Mac the host, the pid and the process start time (the Low).
    • Trust is unchanged: only statuses by the driver's identity or LAND_PROOF_WRITERS count.
  • The check: seeded property tests (mulberry32, failing seed reported) drive random driver runs through the real ledger, writeOutage and outageStreak. Each run has admission with eject and solo, each member's phases (prepared regen, build, land-checks, devices), the full test and bisect prefixes, and publish. Each phase gets a random result: pass, verdict-fail, outage or killed. Outages include batch outages that mark several PRs.

    • 3,000 cases with random results assert that the stored count equals the rule after every run (a no-verdict outage or killed run adds exactly 1, a verdict resets to 0), and that admission ejects exactly at LAND_OUTAGE_EJECT.
    • 2,000 cases with one culprit that stops only one phase, as an outage or by killing the run, assert that the culprit and only the culprit is ejected, and that every other PR gets its verdict and lands. This covers solo rebuilds pinning the culprit, and others' outages never ejecting a PR that gets a verdict.
    • Two mutations of the rule fail both properties: letting a neutral run reset the streak, and not counting a killed run.

Seventh round: the re-review of fb002b1 (the CI waits were not attributed): attribution in the type

  • Attribution is part of the outage. CiOutage(message, prs) now requires the PRs whose code the failing step ran, so TypeScript refuses an unattributed outage. Every site takes them from attribute(stage, …) (land-lib.ts), using a Trees map of the PRs merged into each position this run built. A builder's outage carries its PRs through serializeFatal and parsePrepared. The stages are CI_STAGES:

    Stage What it is Attributed to
    admission-ci the PR head's ci.yml wait that PR
    prepared-regen a prepared position's CI regen the batch's first k PRs
    regen, tree-checks, devices a position build the tree's PRs (the PRs before it, plus it)
    full-test the batch top, a bisect prefix or master the PRs in that tree
    publish-ci the landing commit's ci.yml wait the PRs in that tree
  • Outages are recorded in one place. runBatches returns the outage's PRs, and the driver records them once, at the run's end (recordOutage), before the ledger's end. The per-phase calls are gone.

  • Two consequences of the real wiring:

    • A passed full test is no longer a verdict. The landing commit's CI wait comes after it, and an outage there must count. Otherwise a tree that hangs only in ci.yml at publish was never counted, which was this finding. The verdicts are now only: it failed, or it merged.
    • The asking run is excluded from its own count. The property test found a real bug: a PR requeued in a run and admitted again read its own run's pending in as an outage. Admission and solo now leave the asking run out.
  • The tests use the real wiring:

    • Injection test: an outage is injected at each CI stage, chosen by call order rather than by the PRs passed, through the real runBatches, attribute, Trees and recordOutage. For a batch MIT LICENSE, and ship README and LICENSE in the dragon package #1 Land work through reviewed PRs (Macroscope, PR CI, pr:review) #2 README: Sponsors section #3, it must land on exactly the expected PRs (by hand: admission-ci [2], prepared-regen/regen/tree-checks/devices/publish-ci [1, 2], full-test [1, 2, 3]), minus any PR that merged before the outage.
    • Property tests: the seeded property tests (3,000 random cases and 2,000 culprit cases) run on the same wiring.
    • Static check: an AST check over land.ts derives the CI primitives (every function of land-devices-ci.ts that dispatches or waits on a run, plus checkRuns). It fails unless every call to them, every new CiOutage and every call to a function that takes prs is attributed through attribute(), a forwarded prs, or a builder's round.prs. It also fails unless the stages used are exactly CI_STAGES, which the injection test iterates.
    • Mutations caught: wrongly attributing publish-ci fails the injection and culprit tests, and passing [] at the publish wait fails the AST check.
  • ci.yml's checks job has timeout-minutes: 30. Recent runs take 3 to 5 minutes, so that is six times the slowest, and it is well inside LAND_CI_WAIT (90 min). A hang becomes a failed check, which is a verdict. A test keeps the limit at least 15 minutes and at most half of LAND_CI_WAIT.

Owner-only setup (exactly)

  1. Create the land environment, restricted to master, and put the LAND_TOKEN secret in it, not at repository level. The proof run below auto-created the environment with no rules (deployment_branch_policy: null). Restricting it to master is required: until then, any branch's workflow could reference the environment.
  2. Prefer a GitHub App with contents, pull-requests, actions and statuses write, with its token re-minted per job (App tokens last an hour; the job has timeout-minutes: 350). Otherwise use a fine-grained PAT with the same permissions. For an App token, set the variable LAND_TOKEN_USER_ID to the App bot's user id, because GET user fails for App tokens.
  3. Add a master ruleset requiring pull requests, with no bypass for that App, so a leaked token can't push to master.
  4. Use a reviewer identity separate from the token's. Set the variable LAND_REVIEWERS (ids or logins) to the review agent's own account. The driver refuses to start if the token's identity is on that list.
  5. Set the variable LAND_STOP_ISSUE to the tracking issue for the land-stop label. A queue is refused without it. Optionally set LAND_PROOF_WRITERS, extra trusted ids for land/proof.

Proof without the secret

I dispatched an empty-queue run on this branch at the current head 7e31e21 (gh workflow run land.yml --ref land-yml -f queue=''): https://github.com/compiled-run/dragoncss/actions/runs/37761499936. Earlier rounds' runs: 37759248154 at fb002b1, 37757864100 at b8fc65a, 37756732643 at 5f86880, 37754892663 at 1bb56c5 and 37753380953 at 350af11.

  • The trusted checkout was made read-only, and the step that checks it found nothing writable.
  • The driver logged no "set git config" line, so the merge drivers the workflow set are exactly the ones the driver expects.
  • The rest of start-up ran as in the earlier run: https://github.com/compiled-run/dragoncss/actions/runs/37750849769 at 18ff0b9, which is described below. It exited cleanly with "nothing to land": no LAND_TOKEN is set, so the token check only warns for an empty queue. Under LAND_TRUSTED=1, on ubuntu:
  • the driver set the absolute merge drivers ("set git config merge.dragon-floor.driver", "… merge.dragon-sorted.driver");
  • GET user returned 403 for GITHUB_TOKEN, so it took the github-actions[bot] identity (41898282);
  • it created the driver worktree;
  • it read land/proof: "no record on master or the 20 commits it rests on; under LAND_CI=only that counts as unproved";
  • it printed "nothing to land: the queue is empty; master is unproved …, so the next run with a queue proves it first";
  • the redispatch job was skipped, since the count was 0.

land-checks.yml can only be dispatched from master, so it first runs once this lands.

Tests (no network)

All in the new packages/parity/test/land-trusted.test.ts.

  • The token job runs no tree code:
    • land.yml's token job has environment: land.
    • secrets.LAND_TOKEN appears exactly twice, in the two steps named above, and in no job env.
    • The job's permissions are all read.
    • There is one checkout, at github.sha with persist-credentials: false.
    • Every action and every run: command matches an exact allowlist, and no step changes directory.
    • Concurrency is job-level, and the redispatch job holds no secret.
    • land-checks.yml has no secret, only read permissions and no cache.
  • The driver's ci-only paths:
    • trustedGitConfig rewrites each driver to the trusted checkout and refuses anything else.
    • treeCodeRefusal refuses commands outside the trusted checkout.
    • A source check of land.ts: run refuses first; the lane judgement, the installs and the pull are guarded; the only processes started are git, gh, ps, the trusted scripts and the local-only bash and pnpm (each behind its guard); and the token is read once and deleted.
  • Tree checks: the workflow inputs, the artifact shape, run ids, the scratch branch, and the results validation, including a failed install and a refused merge.
  • Lock: against a real scratch bare repository, covering taking the lock, refusing a live holder, taking over from a gone one, release only while it is still ours, and a land.yml holder. Also the holder parsing and staleness rules.
  • Dispatch: with a fake GitHub: dispatching at once, waiting for a pending run, re-dispatching after a race, giving up past the wait, and parsing runs and jobs.
  • land-actions CLI: queue, redispatch (dispatch and the stop label) and dispatch, against a fake gh on PATH.

Existing tests changed

Each change follows a code change; no check is loosened.

  • land-devices-ci.test.ts: the scratchRef refusal message now lists land-checks/c-<sha>.
  • land.test.ts "never falls back to a local run":
    • It now counts 4 CiUnavailable catches, not 3. The tree-checks catch always stops as a CI outage, which is stricter than the ci-only checks.
    • It expects 3 CI waits passing the queue wait, not 2.
    • Its judgeDevices source strings follow the new lanesRan argument.

What passed

  • pnpm typecheck.
  • pnpm exec vitest run on land-trusted, land, land-devices-ci, land-parallel, land-supervisor, chrome-ports, both registry-claims, macroscope-ignore, cloud-setup, floor-merge and pr-review: 12 files, 338 tests.
  • actionlint 1.7.12 on land.yml, land-checks.yml and ci.yml. ci.yml now lints the two landing workflows.

🤖 Generated with Claude Code

…-checks.yml, the cross-host lock and safe dispatch

Stacked on #225. The token job checks out its own commit without credentials, installs and runs only that checkout, points the
merge drivers at it by absolute path and never pulls; a tree's install, typecheck, evidence stamp, lane judgement and
device-outcome merge run in land-checks.yml (no secrets) and come back as data. LAND_TOKEN is in the land environment and the
driver step only; the driver hands it per call to git push and its own gh calls. refs/tags/land-lock (LAND_GLOBAL_LOCK) keeps a
Mac and an Actions driver apart; pnpm land:dispatch and the re-dispatch never queue a second pending land run.
…er in the token job, the lock read on every run, a required stop channel, land-checks timeouts

- Records are written, read and removed in a tree's worktree only without following a symlink (writeInTree, readInTree,
  removeInTree), after refusing a mode-120000 entry at or above them in the trees they come from; core.symlinks=false under
  LAND_TRUSTED; a regen or checks patch with a symlink is refused; artifacts are read without following links.
- The token job runs no pnpm: its scripts import only node: built-ins (tested by import closure); node scripts/land.ts, the merge
  drivers set by absolute path in the workflow, pr-review.ts run with node; treeCodeRefusal refuses any package manager. The
  trusted checkout is read-only (but .git) before the driver runs. package.json leaves .macroscope/ignore.md, so it is reviewed.
- A run that does not take refs/tags/land-lock still reads it and refuses to start while a running land.yml holds it.
- land.yml refuses a queue without LAND_STOP_ISSUE.
- land-checks.yml: each tree command under timeout; results uploaded always; a failed Land checks step or a completed run with no
  result.json is the tree's failure.
- land/proof: a status with no creator id is ignored as untrusted.
@thejackshelton
thejackshelton changed the base branch from land-on-actions to master October 8, 2026 09:04
- After a dispatched run completed and passed, an artifact that is missing, not the workflow's shape, unreadable, or a patch that
  is refused (PatchRefused: a symlink, not a regular file, git apply failing on the tree it was made from) fails the PR at its
  step; the driver's own precondition and network errors stay CI trouble. Same for the regen patch and the device outcomes.
- Each position build under LAND_CI=only is marked on the PR's head (land/outage: pending, error on a CI outage, success on a
  verdict); admission ejects a PR whose builds ended without a verdict LAND_OUTAGE_EJECT times in a row (default 2), naming the runs.
- land-checks.yml kills every process the tree's commands left before it puts the results and the records patch together.
- The symlink check of a patch normalises CRLF; the Mac's read of the lock tag retries transient errors.
… the PRs whose trees ran (#226 re-review)

- A prepared position's CI regen is marked on its PR's head (land/outage pending, error on an outage).
- A position's full test is marked on every PR its tree holds; a PR with any build or proof since its last verdict ended without
  one is built and proved alone (runBatches solo), so the next outage is pinned on it and the streak ejects the culprit.
- land-checks.yml runs each tree command in a session of its own (setsid timeout) and kills that session (pkill -s) once it ends,
  instead of sweeping the user's processes.
- pnpm land:clear-outage <pr> ends a streak at the current head after a real GitHub outage (trusted like every status).
Each build writes its statuses with its own id ([build <id>]: the land.yml run and attempt, or the host and pid, and a counter);
a build is judged by its newest status, so a pending followed by its outage is one, and a pending with nothing after it (a killed
run) is one. One real outage now leaves a count of 1 (the PR is rebuilt alone), and two in a row eject it.
… property tests (#226 re-review)

The streak counts driver runs, not phases: each run records on a PR's head that it admitted it (in), an outage in a step that ran
its tree (outage), a verdict about it (it failed, a full test passed on a tree holding it, it merged), or, at a normal end with
neither, a neutral left. A run with an outage, or killed with only in, counts 1; a verdict ends the streak; passed phases write
nothing. outageLedger keeps it; seeded property tests drive thousands of random runs of random phases through the real ledger,
writers and streak, and check the count rule, the resets, the eject limit, solo pinning the culprit and every other PR landing.
The Mac's run id holds the process start time.
… the real wiring tested (#226 re-review)

- CiOutage takes the PRs whose code the failing step ran as a required argument; every site takes them from attribute(stage),
  over the PRs merged into each built position (Trees). The CI stages are CI_STAGES: the admission and landing-commit CI waits,
  the prepared regen, the regen, tree checks and devices of a position build, and the full test. A builder's outage carries them.
- runBatches returns the outage's PRs; the driver records them once, at the run's end (recordOutage). The per-phase calls are gone.
- A passed full test is no longer a verdict: the landing commit's CI wait can still hit an outage that must count. The run asking
  is left out of its own count (a PR requeued in a run is admitted again).
- Tests: an outage injected at each CI stage through the real runBatches, attribute and recordOutage lands on exactly the right
  PRs; the seeded property tests run on the same; a TypeScript-AST check of land.ts finds every CI dispatch, wait and CiOutage
  and fails unless each is attributed through attribute() to a stage the injection test covers.
- ci.yml's checks job has timeout-minutes: 30, so a hang is a failed check (a verdict), well inside LAND_CI_WAIT (90 min).
@thejackshelton
thejackshelton merged commit 69b24ec into master Oct 8, 2026
4 checks passed
@thejackshelton
thejackshelton deleted the land-yml branch October 8, 2026 10:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant