Repository navigation
ci(ingest): run each source in its own job and lift the timeout - #239
amanthanvi wants to merge 5 commits into
Conversation
The weekly Ingest workflow ran every source in one 30-minute job. Since the polite crawler landed (250 ms per-host floor), the NIST crawl alone needs about 42 minutes, so every run since 2026-09-14 was cancelled. - plan job lists ingestable sources via a new `ingest --list` mode - one regenerate job per source (fail-fast off, 90-minute timeout), each validating with content:check and uploading its own bundle diff - combine job folds the passing diffs, re-checks them together, and hands one patch to the unchanged write-scoped pull request job - concurrency group keeps two runs from crawling the same hosts at once Co-authored-by: Aman Thanvi <amanthanvi@users.noreply.github.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
@greptileai review |
|
@coderabbitai review |
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Reviewer's GuideThe PR replaces the single timed ingest job with a concurrency-protected pipeline that validates and fans out source selection, crawls each source independently for up to 90 minutes, combines only successful validated patches, and opens one PR with source-level outcome visibility. Sequence diagram for source planning and aggregated pull requestsequenceDiagram
participant Workflow
participant Plan
participant SourceJobs
participant Combine
participant PullRequest
Workflow->>Plan: ingest --list [--source slug]
Plan-->>Workflow: JSON source slugs
Workflow->>SourceJobs: Run one job per source
par Each source job
SourceJobs->>SourceJobs: ingest --source SOURCE
SourceJobs->>SourceJobs: content:check
SourceJobs-->>Combine: Upload bundle-patch-SOURCE
end
Combine->>Combine: Download bundle patches
Combine->>Combine: git apply patches
Combine->>Combine: content:check
alt Validated changes exist
Combine-->>PullRequest: Upload content-patch and sources
PullRequest->>PullRequest: Open pull request
else No validated changes
Combine-->>PullRequest: No pull request
end
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
|
`ingest --list --source` with no value fell through to an unfiltered list. Fail with a usage error instead, which also stops `--source --all` from treating the next flag as a slug. Co-authored-by: Aman Thanvi <amanthanvi@users.noreply.github.com>
|
@coderabbitai review |
|
@greptileai review |
There was a problem hiding this comment.
Hey - I've found 1 issue
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path=".github/workflows/ingest.yml" line_range="20-22" />
<code_context>
+# Overlapping runs would crawl the same hosts twice and race on the pull
+# request branch, so a second run waits for the first.
+concurrency:
+ group: ingest
+ cancel-in-progress: false
+
jobs:
</code_context>
<issue_to_address>
**Pending dispatches are dropped**
When one ingest run is active, one is pending, and another scheduled or dispatched run arrives, the `ingest` concurrency group replaces its pending run with the newer run despite `cancel-in-progress: false`; the replaced dispatch never crawls its requested source, and there is no replay mechanism.
Use a queue or replay mechanism that preserves every scheduled and manual dispatch instead of relying on the concurrency group's single pending slot.
</issue_to_address>Sourcery assessment
Needs a human reviewer. 1 finding to address first, and a workflow or matrix/patch-combination defect could produce an incomplete or incorrect generated-content pull request, or leave runs queued behind the new concurrency group; those effects are bounded and can be corrected by closing the PR, rerunning ingestion, and reverting the workflow. Reverting restores the prior behavior, though already-created PRs, artifacts, and notifications require cleanup.
Blocking findings: .github/workflows/ingest.yml:22
| concurrency: | ||
| group: ingest | ||
| cancel-in-progress: false |
There was a problem hiding this comment.
🟡 Medium · Pending dispatches are dropped
When one ingest run is active, one is pending, and another scheduled or dispatched run arrives, the ingest concurrency group replaces its pending run with the newer run despite cancel-in-progress: false; the replaced dispatch never crawls its requested source, and there is no replay mechanism.
Use a queue or replay mechanism that preserves every scheduled and manual dispatch instead of relying on the concurrency group's single pending slot.
Prompt for AI agents
In `.github/workflows/ingest.yml` at lines 20-22:
**Pending dispatches are dropped**
When one ingest run is active, one is pending, and another scheduled or dispatched run arrives, the `ingest` concurrency group replaces its pending run with the newer run despite `cancel-in-progress: false`; the replaced dispatch never crawls its requested source, and there is no replay mechanism.
Use a queue or replay mechanism that preserves every scheduled and manual dispatch instead of relying on the concurrency group's single pending slot.There was a problem hiding this comment.
Partly valid, so I changed the design in 0d1e2f0 instead of adding a queue. GitHub Actions has no native queue, and building a replay mechanism is out of scope for this item. The concurrency is now per job: regenerate queues on ingest-<source>, and pull-request queues on ingest-pull-request. A pending job can then only be replaced by a newer crawl of the same source, which does the same work with fresher data. Dispatches for different sources no longer block or drop each other. A replaced job also still shows as cancelled in the run list, so it isn't silent.
There was a problem hiding this comment.
Update: I reverted the per-job locks in 09432fe. Greptile pointed out that they let an older run publish after a newer one and roll the PR branch back, which is worse than this case. I'm declining a queue here: the workflow-wide group keeps runs ordered, and the only loss is a pending run replaced by a third overlapping one. That run still shows as cancelled and can be dispatched again. The trade-off is documented in the workflow comment, and a real queue or replay mechanism belongs with the ingest follow-ups (M2 or A2).
A workflow-wide concurrency group holds one pending run, so a third overlapping run could silently drop a queued single-source dispatch. Queue the regenerate jobs per source and the pull request job on its own group: a pending job is then only ever replaced by a newer crawl of the same source, and a dispatch for one source no longer waits on the NIST crawl. Co-authored-by: Aman Thanvi <amanthanvi@users.noreply.github.com>
|
@greptileai review |
Per-job locks let an older run publish after a newer one and roll the pull request branch back to stale bundles. Restore the workflow-wide group and document its cost: a third overlapping run replaces the pending one, which shows as cancelled and can be dispatched again. Co-authored-by: Aman Thanvi <amanthanvi@users.noreply.github.com>
|
@greptileai review |
…rce-jobs-956c Co-authored-by: Aman Thanvi <amanthanvi@users.noreply.github.com>
What
Audit item E4. Restructures
.github/workflows/ingest.ymlso each source is ingested in its own job, under a timeout that fits the NIST crawl.planlists the ingestable sources with a newingest --list [--source <slug>]mode. A mistyped or disabledworkflow_dispatchslug now fails here, before anything is crawled.regenerateis a matrix with one job per source (fail-fast: false,timeout-minutes: 90). Each job runs the adapter, thencontent:check, then uploads its ownbundle-patch-<slug>artifact if its bundle changed.combineruns after all source jobs, whether they passed or failed (!cancelled()). It applies the patches that passed, re-runscontent:checkon them together, and uploads the singlecontent-patchartifact.pull-requestis unchanged, except for its gate and two new lines in the PR body: the sources included, and a link to the run. It is still the only job with write permissions.concurrency: ingestgroup serializes whole runs. Overlapping runs would otherwise crawl the same hosts twice, and an older run could publish after a newer one and rollingest/auto-updateback.The crawler is untouched: robots.txt handling and the 250 ms per-host floor stay as they are.
Why
Since 2026-09-14, all four scheduled runs (most recently run 37275732452) were cancelled at the 30-minute timeout during the NIST crawl. The logs show exactly 100 pages every 25 s, which is the 250 ms floor. At that rate the 10,023 term pages need about 42 minutes. The floor arrived with v0.2.0 (#227) on 2026-09-12. Before that, the 2026-09-07 run crawled the same pages in about 12.5 minutes with 8 unpaced workers.
Why 90 minutes: NIST's
maxItemscap is 15,000 pages, and 15,000 × 250 ms ≈ 63 minutes, plus index pages and setup. MITRE, NICCS, OWASP and RFC 4949 each take under a minute.Parallel jobs stay polite because the crawler paces requests per host, within one process. The five sources use five different hosts, and a new test (
gives each ingestable source its own host) fails if two enabled sources ever share one.Heads-up: jobs will now fail later, at
content:check(M1)This PR gets the jobs past the timeout, but any source whose entries changed will still fail at
content:check. The cause is the whole-corpus tag hash (audit M1). Today, in the live run below:Validate contentstep. NIST alone reports 1,374 errors: the corpus-hash mismatch, tags pointing at entries the v2 adapter no longer emits, and stale tag examples.content/generated/mitre-attack-cti.json, 7 lines). Because of M2, CI won't run on that PR, since it is opened withGITHUB_TOKEN.Out of scope, as agreed: making NIST resumable and switching the PR token to a GitHub App (M2), and the tag gate itself (M1).
How to test
pnpm gate: passes locally (Node 24), both before and after mergingmain. The ingest suite has a newsources.test.tscovering:--listrun as a subprocess, to prove stdout is pure JSON and that a bad or missing--sourceexits non-zeroactionlint1.7.12 with ShellCheck 0.11.0: no findings iningest.yml.run:scripts, extracted withyq. It covers a changed bundle, a brand-new untracked bundle, an unchanged source, combine with zero artifacts, combine with two patches, and the pull-request job applying the combined patch.combineapplied MITRE, re-validated it (content ok), and the stubbed PR job reportedWould open a PR for: mitre-attack-cti.cancelled), one failed, and one produced a patch.combineand the PR job still ran with the good patch. This confirms that!cancelled()stays true when a source job times out.Notes for the reviewer
--sourceflag with no slug, and per-job locks letting runs publish out of order.Checklist
pnpm gatepasses locallyfeat:,fix:,docs:,chore:, ...)README.md,docs/**,content/README.md) if behavior changed.env*stays untracked and placeholders onlySummary by Sourcery
Restructure ingestion CI to run validated sources independently, tolerate partial failures, and publish successful changes together.
New Features:
ingest --list, including early rejection of invalid or disabled manual selections.Bug Fixes:
Enhancements:
CI:
Documentation:
Tests:
ingest --listCLI output.