Repository navigation
ci(judge-e2e): skip the engine on JS-only changes and cut its build from ~37 min - #1359
Conversation
packages/console-ui is a type-only surface and the root pnpm files only reach the judge UIs, yet a PR touching just those ran the full ~37 min real-engine suite (e.g. #1353, one .d.mts file). Keep the UI Vitest and tsc builds on those paths and gate the toolchain, cache, llama.cpp tools, engine install and engine tests on a change the engine suite builds. Claude-Session: https://claude.ai/code/session_01WQhvPNhJJpri8PUFLGuQ5s
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
skill-check — worker0 verified, 83 skipped (no docs/).
Four for four. Nicely done. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 2 included reviews per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe workflow checks changed paths and runs the Rust engine suite only when they match engine-related patterns. Those tests use a shared temporary Cargo target directory and ChangesReal-engine test workflow
Priority: ⬇️ Low Estimated code review effort: 2 (Simple) | ~10 minutes Change: Other Merge Risk: ⚪ Minimal · up to The workflow still selects real-engine tests for this change. No issue identified here prevents merging after normal checks. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. A rabbit checks the paths at dawn, Comment |
The engine step spent ~36 of 37 min compiling: llama.cpp from scratch for each of semif, decider, laya and clef (every CPU variant plus the Vulkan shaders), and eight separate workspaces rebuilding the same deps. The 11 cases themselves take ~27 s. - crates/llama-native: III_LLAMA_CPP_MINIMAL=1 builds one CPU module and no Vulkan (test-only; published builds are unchanged). - judge-e2e: one CARGO_TARGET_DIR for all suites plus the minimal build; drop rust-cache (PR-scoped, never reaches another PR, and it purges path crates) and the Vulkan apt step. Local mirror run: 11/11 pass in 245 s, llama.cpp built twice (judge-laya keeps its own opt-level 3 build). Claude-Session: https://claude.ai/code/session_01WQhvPNhJJpri8PUFLGuQ5s
Why
Judge E2E took ~38 min per run, 37.1 of it in the real-engine step, while its 11 cases take ~27 s. The rest was compilation (run 37926568171):
pull_request, so caches are PR-scoped and no other PR can read them (chore: clean up remaining iii-stream and pubsub references (MOT-3619) #1353 got "No cache found" while refactor(judge-provider,judge-typesafe,judge-clef,judge-decider,judge-semif,judge-laya): MOT-5333 follow-ups (MOT-5342) #1356's 1 GB cache existed). semif/decider were not in its list, and it purges path crates, llama.cpp included.The job also ran in full for PRs that touch only
packages/console-ui(types only) or the root pnpm files, which the engine suite never exercises. #1353 paid 38 min for one.d.mts. Those paths changed 54 times in the last month.What
scopestep diffs the PR merge commit against its first parent. The Rust/engine steps run only when an engine-relevant path changed (on.pull_request.pathsminus the four JS-only entries). The UI Vitest andtscbuilds still run on every trigger.crates/llama-native:III_LLAMA_CPP_MINIMAL=1builds one CPU module and no Vulkan. It is meant for test builds only; published builds are unchanged.CARGO_TARGET_DIRfor all eight suites, so each shared unit builds once. judge-laya'sopt-level = 3still gets its own llama.cpp.copy_intoalready removes before copying, so the shared profile dir is safe.Verification
libggml-cpu.so; decider reuses semif's build (17.9 s).on.pull_request.pathsentry was checked against the scope regex (22/22 as expected).grep -Ehere-string was run under-eo pipefaillike the runner.https://claude.ai/code/session_01WQhvPNhJJpri8PUFLGuQ5s
Summary by CodeRabbit