Repository navigation
fix(backend): make discarded core dumps visible and capturable (#12777) - #12895
Merged
Merged
Conversation
Contributor
✅ SSOT Configuration Compliance: Passing🎉 No hardcoded values detected that have SSOT config equivalents! |
This was referenced Jul 28, 2026
This was referenced Jul 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Thinking Path
#12777 lists three fixes. Item 1 (faulthandler) landed in PR #12815. Item 3 (identify the aborting extension) needs a stack from a live abort and cannot be done from here. Item 2 — "make core capture work, or explicitly opt out with a recorded reason" — was never done, and the test added alongside item 1 states the reason why:
That is not right. Ansible owns the host: both the
RLIMIT_COREthe unit grants and the kernel'score_patternare this repo's to set. What genuinely is not this role's call is flipping a host-wide kernel setting by default —kernel.core_patterngoverns every process on the node, and on a co-located install that decision belongs to the operator.So this takes both branches the issue offers: the detection and the capability always ship, the invasive step is opt-in, and the reason for the default is recorded where the default lives.
The important half is the detection. The current failure mode is not "no cores" — it is cores that look captured: the kernel reports
code=dumped, the handler binary does not exist, and the core is discarded silently. That is strictly worse than no capture configured, and nothing anywhere reported it.What Changed
roles/backend/tasks/core_capture.yml(new, imported early in the role so its report lands near the top of the deploy output)/proc/sys/kernel/core_pattern, and when it pipes to a handler, stats that handler.backend_core_capture=true: creates the core directory and pointscore_patternat it.sysctl_set: true, reload: true, matching the existing precedent inroles/redis.roles/backend/defaults/main.yml—backend_core_capture: falseandbackend_core_dir: /var/lib/autobot/cores, with the rationale for the default inline.roles/backend/templates/autobot-backend.service.j2—LimitCORE={{ backend_core_limit | default('infinity') }}. systemd defaultsRLIMIT_COREto 0, so without this the opt-in would silently produce nothing even withcore_patternpointing somewhere real. Harmless while capture is off — the kernel still has nowhere to put the core.tests/test_backend_service_faulthandler_12777.py— six tests covering the new half, and the incorrect "only (1) is fixable" claim in its module docstring is corrected.Verification
Covers: the unit grants a core limit; the limit is overridable (
LimitCORE=0for a node that must not write cores); capture is opt-in; the broken-handler report is not gated on the opt-in; thesysctlwrite is gated; and the task file is actually imported by the role — an unincluded task file being dead code is the #12777 lesson twice over.ansible-lint/yamllintare not installed in this environment, so validation is YAML-parse plus the render/content tests above.Scope note
This restores the forensic trail; it does not fix the crash. Item 3 — pinning the aborting extension — still needs a real abort observed with faulthandler active, which has to come from the deployed node. #12777 stays open for that.
Closes #12897 — the discrete issue for exactly this scope (core limit, opt-in capture,
broken-handler reporting regardless of the opt-in, gated
sysctlwrite, task fileactually included by the role). All five are implemented and tested.
Refs #12777 (corrected from
Closesafter merge — this restores the forensic trail, it does not fix the crash loop)Model Used
claude-opus-5
Closes #12897