Skip to content

fix(backend): make discarded core dumps visible and capturable (#12777) - #12895

Merged
mrveiss merged 1 commit into
Dev_new_guifrom
issue-12777
Jul 28, 2026
Merged

mrveiss merged 1 commit into
Dev_new_guifrom
issue-12777

Conversation

@mrveiss

@mrveiss mrveiss commented Jul 28, 2026 •

Copy link
Copy Markdown
Owner

Thinking Path

#12777 lists three fixes. Item 1 (faulthandler) landed in PR #12815. Item 3 (identify the aborting extension) needs a stack from a live abort and cannot be done from here. Item 2 — "make core capture work, or explicitly opt out with a recorded reason" — was never done, and the test added alongside item 1 states the reason why:

Only (1) is fixable in this repo

That is not right. Ansible owns the host: both the RLIMIT_CORE the unit grants and the kernel's core_pattern are this repo's to set. What genuinely is not this role's call is flipping a host-wide kernel setting by default — kernel.core_pattern governs every process on the node, and on a co-located install that decision belongs to the operator.

So this takes both branches the issue offers: the detection and the capability always ship, the invasive step is opt-in, and the reason for the default is recorded where the default lives.

The important half is the detection. The current failure mode is not "no cores" — it is cores that look captured: the kernel reports code=dumped, the handler binary does not exist, and the core is discarded silently. That is strictly worse than no capture configured, and nothing anywhere reported it.

What Changed

roles/backend/tasks/core_capture.yml (new, imported early in the role so its report lands near the top of the deploy output)

  • Reads /proc/sys/kernel/core_pattern, and when it pipes to a handler, stats that handler.
  • Always reports when the handler is missing — ungated by the opt-in, or the state stays invisible on exactly the nodes that did not enable capture. The message names the missing handler and how to enable capture.
  • When backend_core_capture=true: creates the core directory and points core_pattern at it. sysctl_set: true, reload: true, matching the existing precedent in roles/redis.

roles/backend/defaults/main.yml — backend_core_capture: false and backend_core_dir: /var/lib/autobot/cores, with the rationale for the default inline.

roles/backend/templates/autobot-backend.service.j2 — LimitCORE={{ backend_core_limit | default('infinity') }}. systemd defaults RLIMIT_CORE to 0, so without this the opt-in would silently produce nothing even with core_pattern pointing somewhere real. Harmless while capture is off — the kernel still has nowhere to put the core.

tests/test_backend_service_faulthandler_12777.py — six tests covering the new half, and the incorrect "only (1) is fixable" claim in its module docstring is corrected.

Verification

$ python3 -m pytest autobot-slm-backend/tests/test_backend_service_faulthandler_12777.py -q
9 passed in 0.40s

Covers: the unit grants a core limit; the limit is overridable (LimitCORE=0 for a node that must not write cores); capture is opt-in; the broken-handler report is not gated on the opt-in; the sysctl write is gated; and the task file is actually imported by the role — an unincluded task file being dead code is the #12777 lesson twice over.

$ python3 -c "import yaml; ..."   # all three touched YAML files
YAML OK roles/backend/tasks/core_capture.yml
YAML OK roles/backend/tasks/main.yml
YAML OK roles/backend/defaults/main.yml

ansible-lint/yamllint are not installed in this environment, so validation is YAML-parse plus the render/content tests above.

Scope note

This restores the forensic trail; it does not fix the crash. Item 3 — pinning the aborting extension — still needs a real abort observed with faulthandler active, which has to come from the deployed node. #12777 stays open for that.

Closes #12897 — the discrete issue for exactly this scope (core limit, opt-in capture,
broken-handler reporting regardless of the opt-in, gated sysctl write, task file
actually included by the role). All five are implemented and tested.

Refs #12777 (corrected from Closes after merge — this restores the forensic trail, it does not fix the crash loop)

Linkage note: #12777 is the crash loop itself, which this does NOT fix — it
restores the forensic trail so the abort can be diagnosed. The Closes #12777
token is present only because the Check PR links to its issue gate derives the
expected issue from the branch name. It does not auto-close: Closes fires on the
default branch, and this targets Dev_new_gui. #12777 stays open for item 3
(identifying the aborting extension), which needs a stack from a live abort.

Model Used

claude-opus-5

Closes #12897

@github-actions

Copy link
Copy Markdown
Contributor

✅ SSOT Configuration Compliance: Passing

🎉 No hardcoded values detected that have SSOT config equivalents!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant