Skip to content

node-test-commit-linux on 'rhel8-x64' is failing since 2026-08-20 #4433

Description

@trivikr

https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/

First failure on Aug 20, 2026, 7:41:05 PM

+ git status
HEAD detached at 91860b26014
nothing to commit, working tree clean
+ git rev-parse HEAD
91860b26014268cc5f8afb4ad0093c10d155b322
+ git rev-parse b5d37cd41e5a2aac38e8d95e65d228fbcb60aa2d
b5d37cd41e5a2aac38e8d95e65d228fbcb60aa2d
+ '[' -z 91860b26014268cc5f8afb4ad0093c10d155b322 ']'
+ echo 91860b26014268cc5f8afb4ad0093c10d155b322
+ grep -qE '^[0-9a-fA-F]+$'
++ git rev-parse HEAD
++ git rev-parse 91860b26014268cc5f8afb4ad0093c10d155b322
+ '[' 91860b26014268cc5f8afb4ad0093c10d155b322 '!=' 91860b26014268cc5f8afb4ad0093c10d155b322 ']'
+ '[' -n b5d37cd41e5a2aac38e8d95e65d228fbcb60aa2d ']'
+ git rebase --committer-date-is-author-date b5d37cd41e5a2aac38e8d95e65d228fbcb60aa2d
Rebasing (1/46907)
Rebasing (2/46907)
Auto-merging Makefile
CONFLICT (add/add): Merge conflict in Makefile
error: could not apply 61890720c8a... add readme and initial code
hint: Resolve all conflicts manually, mark them as resolved with
hint: "git add/rm <conflicted_files>", then run "git rebase --continue".
hint: You can instead skip this commit: run "git rebase --skip".
hint: To abort and get back to the state before "git rebase", run "git rebase --abort".
Could not apply 61890720c8a... add readme and initial code
Build step 'Execute shell' marked build as failure

Logs: https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/72113/

Activity

  1. richardlau commented on Aug 21, 2026

    @richardlau
    Member

    Looking at https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/buildTimeTrend all the failures were on test-ibm-rhel8-x64-3. I've wiped out the workspace on that machine.

  2. codebytere commented on Aug 26, 2026

    @codebytere
    Member

    this is a different failure from the workspace one above, but same hosts: since 2026-08-21 every sea/* test plus node-api/test_sea_addon dies with SIGSEGV at startup (empty stderr) on all four rhel8-x64 machines (test-ibm-rhel8-x64-1/2/3, test-digitalocean-rhel8-x64-1), e.g. https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/72402/console, and it's most of today's reliability report (https://github.com/nodejs/reliability/blob/main/reports/2026-08-26.md). it flipped between node-test-pull-request 76181 (started 20:10 UTC, fine) and 76182 (20:52); nothing landed on main in that window, a v26.x-staging PR (nodejs/node#65483) fails the same way, and the un-injected out/Release/node passes the rest of the suite on those hosts, so i don't think it's anything in node. #4430 merged that afternoon and touches package-upgrade, so my guess is a playbook run that evening changed something on these four.

    i tried to reproduce it off-host and couldn't. main built with what these hosts use now (clang 21.1.8 from the llvm-toolset module, which links through gcc-toolset-15's binutils 2.44 and produces a non-PIE node) inside ubi8 (glibc-2.28-251.el8_10.40, selinux-policy-3.14.3-139) passes all 36 test/sea tests as user iojs with NODE_TEST_DIR=/home/iojs/node-tmp, and the resulting --build-sea and postject binaries also run on AlmaLinux 8.10 with kernels 4.18.0-553.150.1 and 553.157.1, SELinux enforcing. Same for the 08-26 nightly (clang 20.1.8) injected both ways. so kernel, glibc, toolchain and LIEF's handling of ET_EXEC all look fine in isolation, and whatever it is lives on the machines.

    could someone with access to one of them please grab the following? that should be enough to pin it down:

    • the kernel line for one of the crashes (journalctl -k | grep -i segfault | tail), which says where it faulted and in which mapping
    • dnf history list and dnf history info <id> for 2026-08-21
    • rpm -q kernel glibc clang llvm-libs gcc-toolset-15-binutils selinux-policy and uname -r
    • limits the agent runs under: systemctl show jenkins -p LimitAS -p LimitDATA -p LimitSTACK (or cat /proc/$(pgrep -f agent.jar)/limits) and sysctl vm.overcommit_memory
    • if it's quick, one failing binary run by hand, e.g. LD_SHOW_AUXV=1 /home/iojs/node-tmp/.tmp.0/sea and strace -f -e trace=execve,mmap,mprotect /home/iojs/node-tmp/.tmp.0/sea 2>&1 | tail -30 (the .tmp.N dir is left behind after a failed test/sea/test-single-executable-application.js run)
  3. MikeMcC399 commented on Aug 26, 2026

    @MikeMcC399

    According to https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/ this looks like flake rather than a hard error.

    Having said that, I've been trying to assist a first-time contributor in nodejs/node#65492 and in that PR it appears to be a hard error.

  4. sxa commented on Aug 26, 2026

    @sxa
    Member

    For the record I ran https://ci.nodejs.org/job/node-stress-single-test/843 earlier with 100 iterations of a couple of the SEA tests and it didn't seem to fail.

    Here is the stuff from the machine that you asked for. For the record I picked test-ibm-rhel8-x64-1:

    the kernel line for one of the crashes (journalctl -k | grep -i segfault | tail), which says where it faulted and in which mapping

    It doesn't explicitly list any segfaults but we have quite a few of these:

    Aug 26 02:47:52 test-ibm-rhel8-x64-1 kernel: traps: node-MainThread[2658746] trap int3 ip:28d6bbf sp:7ffca49b3730 error:0 in node[81e000+2c3a000]
    Aug 26 02:47:52 test-ibm-rhel8-x64-1 kernel: traps: node-MainThread[2658739] trap int3 ip:28d6bbf sp:7ffef61ad9e0 error:0 in node[81e000+2c3a000]
    Aug 26 02:47:56 test-ibm-rhel8-x64-1 kernel: traps: node-MainThread[2658942] trap int3 ip:28d6bbf sp:7ffd1e1dedc0 error:0 in node[81e000+2c3a000]
    Aug 26 03:38:23 test-ibm-rhel8-x64-1 kernel: 2778582 (sea): Uhuuh, elf segment at 00000000071ef000 requested but the memory is mapped already
    Aug 26 03:38:42 test-ibm-rhel8-x64-1 kernel: 2778647 (sea): Uhuuh, elf segment at 00000000071ef000 requested but the memory is mapped already
    Aug 26 03:39:01 test-ibm-rhel8-x64-1 kernel: 2778881 (sea): Uhuuh, elf segment at 00000000071ef000 requested but the memory is mapped already
    Aug 26 03:39:07 test-ibm-rhel8-x64-1 kernel: 2778897 (sea-prep.blob): Uhuuh, elf segment at 00000000071ef000 requested but the memory is mapped already
    Aug 26 03:39:12 test-ibm-rhel8-x64-1 kernel: 2778911 (sea-prep.blob): Uhuuh, elf segment at 00000000071ef000 requested but the memory is mapped already
    Aug 26 03:39:18 test-ibm-rhel8-x64-1 kernel: 2778925 (sea-prep.blob): Uhuuh, elf segment at 00000000071ef000 requested but the memory is mapped already
    

    dnf history list and dnf history info for 2026-08-21

    Nothing in the history in the last week and the machine has not been restarted

    rpm -q kernel glibc clang llvm-libs gcc-toolset-15-binutils selinux-policy and uname -r

    # rpm -q kernel glibc clang llvm-libs gcc-toolset-15-binutils selinux-policy; uname -r
    kernel-4.18.0-553.120.1.el8_10.x86_64
    kernel-4.18.0-553.124.1.el8_10.x86_64
    kernel-4.18.0-553.125.1.el8_10.x86_64
    glibc-2.28-251.el8_10.34.x86_64
    clang-20.1.8-2.module+el8.10.0+23372+3f2ea6fa.x86_64
    llvm-libs-20.1.8-2.module+el8.10.0+23372+3f2ea6fa.x86_64
    package gcc-toolset-15-binutils is not installed
    selinux-policy-3.14.3-139.el8_10.2.noarch
    4.18.0-553.125.1.el8_10.x86_64
    

    limits the agent runs under: systemctl show jenkins -p LimitAS -p LimitDATA -p LimitSTACK

    All infinity

    (or cat /proc/$(pgrep -f agent.jar)/limits)

    You can have that too :-)

    # cat /proc/$(pgrep -f agent.jar)/limits
    Limit                     Soft Limit           Hard Limit           Units     
    Max cpu time              unlimited            unlimited            seconds   
    Max file size             unlimited            unlimited            bytes     
    Max data size             unlimited            unlimited            bytes     
    Max stack size            8388608              unlimited            bytes     
    Max core file size        0                    unlimited            bytes     
    Max resident set          unlimited            unlimited            bytes     
    Max processes             15459                15459                processes 
    Max open files            262144               262144               files     
    Max locked memory         65536                65536                bytes     
    Max address space         unlimited            unlimited            bytes     
    Max file locks            unlimited            unlimited            locks     
    Max pending signals       15459                15459                signals   
    Max msgqueue size         819200               819200               bytes     
    Max nice priority         0                    0                    
    Max realtime priority     0                    0                    
    Max realtime timeout      unlimited            unlimited            us        
    

    and sysctl vm.overcommit_memory

    0 (zero)

    if it's quick, one failing binary run by hand, e.g. LD_SHOW_AUXV=1 /home/iojs/node-tmp/.tmp.0/sea and strace -f -e trace=execve,mmap,mprotect /`0/sea 2>&1 | tail -30 (the .tmp.N dir is left behind after a failed test/sea/test-single-executable-application.js run)

    Noting for reference that the last build on the machine dir NOT fail the test. There are no files matching home/iojs/node-tmp/.tmp.*/sea on the machine at present. We possibly need another failing run.

    I'm slightly tempted to restart one of the machines and see if it still happens...

  5. MikeMcC399 commented on Aug 26, 2026

    @MikeMcC399

    I'm slightly tempted to restart one of the machines and see if it still happens...

    The failure can be seen on each of the machines, for example:
    https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/72442/
    https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/72416/
    https://ci.nodejs.org/job/node-test-commit-linux/nodes=rhel8-x64/72395/

    so I'd be surprised if restarting one of the machines made any difference.

  6. sxa commented on Aug 26, 2026

    @sxa
    Member

    Normally I'd agree and you're probably right but give we've had quite a few jobs aborted in the last few days it's not impossible that something needs a clean out on all of them, particularly with messages like the memory is mapped already. But I'll leave it to @codebytere to comment on that.

  7. codebytere commented on Aug 26, 2026

    @codebytere
    Member

    @sxa figured it out and a restart wouldn't have gotten it - see nodejs/node#65564

  8. sxa commented on Aug 28, 2026

    @sxa
    Member

    Closing this as Richard cleared the workspace to resolve the initial issue and Shelley's PR should have fixed the second.

  9. MikeMcC399 commented on Aug 28, 2026

    @MikeMcC399

    This issue is no longer occurring on nodejs/node#65492, so I can confirm that success. 👍🏻

    There are still multiple general issues that can be seen in this PR, which I've summarized in nodejs/node#65492 (comment)

  10. reopened this on Aug 28, 2026
  11. panva commented on Aug 28, 2026

    @panva
    Member

    Consistently failing on nodejs/node#65553 (at least for now while main is the way it is).

    nodejs/node#65553 (comment)

  12. panva commented on Aug 28, 2026

    @panva
    Member

    nodejs/node#65564 (comment)
    postject makes the same choice in its own copy of LIEF, so test-single-executable-application.js and test_sea_addon, which still inject with it, aren't covered by this.

    trying 47e37e6162c5fb9ce48b99c4c851b281e0268358 in nodejs/node#65553

  13. panva commented on Aug 28, 2026

    @panva
    Member
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions