Skip to content

Built-In Node has low disk space #4419

Description

@richardlau

Cloned from nodejs/jenkins-alerts#6348 as that will probably auto close as I free up some space.

⚠️ The machine Built-In Node has low space in Disk (used 98%).

Please refer to the Jenkins Dashboard to check its status.

This issue has been auto-generated by UlisesGascon/jenkins-status-alerts-and-reporting.

Activity

  1. richardlau commented on Aug 7, 2026

    @richardlau
    MemberAuthor

    Confirmed the machine is nearly out of space:

    # df -h
    Filesystem      Size  Used Avail Use% Mounted on
    tmpfs           3.2G  744K  3.2G   1% /run
    /dev/vda1       315G  295G  4.5G  99% /
    tmpfs            16G     0   16G   0% /dev/shm
    tmpfs           5.0M     0  5.0M   0% /run/lock
    tmpfs           3.2G     0  3.2G   0% /run/user/0
    #
  2. richardlau commented on Aug 7, 2026

    @richardlau
    MemberAuthor

    Notable large disk usage (that I have found):

    /var/lib/jenkins/jobs

    262G    jobs
    

    of that these are the biggest:

    root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs# du -hs * | sort -hr
    67G     node-test-riscv64-experimental
    45G     node-test-commit-linux-containered
    45G     node-test-commit-linux
    15G     node-test-binary-windows-js-suites
    13G     node-test-commit-linuxone
    12G     node-test-commit-aix
    12G     node-compile-windows
    11G     node-test-commit-plinux
    11G     node-test-commit-osx
    9.0G    node-test-commit-arm
    5.9G    node-test-commit-arm-debug
    3.8G    node-test-commit-custom-suites-freestyle
    3.6G    git-clean-windows
    2.4G    node-test-linter
    2.3G    node-compile-windows-debug
    1.8G    node-test-commit-smartos
    1.5G    node-test-commit-v8-linux
    1.3G    node-test-binary-windows-native-suites
    1.2G    node-cross-compile
    1012M   citgm-smoker
    416M    node-stress-single-test
    ...
    root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs# du -hs node-test-riscv64-experimental/workspace*
    3.0G    node-test-riscv64-experimental/workspace
    2.4G    node-test-riscv64-experimental/workspace@10
    2.5G    node-test-riscv64-experimental/workspace@11
    2.5G    node-test-riscv64-experimental/workspace@12
    2.4G    node-test-riscv64-experimental/workspace@13
    2.1G    node-test-riscv64-experimental/workspace@14
    2.1G    node-test-riscv64-experimental/workspace@15
    2.1G    node-test-riscv64-experimental/workspace@16
    2.1G    node-test-riscv64-experimental/workspace@17
    2.5G    node-test-riscv64-experimental/workspace@18
    2.3G    node-test-riscv64-experimental/workspace@19
    2.9G    node-test-riscv64-experimental/workspace@2
    2.5G    node-test-riscv64-experimental/workspace@20
    2.1G    node-test-riscv64-experimental/workspace@21
    2.1G    node-test-riscv64-experimental/workspace@22
    2.3G    node-test-riscv64-experimental/workspace@23
    2.4G    node-test-riscv64-experimental/workspace@24
    2.3G    node-test-riscv64-experimental/workspace@25
    2.3G    node-test-riscv64-experimental/workspace@26
    2.5G    node-test-riscv64-experimental/workspace@3
    2.7G    node-test-riscv64-experimental/workspace@4
    2.6G    node-test-riscv64-experimental/workspace@5
    2.7G    node-test-riscv64-experimental/workspace@6
    2.5G    node-test-riscv64-experimental/workspace@7
    2.4G    node-test-riscv64-experimental/workspace@8
    2.4G    node-test-riscv64-experimental/workspace@9

    I've removed all of the numbered workspace@* directories (I have left the unsuffixed workspace directory). That has immediately freed up over 50G of space.

    /var/log/

    syslog and kern.log are quite large (even after rotation) and should probably be looked to check if we have a high number of warnings/errors being logged:

    root@infra-digitalocean-ubuntu14-x64-1:~# du -hs /var/log/* | sort -hr
    4.6G    /var/log/syslog
    4.3G    /var/log/kern.log
    4.1G    /var/log/journal
    3.9G    /var/log/syslog.1
    3.5G    /var/log/kern.log.1
    971M    /var/log/dmesg.0
    192M    /var/log/jenkins
    181M    /var/log/syslog.2.gz
    158M    /var/log/syslog.4.gz
    136M    /var/log/syslog.3.gz
    122M    /var/log/kern.log.2.gz
    102M    /var/log/nginx
    100M    /var/log/kern.log.4.gz
    83M     /var/log/kern.log.3.gz
    29M     /var/log/auth.log.1
    ...
  3. richardlau commented on Aug 7, 2026

    @richardlau
    MemberAuthor

    /var/log/kern.log is full of VPN DROP messages, which might be related to Orka? (cc @ryanaslett)

    /var/log/syslog appears to be more varied and might need more analysis on whether there are common warnings/errors in there that could be looked at/eliminated among the normal looking JNLP connection messages.

  4. sxa commented on Aug 11, 2026

    @sxa
    Member

    Thanks for clearing those workspaces - I hadn't appreciated they were using that much space - the number of them is due to using different machines in the axes of the job which are significantly different in performance so one was often playing catch up. It won't get that bad going forward. Apologies for not spotting that sooner. Unfortunate that it happened while I was on vacation.

  5. richardlau commented on Aug 21, 2026

    @richardlau
    MemberAuthor

    #4434 should reduce the amount of messages in /var/log/syslog.

  6. sxa commented on Aug 24, 2026

    @sxa
    Member

    Hit zero again today. I've cleared up 12G but I think with general increases to the size of node we're going to need a differnet strategy other than just cleaning up what we have as a lot of these job directories are now quite large...

  7. sxa commented on Aug 25, 2026

    @sxa
    Member

    Two things of note:

    • We have had a bit of a surge of node-test-commit jobs recently which has resulted in some spikes in usage, hence the issues we saw with smartos recently. This shows the current state of node-test-commit-linux on debian13-x64 for example:
    Image
    • Bearing the above in mind, even though we're only keeping job logs for about 8 days, each axis of the node-test-commit jobs (which is a "distribution" in the case of the linux one) takes up around 17MB for each run, or just under 3.5MB when compressed (only the console log is compressed which typically drops by ~95% to under 400k, not the junitResult.xml files which are increasing in size as the number of tests increase). At the time of writing this is causing close to 1000 builds to be retained for a total of about 7GB for each axis/platform. We also added the two rhel10 architectures recently which added another 2 and that will be adding to the space pressure. Normally a crontab at 0500 each day will compress logs over 7 days old and it's possible that we've just got an unusually high amount of uncompressed logs at the moment while a bit of a backlog of jobs from Friday is cleared.

    Here is a current breakdown of space by axis:

    jenkins@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs$ du -sh node-test-commit-*linux/configurations/axis-nodes/* | grep -v 0K
    8.0G	node-test-commit-linux/configurations/axis-nodes/alpine-last-latest-x64
    8.0G	node-test-commit-linux/configurations/axis-nodes/alpine-latest-x64
    6.3G	node-test-commit-linux/configurations/axis-nodes/debian12-x64
    7.4G	node-test-commit-linux/configurations/axis-nodes/debian13-x64
    7.7G	node-test-commit-linux/configurations/axis-nodes/fedora-last-latest-x64
    7.6G	node-test-commit-linux/configurations/axis-nodes/fedora-latest-x64
    3.0G	node-test-commit-linux/configurations/axis-nodes/rhel10-x64
    7.5G	node-test-commit-linux/configurations/axis-nodes/rhel8-x64
    7.6G	node-test-commit-linux/configurations/axis-nodes/rhel9-x64
    6.2G	node-test-commit-linux/configurations/axis-nodes/ubuntu2404-x64
    1.9G	node-test-commit-plinux/configurations/axis-nodes/rhel10-ppc64le
    6.2G	node-test-commit-plinux/configurations/axis-nodes/rhel8-power9le
    549M	node-test-commit-plinux/configurations/axis-nodes/rhel8-ppc64le
    6.2G	node-test-commit-plinux/configurations/axis-nodes/rhel9-ppc64le
    435M	node-test-commit-v8-linux/configurations/axis-nodes/benchmark-ubuntu2404-intel-64
    510M	node-test-commit-v8-linux/configurations/axis-nodes/rhel8-power9le
    26M	node-test-commit-v8-linux/configurations/axis-nodes/rhel8-ppc64le
    480M	node-test-commit-v8-linux/configurations/axis-nodes/rhel8-s390x
    jenkins@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs$ 
    

    To alleviate this I have re-run the compression operation to compress logs over 5 days old instead of 7 which has reduced many of the above figures by about 20% which should give us some additional headroom.

    There is now 48GB free.

  8. sxa commented on Aug 26, 2026

    @sxa
    Member

    I've modified logcompressor.sh to use -mtime +5 instead of -mtime +7 for now as we have a significant increase in the number of PR test jobs running at the moment.

    number of node-test-pull-request jobs per day in the last month

    From ls -lart | cut -c39-44 | uniq -c in the NTPR builds directory:

         40 Aug  1
         42 Aug  2
         41 Aug  3
         32 Aug  4
         44 Aug  5
         48 Aug  6
         19 Aug  7
         21 Aug  8
         33 Aug  9
         21 Aug 10
         36 Aug 11
         38 Aug 12
         50 Aug 13
         52 Aug 14
         43 Aug 15
         32 Aug 16
         39 Aug 17
         51 Aug 18
         39 Aug 19
         63 Aug 20
        122 Aug 21
         95 Aug 22
        101 Aug 23
        100 Aug 24
        222 Aug 25
    

    As you can see from the above output we've been hitting 100 for most of the last give whereas previously we'd typically have less than 50 (2/hour) each day.

    The high numbers of PR test jobs seen recently causes increased pressure on the CI. At the time of writing (0900 UTC) the jenkins queue has about 50 jobs in it (all jobs, including subjobs of node-test-pull-request) but I'll note that node-test-pull-request currently has 14 actively running builds which is part of why I have proposed throttling node-test-commit.

    Some axes of note from looking around some of the longer running node-test-commit jobs yesterday and discussed in a slack thread. It's not always consistent which one is most problematic but smartos23-x64 and fedora-latest-x64 seem to have the longest queues this morning:

    Label/trend link Approx time/build queue depth now Live load chart
    smartos23-x64 40 minutes 9 chart
    debian13-x64 45 minutes 5 chart
    rhel10-x64 45 minutes 1 chart
    macos15-x64 95 minutes 1 chart
    fedora-latest-x64 52 minutes 14 chart

    Other than smartos which only has one executor available there is no immediately obvious reason that I've seen why some of these are backed up more than others. at 1500UTC yesterday the top four in the table had between 5 and 10 waiting to run. There is no clear indication that one job caused things to back up.

    List of top 20 initiators of node-test-pull-request jobs

    Based on what is currently retained on jenkins

    $ grep userId */build.xml | cut -d: -f2 | tr -d ' ' | sed 's,</*userId>,,g'| sort | uniq -c | sort -n | tail -20
         12 geeksilva97
         13 stefanstojanovic
         14 atlowchemi
         17 legendecas
         22 juanarbol
         24 richardlau
         25 avivkeller
         28 anonrig
         29 renegade334
         40 rafaelgss
         41 mikemcc399
         47 pimterry
         54 mcollina
         67 codebytere
         72 daeyeon
        126 aduh95
        146 jasnell
        221 panva
        406 trivikr
        762 nodejs-github-bot
    
  9. MikeMcC399 commented on Aug 26, 2026

    @MikeMcC399

    https://openjs-foundation.slack.com/archives/C019Y2T6STH/p1787731505704119 has raised the issue of flake and started some actions.

    Flake is going to cause extra load when collaborators use Jenkins "Resume build" to re-run tests.

  10. sxa commented on Aug 26, 2026

    @sxa
    Member

    Noting also that there have been citgm-smoker runs for release builds in the last day which can use up some of the executors for a few hours increasing the pressure and therefore backlog on some of the machines. I'll also point out that smartos is not included in those citgm runs.

  11. sxa commented on Aug 27, 2026

    @sxa
    Member

    I've modified logcompressor.sh to use -mtime +5 instead of -mtime +7 for now as we have a significant increase in the number of PR test jobs running at the moment.

    Noting that the disk space on the jenkins controller is still quite low at 12GB having been up to about 48GB a couple of days ago. Again this is likely due to the load from the last few days. Once those logs become eligible for compression it should improve the situation. The first day of "high usage" (>100 node-test-pull-request jobs) have already been compressed so I think we'll be safe ...

  12. sxa commented on Aug 27, 2026

    @sxa
    Member

    Also noting that this morning there has been an issue with significant delays with github initiating jobs from PRs labelled with request-ci so we should anticipate lots of jobs appearing once that is resolved.

  13. sxa commented on Aug 27, 2026

    @sxa
    Member

    Noting also that these two PRs which were causing issues with test case failures (this increasing the number of reruns being performed as Mike said earlier) have been merged:

  14. pinned this issue on Aug 27, 2026
  15. sxa commented on Aug 28, 2026

    @sxa
    Member

    I've had to reduce log compressor to -mtime +3 due to getting below 1GB just now. It had already been reduced to 4 and dropping to 3 has given us about 40G of space. Hopefully that will be enough ... I've got various *du-sk* files showing the space used by different jobs in the jobs directory on the server so we can analyse what's chewing the most but I suspect it's for higher load reasons already mentioned.

  16. sxa commented on Sep 4, 2026

    @sxa
    Member

    Server has 320G of disk space available. I had increased the compressor to -mtime +5 which has left the space available at about 15G but I've droped it to -mtime 4 for safety since 95% utilised is a bit too close to the wire for my liking.

  17. richardlau commented on Sep 5, 2026

    @richardlau
    MemberAuthor

    Re. #4419 (comment)

    I happened to noticed that the built-in node was running a job today, which isn't expected:

    Image

    Usually we restrict the parent matrix jobs to jenkins-workspace so that the workspace directories get created on those machines and not on the built-in node.

    There are three workspace directories on the built-in node:

    root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs# find . -maxdepth 2 -type d -name "workspace*" -exec du -hs {} \;
    19M     ./post-build-status-update/workspace@script
    2.4G    ./node-test-riscv64-experimental/workspace@3
    2.3G    ./node-test-riscv64-experimental/workspace@2
    2.3G    ./node-test-riscv64-experimental/workspace
    root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs#

    I've edited node-test-riscv64-experimental to restrict the parent to run on jenkins-workspace and I'll clean up the three workspace directories above to reclaim the 6G of space.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions