Repository navigation
Built-In Node has low disk space #4419
Description
Activity
Confirmed the machine is nearly out of space:
# df -h Filesystem Size Used Avail Use% Mounted on tmpfs 3.2G 744K 3.2G 1% /run /dev/vda1 315G 295G 4.5G 99% / tmpfs 16G 0 16G 0% /dev/shm tmpfs 5.0M 0 5.0M 0% /run/lock tmpfs 3.2G 0 3.2G 0% /run/user/0 #
Notable large disk usage (that I have found):
/var/lib/jenkins/jobs262G jobsof that these are the biggest:
root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs# du -hs * | sort -hr 67G node-test-riscv64-experimental 45G node-test-commit-linux-containered 45G node-test-commit-linux 15G node-test-binary-windows-js-suites 13G node-test-commit-linuxone 12G node-test-commit-aix 12G node-compile-windows 11G node-test-commit-plinux 11G node-test-commit-osx 9.0G node-test-commit-arm 5.9G node-test-commit-arm-debug 3.8G node-test-commit-custom-suites-freestyle 3.6G git-clean-windows 2.4G node-test-linter 2.3G node-compile-windows-debug 1.8G node-test-commit-smartos 1.5G node-test-commit-v8-linux 1.3G node-test-binary-windows-native-suites 1.2G node-cross-compile 1012M citgm-smoker 416M node-stress-single-test ...
root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs# du -hs node-test-riscv64-experimental/workspace* 3.0G node-test-riscv64-experimental/workspace 2.4G node-test-riscv64-experimental/workspace@10 2.5G node-test-riscv64-experimental/workspace@11 2.5G node-test-riscv64-experimental/workspace@12 2.4G node-test-riscv64-experimental/workspace@13 2.1G node-test-riscv64-experimental/workspace@14 2.1G node-test-riscv64-experimental/workspace@15 2.1G node-test-riscv64-experimental/workspace@16 2.1G node-test-riscv64-experimental/workspace@17 2.5G node-test-riscv64-experimental/workspace@18 2.3G node-test-riscv64-experimental/workspace@19 2.9G node-test-riscv64-experimental/workspace@2 2.5G node-test-riscv64-experimental/workspace@20 2.1G node-test-riscv64-experimental/workspace@21 2.1G node-test-riscv64-experimental/workspace@22 2.3G node-test-riscv64-experimental/workspace@23 2.4G node-test-riscv64-experimental/workspace@24 2.3G node-test-riscv64-experimental/workspace@25 2.3G node-test-riscv64-experimental/workspace@26 2.5G node-test-riscv64-experimental/workspace@3 2.7G node-test-riscv64-experimental/workspace@4 2.6G node-test-riscv64-experimental/workspace@5 2.7G node-test-riscv64-experimental/workspace@6 2.5G node-test-riscv64-experimental/workspace@7 2.4G node-test-riscv64-experimental/workspace@8 2.4G node-test-riscv64-experimental/workspace@9
I've removed all of the numbered
workspace@*directories (I have left the unsuffixedworkspacedirectory). That has immediately freed up over 50G of space./var/log/syslogandkern.logare quite large (even after rotation) and should probably be looked to check if we have a high number of warnings/errors being logged:root@infra-digitalocean-ubuntu14-x64-1:~# du -hs /var/log/* | sort -hr 4.6G /var/log/syslog 4.3G /var/log/kern.log 4.1G /var/log/journal 3.9G /var/log/syslog.1 3.5G /var/log/kern.log.1 971M /var/log/dmesg.0 192M /var/log/jenkins 181M /var/log/syslog.2.gz 158M /var/log/syslog.4.gz 136M /var/log/syslog.3.gz 122M /var/log/kern.log.2.gz 102M /var/log/nginx 100M /var/log/kern.log.4.gz 83M /var/log/kern.log.3.gz 29M /var/log/auth.log.1 ...
/var/log/kern.logis full ofVPN DROPmessages, which might be related to Orka? (cc @ryanaslett)/var/log/syslogappears to be more varied and might need more analysis on whether there are common warnings/errors in there that could be looked at/eliminated among the normal looking JNLP connection messages.Thanks for clearing those workspaces - I hadn't appreciated they were using that much space - the number of them is due to using different machines in the axes of the job which are significantly different in performance so one was often playing catch up. It won't get that bad going forward. Apologies for not spotting that sooner. Unfortunate that it happened while I was on vacation.
#4434 should reduce the amount of messages in
/var/log/syslog.Reacted by Stewart X AddisonHit zero again today. I've cleared up 12G but I think with general increases to the size of node we're going to need a differnet strategy other than just cleaning up what we have as a lot of these job directories are now quite large...
Two things of note:
- We have had a bit of a surge of node-test-commit jobs recently which has resulted in some spikes in usage, hence the issues we saw with smartos recently. This shows the current state of
node-test-commit-linuxondebian13-x64for example:
- Bearing the above in mind, even though we're only keeping job logs for about 8 days, each axis of the node-test-commit jobs (which is a "distribution" in the case of the linux one) takes up around 17MB for each run, or just under 3.5MB when compressed (only the console log is compressed which typically drops by ~95% to under 400k, not the
junitResult.xmlfiles which are increasing in size as the number of tests increase). At the time of writing this is causing close to 1000 builds to be retained for a total of about 7GB for each axis/platform. We also added the two rhel10 architectures recently which added another 2 and that will be adding to the space pressure. Normally a crontab at 0500 each day will compress logs over 7 days old and it's possible that we've just got an unusually high amount of uncompressed logs at the moment while a bit of a backlog of jobs from Friday is cleared.
Here is a current breakdown of space by axis:
jenkins@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs$ du -sh node-test-commit-*linux/configurations/axis-nodes/* | grep -v 0K 8.0G node-test-commit-linux/configurations/axis-nodes/alpine-last-latest-x64 8.0G node-test-commit-linux/configurations/axis-nodes/alpine-latest-x64 6.3G node-test-commit-linux/configurations/axis-nodes/debian12-x64 7.4G node-test-commit-linux/configurations/axis-nodes/debian13-x64 7.7G node-test-commit-linux/configurations/axis-nodes/fedora-last-latest-x64 7.6G node-test-commit-linux/configurations/axis-nodes/fedora-latest-x64 3.0G node-test-commit-linux/configurations/axis-nodes/rhel10-x64 7.5G node-test-commit-linux/configurations/axis-nodes/rhel8-x64 7.6G node-test-commit-linux/configurations/axis-nodes/rhel9-x64 6.2G node-test-commit-linux/configurations/axis-nodes/ubuntu2404-x64 1.9G node-test-commit-plinux/configurations/axis-nodes/rhel10-ppc64le 6.2G node-test-commit-plinux/configurations/axis-nodes/rhel8-power9le 549M node-test-commit-plinux/configurations/axis-nodes/rhel8-ppc64le 6.2G node-test-commit-plinux/configurations/axis-nodes/rhel9-ppc64le 435M node-test-commit-v8-linux/configurations/axis-nodes/benchmark-ubuntu2404-intel-64 510M node-test-commit-v8-linux/configurations/axis-nodes/rhel8-power9le 26M node-test-commit-v8-linux/configurations/axis-nodes/rhel8-ppc64le 480M node-test-commit-v8-linux/configurations/axis-nodes/rhel8-s390x jenkins@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs$To alleviate this I have re-run the compression operation to compress logs over 5 days old instead of 7 which has reduced many of the above figures by about 20% which should give us some additional headroom.
There is now 48GB free.
- We have had a bit of a surge of node-test-commit jobs recently which has resulted in some spikes in usage, hence the issues we saw with smartos recently. This shows the current state of
I've modified
logcompressor.shto use-mtime +5instead of-mtime +7for now as we have a significant increase in the number of PR test jobs running at the moment.number of node-test-pull-request jobs per day in the last month
From
ls -lart | cut -c39-44 | uniq -cin the NTPRbuildsdirectory:40 Aug 1 42 Aug 2 41 Aug 3 32 Aug 4 44 Aug 5 48 Aug 6 19 Aug 7 21 Aug 8 33 Aug 9 21 Aug 10 36 Aug 11 38 Aug 12 50 Aug 13 52 Aug 14 43 Aug 15 32 Aug 16 39 Aug 17 51 Aug 18 39 Aug 19 63 Aug 20 122 Aug 21 95 Aug 22 101 Aug 23 100 Aug 24 222 Aug 25As you can see from the above output we've been hitting 100 for most of the last give whereas previously we'd typically have less than 50 (2/hour) each day.
The high numbers of PR test jobs seen recently causes increased pressure on the CI. At the time of writing (0900 UTC) the jenkins queue has about 50 jobs in it (all jobs, including subjobs of node-test-pull-request) but I'll note that node-test-pull-request currently has 14 actively running builds which is part of why I have proposed throttling node-test-commit.
Some axes of note from looking around some of the longer running node-test-commit jobs yesterday and discussed in a slack thread. It's not always consistent which one is most problematic but
smartos23-x64andfedora-latest-x64seem to have the longest queues this morning:Label/trend link Approx time/build queue depth now Live load chart smartos23-x64 40 minutes 9 chart debian13-x64 45 minutes 5 chart rhel10-x64 45 minutes 1 chart macos15-x64 95 minutes 1 chart fedora-latest-x64 52 minutes 14 chart Other than smartos which only has one executor available there is no immediately obvious reason that I've seen why some of these are backed up more than others. at 1500UTC yesterday the top four in the table had between 5 and 10 waiting to run. There is no clear indication that one job caused things to back up.
List of top 20 initiators of node-test-pull-request jobs
Based on what is currently retained on jenkins
$ grep userId */build.xml | cut -d: -f2 | tr -d ' ' | sed 's,</*userId>,,g'| sort | uniq -c | sort -n | tail -20 12 geeksilva97 13 stefanstojanovic 14 atlowchemi 17 legendecas 22 juanarbol 24 richardlau 25 avivkeller 28 anonrig 29 renegade334 40 rafaelgss 41 mikemcc399 47 pimterry 54 mcollina 67 codebytere 72 daeyeon 126 aduh95 146 jasnell 221 panva 406 trivikr 762 nodejs-github-bothttps://openjs-foundation.slack.com/archives/C019Y2T6STH/p1787731505704119 has raised the issue of flake and started some actions.
Flake is going to cause extra load when collaborators use Jenkins "Resume build" to re-run tests.
Noting also that there have been
citgm-smokerruns for release builds in the last day which can use up some of the executors for a few hours increasing the pressure and therefore backlog on some of the machines. I'll also point out that smartos is not included in those citgm runs.I've modified logcompressor.sh to use -mtime +5 instead of -mtime +7 for now as we have a significant increase in the number of PR test jobs running at the moment.
Noting that the disk space on the jenkins controller is still quite low at 12GB having been up to about 48GB a couple of days ago. Again this is likely due to the load from the last few days. Once those logs become eligible for compression it should improve the situation. The first day of "high usage" (>100
node-test-pull-requestjobs) have already been compressed so I think we'll be safe ...Also noting that this morning there has been an issue with significant delays with github initiating jobs from PRs labelled with
request-ciso we should anticipate lots of jobs appearing once that is resolved.Reacted by Mike McCreadyNoting also that these two PRs which were causing issues with test case failures (this increasing the number of reruns being performed as Mike said earlier) have been merged:
- sea: keep ELF segments on separate pages in --build-sea output node#65564 (SEA test failure affecting
rhel8-x64only) - test: deflake fastutf8stream destroy and reopen tests node#65554 (fastutf8)
- sea: keep ELF segments on separate pages in --build-sea output node#65564 (SEA test failure affecting
- pinned this issue
on Aug 27, 2026 I've had to reduce log compressor to
-mtime +3due to getting below 1GB just now. It had already been reduced to 4 and dropping to 3 has given us about 40G of space. Hopefully that will be enough ... I've got various*du-sk*files showing the space used by different jobs in thejobsdirectory on the server so we can analyse what's chewing the most but I suspect it's for higher load reasons already mentioned.Server has 320G of disk space available. I had increased the compressor to
-mtime +5which has left the space available at about 15G but I've droped it to-mtime 4for safety since 95% utilised is a bit too close to the wire for my liking.Re. #4419 (comment)
I happened to noticed that the built-in node was running a job today, which isn't expected:
Usually we restrict the parent matrix jobs to
jenkins-workspaceso that the workspace directories get created on those machines and not on the built-in node.There are three workspace directories on the built-in node:
root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs# find . -maxdepth 2 -type d -name "workspace*" -exec du -hs {} \; 19M ./post-build-status-update/workspace@script 2.4G ./node-test-riscv64-experimental/workspace@3 2.3G ./node-test-riscv64-experimental/workspace@2 2.3G ./node-test-riscv64-experimental/workspace root@infra-digitalocean-ubuntu14-x64-1:/var/lib/jenkins/jobs#
I've edited node-test-riscv64-experimental to restrict the parent to run on
jenkins-workspaceand I'll clean up the three workspace directories above to reclaim the 6G of space.- added a commit that references this issue
on Sep 7, 2026 - added 2 commits that reference this issue
on Sep 16, 2026
Cloned from nodejs/jenkins-alerts#6348 as that will probably auto close as I free up some space.
Built-In Nodehas low space in Disk (used 98%).Please refer to the Jenkins Dashboard to check its status.
This issue has been auto-generated by UlisesGascon/jenkins-status-alerts-and-reporting.