Repository navigation
Reduce CRAM decode memory usage when multi-threading - #2102
Open
jkbonfield wants to merge 1 commit into
Open
jkbonfield wants to merge 1 commit into
jkbonfield wants to merge 1 commit into
Conversation
Between 1.23 and 1.24 we improved multi-threading throughput, especially for short read data. However this also increased memory, particularly on long reads. This change (samtools#2015) moved conversion from CRAM to BAM into the threaded portion, but this needs keeping an array of BAM objects in memory. Hence the increase in memory usage. This is now mitigated by three measures. 1. Long read data has a variable number of records, meaning our arrays of bam objects are differing in length. We now free unused bam->data pointers. This doesn't affect short read data which typically has a constant bam array size. 2. Long read data is also variable length sequence. Reusing bam structs repeatedly means over time they grow to hold the biggest sequence seen so far. This has the impact of gradually bloating the memory usage over time. We check for considerably over-sized objects and free the bam data if we can re-allocate it to be much smaller. 3. Our process queue of decoded "results" (arrays of BAMs) was previous 2 per thread. This compensates for variable speeds of the next step in pipeline, but 2* is overkill. This has been reduced to 1.3* which appears to be a sweet spot that smooths out most speed fluctuations while keeping the memory footprint lower. The combined effects are noticable. Small file checking: ONT: elapsed(s) user(s) system(s) // Memory (Mbyte) glibc -@4 1.23 13.54 50.99 5.92 // 986648 dev 14.26 50.64 6.45 // 2614588 PR 14.16 51.04 6.36 // 1217348 tcmalloc -@4 1.23 12.41 49.57 1.28 // 1209644 dev 13.11 50.72 1.72 // 3068996 PR 13.02 50.37 1.77 // 1812032 jemalloc -@4 1.23 13.83 50.90 7.51 // 942580 dev 14.48 50.90 8.06 // 2682520 PR 14.48 51.26 7.56 // 1317776 mimalloc -@4 1.23 13.74 52.84 4.80 // 1743040 dev 14.14 52.03 5.02 // 3635190 PR 14.50 51.90 6.82 // 1743640 ONT is showing the biggest memory jump between 1.23 and 1.24. The PR is using much less memory, but is still more than the old poorly threaded code. (Although the poor threading doesn't show up on long-read data.) Revio: glibc -@16 1.23 6.95 30.55 1.69 // 817972 dev 2.57 33.12 4.81 // 1417952 PR 2.51 32.20 4.67 // 882760 tcmalloc -@16 1.23 6.83 33.34 1.92 // 1113684 dev 2.20 32.10 1.40 // 1631740 PR 2.33 31.72 1.79 // 1108460 Axelios: glibc 1.23 1.76 8.75 1.35 // 728852 dev 0.83 11.49 2.19 // 1561664 PR 0.86 12.15 2.75 // 1149276 tcmalloc 1.23 1.69 9.94 1.31 // 815064 dev 0.77 17.97 2.08 // 1040580 PR 0.71 16.88 1.74 // 1031224 jemalloc 1.23 1.67 8.74 1.20 // 1170995 dev 0.45 11.35 1.71 // 1249407 PR 0.48 11.55 1.65 // 1199570 mimalloc 1.23 1.72 9.16 1.16 // 1331333 dev 0.54 13.06 1.59 // 1483500 PR 0.56 12.84 1.54 // 1323964 NovaSeq: glibc -@16 1.23 1.08 7.78 0.54 // 358644 dev 0.67 8.15 1.20 // 320784 PR 0.71 8.30 0.92 // 354684 The speed up of threading from 1.23 to 1.24 is very visible for Revio ,Alexios and NovaSeq. This PR chasn't compromised those speed gains. Memory usage is on par with 1.23 for NovaSeq and Revio, but not for Alexios. [Also as an aside, although we're working well on 4 different malloc libraries on both short and long read data, it's clear that jemalloc is a good contender for both multi-threaded speed and memory size.] Tested on a full ONT file (and the file which instigated the samtools stats issue) we see bigger memory changes. This is likely because the file being longer means the cached bam structs are even closer to the extreme end of the range distribution. 1.23 337.81 2233.08 340.79 // 2442560 1.24 329.37 2101.39 370.11 // 11325584 334.14 2132.62 382.73 // 11621448 PR: 313.17 2056.04 341.93 // 3179616 312.87 2055.93 342.68 // 3031012 This shows a 3.6x reduction in memory compared to 1.24, although it's still 1.3x larger than the previous release. Fixes samtools/samtools#2392 Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
jkbonfield
force-pushed
the
cram-mem-opt2
branch
from
October 8, 2026 14:06
aa75845 to
227804f
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Between 1.23 and 1.24 we improved multi-threading throughput, especially for short read data. However this also increased memory, particularly on long reads.
This change (#2015) moved conversion from CRAM to BAM into the threaded portion, but this needs keeping an array of BAM objects in memory. Hence the increase in memory usage.
This is now mitigated by three measures.
Long read data has a variable number of records, meaning our arrays of bam objects are differing in length. We now free unused bam->data pointers. This doesn't affect short read data which typically has a constant bam array size.
Long read data is also variable length sequence. Reusing bam structs repeatedly means over time they grow to hold the biggest sequence seen so far. This has the impact of gradually bloating the memory usage over time. We check for considerably over-sized objects and free the bam data if we can re-allocate it to be much smaller.
Our process queue of decoded "results" (arrays of BAMs) was previous 2 per thread. This compensates for variable speeds of the next step in pipeline, but 2* is overkill. This has been reduced to 1.3* which appears to be a sweet spot that smooths out most speed fluctuations while keeping the memory footprint lower.
The combined effects are noticable.
Small file checking:
ONT: elapsed(s) user(s) system(s) // Memory (Mbyte)
ONT is showing the biggest memory jump between 1.23 and 1.24. The PR is using much less memory, but is still more than the old poorly threaded code. (Although the poor threading doesn't show up on long-read data.)
Revio:
Axelios:
NovaSeq:
The speed up of threading from 1.23 to 1.24 is very visible for Revio ,Axelios and NovaSeq. This PR chasn't compromised those speed gains.
Memory usage is on par with 1.23 for NovaSeq and Revio, but not for Alexios.
[Also as an aside, although we're working well on 4 different malloc libraries on both short and long read data, it's clear that jemalloc is a good contender for both multi-threaded speed and memory size.]
Tested on a full ONT file (and the file which instigated the samtools stats issue) we see bigger memory changes. This is likely because the file being longer means the cached bam structs are even closer to the extreme end of the range distribution.
This shows a 3.6x reduction in memory compared to 1.24, although it's still 1.3x larger than the previous release.
Fixes samtools/samtools#2392