Skip to content

Reduce CRAM decode memory usage when multi-threading - #2102

Open
jkbonfield wants to merge 1 commit into
samtools:developfrom
jkbonfield:cram-mem-opt2
Open

jkbonfield wants to merge 1 commit into
samtools:developfrom
jkbonfield:cram-mem-opt2

Conversation

@jkbonfield

@jkbonfield jkbonfield commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Between 1.23 and 1.24 we improved multi-threading throughput, especially for short read data. However this also increased memory, particularly on long reads.

This change (#2015) moved conversion from CRAM to BAM into the threaded portion, but this needs keeping an array of BAM objects in memory. Hence the increase in memory usage.

This is now mitigated by three measures.

  1. Long read data has a variable number of records, meaning our arrays of bam objects are differing in length. We now free unused bam->data pointers. This doesn't affect short read data which typically has a constant bam array size.

  2. Long read data is also variable length sequence. Reusing bam structs repeatedly means over time they grow to hold the biggest sequence seen so far. This has the impact of gradually bloating the memory usage over time. We check for considerably over-sized objects and free the bam data if we can re-allocate it to be much smaller.

  3. Our process queue of decoded "results" (arrays of BAMs) was previous 2 per thread. This compensates for variable speeds of the next step in pipeline, but 2* is overkill. This has been reduced to 1.3* which appears to be a sweet spot that smooths out most speed fluctuations while keeping the memory footprint lower.

The combined effects are noticable.

Small file checking:

ONT: elapsed(s) user(s) system(s) // Memory (Mbyte)

glibc -@4
1.23             13.54 50.99 5.92 //  986648
dev              14.26 50.64 6.45 // 2614588
PR               14.16 51.04 6.36 // 1217348

tcmalloc -@4
1.23             12.41 49.57 1.28 // 1209644
dev              13.11 50.72 1.72 // 3068996
PR               13.02 50.37 1.77 // 1812032

jemalloc -@4
1.23             13.83 50.90 7.51 //  942580
dev              14.48 50.90 8.06 // 2682520
PR               14.48 51.26 7.56 // 1317776

mimalloc -@4
1.23             13.74 52.84 4.80 // 1743040
dev              14.14 52.03 5.02 // 3635190
PR               14.50 51.90 6.82 // 1743640

ONT is showing the biggest memory jump between 1.23 and 1.24. The PR is using much less memory, but is still more than the old poorly threaded code. (Although the poor threading doesn't show up on long-read data.)

Revio:

glibc -@16
1.23             6.95 30.55 1.69 //  817972
dev              2.57 33.12 4.81 // 1417952
PR               2.51 32.20 4.67 //  882760

tcmalloc -@16
1.23             6.83 33.34 1.92 // 1113684
dev              2.20 32.10 1.40 // 1631740
PR               2.33 31.72 1.79 // 1108460

Axelios:

glibc
1.23             1.76  8.75 1.35 //  728852
dev              0.83 11.49 2.19 // 1561664
PR               0.86 12.15 2.75 // 1149276

tcmalloc
1.23             1.69  9.94 1.31 //  815064
dev              0.77 17.97 2.08 // 1040580
PR               0.71 16.88 1.74 // 1031224

jemalloc
1.23             1.67  8.74 1.20 // 1170995
dev              0.45 11.35 1.71 // 1249407
PR               0.48 11.55 1.65 // 1199570

mimalloc
1.23             1.72  9.16 1.16 // 1331333
dev              0.54 13.06 1.59 // 1483500
PR               0.56 12.84 1.54 // 1323964

NovaSeq:

glibc -@16
1.23             1.08 7.78 0.54 // 358644
dev              0.67 8.15 1.20 // 320784
PR               0.71 8.30 0.92 // 354684

The speed up of threading from 1.23 to 1.24 is very visible for Revio ,Axelios and NovaSeq. This PR chasn't compromised those speed gains.

Memory usage is on par with 1.23 for NovaSeq and Revio, but not for Alexios.

[Also as an aside, although we're working well on 4 different malloc libraries on both short and long read data, it's clear that jemalloc is a good contender for both multi-threaded speed and memory size.]

Tested on a full ONT file (and the file which instigated the samtools stats issue) we see bigger memory changes. This is likely because the file being longer means the cached bam structs are even closer to the extreme end of the range distribution.

1.23
337.81 2233.08 340.79 // 2442560

1.24
329.37 2101.39 370.11 // 11325584
334.14 2132.62 382.73 // 11621448

PR:
313.17 2056.04 341.93 // 3179616
312.87 2055.93 342.68 // 3031012

This shows a 3.6x reduction in memory compared to 1.24, although it's still 1.3x larger than the previous release.

Fixes samtools/samtools#2392

Between 1.23 and 1.24 we improved multi-threading throughput,
especially for short read data.  However this also increased memory,
particularly on long reads.

This change (samtools#2015) moved conversion from CRAM to BAM into the
threaded portion, but this needs keeping an array of BAM objects in
memory.  Hence the increase in memory usage.

This is now mitigated by three measures.

1. Long read data has a variable number of records, meaning our
arrays of bam objects are differing in length.  We now free unused
bam->data pointers.  This doesn't affect short read data which
typically has a constant bam array size.

2. Long read data is also variable length sequence.  Reusing bam
structs repeatedly means over time they grow to hold the biggest
sequence seen so far.  This has the impact of gradually bloating the
memory usage over time.  We check for considerably over-sized objects
and free the bam data if we can re-allocate it to be much smaller.

3. Our process queue of decoded "results" (arrays of BAMs) was
previous 2 per thread.  This compensates for variable speeds of the
next step in pipeline, but 2* is overkill.  This has been reduced to
1.3* which appears to be a sweet spot that smooths out most speed
fluctuations while keeping the memory footprint lower.

The combined effects are noticable.

Small file checking:

ONT:  elapsed(s) user(s) system(s) // Memory (Mbyte)

    glibc -@4
    1.23             13.54 50.99 5.92 //  986648
    dev              14.26 50.64 6.45 // 2614588
    PR               14.16 51.04 6.36 // 1217348

    tcmalloc -@4
    1.23             12.41 49.57 1.28 // 1209644
    dev              13.11 50.72 1.72 // 3068996
    PR               13.02 50.37 1.77 // 1812032

    jemalloc -@4
    1.23             13.83 50.90 7.51 //  942580
    dev              14.48 50.90 8.06 // 2682520
    PR               14.48 51.26 7.56 // 1317776

    mimalloc -@4
    1.23             13.74 52.84 4.80 // 1743040
    dev              14.14 52.03 5.02 // 3635190
    PR               14.50 51.90 6.82 // 1743640

ONT is showing the biggest memory jump between 1.23 and 1.24.
The PR is using much less memory, but is still more than the old poorly
threaded code.  (Although the poor threading doesn't show up on
long-read data.)

Revio:

    glibc -@16
    1.23             6.95 30.55 1.69 //  817972
    dev              2.57 33.12 4.81 // 1417952
    PR               2.51 32.20 4.67 //  882760

    tcmalloc -@16
    1.23             6.83 33.34 1.92 // 1113684
    dev              2.20 32.10 1.40 // 1631740
    PR               2.33 31.72 1.79 // 1108460

Axelios:

    glibc
    1.23             1.76  8.75 1.35 //  728852
    dev              0.83 11.49 2.19 // 1561664
    PR               0.86 12.15 2.75 // 1149276

    tcmalloc
    1.23             1.69  9.94 1.31 //  815064
    dev              0.77 17.97 2.08 // 1040580
    PR               0.71 16.88 1.74 // 1031224

    jemalloc
    1.23             1.67  8.74 1.20 // 1170995
    dev              0.45 11.35 1.71 // 1249407
    PR               0.48 11.55 1.65 // 1199570

    mimalloc
    1.23             1.72  9.16 1.16 // 1331333
    dev              0.54 13.06 1.59 // 1483500
    PR               0.56 12.84 1.54 // 1323964

NovaSeq:

    glibc -@16
    1.23             1.08 7.78 0.54 // 358644
    dev              0.67 8.15 1.20 // 320784
    PR               0.71 8.30 0.92 // 354684

The speed up of threading from 1.23 to 1.24 is very visible for Revio
,Alexios and NovaSeq.  This PR chasn't compromised those speed gains.

Memory usage is on par with 1.23 for NovaSeq and Revio, but not for
Alexios.

[Also as an aside, although we're working well on 4 different malloc
libraries on both short and long read data, it's clear that jemalloc
is a good contender for both multi-threaded speed and memory size.]

Tested on a full ONT file (and the file which instigated the samtools
stats issue) we see bigger memory changes.  This is likely because the
file being longer means the cached bam structs are even closer to the
extreme end of the range distribution.

    1.23
    337.81 2233.08 340.79 // 2442560

    1.24
    329.37 2101.39 370.11 // 11325584
    334.14 2132.62 382.73 // 11621448

    PR:
    313.17 2056.04 341.93 // 3179616
    312.87 2055.93 342.68 // 3031012

This shows a 3.6x reduction in memory compared to 1.24, although it's
still 1.3x larger than the previous release.

Fixes samtools/samtools#2392

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Memory usage of samtools stats on v1.24

1 participant