Skip to content

Parallel cram2bam - #2015

Merged
daviesrob merged 8 commits into
samtools:developfrom
jkbonfield:parallel-cram2bam
Jun 3, 2026
Merged

daviesrob merged 8 commits into
samtools:developfrom
jkbonfield:parallel-cram2bam

Conversation

@jkbonfield

Copy link
Copy Markdown
Contributor

The main thread does a lot of work in cram_to_bam and bam_set1. Io_lib's model for this was to put this code into the worker threads. I thought I'd done that here too (I have for the writer already), but apparently not. This PR fixes that.

A profile of the main thread:

$ ./test/test_view.dev -B -@32  ~/scratch/data/novaseq.10m.cram & perf record -t $! 2>/dev/null;perf report -n 2>/dev/null| head -20 | egrep -v '^#'
[1] 3634487
[1]+  Done                    ./test/test_view.dev -B -@32 ~/scratch/data/novaseq.10m.cram
    49.75%          2100  test_view.dev  test_view.dev      [.] bam_set1
     9.51%           407  test_view.dev  test_view.dev      [.] cram_to_bam
     5.93%           266  test_view.dev  libc.so.6          [.] __memmove_avx512_unaligned_erms
     3.13%           129  test_view.dev  libc.so.6          [.] __strncpy_evex
     2.50%            92  test_view.dev  [unknown]          [k] 0xffffffff82cb9b7e
     2.37%            97  test_view.dev  test_view.dev      [.] cram_get_seq
     2.25%            92  test_view.dev  test_view.dev      [.] sam_read1
     2.11%            92  test_view.dev  test_view.dev      [.] bam_cigar2rqlens
     1.65%            64  test_view.dev  test_view.dev      [.] cram_get_bam_seq

With this PR:

@ qpg-gpu-01[samtools.../htslib]; ./test/test_view.cram2bam -B -@32  ~/scratch/data/novaseq.10m.cram & perf record -t $! 2>/dev/null;perf report -n 2>/dev/null| head -20 | egrep -v '^#'
[1] 3634842
[1]+  Done                    ./test/test_view.cram2bam -B -@32 ~/scratch/data/novaseq.10m.cram
    15.65%           101  test_view.cram2  test_view.cram2bam  [.] cram_get_bam_seq
    11.31%            65  test_view.cram2  libc.so.6           [.] _int_free
     6.69%            38  test_view.cram2  [unknown]           [k] 0xffffffff829076fb
     5.17%            34  test_view.cram2  test_view.cram2bam  [.] cram_get_seq
     5.12%            29  test_view.cram2  [unknown]           [k] 0xffffffff8291f048
     4.73%            32  test_view.cram2  test_view.cram2bam  [.] sam_read1
     4.24%            24  test_view.cram2  [unknown]           [k] 0xffffffff828d0d41
     4.23%            24  test_view.cram2  [unknown]           [k] 0xffffffff8292a32a
     3.35%            19  test_view.cram2  [unknown]           [k] 0xffffffff82929c09

Wall clock benchmarks with 32 cores show a doubling in throughput (making it slightly faster than BAM, but that's due to the same main-thread excessive work problem there too):

$ for i in `seq 1 10`;do /usr/bin/time -f '%e %U %S'  ./test/test_view.dev 
-B -@32  ~/scratch/data/novaseq.10m.cram;done
1.14 10.34 0.76
1.06 9.13 0.69
1.11 9.42 0.72
1.12 10.20 0.62
1.00 8.72 0.57
1.23 8.85 0.77
1.46 9.13 0.79
1.17 8.67 0.63
1.30 8.73 0.69
0.99 9.64 0.54


$ for i in `seq 1 10`;do /usr/bin/time -f '%e %U %S'  ./test/test_view.cram2bam -B -@32  ~/scratch/data/novaseq.10m.cram;done
0.50 11.83 0.98
0.50 10.32 0.82
0.45 10.13 0.82
0.42 10.10 0.78
0.50 10.19 0.87
0.47 10.07 0.81
0.48 10.15 0.98
0.49 10.03 1.00
0.51 10.06 0.85
0.49 10.28 0.76

@daviesrob daviesrob self-assigned this May 21, 2026

@daviesrob daviesrob left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks like this does a good job of moving more work into the thread pool. Would it be worth doing the CG processing there too? I think it would just need bam_tag2cigar() to be made non-static so that it can be called from cram/cram_decode.c.

Comment thread cram/cram_decode.c Outdated
Comment thread cram/cram_decode.c Outdated
This means the creation of BAM objects is done in the worker thread
rather than in the main thread.  However we still incur the cost of
copy that data over
- Cache CG tag as sam_tag2cigar is a significant main thread cost
- Check BAM memory policy

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
- Spare_bams is now bam_list as we use the struct directly in slices too.
- We no longer use a single malloc block of memory for the entire bam list.
- We now copy the struct and swap the pointers in cram_get_bam_seq
  (provided bam_get_mempolicy is ok).  This is now workable as every
  bam data object is its own malloc block.

It's not as efficient as before, but it's safe.
This gives us just one malloc instead of 2 mallocs per BAM object,
which still gives us the ability to do pointer-swapping on the
bam->data but without paying the penalty of the bam struct
allocation.

However it's still slow than d62b978
(even with bam_copy1 always in use) with very large numbers of threads
simply due to the sheer pressure that bam_init1 puts on the malloc
system.  It's faster to memcpy than it is to malloc with lots of
threads.

In glibc it can be partially mitigated with

    export MALLOC_TOP_PAD_=100000000
    export MALLOC_MMAP_MAX_=0

Similarly using libtcmalloc.so mitigates it a bit, but it's still
slower here than earlier when using -@128 on my test file.  -@16 gives
a faster time however, so we need to figure out a better balance
somewhere.

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
This is already done via bam_set1 so it's duplicated effort.

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
@jkbonfield

jkbonfield commented Jun 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Looks like this does a good job of moving more work into the thread pool. Would it be worth doing the CG processing there too? I think it would just need bam_tag2cigar() to be made non-static so that it can be called from cram/cram_decode.c.

Possibly, but I doubt this code is called very often. It's only an issue on ultra-long ONT reads, and maybe not even there it's not that often any more due to improved quality. If we moved it, we'd still need a flag to say whether or not it's been done already, so that's changed.

I think I'd need to find some test data to evaluate. (Most of mine got lost on lustre.) It may be an unnecessary optimisation though.

What's more important appears to be careful malloc handling. As we add more threads, we add more bam-list allocations. These appear to cause malloc heap thrashing. GNU malloc is particularly badly performing here. TCMalloc is double the speed at high thread count as it has minimal slow down with more threads. GNU malloc performs badly as it gives back memory and then has to reallocate it again. Defining MALLOC_MMAP_MAX_=0 MALLOC_TOP_PAD_=100000000 solves this, but in theory it's not meant to be giving the memory back anyway as this is what the bam-list bit was meant to solve. Maybe the memory thrashing is somewhere else. I need to diagnose that.

(Io_lib just used mallopt to work around gnu malloc performance, but it's not ideal in a library)

@ [samtools.../htslib]; for i in `seq 1 5`;do /usr/bin/time -f '%e %U %S'  ./test/test_view -B -@64
  ~/scratch/data/novaseq.10m.31.cram;done
3.27 28.82 6.30
3.17 28.65 6.35
3.30 28.32 6.44
3.32 28.42 6.67
3.27 29.00 6.53

@ [samtools.../htslib]; for i in `seq 1 5`;do MALLOC_MMAP_MAX_=0 MALLOC_TOP_PAD_=100000000 /usr/bin/time -f '%e %U %S'  ./test/test_view -B -@64  ~/scratch/data/novaseq.10m.31.cram;done
1.47 28.74 2.88
1.49 29.27 2.91
1.46 28.18 2.71
1.48 28.72 2.71
1.48 28.13 2.79

@ [samtools.../htslib]; for i in `seq 1 5`;do LD_PRELOAD=~/scratch/opt/lib/libtcmalloc.so /usr/bin/time -f '%e %U %S'  ./test/test_view -B -@64  ~/scratch/data/novaseq.10m.31.cram;done
1.66 35.33 3.25
1.61 34.49 3.24
1.61 34.72 3.07
1.64 34.43 3.29
1.65 34.65 3.35

(Note this machine has 16 cores each with 2 hyperthreads, so this is just a stress test of memory by over-committing on bam-list arrays.)

Edit: It's not the bam_list allocations, but the cram record one:

    if (!(s->crecs = hts_malloc_p(sizeof(*s->crecs), s->hdr->num_records)))     
        return -1;                                                              

This is probably another PR, but I'll investigate

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
@jkbonfield
jkbonfield force-pushed the parallel-cram2bam branch from 09b94f1 to 8c6a406 Compare June 2, 2026 10:55
@jkbonfield

Copy link
Copy Markdown
Contributor Author

I have ideas for reducing malloc pressure, but it may come later depending on how long this sits in review (as it'll need to be based on top of these commits).

I've fixed the reallocs and removed needless code. I've rebased it too. I'll take a look at the bam_tag2cigar change as it may make the interface easier so perhaps it's worth it for that reason alone.

The avoids changing the cram_get_bam_seq function and removes some
conditional checks in the cram main thread.

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
@jkbonfield

Copy link
Copy Markdown
Contributor Author

I can squash when you wish, but figured it's helpful for review to not do that just now so you can see the new commits first. (Or squash it yourself if you wish.)

The bam_tag2cigar was quite simple to do.

@daviesrob daviesrob left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks OK, apart from a couple of potential leaks on realloc().

Comment thread cram/cram_decode.c Outdated
Comment thread cram/cram_decode.c Outdated
Comment thread cram/cram_encode.c Outdated
Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
@jkbonfield
jkbonfield force-pushed the parallel-cram2bam branch from d22a331 to 2b6ddeb Compare June 2, 2026 14:49
@jkbonfield

Copy link
Copy Markdown
Contributor Author

I went through all the reallocs in cram_encode.c and cram_decode.c and found a few more, so I fixed those too.
Now pushed up.

I also have another branch that creates a cram_list struct to go along the bam_list struct and reuses it between cram slices.
However it doesn't really seem to help much and we still need mallopt envs to coerce GNU malloc to not be so rubbish! (Or use tcmalloc). Frustrating. So for now I'm not adding that extra complexity.

I'm also now finding it hard to find any actual speed benefit on our modern machines (but it is faster on the old seq4 node). It's still provably faster by profiling the main thread, but something else is slowing it up now. Such are the vagaries of I/O and shared CPUs, but at least it can be faster when the conditions are favourable.

@daviesrob
daviesrob merged commit d926270 into samtools:develop Jun 3, 2026
9 checks passed
jkbonfield added a commit to jkbonfield/htslib that referenced this pull request Oct 8, 2026
Between 1.23 and 1.24 we improved multi-threading throughput,
especially for short read data.  However this also increased memory,
particularly on long reads.

This change (samtools#2015) moved conversion from CRAM to BAM into the
threaded portion, but this needs keeping an array of BAM objects in
memory.  Hence the increase in memory usage.

This is now mitigated by three measures.

1. Long read data has a variable number of records, meaning our
arrays of bam objects are differing in length.  We now free unused
bam->data pointers.  This doesn't affect short read data which
typically has a constant bam array size.

2. Long read data is also variable length sequence.  Reusing bam
structs repeatedly means over time they grow to hold the biggest
sequence seen so far.  This has the impact of gradually bloating the
memory usage over time.  We check for considerably over-sized objects
and free the bam data if we can re-allocate it to be much smaller.

3. Our process queue of decoded "results" (arrays of BAMs) was
previous 2 per thread.  This compensates for variable speeds of the
next step in pipeline, but 2* is overkill.  This has been reduced to
1.3* which appears to be a sweet spot that smooths out most speed
fluctuations while keeping the memory footprint lower.

The combined effects are noticable.

[ADD: BENCHMARK]

Fixes samtools/samtools#2392

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
jkbonfield added a commit to jkbonfield/htslib that referenced this pull request Oct 8, 2026
Between 1.23 and 1.24 we improved multi-threading throughput,
especially for short read data.  However this also increased memory,
particularly on long reads.

This change (samtools#2015) moved conversion from CRAM to BAM into the
threaded portion, but this needs keeping an array of BAM objects in
memory.  Hence the increase in memory usage.

This is now mitigated by three measures.

1. Long read data has a variable number of records, meaning our
arrays of bam objects are differing in length.  We now free unused
bam->data pointers.  This doesn't affect short read data which
typically has a constant bam array size.

2. Long read data is also variable length sequence.  Reusing bam
structs repeatedly means over time they grow to hold the biggest
sequence seen so far.  This has the impact of gradually bloating the
memory usage over time.  We check for considerably over-sized objects
and free the bam data if we can re-allocate it to be much smaller.

3. Our process queue of decoded "results" (arrays of BAMs) was
previous 2 per thread.  This compensates for variable speeds of the
next step in pipeline, but 2* is overkill.  This has been reduced to
1.3* which appears to be a sweet spot that smooths out most speed
fluctuations while keeping the memory footprint lower.

The combined effects are noticable.

[ADD: BENCHMARK]

Fixes samtools/samtools#2392

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
jkbonfield added a commit to jkbonfield/htslib that referenced this pull request Oct 8, 2026
Between 1.23 and 1.24 we improved multi-threading throughput,
especially for short read data.  However this also increased memory,
particularly on long reads.

This change (samtools#2015) moved conversion from CRAM to BAM into the
threaded portion, but this needs keeping an array of BAM objects in
memory.  Hence the increase in memory usage.

This is now mitigated by three measures.

1. Long read data has a variable number of records, meaning our
arrays of bam objects are differing in length.  We now free unused
bam->data pointers.  This doesn't affect short read data which
typically has a constant bam array size.

2. Long read data is also variable length sequence.  Reusing bam
structs repeatedly means over time they grow to hold the biggest
sequence seen so far.  This has the impact of gradually bloating the
memory usage over time.  We check for considerably over-sized objects
and free the bam data if we can re-allocate it to be much smaller.

3. Our process queue of decoded "results" (arrays of BAMs) was
previous 2 per thread.  This compensates for variable speeds of the
next step in pipeline, but 2* is overkill.  This has been reduced to
1.3* which appears to be a sweet spot that smooths out most speed
fluctuations while keeping the memory footprint lower.

The combined effects are noticable.

Small file checking:

ONT:  elapsed(s) user(s) system(s) // Memory (Mbyte)

    glibc -@4
    1.23             13.54 50.99 5.92 //  986648
    dev              14.26 50.64 6.45 // 2614588
    PR               14.16 51.04 6.36 // 1217348

    tcmalloc -@4
    1.23             12.41 49.57 1.28 // 1209644
    dev              13.11 50.72 1.72 // 3068996
    PR               13.02 50.37 1.77 // 1812032

ONT is showing the biggest memory jump between 1.23 and 1.24.
The PR is using much less memory, but is still more than the old poorly
threaded code.  (Although the poor threading doesn't show up on
long-read data.)

Revio:

    glibc -@16
    1.23             6.95 30.55 1.69 //  817972
    dev              2.57 33.12 4.81 // 1417952
    PR               2.51 32.20 4.67 //  882760

    tcmalloc -@16
    1.23             6.83 33.34 1.92 // 1113684
    dev              2.20 32.10 1.40 // 1631740
    PR               2.33 31.72 1.79 // 1108460

Alexios:

    glibc
    1.23             1.76  8.75 1.35 //  728852
    dev              0.83 11.49 2.19 // 1561664
    PR               0.86 12.15 2.75 // 1149276

    tcmalloc
    1.23             1.69  9.94 1.31 //  815064
    dev              0.77 17.97 2.08 // 1040580
    PR               0.71 16.88 1.74 // 1031224

NovaSeq:

    glibc -@16
    1.23             1.08 7.78 0.54 // 358644
    dev              0.67 8.15 1.20 // 320784
    PR               0.71 8.30 0.92 // 354684

The speed up of threading from 1.23 to 1.24 is very visible for Revio
,Alexios and NovaSeq.  This PR chasn't compromised those speed gains.

Memory usage is on par with 1.23 for NovaSeq and Revio, but not for
Alexios.

Tested on a full ONT file (and the file which instigated the samtools
stats issue) we see bigger memory changes.  This is likely because the
file being longer means the cached bam structs are even closer to the
extreme end of the range distribution.

    1.23
    337.81 2233.08 340.79 // 2442560

    1.24
    329.37 2101.39 370.11 // 11325584
    334.14 2132.62 382.73 // 11621448

    PR:
    313.17 2056.04 341.93 // 3179616
    312.87 2055.93 342.68 // 3031012

This shows a 3.6x reduction in memory compared to 1.24, although it's
still 1.3x larger than the previous release.

Fixes samtools/samtools#2392

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
jkbonfield added a commit to jkbonfield/htslib that referenced this pull request Oct 8, 2026
Between 1.23 and 1.24 we improved multi-threading throughput,
especially for short read data.  However this also increased memory,
particularly on long reads.

This change (samtools#2015) moved conversion from CRAM to BAM into the
threaded portion, but this needs keeping an array of BAM objects in
memory.  Hence the increase in memory usage.

This is now mitigated by three measures.

1. Long read data has a variable number of records, meaning our
arrays of bam objects are differing in length.  We now free unused
bam->data pointers.  This doesn't affect short read data which
typically has a constant bam array size.

2. Long read data is also variable length sequence.  Reusing bam
structs repeatedly means over time they grow to hold the biggest
sequence seen so far.  This has the impact of gradually bloating the
memory usage over time.  We check for considerably over-sized objects
and free the bam data if we can re-allocate it to be much smaller.

3. Our process queue of decoded "results" (arrays of BAMs) was
previous 2 per thread.  This compensates for variable speeds of the
next step in pipeline, but 2* is overkill.  This has been reduced to
1.3* which appears to be a sweet spot that smooths out most speed
fluctuations while keeping the memory footprint lower.

The combined effects are noticable.

Small file checking:

ONT:  elapsed(s) user(s) system(s) // Memory (Mbyte)

    glibc -@4
    1.23             13.54 50.99 5.92 //  986648
    dev              14.26 50.64 6.45 // 2614588
    PR               14.16 51.04 6.36 // 1217348

    tcmalloc -@4
    1.23             12.41 49.57 1.28 // 1209644
    dev              13.11 50.72 1.72 // 3068996
    PR               13.02 50.37 1.77 // 1812032

    jemalloc -@4
    1.23             13.83 50.90 7.51 //  942580
    dev              14.48 50.90 8.06 // 2682520
    PR               14.48 51.26 7.56 // 1317776

    mimalloc -@4
    1.23             13.74 52.84 4.80 // 1743040
    dev              14.14 52.03 5.02 // 3635190
    PR               14.50 51.90 6.82 // 1743640

ONT is showing the biggest memory jump between 1.23 and 1.24.
The PR is using much less memory, but is still more than the old poorly
threaded code.  (Although the poor threading doesn't show up on
long-read data.)

Revio:

    glibc -@16
    1.23             6.95 30.55 1.69 //  817972
    dev              2.57 33.12 4.81 // 1417952
    PR               2.51 32.20 4.67 //  882760

    tcmalloc -@16
    1.23             6.83 33.34 1.92 // 1113684
    dev              2.20 32.10 1.40 // 1631740
    PR               2.33 31.72 1.79 // 1108460

Alexios:

    glibc
    1.23             1.76  8.75 1.35 //  728852
    dev              0.83 11.49 2.19 // 1561664
    PR               0.86 12.15 2.75 // 1149276

    tcmalloc
    1.23             1.69  9.94 1.31 //  815064
    dev              0.77 17.97 2.08 // 1040580
    PR               0.71 16.88 1.74 // 1031224

    jemalloc
    1.23             1.67  8.74 1.20 // 1170995
    dev              0.45 11.35 1.71 // 1249407
    PR               0.48 11.55 1.65 // 1199570

    mimalloc
    1.23             1.72  9.16 1.16 // 1331333
    dev              0.54 13.06 1.59 // 1483500
    PR               0.56 12.84 1.54 // 1323964

NovaSeq:

    glibc -@16
    1.23             1.08 7.78 0.54 // 358644
    dev              0.67 8.15 1.20 // 320784
    PR               0.71 8.30 0.92 // 354684

The speed up of threading from 1.23 to 1.24 is very visible for Revio
,Alexios and NovaSeq.  This PR chasn't compromised those speed gains.

Memory usage is on par with 1.23 for NovaSeq and Revio, but not for
Alexios.

[Also as an aside, although we're working well on 4 different malloc
libraries on both short and long read data, it's clear that jemalloc
is a good contender for both multi-threaded speed and memory size.]

Tested on a full ONT file (and the file which instigated the samtools
stats issue) we see bigger memory changes.  This is likely because the
file being longer means the cached bam structs are even closer to the
extreme end of the range distribution.

    1.23
    337.81 2233.08 340.79 // 2442560

    1.24
    329.37 2101.39 370.11 // 11325584
    334.14 2132.62 382.73 // 11621448

    PR:
    313.17 2056.04 341.93 // 3179616
    312.87 2055.93 342.68 // 3031012

This shows a 3.6x reduction in memory compared to 1.24, although it's
still 1.3x larger than the previous release.

Fixes samtools/samtools#2392

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
jkbonfield added a commit to jkbonfield/htslib that referenced this pull request Oct 8, 2026
Between 1.23 and 1.24 we improved multi-threading throughput,
especially for short read data.  However this also increased memory,
particularly on long reads.

This change (samtools#2015) moved conversion from CRAM to BAM into the
threaded portion, but this needs keeping an array of BAM objects in
memory.  Hence the increase in memory usage.

This is now mitigated by three measures.

1. Long read data has a variable number of records, meaning our
arrays of bam objects are differing in length.  We now free unused
bam->data pointers.  This doesn't affect short read data which
typically has a constant bam array size.

2. Long read data is also variable length sequence.  Reusing bam
structs repeatedly means over time they grow to hold the biggest
sequence seen so far.  This has the impact of gradually bloating the
memory usage over time.  We check for considerably over-sized objects
and free the bam data if we can re-allocate it to be much smaller.

3. Our process queue of decoded "results" (arrays of BAMs) was
previous 2 per thread.  This compensates for variable speeds of the
next step in pipeline, but 2* is overkill.  This has been reduced to
1.3* which appears to be a sweet spot that smooths out most speed
fluctuations while keeping the memory footprint lower.

The combined effects are noticable.

Small file checking:

ONT:  elapsed(s) user(s) system(s) // Memory (Mbyte)

    glibc -@4
    1.23             13.54 50.99 5.92 //  986648
    dev              14.26 50.64 6.45 // 2614588
    PR               14.16 51.04 6.36 // 1217348

    tcmalloc -@4
    1.23             12.41 49.57 1.28 // 1209644
    dev              13.11 50.72 1.72 // 3068996
    PR               13.02 50.37 1.77 // 1812032

    jemalloc -@4
    1.23             13.83 50.90 7.51 //  942580
    dev              14.48 50.90 8.06 // 2682520
    PR               14.48 51.26 7.56 // 1317776

    mimalloc -@4
    1.23             13.74 52.84 4.80 // 1743040
    dev              14.14 52.03 5.02 // 3635190
    PR               14.50 51.90 6.82 // 1743640

ONT is showing the biggest memory jump between 1.23 and 1.24.
The PR is using much less memory, but is still more than the old poorly
threaded code.  (Although the poor threading doesn't show up on
long-read data.)

Revio:

    glibc -@16
    1.23             6.95 30.55 1.69 //  817972
    dev              2.57 33.12 4.81 // 1417952
    PR               2.51 32.20 4.67 //  882760

    tcmalloc -@16
    1.23             6.83 33.34 1.92 // 1113684
    dev              2.20 32.10 1.40 // 1631740
    PR               2.33 31.72 1.79 // 1108460

Axelios:

    glibc
    1.23             1.76  8.75 1.35 //  728852
    dev              0.83 11.49 2.19 // 1561664
    PR               0.86 12.15 2.75 // 1149276

    tcmalloc
    1.23             1.69  9.94 1.31 //  815064
    dev              0.77 17.97 2.08 // 1040580
    PR               0.71 16.88 1.74 // 1031224

    jemalloc
    1.23             1.67  8.74 1.20 // 1170995
    dev              0.45 11.35 1.71 // 1249407
    PR               0.48 11.55 1.65 // 1199570

    mimalloc
    1.23             1.72  9.16 1.16 // 1331333
    dev              0.54 13.06 1.59 // 1483500
    PR               0.56 12.84 1.54 // 1323964

NovaSeq:

    glibc -@16
    1.23             1.08 7.78 0.54 // 358644
    dev              0.67 8.15 1.20 // 320784
    PR               0.71 8.30 0.92 // 354684

The speed up of threading from 1.23 to 1.24 is very visible for Revio
,Alexios and NovaSeq.  This PR chasn't compromised those speed gains.

Memory usage is on par with 1.23 for NovaSeq and Revio, but not for
Alexios.

[Also as an aside, although we're working well on 4 different malloc
libraries on both short and long read data, it's clear that jemalloc
is a good contender for both multi-threaded speed and memory size.]

Tested on a full ONT file (and the file which instigated the samtools
stats issue) we see bigger memory changes.  This is likely because the
file being longer means the cached bam structs are even closer to the
extreme end of the range distribution.

    1.23
    337.81 2233.08 340.79 // 2442560

    1.24
    329.37 2101.39 370.11 // 11325584
    334.14 2132.62 382.73 // 11621448

    PR:
    313.17 2056.04 341.93 // 3179616
    312.87 2055.93 342.68 // 3031012

This shows a 3.6x reduction in memory compared to 1.24, although it's
still 1.3x larger than the previous release.

Fixes samtools/samtools#2392

Signed-off-by: James Bonfield <jkb@sanger.ac.uk>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants