Summary
When a recording contains one transition that is much stronger than the rest, detectBoundaries returns no boundaries after it except peaks almost as strong as the global maximum. The novelty curve is normalized by its single global maximum, and the relative threshold (default 0.3) is applied to that scale, so the gate effectively becomes rawNovelty >= 0.3 * globalMax. Genuine later changes that pass the absolute absoluteThreshold (0.005) are rejected by the relative gate alone.
On the recording below, one transition at 17.44 s has raw novelty 0.0656; the standard call returns six boundaries, none between 125.23 s and 335.11 s — although the recording has a ~25 dB energy collapse at 139.7 s and a drastic texture change at 216.2 s.
Environment
@libraz/libsonare 1.8.0, also reproduced with 1.8.1
- Analysis-only WASM entry (
@libraz/libsonare/analysis, sonare-analysis.wasm)
- Call:
detectBoundaries({ samples, sampleRate: 44100, peakDistance: 3 }), all other defaults (threshold: 0.3, absoluteThreshold: 0.005, kernelSize: 64, hop 512)
Input
Public YouTube video xX1Y0cxstBw ("A Klezmer Karnival", Philip Sparke), 339.63 s. Reproduce the PCM with:
yt-dlp -f 140 -o "%(id)s.%(ext)s" "https://www.youtube.com/watch?v=xX1Y0cxstBw"
ffmpeg -itsoffset -0.0362812 -i xX1Y0cxstBw.m4a -c:a copy xX1Y0cxstBw.corrected.m4a
ffmpeg -i xX1Y0cxstBw.corrected.m4a -f f32le -acodec pcm_f32le -ac 2 -ar 44100 pcm.f32
I've called detectBoundaries on the mono downmix created in JavaScript via
function monoDownmix(channels) {
if (channels.length === 0)
throw new Error('monoDownmix needs at least one channel');
if (channels.length === 1) return Float32Array.from(channels[0]);
const length = channels[0].length;
const mono = new Float32Array(length);
for (const channel of channels) {
if (channel.length !== length) {
throw new Error('monoDownmix channels must have equal length');
}
for (let i = 0; i < length; i++)
mono[i] += channel[i] / channels.length;
}
return mono;
}
Unfortunately, it wasn't until after I filed this issue that I realized that the m4a file I used locally were priming-corrected (see chkpnt/tutti-attacca-yt-audio-analysis for some background). Therefore I added the correction (-itsoffset -0.0362812) to the commands above, so you should get the same timestamps.
Actual result
returned boundaries (peakDistance 3):
1.32 s rel 0.3140 raw 0.02061
4.74 s rel 0.5303 raw 0.03481
17.44 s rel 1.0000 raw 0.06564 <- global maximum
93.81 s rel 0.3127 raw 0.02052 <- clears the gate by 4 %
125.23 s rel 0.6248 raw 0.04101
335.11 s rel 0.3359 raw 0.02205
The second half of the recording is not featureless. At least I would expect this two independently measurable events are in the 125.23-335.11 s gap:
- 139.7 s: the full band drops from about -24.6 dB (136 s) / -26.1 dB (138 s) to -51.8 dB (140 s) and stays 15-25 dB below the previous level for more than a minute. Any rehearsal-sections tool wants a boundary here.
- 216.2 s: the texture switches from sparse quiet material to a dense, fast pulse pattern. This is the strongest chroma self-similarity change of the whole recording (8 s checkerboard,
z = 5.3, value 0.1277), computed independently of detectBoundaries.
detectBoundaries reports none of them.
Root cause
compute_novelty_curve() (src/analysis/boundary_detector.cpp:370-392) divides the entire curve by its global maximum and stores it in novelty_peak_. detect_boundaries() (:394-448) then accepts a local maximum only if
novelty_curve_[i] >= config_.threshold && // relative, 0.3
novelty_curve_[i] * novelty_peak_ >= config_.absolute_threshold // absolute, 0.005
Here novelty_peak_ is the 17.44 s transition, 0.06564, so the effective raw gate is 0.3 * 0.06564 = 0.01969. The peaks in the gap all satisfy the absolute floor but not the relative gate:
| time (s) |
raw novelty |
relative |
decision |
| 133.12 |
0.01673 |
0.2549 |
dropped by relative threshold |
| 136.25 |
0.00896 |
0.1364 |
dropped |
| 139.67 |
0.01376 |
0.2096 |
dropped (25 dB energy collapse) |
| 142.97 |
0.00726 |
0.1107 |
dropped |
| 151.12 |
0.00868 |
0.1322 |
dropped |
| 168.67 |
0.00811 |
0.1236 |
dropped |
| 197.14 |
0.00688 |
0.1047 |
dropped |
| 216.22 |
0.00646 |
0.0984 |
dropped (texture change, strongest chroma-SSM peak) |
| 246.29 |
0.00703 |
0.1070 |
dropped |
| 301.07 |
0.00892 |
0.1359 |
dropped |
| 307.97 |
0.00718 |
0.1095 |
dropped |
| 323.71 |
0.01399 |
0.2132 |
dropped |
| 331.98 |
0.01428 |
0.2175 |
dropped |
| 335.11 |
0.02205 |
0.3359 |
kept |
The absolute_threshold documentation says stationary material tops out around 0.003 while genuine changes run 0.008 to 0.95; most dropped peaks sit in that genuine-change range. The relative gate is the only reason they are lost.
Related: the 128.10 s peak (raw 0.03133, relative 0.4774) is dropped by the 3 s peakDistance, 2.88 s after the stronger 125.23 s peak. With the 2 s default it is returned.
Why lowering threshold is not a fix
| threshold |
boundaries |
in 125.23-335.11 s |
still missed |
| 0.3 |
6 |
0 |
everything |
| 0.2 |
15 |
4 |
216.22, 197.14, ... |
| 0.15 |
18 |
4 |
216.22, 197.14, ... |
| 0.1 |
29 |
12 |
216.22 (0.0984, under by 0.0016) |
| 0 |
39 |
21 |
separates nothing: phrase-level peaks everywhere |
A single global fraction cannot both catch a peak at relative 0.098 and reject articulation peaks of the same raw magnitude in loud sections.
Proposed fix
Maybe detectBoundaries works as designed. In this case, it is up to the caller to run detectBoundaries on a sliding window of the samples.
But in my opinion, detectBoundaries should also work when analysing the example piece at whole and libsonare should normalize / threshold locally instead of against the global maximum. My AI agent (DeepSeek V4.1 Flash) is suggesting the following directions:
Option 1: rolling reference for the relative gate
Keep the same ratio idea, but replace the single track-wide maximum with a time-varying reference L(t):
keep peak at t if raw(t) >= k * L(t) and raw(t) >= absoluteThreshold, where L(t) is e.g. the maximum (or a high percentile) of the raw novelty curve inside a window around t (say ±30-60 s).
- Question it answers: "How big is this peak compared to the strongest change nearby?"
- Why it fixes the bug: the 17.44 s transition no longer sets the scale for the whole track; after ~60 s it has left the window, so the 139.7 s or 216.2 s peaks are compared against the max of their region.
- Shape of the change: small. The peak-picking loop (
detect_boundaries) and the peakDistance/absoluteThreshold logic stay as they are; only the normalization reference becomes a sliding statistic. Cheap and deterministic (sliding-window max is O(n)).
- Cost: one new parameter (window length) plus the existing
threshold. On our recording: k = 0.4, ±60 s → 14 candidates incl. 216.2 s; k = 0.3 → 17 incl. 139.7 s.
- Weakness: it is still a ratio to a reference. A strong and a moderate peak inside the same window still suppress each other; the result depends on where the window boundaries fall relative to the events. It essentially makes the current knob local instead of global.
Option 2: prominence/width-based peak picking
Drop the normalization reference entirely and judge each local maximum by its intrinsic shape:
- Prominence: for a local max, walk left and right until you reach higher curve values; the higher of the two valleys you cross is the base.
prominence = peakHeight - base. It's the height of the bump above the surrounding contour, independent of any track-wide or window-wide scale.
- Width: e.g. the width at half prominence (in frames/seconds), used to reject 1–2-frame spikes. Structural transitions usually have some duration; articulation clicks don't.
- Gate becomes: keep a local max if
prominence >= minProminence (or prominence/peakHeight >= k) and width >= minWidth, plus the existing peakDistance.
- Question it answers: "Does this bump stand out from the curve immediately around it, and is it broad enough to be a section change rather than a transient?"
- Why it fixes the bug: the base is local by construction — the dominant 17.44 s peak cannot change the score of a peak 200 s later. There is no window parameter to place, which removes Fix 1's main arbitrariness.
- Weakness: more invasive — it replaces the threshold logic, needs calibration of an absolute prominence floor in raw novelty units (the same calibration problem as
absoluteThreshold), and on this recording the curve is very spiky at ~1 s scale, so prominence alone would accept many articulation peaks. The width criterion is what makes it usable, and width depends on hop size/smoothing.
Option 1 is a localized version of the existing mechanism (low risk, keeps current semantics); Option 2 is a different mechanism (more principled "is this a real peak?", higher implementation cost and calibration effort). They can also be combined: a rolling reference as the primary gate, prominence/width as a secondary filter against single-frame spikes.
Summary
When a recording contains one transition that is much stronger than the rest,
detectBoundariesreturns no boundaries after it except peaks almost as strong as the global maximum. The novelty curve is normalized by its single global maximum, and the relativethreshold(default0.3) is applied to that scale, so the gate effectively becomesrawNovelty >= 0.3 * globalMax. Genuine later changes that pass the absoluteabsoluteThreshold(0.005) are rejected by the relative gate alone.On the recording below, one transition at 17.44 s has raw novelty
0.0656; the standard call returns six boundaries, none between 125.23 s and 335.11 s — although the recording has a ~25 dB energy collapse at 139.7 s and a drastic texture change at 216.2 s.Environment
@libraz/libsonare1.8.0, also reproduced with 1.8.1@libraz/libsonare/analysis,sonare-analysis.wasm)detectBoundaries({ samples, sampleRate: 44100, peakDistance: 3 }), all other defaults (threshold: 0.3,absoluteThreshold: 0.005,kernelSize: 64, hop 512)Input
Public YouTube video
xX1Y0cxstBw("A Klezmer Karnival", Philip Sparke), 339.63 s. Reproduce the PCM with:I've called detectBoundaries on the mono downmix created in JavaScript via
Unfortunately, it wasn't until after I filed this issue that I realized that the m4a file I used locally were priming-corrected (see chkpnt/tutti-attacca-yt-audio-analysis for some background). Therefore I added the correction (
-itsoffset -0.0362812) to the commands above, so you should get the same timestamps.Actual result
The second half of the recording is not featureless. At least I would expect this two independently measurable events are in the 125.23-335.11 s gap:
z = 5.3, value0.1277), computed independently ofdetectBoundaries.detectBoundariesreports none of them.Root cause
compute_novelty_curve()(src/analysis/boundary_detector.cpp:370-392) divides the entire curve by its global maximum and stores it innovelty_peak_.detect_boundaries()(:394-448) then accepts a local maximum only ifHere
novelty_peak_is the 17.44 s transition,0.06564, so the effective raw gate is0.3 * 0.06564 = 0.01969. The peaks in the gap all satisfy the absolute floor but not the relative gate:The
absolute_thresholddocumentation says stationary material tops out around0.003while genuine changes run0.008to0.95; most dropped peaks sit in that genuine-change range. The relative gate is the only reason they are lost.Related: the 128.10 s peak (raw
0.03133, relative0.4774) is dropped by the 3 speakDistance, 2.88 s after the stronger 125.23 s peak. With the 2 s default it is returned.Why lowering
thresholdis not a fixA single global fraction cannot both catch a peak at relative
0.098and reject articulation peaks of the same raw magnitude in loud sections.Proposed fix
Maybe
detectBoundariesworks as designed. In this case, it is up to the caller to rundetectBoundarieson a sliding window of the samples.But in my opinion,
detectBoundariesshould also work when analysing the example piece at whole and libsonare should normalize / threshold locally instead of against the global maximum. My AI agent (DeepSeek V4.1 Flash) is suggesting the following directions:Option 1: rolling reference for the relative gate
Keep the same ratio idea, but replace the single track-wide maximum with a time-varying reference
L(t):keep peak at
tifraw(t) >= k * L(t)andraw(t) >= absoluteThreshold, whereL(t)is e.g. the maximum (or a high percentile) of the raw novelty curve inside a window around t (say ±30-60 s).detect_boundaries) and thepeakDistance/absoluteThresholdlogic stay as they are; only the normalization reference becomes a sliding statistic. Cheap and deterministic (sliding-window max is O(n)).threshold. On our recording: k = 0.4, ±60 s → 14 candidates incl. 216.2 s; k = 0.3 → 17 incl. 139.7 s.Option 2: prominence/width-based peak picking
Drop the normalization reference entirely and judge each local maximum by its intrinsic shape:
prominence = peakHeight - base. It's the height of the bump above the surrounding contour, independent of any track-wide or window-wide scale.prominence >= minProminence(orprominence/peakHeight >= k) andwidth >= minWidth, plus the existingpeakDistance.absoluteThreshold), and on this recording the curve is very spiky at ~1 s scale, so prominence alone would accept many articulation peaks. The width criterion is what makes it usable, and width depends on hop size/smoothing.Option 1 is a localized version of the existing mechanism (low risk, keeps current semantics); Option 2 is a different mechanism (more principled "is this a real peak?", higher implementation cost and calibration effort). They can also be combined: a rolling reference as the primary gate, prominence/width as a secondary filter against single-frame spikes.