Skip to content

[Debugger] DEBUG-5828 Coordinate snapshot sampling per trace - #9266

Merged
dudikeleti merged 9 commits into
masterfrom
dudik/feat/snapshots-correlation
Sep 25, 2026
Merged

dudikeleti merged 9 commits into
masterfrom
dudik/feat/snapshots-correlation

Conversation

@dudikeleti

@dudikeleti dudikeleti commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary of changes

  • Coordinate Live Debugger sampling once per local trace for full-snapshot and capture-expression probes.
  • Share the first decision across all participating probes in the trace; each probe emits at most one snapshot per trace.
  • Preserve independent sampling for template-only logs, metric and span-decoration probes, Exception Replay, and hits without an active trace.
  • Correlate snapshots with the span context captured when probe processing begins, not when the snapshot is finalized.

Reason for change

  • Independent per-probe sampling produces incomplete parent/child/sibling snapshot chains within the same request, with no signal that anything is missing.
  • Align .NET with dd-trace-java#12452, the "Improving Correlation for Live Debugger Snapshots" RFC (Tier 1: existing trace context), and system-test expectations.

Implementation details

Which probes participate

  • A probe is coordinated when it is a full snapshot or has capture expressions (ShouldCoordinateSampling).
  • Everything else keeps its existing sampling path unchanged.

Flow at probe entry (ProbeProcessor.TryBeginProcess)

flowchart TD
    A[Probe hit] --> B{Has condition?}
    B -- yes --> C[No sampling yet.<br/>Capture span context, create snapshot creator]
    B -- no --> D{Coordinated probe<br/>and active trace?}
    D -- no --> E[Independent sampling:<br/>global limiter, then per-probe sampler]
    D -- yes --> F["traceContext.GetOrCreateDebuggerSamplingCoordinator()<br/>.TrySample(probeId, provider)"]
    F --> G{Decision}
    G -- Keep --> H[Capture span context,<br/>create snapshot creator, capture]
    G -- DropGlobal / DropProbe --> I[Skip before any capture work,<br/>record events.skipped]
    C --> J[Evaluate condition]
    J -- false --> K[Stop]
    J -- true --> F
Loading
  • Unconditional probes are sampled before any capture work, so a dropped hit costs a lookup, not a capture.
  • Conditional probes evaluate their condition first and join coordination only when it is true.
  • The span context is captured at entry and used for the snapshot's dd.trace_id / dd.span_id and for the trace lookup, so a snapshot finalized after its span closed still correlates correctly.

Per-trace coordinator (DebuggerSamplingCoordinator)

stateDiagram-v2
    [*] --> Undecided
    Undecided --> Keep: first hit's samplers accept (its probe ID claims a slot)
    Undecided --> DropGlobal: global limiter rejects
    Undecided --> DropProbe: per-probe sampler rejects
    Undecided --> Undecided: sampler threw (nothing published, waiting callers decide)
    Keep --> Keep: later hit claims its probe's slot once, further hits get DropProbe
Loading
  • TraceContext gains one nullable field and a lazy GetOrCreateDebuggerSamplingCoordinator() accessor, the same pattern as the AppSec, IAST and feature-flag per-trace state. All sampling logic lives in the debugger.
  • The coordinator is allocated only for traces that hit a coordinated probe.
  • The coordinator stores a DebuggerSamplingDecision directly. Undecided is the enum's zero value, so a new coordinator starts undecided and an unassigned decision never means Keep.
  • The first hit consults the global limiter and its per-probe sampler once. Once the trace is kept, other probes bypass their samplers and are capped at one snapshot per probe for the trace.
  • A dropped trace is answered by a lock-free volatile read. Kept traces take a short monitor lock to claim the probe's slot in a HashSet.
  • Concurrent first hits wait on the lock for the in-flight decision. The decision is published only after the samplers return, so a sampler exception leaves the trace undecided and a waiting caller decides instead.
  • The lock is reentrant, and the samplers can run customer code on the deciding thread (for example a first-chance exception handler) that hits another coordinated probe in the same trace. Monitor.IsEntered detects that nested call and returns DropProbe instead of re-entering the samplers, which could otherwise recurse without bound.
  • A constrained struct provider (IDebuggerSamplingDecisionProvider) avoids delegate allocation and boxing on the sampling path.

A claimed slot is never released

  • Once a probe claims its slot in a trace, the slot stays used whatever the outcome: a snapshot, a capture expression with no values, or a failed capture.
  • Once a trace is kept, samplers are no longer consulted, so the slot is the only throttle. Releasing it would let every later hit of the probe in that trace (loops, recursion) capture and evaluate again with no throttling.
  • A capture-expression result with no values and no errors depends on the probe definition, not on runtime values, so later hits would be empty too.

Test coverage

  • Unit (CoordinatedSamplingTests): shared Keep/Drop decisions across snapshot and capture-expression probes, per-probe caps, conditions, independent sampling for template logs, decisions local to each trace, concurrent first hits, reentrancy, sampler exception recovery, empty capture expressions keeping the slot (including a mid-trace probe update), skipped-event telemetry, and entry-time span correlation.
  • End-to-end (ProbesTests): CoordinatedSnapshotProbesEmitCompleteChains (a complete three-probe chain across sampling windows) and CoordinatedLineProbesEmitOncePerTrace (one snapshot per line probe per trace).

Other details

  • The global snapshot limiter is consulted once per participating trace; a kept trace may emit one snapshot for each participating probe.
  • Long-lived traces emit at most one snapshot per probe for the lifetime of their TraceContext.
  • Not supported: a probe updated mid-trace (same probe ID) doesn't emit again in that trace; it emits normally starting with the next trace.
  • Hits skipped by the per-trace cap are counted as reason:rateLimitProbe in events.skipped.
  • Condition-error diagnostic snapshots remain independently rate-limited.
  • Exception Replay stays independently sampled and keeps its captured correlation context.
  • dd.trace_id remains the existing 64-bit projection.
  • Enabling the .NET system-test manifest remains follow-up work.

@pr-commenter

pr-commenter Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-09-25 15:19:59

Comparing candidate commit d9f50fc in PR branch dudik/feat/snapshots-correlation with baseline commit 4919c62 in branch master.

📊 Benchmarking dashboard

Found 0 performance improvements and 11 performance regressions! Performance is the same for 61 metrics, 0 unstable metrics, 73 known flaky benchmarks, 53 flaky benchmarks without significant changes.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:Benchmarks.Trace.DbCommandBenchmark.ExecuteNonQuery net472

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+10.133%; +10.144%]
  • 🟥 throughput [-27969.741op/s; -24735.076op/s] or [-7.878%; -6.967%]

scenario:Benchmarks.Trace.DbCommandBenchmark.ExecuteNonQuery net6.0

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+9.750%; +9.761%]

scenario:Benchmarks.Trace.DbCommandBenchmark.ExecuteNonQuery netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+9.830%; +9.841%]
  • 🟥 throughput [-40461.035op/s; -32308.054op/s] or [-10.001%; -7.986%]

scenario:Benchmarks.Trace.HttpClientBenchmark.SendAsync net472

  • 🟥 allocated_mem [+304 bytes; +305 bytes] or [+10.472%; +10.483%]
  • 🟥 throughput [-12255.647op/s; -11646.939op/s] or [-13.991%; -13.296%]

scenario:Benchmarks.Trace.HttpClientBenchmark.SendAsync net6.0

  • 🟥 allocated_mem [+199 bytes; +200 bytes] or [+9.426%; +9.438%]

scenario:Benchmarks.Trace.HttpClientBenchmark.SendAsync netcoreapp3.1

  • 🟥 allocated_mem [+199 bytes; +200 bytes] or [+7.569%; +7.581%]
  • 🟥 throughput [-14613.047op/s; -12545.384op/s] or [-11.596%; -9.955%]

scenario:Benchmarks.Trace.NLogBenchmark.EnrichedLog net472

  • 🟥 throughput [-9036.045op/s; -7288.088op/s] or [-6.922%; -5.583%]

Known flaky benchmarks

These benchmarks are marked as flaky and will not trigger a failure. Modify FLAKY_BENCHMARKS_REGEX to control which benchmarks are marked as flaky.

scenario:Benchmarks.Trace.ActivityBenchmark.StartStopWithChild net472

  • 🟥 throughput [-9227.584op/s; -8370.881op/s] or [-10.941%; -9.925%]

scenario:Benchmarks.Trace.ActivityBenchmark.StartStopWithChild net6.0

  • 🟥 throughput [-9195.990op/s; -6644.860op/s] or [-7.730%; -5.585%]

scenario:Benchmarks.Trace.ActivityBenchmark.StartStopWithChild netcoreapp3.1

  • 🟥 throughput [-12033.006op/s; -10854.565op/s] or [-12.235%; -11.037%]

scenario:Benchmarks.Trace.AgentWriterBenchmark.WriteAndFlushEnrichedTraces net472

  • 🟥 allocated_mem [+1.637KB; +1.637KB] or [+49.731%; +49.747%]
  • 🟥 execution_time [+302.731ms; +304.771ms] or [+150.226%; +151.238%]
  • 🟥 throughput [-51.042op/s; -47.209op/s] or [-9.184%; -8.494%]

scenario:Benchmarks.Trace.AgentWriterBenchmark.WriteAndFlushEnrichedTraces net6.0

  • 🟥 allocated_mem [+1.011KB; +1.011KB] or [+37.485%; +37.498%]
  • 🟥 execution_time [+375.072ms; +376.464ms] or [+296.330%; +297.430%]
  • 🟩 throughput [+68.507op/s; +70.862op/s] or [+9.032%; +9.343%]

scenario:Benchmarks.Trace.AgentWriterBenchmark.WriteAndFlushEnrichedTraces netcoreapp3.1

  • 🟥 allocated_mem [+1.090KB; +1.090KB] or [+40.417%; +40.429%]
  • 🟥 execution_time [+399.714ms; +403.560ms] or [+353.732%; +357.135%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.AllCycleMoreComplexBody net472

  • 🟥 allocated_mem [+4.758KB; +4.758KB] or [+100.159%; +100.174%]
  • 🟥 throughput [-61753.513op/s; -61398.598op/s] or [-48.047%; -47.771%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.AllCycleMoreComplexBody net6.0

  • 🟥 allocated_mem [+3.880KB; +3.880KB] or [+82.058%; +82.072%]
  • 🟩 execution_time [-15.959ms; -11.787ms] or [-7.453%; -5.505%]
  • 🟥 throughput [-61179.857op/s; -58420.147op/s] or [-44.658%; -42.643%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.AllCycleMoreComplexBody netcoreapp3.1

  • 🟥 allocated_mem [+4.608KB; +4.608KB] or [+99.641%; +99.654%]
  • 🟥 throughput [-48919.335op/s; -46672.003op/s] or [-44.229%; -42.197%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.AllCycleSimpleBody net472

  • 🟥 allocated_mem [+1.380KB; +1.380KB] or [+111.648%; +111.662%]
  • 🟥 throughput [-295678.922op/s; -291657.303op/s] or [-30.190%; -29.780%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.AllCycleSimpleBody net6.0

  • 🟥 allocated_mem [+543 bytes; +544 bytes] or [+44.437%; +44.446%]
  • 🟩 execution_time [-26.862ms; -17.530ms] or [-11.979%; -7.818%]
  • 🟥 throughput [-119838.349op/s; -86736.433op/s] or [-12.803%; -9.266%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.AllCycleSimpleBody netcoreapp3.1

  • 🟥 allocated_mem [+1.344KB; +1.344KB] or [+111.253%; +111.267%]
  • 🟥 throughput [-181409.605op/s; -164288.618op/s] or [-26.065%; -23.605%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.ObjectExtractorMoreComplexBody net472

  • 🟥 allocated_mem [+3.378KB; +3.378KB] or [+89.003%; +89.017%]
  • 🟥 throughput [-74109.145op/s; -73330.389op/s] or [-49.875%; -49.351%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.ObjectExtractorMoreComplexBody net6.0

  • 🟥 allocated_mem [+3.336KB; +3.336KB] or [+88.150%; +88.161%]
  • 🟥 throughput [-75008.451op/s; -72129.531op/s] or [-47.727%; -45.895%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.ObjectExtractorMoreComplexBody netcoreapp3.1

  • 🟥 allocated_mem [+3.264KB; +3.264KB] or [+88.493%; +88.506%]
  • 🟥 throughput [-56307.261op/s; -53683.817op/s] or [-44.856%; -42.766%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.ObjectExtractorSimpleBody net6.0

  • 🟩 throughput [+200693.812op/s; +237798.547op/s] or [+6.692%; +7.929%]

scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.ObjectExtractorSimpleBody netcoreapp3.1

  • 🟩 execution_time [-19.040ms; -14.709ms] or [-8.777%; -6.781%]

scenario:Benchmarks.Trace.Asm.AppSecEncoderBenchmark.EncodeArgs net472

  • 🟩 allocated_mem [-13.759KB; -13.757KB] or [-42.326%; -42.318%]
  • 🟥 execution_time [+300.394ms; +301.320ms] or [+150.097%; +150.560%]
  • 🟩 throughput [+1027.912op/s; +1048.354op/s] or [+11.353%; +11.579%]

scenario:Benchmarks.Trace.Asm.AppSecEncoderBenchmark.EncodeArgs net6.0

  • 🟩 allocated_mem [-13.722KB; -13.718KB] or [-42.341%; -42.329%]
  • 🟥 execution_time [+299.391ms; +302.543ms] or [+150.984%; +152.573%]
  • 🟩 throughput [+2260.359op/s; +2471.892op/s] or [+17.288%; +18.906%]

scenario:Benchmarks.Trace.Asm.AppSecEncoderBenchmark.EncodeArgs netcoreapp3.1

  • 🟩 allocated_mem [-13.722KB; -13.718KB] or [-42.341%; -42.329%]
  • 🟥 execution_time [+300.151ms; +302.509ms] or [+151.193%; +152.380%]
  • 🟩 throughput [+1794.479op/s; +1919.419op/s] or [+17.325%; +18.531%]

scenario:Benchmarks.Trace.Asm.AppSecEncoderBenchmark.EncodeLegacyArgs net472

  • 🟥 execution_time [+296.827ms; +298.012ms] or [+145.790%; +146.372%]
  • 🟩 throughput [+592.638op/s; +600.610op/s] or [+15.711%; +15.922%]

scenario:Benchmarks.Trace.Asm.AppSecEncoderBenchmark.EncodeLegacyArgs net6.0

  • 🟥 execution_time [+298.260ms; +300.368ms] or [+145.808%; +146.839%]
  • 🟩 throughput [+2829.520op/s; +2863.705op/s] or [+41.108%; +41.604%]

scenario:Benchmarks.Trace.Asm.AppSecEncoderBenchmark.EncodeLegacyArgs netcoreapp3.1

  • 🟥 execution_time [+300.031ms; +301.036ms] or [+149.955%; +150.458%]
  • 🟩 throughput [+1423.055op/s; +1451.344op/s] or [+28.246%; +28.808%]

scenario:Benchmarks.Trace.Asm.AppSecWafBenchmark.RunWafRealisticBenchmark net472

  • 🟩 execution_time [-142.240µs; -137.312µs] or [-29.204%; -28.192%]
  • 🟩 throughput [+810.117op/s; +844.887op/s] or [+39.456%; +41.150%]

scenario:Benchmarks.Trace.Asm.AppSecWafBenchmark.RunWafRealisticBenchmark net6.0

  • 🟩 execution_time [-132.968µs; -105.875µs] or [-30.496%; -24.282%]
  • 🟩 throughput [+796.290op/s; +926.472op/s] or [+34.620%; +40.279%]

scenario:Benchmarks.Trace.Asm.AppSecWafBenchmark.RunWafRealisticBenchmark netcoreapp3.1

  • 🟩 execution_time [-142.411µs; -120.245µs] or [-30.512%; -25.763%]
  • 🟩 throughput [+773.280op/s; +859.036op/s] or [+35.696%; +39.655%]

scenario:Benchmarks.Trace.Asm.AppSecWafBenchmark.RunWafRealisticBenchmarkWithAttack net472

  • 🟩 execution_time [-125.265µs; -119.743µs] or [-33.821%; -32.330%]
  • 🟩 throughput [+1301.329op/s; +1374.871op/s] or [+48.195%; +50.919%]

scenario:Benchmarks.Trace.Asm.AppSecWafBenchmark.RunWafRealisticBenchmarkWithAttack net6.0

  • 🟩 execution_time [-98.096µs; -74.194µs] or [-31.317%; -23.686%]
  • 🟩 throughput [+1093.927op/s; +1306.753op/s] or [+34.101%; +40.735%]

scenario:Benchmarks.Trace.Asm.AppSecWafBenchmark.RunWafRealisticBenchmarkWithAttack netcoreapp3.1

  • 🟩 execution_time [-138.054µs; -115.674µs] or [-37.766%; -31.644%]
  • 🟩 throughput [+1335.477op/s; +1472.078op/s] or [+47.925%; +52.827%]

scenario:Benchmarks.Trace.AspNetCoreBenchmark.SendRequest net472

  • 🟥 execution_time [+300.633ms; +301.658ms] or [+150.047%; +150.558%]

scenario:Benchmarks.Trace.AspNetCoreBenchmark.SendRequest net6.0

  • unstable execution_time [+410.296ms; +420.072ms] or [+445.804%; +456.426%]

scenario:Benchmarks.Trace.AspNetCoreBenchmark.SendRequest netcoreapp3.1

  • unstable execution_time [+265.939ms; +320.298ms] or [+201.925%; +243.199%]

scenario:Benchmarks.Trace.CIVisibilityProtocolWriterBenchmark.WriteAndFlushEnrichedTraces net472

  • 🟥 allocated_mem [+2.848KB; +2.853KB] or [+5.059%; +5.069%]
  • unstable execution_time [+275.893ms; +336.892ms] or [+126.853%; +154.900%]
  • 🟥 throughput [-538.064op/s; -492.642op/s] or [-48.754%; -44.638%]

scenario:Benchmarks.Trace.CIVisibilityProtocolWriterBenchmark.WriteAndFlushEnrichedTraces net6.0

  • unstable execution_time [+209.861ms; +343.117ms] or [+89.434%; +146.222%]
  • 🟥 throughput [-670.757op/s; -587.267op/s] or [-44.740%; -39.171%]

scenario:Benchmarks.Trace.CIVisibilityProtocolWriterBenchmark.WriteAndFlushEnrichedTraces netcoreapp3.1

  • 🟥 execution_time [+337.661ms; +346.866ms] or [+201.961%; +207.466%]
  • 🟥 throughput [-405.175op/s; -368.121op/s] or [-28.212%; -25.632%]

scenario:Benchmarks.Trace.CharSliceBenchmark.OptimizedCharSliceWithPool netcoreapp3.1

  • unstable throughput [-10.524op/s; +55.243op/s] or [-1.964%; +10.312%]

scenario:Benchmarks.Trace.CharSliceBenchmark.OriginalCharSlice net6.0

  • 🟩 execution_time [-168.129µs; -116.089µs] or [-8.517%; -5.881%]
  • 🟩 throughput [+33.642op/s; +47.433op/s] or [+6.641%; +9.364%]

scenario:Benchmarks.Trace.ElasticsearchBenchmark.CallElasticsearch net472

  • 🟥 allocated_mem [+96 bytes; +97 bytes] or [+12.092%; +12.103%]
  • 🟥 execution_time [+302.102ms; +303.195ms] or [+152.133%; +152.683%]
  • 🟥 throughput [-25375.050op/s; -23571.990op/s] or [-8.165%; -7.585%]

scenario:Benchmarks.Trace.ElasticsearchBenchmark.CallElasticsearch net6.0

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+12.114%; +12.124%]
  • 🟥 execution_time [+299.015ms; +302.185ms] or [+149.837%; +151.425%]
  • 🟥 throughput [-40704.463op/s; -32657.856op/s] or [-6.417%; -5.149%]

scenario:Benchmarks.Trace.ElasticsearchBenchmark.CallElasticsearch netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+12.114%; +12.124%]
  • 🟥 execution_time [+301.252ms; +304.971ms] or [+151.336%; +153.205%]

scenario:Benchmarks.Trace.ElasticsearchBenchmark.CallElasticsearchAsync net472

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+11.165%; +11.177%]
  • 🟥 execution_time [+302.163ms; +303.646ms] or [+151.736%; +152.481%]
  • 🟥 throughput [-20766.754op/s; -19048.015op/s] or [-6.957%; -6.381%]

scenario:Benchmarks.Trace.ElasticsearchBenchmark.CallElasticsearchAsync net6.0

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+12.493%; +12.506%]
  • 🟥 execution_time [+300.135ms; +303.564ms] or [+148.404%; +150.099%]
  • 🟥 throughput [-58022.233op/s; -50753.999op/s] or [-9.349%; -8.177%]

scenario:Benchmarks.Trace.ElasticsearchBenchmark.CallElasticsearchAsync netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+11.422%; +11.434%]
  • 🟥 execution_time [+302.638ms; +306.425ms] or [+153.390%; +155.310%]
  • 🟥 throughput [-36034.633op/s; -26443.863op/s] or [-7.781%; -5.710%]

scenario:Benchmarks.Trace.GraphQLBenchmark.ExecuteAsync net472

  • 🟥 allocated_mem [+96 bytes; +97 bytes] or [+11.197%; +11.208%]
  • 🟥 execution_time [+299.916ms; +303.003ms] or [+150.531%; +152.080%]
  • 🟥 throughput [-31545.049op/s; -27022.726op/s] or [-8.184%; -7.010%]

scenario:Benchmarks.Trace.GraphQLBenchmark.ExecuteAsync net6.0

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+10.618%; +10.631%]
  • 🟥 execution_time [+301.372ms; +303.637ms] or [+150.207%; +151.335%]

scenario:Benchmarks.Trace.GraphQLBenchmark.ExecuteAsync netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+10.618%; +10.631%]
  • 🟥 execution_time [+299.969ms; +303.247ms] or [+149.232%; +150.863%]
  • 🟥 throughput [-32574.417op/s; -27255.306op/s] or [-7.710%; -6.451%]

scenario:Benchmarks.Trace.ILoggerBenchmark.EnrichedLog net472

  • 🟥 allocated_mem [+96 bytes; +97 bytes] or [+6.002%; +6.011%]
  • 🟥 throughput [-17744.650op/s; -15985.767op/s] or [-7.136%; -6.428%]

scenario:Benchmarks.Trace.ILoggerBenchmark.EnrichedLog net6.0

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+5.627%; +5.636%]
  • 🟩 execution_time [-16.183ms; -12.487ms] or [-7.525%; -5.806%]

scenario:Benchmarks.Trace.ILoggerBenchmark.EnrichedLog netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+5.605%; +5.617%]

scenario:Benchmarks.Trace.Iast.StringAspectsBenchmark.StringConcatAspectBenchmark net472

  • unstable execution_time [+10.288µs; +52.471µs] or [+2.541%; +12.961%]

scenario:Benchmarks.Trace.Iast.StringAspectsBenchmark.StringConcatAspectBenchmark net6.0

  • 🟩 allocated_mem [-20.056KB; -20.031KB] or [-7.316%; -7.307%]
  • unstable execution_time [-17.378µs; +39.949µs] or [-3.435%; +7.896%]

scenario:Benchmarks.Trace.Iast.StringAspectsBenchmark.StringConcatAspectBenchmark netcoreapp3.1

  • unstable execution_time [-59.426µs; +2.481µs] or [-10.298%; +0.430%]

scenario:Benchmarks.Trace.Iast.StringAspectsBenchmark.StringConcatBenchmark net6.0

  • unstable execution_time [+8.695µs; +14.168µs] or [+20.553%; +33.488%]
  • 🟥 throughput [-5852.500op/s; -3850.792op/s] or [-24.637%; -16.211%]

scenario:Benchmarks.Trace.Iast.StringAspectsBenchmark.StringConcatBenchmark netcoreapp3.1

  • unstable execution_time [-13.510µs; -5.020µs] or [-20.961%; -7.788%]
  • unstable throughput [+1317.178op/s; +3174.553op/s] or [+8.081%; +19.477%]

scenario:Benchmarks.Trace.Log4netBenchmark.EnrichedLog net472

  • 🟥 execution_time [+303.363ms; +308.194ms] or [+153.337%; +155.778%]

scenario:Benchmarks.Trace.Log4netBenchmark.EnrichedLog net6.0

  • 🟥 execution_time [+305.609ms; +309.953ms] or [+155.554%; +157.765%]

scenario:Benchmarks.Trace.Log4netBenchmark.EnrichedLog netcoreapp3.1

  • 🟥 execution_time [+298.787ms; +301.167ms] or [+149.580%; +150.771%]

scenario:Benchmarks.Trace.RedisBenchmark.SendReceive net472

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+9.197%; +9.208%]
  • 🟥 throughput [-33320.662op/s; -31567.429op/s] or [-9.224%; -8.739%]

scenario:Benchmarks.Trace.RedisBenchmark.SendReceive net6.0

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+9.226%; +9.238%]

scenario:Benchmarks.Trace.RedisBenchmark.SendReceive netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+9.226%; +9.238%]
  • 🟥 throughput [-47313.707op/s; -39049.649op/s] or [-11.198%; -9.242%]

scenario:Benchmarks.Trace.SerilogBenchmark.EnrichedLog net472

  • 🟥 execution_time [+296.905ms; +299.279ms] or [+147.980%; +149.164%]
  • 🟥 throughput [-12145.872op/s; -9760.010op/s] or [-8.020%; -6.445%]

scenario:Benchmarks.Trace.SerilogBenchmark.EnrichedLog net6.0

  • 🟥 allocated_mem [+96 bytes; +96 bytes] or [+6.000%; +6.010%]
  • 🟥 execution_time [+299.965ms; +302.019ms] or [+150.628%; +151.660%]

scenario:Benchmarks.Trace.SerilogBenchmark.EnrichedLog netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+5.819%; +5.828%]
  • 🟥 execution_time [+302.394ms; +304.821ms] or [+153.355%; +154.586%]

scenario:Benchmarks.Trace.SingleSpanAspNetCoreBenchmark.SingleSpanAspNetCore net472

  • 🟥 execution_time [+299.033ms; +300.791ms] or [+149.159%; +150.036%]
  • 🟩 throughput [+60885743.536op/s; +61243821.987op/s] or [+44.341%; +44.602%]

scenario:Benchmarks.Trace.SingleSpanAspNetCoreBenchmark.SingleSpanAspNetCore net6.0

  • unstable execution_time [+345.588ms; +404.198ms] or [+429.800%; +502.692%]

scenario:Benchmarks.Trace.SingleSpanAspNetCoreBenchmark.SingleSpanAspNetCore netcoreapp3.1

  • 🟥 execution_time [+299.570ms; +300.816ms] or [+149.418%; +150.040%]

scenario:Benchmarks.Trace.SpanBenchmark.StartFinishScope net472

  • 🟥 allocated_mem [+31 bytes; +32 bytes] or [+5.244%; +5.256%]
  • 🟥 throughput [-82749.183op/s; -77430.464op/s] or [-9.235%; -8.641%]

scenario:Benchmarks.Trace.SpanBenchmark.StartFinishSpan net472

  • 🟥 allocated_mem [+31 bytes; +32 bytes] or [+6.037%; +6.047%]
  • 🟥 throughput [-67538.378op/s; -63383.561op/s] or [-6.182%; -5.802%]

scenario:Benchmarks.Trace.SpanBenchmark.StartFinishSpan net6.0

  • 🟥 allocated_mem [+31 bytes; +32 bytes] or [+6.057%; +6.065%]

scenario:Benchmarks.Trace.SpanBenchmark.StartFinishSpan netcoreapp3.1

  • 🟥 allocated_mem [+31 bytes; +32 bytes] or [+6.059%; +6.069%]

scenario:Benchmarks.Trace.TraceAnnotationsBenchmark.RunOnMethodBegin net472

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+15.733%; +15.745%]
  • 🟥 throughput [-73368.858op/s; -69064.336op/s] or [-10.738%; -10.108%]

scenario:Benchmarks.Trace.TraceAnnotationsBenchmark.RunOnMethodBegin net6.0

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+14.810%; +14.823%]

scenario:Benchmarks.Trace.TraceAnnotationsBenchmark.RunOnMethodBegin netcoreapp3.1

  • 🟥 allocated_mem [+95 bytes; +96 bytes] or [+14.810%; +14.821%]
  • 🟥 throughput [-67999.054op/s; -53231.250op/s] or [-9.494%; -7.432%]

Known flaky benchmarks without significant changes:

  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_AddEvent_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_AddEvent_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_AddEvent_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_GetContext_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_GetContext_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_GetContext_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_SetAttributes_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_SetAttributes_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_SetAttributes_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_SetStatus_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_SetStatus_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_SetStatus_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_UpdateName_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_UpdateName_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.ActivityBenchmark.StartSpan_UpdateName_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_AddEvent_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_AddEvent_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_AddEvent_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_GetContext_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_GetContext_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_GetContext_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_RecordException_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_RecordException_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_RecordException_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_SetAttributes_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_SetAttributes_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_SetAttributes_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_SetStatus_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_SetStatus_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_SetStatus_Sampled netcoreapp3.1
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_UpdateName_Sampled net472
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_UpdateName_Sampled net6.0
  • scenario:Benchmarks.OpenTelemetry.InstrumentedApi.Trace.TelemetrySpanBenchmark.StartSpan_UpdateName_Sampled netcoreapp3.1
  • scenario:Benchmarks.Trace.Asm.AppSecBodyBenchmark.ObjectExtractorSimpleBody net472
  • scenario:Benchmarks.Trace.CharSliceBenchmark.OptimizedCharSlice net472
  • scenario:Benchmarks.Trace.CharSliceBenchmark.OptimizedCharSlice net6.0
  • scenario:Benchmarks.Trace.CharSliceBenchmark.OptimizedCharSlice netcoreapp3.1
  • scenario:Benchmarks.Trace.CharSliceBenchmark.OptimizedCharSliceWithPool net472
  • scenario:Benchmarks.Trace.CharSliceBenchmark.OptimizedCharSliceWithPool net6.0
  • scenario:Benchmarks.Trace.CharSliceBenchmark.OriginalCharSlice net472
  • scenario:Benchmarks.Trace.CharSliceBenchmark.OriginalCharSlice netcoreapp3.1
  • scenario:Benchmarks.Trace.Iast.StringAspectsBenchmark.StringConcatBenchmark net472
  • scenario:Benchmarks.Trace.SpanBenchmark.StartFinishScope net6.0
  • scenario:Benchmarks.Trace.SpanBenchmark.StartFinishScope netcoreapp3.1
  • scenario:Benchmarks.Trace.SpanBenchmark.StartFinishTwoScopes net472
  • scenario:Benchmarks.Trace.SpanBenchmark.StartFinishTwoScopes net6.0
  • scenario:Benchmarks.Trace.SpanBenchmark.StartFinishTwoScopes netcoreapp3.1

@dd-trace-dotnet-ci-bot

dd-trace-dotnet-ci-bot Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Execution-Time Benchmarks Report ⏱️

Execution-time results for samples comparing This PR (9266) and master.

✅ No regressions detected

📄 View the full report (charts + all metrics) →

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It changes concurrent sampling behavior on a customer-process hot path and warrants final human validation.

Pull request overview

Coordinates Live Debugger snapshot sampling per local trace while preserving independent sampling for excluded probe types.

Changes:

  • Adds trace-scoped sampling state with concurrency-safe per-probe emission caps.
  • Captures entry-time span context for stable snapshot correlation.
  • Adds comprehensive unit and integration coverage.
File summaries
File Description
CoordinatedSamplingTest.cs Adds an end-to-end sampling fixture.
CoordinatedSamplingTests.cs Tests coordination, concurrency, caps, and correlation.
ProbesTests.cs Verifies complete chains and line-probe caps.
TraceContext.cs Stores trace-scoped debugger sampling state.
DebuggerSnapshotCreator.cs Preserves entry-time span context.
IDebuggerSamplingDecisionProvider.cs Defines allocation-free decision abstraction.
DebuggerSamplingCoordinator.cs Implements synchronized trace-level decisions.
ProbeProcessor.cs Integrates coordinated sampling into probe processing.
Review details
  • Files reviewed: 8/8 changed files
  • Comments generated: 0
  • Review effort level: Balanced

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@dudikeleti

Copy link
Copy Markdown
Contributor Author

/codex review

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The coordination lifecycle, failure handling, entry-time correlation, and per-probe limits are coherent and comprehensively tested.

Review details
  • Files reviewed: 8/8 changed files
  • Comments generated: 0 new
  • Review effort level: Balanced

@dudikeleti
dudikeleti requested a review from jpbempel September 18, 2026 11:08
@dudikeleti
dudikeleti marked this pull request as ready for review September 18, 2026 12:25
@dudikeleti
dudikeleti requested review from a team as code owners September 18, 2026 12:25
@dudikeleti
dudikeleti requested review from zacharycmontoya and removed request for a team September 18, 2026 12:25
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-18T12:30:41.376009Z b645573 Draft marked ready
🔒 Security Review ✅ Completed 2026-09-18T12:36:30.809785Z b645573 Draft marked ready

Security findings

Advisory findings (1)

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@dudikeleti
dudikeleti removed the request for review from zacharycmontoya September 18, 2026 12:26
@dudikeleti
dudikeleti force-pushed the dudik/feat/snapshots-correlation branch from b645573 to cc74db2 Compare September 18, 2026 12:29

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b6455734d6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛡️ Codex Security Review · Automatically triggered

Here are some automated security review suggestions for this pull request.

Reviewed commit: b6455734d6

ℹ️ About Codex security reviews in GitHub

This is an experimental Codex feature. Security reviews are triggered when:

  • You comment "@codex security review"
  • A regular code review gets triggered (for example, "@codex review" or when a PR is opened), and you’re opted in so security review runs alongside code review

Once complete, Codex will leave suggestions, or a comment if no findings are found.

Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs Outdated
Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The implementation matches the stated sampling semantics and includes focused coverage for concurrency, failure recovery, and correlation.

Review effort: Balanced
Findings: None

Comment thread tracer/src/Datadog.Trace/Debugger/RateLimiting/DebuggerSamplingCoordinator.cs Outdated
Comment thread tracer/src/Datadog.Trace/Debugger/RateLimiting/DebuggerSamplingCoordinator.cs Outdated
@dudikeleti
dudikeleti removed the request for review from P403n1x87 September 24, 2026 18:14
@dudikeleti
dudikeleti force-pushed the dudik/feat/snapshots-correlation branch from 6493c0b to 0cba76a Compare September 25, 2026 08:23

@andrewlock andrewlock left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TraceContext addition looks fine to me. It'll impact the microbenchmarks due to the extra pointer-size allocation but I don't think there's any way around that

Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs Outdated
Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs
Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs Outdated
Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs Outdated
Comment thread tracer/src/Datadog.Trace/Debugger/Expressions/ProbeProcessor.cs Outdated
Comment thread tracer/src/Datadog.Trace/Debugger/RateLimiting/DebuggerSamplingCoordinator.cs Outdated
Comment thread tracer/src/Datadog.Trace/Debugger/RateLimiting/DebuggerSamplingCoordinator.cs Outdated
Share the first sampling decision across capturing probes and cap each probe to one snapshot per local trace. Preserve independent log and Exception Replay behavior, and stamp correlation IDs at capture start.
Release unused capture-expression reservations and retain rate-limit reasons for coordinated skip telemetry.
…empty

Releasing the slot let every later hit in a kept trace capture and
evaluate again without any sampler throttling, and it gained nothing:
an empty result depends on the probe definition, so later hits are
empty too. A probe updated mid-trace won't emit until the next trace.
TraceContext now only lazily creates the per-trace
DebuggerSamplingCoordinator, like the AppSec, IAST and feature-flag
state. Sampling logic stays in the debugger. The coordinator becomes
the per-trace state class instead of a static wrapper around a nested
State.
Store DebuggerSamplingDecision directly in the coordinator instead of a
separate private state enum, with Undecided as the zero value so an
unassigned decision never means Keep. Nested calls on the deciding
thread are caught with Monitor.IsEntered rather than a Creating state.
The decision is only published after Sample() returns, so a throwing
sampler leaves the trace undecided without a reset.
@dudikeleti
dudikeleti force-pushed the dudik/feat/snapshots-correlation branch from 0cba76a to d9f50fc Compare September 25, 2026 14:26
@dudikeleti
dudikeleti enabled auto-merge (squash) September 25, 2026 14:28
@dudikeleti
dudikeleti merged commit 8e5ba5d into master Sep 25, 2026
147 of 149 checks passed
@dudikeleti
dudikeleti deleted the dudik/feat/snapshots-correlation branch September 25, 2026 22:05
@github-actions github-actions Bot added this to the vNext-v3 milestone Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants