Skip to content

Adjust to sidecar connection bound state - #4299

Open
bwoebi wants to merge 2 commits into
masterfrom
bob/connection-state
Open

bwoebi wants to merge 2 commits into
masterfrom
bob/connection-state

Conversation

@bwoebi

@bwoebi bwoebi commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

We also refactor the background sender to cleanly submit on its own separate state (telemetry for a different service directly), instead of having a dedicated connection and more hacks around that.

This also fixes a bunch of thread-mode sidecar reconnect issues. Also fixing the RC notification issue.

We also cleanup the crash-fd variable a bit to be more stable and less racy.

We also refactor the background sender to cleanly submit on its own separate state (telemetry for a different service directly), instead of having a dedicated connection and more hacks around that.

This also fixes a bunch of thread-mode sidecar reconnect issues.
Also fixing the RC notification issue.

We also cleanup the crash-fd variable a bit to be more stable and less racy.
@bwoebi
bwoebi requested review from a team as code owners October 9, 2026 20:13
@bwoebi
bwoebi requested review from btthomas and hhan2024 and removed request for a team October 9, 2026 20:13
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-09T20:18:03.530230Z 21664dc PR opened
🔒 Security Review ✅ Completed 2026-10-09T20:20:00.134935Z 21664dc PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@datadog-datadog-prod-us1

datadog-datadog-prod-us1 Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Tests

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 95.11%
• Overall Coverage: 68.35% (-0.20%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 4fce3a4 | Docs | View more details | Give us feedback!

@datadog-datadog-prod-us1 datadog-datadog-prod-us1 Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bits Code Review: FAIL

Concurrent PHP thread shutdown and request initialization can erase a replacement signal connection while leaving its owner recorded, disabling termination flushing or macOS crash reporting.

Open Bits AI session

🤖 Bits Code Review · Commit 21664dc

Comment thread ext/sidecar.c Outdated
static void dd_sidecar_disarm_signal_transport(ddog_SidecarTransport *owner) {
#ifndef _WIN32
uintptr_t expected = (uintptr_t)owner;
if (owner && atomic_compare_exchange_strong(&dd_sidecar_signal_owner, &expected, 0)) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Serialize signal teardown with ownership release

In a ZTS process, thread A can release signal ownership here, thread B can claim ownership and publish its connection during RINIT, and then A can clear B’s newly published connection. With Linux signal flushing enabled, this disables SIGTERM/SIGINT flushing and risks losing buffered traces; on macOS, it clears the crash-reporting FD. Subsequent RINITs cannot re-arm a healthy transport because B remains recorded as owner. Serialize connection clearing with ownership release, including the analogous failed-preparation path in dd_sidecar_arm_signal_transport; fixing teardown alone leaves that path vulnerable.

Was this helpful? React 👍 or 👎
🤖 Bits Code Review · Open Bits AI session

@pr-commenter

pr-commenter Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Benchmarks [ appsec ]

Benchmark execution time: 2026-10-09 21:25:18

Comparing candidate commit 4fce3a4 in PR branch bob/connection-state with baseline commit 01f58a4 in branch master.

📊 Benchmarking dashboard

Found 0 performance improvements and 0 performance regressions! Performance is the same for 12 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

- MSVC computed strlen() of the default service name before the NULL check
  in a conditional expression, crashing at connection setup whenever process
  tags are disabled (telemetry/config.phpt on every Windows job). Assign the
  slice in an if statement instead; same for the two other strlen() ternaries
  in exception_serialize.c and ffe.c.
- Clear the signal handlers' published connection before releasing its
  ownership, in the teardown and in the failed-preparation path, so that a
  thread taking over cannot have its connection cleared.
- The fork reconnection FFI moved into datadog-sidecar-ffi; bump libdatadog
  and regenerate the headers.
- Use list() in the orphan test for PHP 7.0, and clang-format appsec.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@pr-commenter

pr-commenter Bot commented Oct 9, 2026

Copy link
Copy Markdown

Benchmarks [ tracer ]

Benchmark execution time: 2026-10-09 22:10:34

Comparing candidate commit 4fce3a4 in PR branch bob/connection-state with baseline commit 01f58a4 in branch master.

📊 Benchmarking dashboard

Found 0 performance improvements and 1 performance regressions! Performance is the same for 193 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:HookBench/benchHookOverheadTraceFunction-opcache

  • 🟥 execution_time [+5.675µs; +11.067µs] or [+3.416%; +6.662%]

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant