Skip to content

Expose FFE serial IDs in flag metadata with span enrichment - #4247

Draft
hhan2024 wants to merge 1 commit into
masterfrom
ffe-serial-id-flag-metadata
Draft

hhan2024 wants to merge 1 commit into
masterfrom
ffe-serial-id-flag-metadata

Conversation

@hhan2024

Copy link
Copy Markdown

Description

Include the selected UFC split's serial ID in evaluation flag metadata under
__dd_split_serial_id when span enrichment is enabled. Preserve zero, omit
missing/null serial IDs, and retain existing metadata and exposure data.

The native evaluator passes the existing span-enrichment gate to ResultMapper.
No native extension changes are needed. Both the native client and OpenFeature
provider use this evaluator.

Companion system-tests draft: DataDog/system-tests#7857
(stacked on #7820 in that repository). Its activation remains pending the SDK
release version; this PR does not predict a release number.

Validation

  • PHP 8.2.28 FeatureFlags unit suite: 39 tests, 111 assertions passed.
  • New positive/zero/camel-case cases fail against the original mapper and pass
    with the fix. Coverage also checks missing/null IDs, disabled enrichment,
    array/object bridge results, and preservation of existing metadata.
  • composer ci-lint: passed.
  • All 5 serial-ID system tests passed using the PHP 1.25.1 native extension
    with this patched PHP source via DD_TRACE_SOURCES_PATH and
    DD_AUTOLOAD_NO_COMPILE=true, plus the companion adapter's empty-map JSON fix.
  • Full multi-version PHP test suite was not run locally.

Reviewer checklist

  • Test coverage seems ok.
  • Appropriate labels assigned.

Preserve zero and omit absent serial IDs while retaining existing flag metadata.

Environment: Datadog workspace

Co-Authored-By: Codex GPT-5 <noreply@localhost>
@datadog-datadog-prod-us1-2

datadog-datadog-prod-us1-2 Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Pipelines  Tests

❌ Errors

Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 2 Pipeline jobs failed

DataDog/apm-reliability/dd-trace-php | framework test: [wordpress_no_ddtrace]

View more details · View in GitLab

DataDog/apm-reliability/dd-trace-php | publish docker image for system tests

View more details · View in GitLab

ℹ️ Info

No other issues found (see more)

🧪 All tests passed
❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 100.00%
• Overall Coverage: 68.26% (-0.01%)

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: e3226c4 | Docs | View more details | Give us feedback!

@pr-commenter

pr-commenter Bot commented Sep 29, 2026

Copy link
Copy Markdown

Benchmarks [ tracer ]

Benchmark execution time: 2026-09-29 19:07:22

Comparing candidate commit e3226c4 in PR branch ffe-serial-id-flag-metadata with baseline commit d98de1e in branch master.

📊 Benchmarking dashboard

Found 0 performance improvements and 1 performance regressions! Performance is the same for 193 metrics, 0 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

scenario:HookBench/benchWithoutHook

  • 🟥 execution_time [+3.185µs; +5.265µs] or [+3.876%; +6.406%]

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant