Skip to content

fix(openfeature): harden agentless EVP fallback - #9988

Closed
leoromanovsky wants to merge 7 commits into
leo.romanovsky/ffl-2446-evp-flagevaluation-nodejsfrom
leo.romanovsky/ffe-agentless-evp-node-hardening
Closed

leoromanovsky wants to merge 7 commits into
leo.romanovsky/ffl-2446-evp-flagevaluation-nodejsfrom
leo.romanovsky/ffe-agentless-evp-node-hardening

Conversation

@leoromanovsky

@leoromanovsky leoromanovsky commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Exposure and flag-evaluation writers currently perform separate /info discovery, retry local discovery before selecting direct intake, and replay a batch directly after ambiguous local transport failures.

Changes

  • Discover EVP support once when Feature Flags activate.
  • Share the selected route with both event writers.
  • Prefer local EVP v4, then v2.
  • Select direct intake when no compatible local route is available and DD_API_KEY is configured.
  • Switch to direct intake after definitive local-route failures.

Decisions

  • Replay the current batch only after connection refusal, DNS failure, missing Unix socket, or HTTP 403, 404, or 405.
  • After connection reset, broken pipe, or timeout, switch only future batches.
  • Keep HTTP 429 and 5xx responses on the selected local route.

Validation

  • 196 focused OpenFeature and EVP discovery tests pass.
  • Relevant ESLint checks pass.
  • npm run verify:config:types passes.
  • DataDog/system-tests#7601 validates flag-evaluation delivery through Agent, agentless direct, and agentless sidecar routes.
  • DataDog/system-tests#7602 validates exposure delivery through the same routes.

@dd-octo-sts

dd-octo-sts Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Overall package size

Self size: 8.45 MB
Deduped: 9.11 MB
No deduping: 9.11 MB

Dependency sizes | name | version | self size | total size | |------|---------|-----------|------------| | import-in-the-middle | 3.3.3 | 125.43 kB | 445.14 kB | | opentracing | 0.14.7 | 194.81 kB | 194.81 kB | | dc-polyfill | 0.1.11 | 25.74 kB | 25.74 kB |

🤖 This report was automatically generated by heaviest-objects-in-the-universe

Comment thread packages/dd-trace/src/evp_proxy/direct.js Fixed
@datadog-prod-us1-5

datadog-prod-us1-5 Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

Pipelines  Tests

✨ Unblock PR with BitsAI

⚠️ Warnings

❌ Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 6 Pipeline jobs failed

All Green | all-green — ❌ 8 tests failed

View more details · View in GitHub Actions

No config file found, ignoring config.

❌ FlaggingProvider Initialization Timeout does not keep the process alive while waiting for configuration from FlaggingProvider Initialization Timeout
Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED')

TypeError: Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED')
    at new FlaggingProvider (packages/dd-trace/src/openfeature/flagging_provider.js:41:29)
    at Context.<anonymous> (packages/dd-trace/test/openfeature/flagging_provider_timeout.spec.js:99:22)
    at /home/runner/work/dd-trace-js/dd-trace-js/packages/dd-trace/test/setup/mocha-hooks.js:56:18
    at Hook.patchedRunHook [as run] (packages/dd-trace/test/setup/mocha-hooks.js:53:30)
    at process.processImmediate (node:internal/timers:483:21)
    at process.callbackTrampoline (node:internal/async_hooks:130:17)
❌ FlaggingProvider Initialization Timeout should allow recovery if configuration is set after timeout from FlaggingProvider Initialization Timeout
Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED')

TypeError: Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED')
    at new FlaggingProvider (packages/dd-trace/src/openfeature/flagging_provider.js:41:29)
    at Context.<anonymous> (packages/dd-trace/test/openfeature/flagging_provider_timeout.spec.js:173:22)
    at /home/runner/work/dd-trace-js/dd-trace-js/packages/dd-trace/test/setup/mocha-hooks.js:56:18
    at Hook.patchedRunHook [as run] (packages/dd-trace/test/setup/mocha-hooks.js:53:30)
    at process.processImmediate (node:internal/timers:483:21)
    at process.callbackTrampoline (node:internal/async_hooks:130:17)
❌ FlaggingProvider Initialization Timeout should call setError with timeout message after 30 seconds from FlaggingProvider Initialization Timeout
Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED')

TypeError: Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED')
    at new FlaggingProvider (packages/dd-trace/src/openfeature/flagging_provider.js:41:29)
    at Context.<anonymous> (packages/dd-trace/test/openfeature/flagging_provider_timeout.spec.js:148:22)
    at /home/runner/work/dd-trace-js/dd-trace-js/packages/dd-trace/test/setup/mocha-hooks.js:56:18
    at Hook.patchedRunHook [as run] (packages/dd-trace/test/setup/mocha-hooks.js:53:30)
    at process.processImmediate (node:internal/timers:483:21)
    at process.callbackTrampoline (node:internal/async_hooks:130:17)
↳ and 5 more — View all
DataDog/apm-reliability/dd-trace-js | validate_supported_configurations_v2_local_file — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

OpenFeature | OpenFeature / unit (node-active) — 🔧 Needs a code fix, caused by this PR

View more details · View in GitHub Actions

8 tests failed due to TypeError: Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED') in FlaggingProvider.

View all 6 failed jobs.

📋 Copy fix prompt
CI on my pull request is failing. Help me find and fix the root cause of each failing job below — they were flagged as caused by changes in this PR, so focus on the diff. For each job, explain the failure and propose a fix.

Before you start, set up the Datadog software-delivery tooling so you can
query the CI data yourself:

1. Check whether you already have the Datadog software-delivery MCP tools
   (e.g. a `search_datadog_ci_pipeline_events` tool) and the `unblock-pr` skill.
2. If either is missing, STOP and ask me for permission before installing
   anything. Do not install or run anything until I have said yes.
3. Only with my explicit approval, set up the Datadog software-delivery MCP
   server and skills by following:
     https://docs.datadoghq.com/getting_started/software_delivery_mcp_tools/
   then restart so the skill is picked up.
4. If I decline, skip all of the above and work from the context below alone.

Then run /unblock-pr — it will pull the CI data itself. The job context below is what we already know.

If /unblock-pr is not available — because I declined the setup above, or it did not install — work from the context below instead.

Datadog has already classified this failure as caused by changes in this PR.
Take that as given and work the fix:

1. Locate the change. Diff this branch against its base and find the change
   that produces this error. Explain the mechanism, don't just name a file:
     git fetch origin && git diff $(git merge-base origin/leo.romanovsky/ffl-2446-evp-flagevaluation-nodejs HEAD)...HEAD
2. Reproduce it locally. Run the failing job's command or test before
   proposing anything.
3. Propose the smallest fix that addresses the root cause — not a workaround,
   not a broadened assertion, not a disabled or skipped test.
4. Re-run the same command to confirm, and say exactly what you ran.
5. If the failure turns out to be intermittent rather than deterministic, say
   so plainly instead of "fixing" it — that is a flaky test, and patching it
   hides the problem.

If the right move is to re-run the job rather than change code, use the job
link in the context below. For GitHub Actions: `gh run rerun <run-id> --failed`,
where the run ID is the number after `/runs/` in that URL (not the trailing
number, which is the job ID).

Branch: leo.romanovsky/ffe-agentless-evp-node-hardening

OpenFeature | OpenFeature / unit (node-active)
Commit: bf1ce08d30a54c4123c6a05a7d1c8e77ea04c0b5
Error (code / test):
8 tests failed due to TypeError: Cannot read properties of undefined (reading 'DD_FEATURE_FLAGS_EVALUATION_COUNTS_ENABLED') in FlaggingProvider.
CI job: https://github.com/DataDog/dd-trace-js/actions/runs/33089516358/job/98578232209

Plus 2 more failing jobs not shown here.

ℹ️ Info

No other issues found (see more)

❄️ No new flaky tests detected

🎯 Code Coverage (details)
• Patch Coverage: 88.46%
• Overall Coverage: 95.34% (-0.94%)

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 921adcc | Docs | View more details | Give us feedback!

@pr-commenter

pr-commenter Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-08-27 16:06:51

Comparing candidate commit 921adcc in PR branch leo.romanovsky/ffe-agentless-evp-node-hardening with baseline commit b79691b in branch leo.romanovsky/ffl-2446-evp-flagevaluation-nodejs.

📊 Benchmarking dashboard

Found 0 performance improvements and 0 performance regressions! Performance is the same for 2298 metrics, 12 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:dogstatsd-with-tags-20

  • unstable cpu_user_time [-457.877ms; +215.517ms] or [-9.360%; +4.406%]
  • unstable execution_time [-460.562ms; +216.117ms] or [-9.268%; +4.349%]
  • unstable throughput [-76675.685op/s; +165342.099op/s] or [-4.542%; +9.794%]

scenario:plugin-cassandra-driver-long-query-20

  • unstable cpu_user_time [-550.283ms; +871.793ms] or [-14.943%; +23.674%]
  • unstable execution_time [-548.934ms; +870.570ms] or [-14.884%; +23.604%]
  • unstable throughput [-1001863.993op/s; +633850.444op/s] or [-12.014%; +7.601%]

scenario:plugin-graphql-long-with-depth-and-collapse-off-20

  • unstable max_rss_usage [-25.895MB; +18.981MB] or [-6.532%; +4.787%]

scenario:plugin-graphql-long-with-depth-off-20

  • unstable max_rss_usage [-5.141MB; +7.682MB] or [-4.024%; +6.013%]

scenario:plugin-graphql-long-with-depth-off-26

  • unstable max_rss_usage [-2.385MB; +50.230MB] or [-1.220%; +25.702%]

scenario:plugin-mongodb-core-plain-find-26

  • unstable execution_time [-123.845ms; +180.003ms] or [-5.978%; +8.688%]
  • unstable throughput [-267891.834op/s; +189380.180op/s] or [-6.706%; +4.741%]

scenario:test-optimization-large-suite-20

  • unstable max_rss_usage [-6.182MB; +2.669MB] or [-7.636%; +3.297%]

@codecov

codecov Bot commented Aug 26, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 88.67925% with 6 lines in your changes missing coverage. Please review.
✅ Project coverage is 95.34%. Comparing base (b79691b) to head (921adcc).

Files with missing lines Patch % Lines
packages/dd-trace/src/openfeature/writers/util.js 88.88% 4 Missing ⚠️
packages/dd-trace/src/openfeature/writers/base.js 0.00% 2 Missing ⚠️
Additional details and impacted files
@@                                  Coverage Diff                                  @@
##           leo.romanovsky/ffl-2446-evp-flagevaluation-nodejs    #9988      +/-   ##
=====================================================================================
- Coverage                                              96.28%   95.34%   -0.94%     
=====================================================================================
  Files                                                    989      987       -2     
  Lines                                                 149672   149491     -181     
  Branches                                               11161    11174      +13     
=====================================================================================
- Hits                                                  144106   142533    -1573     
- Misses                                                  5566     6958    +1392     
Flag Coverage Δ
aiguard 57.20% <100.00%> (-5.05%) ⬇️
aiguard-integration 54.97% <100.00%> (-3.95%) ⬇️
apm-bucket-0 57.07% <100.00%> (-4.61%) ⬇️
apm-bucket-1 62.27% <100.00%> (-4.96%) ⬇️
apm-bucket-2 61.09% <100.00%> (-5.69%) ⬇️
apm-bucket-3 58.70% <100.00%> (-4.75%) ⬇️
apm-capabilities-tracing 62.01% <15.09%> (-0.10%) ⬇️
apm-integrations-aerospike 54.76% <100.00%> (-4.79%) ⬇️
apm-integrations-confluentinc-kafka-javascript 60.07% <100.00%> (-5.66%) ⬇️
apm-integrations-couchbase 55.67% <100.00%> (-4.39%) ⬇️
apm-integrations-http 60.67% <100.00%> (-4.82%) ⬇️
apm-integrations-next 58.36% <100.00%> (-4.65%) ⬇️
apm-integrations-tedious 56.08% <100.00%> (-4.35%) ⬇️
appsec-_express_fastify 72.17% <100.00%> (-5.19%) ⬇️
appsec-graphql_kafka_ldapjs 67.04% <100.00%> (-4.51%) ⬇️
appsec-lodash_mongodb-core_mongoose 63.72% <100.00%> (-4.60%) ⬇️
appsec-mysql_node-serialize_passport 65.32% <100.00%> (-4.47%) ⬇️
appsec-next 55.70% <100.00%> (-0.83%) ⬇️
appsec-postgres_sourcing_stripe 65.45% <100.00%> (?)
appsec-sourcing_stripe_template ?
appsec-template 59.53% <100.00%> (?)
debugger 63.18% <100.00%> (-5.76%) ⬇️
instrumentations-bucket-0 50.75% <100.00%> (-4.13%) ⬇️
instrumentations-bucket-1 58.59% <100.00%> (-4.75%) ⬇️
instrumentations-bucket-10 59.69% <100.00%> (-5.20%) ⬇️
instrumentations-bucket-11 60.34% <100.00%> (-0.13%) ⬇️
instrumentations-bucket-12 50.67% <100.00%> (-4.60%) ⬇️
instrumentations-bucket-13 51.53% <100.00%> (-3.15%) ⬇️
instrumentations-bucket-14 50.77% <100.00%> (?)
instrumentations-bucket-2 52.02% <100.00%> (-3.97%) ⬇️
instrumentations-bucket-3 52.65% <100.00%> (-4.15%) ⬇️
instrumentations-bucket-4 57.69% <100.00%> (-5.11%) ⬇️
instrumentations-bucket-5 48.41% <100.00%> (-0.79%) ⬇️
instrumentations-bucket-6 59.19% <100.00%> (-5.46%) ⬇️
instrumentations-bucket-7 50.96% <100.00%> (-11.35%) ⬇️
instrumentations-bucket-8 57.41% <100.00%> (-6.30%) ⬇️
instrumentations-bucket-9 56.24% <100.00%> (+4.05%) ⬆️
instrumentations-instrumentation-couchbase 49.50% <100.00%> (-3.97%) ⬇️
instrumentations-instrumentation-zlib ?
instrumentations-integration-esbuild 34.01% <100.00%> (-0.07%) ⬇️
llmobs-ai_anthropic_bedrock 61.91% <100.00%> (-4.21%) ⬇️
llmobs-bucket-1 60.45% <100.00%> (-3.80%) ⬇️
llmobs-openai ?
llmobs-openai-agents_vertex-ai ?
llmobs-openai_openai-agents_vertex-ai 62.74% <100.00%> (?)
llmobs-sdk 68.22% <100.00%> (-7.10%) ⬇️
master-coverage ?
openfeature 55.26% <88.67%> (-4.40%) ⬇️
platform-core_esbuild_instrumentations-misc 40.66% <100.00%> (+0.14%) ⬆️
platform-integration 59.40% <100.00%> (-4.82%) ⬇️
platform-shimmer_unit-guardrails_webpack 38.38% <100.00%> (+0.11%) ⬆️
plugins-browser-bunyan_bullmq_cassandra 60.49% <100.00%> (-4.96%) ⬇️
plugins-bucket-0 55.95% <100.00%> (-4.20%) ⬇️
plugins-bucket-1 53.03% <100.00%> (-4.39%) ⬇️
plugins-bucket-11 60.39% <100.00%> (-4.96%) ⬇️
plugins-bucket-14 57.74% <100.00%> (-5.05%) ⬇️
plugins-bucket-17 60.19% <100.00%> (-5.71%) ⬇️
plugins-bucket-18 57.13% <100.00%> (-4.50%) ⬇️
plugins-bucket-19 60.32% <100.00%> (-5.36%) ⬇️
plugins-bucket-20 60.64% <100.00%> (-5.29%) ⬇️
plugins-bucket-4 55.31% <100.00%> (-4.97%) ⬇️
plugins-cookie_cookie-parser_crypto 50.28% <100.00%> (-4.15%) ⬇️
plugins-fastify_fetch_fs 59.50% <100.00%> (-5.12%) ⬇️
plugins-generic-pool_google-cloud-pubsub_grpc 63.04% <100.00%> (-5.17%) ⬇️
plugins-handlebars_hapi_hono 57.53% <100.00%> (-5.15%) ⬇️
plugins-ioredis_langgraph_ldapjs 55.92% <100.00%> (-4.78%) ⬇️
plugins-light-my-request_limitd-client_lodash 57.32% <100.00%> (-5.22%) ⬇️
plugins-mariadb_memcached_mercurius 61.86% <100.00%> (-4.50%) ⬇️
plugins-mongodb-core_mongoose_multer 57.87% <100.00%> (-4.99%) ⬇️
plugins-mysql_mysql2_nats 60.13% <100.00%> (-5.51%) ⬇️
plugins-pino_postgres_process 57.58% <100.00%> (-5.07%) ⬇️
plugins-pug_redis_router 60.46% <100.00%> (-5.47%) ⬇️
plugins-url_valkey_vm 56.07% <100.00%> (-4.93%) ⬇️
plugins-winston_ws 58.54% <100.00%> (-5.29%) ⬇️
profiling 60.63% <100.00%> (-4.93%) ⬇️
serverless-aws-sdk-aws-sdk 53.96% <100.00%> (-0.71%) ⬇️
serverless-aws-sdk-base-inject-field 49.99% <100.00%> (-4.03%) ⬇️
serverless-aws-sdk-bedrockruntime 53.72% <100.00%> (-3.61%) ⬇️
serverless-aws-sdk-client 55.21% <100.00%> (-3.93%) ⬇️
serverless-aws-sdk-dynamodb 54.53% <100.00%> (-3.73%) ⬇️
serverless-aws-sdk-eventbridge 56.05% <100.00%> (-0.85%) ⬇️
serverless-aws-sdk-kinesis 58.06% <100.00%> (-4.26%) ⬇️
serverless-aws-sdk-lambda 56.25% <100.00%> (-3.99%) ⬇️
serverless-aws-sdk-s3 54.63% <100.00%> (-3.70%) ⬇️
serverless-aws-sdk-serverless-peer-service 58.60% <100.00%> (-4.29%) ⬇️
serverless-aws-sdk-sns 58.84% <100.00%> (-4.38%) ⬇️
serverless-aws-sdk-sqs 59.25% <100.00%> (-4.43%) ⬇️
serverless-aws-sdk-stepfunctions 54.46% <100.00%> (-3.68%) ⬇️
serverless-aws-sdk-util 50.49% <100.00%> (-4.19%) ⬇️
serverless-bucket-0 52.63% <100.00%> (-4.27%) ⬇️
serverless-bucket-1 ?
serverless-lambda 54.70% <100.00%> (?)
test-optimization-cucumber 63.63% <100.00%> (-0.03%) ⬇️
test-optimization-cypress 60.72% <100.00%> (-3.43%) ⬇️
test-optimization-jest 70.35% <100.00%> (-1.43%) ⬇️
test-optimization-mocha 64.67% <100.00%> (-0.13%) ⬇️
test-optimization-playwright-playwright-atr 59.31% <100.00%> (-0.08%) ⬇️
test-optimization-playwright-playwright-efd 59.24% <100.00%> (-0.82%) ⬇️
test-optimization-playwright-playwright-final-status 59.58% <100.00%> (-0.08%) ⬇️
test-optimization-playwright-playwright-impacted-tests 59.34% <100.00%> (-0.50%) ⬇️
test-optimization-playwright-playwright-reporting ?
test-optimization-playwright-playwright-test-span 59.31% <100.00%> (-0.13%) ⬇️
test-optimization-selenium 58.48% <100.00%> (-0.07%) ⬇️
test-optimization-vitest 67.04% <100.00%> (-4.15%) ⬇️
test-optimization-vitest-browser 58.32% <100.00%> (-0.07%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Keep direct EVP fallback independent from API key fingerprint generation and validation.
… leo.romanovsky/ffe-agentless-evp-node-hardening
@leoromanovsky
leoromanovsky changed the base branch from master to leo.romanovsky/ffl-2446-evp-flagevaluation-nodejs August 26, 2026 21:31
@linear-code

linear-code Bot commented Aug 26, 2026

Copy link
Copy Markdown

FFL-1482

FFL-1483

Carry the reviewed parent changes into the direct EVP hardening stack.

Environment: Datadog workspace
Environment: Datadog workspace
Environment: Datadog workspace
@leoromanovsky

Copy link
Copy Markdown
Contributor Author

Replaced by #10026 for the exposure fallback hardening against master. Flag-evaluation delivery will be handled separately.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants