Skip to content

feat(skills): request-scoped secrets for skills (closes #3861) - #3871

Merged
WillemJiang merged 12 commits into
bytedance:mainfrom
fancyboi999:feat/3861-skill-request-scoped-secrets
Jul 2, 2026
Merged

WillemJiang merged 12 commits into
bytedance:mainfrom
fancyboi999:feat/3861-skill-request-scoped-secrets

Conversation

@fancyboi999

Copy link
Copy Markdown
Collaborator

What & why

Closes #3861. Lets a caller pass per-request, short-lived end-user credentials (e.g. an ERP token, a per-user cloud API key) to a skill's sandbox scripts without the value entering the prompt, tool arguments, the executed command string, traces, the checkpoint, or the run record — the capability the issue asks for, and the skill-side counterpart of the still-open #3322 (MCP per-user credentials).

The design follows LangGraph's "out-of-band runtime dependency injection" pattern (the same runtime.context DeerFlow already uses) and Anthropic Agent Skills' "credentials live in the execution environment, not the prompt" model.

How it works

caller run request
  └─ config.context.secrets = { "ERP_TOKEN": "<short-lived>" }   ← out-of-band, never a message
       └─ build_run_config → runtime.context["secrets"]  (not mirrored into configurable)
            └─ /skill-name activates → SkillActivationMiddleware resolves the skill's
               declared `required-secrets` against context.secrets → per-run injection set
                 └─ bash tool → Sandbox.execute_command(command, env=…)
                      ├─ LocalSandbox: subprocess env (platform creds scrubbed first)
                      └─ AioSandbox:  bash.exec(env=…) on a fresh session
                           └─ skill script reads $ERP_TOKEN  (value never in any message)
  • Declare — a skill lists what it needs in SKILL.md frontmatter: required-secrets: (string list or {name, optional}). name is both the context.secrets key and the env var exposed to the script.
  • Carry — the caller sends values in the run request's context.secrets (out-of-band). Reuses the existing context passthrough; never mirrored into configurable.
  • Bind — only the slash-activated skill's declared secrets are injected, scoped to that run.
  • Inject — Sandbox.execute_command gains an env parameter (Local + AIO). env=None is fully backward-compatible.
  • Scrub — execute_command no longer leaks the Gateway's os.environ to skill subprocesses: secret-looking names (*KEY*/*SECRET*/*TOKEN*/*PASSWORD*/*CREDENTIAL*/*DSN* + a connection-string denylist like DATABASE_URL/REDIS_URL/GH_PAT) are dropped. A skill can only receive what the caller explicitly supplies; a caller-supplied value wins over a scrubbed host value of the same name (the per-user-key-overrides-shared-key case).

Leak surfaces — all sealed (verified)

Surface How
prompt value travels via runtime.context, never a message
trace tracing/metadata.py never copies context; redact_secret_context_keys available defensively
checkpoint secrets live on runtime.context, not graph state
audit the activation journal records secret names only
bash stdout mask_secret_values redacts injected values from tool output before it re-enters context
run record + run API start_run persists redact_config_secrets(body.config), so runs.kwargs_json and RunResponse.kwargs never carry the secret
inherited env platform creds scrubbed from the sandbox subprocess

Verification (evidence, not assertion)

  • Unit/integration tests — new tests/test_skill_request_scoped_secrets.py (parser, carrier, activation binding, env scrub, the leak surfaces, a real-subprocess end-to-end) + updated tests/test_gateway_services.py and tests/test_local_sandbox_encoding.py.
  • Full backend suite (CI-equivalent): green, all regression tests pass; blocking-IO gate passes.
  • Real gateway end-to-end (real LLM + skill activation + bash): a request with context.secrets ran a skill whose script used the injected secret; the value reached the sandbox subprocess and appeared in none of the surfaces above (verified by scanning the run response, the run API, the trace events, and the full per-thread data dir).
  • Real third-party cloud call: a skill called an external cloud API using a request-scoped API key supplied via context.secrets. The call succeeded (HTTP 200); a same-named key that happened to exist in the Gateway env was scrubbed and never used; the caller's key leaked nowhere. (Credential redacted from this report.)

Notes for reviewers

  • This PR also fixes a pre-existing latent bug surfaced during end-to-end testing: InputSanitizationMiddleware wraps user input in --- BEGIN USER INPUT --- markers, but only preserved the original text for upload / IM-channel messages — so slash-skill activation silently failed for plain text messages. The sanitizer now preserves the original text in ORIGINAL_USER_CONTENT_KEY (additive; sanitization behaviour unchanged), which get_original_user_content_text already expected. Slash activation now works for all messages.
  • build_run_config now sets configurable.thread_id on the context path too — the checkpointer requires it, and the prior behaviour produced a KeyError: 'thread_id' for context-driven runs.

Scope / non-goals

Not verified

  • No live Langfuse/LangSmith backend was exercised; trace-cleanliness was verified against DeerFlow's metadata builder and the in-graph callback payloads, not a live tracing UI.
  • The AIO/Docker bash.exec(env=…) path is covered by unit tests against the SDK signature/return shape; it was not run against a live Docker sandbox (local sandbox was used for the live end-to-end).

… skills

Add an env parameter to Sandbox.execute_command (abstract + local + AIO) so request-scoped secrets can be injected into skill subprocesses, and scrub platform credentials (*KEY*/*SECRET*/*TOKEN*/*PASSWORD*/*CREDENTIAL*) from the inherited environment by default so scoped injection is not security theatre. LocalSandbox always passes an explicit scrubbed env; AioSandbox routes env-bearing commands through bash.exec(env=) on a fresh session and leaves the legacy persistent-shell path unchanged. Part of bytedance#3861. BEHAVIOR CHANGE: execute_command no longer inherits the full os.environ; Windows encoding tests updated to assert the scrubbed dict.
Add SecretRequirement and Skill.required_secrets, and parse the required-secrets SKILL.md frontmatter field (a string list or {name, optional} mappings), dropping malformed entries with a warning so one bad declaration does not invalidate the skill. The declared name is both the context.secrets key and the env var injected at activation. Part of bytedance#3861.
Add SECRETS_CONTEXT_KEY + extract_request_secrets, centralising the context.secrets carrier contract. The existing context passthrough (build_run_config -> _build_runtime_context) already carries the sub-key to runtime.context without mirroring it into configurable; characterization tests lock that behaviour. Part of bytedance#3861.
Binding point A: when a skill is slash-activated, SkillActivationMiddleware resolves its declared required-secrets against the request's context.secrets and writes the per-run injection set to runtime.context. The bash tool forwards that set to execute_command(env=). A skill cannot harvest a host platform credential (is_host_platform_secret guard, cf. GHSA-rhgp-j443-p4rf), and injected values are redacted from bash output (mask_secret_values) so an echoed secret never re-enters the prompt/trace. Part of bytedance#3861.
…n helper

Regression tests assert the secret value is absent from all five surfaces: prompt (activation message), checkpoint (graph state vs context separation), audit (journal records names only), trace (metadata builder never copies context; never mirrored to configurable), and stdout (mask_secret_values). Add redact_secret_context_keys as a defensive helper for any context serialization. Part of bytedance#3861.
Add Request-Scoped Secrets subsection (Skills) + env policy note (Sandbox) and the execute_command(env=) signature change, per the doc-sync policy. Part of bytedance#3861.
…coped secrets

Real-gateway e2e + independent review of bytedance#3861 surfaced three defects, now fixed:

1. Slash activation never fired in the live chain. InputSanitizationMiddleware
   wraps user input in BEGIN/END markers before SkillActivationMiddleware sees it,
   and the original text was only preserved when an upload or IM channel set it.
   For a plain text message the slash command became undetectable, so no secret
   was ever resolved. Fix: the sanitizer now setdefaults the pre-wrap text into
   ORIGINAL_USER_CONTENT_KEY (additive; sanitization behaviour unchanged), so
   slash activation works for all messages. Pre-existing latent bug surfaced here.

2. The raw request config (with context.secrets) was persisted to runs.kwargs_json
   and echoed by the run API (RunResponse.kwargs). Fix: redact_config_secrets()
   strips secret-bearing context keys from the persisted/echoed copy in start_run;
   the live config that drives the run keeps them. build_run_config now also sets
   configurable.thread_id on the context path (the checkpointer requires it).

3. Connection-string credentials (DATABASE_URL, REDIS_URL, SENTRY_DSN, GH_PAT, ...)
   were not scrubbed from the inherited sandbox env. Fix: env_policy adds a *DSN*
   pattern plus an explicit connection-string denylist (no blanket *URL* — benign
   service URLs stay readable).

Verified end-to-end via a real gateway run (real LLM + skill activation + bash):
the secret reaches the sandbox subprocess and appears in NONE of prompt, trace,
checkpoint, audit, stdout, runs.kwargs_json, or the run API. Part of bytedance#3861.
…itizer interaction

Sync the Request-Scoped Secrets section with the verification-driven fixes: inherited-env scrub (incl. connection-string denylist), run-record/run-API redaction as the 6th sealed leak surface, and the sanitizer preserving original content so slash activation fires. Part of bytedance#3861.
…ndant host-name guard

A real-world demo (a skill calling a third-party cloud API with a request-scoped
key) exposed that the is_host_platform_secret guard was both wrong and harmful:
it refused to inject a caller-supplied secret whenever a same-named variable
existed in the Gateway env — which is exactly the bytedance#3861 use case (a per-user key
overriding a shared platform key). The guard was also redundant: build_sandbox_env
already scrubs secret-looking names from the inherited env before injection, so a
skill can never read a host credential — it only ever receives the caller's value.

Remove the guard; the injected (caller) value simply wins over the scrubbed host
value. Verified end-to-end: the agent called the real cloud API successfully with
the caller's key, the host's same-named key was scrubbed and never used, and the
caller's key leaked to none of the surfaces. Part of bytedance#3861.
@github-actions github-actions Bot added area:agents Agents, subagents, graph wiring, prompts, langgraph.json area:backend Gateway / runtime / core backend under backend/ area:docs Documentation and Markdown only area:sandbox Sandboxed execution and docker/ area:skills Skills under skills/ or the skills harness needs-validation Touches front/back contract surface; needs real-path validation risk:high High risk: backend API, agents, sandbox, auth, deps, CI size/XL PR changes 700+ lines labels Jun 29, 2026
@yong326

yong326 commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

@fancyboi999 Code conflicts, please help resolve

…ill-request-scoped-secrets

# Conflicts:
#	backend/packages/harness/deerflow/sandbox/local/local_sandbox.py
#	backend/packages/harness/deerflow/sandbox/tools.py

@willem-bd willem-bd left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review findings from an automated pass. Focus: correctness bugs surfaced by the new env kwarg path, the secrets-carrier lifecycle, and output redaction. See inline comments.

Comment thread backend/packages/harness/deerflow/sandbox/tools.py
Comment thread backend/packages/harness/deerflow/community/aio_sandbox/aio_sandbox.py Outdated
Comment thread backend/packages/harness/deerflow/sandbox/tools.py
"""
with self._lock:
try:
result = self._client.bash.exec(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bash.exec is called with no session_id, so consecutive env-bearing bash calls in the same skill activation share no cwd / env exports / venv state.

A typical skill run issues sequential bash calls like:

cd /mnt/user-data/workspace && source .venv/bin/activate && pip install foo
python run.py

Without secrets, both go through shell.exec_command on the persistent session and the second call sees the first call's cwd/PATH. With secrets injected, both go through _execute_with_env and each gets a fresh auto-created session: the second call starts in $HOME, sees no venv, no foo, and fails with No such file or command not found.

I understand the design tension (a reused session_id would keep secrets scoped-to-thread rather than scoped-to-call). Two options worth considering:

  1. Use a stable per-thread session_id for env-bearing calls and rely on the fact that bash.exec's env overrides the session env per command (verify with the SDK — if that guarantee holds, secrets don't persist across calls).
  2. If session reuse isn't safe, at least pass exec_dir (e.g. the thread's workspace) so relative paths behave the same way as on the legacy path.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intentionally keeping the fresh-session choice and documenting why on _execute_with_env.

The core tension is real, but a shared session_id for env-bearing calls would reintroduce the leak surface this PR exists to close: if bash.exec's env merges into the session env rather than being strictly per-command (the SDK does not contractually guarantee per-command isolation), a request-scoped secret would ride the session into later commands in the same skill — exactly the scoped-to-call invariant the design preserves. Without that SDK guarantee, reusing a session_id is a regression risk I'm not willing to take in a security-focused change.

The fresh-session model matches LocalSandbox, where each execute_command is already a fresh subprocess (cwd is recovered via the command-prefix; sourced venv / exports never cross calls either). So AIO-with-secrets is now consistent with the local backend, not an outlier.

Practical mitigation for the cwd/venv case, now stated in the docstring: skills that need setup must fold it into a single command (e.g. cd /mnt/user-data/workspace && source .venv/bin/activate && python run.py). That's the same pattern LocalSandbox skills already need.

If a later SDK version documents per-command env isolation on a reused session, exec_dir (for cwd) or a stable session_id become safe upgrades — happy to revisit then.

@WillemJiang WillemJiang added reviewing A maintainer is reviewing this PR question Further information is requested labels Jul 2, 2026
@WillemJiang WillemJiang added this to the 2.1.0 milestone Jul 2, 2026
@starhobby

Copy link
Copy Markdown

When will 2.1.0 plan to launch?

@WillemJiang

Copy link
Copy Markdown
Collaborator

When will 2.1.0 plan to launch?

We will release it this month.

Review fixes from PR bytedance#3871:

- E2BSandbox.execute_command now accepts env/timeout and routes them to
  commands.run(envs=, timeout=). The bash tool passes env= unconditionally,
  so the prior signature (command only) raised TypeError on every e2b bash
  call and broke e2b deployments entirely. env=None stays backward-compatible.
- SkillActivationMiddleware clears the active-secret set before resolving each
  activation, so a later skill in the same run never inherits an earlier
  skill's injection set (the bytedance#3861 contract: a skill only receives what the
  caller supplied AND that skill declared).
- AioSandbox env path uses a dedicated _DEFAULT_HARD_TIMEOUT — bash.exec exposes
  no idle/no-change timeout, so the prior reuse of the legacy idle constant
  conflated wall-clock vs idle semantics. The env path also retries on the
  ErrorObservation signature now, sharing the legacy persistent-shell recovery
  contract.
- mask_secret_values skips values below a minimum length floor so a short
  declared secret (e.g. "42") cannot shred unrelated bytes (exit codes,
  timestamps, sizes) of tool output. The secret is still injected into the
  subprocess; only the output mask skips it.

session_id reuse on the env path is intentionally NOT added: a shared session
could let request-scoped secrets ride the session env into later commands,
which the SDK does not contractually forbid. The fresh-session choice matches
the LocalSandbox model (each call is a fresh subprocess); the trade-off
(consecutive env-bearing calls do not share cwd/venv/exports) is documented on
_execute_with_env.
@fancyboi999

Copy link
Copy Markdown
Collaborator Author

Review fixes — push 926c73e5

Addressed the review from @willem-bd. CI is green across the board (backend-unit-tests, backend-blocking-io, lint-backend/frontend, Layer 1/2 golden).

Resolved (5)

  • E2B bash breakage (tools.py) — E2BSandbox.execute_command now accepts env/timeout and routes them to commands.run(envs=, timeout=). The bash tool passes env= unconditionally, so the prior command-only signature raised TypeError on every e2b bash call. env=None is fully backward-compatible (no envs kwarg sent). Regression tests added.
  • Stale active secrets (skill_activation_middleware.py) — _apply_skill_secrets now clears the active-secret set before each activation, so a later skill in the same run can never inherit an earlier skill's injection set. Two regression tests cover both reported paths (next skill declares nothing; caller omits the required value).
  • hard_timeout vs no_change_timeout (aio_sandbox.py) — dedicated _DEFAULT_HARD_TIMEOUT for the env path with a docstring noting bash.exec exposes no idle/no-change timeout. Kept at 600 to stay aligned with the legacy idle budget; named independently so the two call sites evolve without silently altering each other.
  • Short-value mask corruption (tools.py) — mask_secret_values skips values below _MIN_MASK_LENGTH (8 chars) so a short declared secret (region code, numeric id) cannot shred exit codes / timestamps / sizes in tool output. The secret is still injected into the subprocess; only the output mask skips short values.
  • ErrorObservation retry (aio_sandbox.py) — the env path now detects _ERROR_OBSERVATION_SIGNATURE and retries on another fresh session, sharing the legacy persistent-shell recovery contract. Refactored into a _run_bash_exec helper so the retry stays on the env-bearing path (secrets never enter a shell session).

Left open (1) — intentional

  • session_id reuse on the env path (aio_sandbox.py) — I kept the fresh-session choice. A shared session_id could let request-scoped secrets ride the session env into later commands, and the SDK does not contractually guarantee per-command env isolation, so reuse is a leak-surface regression I'm not willing to take in a security-focused change. Fresh session matches the LocalSandbox model (each execute_command is already a fresh subprocess there). The trade-off (consecutive env-bearing calls do not share cwd/venv/exports) is documented on _execute_with_env; skills fold setup into a single command. Happy to revisit if a future SDK version documents per-command env isolation on a reused session. Full rationale is in the thread.

@WillemJiang
WillemJiang merged commit 09988ca into bytedance:main Jul 2, 2026
28 of 32 checks passed
jtaynl added a commit to jtaynl/deer-flow that referenced this pull request Jul 3, 2026
Conflict (1): community/aio_sandbox/aio_sandbox.py — resolved TAKE-UPSTREAM.
Upstream bytedance#3871 reworked execute_command to (self, command, env=None, timeout=None),
which SUPERSEDES our local timeout= shim (added 2026-06-30 for bytedance#3864) AND adds the
new env= for request-scoped skill secrets. Dropping our shim — our fork divergence
shrinks by one patch. Normal bash path (env=None) is byte-identical (legacy
shell.exec_command); the timeout= call from tools.py no longer TypeErrors.

The other 12 auto-merged clean. Notable: bytedance#3904 unified DB config (no-op for us —
our checkpointer.type=postgres section takes precedence), bytedance#3903 recursion clamp
(ceiling 1000 = our value → no-op), bytedance#3872 csrf cookie lifetime fix (positive),
#4fcb subagent step history (persists to existing run_events table, no migration),
bytedance#3871 request-scoped secrets. No new alembic migration → head stays 0002.
i18n auto-merged, WRI branding preserved. Login/setup pages untouched.
jtaynl added a commit to jtaynl/deer-flow that referenced this pull request Jul 3, 2026
yogyoho added a commit to yogyoho/eai-flow-main that referenced this pull request Jul 3, 2026
…edance#3665)

09988ca bytedance#3861 skills request-scoped secrets(secret_context 注入沙箱环境变量)
5a699e2 bytedance#3665 guardrail GuardrailRequest 暴露 user_id/role/run_id 上下文
e2b_sandbox DU 取 eai 删; AGENTS.md 取 eai 版; sandbox/services 冲突 HEAD 空+上游新增
marvin9551 pushed a commit to marvin9551/deer-flow that referenced this pull request Aug 21, 2026
… (bytedance#3871)

* feat(sandbox): per-call env injection + platform-secret scrubbing for skills

Add an env parameter to Sandbox.execute_command (abstract + local + AIO) so request-scoped secrets can be injected into skill subprocesses, and scrub platform credentials (*KEY*/*SECRET*/*TOKEN*/*PASSWORD*/*CREDENTIAL*) from the inherited environment by default so scoped injection is not security theatre. LocalSandbox always passes an explicit scrubbed env; AioSandbox routes env-bearing commands through bash.exec(env=) on a fresh session and leaves the legacy persistent-shell path unchanged. Part of bytedance#3861. BEHAVIOR CHANGE: execute_command no longer inherits the full os.environ; Windows encoding tests updated to assert the scrubbed dict.

* feat(skills): parse required-secrets frontmatter declaration

Add SecretRequirement and Skill.required_secrets, and parse the required-secrets SKILL.md frontmatter field (a string list or {name, optional} mappings), dropping malformed entries with a warning so one bad declaration does not invalidate the skill. The declared name is both the context.secrets key and the env var injected at activation. Part of bytedance#3861.

* feat(runtime): request-scoped secret carrier (context.secrets)

Add SECRETS_CONTEXT_KEY + extract_request_secrets, centralising the context.secrets carrier contract. The existing context passthrough (build_run_config -> _build_runtime_context) already carries the sub-key to runtime.context without mirroring it into configurable; characterization tests lock that behaviour. Part of bytedance#3861.

* feat(skills): inject declared secrets at slash-activation into bash env

Binding point A: when a skill is slash-activated, SkillActivationMiddleware resolves its declared required-secrets against the request's context.secrets and writes the per-run injection set to runtime.context. The bash tool forwards that set to execute_command(env=). A skill cannot harvest a host platform credential (is_host_platform_secret guard, cf. GHSA-rhgp-j443-p4rf), and injected values are redacted from bash output (mask_secret_values) so an echoed secret never re-enters the prompt/trace. Part of bytedance#3861.

* test(skills): lock the five secret leak surfaces + add trace redaction helper

Regression tests assert the secret value is absent from all five surfaces: prompt (activation message), checkpoint (graph state vs context separation), audit (journal records names only), trace (metadata builder never copies context; never mirrored to configurable), and stdout (mask_secret_values). Add redact_secret_context_keys as a defensive helper for any context serialization. Part of bytedance#3861.

* docs(backend): document request-scoped secrets for skills

Add Request-Scoped Secrets subsection (Skills) + env policy note (Sandbox) and the execute_command(env=) signature change, per the doc-sync policy. Part of bytedance#3861.

* fix(skills): close gaps found by end-to-end verification of request-scoped secrets

Real-gateway e2e + independent review of bytedance#3861 surfaced three defects, now fixed:

1. Slash activation never fired in the live chain. InputSanitizationMiddleware
   wraps user input in BEGIN/END markers before SkillActivationMiddleware sees it,
   and the original text was only preserved when an upload or IM channel set it.
   For a plain text message the slash command became undetectable, so no secret
   was ever resolved. Fix: the sanitizer now setdefaults the pre-wrap text into
   ORIGINAL_USER_CONTENT_KEY (additive; sanitization behaviour unchanged), so
   slash activation works for all messages. Pre-existing latent bug surfaced here.

2. The raw request config (with context.secrets) was persisted to runs.kwargs_json
   and echoed by the run API (RunResponse.kwargs). Fix: redact_config_secrets()
   strips secret-bearing context keys from the persisted/echoed copy in start_run;
   the live config that drives the run keeps them. build_run_config now also sets
   configurable.thread_id on the context path (the checkpointer requires it).

3. Connection-string credentials (DATABASE_URL, REDIS_URL, SENTRY_DSN, GH_PAT, ...)
   were not scrubbed from the inherited sandbox env. Fix: env_policy adds a *DSN*
   pattern plus an explicit connection-string denylist (no blanket *URL* — benign
   service URLs stay readable).

Verified end-to-end via a real gateway run (real LLM + skill activation + bash):
the secret reaches the sandbox subprocess and appears in NONE of prompt, trace,
checkpoint, audit, stdout, runs.kwargs_json, or the run API. Part of bytedance#3861.

* docs(backend): document the env scrub, persistence redaction, and sanitizer interaction

Sync the Request-Scoped Secrets section with the verification-driven fixes: inherited-env scrub (incl. connection-string denylist), run-record/run-API redaction as the 6th sealed leak surface, and the sanitizer preserving original content so slash activation fires. Part of bytedance#3861.

* fix(skills): inject caller secret over scrubbed host value; drop redundant host-name guard

A real-world demo (a skill calling a third-party cloud API with a request-scoped
key) exposed that the is_host_platform_secret guard was both wrong and harmful:
it refused to inject a caller-supplied secret whenever a same-named variable
existed in the Gateway env — which is exactly the bytedance#3861 use case (a per-user key
overriding a shared platform key). The guard was also redundant: build_sandbox_env
already scrubs secret-looking names from the inherited env before injection, so a
skill can never read a host credential — it only ever receives the caller's value.

Remove the guard; the injected (caller) value simply wins over the scrubbed host
value. Verified end-to-end: the agent called the real cloud API successfully with
the caller's key, the host's same-named key was scrubbed and never used, and the
caller's key leaked to none of the surfaces. Part of bytedance#3861.

* fix(skills): address review on request-scoped secrets (bytedance#3861)

Review fixes from PR bytedance#3871:

- E2BSandbox.execute_command now accepts env/timeout and routes them to
  commands.run(envs=, timeout=). The bash tool passes env= unconditionally,
  so the prior signature (command only) raised TypeError on every e2b bash
  call and broke e2b deployments entirely. env=None stays backward-compatible.
- SkillActivationMiddleware clears the active-secret set before resolving each
  activation, so a later skill in the same run never inherits an earlier
  skill's injection set (the bytedance#3861 contract: a skill only receives what the
  caller supplied AND that skill declared).
- AioSandbox env path uses a dedicated _DEFAULT_HARD_TIMEOUT — bash.exec exposes
  no idle/no-change timeout, so the prior reuse of the legacy idle constant
  conflated wall-clock vs idle semantics. The env path also retries on the
  ErrorObservation signature now, sharing the legacy persistent-shell recovery
  contract.
- mask_secret_values skips values below a minimum length floor so a short
  declared secret (e.g. "42") cannot shred unrelated bytes (exit codes,
  timestamps, sizes) of tool output. The secret is still injected into the
  subprocess; only the output mask skips it.

session_id reuse on the env path is intentionally NOT added: a shared session
could let request-scoped secrets ride the session env into later commands,
which the SDK does not contractually forbid. The fresh-session choice matches
the LocalSandbox model (each call is a fresh subprocess); the trade-off
(consecutive env-bearing calls do not share cwd/venv/exports) is documented on
_execute_with_env.
jihtsan pushed a commit to jihtsan/dnx-deer-flow that referenced this pull request Aug 29, 2026
… (bytedance#3871)

* feat(sandbox): per-call env injection + platform-secret scrubbing for skills

Add an env parameter to Sandbox.execute_command (abstract + local + AIO) so request-scoped secrets can be injected into skill subprocesses, and scrub platform credentials (*KEY*/*SECRET*/*TOKEN*/*PASSWORD*/*CREDENTIAL*) from the inherited environment by default so scoped injection is not security theatre. LocalSandbox always passes an explicit scrubbed env; AioSandbox routes env-bearing commands through bash.exec(env=) on a fresh session and leaves the legacy persistent-shell path unchanged. Part of bytedance#3861. BEHAVIOR CHANGE: execute_command no longer inherits the full os.environ; Windows encoding tests updated to assert the scrubbed dict.

* feat(skills): parse required-secrets frontmatter declaration

Add SecretRequirement and Skill.required_secrets, and parse the required-secrets SKILL.md frontmatter field (a string list or {name, optional} mappings), dropping malformed entries with a warning so one bad declaration does not invalidate the skill. The declared name is both the context.secrets key and the env var injected at activation. Part of bytedance#3861.

* feat(runtime): request-scoped secret carrier (context.secrets)

Add SECRETS_CONTEXT_KEY + extract_request_secrets, centralising the context.secrets carrier contract. The existing context passthrough (build_run_config -> _build_runtime_context) already carries the sub-key to runtime.context without mirroring it into configurable; characterization tests lock that behaviour. Part of bytedance#3861.

* feat(skills): inject declared secrets at slash-activation into bash env

Binding point A: when a skill is slash-activated, SkillActivationMiddleware resolves its declared required-secrets against the request's context.secrets and writes the per-run injection set to runtime.context. The bash tool forwards that set to execute_command(env=). A skill cannot harvest a host platform credential (is_host_platform_secret guard, cf. GHSA-rhgp-j443-p4rf), and injected values are redacted from bash output (mask_secret_values) so an echoed secret never re-enters the prompt/trace. Part of bytedance#3861.

* test(skills): lock the five secret leak surfaces + add trace redaction helper

Regression tests assert the secret value is absent from all five surfaces: prompt (activation message), checkpoint (graph state vs context separation), audit (journal records names only), trace (metadata builder never copies context; never mirrored to configurable), and stdout (mask_secret_values). Add redact_secret_context_keys as a defensive helper for any context serialization. Part of bytedance#3861.

* docs(backend): document request-scoped secrets for skills

Add Request-Scoped Secrets subsection (Skills) + env policy note (Sandbox) and the execute_command(env=) signature change, per the doc-sync policy. Part of bytedance#3861.

* fix(skills): close gaps found by end-to-end verification of request-scoped secrets

Real-gateway e2e + independent review of bytedance#3861 surfaced three defects, now fixed:

1. Slash activation never fired in the live chain. InputSanitizationMiddleware
   wraps user input in BEGIN/END markers before SkillActivationMiddleware sees it,
   and the original text was only preserved when an upload or IM channel set it.
   For a plain text message the slash command became undetectable, so no secret
   was ever resolved. Fix: the sanitizer now setdefaults the pre-wrap text into
   ORIGINAL_USER_CONTENT_KEY (additive; sanitization behaviour unchanged), so
   slash activation works for all messages. Pre-existing latent bug surfaced here.

2. The raw request config (with context.secrets) was persisted to runs.kwargs_json
   and echoed by the run API (RunResponse.kwargs). Fix: redact_config_secrets()
   strips secret-bearing context keys from the persisted/echoed copy in start_run;
   the live config that drives the run keeps them. build_run_config now also sets
   configurable.thread_id on the context path (the checkpointer requires it).

3. Connection-string credentials (DATABASE_URL, REDIS_URL, SENTRY_DSN, GH_PAT, ...)
   were not scrubbed from the inherited sandbox env. Fix: env_policy adds a *DSN*
   pattern plus an explicit connection-string denylist (no blanket *URL* — benign
   service URLs stay readable).

Verified end-to-end via a real gateway run (real LLM + skill activation + bash):
the secret reaches the sandbox subprocess and appears in NONE of prompt, trace,
checkpoint, audit, stdout, runs.kwargs_json, or the run API. Part of bytedance#3861.

* docs(backend): document the env scrub, persistence redaction, and sanitizer interaction

Sync the Request-Scoped Secrets section with the verification-driven fixes: inherited-env scrub (incl. connection-string denylist), run-record/run-API redaction as the 6th sealed leak surface, and the sanitizer preserving original content so slash activation fires. Part of bytedance#3861.

* fix(skills): inject caller secret over scrubbed host value; drop redundant host-name guard

A real-world demo (a skill calling a third-party cloud API with a request-scoped
key) exposed that the is_host_platform_secret guard was both wrong and harmful:
it refused to inject a caller-supplied secret whenever a same-named variable
existed in the Gateway env — which is exactly the bytedance#3861 use case (a per-user key
overriding a shared platform key). The guard was also redundant: build_sandbox_env
already scrubs secret-looking names from the inherited env before injection, so a
skill can never read a host credential — it only ever receives the caller's value.

Remove the guard; the injected (caller) value simply wins over the scrubbed host
value. Verified end-to-end: the agent called the real cloud API successfully with
the caller's key, the host's same-named key was scrubbed and never used, and the
caller's key leaked to none of the surfaces. Part of bytedance#3861.

* fix(skills): address review on request-scoped secrets (bytedance#3861)

Review fixes from PR bytedance#3871:

- E2BSandbox.execute_command now accepts env/timeout and routes them to
  commands.run(envs=, timeout=). The bash tool passes env= unconditionally,
  so the prior signature (command only) raised TypeError on every e2b bash
  call and broke e2b deployments entirely. env=None stays backward-compatible.
- SkillActivationMiddleware clears the active-secret set before resolving each
  activation, so a later skill in the same run never inherits an earlier
  skill's injection set (the bytedance#3861 contract: a skill only receives what the
  caller supplied AND that skill declared).
- AioSandbox env path uses a dedicated _DEFAULT_HARD_TIMEOUT — bash.exec exposes
  no idle/no-change timeout, so the prior reuse of the legacy idle constant
  conflated wall-clock vs idle semantics. The env path also retries on the
  ErrorObservation signature now, sharing the legacy persistent-shell recovery
  contract.
- mask_secret_values skips values below a minimum length floor so a short
  declared secret (e.g. "42") cannot shred unrelated bytes (exit codes,
  timestamps, sizes) of tool output. The secret is still injected into the
  subprocess; only the output mask skips it.

session_id reuse on the env path is intentionally NOT added: a shared session
could let request-scoped secrets ride the session env into later commands,
which the SDK does not contractually forbid. The fresh-session choice matches
the LocalSandbox model (each call is a fresh subprocess); the trade-off
(consecutive env-bearing calls do not share cwd/venv/exports) is documented on
_execute_with_env.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:agents Agents, subagents, graph wiring, prompts, langgraph.json area:backend Gateway / runtime / core backend under backend/ area:docs Documentation and Markdown only area:sandbox Sandboxed execution and docker/ area:skills Skills under skills/ or the skills harness needs-validation Touches front/back contract surface; needs real-path validation question Further information is requested reviewing A maintainer is reviewing this PR risk:high High risk: backend API, agents, sandbox, auth, deps, CI size/XL PR changes 700+ lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

支持调用方将请求级上下文(如用户身份、动态凭据)传递给 skill

5 participants