Skip to content
This repository was archived by the owner on Oct 1, 2026. It is now read-only.
This repository was archived by the owner on Oct 1, 2026. It is now read-only.

Preserve task input semantics while redacting diagnostic output #93

Description

@lukemaj

Outcome

Authorized task text reaches the selected model without semantic changes introduced by diagnostic redaction, while public/logging surfaces continue to protect sensitive data. This advances Model Router's Objective by making delegated instructions reliable across harnesses.

Evidence and Elon decision

Source audit of main7e64edf found that core._durable_run redacts prompt and direction-block metadata before persistence, then the OpenCode supervisor sends that stored prompt as execution input. Patterns resembling password assignments or authorization examples can therefore change the model's instructions. Codex command arguments follow a different storage path. This is a source-confirmed behavior, not a demonstrated secret exposure or a measured production incident.

Question the coupling between execution input and sanitized diagnostics. Separate their purposes at the smallest existing boundary. Do not remove secret protection globally, create another transcript store, or rewrite the harness seam.

Acceptance criteria

  • Synthetic task/code examples containing redaction-shaped text reach the execution harness with their authorized semantics intact.
  • Public status, reports and exported diagnostics remain sanitized and do not leak sensitive fixture values.
  • Private state, file permissions and existing supported harness contracts remain explicit and consistent. Do not change retention or access policy silently.
  • Direction-block content and recorded identity/hash describe what the model actually receives.
  • Existing secret-protection regressions remain valid; amend tests that intentionally encoded prompt mutation only after establishing the correct input/output boundary.

Non-goals

Broad security redesign, credential migration, new logging/storage systems, changing subscriptions, or deployment/install. No live credentials or raw private transcripts in tests or GitHub artifacts.

Blocked by

None for a synthetic reproduction and bounded design. Not dispatched as part of the current two implementation jobs; prioritize after the demonstrated recovery failures.

Required proof

Use synthetic strings through the actual prompt-to-supervisor seam and public serialization path. Assert exact model input and independently sanitized output, unchanged private permissions, supported harness compatibility and full project proof. Independent exact-candidate review required.

Related architecture audit and recovery work: #87.

Activity

  1. lukemaj commented on Sep 25, 2026

    @lukemaj
    ContributorAuthor

    Obsolete: the path that redacted the prompt before sending it (core._durable_run into the OpenCode supervisor) was deleted with the legacy harness (#110, PR #119). On main the task text is stored as given (_canonical_task) and sent unchanged to the dispatcher and workers through T3; redaction applies only to logs, events and reports.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions