Skip to content

Service review requested: bio_policy refusals during public literature applicability review #52843

Description

@joe-hilling-ps

What issue are you seeing?

Requesting service review of two Codex subagent turns that terminated with codex_error_info: "bio_policy" while performing read-only applicability/coverage critique of public primary literature. We cannot determine the actual triggering input from the returned diagnostics. This is a possible unintended refusal report, not a claim that the safety decision is wrong and not a request to disable safeguards.

The reported error was:

This content was flagged for possible biological risk.

The task concerned fluorescent-reporter calibration and nutrient-dependent transcription-factor localization: whether supplied experimental results and methods actually supported the research question. It did not request experimental execution or modification of an organism. A later narrowed request excluded unrelated material, but it was made in the same reviewer thread, so earlier history may still have been present. The trigger remains unknown.

What steps can reproduce the bug?

We are reporting the saved failed invocations rather than attempting a new reproduction. No further rejected-content submission, model/provider/account switch, or filter workaround has been attempted.

  • Reviewer thread: 01a1207c-767f-7592-9640-2a766bbaffeb.
  • First failing turn: 01a1215a-50a9-7983-bc5c-a69118ed25cf; started 2026-10-09 15:48:49 UTC, completed with the refusal at 15:59:08 UTC. Local delegated message: 2,040 UTF-8 bytes; SHA-256 33791b3e2506d6c708e89f9c85e93380634d323396c86beb09d20dc488c5a025.
  • Second failing turn: 01a12165-f313-70c3-8ea8-a94f53dc7298; started 2026-10-09 16:01:31 UTC, completed with the refusal at 16:01:42 UTC. Local delegated message: 1,976 UTF-8 bytes; SHA-256 85ff7814ef294b3ed0beb6369ae189f16df3f9fcc93233f41258dddf538130f6.

Those message sizes are not total serialized inference input sizes; the reviewer had prior history. No HTTP request/correlation ID or triggering passage/rule was returned. Local task-completion envelopes contained an error, not a successful review.

What is the expected behavior?

Please investigate these saved service outcomes, confirm the appropriate supported review/appeal channel, and provide sufficient diagnostic attribution to distinguish a policy refusal from other failures. We are not requesting a new model answer or automatic clearance. We will preserve both failed attempts and await legitimate review before resuming the affected scientific review.

Additional information

  • Observed runtime model/effort: gpt-6.1-sol, high; read-only biology-evidence reviewer task assigned through collaboration.followup_task to a previously existing subagent (agent type default with the explicit model override).
  • Running app-server version at the reviewer session and present diagnostic read: 0.162.0; separately installed shell CLI version: 0.160.0. This version difference is recorded, not asserted as the cause.
  • Environment: Linux x86_64 development workspace, Codex desktop-associated local app-server.
  • No connected paid reader call was involved in these refusals.
  • This report includes only sanitised error/configuration/task metadata. No full history, logs, source excerpts, private repository content, credentials, patient data or attachments are included or authorised for automatic upload. If more material is needed, please specify the minimum required input and a private supported channel first.

Activity

  1. added
    bugSomething isn't working
    appIssues related to the Codex desktop app
    model-behaviorIssues related to behaviors exhibited by the model
    subagentIssues involving subagents or multi-agent features
    on Oct 10, 2026
  2. joe-hilling-ps commented on Oct 10, 2026

    @joe-hilling-ps
    Author

    A recurrence was observed today,10October2026 at approximately18:35UTC (19:35London), during a separate freshly authorised public-literature reference-review task.

    A narrowly delegated Codex agent, advertised as GPT‑6.1Sol/high, failed with: “This content was flagged for possible biological risk.” The task was to verify public-source rights/answerability and prepare an independent reading rubric. Some eligible public reference acquisition had completed, but no completed rubric/ready result was captured. No experimental OpenAIResponses/token-count call or candidate-search request had been dispatched in this new child. This was a hosted Codex agent-turn refusal, not a captured OpenAI API429 or transport timeout.

    The interface supplied no exact request ID, offending span or diagnosed trigger. We preserved the emitted notice, partial artifact identities and unchanged experiment accounting, halted dependent execution, and have not retried, rephrased/resubmitted the rejected task, switched models/providers/accounts, or uploaded paper bodies/payloads here. This report is a sanitised operational follow-up to the existing issue, not a reproduction attempt or claimed false-positive diagnosis.

    Could the supported service investigation identify the request/trigger category or advise the appropriate supported diagnostic route? The existing issue remains open with no service clarification captured as of18:36UTC. Acknowledgement of this report will not be treated as clearance to resume refused material.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    LinuxappIssues related to the Codex desktop appbugSomething isn't workingmodel-behaviorIssues related to behaviors exhibited by the modelsubagentIssues involving subagents or multi-agent features

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions