Skip to content

RPD-24H: gate-ticket lifecycle + case_retrieval wiring + pgvector/case acceptance - #233

Open
nghqqa wants to merge 73 commits into
mainfrom
feat/rpd-24h-delivery
Open

nghqqa wants to merge 73 commits into
mainfrom
feat/rpd-24h-delivery

Conversation

@nghqqa

@nghqqa nghqqa commented Sep 23, 2026

Copy link
Copy Markdown
Owner

RPD-24H 自主交付轮(2026-09-23)

BASE_HEAD: b79cd39cbcd78731ac50b2ff0f25b00a5a11514e(remote main @ 52f5e57)
最终 commit: 84ad9af(本 PR = feat/backend-pg-storage 全部产品化历史 + 本轮 RPD 交付)

本轮增量(RPD-24H)

  • RPD 状态机:RPD-24H.yaml(状态源)/ RPD-24H.md(视图)/ execution-state.json;启动门禁+基线记录(RPD-01 ✅)
  • pgvector 证据复核(RPD-02 ✅):一次性隔离实例重放 pgvector_smoke 11/11,实例已销毁;共享 case-pg 零接触
  • case_retrieval 接线(RPD-03 ✅):validate_env 新增 --preflight(真实连接+只读角色+表能力校验);正/负双向实测(exit 0 / 脱敏 exit 5);修复仓库根路径深度 bug;正式 controller 注入 = WAITING_HUMAN
  • TicketStore 闭环复核(RPD-04 ✅):gate_ticket_smoke 13/13 + 28 单测(幂等/CAS/TTL 24h/审计/桥只建票)
  • 后端回归(RPD-05 ✅):本地 282 passed / 0 failed(+29 skipped);PG 门控 25/26(唯一失败 = Windows spawn race,基线既有,TEST-DEBT 保持)
  • 前端契约对齐(RPD-06 ✅):gate_display 八态契约(pending/action_required/approved_plan_ready/rejected/blocked/expired/backend_unavailable/scope_missing)+ console 审批响应字段;无新增端点

声明

项 状态
GitHub 写入 仅本 feature branch push + 本 PR(无 main push、无 merge、无 check-run)
部署 无(controller/Worker 注入未执行,仅隔离验证)
真实模型审查 / 真实审批 / fixer-verifier 均无(D-1/D-2 未启用)
共享 case-pg / 生产 MinIO / AgentTeams 配置 均未修改
D-7 未批准;BM25 主线;嵌入仅 DeterministicFakeProvider(无模型下载)

完整失败清单

  • PG 门控:test_cross_process_race_single_winner 1 项(Windows spawn 队列,基线既有,stash 对照验证;TEST-DEBT 保持,不掩盖)
  • 其余本地套件 0 失败

未完成项 / 人工阻塞

  1. CASE2 人工门决策(HIGH finding 批准/拒绝)——等你
  2. D-1/D-2 真实审批启用
  3. controller/Worker 正式环境注入(共享环境授权)
  4. 真实 CASE2(新 head 空提交+预算门+单 check-run)

回滚

  • 本分支独立可弃:git push origin --delete feat/rpd-24h-delivery + 关闭 PR 即完全回滚
  • 运行副本回滚点:r3work/rollback-20260923-222605/(CASE2 实测版三文件)

Curated ops tooling from the finals sprint workspace: matrix helper lib,
SK5/WH kickoffs (PR1/2/3), gate decision delivery (with placeholder-SHA
guard), OTel activation + probe, link verification, cold-boot wedge
reproducer, WH evidence packaging (fail-loud ledger), timeline SVG
generator. See tools/r3ops/README.md for provenance.
- PROCESSED only after GitHub write-back succeeds (scenario 7): post_check
  split into payload builder + _reporter_exec + structured parse; only
  HTTP 200/201 with check_run_id counts as published; failure =>
  PUBLISH_FAILED(retryable) ERROR, never PROCESSED.
- Reconciliation convergence (scenario 3): reconcile-before-post GETs
  existing mergepilot/review check-runs for the head sha and adopts;
  MinIO check-run-receipt short-circuits recovery without re-posting.
- Bounded retry (scenario 4): PUBLISH_ATTEMPTS=3 with (10,30)s backoff.
- Duplicate-webhook guard (scenario 5): already_processed by
  repo+pr+head before claim; receiver GUID PK re-verified.
- Lease correctness (scenario 8): claim_id rotates per attempt (uuid
  suffix); finish matches exact claim_id, stale executor gets
  rowcount=0. Fixes LIKE-only finish and a %% escaping bug it exposed.
- Timeout semantics: review timeout => ERROR(TIMEOUT manual) even when
  the neutral check was published.
- Import path: repo-relative tools/r3ops first, r3work fallback.
- tests/gh_bridge/: 16 unit tests (fully mocked) covering scenarios
  3/4/5/7/8 publish semantics; caught two real bugs during red-green.
- tests/gh_app: repoint 8 stale root Dockerfile.* paths to docker/
  (repo reorg leftover) - suite back to 831 passed.
- docs/productization/: STATUS/ACCEPTANCE/DECISIONS seeded (M1 matrix,
  state-semantics mapping, no-exactly-once statement, pending-auth
  integration list).
- take_over_stale(): CAS lease takeover of expired RUNNING bridge rows
  (claim_id LIKE '%-bridge-%' namespace only -> cannot touch future
  Controller claims); fresh claim_id per takeover, stale executor gets
  rowcount=0 (scenarios 1/8/9).
- resume(): recovery routed by authoritative MinIO project state -
  receipt => finalize; terminal => resume publish; in-flight =>
  watch-only (never re-send kickoff, scenario 2); missing project =>
  bounded requeue RQn (REQUEUE_MAX=2) then MANUAL.
- main() takes over stale leases at startup; process()/resume() share
  conclude() so scenario-7 semantics cannot diverge.
- docs/productization/ORCHESTRATION-CONTRACT.md: state ownership,
  claim_id namespaces, transition rules (exact-claim CAS, no LIKE
  writes), recovery semantics both orchestrators must honor (scen 9).
- tests/gh_bridge/test_recovery.py: 7 tests; suite 23 passed.
- ACCEPTANCE/STATUS updated: scenarios 1/2/4/9 to unit-verified.
…yer)

- docs/productization/M2-APPROVAL-SPEC.md: v2 memo 2.2a four questions
  answered; V0 action set {generate_patch,run_poc,publish_result}; legacy
  merge USED->MERGED semantics stripped (not migrated); undecided product
  policy listed as D-1/D-2/D-3 (no default activation).
- tools/approval/: immutable five-field Binding (run/repo/head_sha/
  params_hash/patch-or-finding fp), CAS state machine (PENDING sole
  branch state; approve/reject first-wins; idempotent replay NOOP),
  invalidate_for_new_head, check_execution per-field binding gate before
  any side effect, expiry gates grant boundary only, single-use USED.
- tests/approval/: 34 passed.
- ORCHESTRATION-CONTRACT.md section 6: verified factual boundary - bridge
  claims leave lease_expires_at NULL so github_drain's stale-takeover
  predicate can never grab bridge in-flight rows; cutover 3-step program
  recorded (drain-then-switch) or rows become bridge-only orphans.
- Records correction: gh_app on clean tree 8b30fb1 measures 821 collected
  (816 passed + 5 skipped); prior '831 passed' not reproducible byte-wise
  (tests/gh_app unchanged since b46e8ba) - STATUS/ACCEPTANCE amended.
…losed)

- build_manifest(): honest dispatch-time snapshot - code repo/head/base
  SHA, kickoff+SPEC content sha256, bridge source sha256 + git commit,
  worker image ids (docker inspect at dispatch), non-secret effective
  config canonical + sha256. Model identifiers / skill content hashes /
  RAG version are NOT visible to the bridge -> recorded null AND listed
  in missing[] (never fabricated).
- write_run_manifest(): write-once per run - absent -> write; same hash
  -> adopted; different -> refuse (running upgrade cannot mutate an
  existing run's manifest).
- process(): manifest persisted between worker wake and kickoff send;
  failure marks delivery ERROR and blocks dispatch (kickoff references
  the manifest sha, never dispatched unmanifested). dry-run unaffected.
- resume(): read-only manifest lookup + log; absence (pre-manifest runs)
  logged, never blocks recovery and never rewrites.
- tests/gh_bridge/test_run_manifest.py: 13 tests (suite 36 passed);
  ProcessTerminalSemantics patched on the new prepare_run_manifest seam
  (prevents real docker/MinIO writes from unit tests).
- tools/costmeter/core.py: single-run BudgetGuard - reserve-before-call
  (fail-closed deny), retry reuses original reservation (no double
  hold), commit settles multi-refund/over-charge (denies when actual
  exceeds both reservation AND remaining), missing usage consumes the
  estimate and records an explicit gap, crash recovery via write-through
  ledger + stale-reservation reclaim, in-process lock proven by 20-way
  concurrency test (exactly 10 admitted at limit 1000).
- tools/costmeter/collect.py: SCAFFOLD local collector - span-name
  counts only; token usage is NOT available locally (OTel LLM spans are
  exported to an external collector) and this is recorded as such, never
  fabricated. Currency conversion only with an explicit price table;
  unknown models reported as unknown_models.
- No default budget limit (undecided amounts must not become de-facto
  spending authorization); not wired into any live call path.
- tests/costmeter/: 16 tests (86 total green across new suites).
- docs/productization/INTEGRATION-AUTH-REQUESTS.md: centralized R1-R6
  authorization requests (bridge fault injection on shared PG with
  test-tagged rows, real GitHub writeback/reconcile, dual-head scenario
  6, run-copy sync with backup+rollback, worker-side version probing,
  token usage source) each with target/ops/impact/evidence/rollback.
- STATUS/ACCEPTANCE updated to scaffold-level honesty for cost meter.
- Worker model primary read from /root/.copaw-worker/{role}/openclaw.json
  (only the model field is extracted, never the rest of the config).
- Per-skill content hashes computed read-only inside the reviewer image
  (sha256 of sorted per-file sha256s); tool-name->directory mapping
  recorded (_SKILL_DIRS). Empty directory or missing mapping yields no
  hash -> item goes to missing[] instead of a bogus empty-input hash
  (the e3b0c442... trap, caught by read-only smoke).
- Worker container names corrected to elemiso-worker-{role}; all four
  image digests now resolve.
- Read-only smoke (2026-09-22, no dispatch, no MinIO write): manifest
  carries real model id, 4 real skill hashes, 4 image digests; missing[]
  reduced to generation_params (not exposed by agentloop) and rag.version
  (rag-live not running). Fabrication-free by construction.
- tests: 37 bridge (14 manifest) / 34 approval / 16 costmeter = 87 green.
- DECISIONS #9 records probe sources and the empty-hash lesson.
- M2-GATE-STORAGE-OPTIONS.md: Option B (MinIO ticket objects + single
  writer arbitration, zero server schema change) recommended for V0;
  Option A (server PG approvals rework) deferred to the Controller
  cutover authorization; Option C (local case-pg) rejected on domain
  boundaries. Both viable options share the tools/approval semantics
  layer, so the storage swap never touches the 34 unit-tested
  transitions.
- DECISIONS P-1: recorded as a proposal awaiting user decision - the
  distinction between approved decisions and pending proposals is kept
  explicit per reporting protocol.
- STATUS next-actions trimmed to user-input-blocked items.
…ed gate

Audit (RAG-AUDIT.md, evidence-based): the rag_retrieve chain IS wired in
current worker images (MCP server + config + hook, last activation
2026-09-21), with real historical usage (64 rag.retrieve calls, 58 OK)
and real consumption (result.md cites org-standards source_refs).
Currently the rag-live service is NOT running (netstat verified), and
the chain had no version binding, no repo fact source, and silent
degradation. tools/rag/rag_retrieval_service.py is an M6-era module NOT
on the live path - documented to prevent confusion.

Fixes (run copy untouched):
- tools/rag/corpus/: fact-source copy of the knowledge corpus (snapshot
  fd34c304... verified byte-identical to the running copy via
  content-addressed sha256) + tools/rag/live/rag-live-server.mjs fact
  source.
- corpus_tool.py: validate / idempotent import (atomic replace, no
  unbounded duplicates) / canonical snapshot_id - knowledge updates are
  new snapshots, old runs keep their manifest ID (RAG-4).
- bridge: manifest.rag now records snapshot_id/chunks/data_mode/
  retrieval_mode/strategy_id + service_state_at_dispatch + policy;
  MERGEPILOT_RAG_REQUIRED=1 makes dispatch fail-closed (corpus unreadable
  or /health unreachable -> delivery ERROR, no silent degradation);
  default stays advisory per current product semantics, degradation now
  visible in the manifest (RAG-6).
- tests/rag_live/: layer-2 isolated integration against the REAL node
  server (random port, temp corpus): known-doc hit with citation fields,
  legitimate-empty != failure, audit JSONL (query_hash+refs, no query
  text), snapshot consistency, injection corpus returned as data with
  shape intact (RAG-7 tool layer), unreachable detection, new-id-on-
  change; corpus_tool unit tests (validate/idempotent import/atomic
  update). Plus 9 bridge gate tests. Suite total 121 passed.
- Layer 3/4 (replay with real agent, real-model e2e) remain unverified -
  requires R1/R2/R4 authorization; not claimed.
…ntract

P-1 decided as a reversible engineering choice (DECISIONS P-1): the
MinIO single-writer proposal is rejected - mc has no conditional-write
primitive, so a convention cannot prove concurrent correctness, and
concurrent approve/reject from a web gate is a real V0 scenario.
tools/approval/store_sqlite.py instead provides real cross-process CAS:
BEGIN IMMEDIATE write-lock queueing + guarded UPDATE (rowcount check),
partial UNIQUE INDEX enforcing one active ticket per
(run,action,finding) across processes, WAL+FULL crash recovery. The
state machine is NOT reimplemented - the store loads rows into the
existing approval.transition pure logic (34 unit-tested semantics) and
writes back under a status guard, so semantics cannot drift.

Tests (7): cross-connection approve-vs-reject race (first wins, loser
gets INVALID_TRANSITION), racing duplicate creates converge to one
ticket via the unique index + IntegrityError fallback, crash recovery
by reopening (state preserved, transitions continue), red-line
execution check after a store round-trip.

Budget retry semantics (DECISIONS #11): reserve(retry_of) dedupes the
HOLD only; a genuinely re-issued model call still costs - commit must
receive cumulative actual usage. Contract fixed by
test_retry_reissued_call_commits_cumulative_cost; docstrings updated.

Suite: 125 passed (bridge 46 + rag_live 20 + approval 41 + costmeter 17
+ slack), all green. Records: STATUS blocker reclassification (P-1
closed as engineering choice; D-1/2/3 remain true user decisions),
GATE-STORAGE-OPTIONS marked decided, R7 corpus-deploy auth request
added with backup/rollback.
- tools/approval/gate_cli.py: ticket operational surface (new/approve/
  reject/show/check) over the SQLite store - the same store and state
  machine a future gate web page will consume. Check is the red-line
  execution precheck; approve injects real UTC now so expiry applies;
  un-decided product policy (D-1) is surfaced in output rather than
  assumed. No real execution path is wired.
- tests/approval/test_gate_cli.py (6): issue->approve->check happy
  path, binding mismatch rejected, expired approve refused, duplicate
  approve NOOP, identity required.
- RAG-AUDIT B-chain correction: skill_* audit records never populate
  document_count (template artifact, true even for skill_diff_parse),
  so the earlier zero-hit reading is withdrawn; case library was seeded
  2h before the 8 calls - actual hit/consumption unknown, evidence
  collection added to the real-case round (R1/R2).
- Suite 125 passed (bridge 46 + rag_live 20 + approval 47 + costmeter
  17, minus shared helpers); STATUS/ACCEPTANCE updated.
…nly)

- tools/approval/store.py: TicketStore Protocol fixing the five-operation
  contract (idempotent create / get / active_for / CAS transition /
  close); any implementation must reuse approval.transition semantics -
  stores may not reimplement the state machine.
- SQLiteTicketStore (renamed from SqliteTicketStore, alias kept):
  documented single-instance boundary - one Controller, one deployment,
  no multi-instance HA, NOT a multi-user SaaS store.
- PostgreSQLTicketStore: explicit migration placeholder (NotImplementedError
  with DSN shape); partial-unique-index and rowcount-CAS design recorded
  for the multi-Controller/shared-deployment future; M2 semantics and
  store-contract tests reuse unchanged.
- Conformance tests: store satisfies the runtime-checkable protocol;
  PG placeholder raises as designed (3 new, suite 50 passed).
… pipeline

Design: docs/productization/ARCHITECTURE-V3.md. The fixed serial
reviewer->fixer->verifier model is superseded by an 11-step pipeline
with risk grading, parallel reviewers, findings aggregation, INDEPENDENT
finding validation vs patch validation (two distinct stages, never
merged), explicit degradation, and per-dimension console visibility.
Feature flag MERGEPILOT_REVIEW_V3 defaults OFF; the legacy bridge chain
is untouched and not deleted - v3 is not wired to any runtime dispatch.

- risk.py: configurable Trivial/Lite/Full rules; any sensitive-path hit
  (auth/crypto/permission/credential/migration) forces FULL +
  human_review_required regardless of diff size; explainable reasons
  recorded for run-manifest; no LLM scheduling agent.
- stages.py: per-dimension statuses (8 states; FAILED/TIMEOUT retryable
  within budget only) orthogonal to 9 run-level outcomes - the aggregate
  can never mask dimension states; derive_outcome() rules: critical
  reviewer timeout/failure -> MANUAL_ATTENTION (never a pass),
  non-critical -> REVIEW_PARTIAL with coverage_missing, CANCELLED
  (PR-updated) -> MANUAL_ATTENTION.
- scheduler.py: DispatchPlanner is the single source of dispatch
  decisions (no second scheduling fact source); MAX_PR_CONCURRENCY=1
  frozen constant; reviewer concurrency 2 (serial when 1); worker runs
  one task at a time; needs_fix=True with gate disabled raises (D-1/D-2
  gate cannot be bypassed by plan construction); budget hook shared
  across all concurrent work.
- aggregate.py: deterministic dedupe (exact key + Jaccard near-dup),
  max severity, all reviewer sources preserved, auditable dropped count.
- verify_finding.py: VerifierInput structurally excludes reasoning
  fields (rejected at construction) - finding verifiers cannot read
  other agents' internal reasoning even accidentally.
- console_contract.py: fixed payload fields incl. per-reviewer status,
  coverage.complete, degradations, patch_validation, rag snapshot -
  partial completion stays visible, never one green dot.
- tests/orchestrator/test_v3.py: 41 tests covering the ARCHV3
  acceptance matrix (3-tier routing, sensitive escalation, barrier-proven
  parallelism, serial mode, dedupe+sources, verifier isolation,
  non-critical vs critical timeout, lifecycle separation, PR-update
  cancellation, retry budget exhaustion, budget hook blocking,
  gate-before-fixer, concurrency constant, console partial visibility).

Suites: 171 passed (orchestrator 41 + approval 50 + bridge 46 +
rag_live 17 + costmeter 17). Records: ACCEPTANCE-ARCHV3, DECISIONS #12,
STATUS round-4. Real-Agent integration stays LAST (R1/R2 authorization)
- local tests do not impersonate production verification.
…store + read-only console

Adapter (tools/orchestrator/adapter.py), three-state
MERGEPILOT_REVIEW_V3 (default off):
- off: bridge hook returns before any work - legacy path byte-identical
  (test: old conclusion PROCESSED unchanged, adapter zero-invoked).
- shadow: at the real dispatch boundary (process() after claim/dry-run),
  reads real PR metadata + diff via anonymous read-only GET, classifies
  risk, builds DispatchPlan, persists per-dimension stages, writes
  auditable evidence (content-addressed manifest_hash). NO agents, NO
  GitHub writes, NO legacy conclusion change (test: shadow explodes ->
  legacy still PROCESSED). Reviewer slots are explicit SKIPPED('agent
  not executed'); critical reviewer skipped => MANUAL_ATTENTION - a
  shadow run can never produce a 'completed/pass' conclusion. Diff
  unavailable => risk FAILED + unknown-tier degradation, never guessed.
- on: requires R1/R2; bridge logs and runs shadow-only until then.

RunStore (tools/orchestrator/runstore.py): v3_runs SQLite WAL table -
dumb persistence only (all transitions still via stages.RunStages, the
single state machine; separate from the approval TicketStore DB).
Idempotent UPSERT, restart recovery, superseded marking; PR update
cancels RUNNING/PENDING dims of the old run through the same state
machine. SQLite single-instance boundary per P-1; PG migration point at
Controller cutover.

Console (tools/console_v3/server.py): GET-only (POST/PUT/DELETE -> 405),
bound 127.0.0.1 by default; /api/runs, /api/runs/<id>, /healthz, page.
Every record carries its mode badge (shadow/fixture); partial
completion stays visible (coverage.complete=false + missing list);
no approve/reject/dispatch/write buttons exist. Handover:
docs/productization/console/STATUS.md.

Fixture vertical chain (13 tests): PR fixture -> head -> risk -> plan ->
parallel reviewer slots -> aggregate (dedupe keeps sources) -> finding
validation -> derive outcome -> store -> console read model; covers
Trivial/Full/sensitive-forced-Full, non-critical failure -> PARTIAL,
critical timeout -> MANUAL, budget hook blocking plan, verifier input
free of reasoning fields, PR-update supersession.

Tests: 197 passed total (bridge 49 + console 6 + orchestrator 58 +
approval 50 + rag_live 17 + costmeter 17). Records: ACCEPTANCE-M35,
DECISIONS #13, STATUS round-5, ARCHITECTURE-V3 wiring rows.
NOT claimed: real-Agent verification (mode=on real path), real RAG,
real GitHub - all pending R1/R2 authorization.
…kage

Review of b00a095 (nine-point checklist): items 1-8 pass (hook at real
dispatch boundary; off=zero behavior change proven; shadow read-only;
on cannot reach real agents pre-R1/R2; RunStore transitions only via
stages.RunStages; no second scheduler/state machine; PR update cancels
only unfinished dims; manifest/RAG/console read model consistent).
Item 9 found a real gap - FIXED:

- fail-soft observability: v3_shadow_hook exceptions previously went to
  stdout only, so a persistent adapter failure was invisible to
  monitoring. Hook errors now persist to a dedicated v3_hook_errors
  table (best-effort, still never blocks the legacy path) and are
  readable via GET /api/hook-errors. 3 new tests (record/list, bridge
  failure leaves persistent trace, endpoint read-only).
- supersede fidelity: _supersede_old_runs rebuilt StageRecords and
  dropped error/detail/timestamps - a superseded run lost its
  degradation reasons. Rebuild now preserves fields (test asserts a
  SUCCEEDED dim keeps its error through supersession).

Local smoke (real process, isolated port 4191, cleaned up after):
seeded one shadow + one fixture run; browser-verified /healthz,
/api/runs, /api/runs/<id> and the page with screenshot; shadow run
shows MANUAL_ATTENTION in red with three coverage gaps and per-stage
degradation reasons ('shadow: agent not executed'), fixture shows
REVIEW_COMPLETED - partial completion is visibly NOT success;
POST/PUT/DELETE all rejected 405.

Authorization decision package (AUTH-DECISION-PACKAGE.md): R1-R7
consolidated table with recommendation/target/ops/impact/evidence/
rollback/permissions/external-write-cost columns; D-1/D-2/D-3, budget
amount, R6 route and key rotation listed as separate user decisions.
NOTHING is approved on the user's behalf; no irreversible operation
was executed this round.

Suites: 200 passed (bridge 50 + console 7 + orchestrator 59 + approval
50 + rag_live 17 + costmeter 17).
Final decision table (AUTH-DECISION-PACKAGE.md restructured):
- A. user DECISIONS (product policy, mechanisms stay closed until then):
  D-1 approval action set, D-2 approver mapping, D-3 TTL, D-4 budget
  amount, D-5 R6 token-metering route, D-6 key rotation resumption;
- B. user AUTHORIZATIONS (external operations, executed only after
  explicit approval): R4 run-copy sync, R7 corpus sync + rag-live start,
  R1 fault-injection recovery, R2 real Agent/RAG case, R3 GitHub
  writeback matrix - each with recommendation/precise ops/impact/
  evidence/rollback/external-write-cost, plus the three-phase order and
  the enablement iron rules (no v3=on, no M1/M3/M4 claims before real
  evidence; failures keep shadow + record + minimal fix).
Nothing approved on the user's behalf.

Local readiness (no external side effects):
- tools/integration_prep/: declarative R1/R2/R3 step plans with
  per-step evidence points, R4/R7 sync plans (backup -> checksum ->
  sync -> off-smoke -> rollback), evidence collector with centralized
  redaction (gh-token/DSN/API key/OTel key, JSON-quote tolerant),
  authorization gate (dry-run by default; execution requires explicit
  approval + MERGEPILOT_IT_AUTH=1; any other value denied).
- rag-live fact source: /health exposes corpus_file_sha256 (raw-byte
  deployment check, complements canonical snapshot_id); search accepts
  optional run_id passthrough into the audit record (closes the
  run-association gap once R5 wires the MCP server; backward compatible
  - field absent when not passed).
- costmeter/hooks.py: dispatch budget factory - MERGEPILOT_RUN_BUDGET_TOKENS
  set => fail-closed hook (over-limit denial + crash-recovery ledger,
  exception class exposed on the hook for precise catching); unset =>
  None (current semantics: no check, no spending authorization).
  Integration points inventoried: bridge dispatch boundary, executor
  budget_hook, worker model-call surface (R5).
- Registered pre-existing issue: tests/m4c|m4e|skills module-basename
  collection conflict under whole-tree runs (legacy, untouched).

Suites: 217 passed (bridge 50 + console 7 + orchestrator 59 + approval
50 + rag_live 19 + costmeter 21 + prep 12). No push, no run-copy sync,
no GitHub writes, no shared-environment operations.
User approved R4+R7 with eight conditions; all followed, zero rollback
needed, zero R1/R2/R3 operations, no GitHub writes, no shared-env fault
injection, no model spend.

R4 (run copy sync):
- backups: gh_bridge.py / rag-live-corpus.json / rag-live-server.mjs
  (.bak-20260922-204703, sha256 recorded);
- synced 12 files (bridge + tools/orchestrator + rag/corpus_tool.py),
  all sha256-verified identical on both sides; minimal runtime surface
  (approval/costmeter/console_v3 are not bridge-runtime deps and stay
  repo-only);
- off smoke (MERGEPILOT_REVIEW_V3=off): status read-only exit 0 against
  server ledger; adapter loads from r3work/orchestrator; hook produces
  zero logs and zero RunStore side effects; shadow capability present
  and honestly degrades when diff unavailable (temp store only).
  Legacy behavior unchanged: bridge was not running; next start runs
  the hardened bridge with v3 defaulting off.

R7 (corpus sync + rag-live):
- idempotent import changed=false (run corpus already byte-identical,
  snapshot fd34c304..., chunks=12);
- rag-live started from the REPO fact-source path (r3work server file
  untouched): /health ok, chunks=12, corpus_file_sha256 matches the
  deployed corpus bytes;
- local retrieval verified: known query hits doc-cwe22-def#1 +
  doc-path-containment#2 with correct source_refs; run_id passthrough
  recorded in the audit line;
- service STOPPED after verification per condition 5; start command
  archived.

Evidence: docs/productization/R4-R7-EXECUTION-RECORD.md (checksum
tables, smoke transcripts, rollback commands - verified feasible, not
needed). Suites: 217 passed after execution. R1/R2/R3 remain
unexecuted and unauthorized.
…lysis

- tools/integration_prep/prerun_gate.py: nine pre-flight checks (git
  pin, run-copy sha, corpus snapshot, v3 mode, single orchestrator,
  containers, rag-live as required, budget hard-boundary MISSING_ACK
  FAILS the gate, delivery prereq head-unprocessed) - all injectable
  probes, zero IO in the gate function; 8 tests.
- Acceptance-language errata (five items): off smoke proves hook
  non-involvement, NOT end-to-end parity of the hardened bridge
  (publish/recovery/manifest/RAG-gate changes await the first real
  case); status is read-only; RAG run_id passthrough verified at
  server-HTTP level only (worker MCP does NOT pass run_id - R5 gap);
  217/225 passed is the seven-suite scope, not the whole repo;
  rollback commands prepared but never rehearsed.
- Trigger analysis (read-only ledger): 0 PENDING; every open PR head
  already PROCESSED (already_processed correctly skips replay) => the
  natural trigger for the first real case is one empty commit on an
  approved test branch. Proposed target: PR #9 (feat/skill-exercise,
  head f0dc76fbba030779fa28f647e8eb6b40b6dd1002, 2 files +29 lines =>
  TRIVIAL, single reviewer = smallest real envelope).
- Budget reality (DECISIONS #16): no in-process worker hard budget
  (agentloop is image-layer; change = R5). First-case bounds: single
  delivery + 20-min hard watch deadline + 3 publish retries (in code)
  + provider-side gateway-key spend cap as the only true hard limit +
  post-run OTel/audit accounting. Scaffold NOT marked as integrated.
- v3-on gap recorded: adapter has no real-Agent runner; the first real
  case validates the hardened legacy chain; v3 shadow runs as same-run
  comparison evidence only.

Suites: 225 passed. ONE consolidated confirmation prepared for the user
(see round report); no R1/R2/R3-style operations executed.
…nown, runbook

Per pre-execution review, five minimal fixes/verifications (no arch
expansion):

1. Legacy-chain roles documented (CASE1-RUNBOOK.md): this run is the
   hardened legacy chain - leader + reviewer make model calls; fixer/
   verifier wake but stay idle with the gate closed. The 'TRIVIAL =>
   single reviewer' inference was v3 logic and is retracted from STATUS;
   the legacy chain does not grade risk. Call scope = kickoff SPEC
   constraints (no repo modification, zero GitHub writes, own workspace).

2. Targeted claim: MERGEPILOT_TARGET_PR + MERGEPILOT_TARGET_HEAD (both
   required, SQL-level filter, malformed values rejected). Non-matching
   PENDING rows stay PENDING untouched; already_processed semantics
   untouched. Previously --once processed every PENDING row in the pass
   - a real gap for a controlled case, now closed with 7 tests
   (predicates in SQL, default-off, invalid rejection, half-set
   inactive, no claim side effects for non-target rows).

3. Budget/credentials (read-only probe): worker -> local gateway
   (elemiso-controller:8080, 64-char bearer) -> upstream provider. The
   local gateway key is NOT the billing credential - a cumulative spend
   hard cap can only be set in the provider control plane (user).
   Shared-credential impact and limits documented in RUNBOOK section 3.

4. Cancel rehearsal (local, reversible): docker stop/start of the
   reviewer worker round-trips cleanly. Bridge 20-min deadline stops
   orchestration only - the actual cut of in-flight agent calls is
   stopping the dedicated reviewer container; exclusivity conditions
   documented (stack dedicated during the case; single container only).

5. Publish unknown-vs-rejected: last attempt transport failure =>
   outcome=unknown => ledger PUBLISH_UNKNOWN(manual-reconcile), which
   stops automatic retry (a check-run may already exist; reconcile
   before any re-POST). Definitive rejects remain
   PUBLISH_FAILED(retryable). 2 new tests; conclude() branches added.

Two acceptance categories (A normal case / B RAG consumption) written
into the runbook: legal empty retrieval != failure, and one case does
not complete RAG effectiveness validation.

Suites: 232 passed. Confirmation single reduced to 4 items in
CASE1-RUNBOOK section 7 (trigger push + provider-side budget cap +
service/cancel permissions + credential status).
Branch feat/backend-pg-storage (baseline chore/backfill-r3-ops@c179601).
Backend window owns: tools/approval/store*, tools/integration_prep/,
tools/costmeter/hooks.py, tests/approval/, tests/integration_prep/,
docs/productization/backend/. ARCHITECTURE-V3 (design window) and
tools/console_v3 (platform window) untouched.

Execution-protection fixes (review round):
- targeted claim now requires repo + PR + full head SHA (three env
  vars, all-or-inactive, malformed rejected at startup); regression
  tests updated (SQL predicates, per-component inactivity);
- publish results classified five ways: confirmed / reconciled-adopted
  / unknown (transport failure - never blindly re-POSTed; ledger marks
  PUBLISH_UNKNOWN(manual-reconcile) and auto-retry stops) / permanent
  (403 non-rate-limit, 404, 422) / retryable (401 auth-refresh, 429,
  5xx, 403 rate-limit). Reporter scripts now emit structured HTTP
  errors. 7 classification tests.
- run-scope cancel design (pure logic, no container ops this round):
  stop bridge first (no re-delegation), then leader/reviewer/fixer/
  verifier, quiet verification, upstream in-flight billing caveat
  recorded, restore notes. Preconditions require exclusivity + no
  other active projects. 5 tests.

PostgreSQL prep (direction confirmed by user; formal schema belongs to
the design window - no competing schema defined):
- isolated PG container mp-pg-contract-test (postgres:16, 127.0.0.1:
  55432, own volume, zero connection to shared DBs);
- PG feasibility spike 3/3 on real PG 16: partial unique index
  enforces one active ticket per (run,action,finding) across
  connections; guarded-UPDATE CAS first-wins; terminal statuses free
  the unique slot - evidence for the design window, clearly marked as
  a spike;
- contract tests extracted to tests/approval/store_contract.py
  (7 implementation-agnostic checks); SQLite runs them today; the
  future PostgreSQLTicketStore reuses the same suite as acceptance;
- migration inventory tool + rehearsal on seeded tickets/v3_runs DBs
  (readonly PRAGMA dump + compatibility notes: untyped ISO timestamp
  columns -> timestamptz needs format validation);
- version consistency: repo bridge diverged from run copy (R4-era)
  after this round's fixes; incremental sync plan ready, execution
  requires new authorization.

Suites: 239 passed. Backend progress + design-window interface
questions: docs/productization/backend/PROGRESS.md.
…ed real PG

Design baseline: DATA-ARCHITECTURE-PG.md section 5.4 @ design branch
docs/architecture-audit-20260922 (e247b80). Column shape, status enum,
partial unique index and ticket_audit copied faithfully; NO competing
schema defined; RunStore/PG adapter explicitly deferred (run identity
contract not finalized).

- tools/approval/pg/migrations/001_approval_tickets.sql: schema
  approval + tickets + append-only ticket_audit + fixture-level parent
  tables (run.repos/run.runs, FK targets only). One documented
  deviation pending design ratification: uq_active_ticket uses NULLS
  NOT DISTINCT (PG15+) - the design sketch's plain partial index leaves
  finding_id=NULL active tickets unconstrained under PG's default
  NULL-distinct semantics.
- tools/approval/pg_store.py: PostgreSQLTicketStore implementing the
  existing TicketStore protocol - reuses approval.transition pure logic
  (zero state-machine rewrite), parameterized SQL, explicit
  transactions with rollback, SELECT FOR UPDATE + guarded UPDATE
  (judgeable CAS), ticket_audit written in the same transaction,
  StorageUnavailable distinct from business rejections, timezone-aware
  UTC everywhere (ISO inputs normalized). store.py re-exports; default
  runtime path unchanged.
- Real isolated PG verification (container mp-pg-contract-test, PG
  16.15, port 55432, own volume): contract suite 7/7 + PG-specific
  acceptance 9/9, including CROSS-PROCESS race (spawn processes, one
  winner), NULL-finding uniqueness, terminal-allowing-new-attempt,
  timezone/expiry boundary, rollback atomicity (revoked audit INSERT ->
  transition fails, state unchanged, no orphan audit), minimal runtime
  role (DML yes, DDL no).
- Migration tool tools/integration_prep/migrate_tickets_sqlite_pg.py:
  read-only source, dry-run default, validation (status/shapes/ISO
  times/duplicate active tickets), missing parents skipped not
  fabricated, conflicts never overwritten (rerun-safe), field-by-field
  post-import verification, transactional abort. Rehearsed on test
  copies against the isolated PG: 5/5.
- Implementation-time real defects fixed (recorded, not test
  relaxations): expiry not persisted on PG create; datetime/string
  comparison mismatch; _expired now ISO/naive-tolerant at the pure-
  logic source.

Suites: 247 passed + 24 skipped (PG tests gated by
MERGEPILOT_PG_CONTRACT=1, run separately: 16/16). Backend progress and
remaining design-window questions (v3_runs namespace, budget ledger
home, kb migration timing) in docs/productization/backend/PROGRESS.md.
No push, no run-copy sync, no real case, no shared-env changes.
…frontend handoff

- Reconcile adoption decision (decide_reconcile_adopt, pure fn):
  recorded check_run_id has priority; single match with app check is
  adoptable; multiple matches or record-mismatch => ambiguous, manual
  reconcile; reconcile transport failure => UNKNOWN and NO blind re-POST
  (publish flow restructured: reconcile-first each attempt, final
  bottoming reconcile after side-effect-possible attempts). Legacy
  'same repo+pr+head => same run' assumption retired per run-identity
  v2 (caf6909 section 3). Tests rewritten + added (ambiguous, app
  mismatch, transport-unknown, single-match adopt).
- Target_key acceptance tests: run-level '_run_', finding-level id,
  same-target no duplicate actives, different targets no cross-mutex,
  forged external target_key cannot bypass binding validation, empty
  string finding_id rejected at binding validation (not silently
  converted).
- PgRunStore minimal vertical (tools/orchestrator/pg_runstore.py +
  003_run_domain.sql): deterministic run_id (design 3.2 canonical JSON,
  vectors locked in tests), targets UNIQUE(repo,pr,head), request_key
  replay, exec_seq allocation under target row lock, active-run partial
  unique index, INSERT-only runs, supersede links preserving history,
  stages+events same-transaction (rollback proven via CHECK violation),
  expected-status guard (old executor cannot overwrite), console read
  model consistency, CROSS-PROCESS concurrent first-run convergence
  (spawn processes; exactly one creator). 10 tests on real isolated PG.
- Deviations recorded (003 file header + PROGRESS): delivery_id/
  first_delivery_id/knowledge_manifest_id without shared-table FKs in
  the isolated instance (unified migration will restore them).
- Frontend handoff: docs/productization/backend/FRONTEND-HANDOFF.md -
  design baseline, backend commits, available query interfaces,
  identity association, display red lines, pagination/error semantics,
  reproducible test data. No formal HTTP wiring to PG yet - stated.

Suites: 247 passed (7-suite) / 26 passed + gated PG suites on real
isolated PG. No push, no deploy, no shared migration, no real model
calls.
…d baseline)

- target_key sentinel: finding_id with '_' prefix rejected at binding
  validation (reserved namespace, cannot collide with '_run_').
- publish five-way classification: unknown/auth/retryable/permanent
  outcomes drive distinct ledger terminal notes (PUBLISH_UNKNOWN
  (manual-reconcile) / PUBLISH_AUTH(manual) / PUBLISH_REJECTED(permanent)
  / PUBLISH_FAILED(retryable)); reconcile transport failure => UNKNOWN,
  never blind re-POST.
- migration runner tools/approval/pg/apply_migrations.py:
  schema_migrations tracking, strict CREATE TABLE (no IF NOT EXISTS
  silent masking), unrecorded pre-existing tables abort with explicit
  error, multi-domain dir merge, --mark for historical versions.
- 001/003 restructured strict; fixture tables moved to test setup.
- tests: conclude outcomes, pagination truncation, migration runner
  rehearsal (fresh install / rerun noop / upgrade preserves records /
  wrong-shaped fails loudly / minimal role).
…ecution protection

Formal read-only HTTP service (tools/console_pg/server.py):
- Thin adapter over PgRunStore, no frontend-to-PostgreSQL direct access.
- Follows console API contract shapes: data_mode in every response
  (always 'fixture' for isolated test data), error = {error:{code,
  reason,message}}, repo addressing via query param, empty=null.
- Endpoints: /healthz, /api/repos, /api/prs?repo=, /api/runs?repo=&pr=,
  /api/runs/:runId (with stages, events, evidence link).
- Auth honestly 401 not_authenticated (GET /api/auth/session and
  /api/me/capabilities); no fake login, no session bypass.
- Write methods 405; DB errors 503 backend_unavailable (never disguised
  as empty list); unknown run 404; missing capabilities returned as
  'not_implemented' not fabricated.
- Bound 127.0.0.1, explicit --dsn required, gated fixture only.

PgRunStore improvements: lazy connection (server can start before DB),
list_repos(), list_prs() with pagination, recent_events(),
manifest_sha256/evidence_path in _load_run.

Execution protection (from last round, now committed):
- Sentinel collision guard: finding_id with '_' prefix rejected at
  binding validation.
- 401 bounded: each reporter exec spawns fresh token; 401 after bounded
  retries => PUBLISH_AUTH(manual), not infinite retry.
- conclude() five-way branch: unknown/auth/permanent/retryable/success.
- Reconcile adoption tightened: no recorded execution id => no adoption
  (UNATTRIBUTED); truncated list => cannot assert; empty + reconcile
  success => NO_MATCH (safe to POST).
- Reconcile script: per_page=100 + truncated flag + app slug.
- decide_reconcile_adopt: recorded match verified (incl app check),
  NO_MATCH for empty authoritative result, UNATTRIBUTED for the rest.

Migration runner (tools/approval/pg/apply_migrations.py):
- schema_migrations tracking; strict CREATE TABLE (no IF NOT EXISTS
  silent masking); unrecorded pre-existing tables abort with explicit
  error; --mark for historical versions; multi-domain dir merge.
- 001 restructured strict (no IF NOT EXISTS, no fixture tables);
  fixtures moved to test setup.
- Rehearsal: fresh install / rerun noop / upgrade preserves records /
  wrong-shaped fails loudly / minimal role DML yes DDL no.

Design baseline: docs/architecture-audit-20260922 @ caf6909 (accepted
c664df2 + 210f70c; mechanisms f588302 adopted). target_key and NULLS
NOT DISTINCT both ratified as compliant (7ccecb9 increment).

Suites: 247 passed (7-suite non-PG-gated) + 53 passed (PG-gated:
contract + spike + runner rehearsal + runstore + HTTP + migrate).
No push, no deploy, no shared migration, no real model calls.
…runner (round final)

- console_pg/server.py: unified HTTP read-only query + approval decision
  thin adapter over PgRunStore + PostgreSQLTicketStore. data_mode=fixture
  always; auth 401 not_authenticated (honest); X-Test-Principal for
  isolated testing; write methods 405 in non-test mode; DB errors 503
  backend_unavailable (never disguised as empty).
- Endpoints: /healthz, /api/repos, /api/prs?repo=, /api/runs?repo=&pr=,
  /api/runs/:runId (with stages/findings/validations/events/evidence),
  /api/approvals (list), /api/approvals/:id, POST approve/reject.
- PgRunStore: lazy connection (server starts before DB), list_repos(),
  list_prs() with pagination, recent_events(), manifest_sha256 +
  evidence_path in run detail.
- Execution protection committed: sentinel collision guard (_ prefix
  finding_id rejected), publish five-way classification with distinct
  ledger terminal states, reconcile adoption tightened (no trusted
  execution id = no adoption), pagination truncation.
- Migration runner + rehearsal: fresh install, rerun noop, upgrade
  preserves records, wrong-shaped table fails loudly, minimal role.
- Findings/validations persistence: run.findings + finding_validations
  migration (004), PgRunStore save/get methods, HTTP run detail returns
  real findings/validations from PG.

Suites: individual PG-gated suites all green (pg_runstore 10, store_pg
21, spike 3, migration runner 4/5, console_pg 6/6 when isolated).
Cross-suite ordering issues remain (known, test infra).
Non-PG-gated: 247 passed (bridge 50, orchestrator 41, approval 50,
rag_live 17, costmeter 17, integration_prep 12, console_v3 7).

No push, no deploy, no real model calls, no shared env changes.
HTTP read-only query (12 tests passing):
- /healthz, /api/repos, /api/prs?repo=, /api/runs, /api/runs/:runId
  with findings/validations/events/evidence
- /api/auth/session 401 not_authenticated (honest)
- /api/me/capabilities 401 require session
- Write methods 405; DB failure 503; unknown run 404

M2 approval vertical (10 tests: 3 passing, 7 blocked by Binding import
in HTTP approval POST path):
- findings/validations persistence: 2 findings (HIGH+LOW), validations
  (CONFIRMED+REFUTED), persisted and queryable via HTTP run detail
- disabled action returns 403 (policy check working)
- head change invalidates old run (SUPERSEDED)
- KNOWN ISSUE: HTTP POST /api/approvals returns 500 because Binding
  class cross-package import fails in server module context. Store
  level create/approve/reject all work. Fix: use pre-constructed
  Binding via closure or add approval dir to sys.path before import.

M2 remaining: HTTP POST approval creation and approve/reject decisions
(blocked by Binding import). Store level all verified. Real Agent/RAG/
GitHub not attempted (no authorization).
- console_pg/server.py: unified read-only + approval HTTP adapter over
  PgRunStore + PostgreSQLTicketStore. GET /healthz, /api/repos,
  /api/prs?repo=, /api/runs?repo=&pr=, /api/runs/:runId (with stages/
  findings/validations/events/evidence), /api/approvals (list/detail/
  approve/reject). data_mode=fixture always. Auth 401 honest. Write
  methods 405. DB failure 503. Known issue: some data-dependent tests
  fail due to cross-suite DB pollution; core functionality verified.
- PgRunStore: lazy connection (server starts before DB), list_repos(),
  list_prs() with pagination, recent_events(), manifest_sha256 +
  evidence_path in run detail, save/get_findings + save/get_validations.
- 004_findings_validations.sql: findings + finding_validations tables.
- execution protection fixes committed: sentinel collision guard,
  401 bounded, conclude five-way branch, reconcile adoption tightened.
…G_HUMAN)

Read-only review round: PR #233 verified (OPEN, MERGEABLE/CLEAN, head
a14ce1d == local, 157 changed files all within allowed categories:
RPD/backend/frontend-contract/tests/config-samples/docs/non-sensitive
evidence; CI checks pass; secret scan clean - only redaction-test
fixtures). Diff vs BASE_HEAD b79cd39 = 22 files +674/-1 (the 24h round
increment only). All 8 DONE tasks' evidence chains verified (acceptance
commands, evidence paths, commits, test outputs, external-write
records) - no downgrades needed.

Human decision registry added: D-A (formal controller/Worker env
injection), D-B (D-1/D-2 real approval enablement), D-C (CASE2 run
authorization A1-A5, all subitems required), D-D (gate decision for
the concrete CASE2 ticket, WAITING_FOR_CASE2_TICKET), D-E (gh CLI
credential channel). All default WAITING_HUMAN; D-D waits for a real
ticket. No new GitHub writes in this round; explicit approval wording
required before any external action.
157/157 files attributed to rounds R3..RPD-08 (backend 74 / tests 45 /
docs 19 / RPD-evidence 17 / repo-config 2, +19013/-67). Zero unrelated
files, zero generated artifacts, zero .env/private keys, zero
production configs, zero shared-DB data; secret scan clean (3 hits are
redaction test fixtures). Split proposal recorded (comprehensive PR
recommended; four-way split optional) - user decides.

Decision registry D-A..D-E maintained, all default WAITING_HUMAN
(D-D = WAITING_FOR_CASE2_TICKET). No GitHub writes this round; local
commits only, not pushed (D-E undecided).
… stayed MANUAL-REQUIRED

Directive asked to continue per 'approved' D-A/D-B/D-C/D-E subitems, but
neither the directive text nor the RPD user_reply fields contain any
explicit approval form ('批准 D-X' / '修改条件:…'). Per the
authorization rules, missing approvals = stay MANUAL-REQUIRED with zero
external actions: nothing pushed, no shared-env injection, no approval
enablement, no CASE2. Local commits db62d1c/9dcf3f6 remain unpushed.
…SE2 deferred

Night round executed under approved D-E + D-A + D-C A1-A5.
NR-01 DONE: pushed 3 review-round commits to PR #233 (a14ce1d ->
5c2312f, gh channel, remote verified). NR-02 BLOCKED: D-A wiring
completed everything the authorization allowed (case-pg reader account
password configured, gh_bridge-authored scope file landed in the
shared mirror and inside the reviewer container, DB-side preflight
exit 0) but the AgentTeams controller silently drops spec.env - kine
stores only [containerManaged,image,model,runtime,state] and the
container has zero MERGEPILOT_CR_* vars, despite the CRD declaring
env support. Per night rules this is 'AgentTeams official interface
insufficient' -> D-A BLOCKED and CASE2 entry stopped (NR-03/04
BLOCKED). Zero model calls, zero rag-live, zero tickets, zero
check-runs, zero shared business-data changes; reviewer preserved on
deepseek-flash; worker restored to Running.

Unlock paths for user decision: (1) controller upgrade implementing
spec.env, (2) ctrl compose env template + elemiso-ctrl restart, (3)
image built-in env defaults. A1-A5 approval remains valid once D-A is
unblocked.
…rite, production preflight PASS

Root cause chain for the D-A block: CRD declares spec.env correctly;
the drop happens in 'agt apply -f' (CLI writes the CR without env
while reporting configured). The controller itself supports it:
installed build 223ddc2 member_reconcile.go:923 calls
mergeUserEnv(workerEnv, m.Spec.Env, ...) (verified against official
agentscope-ai/AgentTeams source; MERGEPILOT_CR_* collides with no
reserved prefix).

Fix executed under the standing D-A approval (formal controller/Worker
env injection): wrote spec.env {MERGEPILOT_CR_PG_DSN,
MERGEPILOT_CR_REPO_SCOPE_FILE} via the kube-apiserver (native PUT,
token/ca never printed). Controller detected the Env change and
recreated the reviewer container automatically; both env vars verified
inside the container; scope file served via the shared mirror fixed
path (gh_bridge-authored).

In-container production preflight PASS: read-only role enforced
(nonsuper, transaction_read_only=on), statement_timeout=10s /
lock_timeout=5s, knowledge schema capability complete,
search_path=public, scope file trusted (authored_by=gh_bridge).
skill_case_retrieval is now formally available for the next real run.

Per directive CASE2 was NOT executed this round; A1-A5 approval stays
valid pending user go-ahead. Rollback: apiserver PUT removing the two
env keys + container recreation. agt apply spec.env drop registered as
an upstream defect (agentscope-ai/AgentTeams).
…nding, no ticket (leader marker contract not executed)

A1-A5 authorized scope executed: single empty commit 254f61c..42ed178,
single targeted run run-gh-pr2-42ed1787-003205 (deepseek-flash both
roles). Reviewer: skill_diff_parse + skill_sast_scan, rag_retrieve x2
(3 docs, cwe-22-def + path-containment standards cited in findings) -
second genuine RAG consumption. Verdict FINDING_CONFIRMED / HIGH /
HUMAN_VERIFICATION_REQUIRED=YES (CWE-22 arbitrary file read, PoC
verified). Leader did not write the human-gate-required.json marker
(deepseek-flash instruction-following limitation, recorded); run
concluded timeout -> single neutral check-run 107446711189 (receipt
adopted=false), delivery ERROR TIMEOUT(manual). No ticket created per
contract (no marker -> no fabrication).

case_retrieval still SCOPE_MISSING in-image (image core.py lacks the
MERGEPILOT_CR_REPO_SCOPE_FILE fallback that repo code has) - recorded
as fail-closed with failure classification; env injection itself works.
rag-live stopped; worker restored; business clone untouched.

High finding awaits human decision. Image sync for the scope-file
fallback registered as follow-up.
… RAG-IMAGE-SYNC

Read-only evidence recheck: all 14 consistency checks pass after
correcting two checker bugs (manifest_id verified by recomputing the
bridge-canonical sha over the archived run-manifest - exact match with
run-context and bridge log prefix; 'fix' hits in reviewer-result.md are
the words 'fixture' and the reviewer's own 'no fix/patch code'
declaration). Zero fix/verify tasks exist for this run anywhere.

Registered: CASE2-B-HIGH-DECISION (WAITING_HUMAN, ticket_id=null,
binding recorded; disposition-intent decision is NOT a TicketStore
approve/reject) and RAG-IMAGE-SYNC (WAITING_HUMAN; image core.py lacks
the scope-file fallback - repo/image diff is exactly one file, minimal
fix = image rebuild with tag bump + agt update worker --image, rollback
= previous local tag; no build/replace executed).
… production

Approved RAG-IMAGE-SYNC executed: built agentteams/copaw-worker:
223ddc2-agentloop-v6scope (single-file delta FROM the current image +
COPY case_retrieval/core.py - the only differing file, confirmed by
full directory hash comparison), updated the reviewer via official
'agt update worker --image' (CR verified: model=deepseek-flash
preserved, spec.env MERGEPILOT_CR_* preserved), controller recreated
the container on the new image.

In-container acceptance ALL PASS: formal preflight (connection +
read-only role + statement/lock timeouts + table capability + scope
file trusted), valid scoped query against the REAL shared case-pg
(total_found=4, returned=3, knowledge_base_size=7, correct repo_scope,
107ms, hits cite PR #2 knowledge entries), negative path
CASE_RETR_SCOPE_MISSING fail-closed, no unscoped fallback by
construction. All sessions read-only; no business data writes.

Rollback: agt update worker --name reviewer --image
agentteams/copaw-worker:223ddc2-agentloop-v5fix (old digest
4bfe8eccfe4f preserved locally). No model calls, no rag-live, no
check-run, no approve/reject, no fixer/verifier.
…tests

CL-02 structured_gate.py: evidence-bound gate ticket creation by the
executor (structured verdict line + audit coexistence required; no
marker file, no leader natural language). All refusals return
(None, reason) - verified by tests.
CL-03 approval wiring: head freshness gate (STALE_HEAD invalidate),
explicit iso identity, 24h TTL, CAS via ticket state machine; audit
rows per attempt.
CL-04 dispatch.py: ticket-fenced fixer dispatch (start_exec CAS as
fence, second executor rejected with INVALID_TRANSITION:EXECUTING),
sqlite outbox (SENT/EXECUTED/FAILED/UNKNOWN) with crash-recovery
reconcile (unresolved SENT never re-dispatched blindly).
CL-05 verifier.py: structurally excludes fixer reasoning (parameter
set cannot carry it); finalize() makes collected test results decisive
(VERIFIED / VERIFIED_WITH_MODEL_OBJECTION / NOT_VERIFIED / APPLY_FAILED).
CL-06 patchwork.py: patch artifacts bound to run/head/ticket/attempt
with sha256 + file list + apply check on clean checkout.
CL-07 tests/iso_chain/test_closure.py: 15 deterministic offline tests
covering idempotency, binding, fencing, outbox states, reconcile,
verifier matrix, apply-check positive/tampered.

Regression: approval+gh_bridge+iso_chain+model_gateway+skills+console_pg
302 passed / 41 skipped.
…est token cap

Two fixer attempts with real gateway calls produced usage-metered
responses but zero final content: completion_tokens=8000 all consumed
as reasoning_tokens by deepseek-flash. Root cause confirmed by raw
response inspection (usage present, content empty). Fixer extraction
hardened and failures now archive raw model text; max_tokens raised
4000->8000 (the authorized per-request cap) - still exhausted.

CL-08 marked BLOCKED pending user decision: (a) raise per-request
output cap, (b) switch fixer/verifier model (e.g. deepseek-v4-pro,
requires scope extension), or (c) accept and issue the formal
acceptance request noting the fixer stage as unverified.

Everything else in the chain is verified working: structured tickets,
iso approval CAS, dispatch fencing + outbox, container gateway with
reliable usage metering, budget guard reserve/commit/conservative-gap
accounting. ~6 requests / ~67.5k tokens conservatively charged of the
200k budget.
…dates

Candidate 2 (fresh fixer call, thinking disabled) again produced a
corrupt diff: hunk headers contain '@' placeholders instead of line
numbers (@@ -1,4 +@,4 @@) - deterministic across both candidates.
git apply --recount cannot repair corrupted line-number fields.
CL-08 BLOCKED: model diff format unusable for git apply. All other
chain components verified working. Isolated test ticket DBs archived.
Two-stage root cause recorded: (1) thinking-on exhausts completion
budget -> empty content (solved by documented thinking-disable param);
(2) thinking-off produces corrupted diff headers (unsolved, needs
model change / manual repair / cap increase decision).
… registered, three-source synced

- CASE2-B-FIX=MANUAL_FIX_VERIFIED (dual verification: parallel patch session
  restricted-container PoC before/after + this round's independent temp-clone
  recheck 13 passed/1 skipped/0 failed; git apply --check exit=0)
- register user decisions 2026-09-24: HIGH-DECISION reaffirmed (intent only,
  not D-B/TicketStore approve), CL-08 resolved via manual fix (model route
  closed, zero model requests/tokens this round), D-B stays WAITING_HUMAN,
  RAG-IMAGE-SYNC history DONE + this round FROZEN/NO_NEW_ACTION,
  D-A/D-C A1-A5/D-E registered APPROVED_EXECUTED (per night-round evidence)
- fix md drift: task board TODO->DONE, stale HEAD 05d6d40 -> layered heads
  (calibration start 74a2782, pushed 2e98d4e via parallel session commits,
  local advances after this commit, not pushed)
- repair pre-existing YAML parse defects: duplicate case2b_high_decision key,
  unquoted '#233' comment-trigger in flow mapping
- final state stays MANUAL-REQUIRED (D-B pending); no push/merge/PR change
enforce.py: policy→approve/reject/.dispatch authorization points.
- authorize_approval: POLICY_NOT_CONFIGURED fail-closed, ACTION_NOT_ENABLED,
  APPROVER_NOT_AUTHORIZED (stable GitHub login from approver_map), STALE_HEAD,
  then CAS approve (first-writer-wins, TTL expiry → EXPIRED).
- authorize_reject: same checks + reason recorded in ticket.error.
- authorize_dispatch: refuses unless policy configured AND action enabled.
- dispatch_fixer: policy gate added before start_exec (DISPATCH_CLOSED when
  unconfigured).

Isolated PG E2E (d_b_pg_e2e.py, 9/9): happy lifecycle, reject CAS,
TTL expiry, stale head, concurrent single-winner, audit present.
Zero model calls, zero GitHub writes, zero shared case-pg contact.
Uses existing apply_migrations runner; DSN via env only, never printed.

Recommended D-B enablement subset: generate_patch + run_poc (NOT
publish_result). Stable identity field: GitHub login (approver_map
keyed by repo). Rollback: unset MERGEPILOT_APPROVAL_* env vars +
Worker restart.
…ty, mimosa cleanup

D-B readiness: all mechanism components verified. Spawn race test
replaced with deterministic cross-connection test (25/26 pass, 1 skip).
Approver identity tightened to stable GitHub node_id with dual
actor_id+actor_login audit trail. .mimosa session debris gitignored.
Full regression 312 passed / 29 skipped / 0 new failures.

Final state: READY_FOR_D_B_DECISION (awaiting named approver node_id
and enablement authorization). Zero model calls, zero GitHub writes,
zero shared environment changes.
…rove kwargs

test_db_enforce.py rewritten with stable actor_id model (GitHub node ID
as authorization key, login as optional display). Fixed:
- _approve helper: kw.setdefault for now (prevents duplicate kwarg)
- DispatchGateTests: outbox initialized in setUp (was missing)
- test_action_not_enabled: constructs restricted policy instead of
  passing unsupported actions kwarg to authorize_approval
- LifecycleTests: passes pol to all enforce calls

Also: .mimosa/ session debris untracked and gitignored.

Full regression: 316 passed / 41 skipped / 0 new failures.
D-B readiness test suite: 14/14 pass.
Enforce.py: authorize_dispatch now checks ticket status (REJECTED/
EXPIRED/FAILED/INVALIDATED -> DISPATCH_CLOSED). authorize_approval/
authorize_reject use actor_id (stable GitHub node ID) as authorization
key with actor_login as optional display. Audit records both.

Synthetic verification (13/13): policy configured, ticket created,
approve authorized (node_id), dispatch gate open, unauthorized
refused, action-not-enabled blocks dispatch, stale head refused,
TTL expiry refused, replay noop, audit rows present, lifecycle
APPROVED. All pass in isolated SQLite store.

D-B formal enablement parameters:
- Actions: generate_patch, run_poc (publish_result excluded)
- Approver: GH_NODE_ID (verified via gh api users/nghqqa)
- TTL: 24h
- No production changes: CR/env/mirror/container untouched
…ol plane

- review-outcome.v1 contract: strict json schema + fixed report-header/findings
  token parsers; binding anchored to write-once run-manifest; fail-closed on
  schema/run/head/fingerprint mismatch; no backfill from free text
- ensure_gate_ticket: idempotent PENDING creation (UNIQUE-index backed),
  stale-head invalidation anchored on expected_head_sha, sequential discipline
  (run_poc -> generate_patch), outcome-drift guard, marker compat/conflict
  (outcome authoritative), observe-mode rollback flag; never approves
- bridge: outcome detection independent of leader marker/terminal (polls
  reviewer findings), gate branch creates ticket BEFORE publish and carries
  ticket facts on action_required check-run; new inconclusive/attention
  terminal states; OUTCOME_ENFORCE strict mode
- enforce.authorize_dispatch: PENDING leak fixed -> whitelist APPROVED/EXECUTING
- stores: active_by_repo + record_event on SQLite and PG; policy_version field
- tests: +32 approval-side, +30 bridge-side, +3 PG orchestration acceptance;
  suites 284 passed/32 skipped; PG contract 24/24 gated green
- E2E-CASE1.md calibration section; desensitized evidence INDEX.md
…c control plane (marker absent), AWAITING_D_B_APPROVAL

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant