Repository navigation
Resume Builder: native-chat app, three-workspace starter pack, eval-hardened instructions - #292
Draft
davewaring wants to merge 85 commits into
Draft
davewaring wants to merge 85 commits into
davewaring wants to merge 85 commits into
Conversation
- Profile trigger: intent + offer (natural phrasing, proactive offer when complete, one confirming question when unsure), never silent — replaces the explicit-ask-only rule - Unresolved items become visible [gap: ...] markers in the Profile; the template's placeholders now use the same convention as the app's incomplete-Profile check - Chat-driven Profile updates: re-read file, apply, say what changed - resume-profile.md is now self-describing (each section states what a strong entry looks like); run-interview.md carries the how, with persona-adaptive guidance; AGENT.md reads the Profile up front - 'the host' -> 'BrainDrive' (runtime model has no host vocabulary) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
- base/AGENT.md: workspace model (Your Agent / page / installed app capability, all app-generic), boot rule (read the active workspace's AGENT.md first; workspaces own their documents and workflow); interview craft stays in per-page run-interview.md where it already lives - career/AGENT.md: Career Workspace framing + generic Available App Workspaces section (never assume an app exists; handoff when installed, describe/offer install when available; apps may use owner-authorized Career context but own their artifacts) - 'the host' -> 'BrainDrive' in both (runtime model has no host vocabulary) - New mirror test guards starter-pack vs your-memory instruction-file drift (skips where the gitignored fixture is absent) Content authored in the 2026-08-16 fixture sessions; this makes the versioned starter pack the durable home per the AGENTS.md pairing rule. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
… absolute dates
Dial-in round 1 (P1 fixture, GLM 5.2) surfaced two grounding failures:
the drafted summary rounded an in-progress degree up to 'recent graduate',
and relative dates ('last June') were resolved by guessing instead of
confirmation, producing wrong years twice. Both fixes are generic
interview-craft rules, not fixture-specific guidance.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
… dates Round 2 of the dial-in showed the model asserting an unsourced current date and pressing a date correction three times. Owner-confirmed dates are final and recorded without lingering confirmation notes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Dial-in round 3: the model wrote 'B.S.' when the owner only said 'Environmental Studies' (round 2 had correctly gap-marked the same unknown). Specifics must come from the owner — ask or gap-mark. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
… holds only resume content Dial-in P2: the model never asked whether the owner had an existing resume (spec's existing-material-first principle) — the owner had to offer the paste. Also standardizes that a written Profile replaces all template guidance text, since everything in the Profile flows into the formatted resume; whether the Profile should instead stay self-describing with a renderer that strips prose is queued as a product decision. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Dial-in P3: the owner's hybrid/no-weekly-travel constraint landed in the resume summary because the instructions gave the model no other home. Implements the shared-memory contract (RB-OQ-8/L-8): grounded reusable facts update the appropriate owner memory, the Resume Profile stays resume-private, no guesses recorded. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Seen twice in the dial-in (P4 'monthly-ish' recorded as 'monthly'; P5 'maybe 2019' proposed as the year): uncertainty firms up only when the owner confirms it, otherwise it stays a visible gap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…ording, unnamed-but-confirmed items become gaps
Judge pass on the confirmation sweep failed one cell: P2 groundedness 3/5
for 'reducing repeat inquiries' — an effect the owner never stated —
plus register inflation citations on P1/P5 ('co-lead' for 'help run',
'fast-paced' for 'pretty routine') and a confirmed-but-unnamed tool
silently dropped. All three are generic no-unbacked-claims classes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
An EISDIR from memory_read (model reads a directory path) was mapped to execution_failed/recoverable:false, and engine/loop.ts aborts the turn on any non-recoverable tool error — the owner sees a silently dead chat. Hit twice in Resume Builder dial-in runs. Map EISDIR/ENOTDIR to a recoverable path_invalid at both seams (local errno mapping and the MCP client boundary, since the packaged memory server reports it as fatal) with a message that steers the model to memory_list. Backend review note for DJ: the deeper fixes — the packaged MCP server's own error mapping and whether loop.ts should ever end a turn without a user-visible error — are left untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Two judge sweeps show the same shape: interview behavior is clean, but resume-idiom pattern-fill slips in at write time (sweep 1: an invented outcome clause; sweep 2: an asserted store title, an unasked duty, an inferred employer city). Adds a pre-announce re-read requiring every specific to trace to the owner or become a gap. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…erred from the owner's own city Adds one generic grounding line to the interview instructions: an employer's location is a Profile specific like any other — ask for it or leave a gap. Closes the dominant residual inference class from the 2026-08-18 evaluation sweeps (owner's home city written onto employer lines) at the source. Spec-derived, not persona-specific. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…, curated) Experiment: GLM 5.2 reviewed its own operating instructions against the spec-derived goals and its observed general failure classes, and proposed edits to maximize its own compliance. All six proposals passed curation (no fixture/persona specifics, no rule weakened) and are applied verbatim: a five-step pre-announce self-check gate (count-and-name gap markers; 'ready to go' forbidden while gaps remain unnamed), a complete-read rule before any 'no stored career context' claim, per-turn re-reads of the two most-violated rules, a forbidden-transformations list, and a compliance note at the Profile template read point. Note for renderer work: two lines state the current renderer prints gap markers verbatim — update them when the renderer defect bundle lands. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…e endpoint) Additive profile alongside openrouter/ollama for first-party Z.ai API access — new GLM releases land there before OpenRouter. Default model glm-5.2; per-account provider_default_models selects newer versions. No BrainDrive-owned key involved; BYOK like openrouter. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…t Profile write waits for a direct go-ahead (1) A section the owner resolves as none (no certifications, no other roles) is omitted from the Profile rather than recorded as a gap marker — a resolved none is not a gap, and everything in the file renders as written. Kills the recurring certifications-section false-prediction class across three models. (2) Even on an intent-path opener, the first Profile write waits for an offer and a direct yes; announcing after the fact is not a substitute. Clarifies the spec Q1 reading judges split on across cells. Both generic and spec-derived; no persona-specific content. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…-full on contentOverride section) The document section injected into ChatPanel's absolute-positioned content box had no height constraint (flex-1 without a flex parent), so it grew to content height and the outer overflow-hidden clipped everything below the fold; the inner overflow-y-auto never engaged. MessageList works because it sets height:100% explicitly — the section now does the equivalent via h-full. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…wn, matching DocumentView Profile, Agent Instructions, and Interview Guide views showed raw markdown in a <pre>; they now use the same prose-bd + MarkdownContent pattern as the document pages, with markdown visible only in edit mode. Resume view keeps its dedicated ResumePreview. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft for Dave J's review per the 8/17 division of labor (D344) — nothing merges until you've been through it. This is the front-end/starter-pack half: the Resume Builder rebuilt on native chat, the D342 three-workspace model ported into the starter pack, and the instruction pack that survived four days of adversarial evals.
What's in it
1. Starter pack (the product surface — review first)
base/AGENT.mdrewritten to the three-workspace model (Your Agent / page / app) per D342 — thinner base, behavior delegated to per-workspace filesresume-builderapp workspace: AGENT.md, interview guide, self-describing Profile template — the instruction pack as eval-hardened through 15+ judged persona cells2. Resume Builder native app
3. Backend fixes (tested)
4. Additive:
z-aiprovider profileEvidence
starter-pack-testing-harness/runs/2026-08-19-t827-pack-ab/Known open issues — deliberately not masked
The dial-in surfaced 13 code-layer findings that instructions cannot and should not paper over (renderer gap-marker/empty-section behavior, no incomplete-Profile gate on Create Resume, dead-stream delivery/idempotency, model-asserted counts, and more). These are the companion backend work list, not regressions from this PR — the full enumerated list with evidence lives in the Library under T-829 (
pulse/backlog.md) and the Resume Builder eval kit run log.🤖 Generated with Claude Code
https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ