Skip to content

Resume Builder: native-chat app, three-workspace starter pack, eval-hardened instructions - #292

Draft
davewaring wants to merge 85 commits into
devfrom
codex/conversational-resume-intake
Draft

davewaring wants to merge 85 commits into
devfrom
codex/conversational-resume-intake

Conversation

@davewaring

Copy link
Copy Markdown
Member

Draft for Dave J's review per the 8/17 division of labor (D344) — nothing merges until you've been through it. This is the front-end/starter-pack half: the Resume Builder rebuilt on native chat, the D342 three-workspace model ported into the starter pack, and the instruction pack that survived four days of adversarial evals.

What's in it

1. Starter pack (the product surface — review first)

  • base/AGENT.md rewritten to the three-workspace model (Your Agent / page / app) per D342 — thinner base, behavior delegated to per-workspace files
  • Career template aligned to the workspace model, including the app-handoff contract
  • New resume-builder app workspace: AGENT.md, interview guide, self-describing Profile template — the instruction pack as eval-hardened through 15+ judged persona cells

2. Resume Builder native app

  • Chat-first app on native chat (no sandboxed frame): conversation + Profile/Resume/Advanced document views, deterministic render, PDF export
  • Client: workspace views wired into ChatPanel; two fixes from today's owner run-through (unscrollable document views; Profile now renders markdown like the other document pages)

3. Backend fixes (tested)

  • Directory-read tool errors recoverable instead of turn-fatal (EISDIR class)
  • Resume chat turn-delivery and stale-projection fixes

4. Additive: z-ai provider profile

  • api.z.ai OpenAI-compatible endpoint, BYOK; added so we could evaluate GLM 5.3 there before its OpenRouter listing
  • We are sticking with GLM 5.2 as the default for now; 5.3 measured better in the evals (accurate counts, exemplary edit-respect) and we may switch once it lists on OpenRouter — with a re-run of the eval suite before any flip
  • No change to existing provider profiles

Evidence

  • Eval record: three five-persona sweeps + a confirmation sweep against the versioned fixture packet — zero fabricated material content in any artifact in any cell; full run logs + judge verdicts in the Library eval kit
  • Model: validated on GLM 5.2 (current default); see the z-ai note above for the 5.3 path
  • Page regression: A/B of dev pack vs. this branch on the starter-pack harness (5 cells/side, Opus judge) — no regression; evidence in the Library at starter-pack-testing-harness/runs/2026-08-19-t827-pack-ab/

Known open issues — deliberately not masked

The dial-in surfaced 13 code-layer findings that instructions cannot and should not paper over (renderer gap-marker/empty-section behavior, no incomplete-Profile gate on Create Resume, dead-stream delivery/idempotency, model-asserted counts, and more). These are the companion backend work list, not regressions from this PR — the full enumerated list with evidence lives in the Library under T-829 (pulse/backlog.md) and the Resume Builder eval kit run log.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ

davewaring and others added 30 commits August 16, 2026 11:29
- Profile trigger: intent + offer (natural phrasing, proactive offer when
  complete, one confirming question when unsure), never silent — replaces
  the explicit-ask-only rule
- Unresolved items become visible [gap: ...] markers in the Profile; the
  template's placeholders now use the same convention as the app's
  incomplete-Profile check
- Chat-driven Profile updates: re-read file, apply, say what changed
- resume-profile.md is now self-describing (each section states what a
  strong entry looks like); run-interview.md carries the how, with
  persona-adaptive guidance; AGENT.md reads the Profile up front
- 'the host' -> 'BrainDrive' (runtime model has no host vocabulary)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
- base/AGENT.md: workspace model (Your Agent / page / installed app
  capability, all app-generic), boot rule (read the active workspace's
  AGENT.md first; workspaces own their documents and workflow); interview
  craft stays in per-page run-interview.md where it already lives
- career/AGENT.md: Career Workspace framing + generic Available App
  Workspaces section (never assume an app exists; handoff when installed,
  describe/offer install when available; apps may use owner-authorized
  Career context but own their artifacts)
- 'the host' -> 'BrainDrive' in both (runtime model has no host vocabulary)
- New mirror test guards starter-pack vs your-memory instruction-file
  drift (skips where the gitignored fixture is absent)

Content authored in the 2026-08-16 fixture sessions; this makes the
versioned starter pack the durable home per the AGENTS.md pairing rule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
… absolute dates

Dial-in round 1 (P1 fixture, GLM 5.2) surfaced two grounding failures:
the drafted summary rounded an in-progress degree up to 'recent graduate',
and relative dates ('last June') were resolved by guessing instead of
confirmation, producing wrong years twice. Both fixes are generic
interview-craft rules, not fixture-specific guidance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
… dates

Round 2 of the dial-in showed the model asserting an unsourced current date
and pressing a date correction three times. Owner-confirmed dates are final
and recorded without lingering confirmation notes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Dial-in round 3: the model wrote 'B.S.' when the owner only said
'Environmental Studies' (round 2 had correctly gap-marked the same
unknown). Specifics must come from the owner — ask or gap-mark.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
… holds only resume content

Dial-in P2: the model never asked whether the owner had an existing
resume (spec's existing-material-first principle) — the owner had to
offer the paste. Also standardizes that a written Profile replaces all
template guidance text, since everything in the Profile flows into the
formatted resume; whether the Profile should instead stay self-describing
with a renderer that strips prose is queued as a product decision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Dial-in P3: the owner's hybrid/no-weekly-travel constraint landed in the
resume summary because the instructions gave the model no other home.
Implements the shared-memory contract (RB-OQ-8/L-8): grounded reusable
facts update the appropriate owner memory, the Resume Profile stays
resume-private, no guesses recorded.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Seen twice in the dial-in (P4 'monthly-ish' recorded as 'monthly';
P5 'maybe 2019' proposed as the year): uncertainty firms up only when
the owner confirms it, otherwise it stays a visible gap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…ording, unnamed-but-confirmed items become gaps

Judge pass on the confirmation sweep failed one cell: P2 groundedness 3/5
for 'reducing repeat inquiries' — an effect the owner never stated —
plus register inflation citations on P1/P5 ('co-lead' for 'help run',
'fast-paced' for 'pretty routine') and a confirmed-but-unnamed tool
silently dropped. All three are generic no-unbacked-claims classes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
An EISDIR from memory_read (model reads a directory path) was mapped to
execution_failed/recoverable:false, and engine/loop.ts aborts the turn on
any non-recoverable tool error — the owner sees a silently dead chat.
Hit twice in Resume Builder dial-in runs. Map EISDIR/ENOTDIR to a
recoverable path_invalid at both seams (local errno mapping and the MCP
client boundary, since the packaged memory server reports it as fatal)
with a message that steers the model to memory_list.

Backend review note for DJ: the deeper fixes — the packaged MCP server's
own error mapping and whether loop.ts should ever end a turn without a
user-visible error — are left untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Two judge sweeps show the same shape: interview behavior is clean, but
resume-idiom pattern-fill slips in at write time (sweep 1: an invented
outcome clause; sweep 2: an asserted store title, an unasked duty, an
inferred employer city). Adds a pre-announce re-read requiring every
specific to trace to the owner or become a gap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…erred from the owner's own city

Adds one generic grounding line to the interview instructions: an
employer's location is a Profile specific like any other — ask for it
or leave a gap. Closes the dominant residual inference class from the
2026-08-18 evaluation sweeps (owner's home city written onto employer
lines) at the source. Spec-derived, not persona-specific.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…, curated)

Experiment: GLM 5.2 reviewed its own operating instructions against the
spec-derived goals and its observed general failure classes, and proposed
edits to maximize its own compliance. All six proposals passed curation
(no fixture/persona specifics, no rule weakened) and are applied verbatim:
a five-step pre-announce self-check gate (count-and-name gap markers;
'ready to go' forbidden while gaps remain unnamed), a complete-read rule
before any 'no stored career context' claim, per-turn re-reads of the two
most-violated rules, a forbidden-transformations list, and a compliance
note at the Profile template read point.

Note for renderer work: two lines state the current renderer prints gap
markers verbatim — update them when the renderer defect bundle lands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…e endpoint)

Additive profile alongside openrouter/ollama for first-party Z.ai API
access — new GLM releases land there before OpenRouter. Default model
glm-5.2; per-account provider_default_models selects newer versions.
No BrainDrive-owned key involved; BYOK like openrouter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…t Profile write waits for a direct go-ahead

(1) A section the owner resolves as none (no certifications, no other
roles) is omitted from the Profile rather than recorded as a gap marker
— a resolved none is not a gap, and everything in the file renders as
written. Kills the recurring certifications-section false-prediction
class across three models.
(2) Even on an intent-path opener, the first Profile write waits for an
offer and a direct yes; announcing after the fact is not a substitute.
Clarifies the spec Q1 reading judges split on across cells.

Both generic and spec-derived; no persona-specific content.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…-full on contentOverride section)

The document section injected into ChatPanel's absolute-positioned content
box had no height constraint (flex-1 without a flex parent), so it grew to
content height and the outer overflow-hidden clipped everything below the
fold; the inner overflow-y-auto never engaged. MessageList works because it
sets height:100% explicitly — the section now does the equivalent via h-full.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
…wn, matching DocumentView

Profile, Agent Instructions, and Interview Guide views showed raw markdown
in a <pre>; they now use the same prose-bd + MarkdownContent pattern as the
document pages, with markdown visible only in edit mode. Resume view keeps
its dedicated ResumePreview.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHym7ksrU2yyDhGZCQmLxQ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants