Repository navigation
roadmap(product): evolve Maka into a durable Agent workspace #2469
Description
Activity
I’d like to propose a focused child RFC for the Phase 2/3 boundary: an evidence-backed work-outcome projection.
I dogfooded current
main@5d2df91fdthrough the real dev CLI, Runtime Host, SQLite authorities, and Desktop UI on 2026-08-29. Two journeys exposed a gap that the existing adjacent issues do not own.Evidence
Successful journey
- A disposable Git project contained one failing shell verification.
gpt-5.4-minichanged exactly one production file and left the verifier unchanged.- Independent verification passed with
verification passed. - Desktop correctly showed the tool sequence and one changed file.
- The only completion statement was still the model’s final sentence; there was no independent semantic completion result tying the objective, validation, and change evidence together.
Interrupted / contradictory journey
- A direct
maka runinvocation started a foreground Bash command whose delayed write was observable incore_shell_runs. - SIGINT was sent to the exact CLI process after the Shell Run was
running. More than 55 seconds later the Shell Run was still running and no abort fact had appeared. - To keep the disposable test from producing the delayed write, I terminated only that verified test process group.
- Runtime then recorded Bash as failed (143); the following Read reported that the requested result path did not exist.
- The model nevertheless answered
DONE. - The AgentRun was persisted as
completed; the CLI exited 130. - After restart, Desktop showed the failed Bash, failed Read, and final
DONEas an ordinary settled conversation. Its Changes panel showed two files left by earlier runs in the same workspace, not evidence produced by this invocation.
This demonstrates a missing product/read-model distinction:
model/provider turn finished != execution had no unresolved failures != the user’s requested work is verified completeProposed first slice
Working name:
WorkOutcomeProjectionV1.- Define a shared read-only contract and pure projector over existing authoritative facts.
- Keep current AgentRun/Turn terminal semantics unchanged.
- Derive a user-level disposition such as
verified_complete | incomplete | blocked | cancelled | outcome_unknown, with stable reason codes and source/evidence references. - Treat unresolved tool errors, denied permissions, cancellation, missing requested outputs, and unverified validation as facts that prevent a
verified_completeprojection. - Cover contradictory journeys directly: tool error + final success text, missing output + final
DONE, cancellation during a live tool, and a real successful edit + validation.
Explicit non-goals for PR 1
- no new persisted store or execution authority;
- no Task Ledger / SessionTodo changes (RFC: Radically slim the Session Task Ledger #2290 / feat(runtime-host,cli): render Task Ledger mutations as a durable semantic timeline #4179);
- no provider finish-reason ownership changes (Runtime treats output-free provider finish frames as successful completion #3772 / refactor(runtime): single owner for provider finish-reason semantics #3802);
- no Usage presentation changes (Usage hides model attempts without token data and marks interrupted tools as successful #2128);
- no Session Catalog work (perf(desktop): make session data flow incremental and bounded #2913);
- no Desktop/CLI UI or wire change yet;
- no full Completion Packet or Recovery Card yet.
The first PR would be contract + pure projector + focused tests only. UI integration would wait for #2616 to converge, so this should not conflict with the Phase 1 journey.
Does this boundary align with #2469? If so, I’d be happy to open a focused child RFC and take the first contract/projector slice.
Dogfood and this draft were prepared with Codex; I reviewed the evidence and own the proposal.
- addedstaleNo qualifying activity within the lifecycle policy windowNo qualifying activity within the lifecycle policy window
on Sep 29, 2026 This issue has had no human activity for 30 days and has been marked stale. It will be closed in 30 days unless someone comments.
If the issue is still current, please confirm it against the latest
mainand add any information that would help move it forward. Assigned issues and issues labelledpinnedare exempt from this policy.Still current, so this should not close as stale.
Phase 1's readiness projection has landed (#2498, #2519, #2523, #2524), and so have the task storage and archive lifecycle steps from #5776: #5832, #5855, #5884, #5896, #5902 and #5961. Phases 2-6 (task status header, reviewable completion, delegation legibility, evidence and recovery, unified long-running work) have no delivered slice yet. #2560 (Work Board) is the active Phase 6 thread.
@yihanzhu, your Phase 2/3 work-outcome projection proposal above is still unanswered. If you're still interested, please open it as a focused child issue with the acceptance criteria, and I'll review it there.
- removedstaleNo qualifying activity within the lifecycle policy windowNo qualifying activity within the lifecycle policy window
on Oct 9, 2026
Problem
Maka already has a substantial Agent Runtime: local-first state, Runtime and Task event ledgers, permissions, Artifacts, Memory, Automations, Plan/Agent collaboration, Graph/subagent orchestration, Headless execution, and recovery semantics shared across Desktop, CLI, TUI, and Headless.
The product experience does not yet express those capabilities as one understandable work system. Users still have to infer important state from conversations, tool cards, settings, and implementation concepts:
This creates a gap between Maka's backend strengths and its user-facing value. The next product pass should not add more top-level Agent modes or tools. It should turn the existing Runtime into a clear, controllable, evidence-backed task experience.
Product direction
Position Maka as:
The product model should move from:
toward:
Conversation remains an important interaction surface, but it should not be the only representation of work.
Product principles
Conversation is an interface, not the work object
A task needs an explicit goal, state, execution progress, permissions, delegated work, deliverables, evidence, and recovery status. Chat is how the user collaborates with that task.
Every task has an observable state
Users should not need to read raw tool output to distinguish discussion, planning, running, waiting for the user, waiting for an external condition, blocked, recoverable, failed, and completed work.
Every completion has evidence
An Agent's claim that work is complete is not sufficient authority. Completion should reference existing execution facts such as changed files, commands, test results, Artifacts, unresolved items, and known risks.
Every interruption has an honest recovery decision
Maka should distinguish safe retry, safe continuation, required revalidation, unknown external side effects, and cases that require an explicit user decision.
Scope
This roadmap is primarily a product-experience and projection effort. It should reuse current Runtime authorities rather than inventing another task/event store.
Phase 0: reconcile product truth
Phase 1: make first success reliable
Phase 2: make task state legible
Phase 3: make completion reviewable
Phase 4: make delegation understandable
Phase 5: productize evidence and recovery
Phase 6: unify long-running work
Non-goals
This roadmap does not propose:
Delivery constraints
Success criteria
Activation
Comprehension
Trust
Continuity
Initial issue split
The first implementation milestone should stop after readiness and first-success evidence are working. Task status and Completion Packet follow on that stable foundation.
Related