Skip to content

roadmap(product): evolve Maka into a durable Agent workspace #2469

Description

@liugddx

Problem

Maka already has a substantial Agent Runtime: local-first state, Runtime and Task event ledgers, permissions, Artifacts, Memory, Automations, Plan/Agent collaboration, Graph/subagent orchestration, Headless execution, and recovery semantics shared across Desktop, CLI, TUI, and Headless.

The product experience does not yet express those capabilities as one understandable work system. Users still have to infer important state from conversations, tool cards, settings, and implementation concepts:

  • whether Maka is actually ready to accept a task;
  • what the current task is doing and how far it has progressed;
  • whether an Agent is running, waiting, blocked, or recoverable;
  • which delegated worker owns which part of the work;
  • whether completion is supported by execution evidence;
  • what is safe to resume after an interruption.

This creates a gap between Maka's backend strengths and its user-facing value. The next product pass should not add more top-level Agent modes or tools. It should turn the existing Runtime into a clear, controllable, evidence-backed task experience.

Product direction

Position Maka as:

A local-first, recoverable, and auditable Agent workspace for work that must finish.

The product model should move from:

conversation -> messages -> tool logs -> final assistant message

toward:

project -> task -> state -> execution -> evidence -> deliverables -> resume or follow-up

Conversation remains an important interaction surface, but it should not be the only representation of work.

Product principles

Conversation is an interface, not the work object

A task needs an explicit goal, state, execution progress, permissions, delegated work, deliverables, evidence, and recovery status. Chat is how the user collaborates with that task.

Every task has an observable state

Users should not need to read raw tool output to distinguish discussion, planning, running, waiting for the user, waiting for an external condition, blocked, recoverable, failed, and completed work.

Every completion has evidence

An Agent's claim that work is complete is not sufficient authority. Completion should reference existing execution facts such as changed files, commands, test results, Artifacts, unresolved items, and known risks.

Every interruption has an honest recovery decision

Maka should distinguish safe retry, safe continuation, required revalidation, unknown external side effects, and cases that require an explicit user decision.

Scope

This roadmap is primarily a product-experience and projection effort. It should reuse current Runtime authorities rather than inventing another task/event store.

Phase 0: reconcile product truth

  • Audit open product issues against the current implementation.
  • Close issues whose intended behavior has shipped.
  • Rewrite partially completed issues around their actual remaining scope.
  • Mark proposals superseded by newer Runtime or product architecture.
  • Publish a current capability matrix for Desktop, CLI/TUI, Headless, Plan, Graph/subagents, Resume, Memory, Automation, Artifacts, and Evidence.

Phase 1: make first success reliable

  • Define one shared readiness snapshot for Runtime Host, model target, workspace, tools, Git, and permissions.
  • Surface readiness failures before task submission with a cause and a direct repair action.
  • Prevent configuration drift, such as an unavailable default model, from appearing as a late unrelated feature error.
  • Establish one real first-success journey: configure a model, open a project, complete a small change, inspect the diff, run validation, and review the result.
  • Measure time to first verified result on Desktop and cover an equivalent CLI smoke path.

Phase 2: make task state legible

  • Project task status from existing Session, AgentRun, TaskRun, plan, permission, and recovery facts.
  • Add a compact task status header showing phase, progress, active delegated work, waiting conditions, intervention needs, elapsed time, and risk.
  • Present user-facing autonomy levels such as discuss, plan first, execute, and continue until done.
  • Keep Plan, Agent, Graph, Swarm, and Headless as Runtime mechanisms or advanced controls instead of requiring every user to understand them as peer product modes.

Phase 3: make completion reviewable

  • Define an evidence-backed Completion Packet projection.
  • Include completed objective, changed files, validation results, produced Artifacts, unresolved work, known risks, and suggested next steps.
  • Render the same semantics appropriately in Desktop and CLI.
  • Support a portable Markdown export without creating a second source of truth.

Phase 4: make delegation understandable

  • Continue feat(desktop): surface linked child sessions in the session workbar #1457 and move linked child work into the session workbar.
  • Show what each child owns, why it was created, and whether it is running, waiting, failed, or complete.
  • Allow users to inspect, interrupt, supplement, or take over delegated work.
  • Fold Graph into the same delegated-work surface instead of maintaining competing orchestration views.
  • Keep the parent Agent responsible for final synthesis.

Phase 5: productize evidence and recovery

Phase 6: unify long-running work

  • Make Automations create ordinary observable Tasks rather than a separate result model.
  • Treat Daily Review as a task template.
  • Surface Headless TaskRuns in Desktop with the same state, evidence, pause, and recovery semantics.
  • Give Memory entries explicit source, scope, freshness, and context-injection semantics.
  • Add background execution, budgets, and notifications only through the shared task model.

Non-goals

This roadmap does not propose:

  • rewriting the Runtime;
  • adding another event or task authority;
  • a big-bang Desktop redesign;
  • additional top-level Agent modes;
  • more tools or providers as part of this work;
  • a decorative Graph visualization before delegated-work status is understandable;
  • using visual polish as a substitute for information architecture;
  • expanding unattended autonomy before completion and recovery are trustworthy.

Delivery constraints

  • Each child issue and PR must be independently reviewable and reversible.
  • Prefer pure projections over existing authoritative facts.
  • Desktop, CLI, and TUI may use different layouts but must share state semantics.
  • Every task-facing surface must account for loading, empty, running, waiting, failed, recoverable, and completed states where applicable.
  • High-risk flows need behavior or real-journey tests, not only source-text contracts.
  • Do not expose a new top-level module when an existing task or workbar surface can own the experience.

Success criteria

Activation

  • A new user can reach a first verified result within 10 minutes.
  • Invalid model or environment configuration is actionable before task submission.

Comprehension

  • A user can identify whether work is running, waiting, blocked, recoverable, or complete within a few seconds.
  • Primary progress does not require reading raw tool logs.

Trust

  • Completed work includes validation evidence by default.
  • Unknown tool side effects are never presented as safe failure or success.
  • Model self-report does not replace Runtime facts.

Continuity

  • After restart, Maka can explain where work stopped and which continuation choices are safe.
  • Interactive, automated, delegated, and Headless work share task-state and result semantics.

Initial issue split

  1. Audit current product issues and capability status.
  2. Define a shared Maka readiness snapshot.
  3. Surface actionable startup/task-submission readiness in Desktop.
  4. Cover the first successful task journey in Desktop and CLI.
  5. Define a task-status projection over existing execution facts.
  6. Add the task status header.
  7. Define the Completion Packet projection.
  8. Present evidence-backed completion in Desktop and CLI.

The first implementation milestone should stop after readiness and first-success evidence are working. Task status and Completion Packet follow on that stable foundation.

Related

Activity

  1. yihanzhu commented on Aug 29, 2026

    @yihanzhu
    Contributor

    I’d like to propose a focused child RFC for the Phase 2/3 boundary: an evidence-backed work-outcome projection.

    I dogfooded current main@5d2df91fd through the real dev CLI, Runtime Host, SQLite authorities, and Desktop UI on 2026-08-29. Two journeys exposed a gap that the existing adjacent issues do not own.

    Evidence

    Successful journey

    • A disposable Git project contained one failing shell verification.
    • gpt-5.4-mini changed exactly one production file and left the verifier unchanged.
    • Independent verification passed with verification passed.
    • Desktop correctly showed the tool sequence and one changed file.
    • The only completion statement was still the model’s final sentence; there was no independent semantic completion result tying the objective, validation, and change evidence together.

    Interrupted / contradictory journey

    • A direct maka run invocation started a foreground Bash command whose delayed write was observable in core_shell_runs.
    • SIGINT was sent to the exact CLI process after the Shell Run was running. More than 55 seconds later the Shell Run was still running and no abort fact had appeared.
    • To keep the disposable test from producing the delayed write, I terminated only that verified test process group.
    • Runtime then recorded Bash as failed (143); the following Read reported that the requested result path did not exist.
    • The model nevertheless answered DONE.
    • The AgentRun was persisted as completed; the CLI exited 130.
    • After restart, Desktop showed the failed Bash, failed Read, and final DONE as an ordinary settled conversation. Its Changes panel showed two files left by earlier runs in the same workspace, not evidence produced by this invocation.

    This demonstrates a missing product/read-model distinction:

    model/provider turn finished
    !=
    execution had no unresolved failures
    !=
    the user’s requested work is verified complete
    

    Proposed first slice

    Working name: WorkOutcomeProjectionV1.

    • Define a shared read-only contract and pure projector over existing authoritative facts.
    • Keep current AgentRun/Turn terminal semantics unchanged.
    • Derive a user-level disposition such as verified_complete | incomplete | blocked | cancelled | outcome_unknown, with stable reason codes and source/evidence references.
    • Treat unresolved tool errors, denied permissions, cancellation, missing requested outputs, and unverified validation as facts that prevent a verified_complete projection.
    • Cover contradictory journeys directly: tool error + final success text, missing output + final DONE, cancellation during a live tool, and a real successful edit + validation.

    Explicit non-goals for PR 1

    The first PR would be contract + pure projector + focused tests only. UI integration would wait for #2616 to converge, so this should not conflict with the Phase 1 journey.

    Does this boundary align with #2469? If so, I’d be happy to open a focused child RFC and take the first contract/projector slice.

    Dogfood and this draft were prepared with Codex; I reviewed the evidence and own the proposal.

  2. added
    staleNo qualifying activity within the lifecycle policy window
    on Sep 29, 2026
  3. github-actions commented on Sep 29, 2026

    @github-actions

    This issue has had no human activity for 30 days and has been marked stale. It will be closed in 30 days unless someone comments.

    If the issue is still current, please confirm it against the latest main and add any information that would help move it forward. Assigned issues and issues labelled pinned are exempt from this policy.

  4. liugddx commented on Oct 8, 2026

    @liugddx
    MemberAuthor

    Still current, so this should not close as stale.

    Phase 1's readiness projection has landed (#2498, #2519, #2523, #2524), and so have the task storage and archive lifecycle steps from #5776: #5832, #5855, #5884, #5896, #5902 and #5961. Phases 2-6 (task status header, reviewable completion, delegation legibility, evidence and recovery, unified long-running work) have no delivered slice yet. #2560 (Work Board) is the active Phase 6 thread.

    @yihanzhu, your Phase 2/3 work-outcome projection proposal above is still unanswered. If you're still interested, please open it as a focused child issue with the acceptance criteria, and I'll review it there.

  5. removed
    staleNo qualifying activity within the lifecycle policy window
    on Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requesttrackingTracking or umbrella issue

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions