Skip to content

Reclaim functional ACP Jobs and expose Agentlet capacity controls #235

Description

@mydmdm

Context

Focused follow-up to #203 and implementation issue for the functional-Agent lifecycle/capacity milestone in #253.

External functional text workflows use ACP Jobs. Each Job receives an isolated live session identity and may spawn its own Agentlet-managed process. The shipped incremental implementation intentionally reused Deployment session/prompt behavior so #203 was not blocked on a broader lifecycle change, but Job completion does not currently imply prompt resource release.

Field use produced this failure while starting an interactive Copilot Profile:

Failed to spawn external agent 'GitHub Copilot (huabu) [copilot]': Max agents reached (10)

Agentlet defaults to ten concurrently managed agents. Retained functional Job processes can consume those slots until idle suspension, especially when several functional operations run in a short period, idle suspension is disabled, cleanup fails, or a harness does not terminate as expected.

Problem

ACP Jobs do not have an explicit, end-to-end ownership and release contract for their:

  • Agenetes Job handle;
  • ACP client/session entry;
  • spawn/session mapping;
  • Agentlet process and daemon slot;
  • pending startup, prompt, cancellation, and teardown work.

handle.close() removing a local client/session entry does not by itself prove that the remote Agentlet process stopped or that its slot became reusable. Cancelling a prompt or bounding the caller's wait also does not establish resource reclamation.

As a result, completed or abandoned functional workflows may retain capacity longer than their useful lifetime, and users receive only the upstream Max agents reached (10) spawn error without enough information to distinguish active interactive sessions from reclaimable functional Jobs.

Required behavior

1. Functional Job lifecycle and automatic cleanup

A functional ACP Job must enter bounded, automatic cleanup after:

  • successful completion;
  • prompt or workflow failure;
  • explicit cancellation;
  • host timeout/deadline;
  • caller abandonment or disconnect where ownership ends;
  • startup/initialization failure, including partial spawn;
  • cleanup invoked more than once or racing with process exit/suspension.

Cleanup must release every resource owned by that Job and confirm or explicitly report when the Agentlet process/slot could not be reclaimed. A successful workflow response must not leave an unowned live process behind merely to reuse Deployment retention behavior.

2. Isolation and durable-state boundaries

  • Cleaning up one Job must not stop another Job or an interactive Deployment, including executions that share a Profile, thread input, machine, or harness.
  • Interactive multi-turn Deployments retain their existing recovery and idle-suspension behavior.
  • Threaded Job history, Agenetes durable records, and harness-native session files are separate from live process ownership; preserve or delete them only according to their existing contracts.
  • Empty-thread and same-thread Jobs still require isolated live session identities.
  • Cleanup must remain correct across daemon reconnects, late exits, duplicate notifications, server shutdown, and retry/recovery paths.

3. Capacity diagnostics

Make capacity exhaustion actionable rather than exposing only the raw spawn failure:

  • Preserve a stable, bounded error classification for Agentlet capacity exhaustion.
  • Report the configured limit and available safe usage information without leaking commands, tokens, prompts, or unrelated session identities.
  • Make it possible to distinguish active interactive sessions from functional Jobs pending cleanup at the appropriate operator/debug surface.
  • Record cleanup failures and final disposition with bounded structured diagnostics.
  • Verify that a released functional Job slot can immediately be reused by a new interactive Agent.

Do not silently kill an arbitrary existing interactive session to make room.

4. Configurable Agentlet concurrent-agent limit

Evaluate and, if supported by the existing daemon lifecycle, expose the Agentlet max-agents limit through Huabu configuration rather than requiring an undocumented command-line change.

The implementation must define:

  • the persisted setting and safe numeric bounds;
  • the default value (preserve the current effective default unless evidence supports a deliberate product change);
  • whether the value is global or per managed daemon;
  • when a change takes effect and whether reconnect/restart is required;
  • behavior when lowering the limit below current usage;
  • Settings copy that explains the resource trade-off and makes clear that raising the limit is not a substitute for Job cleanup;
  • validation and stable errors for invalid values.

If investigation shows that safely configuring the limit requires a distinct implementation boundary, split that work into a linked child issue before implementation rather than silently dropping it from this issue.

Investigation requirements

Before modifying code:

  1. Reproduce slot consumption with functional ACP Jobs and verify which resources remain after success, failure, cancellation, timeout, and partial startup.
  2. Trace ownership through the functional caller, Agenetes Job/handle, ACP driver session/spawn caches, Agentlet Gateway/control RPC, and local Agentlet agents map.
  3. Establish which existing stop/close/suspend operations actually release the daemon slot and what acknowledgements/races exist.
  4. Distinguish resource leaks/retention caused by Huabu/Agenetes from harnesses that ignore graceful shutdown.
  5. Confirm the daemon lifecycle needed for applying max-agents configuration.
  6. Present the minimal implementation boundary and any required protocol/API changes for approval.

Acceptance criteria

Lifecycle

  • Successful functional external-Agent Jobs release their handle, ACP client/session, spawn mapping, process, and Agentlet slot within a documented bounded period.
  • Equivalent cleanup occurs after errors, cancellation, timeout, caller abandonment, and every partial-startup failure stage.
  • Cleanup is idempotent and race-safe across duplicate calls, late process exit, suspension, reconnect, and server shutdown.
  • Cleanup failure is explicit and observable; success is not reported as resource reclamation when the slot remains occupied.
  • One Job's cleanup cannot affect another Job or an interactive Deployment.
  • Threaded Job history and Deployment multi-turn/recovery behavior remain intact.

Capacity

  • Repeating functional workflows does not exhaust the default ten Agentlet slots after completed Jobs have been reclaimed.
  • A slot released by a functional Job is immediately reusable by a new interactive Agent launch.
  • Capacity exhaustion has a stable classification and actionable, redacted diagnostics.
  • Huabu does not automatically evict an arbitrary interactive session to satisfy a spawn.

Configuration

  • If shipped in this issue, max-agents is configurable with validated safe bounds and a documented default, ownership scope, and apply/restart lifecycle.
  • Lowering the limit below current usage has defined non-destructive behavior.
  • Settings explains that increasing capacity raises local resource usage and does not replace cleanup.
  • If capacity configuration is split out after investigation, a linked issue records the complete agreed scope and Follow-up: Agent lifecycle, unified Settings, Windows discovery, and multi-Agentlet enrollment #253 is updated accordingly.

Validation and documentation

  • Focused tests cover success, failure, cancellation, timeout, abandonment, partial spawn, duplicate cleanup, cross-Job isolation, interactive-session isolation, daemon reconnect, and slot reuse.
  • Integration coverage verifies the Agentlet agents map/limit behavior rather than only local handle deletion.
  • Architecture and operator documentation describe Job ownership, cleanup guarantees, idle suspension versus Job release, capacity diagnostics, and any limit configuration.

Non-goals

  • Do not remove the Job/Deployment distinction or introduce a new workflow framework.
  • Do not make functional Jobs share a live session to reduce slot usage unless separately investigated and approved; isolation remains the current contract.
  • Do not delete durable history or harness-native session files merely to reclaim a process.
  • Do not change interactive Deployment retention/idle policy except where required to preserve cleanup isolation.
  • Do not solve exhaustion only by raising the default limit.

Related

Activity

  1. changed the title [-]ACP driver Job 的资源回收和释放[/-] [+]Reclaim functional ACP Jobs and expose Agentlet capacity controls[/+] on Sep 30, 2026
  2. mydmdm commented on Oct 5, 2026

    @mydmdm
    ContributorAuthor

    Delivered to the Alpha integration branch by #261 at 718f758dea5fad50521794e421f036c47d05aaaa. The complete approved issue scope is present in the current alpha history. Further Canary regressions can reopen this issue or be tracked through a focused follow-up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions