This is the public, secret-free starting point for an organization's private ci-fleet configuration repository. It records which trusted projects may use each CI pool, the reviewed desired state for controller machines, infrastructure capacity budgets, logical deployment environments, and the standardized commands every project must expose.
It does not contain runner registration tokens, deploy credentials, private keys, host addresses, VM IDs, storage names, backup identifiers, or .env files.
One explicit exception permits reviewed operational Docker
default_address_pools[].base CIDRs and the optional default_bridge_cidr
gateway/interface CIDR in this private Git-authored policy. The controller must
render and inspect those exact Docker network ranges. They are not credentials,
host identity, or routable service endpoints. All VM, storage, backup, SSH,
rendered runtime, and unrelated infrastructure details stay outside Git.
flowchart LR
E[Public ci-fleet engine] -->|pinned engine commit| C[Private configuration]
C -->|pool policy and capacity budget| A[Controller at site A]
C -->|pool policy and capacity budget| B[Controller at site B]
P[Authorized project repositories] -->|all independent jobs| G[GitHub runner group]
A --> G
B --> G
G --> R[Ephemeral Docker workers]
P -->|approved image digest| D[Development hosts]
P -->|manual approval and same digest| X[Production hosts]
S[GitHub Environments or host secret store] -. secret values .-> P
classDef public fill:#dff4ff,stroke:#1570a6,color:#102a43
classDef private fill:#fff3cd,stroke:#9a6700,color:#3d2b00
class E public
class C,A,B,P,G,R,D,X,S private
-
Create a private repository from this public template.
-
Clone it and initialize the first project and controller:
./scripts/init.sh \ --organization your-org \ --project your-app \ --controller ci-01 \ --location primary-site \ --capacity-budget 1 \ --max-runners 1 \ --networks-per-runner 1 \ --default-bridge-cidr 192.0.2.1/24 \ --engine-ref <reviewed-ci-fleet-commit>
-
Edit
fleet.jsonto add the organization's real logical mappings and replace the generated RFC 5737 Docker network values with reviewed operational values. -
Run the strict policy check:
./scripts/validate.sh --strict
-
Configure secret values in GitHub Environments, root-owned host files, or an external secret manager. The repository stores only names such as
DEPLOY_AUTH.
The initializer refuses to replace a configured file unless --force is explicit. The product of --max-runners and --networks-per-runner cannot exceed 30, so its fictional /24 pool never produces networks smaller than /29 after reserving the controller and operator headroom. It runs non-strict validation because the generated documentation network values are intentionally not deployable. Run ./scripts/init.sh --help for repository, registry, runner-group, controller, location, capacity, network, bridge, resource, and output options.
fleet.json is the reviewed authority for logical controller state. Each entry in controllers has a unique ID and declares:
- its runner pool and logical location;
- whether it is
active,drained, ordisabled; - its unique GitHub scale-set name;
- an
experimental,stable, orretiringlifecycle; - the full reviewed ci-fleet commit SHA it runs;
- a zero managed minimum and reviewed maximum runner capacity;
- CPU and memory available to each ephemeral runner;
- Docker default-address pools, a positive reviewed maximum number of Compose networks per runner, and a reserved subnet count for health inspection.
status_reporting is deliberately omitted from initialized and reference
configurations. For an existing controller, roll out schema support in three
separately reviewed, integrated changes: first update only engine_ref and prove
routine reconciliation has activated that engine; then record the controller ID,
proven active ref, and reviewed status_reporting_config and
required_status_reporting capability booleans in
engine-rollout-evidence.json; only then add the optional status_reporting
object without changing engine_ref. Enabling required delivery also requires the
prior evidence to record both capabilities as true. Transition validation
rejects introducing or enabling the property without the corresponding evidence.
This staging prevents an older active manager from rejecting the new property before it
can upgrade itself. Endpoint and key values remain host-local and never enter
Git.
The same per-controller evidence record may declare
docker_network_policy_config and docker_default_bridge_cidr_config. A
controller may omit the first boolean only while it omits docker_network_policy.
It may omit the second while it omits default_bridge_cidr. Retained fields
require current evidence for the selected engine, and the selected engine
manifest must advertise the matching capabilities plus
docker_network_policy_adapter. Remove unsupported fields while retaining the
current engine, reconcile that release, and select an older engine only in a
later commit. A complete record has this shape after an operator has verified
the named engine is active:
{
"engine_ref": "1111111111111111111111111111111111111111",
"status_reporting_config": false,
"required_status_reporting": false,
"docker_network_policy_config": true,
"docker_default_bridge_cidr_config": true
}The controller ID is how a target host selects its declaration. A location is a non-sensitive logical slug such as primary-site or remote-site, never an address. Runtime-generated configuration and credentials remain host-local.
Each runner pool has a capacity_budget and a runner group that must not be assigned to any other pool. Unique runner-group assignment keeps routing and repository authorization unambiguous. The semantic validator enforces this cross-object rule because JSON Schema cannot compare values stored in separate object properties.
The validator totals the maximum capacity of every active or drained controller assigned to the pool and rejects overcommit. Drained capacity remains reserved so an undrain cannot silently exceed the reviewed budget. Disabled controllers do not reserve capacity.
For docker_network_policy, each pool has an IPv4 CIDR base and Docker
subnet prefix size. The size must be no longer than /29 and cannot be
broader than its base. A policy may declare at most 64 pools. Pools must not
overlap, and active or drained controllers must provide at least
max_runners * networks_per_runner + reserve_subnets + 1
subnets. The final subnet is reserved for the persistent controller Compose
network. The optional default_bridge_cidr is an IPv4 interface address and
prefix such as 192.0.2.1/28, not a canonical network base. Its address must be
a usable gateway, its prefix must leave room for containers, and its subnet must
not overlap the default-address pools. After excluding the network, broadcast,
and bridge gateway addresses, it must provide at least
max_runners + reserve_subnets container addresses. Size this bridge for
containers that use Docker's default network. Size the pools separately for job and controller
networks. The pool-capacity arithmetic above does not change when the bridge
field is present.
Real values belong in the private configuration; this template uses RFC 5737 documentation ranges only. Non-strict validation accepts those public examples. Strict validation rejects them until the private configuration uses a reviewed operational Docker pool CIDR. This is the narrow capacity-policy exception described above, not permission to commit host or service addresses.
The field is optional solely for staged upgrades from older schema-v3 engines.
First pin and activate a compatible engine without adding the field. In a second
reviewed desired-state commit, record the active ref and
docker_network_policy_config: true in engine-rollout-evidence.json. Add the
reviewed policy in a third commit, retaining the same ref and evidence. The older
engine rejects the new key, so skipped commits must not satisfy the gate.
Transition validation requires the activation evidence to exist in the previous
integrated state when introducing the policy. It also requires matching current
evidence whenever a controller retains the policy, including across engine
changes or rollbacks. Adding default_bridge_cidr later repeats the prior-state
gate with docker_default_bridge_cidr_config: true.
After the three-commit engine activation gate above, the pinned public engine
can apply or remove a validated policy with at most 64 pools. The executable
stage holds the installer lock, drains the controller and managed runners, and
changes Docker's default-address-pools key plus bip only when
default_bridge_cidr is configured. It preserves each prior key value, restarts
Docker, verifies the effective bridge subnet and gateway, probes capacity,
resumes the intended controller state, and checks health. Incompatible daemon
settings such as fixed-cidr fail before the drain. The transaction uses prior
key provenance and the prior rendered environment to roll back a failed or
interrupted operation, and retains recovery data if rollback cannot be verified.
A merged or configured policy does not authorize host mutation; rollout still
requires separate operator approval, exact-head CI, and proof for
the reviewed engine and desired-state commits. This stage does not create or
delete networks, prune resources, change controller scale or consumer labels,
or authorize application deployment.
Application repositories do not encode the number of available workers. They submit all independent tasks and shards. Do not use GitHub Actions strategy.max-parallel to model fleet size; controllers and the private configuration decide how many jobs run simultaneously. An application may limit concurrency only for a separately documented external-system constraint, not worker availability.
This separation lets one infrastructure change add, remove, drain, or resize controllers without editing every project workflow.
The public ci-fleet engine owns the host installer. Its intended interface consumes one logical controller from a pinned private configuration revision:
sudo ./scripts/install-worker-controller.sh \
--config-repo example-org/example-fleet-config \
--controller example-ci-01 \
--ref <reviewed-config-commit> \
--installUse --adopt instead of --install to bring an existing controller under Git-authored desired state. The engine contract also provides --check, --upgrade, --rollback, and --uninstall modes.
The command runs on the target Linux Docker machine. It validates the pinned configuration, renders host-local runtime state, preserves root-owned secrets, drains before disruptive changes, verifies a recoverable checkpoint, installs maintenance services, checks health, and reports drift without exposing credentials. OpenClaw or another agent may invoke it, but no agent is required.
The engine-side implementation is scripts/install-worker-controller.sh in the parent ci-fleet repository. Its accepted scope is isolated ordinary-CI fleet hosts under reviewed schema-v3 desired state; this configuration repository never installs a controller by itself.
Host retirement is an explicit reviewed transition:
- Change the controller state to
drainedand merge the private configuration change. - Converge the host and verify that it accepts no new work and has no active runner.
- Remove only fleet-owned residue and verify replacement capacity.
- Unregister its scale set and revoke that host's credentials.
- Change the declaration to
disabledor remove it in a later reviewed change. - Delete or repurpose the machine according to the installation's declared infrastructure policy.
Deleting one generic controller must not require application workflow changes. Legacy project-specific hosts should remain only until CI, promotion, and deployment no longer reference them.
- Public repositories never receive direct access to the trusted self-hosted runner pool.
- Every project publishes
scripts/ci/plan.jsonand implements./scripts/ci/run.sh <task> --shard INDEX/TOTALin its own Docker-defined test environment. ./scripts/ci/run.sh fastandfullremain aggregate developer commands; fleet scheduling expands their named tasks across available workers.- Every matrix job has a five-minute hard timeout, while expected test payload targets four minutes or less to reserve startup and reporting time.
- Application workflows submit all independent jobs; infrastructure configuration alone controls worker capacity.
- A GitHub runner group is assigned to exactly one runner pool.
- CI runner pools and deployment host groups are separate trust roles.
- Production deployment is manual and requires GitHub Environment approval.
- Controller engine revisions, reusable workflows, and third-party actions are pinned to immutable commits.
- Configuration contains logical identifiers only. Secret values, private host details, and credentials never enter Git.
- Promoted artifacts are container image digests; production does not rebuild a different image.
fleet.schema.json provides editor completion and structural documentation. scripts/validate.py is the authoritative dependency-free policy check, including cross-object relationships JSON Schema cannot express clearly.
Projects divide their total test-minutes into independent named tasks and deterministic shards. Forty-five test-minutes require at least nine perfectly balanced five-minute jobs in theory. In practice, projects should create additional shards targeting four minutes of test payload so checkout, image preparation, and reporting remain inside the five-minute job ceiling.
flowchart LR
P[plan.json] --> M[GitHub matrix]
M --> A[lint]
M --> B[unit 1 of 4]
M --> C[unit 2 of 4]
M --> D[integration 1 of 3]
M --> E[other independent shards]
Adding workers reduces wall-clock time only while independent shards remain queued. A genuinely indivisible test longer than five minutes must be optimized, split, or moved into an explicitly slower scheduled class outside ordinary CI.
| Path | Purpose |
|---|---|
fleet.json |
Fictional, valid schema-v3 configuration with one controller |
fleet.schema.json |
JSON Schema draft 2020-12 editor contract |
scripts/init.sh |
Safe first-project and first-controller initializer |
scripts/validate.sh |
Structural, relational, capacity, and secret-boundary validation |
scripts/test_policy.py |
Regression tests proving unsafe configurations fail closed |
examples/multi-host/fleet.json |
Fictional two-project, two-location controller topology |
SECURITY.md |
Secret handling and vulnerability reporting |
AGENTS.md |
Non-negotiable rules for humans and coding agents |
| Safe in this public template | Belongs in the private config repo | Belongs only outside Git |
|---|---|---|
| Schema, validator, fictional examples | Real repository names and runner-group policy | Tokens, passwords, private keys |
| Standard CI entrypoint names | Logical controller IDs and locations | Host addresses, VM IDs, SSH material |
| Controller state and lifecycle vocabulary | Capacity budgets and per-runner limits | Rendered runtime configuration |
| Reusable engine interface | Required secret names | Secret values and application credentials |
The public engine and this template use the Unlicense. See THIRD_PARTY_NOTICES.md before copying third-party material into a derived repository.