Fleet hosts are generic Docker infrastructure. They must receive operating-system security fixes automatically, report health, and clean expired fleet resources and empty job networks in their configured Docker address pools.
- Enable the distribution's unattended security upgrades.
- Do not automatically replace controller, runner, Docker major-version, kernel, or project images without repository validation.
- Schedule reboots only after draining the controller.
- Run health checks every five minutes.
- Run scoped cleanup daily.
- Compare installed state with its pinned Git configuration every fifteen minutes.
- Never schedule
docker system prune. - Alert before Docker storage reaches 80%; treat 90% as critical.
Install and enable Debian's supported security update mechanism:
sudo apt-get update
sudo apt-get install -y unattended-upgrades apt-listchanges
sudo dpkg-reconfigure -plow unattended-upgrades
sudo unattended-upgrade --dry-run --debugReview /etc/apt/apt.conf.d/50unattended-upgrades and confirm only the intended Debian security origins are enabled. Keep automatic reboot disabled; a generic CI host may still be executing a job when a package requests reboot.
scripts/install-worker-controller.sh installs and enables all three timer pairs:
ci-fleet-health.timerruns the complete fleet health contract;ci-fleet-cleanup.timerremoves expired inactive fleet-owned resources and empty networks in the configured Docker default address pools;ci-fleet-drift.timercompares the installation with the exact pinned configuration commit without applying changes;ci-fleet-reconcile.timerfetches the latest reviewed desired-state commit from the private repository over authenticated HTTPS and applies it automatically.
Run each service manually once before relying on its timer:
sudo systemctl start ci-fleet-health.service
sudo systemctl start ci-fleet-drift.service
sudo journalctl -u ci-fleet-health.service --since today
sudo journalctl -u ci-fleet-drift.service --since today
sudo systemd-run --wait --pipe --collect \
--property=User=root \
--property=EnvironmentFile=/etc/ci-fleet/ci-fleet.env \
--property=WorkingDirectory=/opt/ci-fleet/current \
/opt/ci-fleet/current/scripts/cleanup.shThis transient service runs a dry-run with the same root-owned rendered
environment and working directory as the applying cleanup service. After
reviewing its candidates, run sudo systemctl start ci-fleet-cleanup.service.
The cleanup service reads the default address pools from its rendered
/etc/ci-fleet/ci-fleet.env. The documented dry-run loads that file through
systemd, exactly as the applying service does. Without rendered pool values,
cleanup uses only the existing fleet-label expiry rule. Pool reclamation removes
a network with zero attached containers once it is at least ten minutes old
and every allocated subnet lies inside those pools, including project Compose
networks without fleet ownership or expiry labels. It does not infer ownership
from a project name or require a changed desired-state commit.
The fixed ten-minute grace is twice the ordinary job timeout and protects the
gap between network creation and container attachment. Cleanup preserves fresh
networks and networks whose Created timestamp is missing, invalid, lacks a
timezone, or lies in the future. Explicit fleet-label expiry cleanup keeps its
existing age policy.
Cleanup always preserves ci-fleet_default, networks with the controller's
com.docker.compose.project=ci-fleet identity, and the daemon's default bridge.
It preserves networks with any attached container, including stopped containers.
It reports unlabeled networks outside the configured pools without deleting them.
Mixed allocations and networks without allocation information do not qualify for
pool reclamation. Existing instance-scoped expiry cleanup still applies to other
fleet-labeled networks.
Cleanup checks network endpoints and all container references with
docker ps -aq with filters for both the exact network ID and name, including
stopped and created containers. It repeats those checks before each individual
docker network rm. Docker's endpoint check alone cannot protect a container
created between inspection and removal, so ordinary fleet Docker requests use
an independent docker-socket-proxy Compose service. The controller starts only
after that service is healthy. The proxy has no GitHub credentials. It gates
container creation, network connection, and other operations that can introduce
container references through /run/lock/ci-fleet/docker-maintenance.lock.
Cleanup holds that same lock exclusively while it inspects and removes networks.
The proxy retains its shared lock until Docker completes a dispatched mutation,
including when the caller disconnects. Nested raw socket bind requests are
redirected to the proxy socket without project-specific changes.
For /build, the proxy accepts contexts up to 1 GiB and spools at most two
uploads concurrently, bounding temporary context storage to 2 GiB. Larger
contexts receive HTTP 413; use .dockerignore to exclude unnecessary files.
Additional uploads wait for a slot and can cancel while waiting. Spool files
are unlinked before writing, so a proxy crash leaves no context files behind.
The proxy closes each spool and releases its slot when the upload to Docker
ends, while retaining the maintenance lock until the build response completes.
Before removal, cleanup waits up to ten minutes for runnable noncontroller
containers to finish. It releases the exclusive lock between checks so jobs can
finish their Docker work. It requires two idle checks separated by one second.
It never pauses or removes active work. It exempts only current-protocol
controller and proxy containers with the expected shared-directory mount and
Docker endpoint. Older controllers and unknown active workloads defer cleanup.
The applying service also takes the installer lock nonblockingly. An active
installer, including uninstall's inherited lock, produces DEFER without
network deletion. The service has a fifteen-minute start timeout.
Proxy mutations and raw cleanup removals create empty, durable markers under
/run/lock/ci-fleet/inflight/<boot-id>/ and removals/<boot-id>/ before dispatch.
Markers contain no request bodies or credentials. A complete response clears
the marker. A failed connection before any upstream connection is acquired also
clears that request's marker, because Docker received no request. A killed
process or uncertain response preserves it. Cleanup defers
while a current-boot or unknown marker remains, and the proxy refuses new
reference mutations while an unresolved cleanup-removal marker remains.
Restarting the proxy or controller does not clear uncertainty. A host reboot
changes the kernel boot ID; cleanup may then clear older boot directories and
recover capacity automatically. Never delete current-boot markers to bypass
this guard.
This coordination applies to an isolated ordinary Docker/Compose fleet. It
does not make direct host-root Docker requests outside the proxy atomic.
Cleanup defers on an active or unknown Swarm state because Swarm can create
tasks asynchronously after an API response. Do not use this reclamation policy
on a daemon with independent API writers or other asynchronous orchestrators.
The direct runner uses the shared directory and DOCKER_HOST so it can reconnect
after the proxy rebinds its socket. /var/run/docker.sock remains available for
explicit-path clients. Such clients and nested socket-file mounts can retain an
old inode across a proxy restart, which can interrupt their job. Controller-only
restarts leave the independent proxy running.
Cleanup rechecks the network after a confirmed removal failure and retries a
clean network once.
Repeated transient endpoint conflicts defer that network and let later cleanup
candidates proceed. Other persistent Docker errors fail the cleanup service;
uncertain transport failures also preserve the removal marker.
Cleanup never invokes docker network prune or docker system prune. The controller's
low-water gate remains a backstop while the existing daily timer restores leaked
subnet capacity.
- Merge and apply a reviewed desired-state change setting the controller to
drained, or otherwise pause the controller and wait for zero managed runners. - Confirm no managed runner container is active.
- Apply updates and reboot.
- Confirm Docker, disk, time synchronization, DNS, and outbound GitHub connectivity.
- Run
scripts/install-worker-controller.sh --checkagainst the installed configuration commit. - Apply a reviewed
activedesired-state commit and confirmMIN=0produces no idle container.
Dependabot proposes updates for GitHub Actions, Go modules, and both Dockerfiles. Those pull requests must pass inert validation and be reviewed before merge. This preserves unattended host security patching without silently changing the runner control plane.
The Update GitHub Actions runner release workflow checks the official actions/runner releases each Monday. It ignores drafts and prereleases, validates the stable tag and exact Linux x64 and arm64 assets, downloads both archives, and calculates their SHA-256 checksums. When an update exists, it updates the three pins in runner/Dockerfile and the CI_FLEET_RUNNER_VERSION default in deploy/compose.yaml on the machine-managed automation/update-actions-runner branch. It creates or refreshes one pull request for normal review and validation. If all four pins are current, it makes no branch or pull-request change.
Run the same check on demand from Actions → Update GitHub Actions runner release → Run workflow. The workflow only proposes a repository change. It does not build or deploy a host update.
Controller and runner engine updates are also pinned. A merged private configuration change is applied with install-worker-controller.sh --upgrade; the host never follows a moving engine or configuration branch automatically. See Git-authored controller desired state.