Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions benchmarking/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,12 @@ scenario ladder, read [observability.md](observability.md).
> [!IMPORTANT]
> Source the environment configuration file (e.g., `source .ate-dev-env.sh`)
> first so `PROJECT_ID`, `BUCKET_NAME`, etc. are set.
>
> **For Local Kind Clusters**: Set `KO_DOCKER_REPO` to the local registry so images are pulled without remote cloud credentials:
> ```bash
> export KO_DOCKER_REPO="localhost:5001"
> export BUCKET_NAME="ate-snapshots"
> ```

Note that deploying the benchmarks does not run them. You must visit Locust's
web UI to start a test.
Expand Down Expand Up @@ -113,6 +119,11 @@ and state restoration latency when a durable directory is attached to the actor.
* `--durdir-template`: ActorTemplate name:
* `glutton-durdir-data` (default): Attaches a durable data directory without memory snapshot restore.
* `glutton-durdir-full`: Attaches a durable data directory and performs a full memory snapshot restore.
* `glutton-storage`: Attaches an external CSI volume instead of a durable
directory. `deploy_locust.sh` and `workloads/deploy.sh` deploy it only
when `--storage-class-name` names an existing StorageClass. To run it on
NFS, install with `hack/install-ate.sh --setup-csi=nfs` and pass
`--storage-class-name csi-nfs-sc`, the class the e2e tests also use.

#### DurDir Reported Metrics

Expand Down
25 changes: 22 additions & 3 deletions benchmarking/automation/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,9 +65,10 @@ name (the routing key in `tests.yaml`) and the GCP project ID of the
orchestrator image registry. It then builds + pushes the orchestrator image
to `gcr.io/<ORCH_PROJECT_ID>/ate-images/substrate-benchmark-orchestrator:<short-commit>`
(with a `-dirty` suffix if `benchmarking/automation/` has uncommitted
changes) and renders `scratch/cronjob.yaml`, `scratch/test-list.yaml`, and
`scratch/target-clusters.yaml`. You can edit the `.ate-dev-env.sh` for your
workload cluster directly in the config map.
changes) and renders `scratch/cronjob.yaml`, `scratch/test-list.yaml`,
`scratch/target-clusters.yaml`, and `scratch/test-manifests.yaml`. You can
edit the `.ate-dev-env.sh` for your workload cluster directly in the config
map.

Then edit the `--repo / --branch / --dest` args in `scratch/cronjob.yaml` and
apply:
Expand Down Expand Up @@ -168,3 +169,21 @@ gcloud storage buckets add-iam-policy-binding gs://<DEST_BUCKET> \
`tests.yaml` is delivered to the orchestrator via a ConfigMap mounted at
`/etc/orchestrator/tests.yaml`, so the image doesn't need to be rebuilt when
the test list changes. Just reapply the config map.

## Extra manifests for a test

A test can list Kubernetes manifest files in `additionalManifests`. The
orchestrator applies them after it deploys substrate and before the
workloads, and deletes them after the test. Use it for anything the test
needs on the cluster that the substrate install does not create, such as a
StorageClass for `storageClassName` on storage the install does not set up.

Relative paths resolve against the directory holding `tests.yaml`. `setup.sh`
builds the `substrate-benchmark-test-manifests` ConfigMap from the YAML files
in `test-manifests/`, and the CronJob sample mounts it at
`/etc/orchestrator/test-manifests`, next to `tests.yaml`, so a test lists them
as `test-manifests/<file>`. A test that lists a file that is not there fails.

The commented-out `durdir_external_volume_example` entry at the end of
`tests.yaml` runs the DurDir load on `csi-nfs-sc`, the NFS class that
`--setup-csi=nfs` creates, so it needs no extra manifest.
16 changes: 13 additions & 3 deletions benchmarking/automation/manifests/cronjob.yaml.tmpl
Original file line number Diff line number Diff line change
Expand Up @@ -73,16 +73,20 @@ spec:
- name: DOCKER_HOST
value: tcp://localhost:2375
volumeMounts:
# tests.yaml is mounted via subPath so the target-clusters
# ConfigMap can own the /etc/orchestrator/target-clusters
# directory next to it.
# tests.yaml is mounted via subPath so the target-clusters and
# test-manifests ConfigMaps can own directories next to it.
- name: tests-config
mountPath: /etc/orchestrator/tests.yaml
subPath: tests.yaml
readOnly: true
- name: target-clusters
mountPath: /etc/orchestrator/target-clusters
readOnly: true
# Files that tests.yaml entries list in additionalManifests, as
# test-manifests/<file>.
- name: test-manifests
mountPath: /etc/orchestrator/test-manifests
readOnly: true
resources:
requests:
cpu: "1"
Expand All @@ -96,3 +100,9 @@ spec:
- name: target-clusters
configMap:
name: substrate-benchmark-target-clusters
# Optional, so the CronJob runs without it; a test that lists a
# missing file fails.
- name: test-manifests
configMap:
name: substrate-benchmark-test-manifests
optional: true
61 changes: 55 additions & 6 deletions benchmarking/automation/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,13 +26,14 @@
the runner image build [hook: build_image] are cached per
target cluster.
b. Sweep leftovers, deploy substrate, let the type shape the
cluster [hook: pre_test], deploy workloads (+ microvm deps when
the sandbox class needs them).
cluster [hook: pre_test], apply the test's additionalManifests,
deploy workloads (+ microvm deps when the sandbox class needs
them).
c. Render the type's Job template [hooks: job_tmpl, job_subs],
submit it, wait, tail logs, delete the Job.
d. Tear substrate + workloads down again so tests don't pollute
each other, then drop the env file so the next test can't
inherit this one's cluster.
d. Tear substrate + workloads + additionalManifests down again so
tests don't pollute each other, then drop the env file so the
next test can't inherit this one's cluster.
4. Exit non-zero if any test didn't complete.
"""

Expand Down Expand Up @@ -307,6 +308,14 @@ def validate_and_normalize_tests(tests: list[dict[str, Any]]) -> None:
f"test {name!r} has invalid sandboxClass {sandbox_class!r} "
f"(want one of {list(SANDBOX_CLASSES)})"
)
manifests = t.get("additionalManifests", [])
if not isinstance(manifests, list) or not all(
isinstance(m, str) and m for m in manifests
):
raise ValueError(
f"test {name!r} has invalid additionalManifests {manifests!r} "
"(want a list of file paths)"
)
if "type" not in t:
raise ValueError(
f"test {name!r} missing required 'type' field "
Expand Down Expand Up @@ -349,6 +358,7 @@ def deploy_workloads(
actor_memory: str = "",
wait_timeout_secs: int | str = "",
worker_memory: str = "",
storage_class_name: str = "",
) -> None:
cmd = [
"benchmarking/workloads/deploy.sh",
Expand All @@ -366,6 +376,11 @@ def deploy_workloads(
# control how many workers the scheduler packs onto a node.
if worker_memory:
cmd += ["--worker-memory", worker_memory]
# Storage suites set storageClassName in tests.yaml, which makes deploy.sh
# add the glutton-storage template; without it, the template is not
# deployed.
if storage_class_name:
cmd += ["--storage-class-name", storage_class_name]
# Empty keeps deploy.sh's own default; large fleets set workerWaitTimeout
# (whole seconds).
if wait_timeout_secs != "":
Expand All @@ -378,6 +393,30 @@ def teardown_workloads() -> None:
run_no_check(["benchmarking/workloads/deploy.sh", "--delete"])


def additional_manifests(test: dict[str, Any], tests_dir: Path) -> list[Path]:
"""The test's additionalManifests. Relative paths resolve against
tests_dir, the directory holding tests.yaml."""
return [tests_dir / m for m in test.get("additionalManifests", [])]


def apply_additional_manifests(paths: list[Path]) -> None:
"""Check every file first, so a missing one fails the test before
anything is applied."""
for p in paths:
if not p.is_file():
raise FileNotFoundError(f"additional manifest {p} not found")
for p in paths:
run(["kubectl", "apply", "-f", str(p)])


def delete_additional_manifests(paths: list[Path]) -> None:
"""Delete in the reverse order of apply. A missing file has nothing on
the cluster to delete, because apply checks every file first."""
for p in reversed(paths):
if p.is_file():
run_no_check(["kubectl", "delete", "--ignore-not-found", "-f", str(p)])


def run_test(
test: dict[str, Any],
image: str,
Expand Down Expand Up @@ -476,6 +515,7 @@ def main() -> None:
print(f"Building commit {commit}", flush=True)

tests = yaml.safe_load(Path(args.tests).read_text())["tests"]
tests_dir = Path(args.tests).absolute().parent
print(f"Running {len(tests)} test(s)", flush=True)
try:
validate_and_normalize_tests(tests)
Expand Down Expand Up @@ -524,8 +564,11 @@ def main() -> None:
# into this run. All teardowns use --ignore-not-found, so
# this is cheap on a clean cluster. Order matters:
# microvm-deps deletes a SandboxConfig CR, which requires the
# SandboxConfig CRD that teardown_substrate removes.
# SandboxConfig CRD that teardown_substrate removes. The same
# goes for a custom resource in additionalManifests.
manifests = additional_manifests(test, tests_dir)
teardown_workloads()
delete_additional_manifests(manifests)
teardown_microvm_deps()
teardown_substrate()

Expand All @@ -540,12 +583,17 @@ def main() -> None:
# deploy_workloads needs the microvm SandboxConfig.
if sandbox_class == "microvm":
install_microvm_deps()
# After substrate, so a manifest can use its CRDs; before the
# workloads, so a StorageClass exists when glutton-storage is
# created.
apply_additional_manifests(manifests)
deploy_workloads(
test.get("workerCount", 1),
sandbox_class,
test.get("actorMemory", ""),
test.get("workerWaitTimeout", ""),
test.get("workerMemory", ""),
storage_class_name=test.get("storageClassName", ""),
)
try:
status = run_test(
Expand All @@ -569,6 +617,7 @@ def main() -> None:
# microvm-deps must go before substrate for the same reason
# as above.
teardown_workloads()
delete_additional_manifests(manifests)
teardown_microvm_deps()
teardown_substrate()
duration = time.time() - start_time
Expand Down
26 changes: 23 additions & 3 deletions benchmarking/automation/setup.sh
Original file line number Diff line number Diff line change
Expand Up @@ -22,13 +22,14 @@
# .ate-dev-env.sh again if needed) to add another target cluster.
# 2. Build & push the orchestrator image to the orchestrator GCP project's
# registry (the only value not derivable from .ate-dev-env.sh).
# 3. Generate three independent manifests in scratch/ so each can be
# 3. Generate four independent manifests in scratch/ so each can be
# re-applied without touching the others:
# - scratch/cronjob.yaml (Namespace + ServiceAccount + CronJob)
# - scratch/test-list.yaml (substrate-benchmark-tests ConfigMap)
# - scratch/target-clusters.yaml (substrate-benchmark-target-clusters ConfigMap)
# - scratch/test-manifests.yaml (substrate-benchmark-test-manifests ConfigMap)
#
# Re-running this script overwrites those three files and overwrites
# Re-running this script overwrites those four files and overwrites
# scratch/target-clusters/<name>.sh for the prompted name, but leaves other
# files in scratch/target-clusters/ intact.
#
Expand Down Expand Up @@ -140,6 +141,24 @@ kubectl create configmap substrate-benchmark-target-clusters \
> "${SCRATCH_DIR}/target-clusters.yaml"
echo "Wrote ${SCRATCH_DIR}/target-clusters.yaml"

# 4. test-manifests.yaml: the files tests.yaml entries list in
# additionalManifests (one key per test-manifests/*.yaml, if any). The
# CronJob mounts it next to tests.yaml, so test-manifests/<file> resolves
# the same way it does in this repo. Update + re-apply this file alone to
# change the manifests without touching tests or the CronJob.
manifest_args=()
shopt -s nullglob
for f in "${AUTOMATION_DIR}"/test-manifests/*.yaml; do
manifest_args+=(--from-file="$(basename "${f}")=${f}")
done
shopt -u nullglob
kubectl create configmap substrate-benchmark-test-manifests \
--namespace=substrate-benchmark \
${manifest_args[@]+"${manifest_args[@]}"} \
--dry-run=client -o yaml \
> "${SCRATCH_DIR}/test-manifests.yaml"
echo "Wrote ${SCRATCH_DIR}/test-manifests.yaml"

echo
echo "=== Next steps ==="
echo "1. Edit ${SCRATCH_DIR}/cronjob.yaml and fill in the --repo / --branch / --dest args."
Expand All @@ -149,7 +168,8 @@ echo " ${TARGET_CLUSTERS_DIR}/${TARGET_CLUSTER_NAME}.sh and ${SCRATCH_DIR}/tar
echo "3. Apply everything to the orchestration cluster:"
echo " kubectl --context=<orchestration-cluster> apply -f ${SCRATCH_DIR}/cronjob.yaml \\"
echo " -f ${SCRATCH_DIR}/test-list.yaml \\"
echo " -f ${SCRATCH_DIR}/target-clusters.yaml"
echo " -f ${SCRATCH_DIR}/target-clusters.yaml \\"
echo " -f ${SCRATCH_DIR}/test-manifests.yaml"
echo
echo "To update just the test list (no image rebuild):"
echo " kubectl --context=<orchestration-cluster> apply -f ${SCRATCH_DIR}/test-list.yaml"
Expand Down
Loading
Loading