Skip to content

feat(preview): expire preview workspaces and prune expired ones on cleanup - #132

Open
toiroakr wants to merge 5 commits into
mainfrom
feat/preview-ttl-prune
Open

toiroakr wants to merge 5 commits into
mainfrom
feat/preview-ttl-prune

Conversation

@toiroakr

@toiroakr toiroakr commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A preview workspace whose close-time cleanup never ran (the job failed or was disabled, or the PR stayed open) was never deleted. This lets preview workspaces expire and lets every PR close sweep the expired ones.

  • preview-deploy takes ttl (e.g. 7d). The workspace records an expiry when created, and every later push restarts it, so it counts from the last push.
  • preview-cleanup takes prune-expired, organization-id and folder-id. After deleting the per-PR workspace it runs tailor workspace prune --expired.
  • A PR that sat idle past ttl has its workspace swept while it is still open. The next push now creates a fresh workspace under the same name instead of deploying to the deleted one, and closing such a PR no longer fails.
  • Separately, the workspace ID lookup in both actions used (?P<id>…) in a jq capture, which the jq 1.7.1 and 1.8.2 release binaries reject. try hid the error, so an existing workspace went unrecognized. It is now (?<id>…).

Both new inputs are off by default, so existing workflows behave as before. Related to issue #1940 in platform-planning.

Usage

- uses: tailor-platform/actions/preview-deploy@v2
  with:
    workspace-name-prefix: my-app
    region: us-west
    folder-id: ${{ vars.TAILOR_PLATFORM_FOLDER_ID }}
    ttl: 7d
    # ...

- uses: tailor-platform/actions/preview-cleanup@v2
  with:
    workspace-name-prefix: my-app
    prune-expired: "true"
    folder-id: ${{ vars.TAILOR_PLATFORM_FOLDER_ID }}
    # ...

Behavior worth reviewing

  • Scope of the sweep. Only workspaces in the given folder (or, without folder-id, directly under the organization) whose whole name is {prefix}-pr-{number}, with the prefix escaped. Other apps sharing the location are untouched. The folder wins over the organization, as in preview-deploy. A workspace without a recorded expiry is never deleted.
  • --limit 0. workspace prune aborts without deleting anything when more than its default 20 workspaces match. With that, a backlog of 21 would fail every later PR close. The name and location filters keep the sweep narrow.
  • No location known. With none of folder-id, organization-id, TAILOR_PLATFORM_FOLDER_ID or TAILOR_PLATFORM_ORGANIZATION_ID (the two variables workspace create itself reads), the sweep warns and is skipped instead of failing, because prune --expired rejects an empty location and generated workflows pass an empty string for an unset variable. This is a provisional rule.
  • Order and failure. The sweep is the last step and runs under !cancelled(), so a failed deletion is still swept and a failed sweep cannot undo the deletion or the comment update.
  • Recreating a pruned workspace. preview-deploy runs workspace get on the recorded ID. Not found means pruned, so it creates a new workspace; the PR comment is updated with the new ID. Any other failure stops the step so a transient error cannot create a duplicate. If workspace create --ttl prints the new workspace and then exits nonzero because it could not confirm the expiry, the ID is still taken from its output and the expiry is set again, so the workspace is not left untracked. Workspace names are not unique on the platform, so reusing the name is fine.
  • Not found is matched on CLI output (not_found or workspace … not found), because workspace get surfaces the raw Connect error without a stable code. A permission loss also reports not found.

Verification

  • pnpm test:preview runs the real step scripts from the action files with a mocked gh and tailor, and is added to CI. Existing test:cli-bin-contract and test:deploy-outputs still pass; zizmor and ghalint are clean.
  • The jq fix was checked against the jq 1.7.1 and 1.8.2 release binaries; distribution builds were not checked.
  • The sed that escapes the prefix was only run with BSD sed locally. The CI run on ubuntu covers GNU sed.
  • Pruning against a real workspace was not run.

preview-deploy and preview-cleanup extract the workspace ID from the PR
comment with `capture("...(?P<id>...)")` inside `try`. jq 1.8 rejects the
`(?P<name>)` group syntax ("undefined group option"), and `try` turns that
into an empty result, so the comment looked like it held no ID: every push
would create another workspace and closing the PR would skip deleting it.

`(?<name>)` is accepted by both jq 1.7 and 1.8. The tests run the real
step scripts with a mocked gh and tailor, and are wired into CI as
`test:preview`.
…n cleanup

A preview workspace whose close-time cleanup never ran (the job failed
or was disabled, or the PR stayed open) was never deleted. preview-deploy
now takes `ttl`, passed to `workspace create --ttl` so the workspace
records an expiry, and preview-cleanup takes `prune-expired`, which runs
`workspace prune --expired` after deleting the per-PR workspace.

The sweep is scoped to the folder (or, without one, the organization
root) the previews are created in, and to names matching
`<prefix>-pr-<number>` with the prefix escaped, so other apps sharing the
location are untouched. It passes `--limit 0`: prune aborts without
deleting anything when more than its default 20 workspaces match, which
would fail every later PR close once a backlog that size built up.

It is the last step and runs under `!cancelled()`, so a failed deletion
still gets swept and a failed sweep cannot undo the deletion or the
comment update. With no location known it warns and skips rather than
failing every PR close, because `prune --expired` rejects an empty
location.

Redeploys do not extend the expiry; it counts from creation. Both inputs
default to off, so existing workflows behave as before.
The changeset said jq 1.8 rejects `(?P<id>)`. The jq 1.7.1 release binary
rejects it too, so name both versions, and rename the file to match.
With `ttl` and `prune-expired`, a PR that stays open past `ttl` has its
workspace deleted by the sweep. The PR comment still records the old ID,
and preview-deploy reused it without checking, so every later push
deployed to a deleted workspace. Closing such a PR also failed in
preview-cleanup, because deleting the missing workspace is NotFound.

preview-deploy now checks that the recorded workspace exists. If so it
restarts the expiry with `workspace ttl set`, so the expiry counts from
the last push; a failure to do that only warns. If `workspace get` says
not found it creates a new workspace under the same name (names are not
unique on the platform) and the comment is updated with the new ID. Any
other failure of the check stops the step instead of risking a duplicate.

preview-cleanup treats a not-found deletion as already deleted, warns,
and still reports the workspace so the comment is updated.

NotFound is matched on the CLI output because `workspace get` surfaces the
raw Connect error without a stable code.

This comment was marked as resolved.

…, and read the folder env

`workspace create --ttl` prints the new workspace as JSON and then exits
nonzero with WORKSPACE_TTL_WRITE_FAILED when it cannot confirm the expiry.
preview-deploy read the ID through a pipeline under pipefail, so that exit
aborted the step before the ID was saved: the workspace was left untracked
and the next push created another one. It now keeps the exit status and
the output separately, takes the ID from the output, warns, and sets the
expiry again with `workspace ttl set`. With no ID in the output it still
fails.

`workspace create` reads its folder from TAILOR_PLATFORM_FOLDER_ID when no
folder is passed, so preview-deploy can place a preview in a folder while
preview-cleanup saw no location and skipped the sweep. preview-cleanup now
falls back to the same variable, as it already did for the organization.

This comment was marked as off-topic.

@toiroakr
toiroakr marked this pull request as ready for review October 5, 2026 13:14
@toiroakr
toiroakr requested review from a team as code owners October 5, 2026 13:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants