Skip to content

feat: let a relaunch from failed force failed nodes as successful - #1075

Merged
cigamit merged 2 commits into
ctrliq:mainfrom
fernandorocagonzalez:feat-wf-force-node-success
Oct 9, 2026
Merged

cigamit merged 2 commits into
ctrliq:mainfrom
fernandorocagonzalez:feat-wf-force-node-success

Conversation

@fernandorocagonzalez

Copy link
Copy Markdown
Contributor

What this is for

Relaunch from the failed node reruns whatever failed. Most of the time that is exactly right. But there is a case we run into in real operations where rerunning the failed node is not an option, and right now the only ways around it leave no record of what was done.

A few examples from our own patching workflows:

  1. A non critical step fails at the end. The machine got patched fine, but the step that updates the Confluence page (or posts to Slack, or closes the ticket) failed because a token expired. The patching itself is done. The right long term fix is in the workflow design, but right now the run is red and the notification node after it never ran.
  2. The fix exists but cannot be used yet. A node is broken, the fix is written, and it is sitting in a bugfix branch waiting for review. Meanwhile the machine is half patched and the rest of the workflow has to go on.
  3. The fix is too costly for the moment. An external system the node depends on is down, or fixing the node properly takes hours, and the rest of the run does not depend on it.

What people do today in these cases is edit the workflow to drop the node and put it back later, or launch the remaining job templates by hand. Both work, and both hide the failure: nothing in the run says that a step was skipped, who decided it, or why.

This PR adds a way to do it that leaves that record. It does not replace fixing the node. It makes skipping it explicit, visible and attributable, instead of something that happens off the books.

Other orchestrators have the same thing for the same reason. Airflow has "Mark Success" on a failed task, and it is one of the most requested features for GitHub Actions.

The setting

Workflow job templates get allow_force_node_success_on_relaunch, off by default. Nothing changes for any workflow unless whoever administers the template turns it on. That way each team decides per workflow whether this is acceptable. A compliance workflow can leave it off, and a patching workflow with a notification step at the end can turn it on.

Unlike allow_overwrite_flow_vars_on_relaunch, this one is read from the template at relaunch time, the same way prevent_relaunch is checked in #1067. The usual situation is that a run just failed on a node nobody can fix in time, so the setting has to work for that run as soon as it is turned on. It also has to stop working for every run, old ones included, as soon as it is turned off. The run's own copy of the setting is only used once its template has been deleted.

How it works

POST /api/v2/workflow_jobs/N/relaunch/
{
  "nodes": "failed",
  "force_success_nodes": [123],
  "force_success_reason": "Confluence token expired, the patching itself went fine"
}

The nodes listed are carried forward into the new run as though they had succeeded. They are not run again, and the workflow goes on down their success paths. Everything else behaves exactly like a normal relaunch from failed.

This builds on the prior_run_succeeded carry forward that relaunch from failed already uses, so the scheduler is not touched at all. Parallel branches, ALL convergence and conditional connectors behave as they do for any carried node.

Rules:

  • A reason is required. Blank or missing is a 400.
  • Only nodes whose job failed in that run can be forced. A node that succeeded, was already carried, or belongs to another run is refused.
  • Approval nodes can never be forced. Forcing a denied or timed out approval would be a way around the people who were asked to approve.
  • Only together with nodes=failed.
  • If the template no longer has the node, or its identifier changed since the run, the relaunch is refused instead of silently running the node again.

What stays on record

Each forced node in the new run keeps:

  • forced_success: true
  • forced_success_reason: the reason given
  • forced_success_by: who forced it
  • forced_success_job: the job that actually failed

The node also stays marked if the run is relaunched again later. If a forced run fails further down and you relaunch from that failure, the forced node is still shown as forced, with the original reason and pointing to the original failed job. Several nodes can be forced across relaunches, and each keeps its own reason.

Whatever the failed job managed to publish with set_stats before failing is passed down, as a successful job's would be. Anything it did not get to publish can be supplied with extra_vars in the same request when allow_overwrite_flow_vars_on_relaunch is on (#1050).

In the UI

When the template allows it, the relaunch dropdown gets one more entry:

Relaunch from:
  First node
  Failed node
  Failed node, new variables          (when #1050's setting is on)
  Failed node, forced as successful   <- new, only when the template allows it

It opens a modal that lists the failed nodes of the run (approvals left out, a single failed node comes ticked) and a reason field. Relaunch stays disabled until there is a reason.

In the workflow output a forced node is neither green nor red. Its border is split diagonally half green and half red, and its status icon is split the same way, so it never reads as a plain success. Hovering it shows who forced it and why, and clicking it opens the job that failed.

The template form gets the checkbox under Options, next to the overwrite variables one, and the template detail lists it.

Tests

Backend: forcing a node and the record it leaves, the success path taken instead of the failure path, artifacts handed down, the marker kept across relaunches, a forced run failing further down and being relaunched again (plain and forcing the second node too), and every refusal: setting off, setting read live from the template in both directions, missing or blank reason, no nodes=failed, malformed ids, a node of another run, a node that did not fail, an already forced node, an approval, and a node the template lost. Also forcing together with overwriting variables.

UI: the dropdown entry only when allowed, the modal (reason required, single node preselected, approvals left out, nothing to force), the split node and where it links, the tooltip including when the template or the failed job was deleted, and the template form.

Full suite run in tools_ascender_1 following .github/copilot-instructions.md: ruff clean, check_migrations clean, 4608 passed. Migration 0226 adds the setting to both workflow tables and the four fields to WorkflowJobNode. The new UI strings are translated in every locale.

@fernandorocagonzalez

Copy link
Copy Markdown
Contributor Author

Some screencaps

03-relaunch-menu 04-force-modal 05-template-option 01-run-with-forced-nodes 02-forced-node-tooltip

Sometimes a node fails for a reason that has nothing to do with what comes
after it (a Confluence page that did not update at the end of a patching
run, say) and the fix cannot land in time: it is sitting on a bugfix
branch, or it is simply too costly to do right now. Today the only ways
around that are editing the workflow to drop the node and putting it back
later, or launching the remaining templates by hand. Both leave no trace
of the failure that was skipped.

This adds a third way that does leave one. A relaunch from failed nodes can
name failed nodes to carry forward as though they had succeeded, with a
reason that is always required. The new run keeps them marked as forced,
with who forced them, why, and the job that failed, and they stay marked
if that run is relaunched again.

It is off unless the workflow job template turns on
allow_force_node_success_on_relaunch. Unlike the variables overwrite, the
setting is read from the template at relaunch time, as prevent_relaunch
is, so turning it off also stops it for runs already started. Approvals
can never be forced.

It builds on prior_run_succeeded, so the scheduler is untouched. In the
UI the relaunch menu gets a "Failed node, forced as successful" entry
that asks which nodes and why, and a forced node is drawn half green,
half red, with who and why on hover and a click through to the job that
failed.
@cigamit
cigamit force-pushed the feat-wf-force-node-success branch from 89b9fd9 to 3a1a355 Compare October 7, 2026 19:49
@cigamit cigamit self-assigned this Oct 7, 2026
@cigamit cigamit added the enhancement New feature or request label Oct 7, 2026
@cigamit
cigamit requested a balanced review from Copilot October 7, 2026 20:01

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The modal truncates workflows beyond 200 matching nodes, and canceled-node behavior contradicts the documented failed-only contract.

2 open findings
What changed in this PR

Adds auditable forced-success handling when relaunching failed workflow nodes.

Changes:

  • Adds backend persistence, validation, API support, and migration.
  • Adds template controls, relaunch modal, and forced-node visualization.
  • Adds backend/UI tests, generated types, documentation, and translations.
File Description
ascender/​ui/​src/​util/​jobs.ts Determines whether forced-success relaunch is allowed.
ascender/​ui/​src/​util/​jobs.test.ts Tests current-template and fallback behavior.
ascender/​ui/​src/​types/​api.generated.ts Adds generated API fields.
ascender/​ui/​src/​screens/​Template/​WorkflowJobTemplateEdit/​WorkflowJobTemplateEdit.test.tsx Updates template form expectations.
ascender/​ui/​src/​screens/​Template/​WorkflowJobTemplateDetail/​WorkflowJobTemplateDetail.tsx Displays the enabled option.
ascender/​ui/​src/​screens/​Template/​shared/​WorkflowJobTemplateForm.tsx Adds the template checkbox.
ascender/​ui/​src/​screens/​Template/​shared/​WorkflowJobTemplate.helptext.tsx Documents the option in-form.
ascender/​ui/​src/​screens/​Job/​WorkflowOutput/​WorkflowOutputToolbar.tsx Enables the relaunch choice in workflow output.
ascender/​ui/​src/​screens/​Job/​WorkflowOutput/​WorkflowOutputNode.tsx Renders and links forced nodes.
ascender/​ui/​src/​screens/​Job/​WorkflowOutput/​WorkflowOutputNode.test.tsx Tests forced-node presentation and navigation.
ascender/​ui/​src/​screens/​Job/​WorkflowOutput/​WorkflowOutputNode.css Styles the split-status icon.
ascender/​ui/​src/​locales/​zh/​messages.po Adds Chinese translations.
ascender/​ui/​src/​locales/​en/​messages.po Adds English message catalog entries.
ascender/​ui/​src/​components/​Workflow/​workflowReducer.ts Extends workflow node typing.
ascender/​ui/​src/​components/​Workflow/​WorkflowNodeHelp.tsx Shows forced-success audit details.
ascender/​ui/​src/​components/​Workflow/​WorkflowNodeHelp.test.tsx Tests audit-detail tooltips.
ascender/​ui/​src/​components/​Workflow/​WorkflowNodeHelp.css Styles long reason text.
ascender/​ui/​src/​components/​LaunchButton/​WorkflowRelaunchForceSuccessModal.tsx Adds node selection and reason capture.
ascender/​ui/​src/​components/​LaunchButton/​WorkflowRelaunchForceSuccessModal.test.tsx Tests modal behavior.
ascender/​ui/​src/​components/​LaunchButton/​WorkflowReLaunchDropDown.tsx Adds the forced-success relaunch entry.
ascender/​ui/​src/​components/​LaunchButton/​WorkflowReLaunchDropDown.test.tsx Tests relaunch payload construction.
ascender/​ui/​src/​components/​JobList/​JobListItem.tsx Enables the option from job lists.
ascender/​main/​tests/​functional/​models/​test_workflow.py Tests model-level carry-forward behavior.
ascender/​main/​tests/​functional/​api/​test_workflow_job.py Tests API validation and persistence.
ascender/​main/​models/​workflow.py Implements settings, audit fields, and carry-forward logic.
ascender/​main/​migrations/​0226_workflow_force_node_success.py Adds database fields.
ascender/​api/​views/​workflow.py Validates and processes relaunch requests.
ascender/​api/​templates/​api/​workflow_job_relaunch.md Documents the new parameters.
ascender/​api/​serializers/​workflow.py Exposes settings and audit fields.
ascender/​api/​serializers/​base.py Adds related summary data.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread ascender/main/models/workflow.py
The modal only asked for the first 200 failed nodes, so on a big run the
rest could not be picked, and a first page of nothing but approvals made
it look as though there was nothing to force. It now follows next like
the workflow output loaders do.

The docs said only failed nodes can be forced, but canceled ones always
could and the UI offers them on purpose. Say so in the endpoint doc and
in the docstring.
fernandorocagonzalez added a commit to fernandorocagonzalez/ascender that referenced this pull request Oct 7, 2026
Listing ad hoc commands read the instance groups of every row, then asked
whether the user could use them, to fill in user_capabilities.start. The
list querysets now annotate each command with whether it has a group the
user cannot use, and can_start reads that. django-polymorphic carries
annotations over to the real instances, so the unified jobs list gets it
too. A prefetch would not have helped, since the ordered m2m manager
orders the queryset again and skips the prefetch cache.

Also move the migration to 0227, 0226 is taken by ctrliq#1075.
@cigamit

cigamit commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

One additional thing that was caught, was this, but I don't consider it a blocker so I will go ahead and approve and merge, and can worry about that in a new PR.

  • One extra query per forced node on node lists, for each of the two new FKs. It's minor because forced nodes are rare; adding them to the node list's prefetch_related would remove it.

@cigamit
cigamit merged commit 90de5ac into ctrliq:main Oct 9, 2026
fernandorocagonzalez added a commit to fernandorocagonzalez/ascender that referenced this pull request Oct 10, 2026
Listing ad hoc commands read the instance groups of every row, then asked
whether the user could use them, to fill in user_capabilities.start. The
list querysets now annotate each command with whether it has a group the
user cannot use, and can_start reads that. django-polymorphic carries
annotations over to the real instances, so the unified jobs list gets it
too. A prefetch would not have helped, since the ordered m2m manager
orders the queryset again and skips the prefetch cache.

Also move the migration to 0227, 0226 is taken by ctrliq#1075.
@fernandorocagonzalez

Copy link
Copy Markdown
Contributor Author

Thanks! The prefetch is in #1103. While at it I found retried_jobs was doing one query per node on the same lists, so that one is prefetched there too.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Development

Successfully merging this pull request may close these issues.

3 participants