Restore your workflow to a consistent state after a failed run, including how to identify partial side effects, safely re-run failed steps, and prevent double-processing of already-completed actions.
Assessing the impact of a partial run
When an automation run fails mid-execution, some steps will have already produced side effects — records created, emails sent, API calls made — while later steps did not execute. Before attempting any recovery action, it is essential to understand which steps completed successfully. Open the failed run record in the Activity tab and review the step trace. Steps marked with a green checkmark completed and produced side effects; steps marked with a red icon or shown as 'Not reached' did not execute. Make a list of the side effects produced so far — this is your baseline for deciding how to proceed.
Using re-run from step
Cotonity's 'Re-run from step' feature allows you to resume a failed run starting from a specific step, using the same input data that was recorded during the original execution. This is the safest recovery mechanism because it avoids re-executing steps that already completed successfully, preventing double-processing side effects such as sending duplicate emails or creating duplicate records. To use it, open the failed run, hover over the first step that did not complete, and click 'Re-run from here'. Confirm the resume point is correct before clicking Start. The resumed run will appear as a linked child run in the Activity log so you can track it separately from the original.
Handling non-idempotent completed steps
Some completed steps cannot simply be re-run because their side effects are not safely repeatable. For example, if a 'Send email' step completed successfully before the workflow failed, running the workflow from an earlier step would send the email a second time. For workflows that contain non-idempotent actions, add an idempotency check before each such step: query a log or status field to confirm the action has not already been performed, and skip the step if it has. This pattern should ideally be built into the workflow proactively rather than added after a failure, but it can be retrofitted before re-running the workflow to recover from the current incident.
Rolling back partial state when re-run is not possible
In some scenarios, resuming from the point of failure is not the right approach — for example, if the workflow created a corrupted or incomplete record that would cause downstream problems if the workflow were resumed. In these cases, the best path is to roll back the partial side effects manually and restart the full run from scratch. Use the workflow's linked CRM integration to find and delete or correct the records created during the failed run, then manually retrigger the workflow with the original input. For high-volume workflows where manual cleanup is impractical, consider adding a 'transaction' pattern: write records to a staging area first, and only promote them to the production dataset after all steps have completed successfully.