Agent error handling and retries

Updated: 2026-07-20Reading time: 4 min

Configure per-node retry policies, build a global error handler, ensure idempotent operations, and replay failed runs from the exact step that broke.

Per-node retry configuration

Every action node in Cotonity has a Retry Policy section in its configuration panel. You can set the maximum number of retry attempts (1–10), the initial delay between retries (in seconds), and a back-off multiplier that increases the delay exponentially on each subsequent attempt. For example, a policy of 3 retries, 2-second initial delay, and 2x multiplier will wait 2 s, then 4 s, then 8 s before marking the step as failed. Retries are triggered by any non-success response from the action: HTTP 5xx, a timeout, or a network error. Client errors (4xx) are not retried by default because they typically indicate a configuration problem that will not resolve itself.

Global run-level error handling

In addition to per-node retries, you can configure a Global Error Handler for the entire agent. This is a sub-workflow that Cotonity triggers whenever any step fails and is not handled by per-node retry. Common patterns for the error handler include: logging the error to a monitoring service, sending a Slack or email alert to your on-call team, storing the failed input in a dead-letter queue for manual review, and attempting a simplified fallback action. Configure the error handler in the agent's Settings tab; it receives the failing step's name, error message, and the full run context as inputs.

Idempotency and safe retries

Retrying a failed step is only safe if the step is idempotent — meaning running it twice produces the same result as running it once. For read-only steps (data fetches, lookups), this is almost always true. For write steps (creating records, sending emails, charging a payment), you need to ensure idempotency explicitly: use the Memory Store to record a unique run ID before executing the write, and check for that ID at the start of each retry to skip the write if it already succeeded. Many APIs support idempotency keys in request headers — consult the API's documentation and pass a stable key (such as the agent run ID) when available.

Inspecting and replaying failed runs

Every failed run is preserved in the agent's Run History with a Failed status. Click any failed run to open the Run Inspector, which shows the exact step that failed, the full error message, the inputs the step received, and the run log up to that point. From the Run Inspector, you can replay the failed run from the beginning (using the original trigger input) or from the failing step (using its recorded input). Replaying from a specific step is useful when the failure was caused by a transient issue — a brief API outage or a network timeout — and you want to resume without re-executing all the successful steps that preceded it.