Handle 429 Too Many Requests errors from external APIs called by Cotonity agents, including retry strategies, request throttling, batching, and caching to stay within rate limits.
Understanding rate limit responses
A 429 Too Many Requests response means your workflow has exceeded the number of API calls the external service allows within a given time window. Most APIs that return 429 errors include a Retry-After header specifying how many seconds to wait before retrying, and an X-RateLimit-Remaining header showing how many requests are left in the current window. Cotonity's HTTP request step captures both headers and surfaces them in the step's output panel. Understanding the API's rate limit structure — requests per second, per minute, or per day — is essential for choosing the right mitigation strategy.
Configuring automatic retries with backoff
Enable automatic retries on any HTTP request step that may encounter rate limits. In the step settings, open the Retry section, set the retry count to 3–5, and choose 'Exponential backoff' as the retry strategy. With exponential backoff, each successive retry waits twice as long as the previous one, which gives the external API time to reset its rate limit window. Enable 'Respect Retry-After header' so that when the API specifies an exact wait time, Cotonity uses that value instead of the backoff formula. Note that retries count toward your Cotonity plan's execution time limit, so very long wait times between retries may cause the overall run to time out.
Throttling request volume with delays
If your workflow loops over a large list and calls an external API in each iteration, the total request rate may exceed the API's limit even if individual requests are spaced out naturally. Add a Delay step inside the loop to introduce a fixed pause between iterations. The appropriate delay depends on the API's rate limit: for an API that allows 60 requests per minute, add a 1-second delay between iterations. This ensures you never exceed 60 requests per minute even if the loop runs faster than that. Alternatively, switch to a bulk endpoint if the external API supports it — a single bulk request that processes 50 records typically consumes far less of the rate limit budget than 50 individual requests.
Caching responses to reduce request volume
If multiple workflow steps or multiple runs call the same API endpoint with the same parameters, enabling response caching eliminates redundant requests entirely. Cotonity's HTTP request step supports two caching modes: run-level caching (shared within a single run) and cross-run caching (shared across all runs within a configurable TTL). Cross-run caching is particularly useful for reference data that changes infrequently — for example, fetching a list of product categories or exchange rates that are updated once per day. Enable cross-run caching in the step settings and set a TTL that aligns with the data's freshness requirements. Cached responses do not count against the external API's rate limit.