Every integration eventually retries something. A network call times out, a partner API returns an error, a message broker redelivers a message, or an operator restarts a failed job. Whether that retry is harmless or creates a duplicate order, a second payment or a repeated email depends on one property: idempotency. An idempotent operation has the same effect whether it runs once or many times.
This guide explains how to design integrations that can be retried safely, from HTTP semantics and idempotency keys to idempotent message consumers and sensible retry policies. It draws on RFC 9110, AWS's Builders' Library, Stripe's API documentation and Microsoft's Azure Architecture Center.
The core problem: a timeout is not a failure
When a call times out, the caller does not know what happened. Microsoft's Retry pattern guidance describes the classic case: a service receives the request, processes it successfully, but fails to send a response. The retry logic then re-sends the request, assuming the first one was never received. If the operation is "create an invoice", you now have two.
You cannot eliminate this uncertainty; networks and processes fail between steps. What you can do is design operations so that repeating them is safe, and then retry with confidence.
Start with HTTP semantics
RFC 9110 defines a request method as idempotent when the intended effect on the server of multiple identical requests is the same as for a single request. Of the standard methods, PUT, DELETE and the safe methods (GET, HEAD, OPTIONS, TRACE) are idempotent. POST is not. The RFC notes that idempotent requests can be automatically retried when a connection fails before the client reads the response.
This suggests a first design move: where the business operation allows it, model writes as PUT to a resource identifier chosen by the client ("create or replace order 12345") rather than POST to a collection ("create a new order"). If the caller supplies the identifier, a retry lands on the same resource instead of creating a new one.
Idempotency keys for operations that must be POST
Many operations cannot be expressed as a simple PUT, such as initiating a payment or submitting a job. For these, the widely used pattern is an idempotency key: a unique value generated by the client for each logical operation and sent with every attempt of that operation.
Stripe's idempotent requests documentation is a clear public reference for how this works in practice:
- The client sends an
Idempotency-Keyheader, generated as a V4 UUID or another random string with enough entropy to avoid collisions. - The server saves the status code and body of the first request for that key, whether it succeeded or failed, and returns the same result for later requests with the same key — including
500errors. - The server compares incoming parameters with the original request and returns an error if they differ, to prevent accidental misuse of a key.
- Keys can be removed after they are at least 24 hours old; a reused key after that point is treated as a new request.
- Sensitive data such as email addresses should not be used as keys.
AWS describes the same idea from the service side in Making retries safe with idempotent APIs. Callers provide a client request identifier, and any subsequent response to a retry with the same identifier should have the same meaning as the first response, so client code does not need to know that a retry occurred. If the same identifier arrives with different parameters, AWS treats it as a likely mistake and returns a validation error rather than guessing.
Design rules for idempotency keys
- Generate the key once per logical operation, before the first attempt, and persist it with the pending work so retries after a crash reuse it.
- Store the key and the outcome atomically with the business change on the server side.
- Return the original outcome on a repeat, not an "already exists" error.
- Reject parameter mismatches explicitly.
- Set a retention window longer than your longest realistic retry or replay.
Idempotent consumers for queues and events
Message-driven integrations face the same problem from the other direction. Microsoft's Idempotent Consumer pattern notes that most brokers, including Azure Service Bus, Apache Kafka and RabbitMQ, provide at-least-once delivery, and that exactly-once delivery is impractical to guarantee across a distributed system. Duplicates arise from producer retries, redelivery after a lost acknowledgement, and consumers that crash after writing but before acknowledging.
The pattern's answer is to make the consumer tolerate duplicates:
- Extract a stable deduplication key from each message — a producer-assigned message ID or a business-level key, not a transport identifier that changes on redelivery.
- Check a deduplication store for that key.
- If it is already there, acknowledge the message and stop.
- If not, process the message and record the key in the same transaction as the business change, then acknowledge.
Two details from that guidance are easy to get wrong. First, a separate "check, then process, then record" sequence has a failure window; the marker and the side effect must be committed together. Second, concurrent consumers can both pass an existence check, so enforce uniqueness in the data store with a unique constraint or an atomic insert-if-absent, rather than in application logic.
When a side effect cannot join the transaction, such as a call to a third-party API, the guidance suggests recording the key as in progress, performing the call, and then marking it completed with the outcome. A redelivery that finds an in-progress record needs reconciliation rather than a blind retry.
Finally, keep deduplication records for at least as long as the broker can redeliver a message, and account for messages that operators resubmit from dead-letter queues, which can arrive long after the normal window.
Prefer naturally idempotent operations
Deduplication bookkeeping is useful, but the simplest integrations avoid needing it. Microsoft's guidance recommends designing for natural idempotency where possible:
- Upsert on a business key instead of blind inserts.
- Set absolute values ("status = shipped", "balance = 120") instead of relative changes ("add 10").
- Carry state in events so a consumer can apply the latest state rather than replaying a sequence of deltas.
- Include version or sequence numbers so stale or out-of-order messages can be rejected.
Retry policy: when and how to try again
Idempotency makes retries safe; a retry policy makes them sensible. Microsoft's Retry pattern describes three responses to a failure: cancel if the fault is not transient, retry immediately for rare one-off faults, or retry after a delay for common connectivity or "busy" faults, with increasing (incremental or exponential) delays up to a maximum number of attempts. It recommends spacing retries so that requests from multiple instances are spread out rather than hitting a struggling service at the same moment — which is the purpose of adding random jitter to backoff.
Practical rules for business integrations, most of them drawn from the same pattern:
- Classify errors. The pattern advises cancelling when a fault is not transient. Validation failures and authorisation errors will not succeed on retry; fail fast and route them for attention.
- Honour server hints. If an API returns a
Retry-Afterheader (defined in RFC 9110) with a throttling or unavailable response, wait at least that long before trying again. - Don't stack retries. If a task with a retry policy calls another task that also retries, delays multiply. Let lower layers fail fast and retry at one level that understands the full context.
- Log sensibly. Log early failed attempts as informational and only the final failure as an error, so operators are not flooded with alerts for operations that succeeded on retry.
- Stop hammering. After repeated failures, a circuit breaker should stop sending requests for a period and then test cautiously.
How to test that an integration is retry-safe
Idempotency is easy to claim and hard to verify without deliberate tests. Build these into your integration test suite:
- Lost response: make the server commit the operation, then drop the response. Retry and confirm exactly one side effect.
- Duplicate delivery: deliver the same message twice to the consumer and confirm one outcome.
- Concurrent duplicates: deliver the same message to two consumer instances simultaneously.
- Crash mid-process: kill the consumer after the business write but before acknowledgement, then restart.
- Parameter mismatch: reuse an idempotency key with different data and confirm it is rejected.
- Late replay: resubmit a dead-lettered message after the normal redelivery window.
These same properties are what make backlog replay safe after an outage — a central step in any business-continuity runbook for automated workflows. Credential rotation is another common source of transient authentication failures that retries must handle cleanly; see secrets management for automation scripts.
Checklist
- Writes use client-chosen identifiers or idempotency keys.
- Keys are generated once per logical operation and persisted before the first attempt.
- Servers return the original outcome for repeats and reject parameter mismatches.
- Consumers deduplicate on stable keys, atomically with the business change.
- Uniqueness is enforced in the data store, not only in code.
- Retries use capped exponential backoff with jitter and honour
Retry-After. - Non-transient errors fail fast; retries happen at one layer only.
- Duplicate, concurrent and crash scenarios are covered by tests.
If you are planning a new integration or hardening an existing one, our automation solutions team can help; contact us to discuss it.
Frequently asked questions
Is idempotency the same as exactly-once delivery?
No. Exactly-once delivery across a distributed system is impractical to guarantee. The practical goal is at-least-once delivery combined with idempotent processing, which Microsoft's guidance calls effectively-once processing.
Can I use the message's timestamp as a deduplication key?
No. Keys must be stable across redeliveries. Timestamps, delivery counts and broker-generated identifiers can change between attempts and defeat deduplication.
Does broker-level duplicate detection remove the need for idempotent consumers?
No. Features that discard repeated sends within a time window reduce duplicates from producer retries, but they do not prevent a consumer from processing a redelivered message twice. You still need idempotent consumer logic.
Need Help Hardening Your Automation?
Talk to our team about continuity, security and reliability for the workflows your business depends on.
Contact Us →