Skip to content

Idempotency

Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.

Reliability

Learn it

0 of 2 checks done
  1. Networks drop responses, clients retry, queues redeliver, workers restart. Every retry risks doing something twice: charging, emailing, inserting. You can't reliably prevent the duplicate attempt, so you make it harmless.

    An operation is idempotent if doing it twice has the same effect as once. SET status = 'shipped' is; balance = balance - 10 isn't.

  2. Check

    Which of these is naturally idempotent?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Check-then-insert
Two concurrent requests both see no key and both proceed. The claim must be atomic.
Key per request, not per intent
Generating a new key on each retry defeats the purpose.
Key expires too soon
A late retry, after the record has been purged, executes again.
Effect not covered
The key is recorded, but the side effect (an email, a downstream call) can still repeat.

Instead, consider

Natural keys and unique constraints
The intent already has a unique identity, such as one enrollment per (user, course).
Conditional state transitions
The operation is a state change that can be guarded by the current state.
Accept duplicates and reconcile
Duplicates are cheap and rare, and periodic cleanup is simpler than prevention.

In practice

Idempotency-Key HTTP header
Stripe-style: the server stores the response by key for a retention period.
Unique constraint + ON CONFLICT
The database performs the atomic claim.
Message ID dedupe table
For consumers of at-least-once queues and webhooks.
Conditional writes (compare-and-set, ETags)
Effects that only apply from an expected prior state.

It assumes

  • Callers generate the key once per intent and reuse it across retries.
  • The store that records keys supports an atomic insert-if-absent.
  • Downstream systems either accept idempotency keys or their operations are naturally idempotent.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Delivery guarantees

    At-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.

  • Retries, backoff and jitter

    Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.

  • State machines for business state

    Modelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.

  • Concurrency control

    Making read-decide-write sequences safe when other actors may change the same data in between: locks, conditional writes and constraints.

  • Timeouts and unknown outcomes

    A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.

  • Asynchronous processing

    Accepting a request, recording the work durably, and doing it later in another process, so the work can outlive the request.

  • Webhooks

    HTTP callbacks from another system: delivered at least once, possibly out of order, possibly never. Handle them as hints, not truth.

  • Transactions

    Grouping several reads and writes so they take effect all together or not at all, isolated from concurrent work to a defined degree.

  • Transactional outbox

    Recording outgoing messages in the same database transaction as the state change, then delivering them separately, to avoid the dual-write problem.

  • Conflict resolution and convergence

    When replicas accept concurrent changes, a deterministic rule must merge them so every replica ends in the same state without losing intent.

  • Leases and fencing tokens

    Ownership that expires unless renewed, plus a token that lets the rest of the system reject an owner that has lost its claim without knowing it.