Skip to content

Delivery guarantees

At-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.

Communication

Learn it

0 of 2 checks done
  1. A sender transmits a message and waits for an acknowledgement. None comes. Either the message was lost, or it arrived and the acknowledgement was lost. The sender can't tell which. Every messaging system has to decide what to do about that.

  2. There are two honest choices:

    • At-most-once: never retry. Some messages are lost; none are duplicated. Fine for metric samples or cursor positions, where the next message replaces the last.
    • At-least-once: retry until acknowledged. None are lost; some are duplicated. This is what queues, webhooks and most RPC retries give you.
  3. Check

    Which guarantee suits 'Bob is typing…' indicators?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Dedupe window too short
IDs are forgotten before the sender stops retrying, so a late retry is processed again.
Dedupe not atomic with the effect
A crash between recording the ID and applying the effect loses or duplicates the message.
Assuming order
Retries reorder messages: message 2 can be applied before a retried message 1.

Instead, consider

At-most-once (fire and forget)
Loss is acceptable and newer data supersedes older, as with telemetry, presence and cursor positions.
Transactional exactly-once within one system
Consumer and producer state live in the same system (for example Kafka transactions for read-process-write between topics), and no external side effects occur.

In practice

Unique constraint on message ID
Insert-or-ignore in the same transaction as the effect.
Idempotency keys on APIs
The caller supplies the ID; the server stores the result under it.
Conditional state transitions
The state machine itself rejects repeats.
Broker dedupe windows
e.g. SQS FIFO deduplication: helpful, but bounded in time.

It assumes

  • Receivers can identify a message (an ID or a natural key) across redeliveries.
  • The dedupe record lives at least as long as the sender may keep retrying.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Uber's Real-Time Push Platform

    Uber · Madan Thangavelu and others · Post, Dec 2020

    Why polling was replaced, and the delivery problems push brought with it: resuming after a dropped connection, and knowing what actually arrived.

  • Idempotency

    Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.

  • Message queues

    A durable buffer between producers and consumers that hands each message to one consumer at a time and redelivers it unless acknowledged.

  • Webhooks

    HTTP callbacks from another system: delivered at least once, possibly out of order, possibly never. Handle them as hints, not truth.

  • Retries, backoff and jitter

    Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.

  • Transactions

    Grouping several reads and writes so they take effect all together or not at all, isolated from concurrent work to a defined degree.

  • Transactional outbox

    Recording outgoing messages in the same database transaction as the state change, then delivering them separately, to avoid the dual-write problem.

  • Soft state

    State that expires unless refreshed. It is cheap to keep, safe to lose, and right for presence, sessions and anything that describes the present moment.