Delivery guarantees
At-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.
Communication
Learn it
A sender transmits a message and waits for an acknowledgement. None comes. Either the message was lost, or it arrived and the acknowledgement was lost. The sender can't tell which. Every messaging system has to decide what to do about that.
There are two honest choices:
- At-most-once: never retry. Some messages are lost; none are duplicated. Fine for metric samples or cursor positions, where the next message replaces the last.
- At-least-once: retry until acknowledged. None are lost; some are duplicated. This is what queues, webhooks and most RPC retries give you.
Check
Which guarantee suits 'Bob is typing…' indicators?Exactly-once delivery isn't achievable over an unreliable network, for the reason above. What systems build instead is effectively-once processing: at-least-once delivery plus a receiver that makes repeats harmless, by:
- recording processed message IDs and skipping repeats;
- making the operation idempotent (
SET status = 'paid', notbalance = balance - 10); - conditioning the write on state (
… WHERE status = 'processing').
See Idempotency.
Think first
A receiver records 'message 812 processed', then crashes before applying the effect. The message is redelivered. What happens?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Dedupe window too short
- IDs are forgotten before the sender stops retrying, so a late retry is processed again.
- Dedupe not atomic with the effect
- A crash between recording the ID and applying the effect loses or duplicates the message.
- Assuming order
- Retries reorder messages: message 2 can be applied before a retried message 1.
Instead, consider
- At-most-once (fire and forget)
- Loss is acceptable and newer data supersedes older, as with telemetry, presence and cursor positions.
- Transactional exactly-once within one system
- Consumer and producer state live in the same system (for example Kafka transactions for read-process-write between topics), and no external side effects occur.
In practice
- Unique constraint on message ID
- Insert-or-ignore in the same transaction as the effect.
- Idempotency keys on APIs
- The caller supplies the ID; the server stores the result under it.
- Conditional state transitions
- The state machine itself rejects repeats.
- Broker dedupe windows
- e.g. SQS FIFO deduplication: helpful, but bounded in time.
It assumes
- Receivers can identify a message (an ID or a natural key) across redeliveries.
- The dedupe record lives at least as long as the sender may keep retrying.
Explain it in your own words
Where you practise it
- A reliable video processing pipeline
- A job queue that keeps working when workers fall behind
What 'at least once' commits you to · Write the worker loop · Switch over without an outage
- Product analytics over billions of events
- Notifications across email, push and in-app
What the numbers say · One notification per event, per channel · Why users got two emails · Defend the duplicate policy
- A payment workflow that never double-charges
- A real-time collaborative editor
Further reading
Engineers describing it in systems they run.
- Uber's Real-Time Push Platform
Uber · Madan Thangavelu and others · Post, Dec 2020
Why polling was replaced, and the delivery problems push brought with it: resuming after a dropped connection, and knowing what actually arrived.
Related concepts
- Idempotency
Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.
- Message queues
A durable buffer between producers and consumers that hands each message to one consumer at a time and redelivers it unless acknowledged.
- Webhooks
HTTP callbacks from another system: delivered at least once, possibly out of order, possibly never. Handle them as hints, not truth.
- Retries, backoff and jitter
Retrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.
- Transactions
Grouping several reads and writes so they take effect all together or not at all, isolated from concurrent work to a defined degree.
- Transactional outbox
Recording outgoing messages in the same database transaction as the state change, then delivering them separately, to avoid the dual-write problem.
- Soft state
State that expires unless refreshed. It is cheap to keep, safe to lose, and right for presence, sessions and anything that describes the present moment.