State machines for business state
Modelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.
Storage & state
Learn it
An order, a payment or a job is touched by many actors: the API, workers, webhooks, reconcilers, admins. Each sets a status. Without rules, a late or duplicate actor can move a
paidorder back topending, resurrect a deleted video, or ship something twice.A state machine defines the states and the allowed transitions:
created → processing → succeeded | failed,succeeded → refunded.Enforce each transition where the data lives, with a conditional update:
UPDATE payments SET status = 'succeeded', succeeded_at = now() WHERE id = $1 AND status = 'processing';The
WHERE status = 'processing'makes the database the arbiter. If two actors race, exactly one update matches; the other changes zero rows.Check
Why not read the status in application code, check the transition is allowed, then write the new status?Design notes:
- Make uncertainty a state. "We sent the request and don't know the result" needs its own state with a defined way out (Reconciliation). Guessing turns uncertainty into wrong data.
- Terminal states are terminal. Refunds and reversals are new transitions, not overwrites.
- Side effects attach to transitions, and run only for the actor whose update succeeded, ideally via a Transactional outbox.
- Record why: which event or actor caused each transition.
Think first
A worker's conditional update to 'succeeded' changes zero rows. What does that mean, and what should it do?
Quick reference
The same ideas, condensed for revision.
How it goes wrong
- Unconditional updates
- Last writer wins, regardless of whether its information is current.
- No state for 'unknown'
- Timeouts are recorded as failures and later contradicted by reality.
- Side effects outside the transition
- Two actors both send the email because neither checked whether its transition won.
Instead, consider
- Derive state from an event log
- History matters as much as current state; the state machine is then enforced when appending events.
- Workflow engine
- The lifecycle has many steps with timers, retries and human tasks between them.
In practice
- Status column + conditional UPDATE
- Check affected-row count to know whether you won.
- Check constraints / triggers
- Reject invalid transitions in the database itself.
- Typed state machine in code
- Encode allowed transitions; still enforce atomically in storage.
It assumes
- Every actor that changes the state goes through the same conditional transitions.
- The state lives in one place that supports atomic conditional updates.
Explain it in your own words
Where you practise it
- A reliable video processing pipeline
How does the system learn the upload finished? · The video that kills every worker · Deleting a video mid-transcode
- Notifications across email, push and in-app
When are preferences applied? · Digests and quiet hours · Write the sender
- A payment workflow that never double-charges
Should checkout wait for the provider? · Where does payment state live? · Which transitions can happen? · Write the idempotent checkout · Webhooks arrive twice, and out of order · Adding refunds · Defend the guarantee
Further reading
Engineers describing it in systems they run.
- Stripe's payments APIs: the first ten years
Stripe · Michelle Bu · Post, Dec 2020
How payment methods that confirm asynchronously broke the original API, and why the replacement models a payment as one explicit state machine.
- Implementing Stripe-like Idempotency Keys in Postgres
Stripe · Brandur Leach · Post, Oct 2017
The long version, with code: how to make a multi-step request safe to retry when some of its steps call other services.
Related concepts
- Concurrency control
Making read-decide-write sequences safe when other actors may change the same data in between: locks, conditional writes and constraints.
- Idempotency
Designing an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.
- Transactions
Grouping several reads and writes so they take effect all together or not at all, isolated from concurrent work to a defined degree.
- Reconciliation
Periodically comparing your records with an authoritative source and repairing differences, the backstop for every message that was lost.
- Asynchronous processing
Accepting a request, recording the work durably, and doing it later in another process, so the work can outlive the request.
- Webhooks
HTTP callbacks from another system: delivered at least once, possibly out of order, possibly never. Handle them as hints, not truth.
- Timeouts and unknown outcomes
A timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.