Skip to content

Durability

What has to have happened before a system may say "saved": which failures the data must survive, and where that guarantee is actually made.

Storage & state

Learn it

0 of 2 checks done
  1. A system acknowledges a write, then the process crashes, the machine loses power, or a disk dies. Was the write saved? It depends entirely on what happened before the acknowledgement, and acknowledging earlier is always faster, which is why systems are tempted to.

    Durability is defined relative to a failure: "survives a process crash", "survives power loss", "survives losing the machine".

  2. SurvivesMust have happened before the ack
    process crashdata left the process (written to the OS)
    power lossdata on stable storage: fsync, not just write()
    machine or disk lossdata on another machine: synchronous replication
    region lossdata in another region, with cross-region latency

    write() only puts data in the OS page cache. Databases commit by fsyncing a write-ahead log.

  3. Check

    A primary acknowledges commits before its asynchronous replica confirms them. The primary's disk fails and the replica is promoted. What can be lost?

Quick reference

The same ideas, condensed for revision.

How it goes wrong

Ack before flush
Acknowledged writes vanish on crash.
Async replica failover
The primary dies; the promoted replica lacks the last transactions that clients were told were committed.
Durable but unreachable
Data is safe on a single machine that is down, which is durable but not available.

Instead, consider

Acknowledge before durable, accept loss
The data can be regenerated or its loss is cheap: metrics, caches, presence.
Client-side retention until durable ack
The client keeps unacknowledged data and resends it, so loss on the server side is recoverable.

In practice

fsync / fdatasync on a log
The basic mechanism databases use at commit.
Synchronous replication (synchronous_commit, quorum writes)
Survives machine loss at the cost of a network round trip per commit.
Object stores
Replicate across facilities before acknowledging a PUT.

It assumes

  • Storage hardware honours flushes (some consumer disks and virtualized layers do not).
  • The failure model is explicit: which failures must not lose data, and which are acceptable.

Explain it in your own words

Write at least 60 characters (0 so far). Write it as you would say it in a design review. You will compare it against the points a strong answer makes.

Where you practise it

Further reading

Engineers describing it in systems they run.

  • Transactions

    Grouping several reads and writes so they take effect all together or not at all, isolated from concurrent work to a defined degree.

  • Append-only logs

    Recording changes as an ordered, immutable sequence of facts, from which current state, history and replicas can be derived.

  • Object storage

    A flat namespace of immutable blobs addressed by key, built for durability and size rather than queries or in-place updates.

  • Log-structured storage (LSM trees)

    Storage engines that turn every write into a sequential append and merge files in the background: very fast writes, at the cost of compaction, tombstones and more expensive reads.

  • Replication

    Keeping copies of data on several machines for durability, read capacity and locality, and living with copies that briefly disagree.