System design concepts
The mechanisms the investigations depend on. Each one starts from the problem it solves, explains how it actually works, and ends with how it fails, not with a product name.
Communication
- Asynchronous processingAccepting a request, recording the work durably, and doing it later in another process, so the work can outlive the request.7 stages
- Message queuesA durable buffer between producers and consumers that hands each message to one consumer at a time and redelivers it unless acknowledged.13 stages
- Delivery guaranteesAt-most-once, at-least-once, and why 'exactly-once' is achieved by making duplicates harmless rather than by preventing them.11 stages
- Server push: polling, long polling, SSE, WebSocketsThe ways a server can tell a client that something changed, and how update rate, latency needs and who-knows-what decide between them.4 stages
- Persistent connectionsLong-lived connections such as WebSockets turn a stateless request tier into one that holds per-client state, with consequences for routing, deploys and failure detection.6 stages
- WebhooksHTTP callbacks from another system: delivered at least once, possibly out of order, possibly never. Handle them as hints, not truth.5 stages
- Publish/subscribeDecoupling senders from receivers by topic: a publisher sends once and every current subscriber receives a copy.3 stages
Storage & state
- Object storageA flat namespace of immutable blobs addressed by key, built for durability and size rather than queries or in-place updates.4 stages
- TransactionsGrouping several reads and writes so they take effect all together or not at all, isolated from concurrent work to a defined degree.8 stages
- Append-only logsRecording changes as an ordered, immutable sequence of facts, from which current state, history and replicas can be derived.11 stages
- DurabilityWhat has to have happened before a system may say "saved": which failures the data must survive, and where that guarantee is actually made.3 stages
- State machines for business stateModelling an entity's lifecycle as explicit states and allowed transitions, enforced with conditional updates so concurrent or stale actors cannot corrupt it.13 stages
- Online data migrationsMoving live data to a new schema or store without downtime: write to both, backfill the past, verify, switch reads, then switch writes, with a way back at every step.7 stages
- Log-structured storage (LSM trees)Storage engines that turn every write into a sequential append and merge files in the background: very fast writes, at the cost of compaction, tombstones and more expensive reads.5 stages
- Columnar storageStoring each column of a table separately, so analytical queries read only the columns they use and compress them well, at the cost of slow single-row lookups and updates.5 stages
Reliability
- IdempotencyDesigning an operation so that performing it twice has the same effect as performing it once, which is what makes retries safe.24 stages
- Retries, backoff and jitterRetrying transient failures with growing, randomized delays and a budget, so recovery does not become the next outage.11 stages
- Timeouts and unknown outcomesA timeout bounds how long you wait. It tells you nothing about what happened, so the operation's outcome becomes unknown.12 stages
- ReconciliationPeriodically comparing your records with an authoritative source and repairing differences, the backstop for every message that was lost.3 stages
- Transactional outboxRecording outgoing messages in the same database transaction as the state change, then delivering them separately, to avoid the dual-write problem.6 stages
Concurrency
- Concurrency controlMaking read-decide-write sequences safe when other actors may change the same data in between: locks, conditional writes and constraints.17 stages
- OrderingThere is no global 'now' in a distributed system. Order exists only where something assigns it, so decide which order you need and who assigns it.5 stages
- Conflict resolution and convergenceWhen replicas accept concurrent changes, a deterministic rule must merge them so every replica ends in the same state without losing intent.4 stages
Distribution
- Generating unique identifiersMaking ids that are unique across machines and time, and choosing what else they reveal: order, volume, guessability, length.7 stages
- ReplicationKeeping copies of data on several machines for durability, read capacity and locality, and living with copies that briefly disagree.9 stages
- Leases and fencing tokensOwnership that expires unless renewed, plus a token that lets the rest of the system reject an owner that has lost its claim without knowing it.10 stages
- PartitioningSplitting data or work by key so each part is handled independently: scaling out, and giving each key a single owner.20 stages
- Soft stateState that expires unless refreshed. It is cheap to keep, safe to lose, and right for presence, sessions and anything that describes the present moment.5 stages
- Consistent hashingMapping keys to nodes so that adding or removing a node moves only a small share of keys, instead of reshuffling almost all of them.3 stages
Performance & scale
- Backpressure and capacityWhen work arrives faster than it can be done, something has to give: the queue grows, the producer slows, or work is shed. Choose which on purpose.24 stages
- CachingKeeping a copy of data closer to where it is used, trading freshness and complexity for speed and reduced load on the source.26 stages
- Rate limitingCapping how fast a client may use a resource, to protect capacity, enforce fairness, and stay within the limits of the systems you depend on.17 stages
- Fan-out on write and fan-out on readWhen one write must reach many readers, do the work when it is written (precompute every reader's view) or when it is read (assemble it on demand). Most real feeds do both.10 stages
- Request coalescingWhen many callers ask for the same thing at the same moment, do the expensive work once and give every caller the result. The fix for thundering herds and hot keys.6 stages
- Load sheddingRejecting some work on purpose when a system is overloaded, so the work it does accept still finishes in time. Cheap rejections beat slow failures.2 stages