Design an API Rate Limiter, stage 4 of 10: break it
The limiter that leaks
The buckets are in Redis, as decided. The logic is a direct translation of the token bucket. Find the lines that let a 60/minute key make 4,000 requests, and the line that made the incident worse.
System so far· 5 parts
Select a component to see what it is responsible for and which state it owns.
- 1API clients → Load balancer: Requests with API key
- 2Load balancer → API instances: Round-robin across instances
- 3API instances → Redis: Atomic take-tokens script
- 4API instances → Postgres: Admitted requests; plan lookups (cached)
What you need to know
0 of 2 checks done
A read-modify-write reads a value, computes a new one, and writes it back. If it takes two round trips (a
GET, then aSET), other clients can read the same old value in between.Each of them computes its own update from the same starting point, and the last write wins. The other updates are lost.
Check
A bucket holds 1 token. Three instances each GET it at the same moment, see 1 token, allow their request and SET the bucket to 0. How many requests were admitted, and how many tokens were spent?Two more details decide whether a limiter behaves:
- Whose clock? Refilling needs "seconds since the last refill". If each instance uses its own clock, one running ahead refills tokens that were never earned. One clock for every caller fixes that, for example the store's own clock.
- What does a rejection say? Most HTTP clients retry failed requests. A 429 with no hint about when to come back is usually retried at once, so every rejection produces another request.
Think first
A client retries every failed request immediately, and its limiter rejects 90% of its requests. What happens to the number of requests the limiter has to handle?