Design an API Rate Limiter, stage 9 of 10: change it
Not all requests cost the same
Every customer stayed inside their requests-per-minute plan. The cluster went down anyway.
System so far· 5 parts
Select a component to see what it is responsible for and which state it owns.
- 1API clients → Load balancer: Requests with API key
- 2Load balancer → API instances: Round-robin across instances
- 3API instances → Redis: Atomic take-tokens script
- 4API instances → Postgres: Admitted requests; plan lookups (cached)
What you need to know
0 of 2 checks done
A request-count limit assumes requests cost roughly the same. When one request can take 1,000 times longer than another, counting requests says little about load.
What the backend runs out of is time: how many query-seconds it can serve per second. That's the resource to limit.
Work it out
A customer's plan allows 10 searches a second, and each of their searches takes 2 seconds of cluster time. How many search-seconds of work do they add every second?Three tools, each limiting something different:
- Cost-weighted tokens: a request takes tokens in proportion to its estimated cost, so the bucket measures work rather than requests.
- Concurrency limit per key: at most N requests from one key running at once, which bounds how much of the backend one client can occupy.
- Global load shedding: when the backend's queue grows, reject or defer work from everyone, even clients within their own limits.
Check
Every customer is within their own per-key limits, yet together they exceed what the cluster can serve. Which tool still protects it?