Design an API Rate Limiter, stage 7 of 10: change it
The biggest customer
Partitioning Redis by key spreads 50,000 keys nicely, except when one key is most of the traffic.
System so far· 5 parts
Select a component to see what it is responsible for and which state it owns.
- 1API clients → Load balancer: Requests with API key
- 2Load balancer → API instances: Round-robin across instances
- 3API instances → Redis: Atomic take-tokens script
- 4API instances → Postgres: Admitted requests; plan lookups (cached)
What you need to know
0 of 2 checks done
Redis Cluster splits keys over shards by hashing the key name. 50,000 keys spread evenly, and adding shards spreads them further.
But all commands for one key go to one shard, and each shard runs commands on a single thread. A key that receives most of the traffic stays on one shard however many shards exist. That is a hot key.
Check
One key gets 15,000 script calls a second and its shard is at 90% CPU. What does doubling the number of shards do for that key?Two ways to take load off one key:
- Batch: an instance takes many tokens at once (say 200) and spends them locally. One Redis call now covers 200 requests.
- Split: store the key as N sub-buckets on different shards, each holding 1/N of the limit. Each request picks one sub-bucket.
Both reduce coordination per request. Both make the limit slightly less exact.
Work it out
15,000 requests a second for one key, with instances taking tokens 200 at a time. About how many Redis calls a second does that key need?