Skip to content

Design an API Rate Limiter, stage 5 of 10: break it

Write the atomic bucket

Write the Redis-side logic (Lua, or pseudo-code for "runs atomically on the server") that refills, takes tokens and reports the result, plus the caller that sets the response headers.

System so far· 5 parts
1234CLIENTAPI clientsEDGELoad balancerSERVICEAPI instancesCACHERedisDATABASEPostgres

Select a component to see what it is responsible for and which state it owns.

  1. 1API clients → Load balancer: Requests with API key
  2. 2Load balancer → API instances: Round-robin across instances
  3. 3API instances → Redis: Atomic take-tokens script
  4. 4API instances → Postgres: Admitted requests; plan lookups (cached)

What you need to know

0 of 3 checks done
  1. Redis runs a Lua script as one atomic step: no other command on that shard runs while the script is running. So a script can read a bucket, refill it, take tokens and write it back without anyone else seeing it half-updated.

    Inside a script, redis.call('TIME') returns the server's clock as seconds and microseconds. Every instance that calls the script shares that one clock.

  2. The refill arithmetic, for a bucket with capacity capacity and refill rate tokens per second:

    tokens = min(capacity, tokens + (now − last) × rate)

    If tokens ≥ cost, subtract the cost and allow. Otherwise reject, and the wait until enough tokens exist is (cost − tokens) ÷ rate.

  3. Work it out

    A bucket refills at 10 tokens a second and holds 0.4 tokens. A request costs 1. How many seconds until it can be admitted?
    seconds