Skip to content

Design an API Rate Limiter, stage 10 of 10: defend it

Defend a 429

A Pro customer files a ticket: "Our logs show we sent 540 requests in the last minute, under our 600 limit, and we got 429s. Your limiter is broken."

Respond as the engineer who built it: how can that happen, which cases are bugs, and what would you check?

System so far· 6 parts
12345CLIENTAPI clientsEDGELoad balancerSERVICEAPI instancesCACHERedisDATABASEPostgresSERVICESearch cluster

Select a component to see what it is responsible for and which state it owns.

  1. 1API clients → Load balancer: Requests with API key
  2. 2Load balancer → API instances: Round-robin across instances
  3. 3API instances → Redis: Atomic take-tokens script
  4. 4API instances → Postgres: Admitted requests; plan lookups (cached)
  5. 5API instances → Search cluster: Admitted searches, cost-weighted

What you need to know

0 of 2 checks done
  1. Customers read plans as "600 requests a minute". A token bucket enforces something slightly different: a burst of up to B at once, then a refill of r per second. Over a long period those agree. Over a few seconds they can differ.

  2. Think first

    Pro has B = 50 and r = 10 per second. A client sends 200 requests in the first 5 seconds of a minute, then 340 spread over the remaining 55 seconds. Does it get any 429s?