AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.
View the Markdown
```meta
title: Public API rate limiter
subtitle: How every request to api.example.com is counted, where the count lives, and what a client sees at the limit.
tag: DESIGN
```
The public API serves 12,000 API keys from six gateway pods in one region. Peak traffic is 2,400 requests a second, and one key can send 40% of it during a bulk import. The limiter protects the services behind the gateway from one tenant, not from the internet: the WAF handles floods. Every key gets 100 requests a second sustained with a burst of 200.
Assumptions in this document: the gateway is Envoy with a Go external-auth filter. The count store is a Redis cluster the platform team already runs. Limits are per key, and a key belongs to one plan.
## Where the count lives
```block
id: rl-topology
preset: infra
groups:
- { id: edge, col: 1, row: 1, cols: 1, rows: 2, label: Edge }
- { id: gw, col: 2, row: 1, cols: 2, rows: 2, label: Gateway }
- { id: data, col: 4, row: 1, cols: 1, rows: 2, label: State }
nodes:
- { id: client, col: 1, row: 1, kind: client, name: API client, tech: "X-Api-Key" }
- { id: waf, col: 1, row: 2, kind: waf, name: WAF, tech: Cloudflare }
- { id: envoy, col: 2, row: 1, kind: gateway, name: Envoy, tech: "6 pods", replicas: 6 }
- { id: limiter, col: 3, row: 1, kind: service, name: limiter, tech: "Go ext_authz" }
- { id: plans, col: 3, row: 2, kind: config, name: plan-config, tech: "ConfigMap, 60 s reload" }
- { id: redis, col: 4, row: 1, kind: redis, name: buckets, tech: "Redis cluster, 3 shards" }
- { id: upstream, col: 4, row: 2, kind: service, name: Orders API }
edges:
- client -> waf: HTTPS
- waf -> envoy: filtered traffic
- envoy -> limiter: check key
- limiter -> redis: take a token
- plans --> limiter: limits per plan
- envoy -> upstream: forward if allowed
```
The count lives in Redis, not in the pod. Six pods each hold a local count and a key could send 6× its limit by spreading requests. Redis adds one round trip of about 0.6 ms inside the region, which is 2% of the API's p50.
## One request at the limit
```sequence
id: rl-request
actors:
- { id: Client, name: API client }
- { id: Envoy, name: Envoy }
- { id: Limiter, name: limiter }
- { id: Redis, name: buckets }
- { id: Orders, name: Orders API }
messages:
- Client -> +Envoy: "GET /v1/orders (X-Api-Key: k_7f2)"
- Envoy -> +Limiter: "ext_authz check (key, route)"
- Limiter -> +Redis: "EVALSHA take_token k_7f2 rate=100 burst=200 now"
- alt: tokens remain
- Redis --> -Limiter: "allowed=1 remaining=143 reset=1"
- Limiter --> -Envoy: "200 OK + X-RateLimit-Remaining: 143"
- Envoy -> +Orders: forward
- Orders --> -Envoy: "200"
- Envoy --> -Client: "200 + rate limit headers"
- else: bucket empty
- Redis --> Limiter: "allowed=0 remaining=0 retry_after=0.42"
- Limiter --> Envoy: "429 + Retry-After: 1"
- Envoy --> Client: "429 Too Many Requests"
- else: Redis timeout after 5 ms
- Limiter --> Envoy: "200 OK, header X-RateLimit-Mode: open"
- Envoy --> Client: "200 (limit not enforced)"
- end
foot:
- { label: Redis budget, value: "5 ms, then fail open" }
- { label: Headers, value: "X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, Retry-After" }
```
The Lua script reads, refills, and decrements the bucket in one call, so two pods cannot both take the last token. A Redis timeout fails open: a slow Redis must not become an API outage. The gateway counts open-mode responses, and more than 1% in five minutes pages the platform on-call.
## What Redis has to carry
```envelope
id: rl-envelope
assumptions:
- { label: Peak requests per second, value: "2,400" }
- { label: Redis calls per request, value: "1" }
- { label: Active keys in any hour, value: "12,000" }
- { label: Bucket record, value: "key + tokens + timestamp, about 96 B" }
steps:
- { label: Redis ops per second at peak, calc: "2,400 × 1", result: "2,400 ops/s" }
- { label: Ops per shard, calc: "2,400 / 3 shards", result: "800 ops/s" }
- { label: Memory for all buckets, calc: "12,000 × 96 B", result: "1.2 MB" }
- { label: Headroom to a single shard's limit, calc: "80,000 ops/s / 800", result: "100×" }
result: { label: Provision, value: "the existing 3-shard cluster; no new capacity" }
```
The limiter is not a Redis capacity problem at any traffic the WAF lets through. The risk is latency, not throughput, which is why the timeout is 5 ms and not 50.
## Why a token bucket
```options
id: rl-algorithm
items:
- kicker: Option 1
title: Fixed window
how: One counter per key per second; reset at the boundary.
pros: [One INCR per request, Trivial to explain]
cons: ["A client can send 200 requests in the 10 ms around a boundary", "No burst allowance without a second counter"]
verdict: "REJECTED — the boundary burst is 2× the limit"
tone: rejected
- kicker: Option 2
title: Sliding log
how: Store every request timestamp in a sorted set; count the last second.
pros: [Exact at any instant]
cons: ["Memory grows with the rate: 100 entries per key per second", "ZADD + ZREMRANGEBYSCORE + ZCARD per request"]
verdict: "REJECTED — 240 KB per hot key, and 3 ops per request"
tone: rejected
- kicker: Option 3
title: Token bucket
how: Each key holds tokens and a last-refill time; refill on read, take one per request.
pros: [Burst is the bucket size, Two numbers per key, One Lua call per request]
cons: ["Refill math depends on the Redis clock, so all shards use TIME inside the script"]
verdict: "CHOSEN"
tone: chosen
```
## The contract a client can rely on
```spec
id: rl-contract
accent: teal
rows:
- { label: Scope, value: "One bucket per API key. Plan limits apply to every route the same way." }
- { label: Default limits, value: "100 requests per second sustained, burst 200. Enterprise plan: 1,000 and 2,000." }
- { label: On every response, value: "X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset (seconds to full refill)." }
- { label: At the limit, value: "429 with Retry-After in whole seconds, rounded up. The body is a JSON error with code rate_limited." }
- { label: Redis unavailable, value: "Requests pass with X-RateLimit-Mode: open. Open mode above 1% for 5 minutes pages the platform on-call." }
- { label: Limit change, steps: [Edit plan-config, Merge, "Gateway reloads within 60 s", "Buckets keep their tokens"] }
```
A plan change never resets a bucket, so a client in the middle of a burst keeps its remaining tokens. Per-route limits are out of scope for this version; the bucket key has room for a route suffix when a route needs one.