Skip to content
chiltepin

Generated from: “Design the rate limiter for the public API.”

Public API rate limiter design document

AI-generated example · 2026-09-14. The source passes chiltepin check. System details and measurements are illustrative; review them before adapting this document.

DOCUMENTDESIGN

Public API rate limiter

How every request to api.example.com is counted, where the count lives, and what a client sees at the limit.

The public API serves 12,000 API keys from six gateway pods in one region. Peak traffic is 2,400 requests a second, and one key can send 40% of it during a bulk import. The limiter protects the services behind the gateway from one tenant, not from the internet: the WAF handles floods. Every key gets 100 requests a second sustained with a burst of 200.

Assumptions in this document: the gateway is Envoy with a Go external-auth filter. The count store is a Redis cluster the platform team already runs. Limits are per key, and a key belongs to one plan.

Where the count lives

SECTION 01 · Architecture
INFRA
Block diagram: 7 nodes, 6 connectionsGatewayEdgeStateAPI clientX-Api-KeyCLIENTWAFCloudflareWAFEnvoy6 podsGATEWAY×6limiterGo ext_authzSVCplan-configConfigMap, 60 sreloadCONFIGbucketsRedis cluster, 3shardsCACHEOrders APISVC123456
1HTTPS2filtered traffic3check key4take a token5limits per plan6forward if allowed
LegendCLIENTclientWAFfirewallGATEWAYgatewaySVCserviceCONFIGconfigCACHEcache×Nreplicascallsasync / optionalentry point

The count lives in Redis, not in the pod. Six pods each hold a local count and a key could send 6× its limit by spreading requests. Redis adds one round trip of about 0.6 ms inside the region, which is 2% of the API's p50.

One request at the limit

SECTION 02 · Sequence
SEQUENCE
Sequence diagram: 13 messages between 5 actorsAPI clientEnvoylimiterbucketsOrders APIALT[tokens remain][bucket empty][Redis timeout after 5 ms]1GET /v1/orders (X-Api-Key: k_7f2)2ext_authz check (key, route)3EVALSHA take_token k_7f2 rate=100 burst=200 now4allowed=1 remaining=143 reset=15200 OK + X-RateLimit-Remaining: 1436forward72008200 + rate limit headers9allowed=0 remaining=0 retry_after=0.4210429 + Retry-After: 111429 Too Many Requests12200 OK, header X-RateLimit-Mode: open13200 (limit not enforced)
Legendcallresponsethe answer the caller getsfragment (alt / opt / loop)active
Redis budget: 5 ms, then fail openHeaders: X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, Retry-After

The Lua script reads, refills, and decrements the bucket in one call, so two pods cannot both take the last token. A Redis timeout fails open: a slow Redis must not become an API outage. The gateway counts open-mode responses, and more than 1% in five minutes pages the platform on-call.

What Redis has to carry

SECTION 03 · Capacity math
Peak requests per second
2,400
Redis calls per request
1
Active keys in any hour
12,000
Bucket record
key + tokens + timestamp, about 96 B
Redis ops per second at peak2,400 × 12,400 ops/s
Ops per shard2,400 / 3 shards800 ops/s
Memory for all buckets12,000 × 96 B1.2 MB
Headroom to a single shard's limit80,000 ops/s / 800100×
Provision
the existing 3-shard cluster; no new capacity

The limiter is not a Redis capacity problem at any traffic the WAF lets through. The risk is latency, not throughput, which is why the timeout is 5 ms and not 50.

Why a token bucket

SECTION 04 · Options
Option 1Fixed window
One counter per key per second; reset at the boundary.
  • One INCR per request
  • Trivial to explain
  • A client can send 200 requests in the 10 ms around a boundary
  • No burst allowance without a second counter
REJECTED — the boundary burst is 2× the limit
Option 2Sliding log
Store every request timestamp in a sorted set; count the last second.
  • Exact at any instant
  • Memory grows with the rate: 100 entries per key per second
  • ZADD + ZREMRANGEBYSCORE + ZCARD per request
REJECTED — 240 KB per hot key, and 3 ops per request
Option 3Token bucket
Each key holds tokens and a last-refill time; refill on read, take one per request.
  • Burst is the bucket size
  • Two numbers per key
  • One Lua call per request
  • Refill math depends on the Redis clock, so all shards use TIME inside the script
CHOSEN

The contract a client can rely on

SECTION 05 · Spec
Scope
One bucket per API key. Plan limits apply to every route the same way.
Default limits
100 requests per second sustained, burst 200. Enterprise plan: 1,000 and 2,000.
On every response
X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset (seconds to full refill).
At the limit
429 with Retry-After in whole seconds, rounded up. The body is a JSON error with code rate_limited.
Redis unavailable
Requests pass with X-RateLimit-Mode: open. Open mode above 1% for 5 minutes pages the platform on-call.
Limit change
Edit plan-config→Merge→Gateway reloads within 60 s→Buckets keep their tokens

A plan change never resets a bucket, so a client in the middle of a burst keeps its remaining tokens. Per-route limits are out of scope for this version; the bucket key has room for a route suffix when a route needs one.

View the Markdown
```meta
title: Public API rate limiter
subtitle: How every request to api.example.com is counted, where the count lives, and what a client sees at the limit.
tag: DESIGN
```

The public API serves 12,000 API keys from six gateway pods in one region. Peak traffic is 2,400 requests a second, and one key can send 40% of it during a bulk import. The limiter protects the services behind the gateway from one tenant, not from the internet: the WAF handles floods. Every key gets 100 requests a second sustained with a burst of 200.

Assumptions in this document: the gateway is Envoy with a Go external-auth filter. The count store is a Redis cluster the platform team already runs. Limits are per key, and a key belongs to one plan.

## Where the count lives

```block
id: rl-topology
preset: infra
groups:
  - { id: edge, col: 1, row: 1, cols: 1, rows: 2, label: Edge }
  - { id: gw, col: 2, row: 1, cols: 2, rows: 2, label: Gateway }
  - { id: data, col: 4, row: 1, cols: 1, rows: 2, label: State }
nodes:
  - { id: client, col: 1, row: 1, kind: client, name: API client, tech: "X-Api-Key" }
  - { id: waf, col: 1, row: 2, kind: waf, name: WAF, tech: Cloudflare }
  - { id: envoy, col: 2, row: 1, kind: gateway, name: Envoy, tech: "6 pods", replicas: 6 }
  - { id: limiter, col: 3, row: 1, kind: service, name: limiter, tech: "Go ext_authz" }
  - { id: plans, col: 3, row: 2, kind: config, name: plan-config, tech: "ConfigMap, 60 s reload" }
  - { id: redis, col: 4, row: 1, kind: redis, name: buckets, tech: "Redis cluster, 3 shards" }
  - { id: upstream, col: 4, row: 2, kind: service, name: Orders API }
edges:
  - client -> waf: HTTPS
  - waf -> envoy: filtered traffic
  - envoy -> limiter: check key
  - limiter -> redis: take a token
  - plans --> limiter: limits per plan
  - envoy -> upstream: forward if allowed
```

The count lives in Redis, not in the pod. Six pods each hold a local count and a key could send 6× its limit by spreading requests. Redis adds one round trip of about 0.6 ms inside the region, which is 2% of the API's p50.

## One request at the limit

```sequence
id: rl-request
actors:
  - { id: Client, name: API client }
  - { id: Envoy, name: Envoy }
  - { id: Limiter, name: limiter }
  - { id: Redis, name: buckets }
  - { id: Orders, name: Orders API }
messages:
  - Client -> +Envoy: "GET /v1/orders (X-Api-Key: k_7f2)"
  - Envoy -> +Limiter: "ext_authz check (key, route)"
  - Limiter -> +Redis: "EVALSHA take_token k_7f2 rate=100 burst=200 now"
  - alt: tokens remain
  - Redis --> -Limiter: "allowed=1 remaining=143 reset=1"
  - Limiter --> -Envoy: "200 OK + X-RateLimit-Remaining: 143"
  - Envoy -> +Orders: forward
  - Orders --> -Envoy: "200"
  - Envoy --> -Client: "200 + rate limit headers"
  - else: bucket empty
  - Redis --> Limiter: "allowed=0 remaining=0 retry_after=0.42"
  - Limiter --> Envoy: "429 + Retry-After: 1"
  - Envoy --> Client: "429 Too Many Requests"
  - else: Redis timeout after 5 ms
  - Limiter --> Envoy: "200 OK, header X-RateLimit-Mode: open"
  - Envoy --> Client: "200 (limit not enforced)"
  - end
foot:
  - { label: Redis budget, value: "5 ms, then fail open" }
  - { label: Headers, value: "X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, Retry-After" }
```

The Lua script reads, refills, and decrements the bucket in one call, so two pods cannot both take the last token. A Redis timeout fails open: a slow Redis must not become an API outage. The gateway counts open-mode responses, and more than 1% in five minutes pages the platform on-call.

## What Redis has to carry

```envelope
id: rl-envelope
assumptions:
  - { label: Peak requests per second, value: "2,400" }
  - { label: Redis calls per request, value: "1" }
  - { label: Active keys in any hour, value: "12,000" }
  - { label: Bucket record, value: "key + tokens + timestamp, about 96 B" }
steps:
  - { label: Redis ops per second at peak, calc: "2,400 × 1", result: "2,400 ops/s" }
  - { label: Ops per shard, calc: "2,400 / 3 shards", result: "800 ops/s" }
  - { label: Memory for all buckets, calc: "12,000 × 96 B", result: "1.2 MB" }
  - { label: Headroom to a single shard's limit, calc: "80,000 ops/s / 800", result: "100×" }
result: { label: Provision, value: "the existing 3-shard cluster; no new capacity" }
```

The limiter is not a Redis capacity problem at any traffic the WAF lets through. The risk is latency, not throughput, which is why the timeout is 5 ms and not 50.

## Why a token bucket

```options
id: rl-algorithm
items:
  - kicker: Option 1
    title: Fixed window
    how: One counter per key per second; reset at the boundary.
    pros: [One INCR per request, Trivial to explain]
    cons: ["A client can send 200 requests in the 10 ms around a boundary", "No burst allowance without a second counter"]
    verdict: "REJECTED — the boundary burst is 2× the limit"
    tone: rejected
  - kicker: Option 2
    title: Sliding log
    how: Store every request timestamp in a sorted set; count the last second.
    pros: [Exact at any instant]
    cons: ["Memory grows with the rate: 100 entries per key per second", "ZADD + ZREMRANGEBYSCORE + ZCARD per request"]
    verdict: "REJECTED — 240 KB per hot key, and 3 ops per request"
    tone: rejected
  - kicker: Option 3
    title: Token bucket
    how: Each key holds tokens and a last-refill time; refill on read, take one per request.
    pros: [Burst is the bucket size, Two numbers per key, One Lua call per request]
    cons: ["Refill math depends on the Redis clock, so all shards use TIME inside the script"]
    verdict: "CHOSEN"
    tone: chosen
```

## The contract a client can rely on

```spec
id: rl-contract
accent: teal
rows:
  - { label: Scope, value: "One bucket per API key. Plan limits apply to every route the same way." }
  - { label: Default limits, value: "100 requests per second sustained, burst 200. Enterprise plan: 1,000 and 2,000." }
  - { label: On every response, value: "X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset (seconds to full refill)." }
  - { label: At the limit, value: "429 with Retry-After in whole seconds, rounded up. The body is a JSON error with code rate_limited." }
  - { label: Redis unavailable, value: "Requests pass with X-RateLimit-Mode: open. Open mode above 1% for 5 minutes pages the platform on-call." }
  - { label: Limit change, steps: [Edit plan-config, Merge, "Gateway reloads within 60 s", "Buckets keep their tokens"] }
```

A plan change never resets a bucket, so a client in the middle of a burst keeps its remaining tokens. Per-route limits are out of scope for this version; the bucket key has room for a route suffix when a route needs one.