Home / Blog / Rate-limit design
API Design and Platform Reliability

API Rate Limits Are Product Architecture, Not Just a 429

A rate limit is a product promise about shared capacity: who gets to use it, how bursts are treated, what happens at the boundary and how a client recovers. A blunt “100 requests per minute” can protect a service while punishing normal batch work, hiding expensive endpoints or making retries amplify an outage.

Translate service constraints into a client contract

Enforce policy near ingress, communicate a useful limit and preserve a safe retry path
1 / IdentifyAuthenticate tenant, user, credential and endpoint.
2 / ClassifyEstimate cost, burst and concurrency demand.
3 / EnforceApply scoped budget at a consistent boundary.
4 / RecoverReturn 429 guidance; clients back off with jitter.

Start with the reason for limiting: protect database capacity, isolate tenants, control an expensive model call, stop credential abuse or fit a contractual quota. These are different policies and may need different counters. A user-facing quota should be observable and explainable; a protective circuit limit may be intentionally short and adaptive.

Choose what the budget counts

CounterProtectsCommon failure if used alone
Requests per key/accountFair access and abuse containment.Cheap reads and expensive writes cost the same.
Weighted units per tenantDatabase, compute or AI spend.Cost weights drift from real resource consumption.
Concurrent in-flight workWorker pools, connections and long-running jobs.Does not limit total daily usage.
Global service budgetProtects shared infrastructure during spikes.One noisy tenant can crowd out everyone else.
Business quotaPlan allowance or paid consumption.Must match billing, reporting and reset semantics.

Most production APIs need layered controls: per-credential or tenant fairness, endpoint-specific cost weights, a concurrency guard for scarce resources and a global emergency ceiling. Keep authentication before per-user enforcement; IP limits are useful for unauthenticated abuse but can penalize users behind shared NATs.

Understand common algorithms

A fixed window is simple and cheap but permits boundary bursts: a client can spend the full quota just before and after reset. A sliding log is accurate but stores more events. A sliding-window counter approximates the recent rate with less state. A token bucket allows a defined burst while enforcing a long-term refill rate; leaky-bucket shaping smooths output rather than merely rejecting. Concurrency limits cap simultaneous work and complement request-rate limits. Choose based on desired user behavior, implementation cost and consistency needs.

Make 429 useful and safe

HTTP 429 means too many requests; RFC 6585 says the response should explain the condition and may include Retry-After. Return a stable machine-readable error code, a retry time when meaningful and limit metadata where it helps clients plan. Avoid leaking another tenant's usage or internal capacity details. Make sure the 429 response is not cached as a normal successful representation.

Client libraries should honor Retry-After, use exponential backoff with jitter, cap retries and avoid retrying non-idempotent mutations unless they carry an idempotency key. Synchronized retries can create a thundering herd just as the service recovers. For batch APIs, offer bounded batch size or asynchronous job submission rather than forcing clients to issue thousands of tiny calls.

Distributed enforcement is a consistency choice

A counter stored in one process is fast but resets on restart and diverges across replicas. A centralized counter gives a shared view but adds latency and can become a dependency. Edge-local enforcement scales geographically but requires allowance allocation and accepts bounded overshoot. Decide whether strict global quota or availability matters more during counter-store failure; document fail-open or fail-closed behavior per endpoint and risk.

For a commercial quota, billing-grade usage should come from durable events or an accounting ledger, not only a short-lived throttle counter. Rate limiting protects the service; it is not automatically a trustworthy meter of billable consumption.

Communicate limits before clients hit them

Document per-plan quotas, burst allowances, endpoint cost differences, concurrency caps, reset times, pagination and retry expectations. Expose usage telemetry to account owners and warn before hard limits where practical. Version policy changes and provide a migration window for integrations. Support needs a way to distinguish a policy limit from an outage, authentication failure or exhausted account balance.

Measure fairness and side effects

Track 429 rate by tenant/route, accepted cost units, queue and connection saturation, retry amplification, latency and business completion. Look for clients receiving repeated 429s despite low average use: bursts, clock assumptions or an unfair shared counter may be the real issue. Load-test both synchronized bursts and one noisy neighbor, including recovery after the limiter store is unavailable.

In summary

Rate limits shape how customers integrate with a service. Define the capacity or abuse risk first, choose counters and algorithms that match the workload, apply layered scopes, return actionable 429 guidance and specify behavior when enforcement state fails. A good limit preserves fair access and gives well-behaved clients a predictable path to success.

References