Home / Blog / Backpressure
Distributed Systems

Backpressure in Everyday Systems

When an API, webhook producer or user interface creates work faster than a downstream service can complete it, a queue can absorb the burst but cannot erase the capacity gap. Without a feedback mechanism, the backlog grows until latency, memory, storage or customer patience runs out. Backpressure is how a system makes that mismatch visible and bounded.

Read depth together with arrival and service rates

Illustrative queue behavior; the key signal is the gap between arrival and completion
StableIncoming rate stays below sustained consumer throughput; queue returns to baseline.
OverloadedIncoming rate exceeds service capacity; depth and message age trend upward.
RecoveringAdmission is reduced and spare throughput drains backlog faster than new arrivals.

Queue depth alone is ambiguous. Pair it with enqueue rate, completion rate, oldest-message age, retry volume and consumer saturation. A deep queue may be normal for batch work; an old message in a latency-sensitive workflow may already be an incident. Estimate drain time from backlog divided by spare processing capacity, then include new arrivals rather than assuming the queue is static.

Bound every buffer

BoundaryBackpressure mechanismTrade-off
HTTP ingressRate limit, concurrency cap, 429/503 with retry guidanceProtects service but pushes retry responsibility to clients
Broker deliveryPrefetch or receive batch limit and explicit acknowledgementsBounds in-flight work; too low can reduce throughput
Worker poolBounded concurrency and queue sizePredictable memory; excess work waits or is rejected
DatabaseConnection pool, bulkhead and admission controlPrevents connection collapse; may increase upstream latency
Browser clientDebounce, cancel stale requests, cap parallel fetchesReduces duplicate demand; can delay visible updates

An unbounded in-memory queue is not resilience. It converts overload into memory growth and a later, less controlled failure. Set queue and payload limits, define expiry, and decide whether excess work is rejected, coalesced, persisted or dropped. Different work has different value: stale typing suggestions can be discarded; financial reconciliation cannot.

Use feedback, not only more workers

Scaling consumers helps when the downstream dependency has capacity. If the database or partner API is already saturated, adding workers can increase contention and retries. Apply concurrency limits at the dependency boundary, use exponential backoff with jitter and honor provider retry hints. A circuit breaker can pause calls, but recovery needs a half-open test and a safe queue drain policy.

For ordered workflows, partition by a stable key and avoid allowing one poison message to block unrelated partitions. For at-least-once delivery, acknowledge only after durable success and make side effects idempotent. Set visibility or acknowledgement timeouts longer than typical processing, with a controlled extension path for long jobs.

Tell clients what the system can accept

Backpressure should be part of the API contract. Return a clear overload status, a retry-after signal where appropriate and a correlation identifier. For asynchronous acceptance, respond only after the work is durably recorded and give a status endpoint or callback. Do not report “accepted” for work held only in volatile memory.

Operational signals that matter

Dashboard arrival and completion rates together, queue age percentiles, in-flight count, retries, dead-letter volume, consumer utilization and dependency saturation. Alert on sustained growth and time-to-drain, not a fixed queue-depth threshold divorced from workload. During recovery, communicate whether the queue is draining and avoid a synchronized retry storm.

Related: queue-based architecture and API rate limits as product architecture.

In summary

Backpressure is a capacity policy expressed through bounded buffers, controlled concurrency and explicit overload responses. Measure how quickly work arrives and completes, protect the slowest dependency, and choose what to defer or reject before the system chooses for you.

References