Home / Blog / Webhook reliability
Backend Reliability and SaaS Integrations

Webhook Reliability: The Boring Architecture Behind Modern SaaS

A webhook is an asynchronous HTTP message from another system, not a guaranteed function call. The sender may retry, events can arrive late or out of order, your endpoint may be down, and a slow handler can cause redelivery. A reliable integration assumes duplicate delivery and makes every transition visible and recoverable.

Acknowledge only after durable acceptance

Authenticate at ingress, persist the delivery, then process asynchronously
1 / ReceiveBound body size and retain exact raw bytes.
2 / VerifyValidate signature, timestamp and endpoint identity.
3 / PersistStore delivery ID and payload durably.
4 / AcknowledgeReturn 2xx quickly after durable acceptance.
5 / ProcessIdempotent worker, retries, alert and replay.

Do not perform slow business workflows before the acknowledgment. GitHub, for example, expects a 2xx response within 10 seconds and recommends asynchronous queue processing when work takes longer. That timeout is provider-specific; check each sender's delivery and retry contract.

Ingress is a security boundary

Use HTTPS and verify the provider's signature over the exact raw request body before parsing or mutating state. Store webhook secrets in a secret manager, compare MACs in constant time, and rotate secrets with a controlled overlap if the provider supports it. Validate timestamp windows when the signature scheme includes one. IP allowlists can be defense in depth, not a substitute for signature verification, and provider IP ranges can change.

Enforce a body-size limit, accepted content type, endpoint-specific secret and expected event types. Avoid secrets in URLs. Treat a valid signature as proof of sender authenticity and payload integrity, not proof that every field is safe or the event is relevant to this tenant.

Separate delivery identity from business identity

IdentifierUseDo not assume
Provider delivery IDDeduplicate transport retries and support redelivery tracing.It always identifies a unique business action across providers.
Business object/versionEnforce domain-level transition or compare source version.Events arrive in creation order.
Internal processing IDCorrelate queue attempt, side effects and logs.A retry is a new business event.

Persist the provider delivery identifier with a uniqueness constraint. Some senders reuse the same ID on manual redelivery; retain attempt history separately if support needs it. A distinct provider event may still describe the same business transition, so idempotency also belongs at the operation boundary.

Use a transactional inbox and idempotent effects

At minimum, insert the delivery record before returning success. A queue can be fed through a transactional outbox or a recoverable dispatcher so a crash between database commit and queue publish does not lose work. Consumers should record a processing state and make external effects idempotent too: use an idempotency key with the payment/email/CRM API where supported, or a local operation ledger. A database transaction cannot make an unrelated HTTP call atomic.

Retry with limits and make poison events visible

Retry transient network, rate-limit and dependency failures with exponential backoff and jitter. Honor provider Retry-After guidance where applicable. Do not retry permanent schema or authorization errors forever. After a bounded policy, move the message to a dead-letter or quarantine state with reason, attempts, timestamps and a safe payload reference. Alert on queue age and repeated failures; a dead-letter queue nobody watches is only a hidden failure bucket.

Expect out-of-order and stale events

Use source versions, object update times or domain state checks when order matters. Fetch current source state for consequential transitions if the provider contract allows it. Do not apply an older “subscription canceled” event over a newer reactivation simply because it arrived later. Record both provider event time and your receipt/processing times to diagnose lag.

Build replay and reconciliation into operations

Operators need to find a delivery by provider ID, tenant, event type, object and time; inspect verification and processing states; retry safely; and compare local state with the source system. Replay must preserve original identity while recording a new attempt. Redact personal data in logs and define payload retention. Reconciliation jobs catch events the sender never delivered or retries that exhausted.

Test failure paths, not just a happy webhook

Test invalid signatures, body mutation, duplicates, slow dependencies, database/queue outage, worker crash after commit, rate limiting, redelivery, out-of-order events and poison payloads. Assert that acknowledgment happens only after durable acceptance and that repeating a delivery does not repeat side effects. Use provider test tools and sandbox endpoints before production.

In summary

Reliable webhooks use a short authenticated ingress path, durable acceptance, fast acknowledgment, idempotent asynchronous processing, bounded retries, visible dead letters and safe replay. The boring architecture matters because payment, provisioning and customer state should not depend on one HTTP request arriving at exactly the right moment.

References