Integration reliability
Why webhooks fail quietly—and how to make them recoverable
A webhook response is not proof that the business event finished. Design for retries, duplicates, missing records, and useful recovery.
A successful response can hide an unfinished job
A provider sends an event. Your endpoint returns a successful response. It is tempting to call the integration complete. But the database write, CRM update, notification, or downstream API call may still fail after that response is sent.
This is why webhook reliability is not mainly an HTTP problem. It is a state-tracking problem. The useful question is not “did we receive a request?” It is “can we prove what happened to this business event?”
Separate receipt from processing
The receiving endpoint should do a small amount of work: authenticate the sender, validate the envelope, assign or preserve an event identifier, store the event durably, and acknowledge receipt. Slower business work can happen after that durable handoff.
This boundary makes timeouts less dangerous. It also gives the team an original payload to inspect when a later step fails.
- Record the provider event ID and received time.
- Keep the original payload or a safe, auditable representation.
- Track processing states such as received, applied, failed, and ignored.
- Attach a clear failure reason and attempt count.
Assume retries and duplicates
Providers retry when they do not receive a timely response. Networks also create ambiguous outcomes: your system may finish the work even though the provider never sees the acknowledgement. The same event can therefore arrive more than once.
Processing must be idempotent: applying the same event again should not create a second lead, charge, ticket, or email. A stable event ID, a uniqueness rule, and a transaction around the state change are more reliable than a quick lookup followed by an unprotected write.
Build the recovery path before you need it
A dead-letter queue is only useful when someone can understand and act on it. Show the event, failure reason, last attempt, affected record, and whether replay is safe. Protect replay with the same duplicate controls as normal processing.
The operational test is simple: if an event fails at 2 a.m., can a person find it the next morning, understand the impact, and recover it without editing production data by hand? If not, the integration is not finished.
Better follow-up and less busywork for small businesses. Have a task like this? Tell me what happens today and what you’d like to change.
Book a 20-minute call