All storiesIntegration

Webhook idempotency for document saves: keys, retries, and recovery

Design callback handling around authenticated events, durable acknowledgement, and idempotent effects.

Webhook idempotency for document saves: keys, retries, and recovery: Event receipt, Operation key, Durable decision, Duplicate check.
Integration / Office SDK

A document save webhook can arrive twice even when the first delivery succeeded. Authenticate it, identify the logical operation, and make publication repeatable without creating another version. The deduplication record must also let a crashed worker resume; merely marking an event as seen can permanently lose a save.

Accept only after a durable decision

Determine what the provider expects from the callback response and how its retry policy works. Where asynchronous processing is permitted, authenticate and validate the request, record an event durably, and acknowledge that acceptance. A memory only queue cannot support this promise through a process crash. If the contract requires synchronous completion, make the work bounded and make every retry safe. In either design, distinguish acceptance from successful downstream processing in logs and monitoring so an acknowledged event does not disappear into an invisible failure queue.

A callback is validated and durably recorded before asynchronous acknowledgement and processing.
Figure 1. Use this sequence only when the provider contract permits asynchronous completion.

Choose a stable operation key

Use the provider's documented stable event identifier when available, scoped to the integration and tenant. If no such identifier exists, derive a key only from fields whose semantics you understand, such as document, source version, and event type. A timestamp alone is usually insufficient. A hash of the entire payload may also fail if retries contain refreshed links or changing metadata. Record the payload fingerprint alongside the key so that conflicting requests using the same operation identity can be detected and investigated rather than silently accepted.

Track incomplete work as well as duplicates

A separate check followed by an insert is vulnerable when two workers receive the same event simultaneously. Use a uniqueness constraint or equivalent atomic claim. Track pending, completed, and failed processing states with explicit recovery rules. If file storage and database updates cannot share one transaction, retain intermediate references and reconcile incomplete work. Deduplication should prevent duplicate business effects without permanently suppressing a legitimate retry after a worker crashed between claiming an event and writing its resulting document version.

event identity: integration + tenant + provider event ID
processing state: accepted → retrieving → published
result reference: business document + published version

This is an illustrative event model, not a provider payload. The event identity needs a uniqueness constraint. The processing state needs a recovery policy. The result reference lets a duplicate return or locate the already completed result without replaying secondary effects.

Consider a worker that claims the event and crashes before fetching the saved file. A second delivery must discover incomplete work, not interpret the claim as completion. Use a durable job, a bounded processing lease, or another recovery mechanism consistent with your database and queue. Preserve the original identity when work is retried.

Choose deduplication retention using the provider's delivery and replay behavior plus your own incident recovery process. If an operator can replay an event after the deduplication record expires, a historical replay may create a fresh effect. Either retain the identity long enough, verify the published source version independently, or restrict that replay path. Document which protection applies so maintenance does not accidentally remove it.

Duplicate delivery is handled according to whether work is complete, active, or interrupted.
Figure 2. A deduplication claim is not proof that the business effect finished.

Force a lost response and a worker crash

For a test document, let the callback complete its database change, then intentionally close the connection before sending the response. Deliver the same event again. The expected outcome is one published version and a response consistent with the provider contract. Repeat while two workers process the event concurrently, and once with a changed payload under the same event ID. Use synthetic credentials and a test object. These experiments expose gaps that a happy path test with one successful request cannot reveal.

  • Record authentication failures separately from temporary storage or database failures.
  • Set a retention period for deduplication records that covers the documented delivery window and replay process.
  • Expose counts and age of pending or failed events.
  • Provide a controlled replay operation that preserves the original event identity.
  • Verify that a duplicate never triggers a second external notification or approval transition.

Further reading

Back to all stories

Keep reading.

All stories