Webhook Architecture for Asynchronous Document Extraction APIs
Separate document ingestion from extraction work to handle long-running jobs reliably.

A developer submits a PDF to an extraction API and waits on an open connection. Somewhere between OCR, layout analysis, and field validation, the connection drops or times out, and the work hasn't even finished yet. Document extraction jobs routinely exceed what a synchronous HTTP round-trip can absorb, so building the system around a request-response call is a mismatch from the start, not a performance problem to tune away. If a webhook provider enforces a short response window, extraction on a complex PDF can blow past it, so the delivery gets marked failed even though the server is still working on a result nobody is waiting for anymore. Polling looks like the fix but it isn't: it trades timeouts for repeated outbound requests, a lag between completion and detection, and server load that climbs with every document added to the queue. Extraction needs to be treated as a job with a lifecycle, not a request that either returns in time or doesn't.
The two-layer queue-first pattern that separates ingestion from processing
The fix for extraction's latency is a split: one layer that receives and acknowledges, a second layer that processes, running as separate services rather than two functions bundled into the same handler.
The ingestion endpoint does three things and nothing more: it authenticates the request, pushes the payload onto a message queue, and returns a 202 Accepted. No extraction work happens here. The worker pool is the second layer, and it pulls jobs off that queue to run the actual extraction. If a worker dies mid-job, the message stays in the queue and another worker picks it up, so no event gets lost to a crashed process.
The queue itself can be Apache Kafka, RabbitMQ, or AWS SQS. Each gives the system message durability and fault tolerance: a message sitting in a queue survives a worker crash in a way that a thread holding state in memory never does. The deployment separation is what actually delivers the benefit here. When ingestion and processing run as distinct services, a failure in the worker pool can't take down the endpoint, because it's still accepting new documents from the provider.
That 202 response does real work. It tells the upstream provider the event was received and accepted, which stops the provider's retry clock, but it makes no promise about when the extraction will actually finish. Document extraction APIs are built around this latency reality from the ground up: extraction is a composite job spanning OCR, layout analysis, field validation, and confidence scoring, and that work routinely stretches well past any synchronous HTTP timeout window. The webhook pattern is what the architecture requires once extraction latency is measured in seconds and accuracy guarantees demand complete, auditable processing.
One ingested event can fan out to multiple queues at once, so it drives an audit log, an extraction worker, and a notification service in parallel without wiring those consumers together.
Signature verification on the raw request body before any parsing
Before a payload gets queued, it has to be trusted, and trust starts with signature verification on the raw, unparsed request body. This is a precise technical requirement, and most hand-written implementations get it wrong in the same way.
The standard approach has the provider sign the raw body with a shared secret using HMAC-SHA256, then include the resulting hash in a request header. The receiver recomputes that hash over the same raw bytes and compares it against what arrived. A different hash for a completely legitimate payload appears when a JSON framework parses the body before verification runs: parsing can reorder keys or alter whitespace. In practice, this means capturing the raw body before any parsing middleware touches it, computing the hash over those exact bytes, and comparing the result with a constant-time comparison function rather than a standard string comparison, which can leak timing information an attacker could use to forge a signature.
Timestamp validation belongs in this same step. If the event timestamp falls outside an acceptable window, you can catch a replay attack, where someone resubmits a captured valid request later.
Providers don't all sign the same way, either. Newer providers tend to converge on one of two styles: Standard Webhooks headers (webhook-id, webhook-timestamp, webhook-signature, the approach OpenAI uses), or asymmetric signatures verified against public keys pulled from a JWKS URL. The ingestion layer has to handle provider-specific verification logic, so you can't hard-code it to a single scheme.
Extraction jobs often carry sensitive financial documents or personally identifiable information, so this verification step has to work as a security boundary. If any parsing step touches the byte stream before verification runs, it can mask a maliciously altered payload or hide a silent corruption that verification would otherwise have caught. At high volume, a systematic bug in this step cascades: failed verification triggers provider retries, the bug fails those retries too, and the queue fills with phantom retries while legitimate new submissions get delayed or dropped.
Idempotent processing as the contract between the queue and the worker
Webhooks are delivered at-least-once. Duplicates and out-of-order events aren't edge cases, they're guaranteed to happen at scale, and a worker that isn't built to absorb them will corrupt data the first time a duplicate arrives. Idempotency is the contract that makes the queue-first architecture safe to run.
Three patterns handle this in production, and each comes with its own tradeoff. Fetch-before-process treats the webhook as a notification only: the worker uses the event ID to fetch current state from the source API before acting, which works well when the latest state is what matters but gets expensive if rate limits are tight. Insert-or-update with timestamp checks uses conditional writes: you upsert with a WHERE clause comparing the incoming timestamp against the stored one, so a duplicate event carrying an older timestamp can't overwrite a newer state. Deduplication by event key stores a unique identifier, the webhook's own ID or a combination of document ID and event type, in an atomic store before processing begins; if that key already exists, the worker skips processing and simply acknowledges, and the atomic write is what prevents a race between two workers that both receive the same duplicate at nearly the same moment.
Out-of-order delivery raises a related problem. A "job completed" event can arrive before a "job started" event, and workers need to handle or discard that kind of out-of-order transition instead of assuming events always show up in sequence.
This separation of concerns matters especially in document extraction, where you can have a single submitted PDF fail partway through layout analysis and need a retry. A queue-backed worker pool ensures that failure doesn't lose the job, and if a document gets resubmitted, it gets processed idempotently rather than duplicating the extraction result or firing the downstream webhook twice. At the field level, if a duplicate event causes the same extracted record to post twice into an accounts-payable system, you get a duplicate payment. The cost of skipping idempotency here is financial and specific, not theoretical.
Retry logic, exponential backoff, and dead-letter queues for permanently failed jobs
A webhook architecture without a dead-letter queue has no floor under it. Transient failures turn into permanent data loss, and no record is left of what disappeared or why.
Retry logic needs exponential backoff with jitter built in. Picture a downstream database coming back online after an outage: if every queued job retries on the same fixed interval, every worker converges on that database at once, and the resulting surge of traffic can push the database back into failure right after it recovered. Jitter spreads those retries out so recovery actually sticks.
AWS SQS handles a version of this automatically through its Visibility Timeout: a message pulled from the queue becomes invisible to other consumers for a set duration, and if the worker doesn't delete it within that window, the message reappears and gets retried without any explicit scheduling logic. RabbitMQ takes a more manual approach, using message acknowledgement and rejection: a worker that hits a retriable error rejects the message with a delay before requeue, while a worker that hits a non-retriable error routes the message straight to a dead-letter exchange.
Once messages exhaust every retry, they land in the dead-letter queue. It needs active monitoring and alerting, because a DLQ that fills up silently is data loss with extra steps.
The provider runs its own retry schedule independently of whatever retry logic the consumer builds internally. A consumer that fails to acknowledge an event can end up receiving both a provider-side retry and an internal queue retry for the same event at the same time, and idempotency is what keeps that collision from corrupting state. At high document volume, this is how retry storms happen: a systemic bug, like the signature verification failure described earlier, causes every delivery to fail, every provider retries on its own independent schedule, and the queue fills with a multiplying backlog of the same events. Fixing it takes two separate steps: correcting the underlying bug, and replaying only the messages sitting in the dead-letter queue rather than re-ingesting everything from the provider again.
Payload versioning and schema stability across provider updates
If a provider updates its payload format, a webhook integration that works correctly today can break silently, even when the consumer's own code and infrastructure haven't changed.
Providers can add, rename, or remove fields in a webhook payload without treating it as a breaking change on their end. If a consumer hard-codes exact field names, or assumes one fixed JSON structure, it starts producing wrong results or silent nulls with no error raised anywhere in the pipeline. The defensive pattern is to parse payloads permissively: accept unknown fields, treat missing non-critical fields as absent rather than throwing an error, and validate only the fields the consumer actually depends on. A well-designed extraction API puts a schema version field in every webhook payload, so you can route different versions to different parsing handlers instead of maintaining one handler that has to support every historical shape the provider ever shipped. Building the parsing layer to treat the payload as untrusted structured data, rather than a guaranteed contract, is what keeps a provider's routine update from becoming an outage.
Confidence scores and field-level routing in the extraction result handler
When an extraction result reaches the worker, you need field-level confidence scores to function as routing signals, not as decoration. A high confidence score attached to a wrong value is the most dangerous outcome a document extraction pipeline can produce, because it's the one result nobody goes back to check.
A confidence band near 1.0 is a reasonable candidate for automated acceptance, but that's precisely the band where a wrong value does the most damage: it never gets reviewed, so the error propagates downstream untouched. Mid-range scores call for judgment specific to the field and the use case. Low scores should route to human review or get rejected outright, and anything below a defined floor should trigger a flag for manual reprocessing.
Calibration drift is the failure mode to watch for. A model accurate at deployment can grow less accurate on new document formats over time while its confidence scores stay just as high, because the model has become certain about the wrong value. A confidence score means nothing if the model is reporting certainty about a hallucinated value instead of one traced back to actual pixels on the page. The result handler should also run cheap, field-type-specific checks of its own: string fields for plausible character patterns, numeric fields for range validity, date fields for calendar coherence. These catch a category of extraction error before it ever reaches a downstream system, independent of what the model's own confidence score says.
For extraction APIs that offer a contractual accuracy guarantee, the confidence scores in the webhook payload and the SLA's accuracy definition have to line up. A per-field SLA only means something if the confidence score reported is calibrated against ground truth. Invofox's field-level confidence scores are tied back to the page, region, and source they came from, and its per-field accuracy SLA, the Perfect Docs Guaranteed commitment, gives the routing logic in the worker a known, contractual baseline to operate against.
Observability, alerting, and replay as the operational layer the architecture requires
Every layer described so far, ingestion, queue, worker pool, dead-letter queue, confidence routing, produces events that need to be visible somewhere other than inside the system itself. A queue-first architecture without observability is a black box that happens to be reliable when nothing goes wrong and undiagnosable the moment something does.
Tracking needs to run at the level of the individual event: when it was received, when signature verification passed, when it landed in the queue, which worker picked it up, how long processing took, and what the final status was. Without that trail, a document that silently failed extraction looks identical to one that was never submitted at all, and the two require completely different responses.
Alerting has to sit on top of the dead-letter queue specifically, because a DLQ that fills without anyone noticing is data loss wearing a different name. Alerting also belongs on retry rates: a sudden spike in retries is often the earliest signal of a systemic problem, like the signature verification cascade or a downstream dependency going dark, long before the dead-letter queue itself starts filling up.
Replay is what closes the loop once a bug gets fixed. The correct response to a retry storm or a verification bug is to replay the specific messages sitting in the dead-letter queue, process them against the corrected code, and confirm each one clears idempotently. That distinction, replaying from the consumer's own durable store, is what turns a production incident into a bounded, recoverable event.


