All storiesOperations

Document conversion queues: backpressure, fairness, and retry limits

Use admission limits, fair scheduling, and visible queue states to preserve predictable conversion behavior under bursts.

Document conversion queues: backpressure, fairness, and retry limits: Admission limit, Work classes, Worker capacity, Queue status.
Operations / Office SDK

Limit both accepted work and active conversion workers. Worker concurrency alone cannot protect an unbounded queue from expired source links, growing storage use, or hours of waiting. Use queue age, queued bytes, and per-tenant demand to decide when to accept, defer, or reject a job.

Admission, scheduling, and execution are separate controls

A worker concurrency limit protects active CPU and memory, but it does not limit queued objects, expired credentials, or waiting users. Set explicit limits for accepted jobs, queued bytes, job age, and per tenant demand. Decide which requests can wait and which should be rejected with a retryable response. Admission should account for the source being available long enough to process. A queue full of jobs whose source links have already expired is not useful capacity; it is deferred failure that requires additional cleanup and retries.

There are at least three controls to tune independently. Admission decides whether a new job enters the system. Scheduling decides which accepted job runs next. Execution limits bound the resources consumed while it runs. Combining all three into one worker-count setting makes overload harder to diagnose.

ControlUseful signalFailure it contains
AdmissionQueued bytes, age, tenant quotaUnlimited accumulation of work
SchedulingInteractive versus batch demandOne large batch delaying unrelated users
ExecutionMemory, runtime, downstream limitsOne job exhausting a worker or dependency

Retry traffic must consume the same controlled capacity. A failing document that immediately requeues itself can monopolize processing even when new submissions are limited. Apply bounded attempts and delay, and identify permanent input errors early. During an incident, report the oldest waiting job and the reason workers are unavailable. A queue length of twenty means little without knowing whether those jobs are small previews, enormous presentations, or repeated failures that cannot complete.

Admission, scheduling, execution, and retry controls protect different stages of conversion.
Figure 1. Scaling workers does not replace a policy for work that cannot be accepted.

Keep bulk work from starving interactive previews

Use available signals such as format, source size, and known workbook or presentation characteristics to separate unusually expensive tasks when helpful. These estimates are imperfect, so enforce runtime limits and observe actual resource consumption. Avoid letting one large batch block every interactive preview. Fair scheduling can allocate work among tenants or priority classes, but document the policy and prevent unlimited priority escalation. Reserve enough capacity for recovery and routine interactive work so a bulk import cannot consume every worker and erase the intended service experience.

Waiting and cancellation need honest states

Distinguish accepted, waiting, processing, completed, failed, and canceled states when the backend can support them reliably. Show progress percentages only when there is a meaningful denominator; otherwise explain the current stage. Cancellation should stop queued work promptly and attempt to stop running work according to engine capabilities. Keep already completed outputs consistent with the cancellation result. Do not invite users to click repeatedly when processing is merely waiting, because duplicate requests amplify load and make completion ownership harder to understand.

Reproduce a burst with one failing document

Queue a representative batch of test documents while several users open individual files. Observe queue age, active workers, memory, source expiry failures, and interactive completion times. Introduce one document that consistently fails and confirm that its retries are bounded and delayed. Then stop a worker and verify that claimed jobs are recovered without uncontrolled duplication. This scenario tests the scheduling policy under pressure. It should produce recorded observations for the actual environment, not a generalized throughput claim derived from an artificial best case.

  • Alert on oldest waiting job and queued bytes, not only queue length.
  • Limit retries and distinguish permanent format errors from temporary infrastructure failures.
  • Apply per tenant admission controls to contain large batches.
  • Make rejection and cancellation understandable in the interface.
  • Scale workers only within downstream storage, database, and licensing constraints that actually apply.
Bulk work, interactive requests, and failures are combined in one controlled queue test.
Figure 2. The queue is healthy when accepted work completes predictably, not merely when it accepts more jobs.

Further reading

Back to all stories

Keep reading.

All stories