Trace IDs for document opening, conversion, and save failures
Connect launch, retrieval, conversion, and save evidence without placing document content or credentials in telemetry.

Give the user's document operation a durable identifier and connect each request, queue job, and callback to it. A request trace explains one execution; an operation ID explains the complete attempt, including retries. Keep both relationships without putting document text or signed URLs into telemetry.
Operation identity versus request identity
One user action can produce many requests and several retries. Use a business operation ID for the overall attempt and request or trace context for individual executions. Link them explicitly. A retry may create a new trace while still belonging to the same import or save operation. Document IDs provide useful context but are not substitutes for attempt identifiers, because many operations can affect one document simultaneously. When adopting a standard trace format, validate incoming context and follow the supported instrumentation model rather than treating arbitrary client text as trusted metadata.

Cross queues and provider callbacks explicitly
Record the operation relationship when creating a queue job and restore it when a worker starts. External services may not forward your trace headers, so retain a mapping to their documented job or event identifiers. A callback can then be linked to the initiating operation without relying on timestamp proximity. Do not require every system to share one tracing implementation before gaining value. Consistent structured fields and a reliable mapping table can connect evidence across boundaries that full distributed tracing cannot currently cross.
A practical correlation record
| Field | Question it answers |
|---|---|
| Operation ID | Which user intention are we following? |
| Trace or request ID | Which execution produced this observation? |
| Stage and outcome | Where did work succeed or first fail? |
| Source version | Which content state was involved? |
| External event or job ID | How does a provider callback connect to our operation? |
Use stable, documented field names across your own services. A gateway's request ID and a converter's job ID can both be valuable, but a support operator needs a recorded relationship between them. Asking someone to correlate by filename and approximate time is fragile when several users open the same document together.
Choose where the user sees the support identifier: an error panel, import activity record, or copyable failure detail. It should not require developer tools for every incident. Verify that the identifier can be searched by the people responsible for support, with the correct tenant restrictions. Telemetry that exists only in a system nobody on call can access does not complete the diagnostic path.
Collect evidence without collecting the document
Useful fields include operation type, stage, safe document identifier, tenant scope, outcome, duration, and normalized error category. Exclude access tokens, signed URL query strings, full callback secrets, and document text from routine telemetry. Classify identifiers and apply retention and access controls appropriate to their sensitivity. Log enough to tell whether authorization, source retrieval, conversion, or publication failed. A single generic error message repeated by every component adds volume without helping operators identify the first failing stage or the system responsible for recovery.
Support-readiness checks
- Show a copyable operation or support ID in actionable error states.
- Preserve context across worker restarts, retries, and external event mappings.
- Use consistent names and units for stage durations.
- Redact credentials before logs leave the service boundary.
- Test that an operator can identify the owning component from one synthetic failure.
Follow a denied presentation source to its first failure
For a test presentation, introduce a storage access denial. Start from the operation ID shown in the user facing error, locate the launch request, then follow its source retrieval and worker records. Confirm that the first failure is visible as a storage authorization outcome rather than only a later editor timeout. Repeat through a queued retry and check that the relationship remains intact. The exercise should be possible for an authorized support operator without requesting the presentation contents or asking the user to expose a signed download link.



