All storiesSelf-hosting

Kubernetes document collaboration: sessions, probes, and upgrades

Operate long-lived editing connections and conversion workers with supported topology, graceful drain, and meaningful checks.

Kubernetes document collaboration: sessions, probes, and upgrades: Session routing, Durable state, Probe design, Graceful drain.
Self-hosting / Office SDK

A Kubernetes document service needs a supported topology for durable state, editing sessions, and conversion jobs. Configure ingress, probes, resource budgets, and termination around that behavior. Verify an active document workflow during rollout and recovery; healthy pods alone do not establish a usable collaboration service.

Map supported service boundaries

Identify editor gateways, collaboration workers, conversion workers, databases, queues, caches, and object storage in the documented deployment. Determine which components are stateless and which contain data that must survive pod replacement. Use supported image versions and installation patterns rather than assuming every component can scale independently. Assign owners to external databases and storage even if they live outside the cluster. Record how the business application maps document identities and receives saved results. Cluster health can appear normal while a callback route is broken, so the complete document workflow must remain visible in the operating model.

Respect connection and session behavior

Collaborative editing may use persistent connections and product specific session routing. Configure ingress timeouts and routing according to the supported architecture, then test reconnection and draining. Do not assume a generic sticky session setting resolves every state ownership problem. Observe how an active session behaves when its serving pod terminates and whether users can reconnect without losing acknowledged work. For background conversion, test queued job handling and duplicate execution after a restart. Define which component owns retry and how results are matched to the intended business version.

The service has application, session, worker, and durable layers.
Figure 1. Supported state ownership determines what may scale or restart independently.

Use a concrete rollout rehearsal

Suppose a team plans an evening update while several engineers remain in a shared design workbook. In staging, reproduce that session and start the proposed rollout. Measure connection interruption, user messages, save progress, and reopening behavior. Verify that termination grace periods and any documented drain operation give the service time to stop safely. Read the upgrade notes for schema or state compatibility before allowing mixed versions. A rolling deployment is not automatically a zero downtime upgrade when old and new processes cannot share the same persisted state or communication protocol.

Make probes and limits meaningful

Separate startup, readiness, and liveness behavior. A worker performing a long conversion should not be restarted simply because a probe mistakes resource pressure for a dead process. Follow vendor recommendations, then verify them under representative documents and peak workload. Set resource requests and limits from measurements, including memory spikes during conversion. Monitor queue age, document startup failures, callback failures, and durable save progress alongside pod status. Back up persistent services and configuration, then restore into an isolated environment. A healthy deployment manifest is insufficient if the documents or their identity mappings cannot be recovered.

An illustrative rollout record

For each rehearsal, record the release pair, database compatibility notes, ingress settings, termination behavior, and the sessions active at the start. Use a workbook with at least two collaborators plus a queued conversion. Observe the user's saved milestone before and after the pod replacement.

rollout: release-A to release-B
active_work: shared workbook + queued export
observe: reconnect, save progress, callback delivery
accept: reopened content matches acknowledged work
stop: incompatible state or unexplained lost update

This is a test record, not a vendor manifest or a promise about any particular topology. A deployment that supports reconnection may still need a maintenance pause for a schema migration. Capture that limit explicitly rather than assuming rolling updates cover every release.

Signals that distinguish failure modes

SignalLikely investigation
Pods ready, document open failsSource retrieval, identity, storage, or callback path
Queue age risesWorker availability, job errors, resource demand
Repeated reconnectsIngress timeouts, routing, drain, and service behavior
Memory peaks during exportDocument mix, limits, and conversion concurrency

Investigate with correlated request and job identifiers. Restarting workers without identifying the failing stage can turn a diagnosable problem into repeated execution and additional load.

Active work is considered before, during, and after replacement.
Figure 2. A successful rollout command is not the final acceptance check.

Rollout Decision notes

  • Confirm supported images, topology, state ownership, and external dependencies.
  • Test persistent connection timeouts, reconnection, and graceful draining.
  • Rehearse an upgrade with active editing and queued conversions.
  • Measure resources and verify probes against demanding documents.
  • Restore content, mappings, configuration, and permission relationships together.

Further reading

Back to all stories

Keep reading.

All stories