All storiesOperations

Office document capacity planning: sessions, conversions, and save traffic

Distinguish open sessions, requests, conversions, and save traffic so capacity estimates reflect the real document workload.

Office document capacity planning: sessions, conversions, and save traffic: Workload units, User latency, Resource use, Peak scenarios.
Operations / Office SDK

Model capacity with the work the platform performs, not only the number of users. Idle sessions, first-time conversions, large workbook edits, and save bursts stress different components. Measure latency and correctness alongside queue age and resource use under a stated document mix and concurrency profile.

Choose workload units with clear definitions

Track active editing sessions, open and close rates, source retrieval bytes, conversion arrivals, queued work, save events, and export requests where relevant. Keep concurrency distinct from requests per second and from licensed limits that may use their own definitions. A tenant with many idle sessions can behave differently from a smaller tenant importing a large folder. Record document format and complexity distributions alongside volume. These dimensions make the workload reproducible and prevent a single user count from concealing the operations that actually constrain the system.

Workload unitWhat it describesOutcome to watch
Active sessionsSimultaneous supported collaboration contextsInteraction and reconnect behavior
New opens per intervalInitialization and source-retrieval demandTime to useful display
Conversion arrivalsFresh preview or export processingQueue age and completion time
Save eventsPersistence and publication demandDurable version convergence

Keep these operational units separate from contractual licensing terms. A license may count something called a session or request differently from your dashboard. Use the documented commercial definition for entitlement and the measured system definition for engineering, and explain the relationship instead of substituting one for the other.

When reporting results, include the file mix, cold or warm cache state, client profile, infrastructure, concurrency, and observation window. A number without those conditions cannot be reproduced. Report failures and rejected work as well as completed operations; admission control may protect latency precisely by declining excess demand.

Four operational units describe activity more precisely than total user count.
Figure 1. Record workload definitions so licensing terminology does not obscure engineering measurements.

Measure user outcomes alongside infrastructure

Observe latency, failure rate, and queue age together with CPU, memory, storage throughput, network traffic, and database behavior. High utilization is not automatically a problem if outcomes remain within requirements, while low average utilization can hide a saturated single worker or a long tail of slow documents. Break down opening into retrieval, initialization, and useful display where possible. Report percentile or distribution information with sample size and conditions rather than only an average. Include failed operations so the fastest successful subset does not define the reported result.

Build a representative load profile

Model normal activity, expected peaks, batch imports, and recovery after a temporary outage. Use representative synthetic or authorized documents with the complexity seen in production. Separate cold and warm cache behavior. Increase demand gradually while watching downstream dependencies, and stop according to agreed safety limits in the test environment. Do not extrapolate linearly across a bottleneck without evidence. More workers may shift pressure to object storage, databases, network links, or a shared conversion dependency instead of increasing useful throughput in the same proportion.

Compare cached reading with new workbook processing

Run one scenario with many users reading already prepared previews and another with fewer users opening new complex workbooks while a batch import runs. Keep the environment fixed and record open latency, queue age, save completion, and resource use. The purpose is to identify different limiting stages, not to produce a universal benchmark. If the second scenario overloads conversion, test admission limits or scheduling before assuming every component needs more replicas. Document the resulting operating envelope with the exact workload mix that produced it.

Prepared viewing, new workbook editing, and mixed batch work create distinct demand profiles.
Figure 2. Publish the tested workload mix with the result instead of a universal user-capacity claim.

Acceptance and reporting conditions

  • Define target outcomes for opening, editing, saving, and exporting.
  • Track workload units separately from licensing terminology.
  • Identify the first saturated dependency under each representative peak.
  • Reserve capacity for retries, failover, and maintenance where required.
  • Repeat affected scenarios after changing document limits or infrastructure.

Let the bottleneck determine the next experiment

Use the bottleneck to choose the next experiment. If queue age rises while converters remain busy, test scheduling, admitted load, or converter capacity within downstream limits. If workers wait for object storage, adding more workers may only multiply waiting connections. If browsers run out of memory on one workbook pattern, a backend replica change may not address the user-visible problem.

Reserve a scenario for recovery: restart a worker or briefly interrupt a dependency, then watch retry demand and backlog drain. Normal steady-state headroom is not the same as recovery capacity. Record how long the service takes to return to its agreed operating range and whether current users remain able to save during that period.

Further reading

Back to all stories

Keep reading.

All stories