All storiesOperations

Hardware sizing for a document platform: what to measure first

Translate representative editing, conversion, storage, and failure scenarios into defensible hardware requirements.

Hardware sizing for a document platform: what to measure first: Supported topology, Workload envelope, Resource limits, Failure reserve.
Operations / Office SDK

Size document infrastructure from the supported topology and a measured workload of editing, calculations, opens, conversions, storage, and recovery. Record latency as well as resource use. User counts and file size alone cannot justify cores, memory, disk performance, or reserve capacity for the failure cases you promise.

Define the operating envelope

Record active users, peak sessions, document types, file size distribution, opening rate, conversion jobs, and retained data growth. Add latency and recovery expectations. Obtain current minimum requirements and recommended topology for the actual product version, then treat them as a starting constraint rather than proof of workload capacity. Separate editor, conversion, database, storage, and supporting services according to the supported deployment. If resources are shared, document the competing workload. A machine with adequate nominal specifications can still be unsuitable when other applications consume its memory, disk throughput, or network during the same business peak.

Conversion, editing, persistence, and recovery activities mapped to resource measurements.
Figure 1. Use simultaneous work when it occurs simultaneously in the business.

Write a performance specification, not just a parts list

A procurement request should explain the tested operating envelope and the infrastructure properties it depends on. Disk capacity and processor count are necessary descriptions, but they do not capture database write latency, virtual-machine contention, or transfer performance to the actual storage endpoint.

Compute and memory
Record the simultaneous editing and conversion workload, resource use, and any dedicated-resource assumption.
Storage
Record both retained capacity and the performance needed by the measured database and object operations.
Network
Record the participating zones and source-transfer path, including the file sizes used in the test.
Failure reserve
Record the supported reduced topology, reconnect burst, and allowed degraded behavior.

Compare infrastructure candidates against the same workload and software configuration. If one uses shared resources and another dedicated resources, retain that distinction rather than attributing every difference to nominal core count.

Attach an expansion trigger to the request: a sustained operating metric or workload change that causes review before the service breaches its target. The trigger should allow for procurement lead time. That gives operations a way to act on the sizing evidence after launch instead of treating the initial purchase as permanently sufficient.

Measure each resource in context

Observe processor use during conversion and calculation, memory across active sessions, database wait time, storage latency, and network transfer. Include warm and cold starts because caches can change the result. Record sustained behavior as well as short peaks. Hardware specifications such as core count and disk capacity omit performance characteristics and virtualization contention. Compare the intended infrastructure using a repeatable workload and the same software configuration. Keep the measurement method with the results so the procurement team understands whether an estimate assumes dedicated resources, particular storage performance, or a tested network route.

Build a planning office example

A planning office expects 80 active sessions, a burst of large presentation imports, and a weekly spreadsheet review. Create a harmless corpus reflecting those activities and test the supported deployment on candidate infrastructure. Run imports during the review instead of measuring them separately. Observe opening latency and durable save completion as resources approach saturation. Increase load gradually to identify the first limiting component. The scenario numbers are illustrative, not a universal recommendation. Their purpose is to show how procurement evidence should represent simultaneous work rather than a benchmark built from small empty documents.

Corpus, controlled test, resource limit, and purchase scope sequence.
Figure 2. The test scenario is evidence only for its stated operating envelope.

Reserve for growth and failure

Estimate growth from measured activity and retention policy, then identify the supported scaling path. If the design requires service continuity after a node fails, test remaining capacity and reconnect behavior rather than assuming unused average CPU provides enough reserve. Include test and recovery infrastructure in the purchase scope. Specify how resource expansion is triggered and what lead time procurement needs. Avoid padding every component by an unexplained percentage; targeted headroom is more defensible when linked to a workload or failure case. Record limitations that remain, such as delayed conversions during a reduced capacity period.

Sizing request

  • Attach software version, topology, and workload assumptions.
  • Report observed bottlenecks and user latency at the tested load.
  • Specify storage performance and network needs alongside capacity.
  • Include agreed growth, failure, test, and recovery requirements.
  • Define the measurements that will trigger future expansion.

Further reading

Back to all stories

Keep reading.

All stories