Data-centre infrastructure handbook
How a data centre works
A field guide to the physical systems, operating records and decisions behind Canadian data-centre capacity.
Ten connected systems ยท Reviewed July 27, 2026
Site and building
The site and building set the limits
The first infrastructure decision is not the server. It is the piece of land, the utility service available to it and the building that must protect every system placed inside.
A location screen should separate disqualifying constraints from qualities that can be priced or improved. Flood exposure, wildfire interface, access to several carrier routes, utility connection options, fuel delivery, water availability, emergency access, labour and the permitting path can determine whether a project is workable. A metro label is not enough. Conditions can change from one utility feeder, municipality or parcel to the next.
The building must support equipment movement, floor loading, structural bracing, fire separations, secure circulation and maintainable plant rooms. Roof and yard space can become scarce once generators, cooling equipment, tanks, transformers and service access are laid out. Future phases need protected routes and tie-in points. If expansion requires shutting down the operating hall, the original layout has already spent part of the facility's resilience.
Canadian model codes are a starting reference, not a permit. Provinces, territories and municipalities adopt editions with their own timing and amendments. Record the authority having jurisdiction, adopted edition, occupancy assumptions, permit milestones and any alternate solutions. The public directory can identify a facility or project, but only the project team can establish what applies to a particular site.
Electrical path
Power is a chain, not a headline capacity
A megawatt number says little until the electrical one-line shows how power reaches the critical load, which parts are shared and what happens when one part is unavailable.
The normal path can include the utility point of connection, medium-voltage switchgear, transformers, low-voltage switchboards, UPS input and output gear, static and maintenance bypass, PDUs, remote power panels, busway and rack distribution. Standby generation adds fuel systems, starting systems, paralleling controls and transfer logic. Each boundary has a rating, protective device, control dependency and maintenance state.
Redundancy labels need a defined load and operating condition. An N+1 plant may still have a common control supply, shared fuel path or distribution segment. A dual-corded rack is not protected if both feeds return to the same upstream device. Review the one-line against the physical installation, equipment nameplates, protection study and sequence of operation. Keep normal, maintenance, emergency and black-start states separate.
Capacity planning should use the limiting component in the active path, not the sum of equipment nameplates. Reserve, ambient derating, battery recharge, cooling restart and future load steps all matter. Utility capacity, building electrical capacity and currently saleable IT capacity are different quantities. Public marketing often describes only one of them, which is why project-specific evidence is essential.
UPS and batteries
The UPS buys time for the transfer
A UPS must carry the defined critical load through a disturbance, support the transfer strategy and produce enough warning to act before stored energy is exhausted.
The operating record should identify topology, model, firmware, rated power, actual loading, bypass arrangement, battery chemistry, string configuration and environmental limits. Runtime is a condition-based result, not a permanent specification. Load, cell temperature, age, connection resistance, charger settings and the weakest unit in the string all affect it. A battery fleet needs traceable readings and exceptions at unit or jar level.
Preventive maintenance starts before the technician arrives. Confirm the approved maintenance state, switching procedure, bypass capacity, redundancy available, change window, access rules and escalation contacts. The visit should include visual and environmental inspection, alarm and event review, fan and capacitor condition where applicable, battery measurements, connection checks within the approved scope and restoration evidence. Intrusive tests and transfers require a specific method of procedure.
Trend comparable readings rather than collecting isolated pass marks. A unit that still supports the load may be deteriorating faster than its peers. Track temperature, conductance or impedance, float voltage, connection condition, discharge history and recurring alarms. Define replacement criteria, approved parts, full-string policy, disposal route and the temporary risk created during replacement. Monitoring helps find drift, but it does not replace physical inspection or a tested response plan.
Cooling
Cooling is heat transport with controls attached
Cooling capacity has to reach the rack, carry heat out of the room and reject it outdoors under the actual weather, water and failure conditions.
The thermal path may include rack airflow, containment, room units, chilled-water loops, pumps, heat exchangers, chillers, cooling towers, dry coolers or direct liquid-cooling loops. Nameplate tonnage does not prove usable IT capacity. Air distribution, coil approach, water temperatures, fouling, pump head, control stability and simultaneous maintenance determine what the plant can carry.
Start with the equipment environmental envelope and the rack density plan. Measure temperature where IT equipment takes in air, not only at a wall sensor. Look for recirculation, bypass air and pressure imbalance. High-density or liquid-cooled deployments need clear ownership of cooling distribution units, secondary loops, leak detection, water quality, isolation valves and the boundary between facility and IT teams.
Efficiency changes must be tested against resilience. Raising a setpoint, widening a control band or changing fan logic can save energy but also reduce recovery margin or expose an unstable sequence. Record operating modes for normal load, one component out, loss of utility, restart and extreme weather. Water use, freeze protection, plume, noise and local restrictions belong in the design file as early as energy use.
White space
The rack is where every infrastructure promise meets
White-space design converts upstream power, cooling and connectivity into a rack position that can actually carry the intended equipment.
A useful rack schedule records cabinet dimensions, usable units, floor or slab loading, power feeds, receptacles, metering, target and maximum load, network paths and cooling method. It should also show reserved positions, staging areas and the route used to move equipment without crossing secure or live-work boundaries. Total room area does not reveal how much deployable capacity remains.
Rack power needs a phase and circuit plan. Dual feeds should be tested back to their independent source paths, and each feed must be able to carry the agreed failure state. Metering at branch, PDU or busway level makes stranded capacity visible. Cable routes need bend radius, separation, support, grounding and firestopping discipline. Poor patching and unmanaged blanking panels become operating problems long before they appear in a high-level availability claim.
Changes at the rack should update the capacity model. A new GPU platform can alter weight, inrush, harmonics, airflow, liquid demand, rear clearance and service procedure at once. Use a pre-install review that joins facility, network, security and IT owners. The decision is not only whether the rack can be powered today, but whether it can be installed, cooled, maintained and removed safely.
Connectivity
Two carriers do not always mean two routes
Resilient connectivity depends on physical diversity, service design and tested recovery, not the number of logos shown on a sales sheet.
Trace each service from the application through the network edge, meet-me room, entrance facility, outside plant and remote endpoint. Two carriers can share a conduit, street crossing, bridge, pole line or upstream building. A second entrance can still return to the same metro ring. Route evidence should state what has been verified, by whom and on what date without publishing sensitive details.
Cloud on-ramps and internet exchanges are useful connection points, not automatic relationships. Proximity does not prove that a facility has a specific service, carrier or point of presence. Confirm the ordered product, demarcation, cross-connect, port, protection class, latency test method, support boundary and repair commitment. Keep carrier diversity, cloud diversity and application recovery as separate tests.
The management network needs its own failure analysis. UPS, generators, cooling controls, security systems and remote hands may depend on authentication, name resolution, time, alerting and out-of-band access. Network zones and controlled flows help contain an incident, but operators still need a documented path for emergency access and manual operation when normal corporate services are unavailable.
Protection
Fire, access and life safety have to work together
A data centre needs rapid detection and controlled response without putting occupants, emergency crews or the electrical recovery path at risk.
The protection strategy can include early-warning smoke detection, spot detection, pre-action sprinklers, clean-agent systems, portable extinguishers, compartmentation and firestopping. The adopted fire code, building construction, occupancy, insurer requirements and authority having jurisdiction shape the final system. A technology standard can inform the design but does not replace local code review.
Cause-and-effect records should show what each alarm does to air handling, dampers, doors, suppression release, power and monitoring. Tests must include interfaces, not only individual devices. Changes to cable penetrations, containment, battery rooms or cooling equipment can affect the fire strategy. Keep impairment permits, isolation limits, fire watch requirements and restoration checks with the maintenance procedure.
Physical security should create clear zones from the property line to the rack. Visitor control, loading access, cameras, credentialing, anti-tailgating measures, key control and audit logs need an owner and retention rule. Emergency access must remain possible during a network or power event. Security controls that prevent safe maintenance or emergency response are not resilient controls.
Controls and alarms
An alarm is useful only when somebody can act on it
BMS, EPMS, DCIM and device monitoring should turn equipment states into a clear operating response without creating a second, poorly secured control plane.
Define which platform is authoritative for electrical, mechanical, environmental and capacity data. Tags need consistent equipment names, units, limits and time. Time synchronization matters when an operator reconstructs a transfer that lasted seconds. Keep alarm rationalization records so a real fault is not buried under repeated nuisance events or a flood of downstream symptoms.
Every important alarm needs severity, delay, destination, acknowledgement expectation, escalation route and a written response. Test the complete path from sensor or device through the network and notification service to the person on call. Include loss of the monitoring platform, loss of a gateway and stale data. A green dashboard can be wrong if communications stopped and the last state was retained.
Remote monitoring interfaces need inventory, supported firmware, protected credentials, segmented access, logging and a retirement plan. Disable unused services and avoid exposing equipment management directly to the internet. Changes to firmware, cards, polling or thresholds belong in change control because they can affect both cyber risk and the control sequence observed by operators.
Commissioning
Commissioning proves sequences, not just equipment
Factory tests show what a product can do. Site and integrated systems tests show whether the installed facility can deliver the owner's operating requirements.
The commissioning plan should begin with written owner project requirements and a basis of design. Submittal review, factory witness testing, installation checks, pre-functional verification and functional tests then build evidence in stages. Deficiencies need an owner, target date and retest record. Accepted exceptions should be explicit rather than disappearing into a final report.
Integrated systems testing exercises dependencies and transitions: loss of a utility source, generator start, UPS ride-through, transfer, cooling response, alarm delivery, partial failure and restoration. Tests need safe prerequisites, defined loads, expected observations, hold points and abort criteria. A scripted success path is not enough if recovery from a failed start or stuck device has never been rehearsed.
Commissioning continues through turnover. Operators need current drawings, sequences, settings, training, spares, warranties and baseline readings. Seasonal and deferred tests should remain on a controlled schedule. When a facility adds a new load block or changes a major sequence, recommission the affected chain rather than assuming the original result still applies.
Operations
Reliability is maintained one controlled change at a time
Operating discipline keeps a good design from being weakened by undocumented settings, deferred defects, temporary workarounds and untested assumptions.
The operating model should name roles for facilities, IT, security, vendors and management. Critical procedures include standard operations, methods of procedure for planned work and emergency operations for abnormal events. Each needs scope, prerequisites, responsibilities, communication, hold points, rollback and closeout. Copying an old procedure without verifying the current one-line and site condition creates hidden risk.
Maintenance plans should combine manufacturer requirements, code obligations, observed condition, operating criticality and available redundancy. Schedule the whole system, not isolated assets. UPS service may depend on bypass health, cooling service on electrical redundancy, and generator work on fuel and controls. Record defects in one controlled backlog with risk, interim control, owner and due date.
Measure what supports decisions. Power usage effectiveness can help track facility energy, but boundary and measurement period must stay consistent. Also watch capacity reserve, alarm burden, maintenance completion, battery exceptions, failed starts, water use where relevant and repeat incidents. After a real event or test, update procedures, training and design assumptions. The value is in closing the loop.