Canadian data centres, planning guides, power, cloud regions and internet exchanges.

Planning guide 04

Verify data-centre connectivity and recovery paths

Connectivity is a set of contracted services and physical and logical paths, not a collection of logos. A resilient design names the applications, traffic flows, service handoffs, route owners, security boundaries, failure cases and recovery sequence. Public exchange and facility lists can identify questions. They do not prove a cross-connect, a diverse entrance, a cloud on-ramp or an available circuit for a particular customer.

01

Start with applications, traffic and recovery

List the applications, users, sites, cloud services, partners and administrative systems that exchange traffic with the data centre. Record bandwidth, latency, jitter, packet-loss tolerance, encryption, addressing, name resolution, time synchronization and expected growth. Separate normal traffic from backup replication, software distribution, monitoring and recovery. A link sized only from today's internet use can fail when replication or restoration becomes the dominant flow.

Sources: National Institute of Standards and Technology

For every important service, state the recovery objective and the dependency that must be available first. Authentication, DNS, network control, firewalls, cloud consoles and out-of-band access can become prerequisites for restoring the application. Test whether administrators can reach the alternate path when the primary identity or management service is unavailable. Keep manual recovery steps current and accessible under the organization's security controls.

Sources: National Institute of Standards and Technology, Canadian Centre for Cyber Security

02

Prove each service from provider to demarcation

For each carrier or cloud service, record the legal provider, reseller if any, ordered product, bandwidth, handoff, address, demarcation point, cross-connect, installation interval, service level and support route. Confirm whether the service is installed, orderable, planned or simply marketed in the metro. Provider and facility names change, so keep the contract and current service record beside the planning diagram.

Sources: Equinix, Cologix

An internet exchange creates a place where participating networks can exchange traffic. It does not automatically provide transit, private cloud access, customer connectivity or a path into every nearby facility. CIRA lists community exchanges across Canada and links to their operators. Use that list to identify the exchange and then confirm membership, location, port, transport and route policy directly with the parties involved.

Sources: Canadian Internet Registration Authority

03

Trace physical diversity until the paths separate

Ask where entrances, conduits, splice points, meet-me rooms, risers, provider equipment, poles, ducts and upstream routes are shared. Two circuits can use different provider names and still converge on one local route or common carrier. Two entrances can join outside the property. Obtain route evidence through the appropriate controlled process and record the point at which the paths become independent.

Sources: Equinix, Cologix

Define the failure the diversity is meant to survive. Building work, a local cut, provider maintenance, regional backbone loss and control-plane failure require different answers. Include cross-connect panels, power to provider equipment and access to the meet-me room. A physically diverse path may still fail with the primary service if both depend on one routing policy, firewall cluster, DNS service or administrative account.

Sources: National Institute of Standards and Technology, Canadian Centre for Cyber Security

04

Design network zones and management access explicitly

Group services by security policy and control traffic at defined boundaries. The Canadian Centre for Cyber Security describes security zones as logical groupings under common policy constraints, separated by controlled interfaces. Map public, user, application, data, management, monitoring, backup and vendor-access flows in terms that match the organization's security design. Do not rely on rack location alone to create separation.

Sources: Canadian Centre for Cyber Security

Protect the management path from the failures and access problems it is expected to diagnose. Record how authorized staff reach console servers, power controls, network devices and monitoring when the production network or identity service is down. Limit access, log activity, maintain unique accounts and remove supplier access when it is no longer required. Keep credentials and sensitive topology out of public planning tools.

Sources: Canadian Centre for Cyber Security, National Institute of Standards and Technology

05

Keep cloud access and data movement measurable

For cloud-connected workloads, identify the cloud region, service type, on-ramp or internet path, provider, routing ownership, encryption, bandwidth, quotas and transfer costs. A cloud region in the same country does not prove low latency or a direct route from the chosen facility. Measure from the intended network and include the result, time, path and test method in the decision file.

Sources: Amazon Web Services, Google Cloud, Oracle Cloud

Estimate backup, replication, migration and recovery traffic with real data volumes and change rates. Test how long a restore or bulk transfer takes under the contracted service and competing traffic. If a recovery plan depends on moving data to another region or site, include the time to make the compute, storage, keys, network policy and application dependencies ready. Network capacity alone does not complete recovery.

Sources: National Institute of Standards and Technology

06

Test failover from the user point of view

Test circuit failure, device failure, provider loss, route withdrawal, firewall state, DNS change, cloud-path loss and loss of the management network. Observe the application from representative user locations. Record detection time, route convergence, session behaviour, name resolution, security controls and the manual action required. A green interface does not prove that users reached the correct application and data.

Sources: National Institute of Standards and Technology, Canadian Centre for Cyber Security

Include the return to normal service. Failback can create a second interruption, asymmetric routing, stale sessions or unexpected replication. Define who approves the change and what must be checked before traffic moves. Keep test data, expected results, actual results, defects and retests. A successful annual exercise does not remove the need to test after material routing, provider, firewall, identity or application changes.

Sources: National Institute of Standards and Technology

07

Make contracts and operations match the diagram

Confirm order ownership, support contacts, authorization lists, maintenance notice, escalation, monitoring responsibility, cross-connect charges, service levels, renewal, cancellation and removal. Ensure that a provider change or circuit move triggers diagram and recovery-plan updates. If one supplier delivers both paths, state how the contract preserves the required separation and how the customer will be told when a route changes.

Sources: National Institute of Standards and Technology

Review the service map with network, security, application, facilities and procurement owners. Mark installed service, contracted change, assumption and unresolved claim differently. Keep public directory information as a lead, not as proof. The final record should let an operator identify what failed, who owns it, which alternate path is expected to work and how to verify that the business service has recovered.

Sources: Canadian Internet Registration Authority, Canadian Centre for Cyber Security

Review file

  • Applications, users, traffic, performance limits and recovery objectives
  • Provider, product, handoff, demarcation, cross-connect and support record
  • Physical route, shared points, power and meet-me-room dependencies
  • Security zones, controlled interfaces and out-of-band management
  • Cloud region, path, bandwidth, quotas, transfer and measured latency
  • Failure, failover, failback and user-level acceptance tests
  • Contract owners, escalation, maintenance notice and change records

Next action

A resilient path is one the team can prove and operate

Use the internet-exchange and cloud-region directories to identify relevant infrastructure, then confirm the actual service and route with the facility, network and cloud providers. Keep sensitive route and access information in the organization's controlled records.