The current HIPAA Security Rule requires a disaster recovery plan but sets no clock on it. The proposed update sets two: exact retrievable copies of electronic protected health information (ePHI) no more than 48 hours old, and written procedures to restore critical electronic information systems and data within 72 hours of a loss, with an annual test and a criticality analysis to decide what "critical" means. Those are a recovery point objective and a recovery time objective written into regulation. This post is for the infrastructure lead who has to turn them into an architecture: what to replicate, what to back up, where to put it, how to order the restoration, and how to run the rehearsal so that the report in the evidence file has real timestamps on it.

Start with the criticality analysis, not the storage

The proposed rule makes the applications and data criticality analysis mandatory, and it is the right first step for engineering reasons as well as legal ones. A health system runs dozens of systems that touch ePHI; only some of them need to be back within 72 hours, and the ones that do need to come back in a specific order because they depend on each other.

A workable analysis answers three questions per system: what happens to patient care and to the organisation if it is unavailable for one hour, one day and three days; which other systems must be running before it can be restored; and how much data loss is tolerable. The output is a ranked list with a recovery time objective (RTO) and a recovery point objective (RPO) per system. For a mid-sized system the top of the list usually reads:

TierSystemsRPORTODepends on
0Identity provider, DNS, network, backup catalogueMinutes2 to 4 hoursNothing; these are restored first
1EHR and its database, imaging (PACS), pharmacy, lab interfacesMinutes to 1 hour8 to 24 hoursTier 0
2Email, shared files, document collaboration, internal messaging, telehealthUnder 24 hours24 to 48 hoursTier 0
3Analytics, reporting, training, intranet48 hoursBeyond 72 hoursTiers 0 to 2

Two observations. The 72-hour clause applies to tiers 0 to 2, and those tiers have to be restored sequentially, so the individual objectives have to add up to under 72 hours with margin. And the 48-hour backup clause applies to every tier that holds ePHI, including tier 3, because the rule limits the age of the copy, not the speed of its restoration.

Separate the common failures from the rare ones

Most outages are not disasters. A disk fails, a host fails, a network link flaps, a virtual machine locks up. An architecture that handles those with no restoration at all does two things for the compliance programme: it keeps the restoration runbook for the events that actually need it, and it makes the measured RTO for ordinary incidents close to zero, which the risk analysis can cite.

MassiveGRID's high-availability platform is built for the common case. Workloads run on Proxmox clusters with Ceph storage, where every block is written to multiple physical nodes before the write is acknowledged, and a failed host triggers an automatic restart of its virtual machines on a healthy one within seconds. Nothing is restored from backup; the data was never on only one machine. This is the layer that makes a 100% uptime SLA possible, and it is also the layer that satisfies the proposed rule's requirement for exact copies, because replication keeps the copies current to the second.

Replication is not a backup, though, and the rule knows the difference. A corrupted database is replicated faithfully. A ransomware payload that encrypts a file share is replicated faithfully. An administrator who drops the wrong table has the drop replicated. The 72-hour clause is about those events, and they need point-in-time copies that live somewhere the failure cannot reach.

The three recovery tiers

The architecture that serves both clauses has three layers, each with a different RPO and a different restore mechanism. Every ePHI-holding system sits in all three.

Layer 1: synchronous replication for availability

Ceph replication across nodes in the cluster, with live migration for maintenance. RPO zero, recovery automatic, no runbook. Covers hardware failure and planned work. This is where the HA guarantee comes from, and our post on why collaboration platforms need high availability explains the failure modes it removes.

Layer 2: point-in-time copies on separate storage

Snapshots and dumps that can be rolled back to a chosen moment, held on a storage system that is logically and administratively separate from production, with credentials that production hosts do not hold. For databases this means continuous archiving of the write-ahead log, which gives an RPO of minutes and the ability to restore to the second before the bad transaction. For file stores it means daily snapshots at minimum, hourly where the change rate justifies it. The 48-hour clause is satisfied with margin by anything daily; the reason to go further is that a 24-hour loss of referral paperwork or clinical notes is still a reportable event internally, whatever the rule allows.

Layer 3: off-site, immutable copies for the disaster case

A copy of layer 2 replicated to a second datacenter in a different region, written to storage with object lock or an equivalent immutability control so that a compromised administrator credential cannot delete it. This is the copy that answers "what if the site is gone" and "what if the attacker got everything", and it is the one the restoration rehearsal should start from, because if the rehearsal works from the worst copy it works from any of them. MassiveGRID's backup services provide the AES-256 encrypted, off-site replicated store for this layer, and the disaster recovery service adds a standby environment in a second datacenter for customers whose tier 1 RTO is measured in hours rather than days.

Sizing the restoration against the clock

The 72 hours are consumed by three things: detection and decision, data transfer, and application recovery. The architecture controls the second and the runbook controls the third; the first is an incident response matter and the proposed rule's separate requirement for a written, tested incident response plan covers it.

Transfer time is the one people underestimate. A 40 TB document store restored over a 1 Gbps link takes about four days at full rate, which fails the clause before anyone has typed a command. The fixes are structural: keep layer 2 in the same datacenter as production so that a point-in-time restore moves data across a local network rather than a WAN; keep the layer 3 copy in a datacenter with a standby environment next to it, so that a disaster restore is also local; and restore in criticality order so that tier 1 is serving clinicians while tier 2 is still copying. MassiveGRID operates datacenters in New York, Ashburn, Chicago, Dallas, Miami, Phoenix, Los Angeles and Seattle, so a US health system can keep both copies in-country and still in different regions.

Application recovery has to be scripted. The runbook for each tier 0 to 2 system should be an ordered list of commands and checks, not a description, with the restore of configuration before data, the identity provider reconnected before any user-facing test, and a defined acceptance test (a clinician can log in and open a chart; a user can send and receive mail) that marks the system as restored. Each step gets a target duration, and the sum of the targets across the tiers is the planned RTO. If the sum is over 60 hours, there is no margin and the design needs work.

The rehearsal is the deliverable

Under the current rule, testing and revision of the contingency plan is addressable; under the proposal it is required at least every 12 months. The compliance value of the test is the evidence it produces, so run it in a way that produces evidence.

  1. Restore from the layer 3 copy into an isolated environment, not into production. The isolated environment should be a real deployment, sized like production, on the standby platform. Cloning the production environment is not a test of restoration; starting from a backup is.
  2. Follow the runbook as written, by someone other than its author if possible, and log every deviation. Deviations are the output; the runbook gets revised afterwards, which is the "revision" half of the requirement.
  3. Time every step and record the start and end of each tier. The report states the planned and the measured RTO per system.
  4. Run the acceptance tests with the business present. The compliance officer, or a clinician delegate, confirms that the restored system holds the expected data and that the newest record is within the RPO. That confirmation is the line in the report an investigator reads first.
  5. Verify the backup chain itself at the same time: the age of the newest copy of every ePHI store, the integrity check results, and the immutability settings on the off-site copy. This closes the 48-hour clause with a dated artefact.

MassiveGRID runs this rehearsal with healthcare customers as part of the managed service, and the report, with timings, deviations and sign-off, is one of the documents that back the annual written verification the proposed rule requires from a business associate. The contingency plan checklist lists where it sits among the other evidence.

Where SaaS fits, and where it does not

Nothing above requires the covered entity to run its own datacenter; MassiveGRID runs the platform, the backups, the replication and the rehearsal. What it does require is that the systems be restorable by someone who is contractually bound to the entity and can show their work. That is the line between a managed single-tenant deployment and a hyperscale SaaS tenant. For the EHR and for the file, mail and collaboration stack that most criticality analyses put in tiers 1 and 2, the restoration has to be something the entity can rehearse. Our post on Microsoft 365 under the proposed rule examines what happens to the 72-hour objective when it cannot, and the Nextcloud Hub Enterprise replacement guide shows the tier 2 stack deployed in the architecture described here.

Frequently Asked Questions

Does the current HIPAA Security Rule require a 72-hour restoration?

No. The current rule at 45 CFR 164.308(a)(7) requires a disaster recovery plan, a data backup plan and an emergency mode operation plan, with testing and criticality analysis as addressable specifications, and it sets no time limit. The 72-hour restoration objective and the 48-hour backup currency limit are in the proposed rule published in January 2025, which has not been finalised as of October 2026.

Is storage replication enough to satisfy the backup requirement?

No. Replication keeps copies current but replicates corruption, deletion and ransomware encryption faithfully. The rule requires retrievable exact copies, which in practice means point-in-time copies on separate storage, with an off-site immutable copy for the disaster case. Replication is what makes ordinary hardware failures a non-event; backups are what the 72-hour clause is about.

What RPO and RTO should a health system set?

The proposed rule sets the ceiling: copies no more than 48 hours old and critical systems restored within 72 hours. The criticality analysis sets the real targets, which are usually much tighter for the EHR and identity systems (minutes of RPO, hours of RTO) and closer to the ceiling for reporting and training systems. MassiveGRID agrees RPO and RTO per deployment and writes them into the disaster recovery plan.

How often should the restoration be rehearsed?

At least every 12 months under the proposed rule, and after any material change to the systems or the backup design. MassiveGRID conducts the rehearsal with healthcare customers as part of the managed service and provides the timed report for the compliance file.

Does MassiveGRID sign a Business Associate Agreement for disaster recovery services?

Yes. MassiveGRID signs a Business Associate Agreement covering the hosting, backup and disaster recovery services it provides for ePHI, including the 24-hour contingency plan activation notice the proposed rule would require from business associates.

A disaster recovery plan you can rehearse

MassiveGRID designs and operates DR for ePHI workloads: Proxmox HA with Ceph replication for the common failures, point-in-time backups on separate storage, immutable off-site copies in a second US datacenter, and an annual restoration rehearsal with a timed report. Under a Business Associate Agreement, with a 100% uptime SLA.

Disaster recovery services

Further Reading