UK LTO tape specialists since 1994

Picture it: your business is running smoothly, and then disaster strikes. Systems crash, data is lost and staff can’t work. Whether the cause is a cyber attack, a flood, a fire, a power failure or a simple human mistake, the question is the same: how quickly can you get back to business? Disaster recovery is how you answer it.

Key points

  • Disaster recovery (DR) is the plan and technology for restoring IT systems and data after a serious disruption.
  • Set an RTO (maximum downtime) and RPO (maximum data loss) for each system.
  • DR is different from high availability. You need both, for different reasons.
  • An offline copy, such as LTO tape, protects your last line of recovery from ransomware.
  • A plan that has never been tested is not a plan.

What is disaster recovery?

Disaster recovery is the process of planning and putting in place the strategies, people and technology that let an organisation restore its critical systems and data after an unexpected event. It is part of wider business continuity planning, which also covers people, premises and suppliers. DR focuses on IT.

Why DR planning matters

Almost every organisation now depends on IT to operate. Without a DR plan, an outage can mean:

  • Extended downtime and lost productivity
  • Lost revenue, missed orders and contractual penalties
  • Permanent loss of data that cannot be recreated
  • Regulatory consequences, including under UK GDPR
  • Lasting damage to your reputation and customer trust

A good DR plan keeps downtime short and predictable, and gives everyone a clear process to follow under pressure.

The key goals of disaster recovery

Recovery Time Objective (RTO)

The maximum acceptable time a system can be unavailable. For example, “email must be back within 4 hours”.

Recovery Point Objective (RPO)

The maximum acceptable data loss, measured in time. For example, “we can lose no more than 15 minutes of orders”.

RTO and RPO should be agreed with the business for each system. They determine the technology you need: tight targets mean replication and standby systems, while relaxed ones may be met by restoring from backup.

Types of outage

Unplanned outages

Sudden, unexpected disruptions: hardware failure, software bugs, power cuts, network failures, cyber attacks, fire, flood and human error. These are what DR plans are mainly designed for.

Planned outages

Deliberate downtime for maintenance, upgrades or data centre moves. Good DR tools, such as a controlled switchover to a secondary site, can also keep services running during planned work.

Common pain points in disaster recovery

  • Untested plans that fail when they are finally needed
  • Backups compromised by the same ransomware that hit production
  • Unrealistic RTOs: restoring hundreds of terabytes over a WAN or from cloud takes far longer than expected
  • Missing dependencies: restoring an application without its database, identity service or DNS
  • Knowledge held by one person, who may not be available during an incident
  • Cost: duplicating everything at a second site is expensive, so priorities are needed

Disaster recovery vs high availability

High availability (HA)Disaster recovery (DR)
PurposePrevent downtime from component failuresRecover from major incidents that take out a system or site
HowRedundant hardware, clustering, automatic failover within a siteBackups, replication to another site, documented recovery procedures
Protects againstFailed disks, servers, power suppliesSite loss, ransomware, corruption, major human error
Typical recoverySeconds to minutesMinutes to days, depending on RTO

In summary: HA keeps things running through everyday failures, and DR brings them back after a disaster. HA replicates corruption and ransomware instantly, so it is never a substitute for point-in-time backups.

Disaster recovery strategies

  • Backup and restore: the simplest and cheapest. Suitable for systems with RTOs of hours to days. Tape and disk backups are restored to replacement hardware.
  • Cold standby: infrastructure is available at a secondary site but not running. Restore from backup when needed.
  • Warm standby: systems at the secondary site are running with regularly replicated data, ready to take over within hours.
  • Hot standby / active-active: fully replicated systems that can take over in minutes. This is the most expensive option, used for the most critical services.
  • DR to cloud (DRaaS): replicate or back up into a cloud platform and start systems there if your primary site is lost.

DR operations: switchover vs failover

Switchover

A planned, controlled role change. The primary is shut down cleanly and the secondary takes over with no data loss. It is used for maintenance, and it is a great way to test DR.

Failover

An unplanned move to the secondary site after the primary has failed. Some data loss, up to the RPO, may occur, and failing back afterwards has to be planned.

The role of tape in disaster recovery

Replication and snapshots give fast recovery, but they are online, so they can be corrupted or encrypted along with production. LTO tape gives you an offline, air-gapped copy that survives when everything else is compromised. It is also the most cost-effective way to keep long retention for recovering from slow-burn corruption or attacks discovered weeks later. A local tape library can restore large data sets at full streaming speed. Read more about tape and ransomware.

Building your DR plan

  1. List your systems and their dependencies.
  2. Agree RTO and RPO for each with the business.
  3. Choose strategies and technology to meet them within budget.
  4. Document step-by-step recovery procedures and contact lists.
  5. Keep at least one offline or immutable copy of critical data.
  6. Test regularly, then learn and update.

Our data backup guide covers the backup side in more detail. For help designing or reviewing your DR storage, contact our team.

Free, no-obligation advice

Planning a tape purchase or upgrade?

Tell us your data volumes, retention needs and existing hardware. We will recommend the right generation, media and hardware — and quote it the same working day.