Backup and Disaster Recovery
Almost every organisation has backups. Far fewer have ever restored one. The gap between a backup job reporting success and a business actually recovering is where most disaster recovery plans fail, and it is only discovered at the worst possible moment. We design against recovery objectives the business has agreed, then prove them by restoring.
1
Assess
What downtime actually costs
2
Design
RPO and RTO per system
3
Implement
Backup and replication
4
Test
Restore, do not assume
5
Report
Runbooks and evidence
How We Approach It
01
Business Impact and Recovery Objectives
Recovery design starts with the business, not the technology. We establish what each system costs the organisation per hour of downtime and how much data loss is tolerable, which produces the recovery objectives everything else is built against.
- Business impact assessment per system
- Recovery time objective agreed per workload
- Recovery point objective agreed per workload
- Prioritised recovery sequence for a full outage
02
Backup Design and Implementation
Design follows the objectives. Systems with a one hour recovery point need a different approach from those where a nightly job is sufficient, and paying for the former across the whole estate is waste.
- Backup schedule aligned to agreed objectives
- Application aware backup for databases and mail
- Retention tiers by system criticality
- Capacity planning and growth allowance
03
Offsite and Immutable Copies
A backup on the same infrastructure as the system it protects is not a backup. Ransomware specifically targets backup repositories, which is why at least one copy has to be immutable and out of reach of a compromised domain.
- Offsite replication to a separate location
- Immutable copies that cannot be altered or deleted
- Separation from production authentication
- Encryption in transit and at rest
04
Restore Testing
This is the part that gets skipped. We restore on a schedule and record the result, because a job reporting success proves the job ran, not that the data is usable.
- Scheduled restore testing with recorded evidence
- Full system and granular file level recovery
- Verification that restored data is usable
- Time to recover measured against the objective
05
Runbooks and Reporting
Recovery should not depend on who is available. We document the runbook step by step, keep it current, and report backup health monthly so problems surface before they matter.
- Documented recovery runbooks per system
- Roles, escalation and contact details
- Monthly backup health and failure reporting
- Annual disaster recovery exercise
What You Receive
- Business impact assessment with agreed RPO and RTO per system
- Backup design document and implemented configuration
- Offsite and immutable copy arrangement
- Scheduled restore testing with recorded evidence
- Documented recovery runbooks
- Monthly backup health reporting
Indicative Timeline
Design and implementation normally runs three to six weeks depending on data volumes and how much has to be seeded offsite. The initial full copy to an offsite location is frequently the longest single step.
- Business impact assessment: one week
- Design and approval: one week
- Implementation and initial seeding: two to four weeks
- First restore test and runbook handover: one week
Technology We Deploy
We hold partner accreditation with the vendors below and deploy their products in client environments.
Veeam
Backup, replication and recovery orchestration, including immutable repositories and automated restore verification.
HPE
Server and storage infrastructure underpinning on premise backup targets and hybrid recovery.
Immutable Storage
Write once repositories that a compromised administrator account cannot alter or delete.
Cloud Replication
Offsite copies held in a separate cloud region so a site level event does not take the recovery copy with it.
Microsoft 365
Backup of mail, files and collaboration data, which the platform retention policy alone does not fully cover.
Monitoring
Alerting on failed and missed jobs, because a backup that silently stopped is worse than none at all.
Frequently Asked Questions
We already have backups. Why would we need this?
Because having backups and being able to recover are different things. The common findings are backups sitting on the same infrastructure as production, retention that does not match any agreed objective, and a restore that has never once been tested.
What are RPO and RTO?
Recovery point objective is how much data you can afford to lose, measured in time. Recovery time objective is how long you can afford to be down. They are business decisions, not technical ones, and every design choice follows from them.
Does this protect against ransomware?
It is a major part of it. Modern ransomware deliberately targets backup repositories first, which is why immutable copies separated from production authentication matter more than backup frequency.
Do we need to back up Microsoft 365?
Generally yes. The platform protects against its own infrastructure failure, not against a user deleting a mailbox, a malicious insider, or a retention policy that expires data you later need.
How often do you test restores?
On an agreed schedule, typically quarterly for critical systems with a full exercise annually. Every test is recorded with the time taken, so you can demonstrate the objective is actually met rather than merely documented.
Can you work with our existing backup product?
Usually. Where the existing product meets the objectives we configure and manage it rather than replace it. Replacement is proposed only where the current tool genuinely cannot meet an agreed recovery objective.
Related Services
This sits inside our Managed ICT Services practice. Related work: ICT Audit where recovery capability has to be independently tested, and Cyber Security Assessments for the ransomware exposure backups exist to survive.
