INSIGHTS
Technology Fundamentals

ISACA CISA: Business Resilience from an Auditor’s Perspective

In this article
  1. Distinguish resilience, continuity, and disaster recovery
  2. Use business impact analysis to define recovery priorities
  3. Validate RTO and RPO against architecture and evidence
  4. Audit backup as a recovery capability, not a storage task
  5. Review continuity plans for usable decision authority
  6. Test disaster recovery under conditions that can actually fail
  7. Include providers and concentration risk in resilience audits
  8. Evaluate cyber resilience and the integrity of recovery
  9. Report resilience in terms of proven capability

Business resilience is the ability to continue delivering critical outcomes when technology, facilities, suppliers, people, or data are disrupted. From an auditor’s perspective, resilience is not proven by the existence of a continuity plan or a successful annual tabletop alone. The auditor needs evidence that critical processes have been identified, recovery objectives reflect business need, dependencies are understood, backup and failover controls work, people know their roles, providers can support recovery, and lessons from testing are closed.

The current CISA outline devotes major attention to information-systems operations and business resilience, including business impact analysis, system and operational resilience, data backup and restoration, business continuity, and disaster recovery. The ISACA certifications also connects resilience to governance, risk, security management, and vendor oversight. A strong audit therefore evaluates whether the organization can recover the business service, not merely whether IT can restart individual servers.

Distinguish resilience, continuity, and disaster recovery

Resilience is the broader outcome. Business continuity focuses on keeping or restoring critical business activities, while disaster recovery focuses more specifically on recovering technology and data after disruption. The three overlap, but the distinction matters because a technically successful recovery may still fail the business if people, facilities, suppliers, or manual procedures are unavailable.

An audit should determine whether governance reflects this breadth. Continuity plans owned only by infrastructure teams may omit customer communications, regulatory obligations, alternate work arrangements, vendor dependencies, or business decision authority. Conversely, business plans that assume “IT will restore everything” may ignore realistic recovery sequencing and capacity limits.

The concepts in business continuity and disaster recovery planning are most useful when mapped to actual services and dependencies. Generic templates provide structure, but evidence of resilience comes from specific recovery capabilities.

Auditors should also examine whether resilience includes cyber scenarios. Ransomware, identity compromise, destructive administrator activity, and supply-chain incidents can make normal failover unsafe if the backup or secondary environment shares the same compromised trust.

Scope should include the minimum business capability needed during disruption, not only the full steady-state service. Manual workarounds, reduced transaction limits, alternate communication channels, and temporary staffing arrangements may be acceptable for a defined period. Auditors should verify that these degraded modes are documented and tested.

Use business impact analysis to define recovery priorities

A business impact analysis should identify critical processes, required resources, tolerable outage periods, upstream and downstream dependencies, regulatory or contractual obligations, and the consequences of disruption. The auditor should verify that the BIA is current enough to reflect acquisitions, new cloud services, changed customer commitments, and retired applications.

Criticality should be based on business impact rather than technical prestige. A small authentication component or batch interface may be more important than a large application if many services depend on it. Auditors should look for dependencies that cause lower-tier systems to become single points of failure for higher-tier processes.

BIA assumptions should be supported by process owners. If a system is assigned a four-hour recovery target, the business should be able to explain what happens after four hours and why that period is tolerable. Arbitrary targets copied from previous plans make recovery design difficult to justify.

The risk perspective of CRISC helps connect recovery investment to business consequence. Not every system needs zero downtime, and forcing unrealistic objectives can waste resources while genuinely critical dependencies remain underprotected.

Business impact assumptions need evidence and periodic challenge. Revenue dependency, regulatory deadlines, customer tolerance, safety implications, and upstream obligations can change as products and operating models evolve. Stale BIA rankings can direct recovery investment toward yesterday’s priorities.

Validate RTO and RPO against architecture and evidence

Recovery Time Objective describes how quickly a service should be restored, while Recovery Point Objective describes the acceptable amount of data loss measured in time. Auditors should verify that architecture and operations can meet both. A documented one-hour RPO is meaningless if backups run once per day or replication is not monitored.

Recovery objectives should be assessed end to end. Restoring a database within two hours may not meet the application RTO if identity, DNS, certificates, network routes, queues, integration partners, or application configuration take another day. The critical path determines the real recovery time.

Availability metrics can provide supporting evidence, but high availability is not the same as recoverability. The discussion around five-nines availability illustrates how uptime targets shape design, yet an always-on service can still be vulnerable to corruption, ransomware, or administrative error.

Auditors should compare stated objectives with test results and actual incidents. Repeatedly missing RTO in exercises without revising the plan or architecture is a control failure even if each individual test is documented thoroughly.

RTO and RPO should be reconciled with contractual and technical dependencies. A service cannot meet a one-hour recovery objective if a critical supplier restores in four hours or if required data is backed up only once per day. Dependency analysis makes conflicting objectives visible before a crisis.

Audit backup as a recovery capability, not a storage task

Backup controls should address scope, frequency, retention, encryption, immutability, access, monitoring, and restoration. The most important evidence is not that backup jobs show “success,” but that required data can be restored to a usable and trusted state within the recovery objective.

Coverage should include more than primary databases. Configuration, infrastructure templates, encryption keys, certificates, application secrets, SaaS exports, identity configuration, and critical file stores may all be necessary to rebuild a service. An application backup without the identity or key material needed to access it may be unusable.

The variety of cloud backup approaches means auditors should understand which copies are logically or administratively separated. Backups controlled by the same compromised administrator account as production may be vulnerable to deletion during an attack.

Restore testing should use representative data and document duration, errors, manual steps, and business validation. A backup that can technically be mounted but produces inconsistent application state does not provide the intended assurance.

Restore testing should validate both data and usability. A backup may complete successfully yet contain corrupted data, missing keys, incompatible versions, or incomplete configuration. Recovery evidence should therefore include application startup, integrity checks, representative transactions, and business-owner acceptance where appropriate.

Review continuity plans for usable decision authority

A continuity plan should tell people what to do when normal assumptions fail. Roles, escalation paths, declaration criteria, communication channels, alternate procedures, decision authority, contact information, and dependencies should be clear enough to use under pressure. Auditors should challenge plans that contain extensive policy text but little operational guidance.

Plans should account for unavailable personnel. If only one administrator knows how to restore a critical service, resilience depends on that person’s availability. Cross-training, documented procedures, and delegated authority can reduce key-person risk.

Communication is a control. Internal teams, customers, regulators, suppliers, executives, and the media may require different information at different stages. The plan should identify who approves external statements and how communication continues if normal email, collaboration, or telephony is unavailable.

Manual workarounds should be realistic. A plan that says “process transactions manually” should identify the volume that can be handled, how records will be reconciled later, and what controls prevent fraud or duplication during the workaround period.

Decision authority should be explicit for ambiguous scenarios. Teams need to know who can declare a disaster, activate alternate sites, invoke emergency spending, notify regulators, or accept temporary risk. Unclear authority causes delay precisely when normal governance channels may be unavailable.

Test disaster recovery under conditions that can actually fail

Testing should validate more than the happy path. A useful exercise can include unavailable administrators, corrupted backups, failed DNS, loss of a cloud region, compromised identity, or a vendor that does not respond on schedule. These conditions reveal assumptions that a scripted demonstration may never challenge.

Disaster recovery testing can range from walkthroughs and simulations to parallel processing and full failover. The appropriate method depends on risk, but critical services should eventually prove that the actual recovery mechanism works, not just that people can discuss it.

Auditors should inspect the gap between test scope and production reality. If the test restores a small data set, uses a prebuilt environment, or excludes integrations, the result may not support the claimed enterprise RTO. Limitations should be documented and addressed through additional evidence.

Post-test findings require ownership, due dates, and retesting. Repeating the same recovery failure across annual exercises indicates weak governance even when every test is technically completed.

Exercises should generate measurable lessons. Observed recovery time, failed steps, communication delays, missing credentials, dependency surprises, and manual workarounds should become tracked corrective actions with owners and deadlines. Repeating the same finding across exercises is evidence that testing is not driving improvement.

Include providers and concentration risk in resilience audits

Critical business services often depend on cloud providers, SaaS platforms, telecommunications, payment processors, identity services, DNS, and managed security vendors. Auditors should determine whether continuity plans include those providers and whether contracts support required recovery, support, notification, and data-access needs.

Multiple vendors do not necessarily eliminate concentration risk. Several SaaS providers may run in the same cloud region or depend on the same identity platform. The audit should look for common dependencies that can fail simultaneously even when supplier names differ.

Service-level agreements should be compared with business objectives. A vendor’s contractual recovery target may be slower than the customer’s required RTO. In that case, the organization needs an alternate design, compensating process, or explicit risk acceptance rather than assuming the contract solves the gap.

The security-management lens of CISM is useful because provider resilience is part of the organization’s own program. Management cannot delegate accountability for continuity simply because a third party operates the technology.

Supplier resilience should be tested through evidence rather than questionnaire promises. Organizations can examine provider recovery reports, regional architecture, dependency maps, outage history, continuity exercises, and contractual recovery commitments. Critical suppliers should also be included in scenario planning when coordination is necessary.

Evaluate cyber resilience and the integrity of recovery

Cyber incidents create a different recovery problem from equipment failure. A replica can copy malicious changes perfectly, and a backup can contain persistent malware or corrupted configuration. Auditors should assess whether recovery plans include a method for identifying a trusted restore point and validating the integrity of systems before reconnecting them.

Identity recovery is particularly important. If privileged accounts or federation systems are compromised, the organization may need isolated emergency identities, protected credentials, and a way to rebuild administrative trust. Restoring servers while leaving attacker-controlled identities active can cause immediate reinfection.

Immutable backups, offline copies, protected logging, golden images, infrastructure templates, and known-good configuration can support cyber recovery, but each needs access controls and testing. An immutable backup policy that administrators can disable using the same compromised credentials may provide less protection than expected.

Audit should also consider evidence preservation. Rapid rebuilding should not destroy the only artifacts needed to determine cause, scope, or legal impact. Recovery and forensics should be coordinated so the organization can restore service while retaining evidence.

Cyber recovery should preserve a trusted path back to operation. Clean-room procedures, known-good images, privileged identity recovery, immutable backups, malware scanning, and staged reconnection can reduce the chance of restoring compromised state. The sequence of recovery is a security control, not only an operations plan.

Report resilience in terms of proven capability

Resilience metrics should distinguish documented intent from demonstrated performance. Useful measures include percentage of critical services with tested recovery, achieved RTO and RPO, restore success, unresolved exercise findings, age of BIA data, untested provider dependencies, and time to mobilize the continuity team.

Audit findings should connect the weakness to business consequence. “DR plan not updated” is less useful than explaining that the recovery plan references a retired data center and omits the cloud identity platform now required to access all production systems. Specific impact helps management prioritize remediation.

Evidence from real incidents can be especially valuable. If a recent outage took eight hours to recover a service with a two-hour RTO, management should understand why and whether the target, architecture, or process needs to change. A paper plan should not override operational reality.

Business resilience is mature when critical outcomes can continue or recover under realistic disruption, and when management can prove that capability through current analysis, tested controls, accountable ownership, and closed lessons. The auditor’s role is to challenge assumptions until the difference between “we have a plan” and “we can recover” becomes visible.

Board and executive reporting should distinguish capability from documentation. Percent of critical services with tested recovery, actual versus target recovery times, unresolved exercise findings, single points of failure, and concentration exposure provide a clearer view than the number of continuity plans on file.

A mature resilience report also explains uncertainty. Recovery objectives may depend on untested suppliers, manual steps, or assumptions about staff availability, so auditors should distinguish demonstrated capability from estimated capability. Management can then prioritize the gaps that create the largest difference between target and proven recovery. This framing avoids false precision and makes investment decisions clearer: a service that has repeatedly met its objective under realistic exercises is fundamentally different from one whose recovery time exists only in a plan.

Evidence of improvement matters too. When later exercises close earlier gaps and recovery performance becomes more predictable, auditors can show that resilience governance is learning from tests rather than simply accumulating plans, findings, and remediation tickets.

Filed under Technology Fundamentals