06 Availability
The Availability Trust Services category focuses on whether systems are available for operation and use as committed or agreed.
Availability does not simply mean:
The system is online.A stronger availability program asks:
What availability did we promise? ↓What infrastructure supports it? ↓What could cause disruption? ↓How do we detect failure? ↓How do we recover? ↓Can we prove recovery works?For SOC 2, Availability is especially relevant for services where downtime can materially affect customers.
Examples include:
Cloud Platforms
SaaS Applications
Payment Platforms
Healthcare Systems
Identity Providers
Security Services
Customer-Facing ApplicationsThe category may include controls supporting:
-
Service availability.
-
Capacity.
-
redundancy.
-
backup.
-
disaster recovery.
-
monitoring.
-
incident management.
-
environmental protections.
-
recovery testing.
Learning Objectives
Section titled “Learning Objectives”By the end of this lesson, you will be able to:
-
Explain the SOC 2 Availability category.
-
Understand availability commitments.
-
Distinguish availability from security.
-
Understand SLA and availability objectives.
-
Build capacity-management controls.
-
Understand resilience and redundancy.
-
Govern single points of failure.
-
Build availability-monitoring controls.
-
Define backup controls.
-
Understand RTO and RPO.
-
Design disaster-recovery controls.
-
Understand business-continuity dependencies.
-
Test recovery procedures.
-
Manage availability incidents.
-
Assess third-party availability dependencies.
-
Define Availability evidence.
-
Test design and operating effectiveness.
-
Build practical SOC 2 Availability artifacts.
1. What Is Availability?
Section titled “1. What Is Availability?”Availability concerns whether information and systems are available for operation and use according to the organization’s commitments and system requirements.
Conceptually:
Service Commitment ↓Availability Requirement ↓Architecture ↓Monitoring ↓Recovery ↓Evidence2. Availability Is Based on Commitments
Section titled “2. Availability Is Based on Commitments”The organization should first understand what it has promised.
Examples:
99.9% Uptime
24 × 7 Service
4-Hour Recovery
Daily Backup
Regional ResilienceThese commitments may come from:
Customer Contracts
SLAs
Policies
Architecture Standards
Business Requirements3. Availability vs Security
Section titled “3. Availability vs Security”Security asks:
Is the system protected from unauthorized access and threats?
Availability asks:
Can authorized users access the system when required?
These overlap.
Example:
DDoS Attack ↓Security Event +Availability Impact4. Availability vs Business Continuity
Section titled “4. Availability vs Business Continuity”Availability often focuses on the service or system.
Business continuity is broader.
Availability→ Can the technology service operate?
Business Continuity→ Can the business continue critical operations?The two should support one another.
5. Service Availability Requirements
Section titled “5. Service Availability Requirements”For each critical service document:
Service
Business Criticality
Availability Requirement
RTO
RPO
Support Hours
Owner6. Availability Commitment Example
Section titled “6. Availability Commitment Example”Example:
Service:CloudCRM
Availability Commitment:99.9% monthly uptimeControls should support that commitment.
7. SLA
Section titled “7. SLA”A Service Level Agreement may define:
Availability Percentage
Support Response
Service Credits
Maintenance Windows
ExclusionsSOC 2 Availability assurance evaluates controls that support these commitments.
8. SLA Is Not an Availability Control by Itself
Section titled “8. SLA Is Not an Availability Control by Itself”Important:
SLA≠ControlThe SLA defines what should be achieved.
Controls help achieve it.
9. Availability Percentage
Section titled “9. Availability Percentage”A service may define:
99.9%But the organization should also define:
How measured?
What period?
Which components?
What exclusions?10. Planned Maintenance
Section titled “10. Planned Maintenance”Availability measurement may exclude approved maintenance windows.
Example:
Monthly Maintenance Window:Sunday 02:00–04:00This should be defined and communicated.
11. System Criticality
Section titled “11. System Criticality”Not every system requires the same availability.
Example:
| System | Criticality | Target |
|---|---|---|
| Customer Portal | Critical | 99.99% |
| Internal Wiki | Medium | 99.5% |
| Test Environment | Low | Best Effort |
12. Business Impact Analysis
Section titled “12. Business Impact Analysis”Availability requirements should be informed by business impact.
Ask:
If this service stops:
What customer impact?
What financial impact?
What operational impact?
What regulatory impact?13. BIA Output
Section titled “13. BIA Output”A Business Impact Analysis may define:
Critical Services
Maximum Tolerable Downtime
RTO
RPO
Dependencies14. Recovery Time Objective
Section titled “14. Recovery Time Objective”RTO means:
Recovery Time ObjectiveIt defines how quickly a system should be restored after disruption.
Example:
RTO:4 Hours15. RTO Example
Section titled “15. RTO Example”Incident begins:
10:00RTO:
4 HoursTarget recovery:
By 14:0016. Recovery Point Objective
Section titled “16. Recovery Point Objective”RPO means:
Recovery Point ObjectiveIt defines the acceptable amount of data loss measured in time.
Example:
RPO:1 Hour17. RPO Example
Section titled “17. RPO Example”Failure occurs:
16:00RPO:
1 HourThe organization should be able to recover data to approximately:
15:00or later.
18. RTO vs RPO
Section titled “18. RTO vs RPO”RTO→ How long can service be unavailable?
RPO→ How much data can be lost?Both should drive architecture and backup design.
19. Build RTO/RPO Register
Section titled “19. Build RTO/RPO Register”Create:
01 RTO-RPO RegisterUse:
| Service | Criticality | RTO | RPO | Owner |
|---|
20. Capacity Management
Section titled “20. Capacity Management”Availability can fail because the system lacks resources.
Examples:
CPU Exhaustion
Memory Exhaustion
Disk Capacity
Connection Limits
Database Capacity
Network Bandwidth21. Capacity Monitoring
Section titled “21. Capacity Monitoring”A basic model:
Resource Usage ↓Threshold ↓Alert ↓Action22. Capacity Control
Section titled “22. Capacity Control”Example:
Critical infrastructure capacity is monitored against defined thresholds, and alerts are investigated before resource exhaustion affects service availability.
23. Capacity Evidence
Section titled “23. Capacity Evidence”Examples:
Monitoring Dashboard
Capacity Alerts
Scaling Events
Capacity Reviews24. Scaling
Section titled “24. Scaling”Modern platforms may support:
Auto Scaling
Load Balancing
Elastic CapacityThese can support Availability controls.
25. Capacity Forecasting
Section titled “25. Capacity Forecasting”Organizations should also consider anticipated demand.
Examples:
Customer Growth
Seasonal Peaks
Product Launch
Marketing Campaign
Acquisition26. Infrastructure Resilience
Section titled “26. Infrastructure Resilience”Resilience means designing systems to continue or recover when components fail.
A resilient design may include:
Redundancy
Failover
Multiple Zones
Multiple Regions
Backup
Alternative Connectivity27. Single Point of Failure
Section titled “27. Single Point of Failure”A single point of failure exists when one component can disrupt the whole service.
Example:
Application ↓Single Database ServerIf the database fails:
Entire Service Fails28. Identify Single Points of Failure
Section titled “28. Identify Single Points of Failure”Review:
Compute
Database
Storage
Network
Identity
DNS
Cloud Region
Third Parties29. Redundancy
Section titled “29. Redundancy”Redundancy may include:
Multiple Servers
Multiple Network Paths
Replicated Databases
Multiple Availability Zones30. High Availability Architecture
Section titled “30. High Availability Architecture”Example:
Load Balancer / \ / \ App Zone A App Zone B \ / \ / Replicated DatabaseThis reduces dependency on one component.
31. Availability Zones
Section titled “31. Availability Zones”Cloud environments may distribute workloads across multiple availability zones.
Example:
Zone A+Zone BIf one zone fails:
Service Continuesif designed appropriately.
32. Multi-Region Architecture
Section titled “32. Multi-Region Architecture”For higher-criticality systems:
Primary Region ↓Secondary Regionmay support recovery from regional failure.
33. Multi-Region Does Not Automatically Mean Resilient
Section titled “33. Multi-Region Does Not Automatically Mean Resilient”If both regions depend on:
One Identity Provider
One DNS Provider
One Database Dependencysignificant concentration risk remains.
34. Dependency Mapping
Section titled “34. Dependency Mapping”For every critical service identify:
Application
Database
Network
Identity
Cloud Provider
DNS
CDN
SaaS Dependencies35. Availability Dependency Map
Section titled “35. Availability Dependency Map”Create:
02 Availability Dependency MapExample:
Customer Portal ↓DNS ↓CDN ↓Cloud Load Balancer ↓Application ↓Database ↓Identity Provider36. Availability Monitoring
Section titled “36. Availability Monitoring”Organizations should detect outages quickly.
Monitoring may include:
Application Health
Infrastructure Health
Network Health
Synthetic Tests
Transaction Monitoring37. External Monitoring
Section titled “37. External Monitoring”Internal monitoring alone may miss certain failures.
Example:
Server Says:Healthybut:
Customers Cannot Reach ServiceExternal synthetic monitoring can identify this.
38. Synthetic Monitoring
Section titled “38. Synthetic Monitoring”Example:
Every Minute ↓Login ↓Perform Test Transaction ↓Measure ResponseThis validates customer-facing availability.
39. Availability Monitoring Control
Section titled “39. Availability Monitoring Control”Example:
Critical production services are continuously monitored for availability, performance, and infrastructure health, with defined escalation thresholds.
40. Availability Monitoring Evidence
Section titled “40. Availability Monitoring Evidence”Examples:
Uptime Reports
Health Checks
Synthetic Monitoring
Alert Records
Incident Tickets41. Monitoring Coverage
Section titled “41. Monitoring Coverage”Build:
| Service | Availability Monitoring | Alerting | Owner |
|---|
Every critical service should be represented.
42. Alert Thresholds
Section titled “42. Alert Thresholds”Examples:
Service Down
Error Rate >5%
Latency >2 Seconds
CPU >90%Thresholds should align with service risk.
43. Alert Escalation
Section titled “43. Alert Escalation”A basic workflow:
Alert ↓On-Call Engineer ↓Incident ↓Escalation44. 24×7 Services
Section titled “44. 24×7 Services”If a service commitment is:
24 × 7but support monitoring only operates:
09:00–17:00there may be a control-design issue.
45. Incident Management
Section titled “45. Incident Management”Availability incidents should follow a structured process.
Detect ↓Triage ↓Severity ↓Restore Service ↓Communicate ↓Root Cause ↓Corrective Action46. Availability Incident Examples
Section titled “46. Availability Incident Examples”Examples:
Database Outage
Cloud Region Failure
DNS Failure
Network Failure
Capacity Exhaustion
Deployment Failure47. Incident Severity
Section titled “47. Incident Severity”Example:
SEV-1Critical Service Unavailable
SEV-2Major Degradation
SEV-3Limited Impact48. Availability Incident Control
Section titled “48. Availability Incident Control”Critical service disruptions are documented, prioritized, escalated, restored, and reviewed according to the incident-management process.
49. Mean Time to Detect
Section titled “49. Mean Time to Detect”MTTD measures:
Failure ↓Detection50. Mean Time to Restore
Section titled “50. Mean Time to Restore”MTTR may measure:
Incident Begins ↓Service Restoreddepending on organizational definition.
51. Availability Metrics
Section titled “51. Availability Metrics”Useful metrics include:
Uptime
MTTD
MTTR
Number of Major Incidents
SLA Breaches
Recovery Success52. Backup
Section titled “52. Backup”Backup is a key Availability control.
But:
Backup Exists≠System Can RecoverBackup is only one part of recovery assurance.
53. Backup Scope
Section titled “53. Backup Scope”Identify:
Databases
Object Storage
Virtual Machines
Configurations
SaaS Data
Critical Files54. Backup Frequency
Section titled “54. Backup Frequency”Examples:
Continuous
Hourly
Daily
WeeklyFrequency should support RPO.
55. RPO and Backup Frequency
Section titled “55. RPO and Backup Frequency”If:
RPO = 1 Hourbut:
Backup = Once Per Daythe design does not support the objective.
56. Backup Retention
Section titled “56. Backup Retention”Define:
7 Days
30 Days
1 Yearbased on requirements.
57. Backup Encryption
Section titled “57. Backup Encryption”Backups may contain the same sensitive data as production.
Therefore controls may include:
Encryption
Access Restrictions
Logging58. Backup Isolation
Section titled “58. Backup Isolation”Ransomware risk may require:
Immutable Backups
Offline Copies
Separate Accounts
Restricted Credentials59. Backup Failure Monitoring
Section titled “59. Backup Failure Monitoring”A failed backup should generate:
Alert ↓Investigation ↓Correction60. Backup Control
Section titled “60. Backup Control”Example:
Critical systems and data are backed up according to defined schedules and retention requirements, and failed backup jobs are monitored and remediated.
61. Backup Evidence
Section titled “61. Backup Evidence”Examples:
Backup Job Reports
Failure Alerts
Retention Configuration
Encryption Configuration62. Restore Testing
Section titled “62. Restore Testing”A critical control is:
Can We Restore?A backup can exist but be:
Corrupted
Incomplete
Misconfigured
Unavailable63. Restore Test
Section titled “63. Restore Test”Example:
Backup ↓Restore to Test Environment ↓Validate Data ↓Validate Application64. Restore Test Evidence
Section titled “64. Restore Test Evidence”Include:
Test Date
System
Backup Used
Recovery Time
Data Recovered
Result
Issues65. Build Backup Evidence Register
Section titled “65. Build Backup Evidence Register”Create:
03 Backup Evidence RegisterUse:
| System | Backup | Frequency | Last Success | Restore Test |
|---|
66. Disaster Recovery
Section titled “66. Disaster Recovery”Disaster Recovery focuses on restoring technology services after major disruption.
Scenarios may include:
Regional Cloud Failure
Data Center Failure
Ransomware
Major Network Failure
Storage Corruption67. Disaster Recovery Plan
Section titled “67. Disaster Recovery Plan”The plan should define:
Systems
Priority
Dependencies
RTO
RPO
Recovery Steps
Roles
Communication68. DR Architecture
Section titled “68. DR Architecture”Example:
Primary Production Region ↓Replication ↓Recovery Region69. Recovery Runbooks
Section titled “69. Recovery Runbooks”Technical teams should know:
How to fail over?
How to restore?
How to validate?
How to fail back?70. Disaster Recovery Testing
Section titled “70. Disaster Recovery Testing”DR plans should be tested.
A test may include:
Tabletop
Partial Technical Recovery
Full Failover
Full Regional Recovery71. Tabletop Exercise
Section titled “71. Tabletop Exercise”A tabletop asks:
What would each team do if the primary region failed?
It tests:
Roles
Decision Making
Communication
Dependenciesbut may not prove technical recovery.
72. Technical Recovery Test
Section titled “72. Technical Recovery Test”A stronger test may:
Recover Database
Start Application
Validate Connectivity
Validate Customer Workflow73. Full Failover Test
Section titled “73. Full Failover Test”For mature environments:
Production / Controlled Environment ↓Fail to Secondary Region ↓Validate ServiceThis provides stronger assurance but carries operational risk.
74. Recovery Test Register
Section titled “74. Recovery Test Register”Create:
04 Disaster Recovery Test RegisterUse:
| Service | Test Type | Date | RTO | Actual | Result |
|---|
75. Compare Actual Recovery to RTO
Section titled “75. Compare Actual Recovery to RTO”Example:
RTO:4 Hours
Actual Recovery:2 Hours 45 Minutes
Result:Pass76. Recovery Failure
Section titled “76. Recovery Failure”Example:
RTO:2 Hours
Actual:6 HoursResult:
Control Gap77. Compare Recovered Data to RPO
Section titled “77. Compare Recovered Data to RPO”Example:
RPO:30 Minutes
Recovered Data Loss:15 Minutes
Pass78. Recovery Test Findings
Section titled “78. Recovery Test Findings”Common issues include:
Outdated Runbook
Missing Credentials
Broken Replication
DNS Failure
Unknown Dependency
Insufficient Capacity79. Post-Test Review
Section titled “79. Post-Test Review”After each major test:
Test ↓Findings ↓Root Cause ↓Corrective Action ↓Retest80. DR Control
Section titled “80. DR Control”Example:
Disaster-recovery capabilities for critical systems are periodically tested against established recovery objectives, and deficiencies are tracked to remediation.
81. Business Continuity Dependencies
Section titled “81. Business Continuity Dependencies”A technically recovered system may still be unusable if:
Employees Cannot Work
Identity Is Unavailable
Critical Supplier Is Down
Communication Channels FailAvailability planning should understand these dependencies.
82. People Dependencies
Section titled “82. People Dependencies”Ask:
Who performs recovery?
Are they available 24×7?
Is knowledge concentrated in one person?83. Documentation Dependencies
Section titled “83. Documentation Dependencies”Recovery should not rely only on:
One Engineer Remembers HowUse documented runbooks.
84. Credential Dependencies
Section titled “84. Credential Dependencies”Recovery credentials should be:
Available
Protected
Tested85. Third-Party Availability
Section titled “85. Third-Party Availability”Critical services may depend on vendors.
Examples:
Cloud Provider
DNS
CDN
Identity Provider
Payment Provider
Monitoring SaaS86. Provider Assurance
Section titled “86. Provider Assurance”For critical providers review:
SOC 2 Availability Scope
SLA
Resilience Architecture
Incident History
DR Commitments87. Provider Availability Is Not Your Availability
Section titled “87. Provider Availability Is Not Your Availability”Important:
Cloud Provider Availabledoes not guarantee:
Your Application AvailableCustomer architecture still matters.
88. Example
Section titled “88. Example”Provider offers:
99.99% Infrastructure AvailabilityCustomer deploys:
Single InstanceSingle ZoneCustomer architecture may remain fragile.
89. Availability Shared Responsibility
Section titled “89. Availability Shared Responsibility”Provider may own:
Physical Infrastructure
Platform AvailabilityCustomer may own:
Application Architecture
Failover
Backup Configuration
Recovery Testing90. SaaS Availability
Section titled “90. SaaS Availability”For SaaS providers assess:
Availability Commitment
Historical Outages
Provider Monitoring
Backup
Recovery
Customer Communication91. SaaS Customer Responsibilities
Section titled “91. SaaS Customer Responsibilities”Customer may still need:
Data Export
Alternative Business Process
Manual Workaround
Exit Strategyfor critical SaaS.
92. Availability Risk Register
Section titled “92. Availability Risk Register”Typical risks include:
Cloud Region Outage
Database Failure
Capacity Exhaustion
DNS Failure
Ransomware
Backup Failure
Provider Outage
Recovery Failure93. Risk-to-Control Mapping
Section titled “93. Risk-to-Control Mapping”Example:
Risk:Regional Cloud Failure ↓Controls:Multi-Region RecoveryBackupsDR TestIncident Response94. Availability Control Matrix
Section titled “94. Availability Control Matrix”Create:
05 Availability Control MatrixUse:
| Control ID | Risk | Control | Owner | Evidence |
|---|
95. Example Availability Controls
Section titled “95. Example Availability Controls”AVL-001Availability Commitments
AVL-002Service Monitoring
AVL-003Capacity Monitoring
AVL-004Infrastructure Resilience
AVL-005Backup
AVL-006Backup Failure Monitoring
AVL-007Restore Testing
AVL-008Disaster Recovery
AVL-009Recovery Testing
AVL-010Availability Incident Management96. Availability Evidence Matrix
Section titled “96. Availability Evidence Matrix”Create:
06 Availability Evidence MatrixUse:
| Control | Evidence | Frequency | Source | Owner |
|---|
97. Evidence — Availability Monitoring
Section titled “97. Evidence — Availability Monitoring”Examples:
Uptime Reports
Synthetic Monitoring
Alerts
Incident Tickets98. Evidence — Capacity
Section titled “98. Evidence — Capacity”Examples:
Resource Dashboard
Threshold Alerts
Capacity Review
Scaling Records99. Evidence — Backup
Section titled “99. Evidence — Backup”Examples:
Backup Reports
Failure Logs
Retention Settings
Encryption100. Evidence — Recovery
Section titled “100. Evidence — Recovery”Examples:
DR Test
Restore Test
RTO/RPO Results
Corrective Actions101. Availability Control Testing
Section titled “101. Availability Control Testing”For each control define:
Requirement
Population
Evidence
Test
Exceptions
Conclusion102. Test Availability Monitoring
Section titled “102. Test Availability Monitoring”Population:
Critical Production ServicesVerify:
Monitoring Enabled?
Alerting Enabled?
Owner Defined?103. Test Backup
Section titled “103. Test Backup”Population:
Critical SystemsVerify:
Backup Enabled
Correct Frequency
Retention
Encryption
Success104. Test Restore
Section titled “104. Test Restore”Sample restore tests.
Validate:
Completed?
Successful?
Within RTO?
Within RPO?105. Test DR
Section titled “105. Test DR”For each critical service:
Required Test Frequency
Latest Test
Results
Findings
Remediation106. Test Incident Response
Section titled “106. Test Incident Response”Sample major availability incidents.
Trace:
Detection
Escalation
Restoration
Communication
Root Cause
Closure107. Design Effectiveness
Section titled “107. Design Effectiveness”Ask:
If these controls operate as designed, will the organization reasonably meet its Availability commitments?
Example:
RTO:1 Hour
Recovery Architecture:Manual rebuild from backupestimated at 12 hoursThis is a design gap.
108. Operating Effectiveness
Section titled “108. Operating Effectiveness”Ask:
Did the control operate throughout the period?
Example:
Monthly Backup Review
Jan ✓Feb ✓Mar ✗Apr ✓Result:
Partially Effective109. Complete Population Testing
Section titled “109. Complete Population Testing”Technical evidence may allow testing of:
100% Backup Coverage
100% Monitoring Coverage
100% Replication Status110. Sampling
Section titled “110. Sampling”Manual controls may require sampling:
Recovery Tests
Incident Reviews
Capacity Reviews
Provider Assessments111. Availability Finding Example — Backup
Section titled “111. Availability Finding Example — Backup”The Backup Standard requires daily backups for all critical production databases. Two of twenty critical databases were not included in the automated backup policy during the examination period.
Risk:
Data Loss
Extended Recovery Time112. Availability Finding Example — DR
Section titled “112. Availability Finding Example — DR”The organization requires annual disaster-recovery testing for critical services. Three critical applications had not undergone recovery testing within the required period.
113. Availability Finding Example — Capacity
Section titled “113. Availability Finding Example — Capacity”Capacity alerts were not configured for the primary production database despite recurring utilization above 90%.
114. Availability Finding Example — Single Point of Failure
Section titled “114. Availability Finding Example — Single Point of Failure”The customer-facing application relies on a single production database instance without tested failover despite a four-hour recovery commitment.
115. Root Cause Analysis
Section titled “115. Root Cause Analysis”Example:
DR Test Missing ↓No Scheduled Test ↓No Central Recovery CalendarRoot cause:
Disaster-recovery testing is managed independently by application teams without centralized governance or escalation.
116. Correction vs Corrective Action
Section titled “116. Correction vs Corrective Action”Correction:
Perform overdue DR test.Corrective action:
Implement centralized DR test schedulingand escalation for all critical services.117. Availability Exceptions
Section titled “117. Availability Exceptions”Example requirement:
Multi-Zone Deployment RequiredLegacy platform:
Cannot Support ItException should include:
Risk
Compensating Controls
Owner
Approval
Expiry
Migration Plan118. Possible Compensating Controls
Section titled “118. Possible Compensating Controls”Examples:
Frequent Backups
Warm Standby
Enhanced Monitoring
Reduced Recovery Time
Replacement Roadmap119. Availability Dashboard
Section titled “119. Availability Dashboard”Useful metrics include:
Service Availability
SLA Compliance
Backup Coverage
Restore Test Completion
DR Test Completion
RTO Achievement
RPO Achievement
Major Incidents120. Example Availability Dashboard
Section titled “120. Example Availability Dashboard”| Metric | Target | Current |
|---|---|---|
| Critical Service Monitoring | 100% | 100% |
| Backup Coverage | 100% | 98% |
| Restore Tests Current | 100% | 95% |
| DR Tests Current | 100% | 90% |
| RTO Tests Passed | 100% | 92% |
121. Availability KPI
Section titled “121. Availability KPI”Example:
KPI:Percentage of critical servicesmeeting defined availability objective122. Availability KRI
Section titled “122. Availability KRI”Example:
KRI:Number of critical serviceswithout current recovery testingTolerance:
0123. Backup KRI
Section titled “123. Backup KRI”Example:
Critical workloadswithout successful backupwithin required windowTolerance:
0124. Recovery KRI
Section titled “124. Recovery KRI”Example:
Number of recovery teststhat exceed required RTO125. Audit Walkthrough — Availability Monitoring
Section titled “125. Audit Walkthrough — Availability Monitoring”Auditor selects:
Critical SaaS ServiceTrace:
Availability Requirement ↓Monitoring ↓Alerts ↓Incidents ↓Reports126. Audit Walkthrough — Backup
Section titled “126. Audit Walkthrough — Backup”Auditor selects:
Production DatabaseVerify:
Backup Schedule
Successful Jobs
Retention
Encryption
Restore Test127. Audit Walkthrough — Disaster Recovery
Section titled “127. Audit Walkthrough — Disaster Recovery”Auditor selects:
Critical ApplicationReview:
RTO
RPO
Recovery Plan
Latest Test
Result
Findings
Remediation128. Audit Walkthrough — Availability Incident
Section titled “128. Audit Walkthrough — Availability Incident”Select a major outage.
Trace:
Detection ↓Escalation ↓Recovery ↓Customer Communication ↓Root Cause ↓Corrective Action129. Common Availability Mistakes
Section titled “129. Common Availability Mistakes”Mistake 1 — Availability Equals Uptime Report
Section titled “Mistake 1 — Availability Equals Uptime Report”Supporting controls are ignored.
Mistake 2 — SLA Treated as a Control
Section titled “Mistake 2 — SLA Treated as a Control”An SLA is a commitment, not a control.
Mistake 3 — RTO/RPO Undefined
Section titled “Mistake 3 — RTO/RPO Undefined”Recovery architecture has no measurable target.
Mistake 4 — Backup Equals Recovery
Section titled “Mistake 4 — Backup Equals Recovery”Backups may be unusable.
Mistake 5 — DR Plan Never Tested
Section titled “Mistake 5 — DR Plan Never Tested”Recovery remains theoretical.
Mistake 6 — Monitoring Covers Infrastructure Only
Section titled “Mistake 6 — Monitoring Covers Infrastructure Only”Application or customer experience failures may be missed.
Mistake 7 — Single Points of Failure Not Documented
Section titled “Mistake 7 — Single Points of Failure Not Documented”Architecture risk remains hidden.
Mistake 8 — Provider SLA Equals Customer Resilience
Section titled “Mistake 8 — Provider SLA Equals Customer Resilience”Customer design may remain fragile.
Mistake 9 — Recovery Findings Not Remediated
Section titled “Mistake 9 — Recovery Findings Not Remediated”The same failures appear every test.
Mistake 10 — Availability Scope Too Broad
Section titled “Mistake 10 — Availability Scope Too Broad”Controls do not clearly map to specific service commitments.
130. Weak Availability Model
Section titled “130. Weak Availability Model”Cloud Provider SLA +Daily Backup =Assume Available131. Strong Availability Model
Section titled “131. Strong Availability Model”Business Requirement ↓Availability Commitment ↓RTO / RPO ↓Resilient Architecture ↓Monitoring ↓Backup ↓Recovery ↓Testing ↓Incident Management ↓Evidence132. Practical Activity — Build Availability Control Matrix
Section titled “132. Practical Activity — Build Availability Control Matrix”Create:
01 Availability Control MatrixInclude at least 15 controls across:
Availability Commitments
Capacity
Monitoring
Resilience
Backup
Recovery
Incident Management
Third Parties133. Practical Activity — Build RTO/RPO Register
Section titled “133. Practical Activity — Build RTO/RPO Register”Create:
02 RTO-RPO RegisterAdd at least ten fictional critical systems.
Use:
| Service | Criticality | RTO | RPO | Owner |
|---|
134. Practical Activity — Build Backup Evidence Register
Section titled “134. Practical Activity — Build Backup Evidence Register”Create:
03 Backup Evidence RegisterTrack:
System
Frequency
Retention
Latest Backup
Encryption
Last Restore Test135. Practical Activity — Build Disaster Recovery Test Register
Section titled “135. Practical Activity — Build Disaster Recovery Test Register”Create:
04 Disaster Recovery Test RegisterUse:
| Service | Test | RTO | Actual | RPO | Result |
|---|
136. Practical Activity — Build Availability Monitoring Matrix
Section titled “136. Practical Activity — Build Availability Monitoring Matrix”Create:
05 Availability Monitoring MatrixUse:
| Service | Monitor | Alert | Escalation | Owner |
|---|
137. Practical Activity — Build Dependency Register
Section titled “137. Practical Activity — Build Dependency Register”Create:
06 Availability Dependency RegisterInclude:
Identity Provider
DNS
Cloud Provider
Database
CDN
Network
Critical SaaS138. Practical Activity — Build Availability Control Testing Checklist
Section titled “138. Practical Activity — Build Availability Control Testing Checklist”Create:
07 Availability Control Testing ChecklistTest:
Service Monitoring
Capacity
Backup
Backup Failure Monitoring
Restore Testing
Disaster Recovery
RTO
RPO
Major Incident Management139. Availability Readiness Checklist
Section titled “139. Availability Readiness Checklist”-
Critical services identified.
-
Availability commitments documented.
-
System boundaries defined.
-
Owners identified.
RTO/RPO
Section titled “RTO/RPO”-
RTO defined.
-
RPO defined.
-
Objectives approved.
-
Architecture aligns with objectives.
Capacity
Section titled “Capacity”-
Capacity monitored.
-
Thresholds defined.
-
Alerts configured.
-
Forecasting performed where appropriate.
Resilience
Section titled “Resilience”-
Single points of failure identified.
-
Redundancy implemented.
-
Failover defined.
-
Dependencies documented.
Monitoring
Section titled “Monitoring”-
Critical services monitored.
-
External monitoring used where required.
-
Alerts defined.
-
Escalation established.
-
Uptime measured.
Backup
Section titled “Backup”-
Critical data backed up.
-
Frequency supports RPO.
-
Retention defined.
-
Encryption implemented.
-
Failure alerts configured.
Recovery
Section titled “Recovery”-
Recovery plans documented.
-
Recovery runbooks current.
-
Restore testing performed.
-
DR testing performed.
-
RTO results recorded.
-
RPO results recorded.
-
Findings remediated.
Incidents
Section titled “Incidents”-
Availability incidents classified.
-
Escalation defined.
-
Customer communication defined.
-
Root cause performed.
-
Corrective actions tracked.
Third Parties
Section titled “Third Parties”-
Critical providers identified.
-
Availability commitments reviewed.
-
Provider assurance reviewed.
-
Concentration risk considered.
-
Alternative arrangements considered where required.
140. GRC Analyst Responsibilities
Section titled “140. GRC Analyst Responsibilities”A GRC professional supporting SOC 2 Availability may:
-
Identify in-scope services.
-
Document availability commitments.
-
Maintain RTO/RPO registers.
-
Map availability risks to controls.
-
Review architecture dependencies.
-
Maintain Availability control matrices.
-
Review backup evidence.
-
Review recovery testing.
-
Validate DR test frequency.
-
Review provider availability assurance.
-
Assess control design.
-
Test operating effectiveness.
-
Document findings.
-
Track remediation.
-
Maintain Availability KPIs and KRIs.
-
Support external SOC auditors.
GRC connects:
Business Owners
Engineering
Cloud
IT Operations
SRE
Security
Risk
Business Continuity
Vendors
Auditors141. Availability Maturity Model
Section titled “141. Availability Maturity Model”Level 1 — Reactive
Section titled “Level 1 — Reactive”Outage Occurs ↓Team RespondsLevel 2 — Defined
Section titled “Level 2 — Defined”Backups
Monitoring
Recovery PlansLevel 3 — Risk Based
Section titled “Level 3 — Risk Based”RTO
RPO
Criticality
Recovery TestingLevel 4 — Integrated
Section titled “Level 4 — Integrated”Resilience
Automated Monitoring
Metrics
Provider AssuranceLevel 5 — Continuous Resilience Assurance
Section titled “Level 5 — Continuous Resilience Assurance”Continuous Availability Monitoring
Automated Recovery Validation
Dependency Monitoring
Dynamic Risk Signals142. Availability Mindset
Section titled “142. Availability Mindset”For every critical service ask:
What have we promised customers?
How critical is the service?
What is the RTO?
What is the RPO?
Which components could fail?
Which dependencies could fail?
How quickly would we detect failure?
Do we have sufficient capacity?
Are backups complete?
Can we restore them?
Can we recover the whole service?
When was recovery last tested?
Did the test meet RTO and RPO?
What happened during real outages?
Were root causes corrected?
Which providers do we depend on?
Can we prove all of this to an auditor?When these questions can be answered, Availability becomes demonstrable operational resilience rather than simply an uptime percentage.
Key Takeaways
Section titled “Key Takeaways”-
Availability evaluates whether systems are available for use according to commitments and requirements.
-
Availability commitments should be clearly defined before controls are designed.
-
SLAs express commitments but are not themselves availability controls.
-
System criticality should drive availability requirements.
-
RTO defines how quickly service must recover.
-
RPO defines acceptable data loss.
-
Capacity management helps prevent resource-exhaustion outages.
-
Resilient architectures reduce single points of failure.
-
Availability monitoring should include customer-facing service health where appropriate.
-
Backup success does not prove recoverability.
-
Restore and disaster-recovery testing are essential.
-
Recovery tests should be evaluated against RTO and RPO.
-
Availability incidents should be detected, escalated, restored, reviewed, and remediated.
-
Third-party dependencies and concentration risk are important Availability considerations.
-
Provider availability does not automatically guarantee application availability.
-
GRC connects Availability commitments to risks, controls, evidence, testing, findings, and assurance.
Knowledge Check
Section titled “Knowledge Check”Before continuing, make sure you can answer:
-
What does the SOC 2 Availability category address?
-
Why should availability commitments be defined first?
-
How is an SLA different from a control?
-
What is RTO?
-
What is RPO?
-
How does backup frequency relate to RPO?
-
What is capacity management?
-
What is a single point of failure?
-
How does redundancy improve availability?
-
Why does multi-region architecture not automatically guarantee resilience?
-
What should availability monitoring include?
-
Why can external synthetic monitoring be useful?
-
Why does backup success not prove recovery?
-
What should a restore test validate?
-
What should a DR plan contain?
-
What is the difference between a tabletop and a technical recovery test?
-
Why should recovery results be compared with RTO and RPO?
-
How can third-party providers affect availability?
-
What evidence can support Availability controls?
-
What role does GRC play in SOC 2 Availability?
What’s Next?
Section titled “What’s Next?”➡️ Next: 07 — Processing Integrity
In the next lesson, you will move into the Processing Integrity Trust Services category and learn how organizations demonstrate that system processing is complete, valid, accurate, timely, and authorized.
You will work through:
Input Validation ↓Transaction Authorization ↓Processing Logic ↓Completeness ↓Accuracy ↓Error Handling ↓Reconciliation ↓Output Validation ↓Exception Management ↓Evidence & TestingYou will also build practical artifacts including a Processing Integrity Control Matrix, Transaction Flow Map, Input/Output Control Register, Reconciliation Register, Processing Exception Register, and Processing Integrity Control Testing Checklist.