Cloud Environment Health Assessment Lab
A cloud environment can be online and still be unhealthy. Operational health means the environment is available, secure, resilient, observable, recoverable, supportable, and appropriately governed.
Welcome to Lab 26 of the CompTIA Cloud+ practical lab sequence.
In the previous lab, you learned how to troubleshoot individual cloud failures:
Symptom βScope βEvidence βHypothesis βRoot Cause βRemediationNow you will move from:
Reactive Troubleshootingto:
Proactive Cloud Health AssessmentInstead of waiting for something to fail, you will systematically evaluate the environment and identify weaknesses before they become incidents.
π― Mission Information
Section titled βπ― Mission Informationβ| Item | Details |
|---|---|
| Lab | 26 β Cloud Environment Health Assessment Lab |
| Difficulty | Intermediate |
| Estimated Time | 180β240 Minutes |
| Certification Alignment | CompTIA Cloud+ |
| Primary Focus | Cloud Operational Health Assessment |
| Previous Lab | 25 β Cloud Troubleshooting Lab |
| Career Alignment | Cloud Engineer, Cloud Administrator, Cloud Operations Engineer, Cloud Consultant |
| Major Skills | Assessment, Monitoring, Security, Resilience, Backup, Cost, Operations |
| Deliverable | Cloud Health Assessment Report + Findings Register + Remediation Plan |
π’ Scenario
Section titled βπ’ ScenarioβYour organization operates a production cloud environment.
Management asks:
βIs our cloud environment healthy?β
A simple check shows:
Application:ONLINE
Virtual Machines:RUNNING
Database:RUNNINGAn inexperienced engineer might conclude:
Environment HealthyBut further investigation reveals:
CPU frequently reaches 95%
Storage is 92% full
Backups have not been tested
Several users have excessive permissions
One administrative port is publicly exposed
Monitoring covers only some workloads
Certificates expire soon
No disaster recovery test exists
Unused resources are increasing costsThe environment is:
RUNNINGbut not necessarily:
HEALTHYYour mission is to perform a structured cloud environment health assessment.
π― Lab Objectives
Section titled βπ― Lab ObjectivesβBy completing this lab, you should be able to:
-
define cloud operational health
-
establish assessment scope
-
inventory cloud resources
-
understand architecture dependencies
-
assess compute health
-
assess network health
-
assess storage health
-
assess identity and access
-
assess security controls
-
assess availability
-
assess scalability
-
assess monitoring
-
assess logging
-
assess backup
-
assess disaster recovery
-
assess automation
-
assess configuration management
-
identify configuration drift
-
assess operational readiness
-
assess cloud costs
-
identify unused resources
-
review quotas and limits
-
assess certificates and secrets
-
review incident readiness
-
classify findings
-
prioritize remediation
-
create a cloud health scorecard
-
create an executive health report
-
build a reusable health-assessment runbook
01 β Understand Cloud Environment Health
Section titled β01 β Understand Cloud Environment HealthβCloud health is broader than:
Resource Running?A useful health model is:
CLOUD HEALTH | +-----------------+-----------------+ | | | v v v Availability Security Performance | | | +-----------------+-----------------+ | +--------+--------+ | | v v Resilience Observability | | +--------+--------+ | v Operations02 β Define the Assessment Domains
Section titled β02 β Define the Assessment DomainsβAssess:
Architecture
Compute
Networking
Storage
Identity
Security
Availability
Performance
Scalability
Monitoring
Logging
Backup
Disaster Recovery
Automation
Configuration
Cost
Governance
Operations03 β Establish Assessment Scope
Section titled β03 β Establish Assessment ScopeβBefore starting, define:
Cloud Account / Subscription / Project
Region
Environment
Applications
Resources
Business Services
Assessment PeriodExample:
Environment:Production
Region:Primary Region
Application:Customer Portal
Assessment:Operational Health Review04 β Build the Assessment Scope Sheet
Section titled β04 β Build the Assessment Scope SheetβAssessment Name:
Cloud Platform:
Account / Subscription / Project:
Environment:
Region(s):
Applications:
Business Services:
Resource Groups / Projects:
Assessment Owner:
Technical Contacts:
Start Date:
Assessment Date:05 β Understand Business Context
Section titled β05 β Understand Business ContextβNot every workload has the same requirements.
Example:
Development Sandbox
Availability Requirement:LOWERversus:
Payment Application
Availability Requirement:HIGHAssessment findings should consider business importance.
06 β Identify Critical Services
Section titled β06 β Identify Critical ServicesβClassify workloads:
| Service | Criticality | Availability Requirement | Owner |
|---|---|---|---|
| Customer Portal | High | High | |
| Internal Wiki | Medium | Medium | |
| Development VM | Low | Low | |
| Production Database | Critical | Very High |
07 β Build the Resource Inventory
Section titled β07 β Build the Resource InventoryβYou cannot assess what you do not know exists.
Inventory:
Virtual Machines
Containers
Databases
Storage
Networks
Load Balancers
Public IPs
DNS
Identity Resources
Security Services
Monitoring
Backup Resources08 β Resource Inventory Template
Section titled β08 β Resource Inventory Templateβ| Resource | Type | Environment | Region | Owner | Criticality |
|---|---|---|---|---|---|
09 β Identify Resource Ownership
Section titled β09 β Identify Resource OwnershipβEvery important resource should ideally have:
Technical Owner
Business Owner
Purpose
Environment
CriticalityResources without owners create operational risk.
10 β Review Resource Metadata
Section titled β10 β Review Resource MetadataβWhere supported, review:
Tags
Labels
Naming
Environment
Application
Owner
Cost Center
Data Classification11 β Assess Architecture
Section titled β11 β Assess ArchitectureβDocument:
Users βDNS βLoad Balancer βApplication βDatabase βStorageThen identify dependencies.
12 β Build the Dependency Map
Section titled β12 β Build the Dependency MapβExample:
Internet | v DNS | v Load Balancer | +-------+-------+ | | v v App-01 App-02 | | +-------+-------+ | v Database | v Storage13 β Identify Single Points of Failure
Section titled β13 β Identify Single Points of FailureβLook for:
Single VM
Single Database
Single Network Path
Single Availability Zone
Single Authentication Dependency
Single Storage DependencyAsk:
What happens if this component fails?
14 β Assess Compute Health
Section titled β14 β Assess Compute HealthβReview:
Resource State
Platform Health
CPU
Memory
Disk
Network
Instance Size
Operating System
Patching
Availability
Scaling15 β Build the Compute Health Table
Section titled β15 β Build the Compute Health Tableβ| Compute Resource | CPU | Memory | Disk | Health | Finding |
|---|---|---|---|---|---|
| App-01 | |||||
| App-02 |
16 β Review CPU Utilization
Section titled β16 β Review CPU UtilizationβLook for:
Sustained High CPU
Frequent Spikes
Very Low Utilization
Unexpected ChangesHigh utilization may indicate:
Capacity Issue
Application Issue
Traffic Increase
Scaling Failure17 β Review Memory
Section titled β17 β Review MemoryβLook for:
Memory Pressure
Swap Usage
OOM Events
Memory Leaks
Incorrect Limits18 β Review Compute Sizing
Section titled β18 β Review Compute SizingβPossible findings:
Under-Provisionedor:
Over-ProvisionedExample:
VM Capacity:16 vCPU
Average CPU:4%This may represent an optimization opportunity.
19 β Review Compute Availability
Section titled β19 β Review Compute AvailabilityβAsk:
Multiple Instances?
Multiple Availability Zones?
Autoscaling?
Health Checks?
Automatic Replacement?20 β Review Operating System Health
Section titled β20 β Review Operating System HealthβAssess:
Supported OS?
Patching Current?
Disk Healthy?
Agents Running?
Logging Working?
Monitoring Working?21 β Assess Network Health
Section titled β21 β Assess Network HealthβReview:
Virtual Networks
Subnets
Routes
Firewalls
Security Rules
Public IPs
DNS
Load Balancers
Private Connectivity
Network Logs22 β Build the Network Architecture
Section titled β22 β Build the Network ArchitectureβInternet | vPublic Endpoint | vFirewall / Security Controls | vLoad Balancer | vApplication Subnet | vDatabase Subnet23 β Review Public Exposure
Section titled β23 β Review Public ExposureβInventory:
Public IP Addresses
Internet-Facing Load Balancers
Public Storage
Public Databases
Administrative PortsAsk:
Does this resource actually need to be publicly accessible?
24 β Review Administrative Access
Section titled β24 β Review Administrative AccessβPay particular attention to:
SSH β TCP 22
RDP β TCP 3389Check whether access is:
Restrictedor unnecessarily:
Internet-Wide25 β Review Firewall Rules
Section titled β25 β Review Firewall RulesβLook for:
Any β Any
0.0.0.0/0
Overly Broad Ports
Unused Rules
Duplicate Rules
Temporary Rules
Unowned Rules26 β Review Network Segmentation
Section titled β26 β Review Network SegmentationβDetermine whether:
Web
Application
Database
Managementworkloads are appropriately separated.
27 β Review Routes
Section titled β27 β Review RoutesβCheck for:
Incorrect Routes
Unexpected Internet Routes
Missing Private Routes
Asymmetric Routing
Obsolete Routes28 β Review DNS Health
Section titled β28 β Review DNS HealthβAssess:
Records
Resolution
TTL
Private Zones
Public Zones
Ownership
Redundancy29 β Review Load Balancers
Section titled β29 β Review Load BalancersβCheck:
Listeners
Backend Health
Health Checks
Certificates
Logging
Availability
Session Configuration30 β Assess Storage Health
Section titled β30 β Assess Storage HealthβInventory:
Block Storage
Object Storage
File Storage
Database Storage
Snapshots
Backup Storage31 β Review Storage Capacity
Section titled β31 β Review Storage CapacityβCheck:
Used Capacity
Available Capacity
Growth Rate
Thresholds
AlertsExample:
Capacity:1 TB
Used:920 GB
Utilization:92%Finding:
Capacity risk.
32 β Review Storage Performance
Section titled β32 β Review Storage PerformanceβAssess:
IOPS
Throughput
Latency
Queue DepthCompare performance with workload requirements.
33 β Review Storage Availability
Section titled β33 β Review Storage AvailabilityβDetermine whether storage provides the required:
Redundancy
Replication
Durability
Availability34 β Review Storage Security
Section titled β34 β Review Storage SecurityβCheck:
Encryption at Rest
Encryption in Transit
Access Policies
Public Access
Logging
Versioning35 β Review Storage Lifecycle
Section titled β35 β Review Storage LifecycleβLook for:
Old Logs
Old Backups
Unused Snapshots
Temporary Data
Old Object VersionsApply lifecycle policies where appropriate.
36 β Assess Identity Health
Section titled β36 β Assess Identity HealthβInventory:
Users
Groups
Roles
Service Accounts
Managed Identities
Applications
Privileged Accounts37 β Review Privileged Access
Section titled β37 β Review Privileged AccessβIdentify:
Administrators
Owners
Global Roles
Root-Level Access
Highly Privileged Service IdentitiesAsk:
Does each identity still require this level of access?
38 β Review Least Privilege
Section titled β38 β Review Least PrivilegeβPoor:
Application βAdministrator RoleBetter:
Application βRequired Permissions Only39 β Review Dormant Identities
Section titled β39 β Review Dormant IdentitiesβLook for:
Inactive Users
Former Employees
Unused Service Accounts
Old API Credentials
Unused Application Identities40 β Review MFA
Section titled β40 β Review MFAβVerify appropriate MFA coverage for:
Administrators
Privileged Users
Remote Access
Sensitive Operations41 β Review Service Identities
Section titled β41 β Review Service IdentitiesβCheck:
Purpose
Owner
Permissions
Credential Type
Credential Age
Last Used42 β Review Long-Lived Credentials
Section titled β42 β Review Long-Lived CredentialsβIdentify:
Static Access Keys
API Keys
Passwords
Certificates
TokensPrefer managed or short-lived identity mechanisms where supported.
43 β Assess Security Health
Section titled β43 β Assess Security HealthβReview:
Identity
Network
Data
Compute
Configuration
Logging
Vulnerability Management
Threat Detection44 β Review Security Baseline
Section titled β44 β Review Security BaselineβCreate:
[ ] MFA enabled[ ] Least privilege implemented[ ] Public exposure reviewed[ ] Encryption enabled[ ] Logging enabled[ ] Security monitoring enabled[ ] Vulnerability scanning enabled[ ] Patch management defined[ ] Backup protected[ ] Security alerts reviewed45 β Review Encryption
Section titled β45 β Review EncryptionβAssess:
Data at Rest
Data in Transit
Database Encryption
Storage Encryption
Backup Encryption
Key Management46 β Review Vulnerability Management
Section titled β46 β Review Vulnerability ManagementβCheck whether:
VMs
Containers
Images
Dependencies
Applicationsare appropriately scanned.
47 β Review Patching
Section titled β47 β Review PatchingβDetermine:
Patch Policy
Patch Frequency
Current Patch Status
Exceptions
Unsupported Systems48 β Review Security Findings
Section titled β48 β Review Security FindingsβInspect available security-management services for:
Critical Findings
High Findings
Misconfigurations
Threat Detections
Exposure49 β Assess Availability
Section titled β49 β Assess AvailabilityβAsk:
What happens if a VM fails?
What happens if an availability zone fails?
What happens if the database fails?
What happens if a network dependency fails?50 β Review High Availability
Section titled β50 β Review High AvailabilityβCheck:
Multiple Instances
Multiple Zones
Load Balancing
Database Redundancy
Storage Redundancy
Automatic Failover51 β Identify Availability Risk
Section titled β51 β Identify Availability RiskβExample:
Internet βLoad Balancer βSingle VMFinding:
Single Point of Failure52 β Review Autoscaling
Section titled β52 β Review AutoscalingβAssess:
Minimum Capacity
Maximum Capacity
Scaling Trigger
Cooldown
Health
Quota53 β Test Scaling Assumptions
Section titled β53 β Test Scaling AssumptionsβAsk:
Can the environment actually scale when required?
A configured scaling policy may still fail because of:
Quota
Permissions
Capacity
Configuration
Network Limits54 β Assess Monitoring Health
Section titled β54 β Assess Monitoring HealthβMonitoring should answer:
Is It Available?
Is It Performing Normally?
Is Capacity Healthy?
Are Errors Increasing?
Are Dependencies Healthy?55 β Review Monitoring Coverage
Section titled β55 β Review Monitoring CoverageβInventory:
| Resource | Metrics | Logs | Alerts | Dashboard |
|---|---|---|---|---|
| VM | ||||
| Database | ||||
| Load Balancer | ||||
| Storage |
56 β Review Alerting
Section titled β56 β Review AlertingβCheck:
CPU
Memory
Disk
Availability
Latency
Errors
Failed Backups
Certificate Expiry
Security Events57 β Review Alert Quality
Section titled β57 β Review Alert QualityβDetermine whether alerts are:
Actionable
Owned
Prioritized
Tested
Escalated58 β Identify Alert Fatigue
Section titled β58 β Identify Alert FatigueβToo many alerts:
AlertAlertAlertAlertAlert βNoise βIgnored AlertsTune alerts appropriately.
59 β Assess Logging Health
Section titled β59 β Assess Logging HealthβCheck:
Application Logs
System Logs
Security Logs
Network Logs
Audit Logs
Database Logs60 β Review Centralized Logging
Section titled β60 β Review Centralized LoggingβPrefer:
VM Logs --------\Application -----\Firewall ---------> Central Log PlatformDatabase --------/Cloud Audit -----/rather than isolated logs across resources.
61 β Review Log Retention
Section titled β61 β Review Log RetentionβDetermine:
Retention Period
Security Requirements
Compliance Requirements
Cost
Archive Requirements62 β Review Audit Logging
Section titled β62 β Review Audit LoggingβEnsure important administrative activities are recorded.
Examples:
Resource Creation
Resource Deletion
IAM Changes
Network Changes
Security Changes
Policy Changes63 β Assess Backup Health
Section titled β63 β Assess Backup HealthβA backup configuration is not enough.
Review:
Backup Enabled?
Backup Successful?
Retention Correct?
Encrypted?
Protected?
Restore Tested?64 β Review Backup Status
Section titled β64 β Review Backup Statusβ| Resource | Backup | Last Backup | Retention | Restore Tested |
|---|---|---|---|---|
| Database | ||||
| VM | ||||
| Storage |
65 β Look for Backup Failures
Section titled β65 β Look for Backup FailuresβExamples:
Backup Job Failed
Backup Agent Offline
Storage Full
Permission Failure
Retention Misconfigured66 β Understand the Restore Principle
Section titled β66 β Understand the Restore PrincipleβRemember:
A backup that has never been successfully restored is an unverified recovery capability.
67 β Perform a Safe Restore Test
Section titled β67 β Perform a Safe Restore TestβIn an approved non-production location:
Backup βRestore βValidate Data βValidate ApplicationDo not overwrite production data during a lab restore test.
68 β Review RPO
Section titled β68 β Review RPOβRPO asks:
How much data can the organization tolerate losing?
Example:
RPO:1 HourThe backup/recovery design should support that requirement.
69 β Review RTO
Section titled β69 β Review RTOβRTO asks:
How long can the service remain unavailable?
Example:
RTO:4 Hours70 β Compare Recovery Requirements
Section titled β70 β Compare Recovery RequirementsβBusiness Requirement βRPO + RTO βBackup / Replication / DR Design71 β Assess Disaster Recovery
Section titled β71 β Assess Disaster RecoveryβReview:
Recovery Region
Data Replication
Infrastructure Recovery
Application Recovery
DNS Failover
Identity Dependencies
Runbooks
Testing72 β Review DR Documentation
Section titled β72 β Review DR DocumentationβAsk:
Who Declares Disaster?
Who Performs Recovery?
Where Is Infrastructure Recreated?
How Is Data Restored?
How Is Traffic Redirected?
How Is Recovery Validated?73 β Review DR Testing
Section titled β73 β Review DR TestingβA documented plan without testing may contain unknown assumptions.
Record:
Last DR Test:
Result:
RTO Achieved:
RPO Achieved:
Failures:
Actions:74 β Assess Automation
Section titled β74 β Assess AutomationβReview:
Provisioning
Patching
Backup
Scaling
Monitoring
Deployment
RemediationIdentify unnecessarily manual processes.
75 β Review Infrastructure as Code
Section titled β75 β Review Infrastructure as CodeβDetermine:
What Is Managed by IaC?
What Is Manual?
Where Is State Stored?
Who Can Change Infrastructure?
Is Code Reviewed?76 β Assess Configuration Drift
Section titled β76 β Assess Configuration DriftβCompare:
Desired State βActual StateExample:
IaC:Port 443
Actual:Port 443 + Port 22 PublicFinding:
Configuration drift.
77 β Review Automation Safety
Section titled β77 β Review Automation SafetyβAutomation should include:
Validation
Logging
Error Handling
Permissions
Rollback
Approval Where Required78 β Assess CI/CD Health
Section titled β78 β Assess CI/CD HealthβReview:
Source Control
Build
Testing
Security Scanning
Artifacts
Deployment
Approval
Rollback
Monitoring79 β Review Pipeline Identity
Section titled β79 β Review Pipeline IdentityβAsk:
Does the pipeline use administrator access?If yes:
High-Risk FindingApply least privilege.
80 β Review Secrets Management
Section titled β80 β Review Secrets ManagementβLook for secrets in:
Source Code
Configuration Files
Scripts
Pipeline Files
Logs
Images81 β Assess Certificate Health
Section titled β81 β Assess Certificate HealthβInventory:
TLS Certificates
Application Certificates
Client Certificates
Signing CertificatesRecord:
Owner
Expiration
Renewal Method
Monitoring82 β Identify Certificate Risk
Section titled β82 β Identify Certificate RiskβExample:
Certificate Expiration:14 Days
Alert:NONEFinding:
Potential service outage.
83 β Review Secret Expiration
Section titled β83 β Review Secret ExpirationβCheck:
Client Secrets
API Tokens
Keys
Passwords
CertificatesIdentify credentials approaching expiration.
84 β Assess Cloud Cost Health
Section titled β84 β Assess Cloud Cost HealthβOperational health also includes financial efficiency.
Review:
Compute
Storage
Data Transfer
Snapshots
Public IPs
Databases
Licensing
Managed Services85 β Identify Idle Compute
Section titled β85 β Identify Idle ComputeβExample:
VM:8 vCPU
Average CPU:2%
Business Use:UnknownInvestigate whether the resource is:
Oversized
Unused
Required for Standby86 β Identify Orphaned Resources
Section titled β86 β Identify Orphaned ResourcesβLook for:
Unused Disks
Unused Public IPs
Old Snapshots
Old Load Balancers
Unused Network Interfaces
Abandoned Development Resources87 β Do Not Delete Based on Cost Alone
Section titled β87 β Do Not Delete Based on Cost AloneβBefore removing:
Confirm Owner
Confirm Purpose
Confirm Dependency
Confirm Backup
Confirm Change Approval88 β Review Storage Cost
Section titled β88 β Review Storage CostβInvestigate:
Old Snapshots
Old Backups
Unused Object Versions
Incorrect Storage Tier
Excessive Log Retention89 β Review Data Transfer
Section titled β89 β Review Data TransferβUnexpected transfer costs may indicate:
Architecture Inefficiency
Cross-Region Traffic
Internet Egress
Unexpected Workload Behavior90 β Review Reserved Capacity and Commitments
Section titled β90 β Review Reserved Capacity and CommitmentsβWhere applicable, evaluate whether predictable workloads could benefit from appropriate pricing commitments.
Do not optimize cost at the expense of required flexibility.
91 β Review Quotas
Section titled β91 β Review QuotasβAssess:
Compute Quota
Storage Quota
IP Limits
API Limits
Database Limits92 β Identify Quota Risk
Section titled β92 β Identify Quota RiskβExample:
vCPU Quota:100
Current:96Risk:
Autoscaling / Deployment Failure93 β Assess Operational Documentation
Section titled β93 β Assess Operational DocumentationβCheck whether teams have:
Architecture Diagrams
Runbooks
Recovery Procedures
Deployment Procedures
Escalation Procedures
Ownership Information94 β Assess Incident Readiness
Section titled β94 β Assess Incident ReadinessβAsk:
Who receives alerts?
Who is on call?
How are incidents declared?
Where are logs?
How is escalation performed?
Who communicates business impact?95 β Review Runbooks
Section titled β95 β Review RunbooksβImportant services should have runbooks for:
Application Failure
VM Failure
Database Failure
Network Failure
Backup Failure
Security Incident
Capacity Issue
Certificate Expiry96 β Review Change Management
Section titled β96 β Review Change ManagementβAssess:
How Changes Are Requested
How Changes Are Reviewed
How Changes Are Approved
How Changes Are Deployed
How Changes Are Rolled Back97 β Review Manual Changes
Section titled β97 β Review Manual ChangesβUncontrolled console changes can cause:
Configuration Drift
Security Exposure
Audit Gaps
Deployment Inconsistency98 β Assess Resource Lifecycle
Section titled β98 β Assess Resource LifecycleβDetermine how resources are:
Created
Owned
Monitored
Updated
Backed Up
Decommissioned99 β Identify Abandoned Resources
Section titled β99 β Identify Abandoned ResourcesβPossible indicators:
No Owner
No Tags
No Traffic
No Recent Use
Old Creation Date
No MonitoringInvestigate before removal.
100 β Assess Governance
Section titled β100 β Assess GovernanceβReview whether standards exist for:
Naming
Tagging
Regions
Identity
Networking
Encryption
Logging
Backup
Resource Creation101 β Review Policy Enforcement
Section titled β101 β Review Policy EnforcementβWhere supported, cloud policies can help prevent:
Public Storage
Unapproved Regions
Missing Encryption
Missing Tags
Prohibited Resource Types102 β Assess Security Governance
Section titled β102 β Assess Security GovernanceβDetermine whether the organization has defined requirements for:
Privileged Access
MFA
Public Exposure
Encryption
Secrets
Vulnerability Management
Logging
Incident Response103 β Build the Findings Register
Section titled β103 β Build the Findings RegisterβCreate:
| ID | Domain | Finding | Risk | Priority | Recommendation |
|---|---|---|---|---|---|
| F-001 | Compute | ||||
| F-002 | Network | ||||
| F-003 | Identity |
104 β Classify Findings
Section titled β104 β Classify FindingsβUse:
Critical
High
Medium
Low
Informational105 β Understand Critical Findings
Section titled β105 β Understand Critical FindingsβExamples may include:
Exposed Administrative Interface
Compromised Credential
Public Sensitive Data
No Recovery for Critical System
Critical Security VulnerabilityThese may require immediate action.
106 β Understand High Findings
Section titled β106 β Understand High FindingsβExamples:
Excessive Privileges
Missing MFA for Privileged Access
No Tested Backup
Single Point of Failure
Missing Critical Monitoring107 β Understand Medium Findings
Section titled β107 β Understand Medium FindingsβExamples:
Missing Tags
Over-Provisioned VM
Incomplete Documentation
Old Snapshots
Alert Tuning Required108 β Prioritize by Risk
Section titled β108 β Prioritize by RiskβDo not simply fix:
Easiest Finding FirstPrioritize using:
Likelihood +Impact +Business Criticality109 β Build the Risk Matrix
Section titled β109 β Build the Risk Matrixβ| Likelihood | Impact | Priority |
|---|---|---|
| High | High | Critical |
| Medium | High | High |
| High | Medium | High |
| Medium | Medium | Medium |
| Low | Low | Low |
110 β Consider Business Context
Section titled β110 β Consider Business ContextβExample:
Single VMmay be:
LOW RISKfor a disposable development system.
But:
CRITICAL RISKfor a revenue-generating production application.
111 β Create Remediation Recommendations
Section titled β111 β Create Remediation RecommendationsβRecommendations should be:
Specific
Actionable
Prioritized
Owned
MeasurablePoor:
Improve security.
Better:
Remove direct internet SSH access from production application servers and require approved administrative connectivity.
112 β Build the Remediation Plan
Section titled β112 β Build the Remediation Planβ| Finding | Action | Owner | Priority | Target Date | Status |
|---|---|---|---|---|---|
113 β Separate Quick Wins
Section titled β113 β Separate Quick WinsβSome findings may be:
High Value+Low ComplexityExamples:
Enable Missing Alert
Remove Unused Public IP
Add Resource Owner Tag
Configure Certificate Expiry Alert114 β Identify Strategic Improvements
Section titled β114 β Identify Strategic ImprovementsβOther findings require longer projects:
Multi-Region DR
Identity Redesign
Network Segmentation
Centralized Logging
IaC Migration115 β Create a Cloud Health Scorecard
Section titled β115 β Create a Cloud Health ScorecardβExample:
| Domain | Score | Status |
|---|---|---|
| Compute | 8/10 | Good |
| Networking | 6/10 | Needs Improvement |
| Storage | 7/10 | Good |
| Identity | 5/10 | Needs Improvement |
| Security | 6/10 | Needs Improvement |
| Availability | 7/10 | Good |
| Monitoring | 5/10 | Needs Improvement |
| Backup | 4/10 | High Risk |
| Automation | 7/10 | Good |
| Cost | 6/10 | Needs Improvement |
116 β Do Not Treat Scores as Absolute Truth
Section titled β116 β Do Not Treat Scores as Absolute TruthβThe scorecard should:
Summarize Findingsnot replace:
Detailed Technical Analysis117 β Build Overall Health Status
Section titled β117 β Build Overall Health StatusβYou may use:
HEALTHY
NEEDS IMPROVEMENT
HIGH RISK
CRITICALExample:
Overall Status:NEEDS IMPROVEMENT118 β Create an Executive Summary
Section titled β118 β Create an Executive SummaryβExecutives generally need:
Overall Health
Business Risk
Critical Findings
High Findings
Major Recommendations
Remediation Prioritiesnot hundreds of raw technical observations.
119 β Example Executive Summary
Section titled β119 β Example Executive SummaryβOverall Environment Health:Needs Improvement
Critical Findings:0
High Findings:4
Primary Risks:- Untested database recovery- Excessive privileged access- Single application availability zone- Incomplete production monitoring
Immediate Priorities:1. Validate backup restoration2. Reduce privileged access3. Improve workload redundancy4. Expand monitoring coverage120 β Create the Technical Assessment Report
Section titled β120 β Create the Technical Assessment ReportβUse:
Cloud Environment Health Assessment
Executive Summary:
Scope:
Architecture:
Resource Inventory:
Compute Assessment:
Network Assessment:
Storage Assessment:
Identity Assessment:
Security Assessment:
Availability Assessment:
Monitoring Assessment:
Logging Assessment:
Backup Assessment:
Disaster Recovery Assessment:
Automation Assessment:
CI/CD Assessment:
Cost Assessment:
Governance Assessment:
Operational Readiness:
Critical Findings:
High Findings:
Medium Findings:
Low Findings:
Recommendations:
Remediation Roadmap:
Overall Health:121 β Build the Health Assessment Runbook
Section titled β121 β Build the Health Assessment RunbookβRunbook:Cloud Environment Health Assessment
Assessment Owner:
Environment:
Cloud Platform:
Business Services:
Criticality:
Step 1:Confirm scope
Step 2:Inventory resources
Step 3:Document architecture
Step 4:Map dependencies
Step 5:Assess compute
Step 6:Assess networking
Step 7:Assess storage
Step 8:Assess identity
Step 9:Assess security
Step 10:Assess availability
Step 11:Assess monitoring and logging
Step 12:Assess backup and DR
Step 13:Assess automation and IaC
Step 14:Assess CI/CD
Step 15:Assess cost
Step 16:Assess governance
Step 17:Document findings
Step 18:Prioritize risk
Step 19:Build remediation plan
Step 20:Create executive report122 β Perform Final Integrated Assessment
Section titled β122 β Perform Final Integrated AssessmentβAssume the following environment:
Internet | v DNS | v Load Balancer | +------+------+ | | v v App-01 App-02 | | +------+------+ | v Database | v Storage123 β Initial Observations
Section titled β123 β Initial ObservationsβYou discover:
App-01 CPU:42%
App-02 CPU:45%
Database:Healthy
Storage:92% Full
Backup:Enabled
Last Restore Test:Never
SSH:Publicly Accessible
Admin MFA:Enabled
Application Monitoring:Enabled
Database Monitoring:Enabled
Storage Alert:Missing
Certificate:Expires in 21 Days
Autoscaling:Not Configured124 β Identify Findings
Section titled β124 β Identify FindingsβPossible findings:
F-001Storage capacity approaching threshold
F-002Backup restoration has never been validated
F-003Administrative SSH publicly exposed
F-004Storage capacity alert missing
F-005Certificate approaching expiration
F-006Application lacks autoscaling125 β Prioritize
Section titled β125 β PrioritizeβExample:
| Finding | Priority |
|---|---|
| Public Administrative Access | High |
| Untested Recovery | High |
| Storage 92% Full | High |
| Certificate Expiry | High |
| Missing Storage Alert | Medium |
| No Autoscaling | Medium |
Actual priority should reflect business context and existing compensating controls.
126 β Recommend Remediation
Section titled β126 β Recommend RemediationβPublic SSH βRestrict Administrative Access
Untested Backup βPerform Restore Test
Storage 92% βIncrease / Optimize Capacity
Certificate βRenew + Configure Expiry Monitoring
Missing Alert βConfigure Capacity Alert
No Autoscaling βEvaluate Workload Scaling Requirements127 β Validate Remediation
Section titled β127 β Validate RemediationβAfter remediation:
[ ] Public administrative exposure removed[ ] Restore successfully tested[ ] Storage capacity healthy[ ] Capacity alert working[ ] Certificate renewed[ ] Certificate monitoring enabled[ ] Scaling requirement reviewed128 β Understand Continuous Health Assessment
Section titled β128 β Understand Continuous Health AssessmentβA health assessment should not necessarily happen only once.
Cloud environments continuously change:
Deployments
New Resources
New Users
New Permissions
Scaling
Certificates
Costs
Security FindingsTherefore:
Assess βRemediate βMonitor βReassess129 β Define Assessment Frequency
Section titled β129 β Define Assessment FrequencyβDepending on organizational requirements:
Continuous Monitoring
Weekly Operational Review
Monthly Health Review
Quarterly Architecture Review
Annual DR ExerciseDifferent controls may require different frequencies.
130 β Build Health Indicators
Section titled β130 β Build Health IndicatorsβExamples:
Availability
Error Rate
CPU
Memory
Storage Capacity
Backup Success
Restore Success
Security Findings
Certificate Expiration
Cost Variance131 β Build the Cloud Health Dashboard
Section titled β131 β Build the Cloud Health DashboardβConceptually:
+----------------------------------------+| CLOUD HEALTH DASHBOARD |+----------------------------------------+| Availability HEALTHY || Compute HEALTHY || Network HEALTHY || Storage WARNING || Identity HEALTHY || Security WARNING || Monitoring WARNING || Backup HIGH RISK || Cost HEALTHY |+----------------------------------------+132 β Create the Final Lab Report
Section titled β132 β Create the Final Lab ReportβUse:
Lab:Cloud Environment Health Assessment Lab
Cloud Platform:
Environment:
Business Service:
Assessment Date:
Assessment Scope:
Architecture:
Resource Inventory:
Compute Health:
Network Health:
Storage Health:
Identity Health:
Security Health:
Availability:
Scalability:
Monitoring:
Logging:
Backup:
Restore Validation:
RPO:
RTO:
Disaster Recovery:
Automation:
Infrastructure as Code:
CI/CD:
Configuration Drift:
Certificates:
Secrets:
Cloud Costs:
Unused Resources:
Quotas:
Operational Documentation:
Incident Readiness:
Governance:
Critical Findings:
High Findings:
Medium Findings:
Low Findings:
Quick Wins:
Strategic Improvements:
Overall Health:
Remediation Plan:
Lessons Learned:π§ͺ Final Validation Checklist
Section titled βπ§ͺ Final Validation Checklistβ| Validation | Status |
|---|---|
| Assessment scope defined | |
| Business criticality documented | |
| Resource inventory created | |
| Resource ownership reviewed | |
| Architecture documented | |
| Dependencies mapped | |
| Single points of failure identified | |
| Compute health assessed | |
| CPU reviewed | |
| Memory reviewed | |
| Compute sizing reviewed | |
| Network health assessed | |
| Public exposure reviewed | |
| Administrative access reviewed | |
| Firewall rules reviewed | |
| Segmentation reviewed | |
| DNS reviewed | |
| Load balancers reviewed | |
| Storage capacity reviewed | |
| Storage performance reviewed | |
| Storage security reviewed | |
| Identity inventory completed | |
| Privileged access reviewed | |
| Least privilege reviewed | |
| Dormant identities reviewed | |
| MFA reviewed | |
| Service identities reviewed | |
| Security baseline reviewed | |
| Encryption reviewed | |
| Vulnerability management reviewed | |
| Patching reviewed | |
| Availability reviewed | |
| Autoscaling reviewed | |
| Monitoring coverage reviewed | |
| Alerting reviewed | |
| Logging reviewed | |
| Audit logging reviewed | |
| Backup health reviewed | |
| Restore tested | |
| RPO reviewed | |
| RTO reviewed | |
| DR readiness reviewed | |
| Automation reviewed | |
| IaC reviewed | |
| Configuration drift reviewed | |
| CI/CD reviewed | |
| Secrets reviewed | |
| Certificates reviewed | |
| Cost health reviewed | |
| Idle resources reviewed | |
| Quotas reviewed | |
| Documentation reviewed | |
| Incident readiness reviewed | |
| Governance reviewed | |
| Findings classified | |
| Remediation prioritized | |
| Health scorecard created | |
| Executive summary created | |
| Assessment runbook completed |
π― Certification Connection
Section titled βπ― Certification ConnectionβA Cloud+ scenario may say:
All production VMs are running, but storage utilization has reached 95%.
Think:
The environment may be available but has a capacity-health risk.
Another:
Backups complete successfully every night, but no restore has ever been performed.
Think:
Recovery capability has not been validated.
Another:
Production administrative ports are accessible from anywhere on the internet.
Think:
Reduce public exposure and restrict administrative access.
Another:
A workload uses 3% CPU but runs on a very large instance.
Think:
Review rightsizing and cost optimization.
Another:
Autoscaling is configured, but the account is close to its compute quota.
Think:
Capacity and quota risk could prevent scaling.
Another:
A critical TLS certificate expires next week.
Think:
Certificate lifecycle and operational availability risk.
Another:
Infrastructure differs from the approved IaC configuration.
Think:
Configuration drift.
π€ Interview Questions
Section titled βπ€ Interview QuestionsβPractice without notes.
1. What makes a cloud environment healthy?
Section titled β1. What makes a cloud environment healthy?β2. How would you perform a cloud health assessment?
Section titled β2. How would you perform a cloud health assessment?β3. Why create a resource inventory?
Section titled β3. Why create a resource inventory?β4. Why map application dependencies?
Section titled β4. Why map application dependencies?β5. What is a single point of failure?
Section titled β5. What is a single point of failure?β6. How would you assess compute health?
Section titled β6. How would you assess compute health?β7. How would you assess network health?
Section titled β7. How would you assess network health?β8. How would you assess storage health?
Section titled β8. How would you assess storage health?β9. How would you review IAM health?
Section titled β9. How would you review IAM health?β10. Why review privileged access?
Section titled β10. Why review privileged access?β11. How would you assess cloud security?
Section titled β11. How would you assess cloud security?β12. How would you evaluate high availability?
Section titled β12. How would you evaluate high availability?β13. Why review autoscaling and quotas together?
Section titled β13. Why review autoscaling and quotas together?β14. How would you assess monitoring?
Section titled β14. How would you assess monitoring?β15. Why centralize logging?
Section titled β15. Why centralize logging?β16. Why is successful backup completion not enough?
Section titled β16. Why is successful backup completion not enough?β17. What are RPO and RTO?
Section titled β17. What are RPO and RTO?β18. How would you assess disaster recovery?
Section titled β18. How would you assess disaster recovery?β19. What is configuration drift?
Section titled β19. What is configuration drift?β20. How would you assess CI/CD health?
Section titled β20. How would you assess CI/CD health?β21. How would you identify cloud cost waste?
Section titled β21. How would you identify cloud cost waste?β22. Why review certificates?
Section titled β22. Why review certificates?β23. How do you prioritize findings?
Section titled β23. How do you prioritize findings?β24. What belongs in an executive cloud-health report?
Section titled β24. What belongs in an executive cloud-health report?β25. How frequently should cloud health be reviewed?
Section titled β25. How frequently should cloud health be reviewed?βπ¨ Scenario Interview Question 1
Section titled βπ¨ Scenario Interview Question 1βManagement asks whether production is healthy because every VM is running.
A strong answer:
VM state is only one health indicator.
Review:
Availability
Performance
Capacity
Security
Monitoring
Backup
Recovery
Dependencies
Cost
Operationsπ¨ Scenario Interview Question 2
Section titled βπ¨ Scenario Interview Question 2βProduction storage is 92% utilized but the application is still functioning normally.
Think:
Proactive capacity finding.
Do not wait for:
100% βApplication Failureπ¨ Scenario Interview Question 3
Section titled βπ¨ Scenario Interview Question 3βNightly database backups show successful status.
Is recovery validated?
Not until restore capability has been appropriately tested.
π¨ Scenario Interview Question 4
Section titled βπ¨ Scenario Interview Question 4βA production application has two healthy VMs, but both are in the same availability zone.
Think:
Zone-level availability risk.
π¨ Scenario Interview Question 5
Section titled βπ¨ Scenario Interview Question 5βThe organization has monitoring dashboards but nobody receives alerts.
Think:
Visibilityβ Operational ResponseReview alert ownership and escalation.
π¨ Scenario Interview Question 6
Section titled βπ¨ Scenario Interview Question 6βA CI/CD pipeline has full administrator access to the cloud environment.
Think:
Excessive privilege.
π¨ Scenario Interview Question 7
Section titled βπ¨ Scenario Interview Question 7βAn unused disk has been generating charges for six months.
Before deletion:
Identify Owner βConfirm Dependency βConfirm Data Requirement βFollow Change Process βRemove if Approvedπ¨ Scenario Interview Question 8
Section titled βπ¨ Scenario Interview Question 8βA certificate expires in ten days and no engineer owns the renewal process.
Think:
Operational ownership and certificate lifecycle risk.
π¨ Scenario Interview Question 9
Section titled βπ¨ Scenario Interview Question 9βIaC defines one firewall configuration, but production contains additional manually created rules.
Think:
Configuration drift.
π¨ Scenario Interview Question 10
Section titled βπ¨ Scenario Interview Question 10βYou find 40 assessment issues. Which do you fix first?
Prioritize based on:
Likelihood+Impact+Business Criticality+Existing Controlsnot simply the order discovered.
π§ Cloud Health Assessment Framework
Section titled βπ§ Cloud Health Assessment FrameworkβRemember:
SCOPE βINVENTORY βARCHITECTURE βDEPENDENCIES βCOMPUTE βNETWORK βSTORAGE βIDENTITY βSECURITY βAVAILABILITY βMONITORING βBACKUP & DR βAUTOMATION βCOST βGOVERNANCE βFINDINGS βRISK βREMEDIATION βREASSESSπ¬ Interview Tip
Section titled βπ¬ Interview TipβAvoid:
βI would check whether all the cloud resources are running.β
A stronger answer is:
βI would begin by defining the assessment scope, business criticality, architecture, resource inventory, and dependencies. I would then assess compute capacity and health, network exposure and connectivity, storage capacity and performance, identity and privileged access, security controls, availability, monitoring and logging, backup and recovery, automation, configuration drift, CI/CD, quotas, certificates, cost efficiency, and operational readiness. I would document findings, prioritize them based on likelihood, impact, and business criticality, create a remediation roadmap, and establish continuous monitoring and reassessment.β
That demonstrates Cloud Engineer and Cloud Operations thinking.
π Portfolio Deliverables
Section titled βπ Portfolio DeliverablesβKeep sanitized versions of:
1. Cloud Health Assessment Report
Section titled β1. Cloud Health Assessment ReportβInclude:
Scope
Architecture
Assessment Domains
Findings
Risk
Recommendations2. Cloud Architecture Diagram
Section titled β2. Cloud Architecture DiagramβShow:
Users βNetwork βApplication βDataplus major dependencies.
3. Resource Inventory
Section titled β3. Resource InventoryβInclude:
Resource
Type
Environment
Owner
Criticality4. Cloud Health Scorecard
Section titled β4. Cloud Health ScorecardβCover:
Compute
Network
Storage
Identity
Security
Availability
Monitoring
Backup
Automation
Cost5. Findings Register
Section titled β5. Findings RegisterβDocument:
Finding
Evidence
Risk
Priority
Recommendation6. Remediation Roadmap
Section titled β6. Remediation RoadmapβSeparate:
Immediate
Short-Term
Medium-Term
Strategicimprovements.
7. Cloud Health Assessment Runbook
Section titled β7. Cloud Health Assessment RunbookβCreate a reusable operational assessment procedure.
π Resume Examples
Section titled βπ Resume ExamplesβInstead of:
Performed cloud health checks.
Use:
Conducted comprehensive cloud environment health assessments across compute, networking, storage, IAM, security, availability, monitoring, backup, disaster recovery, automation, configuration management, and cost optimization.
Or:
Evaluated cloud operational readiness by inventorying resources, mapping service dependencies, identifying single points of failure, reviewing monitoring and recovery controls, assessing configuration drift, and prioritizing remediation based on business risk.
Or:
Developed cloud health scorecards, technical findings registers, executive assessment reports, and prioritized remediation roadmaps for simulated production cloud environments.
β Job-Readiness Check
Section titled ββ Job-Readiness CheckβYou should now be able to:
-
define cloud environment health
-
scope a cloud assessment
-
inventory resources
-
map architecture
-
map dependencies
-
identify single points of failure
-
assess compute
-
assess networking
-
assess storage
-
assess identity
-
assess security
-
assess availability
-
assess scalability
-
assess monitoring
-
assess logging
-
assess backups
-
evaluate restore capability
-
understand RPO
-
understand RTO
-
assess disaster recovery
-
assess automation
-
assess IaC
-
identify configuration drift
-
assess CI/CD
-
review secrets
-
review certificates
-
assess cloud costs
-
identify idle resources
-
assess quotas
-
review documentation
-
assess incident readiness
-
review governance
-
classify findings
-
prioritize risks
-
build remediation plans
-
create cloud health scorecards
-
create executive reports
-
build reusable assessment runbooks
π Mission Complete
Section titled βπ Mission CompleteβYou have progressed from:
Is It Running? βYes βEverything Must Be Fineto:
Inventory βArchitecture βPerformance βCapacity βSecurity βAvailability βObservability βRecoverability βAutomation βOperational Readiness βCloud HealthYou now understand an important Cloud Engineer principle:
Operational health is not the absence of an outage. A healthy cloud environment is observable, secure, appropriately sized, resilient, recoverable, maintainable, and continuously assessed.
π Whatβs Next?
Section titled βπ Whatβs Next?βYou have now completed a broad operational assessment of a cloud environment.
The next step is to move from:
Assessing Current Healthto:
Preparing for Operational FailureYou will build a structured cloud incident-response workflow covering:
Detection
Triage
Severity
Containment
Evidence Collection
Service Restoration
Recovery
Validation
Root-Cause Analysis
Communication
Lessons LearnedYou will learn to manage the complete lifecycle:
ALERT βTRIAGE βCLASSIFY βINVESTIGATE βCONTAIN βRECOVER βVALIDATE βDOCUMENT βIMPROVEβ‘οΈ Next: Lab 27 β Cloud Incident Response Lab