Skip to content

Runbook 02 — Enterprise EKS Review

Item Details
Runbook ID AWS-EKS-RB-02
Type Enterprise Architecture & Security Review
Difficulty Expert
Estimated Time 1–2 Business Days
Platform Amazon EKS
Cloud Provider AWS
Primary Role Principal Cloud Security Engineer
Supporting Roles Cloud Architect, Kubernetes Administrator, DevOps Engineer, DevSecOps Engineer, SOC Analyst, Platform Engineering
Review Type Production Readiness Review
Frequency Quarterly or Before Production Go-Live
Classification Internal Use Only

This runbook provides a structured methodology for performing a complete enterprise review of an Amazon EKS platform.

Unlike a technical security assessment that focuses on individual controls, this review evaluates the entire Kubernetes platform from an enterprise perspective, including:

  • Architecture
  • Governance
  • Operational maturity
  • Identity
  • Networking
  • Infrastructure
  • Workloads
  • Security
  • Compliance
  • Monitoring
  • Business continuity
  • Disaster recovery
  • Operational excellence

The review determines whether the Amazon EKS platform is suitable for production deployment while aligning with organisational security standards and AWS best practices.


CloudNova Technologies has completed the migration of several enterprise applications to Amazon EKS.

Before onboarding customer-facing production workloads, the Executive Technology Review Board has requested a formal Enterprise EKS Review.

The review must determine whether the platform is:

  • Secure
  • Scalable
  • Highly available
  • Governed
  • Recoverable
  • Observable
  • Operationally mature

The outcome will determine whether production deployment is approved.


Validate the enterprise readiness of:

  • Platform Architecture
  • AWS Account Design
  • Networking
  • Identity
  • Kubernetes Configuration
  • Cluster Operations
  • Worker Nodes
  • Application Security
  • Secrets Management
  • Observability
  • Incident Response
  • Compliance
  • Governance
  • Business Continuity
  • Disaster Recovery

AWS Organization
AWS Landing Zone
Production AWS Account
Amazon VPC
┌────────────────────────────────────┐
│ Amazon EKS Platform │
│ │
│ Control Plane │
│ Worker Nodes │
│ Kubernetes │
│ Networking │
│ Storage │
│ Security │
│ Monitoring │
│ Governance │
└────────────────────────────────────┘
│ │
▼ ▼
AWS Native Services Enterprise SOC
│ │
▼ ▼
Compliance Executive Reporting

Preparation
Architecture Review
Infrastructure Review
Identity Review
Networking Review
Workload Review
Operations Review
Governance Review
Compliance Review
Risk Assessment
Executive Decision

Review:

  • Multi-AZ design
  • VPC architecture
  • High availability
  • Scalability
  • Fault tolerance
  • Landing Zone alignment

Review:

  • AWS Accounts
  • IAM
  • Organizations
  • SCPs
  • Route53
  • Load Balancers
  • Auto Scaling
  • VPC Endpoints
  • Transit Gateway

Review:

  • Kubernetes version
  • Control Plane
  • Managed Node Groups
  • Cluster Autoscaler
  • Add-ons
  • CNI
  • CSI Drivers
  • Upgrade process

Validate:

  • IAM
  • IRSA
  • RBAC
  • Service Accounts
  • Federation
  • MFA
  • Least Privilege

Review:

  • VPC
  • Subnets
  • NAT Gateway
  • Internet Gateway
  • Route Tables
  • Security Groups
  • Network ACLs
  • NetworkPolicies
  • DNS
  • Ingress
  • Egress

Review:

  • Security Contexts
  • Pod Security Admission
  • Resource Limits
  • Readiness Probes
  • Liveness Probes
  • Health Checks
  • Deployments
  • StatefulSets
  • DaemonSets
  • CronJobs

Review:

  • Approved Registries
  • Image Signing
  • Image Scanning
  • Immutable Tags
  • Base Images
  • Supply Chain

Review:

  • AWS Secrets Manager
  • KMS
  • IRSA
  • External Secrets Operator
  • Secret Rotation
  • Secret Ownership

Validate:

  • CloudWatch
  • Container Insights
  • Prometheus
  • Grafana
  • Fluent Bit
  • CloudTrail
  • GuardDuty
  • Security Hub

Review:

  • Logging
  • Alerting
  • Runbooks
  • Evidence Collection
  • Isolation Procedures
  • Recovery Procedures

Review:

  • Velero
  • EBS Snapshots
  • Persistent Volumes
  • etcd Protection
  • Restore Procedures
  • Cross-Region Recovery
  • RTO
  • RPO

Review:

  • Naming Standards
  • Labels
  • Tags
  • Ownership
  • Policies
  • Change Control
  • Documentation
  • Operational Standards

Validate alignment with:

  • AWS Well-Architected Framework
  • CIS Kubernetes Benchmark
  • CIS AWS Foundations Benchmark
  • NIST SP 800-53
  • ISO 27001
  • PCI DSS (where applicable)
  • SOC 2
  • Internal Security Standards

Validate:

  • Multi-AZ deployment
  • Fault tolerance
  • High availability
  • Scalability
  • Regional design
  • Future expansion

Assess:

  • VPC
  • IAM
  • Security Groups
  • KMS
  • DNS
  • Load Balancers
  • Storage

Review:

  • Cluster health
  • Node health
  • Add-ons
  • Networking
  • Storage
  • Scheduling
  • Resource management

Assess:

  • IAM
  • IRSA
  • RBAC
  • Secrets
  • Runtime Security
  • Supply Chain
  • Encryption
  • Monitoring

Evaluate:

  • Monitoring
  • Alerting
  • Upgrades
  • Incident response
  • Maintenance
  • Capacity planning

Review:

  • Documentation
  • Ownership
  • Standards
  • Processes
  • Reviews
  • Audits

Domain Status
Architecture Reviewed
AWS Infrastructure Reviewed
EKS Configuration Reviewed
Networking Reviewed
IAM Reviewed
IRSA Validated
RBAC Reviewed
Worker Nodes Reviewed
Workloads Reviewed
Secrets Reviewed
Image Security Reviewed
Logging Reviewed
Monitoring Reviewed
GuardDuty Reviewed
Security Hub Reviewed
Backup Strategy Reviewed
Disaster Recovery Reviewed
Compliance Reviewed
Governance Reviewed
Risks Documented
Executive Report Completed

Examples:

  • Single Availability Zone deployment
  • No disaster recovery strategy
  • Administrator IAM permissions
  • Public Kubernetes API without restrictions
  • No audit logging
  • No secrets management
  • Unsupported Kubernetes version

Examples:

  • Missing NetworkPolicies
  • Weak RBAC
  • No runtime monitoring
  • Missing backup validation
  • Unencrypted storage
  • Missing IRSA

Examples:

  • Incomplete documentation
  • Missing labels
  • Weak naming standards
  • Manual operational procedures

Examples:

  • Documentation improvements
  • Tagging consistency
  • Minor governance issues

Domain Score
Architecture /10
AWS Infrastructure /10
Amazon EKS Platform /10
Identity & Access /10
Networking /10
Security Controls /10
Secrets Management /10
Observability /10
Operations /10
Governance & Compliance /10

Overall Enterprise Score: ____ /100

Score Rating
95–100 Production Ready
85–94 Production Ready with Minor Improvements
70–84 Conditionally Ready
50–69 Significant Remediation Required
Below 50 Not Approved for Production

Finding ID Finding Severity Domain Owner Status
EKS-ENT-001 Open

Risk ID Risk Severity Likelihood Impact Treatment
EKS-RISK-001 Public API Endpoint Exposure Critical Medium High Immediate Remediation
EKS-RISK-002 Excessive IAM Permissions High Medium High Least Privilege Implementation
EKS-RISK-003 Missing Disaster Recovery Testing High Medium High Conduct DR Exercises
EKS-RISK-004 Weak Governance Controls Medium Medium Medium Improve Operational Processes

  • Remove critical security findings.
  • Eliminate excessive IAM permissions.
  • Restrict Kubernetes API access.
  • Enable all required audit logging.
  • Implement missing NetworkPolicies.
  • Validate backup and restore procedures.

  • Standardise cluster governance.
  • Improve monitoring coverage.
  • Complete workload hardening.
  • Enhance incident response procedures.
  • Validate disaster recovery processes.

  • Implement continuous compliance monitoring.
  • Automate security assessments.
  • Improve platform resilience.
  • Expand Zero Trust networking.
  • Mature operational governance.

  • Use managed EKS add-ons where practical.
  • Keep Kubernetes versions supported and current.
  • Apply least-privilege IAM and Kubernetes RBAC.
  • Isolate workloads with namespaces and NetworkPolicies.
  • Use IRSA for workload identity.
  • Store secrets in AWS Secrets Manager with KMS encryption.
  • Enable comprehensive logging and monitoring.
  • Test backup and disaster recovery procedures regularly.
  • Conduct quarterly enterprise platform reviews.
  • Integrate security assessments into CI/CD and operational workflows.

Why is an Enterprise EKS Review different from a standard security assessment?

Answer: It evaluates not only security controls but also architecture, operational maturity, governance, resilience, compliance, and production readiness across the entire Kubernetes platform.

Why should disaster recovery be included in an EKS review?

Answer: A secure platform must also be recoverable. Backup validation, restore testing, and defined RTO/RPO objectives are essential for business continuity.

What is the benefit of a production readiness scorecard?

Answer: It provides stakeholders with a measurable view of platform maturity, identifies priority improvements, and supports informed go-live decisions.

Why should governance be reviewed alongside technical controls?

Answer: Strong governance ensures consistent standards, ownership, change control, documentation, and long-term operational sustainability.

How often should an Enterprise EKS Review be performed?

Answer: At least quarterly, before major production releases, after significant architectural changes, and following major security incidents.


This runbook provides a comprehensive framework for evaluating Amazon EKS from an enterprise perspective.

The review extends beyond technical validation by assessing:

  • Platform architecture
  • AWS infrastructure
  • Kubernetes operations
  • Identity and access management
  • Network security
  • Workload security
  • Secrets management
  • Observability
  • Incident response
  • Backup and disaster recovery
  • Governance
  • Compliance
  • Production readiness

The final deliverable enables leadership teams to determine whether the Amazon EKS platform is operationally mature, secure, resilient, and ready to host critical production workloads while supporting long-term enterprise operations.

Next Runbook: Runbook 03 — Amazon EKS Incident Response & Security Investigation

In the next runbook, you will investigate Amazon EKS security incidents by collecting forensic evidence, analyzing Kubernetes audit logs, CloudTrail events, GuardDuty findings, workload activity, IAM changes, network traffic, and runtime alerts to determine root cause, assess impact, and coordinate containment, eradication, and recovery.