Skip to content

Runbook 01 — Kubernetes Logging & Monitoring Health Assessment

Logging and monitoring are the foundation of Kubernetes Security Operations.

Without complete visibility, organizations cannot:

  • Detect attacks
  • Investigate incidents
  • Meet compliance requirements
  • Troubleshoot production issues
  • Perform digital forensics
  • Measure operational health

This runbook provides a structured assessment that Cloud Security Engineers, Platform Engineers and SOC Analysts can use to evaluate the health of Kubernetes logging and monitoring capabilities across Amazon EKS environments.

The assessment aligns with enterprise operational practices used during:

  • Security Health Checks
  • Production Readiness Reviews
  • Compliance Audits
  • SOC Maturity Assessments
  • Cloud Security Assessments

Evaluate the effectiveness, completeness and operational readiness of Kubernetes logging and monitoring across Amazon EKS.


2–3 Hours


  • Operational Assessment
  • Security Assessment
  • Infrastructure Review
  • Compliance Validation
  • SOC Readiness Assessment

Review the following components:

  • Kubernetes Audit Logging
  • Amazon CloudWatch Logs
  • Prometheus
  • Grafana
  • Falco
  • AWS CloudTrail
  • Amazon GuardDuty
  • AWS Security Hub
  • Alerting
  • SIEM Integration
  • Operational Dashboards
  • Monitoring Governance

A financial organisation operates:

  • 420 Amazon EKS clusters
  • 9 AWS Accounts
  • Multiple AWS Regions
  • 24×7 Security Operations Centre

Senior management requests an enterprise-wide health assessment after an external audit identifies inconsistent monitoring configurations across several production clusters.

You are assigned to assess the monitoring platform and identify operational risks before the next compliance audit.


Amazon EKS
Audit Logs
Falco
Prometheus
Grafana
CloudWatch
CloudTrail
GuardDuty
Security Hub
SIEM
SOC

Component Status Notes
Audit Logging Enabled
CloudWatch Log Groups
Falco Running
Prometheus Healthy
Grafana Accessible
Alertmanager Configured
CloudTrail Enabled
GuardDuty Enabled
Security Hub Enabled
SIEM Receiving Logs
Alerting Operational
Dashboards Reviewed

Phase 1 — Kubernetes Audit Logging Review

Section titled “Phase 1 — Kubernetes Audit Logging Review”

Confirm that Kubernetes API activity is being recorded.

Describe the cluster.

Terminal window
aws eks describe-cluster \
--name production-cluster \
--query "cluster.logging"

Verify:

  • Audit Logs enabled
  • API Logs enabled
  • Authenticator Logs enabled
  • Scheduler Logs enabled
  • Controller Manager Logs enabled

Check Pass Fail
Audit Logs Enabled
API Logs Enabled
Log Delivery Working
Log Retention Configured

Verify Log Groups.

Terminal window
aws logs describe-log-groups

Review:

  • Log retention
  • Encryption
  • Naming standards
  • Access permissions

Check Pass Fail
Log Groups Present
KMS Encryption Enabled
Retention Policy Applied
IAM Access Restricted

Verify Falco.

Terminal window
kubectl get pods -n falco

Verify DaemonSet.

Terminal window
kubectl get daemonset -n falco

Review:

  • Running Pods
  • Failed Pods
  • Runtime alerts
  • Rule configuration

Check Pass Fail
DaemonSet Running
Alerts Generated
Rules Updated
Runtime Coverage Complete

Verify monitoring components.

Terminal window
kubectl get pods -n monitoring

Review:

  • Prometheus
  • Node Exporter
  • kube-state-metrics
  • Alertmanager

Check Pass Fail
Prometheus Running
Targets Healthy
Metrics Collected
Storage Healthy

Access Grafana.

Review dashboards.

Confirm:

  • Cluster Health
  • Node Health
  • Namespace Metrics
  • API Server Metrics
  • Pod Health

Dashboard Healthy
Cluster Overview
Node Dashboard
Namespace Dashboard
Kubernetes API
Resource Utilisation

Verify CloudTrail.

Terminal window
aws cloudtrail describe-trails

Verify GuardDuty.

Terminal window
aws guardduty list-detectors

Verify Security Hub.

Terminal window
aws securityhub get-enabled-standards

Service Enabled
CloudTrail
GuardDuty
Inspector
Security Hub

Review:

  • Log forwarding
  • Event correlation
  • Detection rules
  • Alert routing

Confirm telemetry sources.

Source Received
Audit Logs
Falco
CloudTrail
GuardDuty
Inspector
Security Hub

Generate a test event.

Example:

Terminal window
kubectl create namespace monitoring-test

Confirm:

  • Audit Log generated
  • Alert received
  • SIEM correlation completed
  • Notification delivered

Validation Pass
Alert Generated
Alert Routed
Alert Investigated
Evidence Preserved

Review operational dashboards.

Evaluate:

  • CPU utilisation
  • Memory utilisation
  • Pod restarts
  • Node availability
  • Namespace health
  • Runtime alerts
  • Audit events

Questions:

  • Are dashboards useful?
  • Are dashboards current?
  • Are alerts actionable?
  • Are KPIs visible?

Evaluate the following areas.

Area Rating (1–5)
Logging Coverage
Monitoring Coverage
Runtime Detection
Dashboard Quality
Alert Accuracy
SIEM Integration
SOC Readiness
Incident Detection
Automation
Operational Maturity

Document identified risks.

Risk Severity Recommendation
Missing Audit Logs High Enable Audit Logging
Missing Runtime Detection High Deploy Falco
No SIEM Integration Medium Centralize telemetry
Weak Alerting Medium Improve detection rules
Dashboard Gaps Low Enhance Grafana dashboards

Summarize the assessment.

Category Status
Logging
Monitoring
Detection
Investigation
Alerting
Automation
Governance
Compliance

Overall Assessment:

  • ☐ Excellent
  • ☐ Good
  • ☐ Requires Improvement
  • ☐ Critical Issues Identified

Prioritize corrective actions.

Priority Action Owner Target Date
High Enable missing Audit Logs Platform Team
High Deploy Falco on all nodes Security Team
High Integrate SIEM SOC Team
Medium Improve dashboards Platform Team
Medium Tune detection rules Detection Engineering
Low Review retention policies Compliance Team

As a Cloud Security Engineer:

  • Enable Audit Logging on every production cluster.
  • Monitor Kubernetes continuously using Prometheus and Grafana.
  • Deploy Falco across every worker node.
  • Forward all security telemetry to a centralized SIEM.
  • Regularly validate alerting and detection workflows.
  • Encrypt all log storage using AWS KMS.
  • Restrict access to monitoring platforms using IAM and RBAC.
  • Review dashboards daily as part of SOC operations.
  • Test monitoring capabilities during incident response exercises.
  • Conduct quarterly health assessments to maintain operational readiness.

A multinational healthcare provider experienced delayed detection of a ransomware attack because several production EKS clusters had disabled Audit Logging and incomplete runtime monitoring.

During the post-incident review, engineers discovered:

  • Missing Kubernetes Audit Logs
  • Falco deployed on only half of the worker nodes
  • Inconsistent Prometheus metric collection
  • Security Hub findings were not forwarded to the SIEM

Following a comprehensive health assessment, the organisation:

  • Standardized logging across every cluster.
  • Enabled runtime monitoring on all nodes.
  • Centralized telemetry into the enterprise SIEM.
  • Implemented automated health checks for monitoring components.
  • Reduced Mean Time to Detect (MTTD) by more than 60%.

Congratulations!

You have completed an enterprise Kubernetes Logging & Monitoring Health Assessment.

During this runbook you evaluated:

  • Kubernetes Audit Logging
  • CloudWatch Logs
  • Falco Runtime Security
  • Prometheus Monitoring
  • Grafana Dashboards
  • CloudTrail
  • GuardDuty
  • Security Hub
  • SIEM Integration
  • Alerting
  • SOC Readiness

This assessment provides a repeatable framework for evaluating the operational maturity of enterprise Kubernetes monitoring environments.


  • Effective monitoring depends on complete telemetry collection.
  • Runtime detection, logging and metrics must work together.
  • Regular health assessments identify monitoring gaps before they become security incidents.
  • Centralized monitoring improves detection, investigation and compliance.
  • Continuous assessment is essential for maintaining a mature Kubernetes security programme.

The next runbook focuses on a broader enterprise review of logging and monitoring architecture, governance and operational maturity across multiple Kubernetes clusters.

➡️ Next Runbook: Runbook 02 — Enterprise Logging & Monitoring Assessment