Skip to content

Lesson 01 — Kubernetes Incident Response

By the end of this lesson, you will be able to:

  • Understand Kubernetes incident response fundamentals
  • Explain the Kubernetes Incident Response lifecycle
  • Identify common Kubernetes security incidents
  • Detect indicators of compromise (IOCs)
  • Investigate security alerts
  • Contain Kubernetes attacks
  • Recover compromised workloads
  • Perform post-incident analysis
  • Build Kubernetes incident response playbooks
  • Integrate Kubernetes incident response with enterprise SOC operations

No Kubernetes cluster is completely immune to attacks.

Even with:

  • Strong IAM
  • RBAC
  • Network Policies
  • Secure Images
  • Runtime Monitoring
  • Admission Controllers

Security incidents can still occur.

Examples include:

  • Compromised Pods
  • Stolen Service Account Tokens
  • Malicious Images
  • Container Escape
  • Cryptomining
  • Data Exfiltration
  • Kubernetes API Abuse
  • Insider Threats

The goal is not only to prevent attacks but to detect and respond quickly.


Incident Response (IR) is the structured process of identifying, investigating, containing, eradicating and recovering from security incidents.

Detect
Investigate
Contain
Eradicate
Recover
Lessons Learned

Preparation
Detection
Analysis
Containment
Eradication
Recovery
Post Incident Review

This lifecycle closely aligns with industry frameworks such as the NIST Computer Security Incident Handling Guide (SP 800-61).


Examples include:

  • Pod compromise
  • Privileged container deployment
  • Container escape
  • Malware execution
  • Reverse shell
  • Cryptomining
  • Service Account compromise
  • Secret theft
  • Kubernetes API abuse
  • Unauthorized namespace creation
  • Malicious image deployment
  • Supply chain compromise
  • Node compromise
  • Denial-of-Service (DoS)
  • Data exfiltration

Amazon EKS Cluster
CloudTrail
Kubernetes Audit Logs
Falco
GuardDuty
CloudWatch
Security Hub
SIEM
SOC Analyst
Incident Response Team

Severity Example
Critical Cluster compromise
Critical Container Escape
High Privileged Pod
High Secret Theft
High Node Compromise
Medium Malware
Medium Unauthorized API Access
Low Policy Violation

Role Responsibility
SOC Analyst Detect and triage alerts
Cloud Security Engineer Investigate Kubernetes security events
Platform Engineer Restore platform services
DevOps Team Redeploy workloads
Incident Manager Coordinate response
Compliance Team Regulatory reporting
Management Business communication

Preparation occurs before an incident happens.

Activities include:

  • Enable audit logging
  • Configure CloudTrail
  • Deploy runtime monitoring
  • Maintain asset inventory
  • Create incident runbooks
  • Conduct tabletop exercises
  • Test backup restoration
  • Define escalation paths

✔ Incident Response Plan

✔ Contact List

✔ Escalation Matrix

✔ Runbooks

✔ Monitoring

✔ Backup Strategy

✔ Logging

✔ Threat Intelligence

✔ Recovery Procedures


Detection identifies suspicious activity.

Sources include:

  • Falco
  • GuardDuty
  • CloudTrail
  • Kubernetes Audit Logs
  • Security Hub
  • SIEM
  • Prometheus Alerts
  • IDS/IPS
  • Threat Intelligence

Malicious Pod
Reverse Shell
Falco Alert
SIEM
SOC Investigation

Common Kubernetes IOCs include:

  • New privileged Pods
  • Unexpected shell execution
  • High CPU usage
  • Unknown container images
  • Unauthorized Secrets access
  • Unusual API requests
  • Unexpected outbound traffic
  • Failed authentication attempts
  • Disabled logging
  • Suspicious DNS queries

Source Purpose
CloudTrail AWS API activity
Kubernetes Audit Logs Kubernetes API activity
Falco Runtime detection
GuardDuty Threat detection
Security Hub Aggregated findings
Prometheus Infrastructure metrics
CloudWatch Operational logs
SIEM Security correlation

The investigation begins after an alert.

Questions include:

  • What happened?
  • When did it occur?
  • Which cluster?
  • Which namespace?
  • Which Pod?
  • Which user?
  • Which image?
  • What data was accessed?
  • What systems are affected?

Alert
Collect Evidence
Validate Alert
Determine Scope
Assess Impact
Decide Response

List Pods:

Terminal window
kubectl get pods -A

Describe Pod:

Terminal window
kubectl describe pod <pod-name>

View Logs:

Terminal window
kubectl logs <pod-name>

View Events:

Terminal window
kubectl get events -A

The objective is to stop the attack from spreading.

Containment actions include:

  • Isolate namespace
  • Apply Network Policies
  • Scale deployment to zero
  • Remove public access
  • Block malicious image
  • Revoke credentials
  • Disable compromised Service Account
  • Remove compromised node

Compromised Pod
Block Network
Capture Evidence
Scale Deployment
Remove Pod

Short-Term Long-Term
Block network access Patch application
Stop compromised Pods Rebuild workloads
Revoke tokens Rotate credentials
Isolate namespace Improve security controls

After containment, remove the root cause.

Tasks include:

  • Remove malware
  • Delete malicious Pods
  • Remove unauthorized users
  • Patch vulnerabilities
  • Rotate Secrets
  • Replace compromised nodes
  • Remove backdoors

Recovery restores normal operations.

Recovery activities include:

  • Restore workloads
  • Recover backups
  • Redeploy applications
  • Validate integrity
  • Monitor closely
  • Confirm business functionality

Restore
Validate
Monitor
Production

Every incident should improve security.

Review:

  • Root cause
  • Timeline
  • Response effectiveness
  • Detection gaps
  • Monitoring improvements
  • Documentation updates
  • Training needs

Collect evidence before modifying compromised systems.

Evidence includes:

  • Pod logs
  • Node logs
  • Kubernetes events
  • Audit logs
  • CloudTrail logs
  • Runtime alerts
  • Network flows
  • Container image digest
  • Deployment YAML
  • Service Account configuration

When handling forensic evidence:

  • Record who collected the evidence.
  • Record when it was collected.
  • Protect evidence integrity.
  • Store evidence securely.
  • Limit access.
  • Maintain documentation.

Notify:

  • SOC
  • Platform Team
  • Application Owner
  • Management
  • Compliance Team
  • Customers (if required)

Use approved communication channels.


Alert Received
Validate Alert
Identify Cluster
Identify Namespace
Identify Pod
Collect Evidence
Contain Threat
Eradicate Threat
Recover Services
Document Incident

Detection
SOC Triage
Cloud Security Investigation
Containment
Recovery
Root Cause Analysis
Executive Report
Continuous Improvement

Record:

  • Incident ID
  • Date and Time
  • Severity
  • Affected Systems
  • Timeline
  • Root Cause
  • Impact
  • Actions Taken
  • Evidence Collected
  • Lessons Learned

  • Deleting Pods before collecting evidence
  • Restarting compromised nodes immediately
  • Ignoring audit logs
  • Not preserving logs
  • Failing to rotate credentials
  • Poor communication
  • No documentation
  • No recovery testing

As a Cloud Security Engineer:

  • Maintain updated incident response runbooks.
  • Enable Kubernetes audit logging.
  • Integrate GuardDuty, Security Hub and SIEM.
  • Deploy runtime monitoring such as Falco.
  • Collect evidence before remediation.
  • Rotate compromised credentials immediately.
  • Restore workloads from trusted images.
  • Conduct post-incident reviews.
  • Regularly test incident response procedures through tabletop exercises.
  • Continuously improve detection and response capabilities.

A global e-commerce company operates multiple Amazon EKS clusters.

Falco detects an interactive shell inside a production payment Pod.

Simultaneously:

  • GuardDuty reports suspicious runtime activity.
  • Security Hub creates a high-severity finding.
  • The SIEM correlates Kubernetes audit logs and CloudTrail events.

The Incident Response team:

  1. Confirms the alert.
  2. Identifies the affected Pod and namespace.
  3. Captures logs and runtime evidence.
  4. Applies a restrictive Network Policy to isolate the workload.
  5. Revokes the compromised Service Account permissions.
  6. Removes the malicious Pod after evidence collection.
  7. Rebuilds the application from a signed and trusted container image.
  8. Rotates all associated credentials and Secrets.
  9. Reviews audit logs to determine the initial access vector.
  10. Updates detection rules and incident runbooks.

The attack is contained without impacting other production workloads.


  • Incident response is a structured lifecycle of preparation, detection, analysis, containment, eradication, recovery and continuous improvement.
  • Kubernetes incidents require evidence collection before remediation whenever possible.
  • Multiple telemetry sources—including Kubernetes Audit Logs, CloudTrail, GuardDuty and Falco—provide comprehensive visibility.
  • Effective containment minimizes business impact while preserving forensic evidence.
  • Every incident should result in updated security controls, documentation and operational improvements.

1. What are the main phases of Kubernetes Incident Response?

Section titled “1. What are the main phases of Kubernetes Incident Response?”

Answer: Preparation, Detection, Analysis, Containment, Eradication, Recovery and Lessons Learned.

2. Why should evidence be collected before deleting a compromised Pod?

Section titled “2. Why should evidence be collected before deleting a compromised Pod?”

Answer: Deleting the Pod may destroy valuable forensic evidence needed to determine the root cause, attack timeline and attacker actions.

3. Which AWS and Kubernetes services commonly support incident detection?

Section titled “3. Which AWS and Kubernetes services commonly support incident detection?”

Answer: Kubernetes Audit Logs, AWS CloudTrail, Amazon GuardDuty, AWS Security Hub, Falco, CloudWatch and enterprise SIEM platforms.

Answer: To stop the attack from spreading while preserving evidence and minimizing business impact.

5. Why are post-incident reviews important?

Section titled “5. Why are post-incident reviews important?”

Answer: They identify root causes, improve detection capabilities, strengthen security controls and help prevent similar incidents in the future.

In the next lesson, we will explore Lesson 02 — Container Forensics, where you’ll learn how to preserve, acquire and analyze forensic evidence from compromised Kubernetes containers, including filesystem analysis, image verification, log preservation and evidence handling.