Skip to content
All posts

Chaos Engineering on Kubernetes with Chaos Mesh

When working with distributed systems, one of the biggest questions we usually ask ourselves is:

Will my application really survive when something unexpected happens?

Most of the time, we only get the answer after an issue occurs in production. Unfortunately, by then it's already too late. That's where Chaos Engineering comes in. The goal isn't to break your system for its own sake, but to understand how your applications behave when things go wrong in a controlled environment.

In this article, I'll walk through Chaos Mesh, one of the Chaos Engineering tools I've been using recently. We'll install it and go through a few simple scenarios to see how it works.

 

What is Chaos Mesh?

What I liked when I first started using it was that I didn't have to learn another platform or a new way of defining experiments. If you've written Kubernetes manifests before, creating a chaos experiment feels no different - it's just another YAML resource.

 

Why Chaos Mesh?

Several Chaos Engineering tools are available today, but Chaos Mesh stands out for a few reasons.

Dashboard

One thing I really like is the dashboard. You can:

  • Create new chaos experiments
  • Monitor running experiments
  • Review experiment results
  • Stop experiments whenever needed

If you prefer a graphical interface instead of only using kubectl, the dashboard is a great addition.

 

Supported Chaos Scenarios

Chaos Mesh supports many scenarios, including:

  • PodChaos - Kill pods or inject container failures
  • NetworkChaos - Network latency, packet loss, network partition
  • DNSChaos - DNS resolution failures
  • HTTPChaos - HTTP delays and response injection
  • StressChaos - CPU and Memory stress
  • IOChaos - Disk I/O latency and failures
  • TimeChaos - System clock manipulation
  • KernelChaos - Kernel-level fault injection

This allows you to validate not only your application but also the infrastructure it depends on.

 

RBAC Support

Not everyone should be able to launch chaos experiments, especially in production environments. Chaos Mesh integrates with Kubernetes RBAC, so only authorized users can create and run experiments.

 

Workflow & Schedule

Real incidents rarely consist of a single failure. With Chaos Mesh, you can:

  • Run experiments sequentially
  • Execute multiple experiments simultaneously
  • Schedule experiments to run automatically

This makes it much easier to simulate realistic failure scenarios.

 

Architecture

Chaos Mesh has three core components.

Chaos Dashboard

This is the web interface users interact with. Users create, start, stop, and monitor experiments here.

Chaos Controller Manager

Think of this as Chaos Mesh's brain. It watches the custom resources, schedules experiments, and coordinates all other components.

Chaos Daemon

This is the component that actually injects failures. Running as a DaemonSet on every node, it enters the target pod's namespace whenever necessary and performs operations such as modifying networking, filesystem behavior, or even kernel-level functionality.

Demo

For this demo, I'll use a local Kind Cluster, but the same steps work on Minikube or any Kubernetes cluster.

Install Chaos Mesh

curl -sSL https://mirrors.chaos-mesh.org/v2.8.1/install.sh | bash -s -- --local kind

Install Metrics Server

We'll need Metrics Server to observe CPU and Memory usage.

kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

If you're using Kind or Minikube, don't forget the following configuration.

kubectl edit deployment metric-server -n kube-system
containers:
  args:
    - --kubelet-insecure-tls

Access the Dashboard

kubectl port-forward svc/chaos-dashboard -n chaos-mesh 8080:2333

Then open: http://localhost:8080

 

Deploy the Demo Application

apiVersion: v1
kind: Namespace
metadata:
  name: demo
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: nginx
  namespace: demo
spec:
  replicas: 2
  selector:
    matchLabels:
      app: nginx
  template:
    metadata:
      labels:
        app: nginx
    spec:
      containers:
        - name: nginx
          image: nginx
          resources:
            requests:
              cpu: 100m
              memory: 128Mi
            limits:
              cpu: 200m
              memory: 256Mi
---
apiVersion: v1
kind: Service
metadata:
  name: nginx
  namespace: demo
spec:
  type: ClusterIP
  selector:
    app: nginx
  ports:
    - port: 80
      targetPort: 80
      protocol: TCP
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: nginx-hpa
  namespace: demo
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: nginx
  minReplicas: 2
  maxReplicas: 5
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 50
kubectl apply -f demo.yaml

 

Scenario 1 - Pod Kill

Let's start with the most common scenario.

apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
  name: pod-kill-demo
  namespace: demo
spec:
  action: pod-kill
  mode: all
  selector:
    namespaces:
      - demo
    labelSelectors:
      app: nginx
  duration: "60s"
kubectl apply -f kill-pod.yaml
kubectl get pod -n demo -w

Here we can watch Kubernetes' self-healing capability in action. The pod is deleted, and the Deployment immediately creates a replacement.

 

Scenario 2 - Network Delay

Now let's introduce artificial network latency.

apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: network-delay-demo
  namespace: demo
spec:
  action: delay
  mode: all
  selector:
    namespaces:
      - demo
    labelSelectors:
      app: nginx
  delay:
    latency: "3000ms"
    correlation: "100"
    jitter: "0ms"
  duration: "30s"
kubectl apply -f network-delay.yaml
kubectl run curl --image=curlimages/curl -it --rm -- sh
curl nginx.demo.svc.cluster.local

You'll notice the requests become noticeably slower.

 

Scenario 3 - CPU Stress

apiVersion: chaos-mesh.org/v1alpha1
kind: StressChaos
metadata:
  name: cpu-stress-demo
  namespace: demo
spec:
  mode: one
  selector:
    namespaces:
      - demo
    labelSelectors:
      app: nginx
  stressors:
    cpu:
      workers: 2
      load: 80
  duration: "60s"
kubectl apply -f cpu-stress.yaml
kubectl top pod -n demo

This experiment puts CPU pressure on the target pod, allowing you to observe increased resource consumption through Metrics Server.

 

Scenario 4 - DNS Failure

apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
  name: dns-port-loss
  namespace: demo
spec:
  action: loss
  mode: all
  selector:
    namespaces:
      - demo
  loss:
    loss: "100"
    correlation: "0"
  direction: to
  externalTargets:
    - "10.96.0.10"
  duration: "60s"
kubectl apply -f dns-fail.yaml
kubectl run dns-test \
  --image=busybox:1.35 \
  -n demo \
  -it --rm --restart=Never -- sh
nslookup nginx.demo.svc.cluster.local

The service name will no longer resolve successfully, allowing you to test your application's DNS dependency.

 

Scenario 5 - Scheduled Pod Kill

apiVersion: chaos-mesh.org/v1alpha1
kind: Schedule
metadata:
  name: scheduled-pod-kill
  namespace: demo
spec:
  schedule: "@every 1m"
  type: PodChaos
  podChaos:
    action: pod-kill
    mode: one
    selector:
      namespaces:
        - demo
      labelSelectors:
        app: nginx
    duration: "20s"
kubectl apply -f scheduled-pod-kill.yaml

In this example, our pod will be automatically deleted at specified intervals.

 

Scenario 6 - Workflow

Finally, let's look at one of my favorite features.

apiVersion: chaos-mesh.org/v1alpha1
kind: Workflow
metadata:
  name: cascading-failure-workflow
  namespace: demo
spec:
  entry: chaos-sequence
  templates:
    - name: chaos-sequence
      templateType: Serial
      children:
        - pod-kill-step
        - network-delay-step
        - cpu-stress-step

    - name: pod-kill-step
      templateType: PodChaos
      deadline: 30s
      podChaos:
        action: pod-kill
        mode: one
        selector:
          namespaces:
            - demo
          labelSelectors:
            app: nginx

    - name: network-delay-step
      templateType: NetworkChaos
      deadline: 40s
      networkChaos:
        action: delay
        mode: all
        selector:
          namespaces:
            - demo
          labelSelectors:
            app: nginx
        delay:
          latency: "2000ms"
          correlation: "100"
          jitter: "0ms"

    - name: cpu-stress-step
      templateType: StressChaos
      deadline: 60s
      stressChaos:
        mode: one
        selector:
          namespaces:
            - demo
          labelSelectors:
            app: nginx
        stressors:
          cpu:
            workers: 2
            load: 80
kubectl apply -f workflow.yaml

In real production environments, failures rarely happen one at a time. A network slowdown may be followed by CPU spikes, which may eventually trigger pod restarts. With Workflow, it's very easy to chain these failures together and create realistic resilience tests.

 

Final Thoughts

At first, Chaos Engineering may sound a little scary because you're intentionally breaking your own system. But that's not really the goal.

The real goal is to understand how your applications behave before real users experience those failures.

My recommendation is to start small. Start with Pod Kill, then try Network Delay and CPU Stress, and gradually combine them using Workflows.

If you're running critical workloads on Kubernetes, Chaos Mesh is definitely a tool worth exploring.