Troubleshooting

Step-by-step guides for recovering from common production issues such as non-ready pods, locked Terraform state, full PVCs, and Loki alerts.

These guides cover recovering from problems in a running environment: pods that will not become ready, a Terraform state left locked by a failed release, a PersistentVolumeClaim running out of space, unhealthy Loki ingesters, and Promtail or Loki components being throttled.

Most of them work against the environment's Kubernetes cluster, so they start by downloading its credentials. Rolling back a service to a previously registered image is here too, and rolling back a release itself is covered under Rolling back.