Kubernetes Troubleshooting¶
Nubenetes V2 Elite Portal
You are browsing the AI-Curated V2 Elite Edition. Looking for the exhaustive list of references? Check out the V1 Historical Archive.
Architectural Context
Detailed reference for Kubernetes Troubleshooting in the context of The Container Stack.
Table of Contents¶
- Architectural Foundations
- Kubernetes Tools
- Architecture
- Pod Lifecycle
- Chaos Engineering
- Curated Playbooks
- Cloud-Native Platforms
- Operating Systems
- Container Runtimes
- containerd
- Development Workflow
- IDE Extensions
- Local Development
- Infrastructure
- Compute
- Kubernetes
- Core Mechanics
- Networking
- Observability
- Resource Management
- Runtime Security
- Security
- Troubleshooting
- Troubleshooting Tooling
- Kubernetes Platform Engine
- Cluster Operations
- Observability
- Debugging
- Deployments
- Networking
- Troubleshooting Platforms
- UI Clients
Architectural Foundations¶
Kubernetes Tools¶
General Reference¶
- medium: 5 tips for troubleshooting apps on Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium: 5 tips for troubleshooting apps on Kubernetes in the Kubernetes Tools ecosystem.
- veducate.co.uk: How to fix in Kubernetes β Deleting a PVC stuck in status' βTerminatingβ [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering veducate.co.uk: How to fix in Kubernetes β Deleting a PVC stuck in status' βTerminatingβ in the Kubernetes Tools ecosystem.
- levelup.gitconnected.com: 5 tips for troubleshooting apps on Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering levelup.gitconnected.com: 5 tips for troubleshooting apps on Kubernetes in the Kubernetes Tools ecosystem.
- medium: Better Debugging Environment for your Micro-Services [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium: Better Debugging Environment for your Micro-Services in the Kubernetes Tools ecosystem.
- medium.com/@andrewachraf: Detect crashes in your Kubernetes cluster using' kwatch and Slack π [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@andrewachraf: Detect crashes in your Kubernetes cluster using' kwatch and Slack π in the Kubernetes Tools ecosystem.
- pauldally.medium.com: Kubernetes β Debugging NetworkPolicy (Part 1) [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering pauldally.medium.com: Kubernetes β Debugging NetworkPolicy (Part 1) in the Kubernetes Tools ecosystem.
- medium.com/geekculture: Common Pod Errors in Kubernetes to Watch Out For [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/geekculture: Common Pod Errors in Kubernetes to Watch Out For in the Kubernetes Tools ecosystem.
- faun.pub: Kubernetes β Debugging NetworkPolicy (Part 1) [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering faun.pub: Kubernetes β Debugging NetworkPolicy (Part 1) in the Kubernetes Tools ecosystem.
- pauldally.medium.com: Kubernetes β Debugging NetworkPolicy (Part 2) [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering pauldally.medium.com: Kubernetes β Debugging NetworkPolicy (Part 2) in the Kubernetes Tools ecosystem.
- tratnayake.dev: Oncall Adventures - When your Prometheus-Server mounted' to GCE Persistent Disk on K8s is Full [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering tratnayake.dev: Oncall Adventures - When your Prometheus-Server mounted' to GCE Persistent Disk on K8s is Full in the Kubernetes Tools ecosystem.
- blog.devgenius.io: All You Need to Know about Debugging Kubernetes Cronjob [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering blog.devgenius.io: All You Need to Know about Debugging Kubernetes Cronjob in the Kubernetes Tools ecosystem.
- saiteja313.medium.com: Tracing DNS issues in Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering saiteja313.medium.com: Tracing DNS issues in Kubernetes in the Kubernetes Tools ecosystem.
- medium.com/@jasonmfehr: Kubernetes Informers: Opening the Mystery Box [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@jasonmfehr: Kubernetes Informers: Opening the Mystery Box in the Kubernetes Tools ecosystem.
- maxilect-company.medium.com: Graceful shutdown in a cloud environment (the' example of Kubernetes + Spring Boot) π [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering maxilect-company.medium.com: Graceful shutdown in a cloud environment (the' example of Kubernetes + Spring Boot) π in the Kubernetes Tools ecosystem.
- madeeshafernando.medium.com: Capturing Heap Dumps of stateless Kubernetes' pods before container termination and export to AWS S3 [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering madeeshafernando.medium.com: Capturing Heap Dumps of stateless Kubernetes' pods before container termination and export to AWS S3 in the Kubernetes Tools ecosystem.
- faun.pub: Troubleshooting Kubernetes nodes storage space shortage on Aliyun' (Alibaba Cloud) [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering faun.pub: Troubleshooting Kubernetes nodes storage space shortage on Aliyun' (Alibaba Cloud) in the Kubernetes Tools ecosystem.
- nicolasbarlatier.hashnode.dev: .NET Core Tip 2: How to troubleshoot Memory' Leaks within a .NET Console application running in a Linux Docker Container in Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering nicolasbarlatier.hashnode.dev: .NET Core Tip 2: How to troubleshoot Memory' Leaks within a .NET Console application running in a Linux Docker Container in Kubernetes in the Kubernetes Tools ecosystem.
- dzone.com: Tackling the Top 5 Kubernetes Debugging Challenges [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering dzone.com: Tackling the Top 5 Kubernetes Debugging Challenges in the Kubernetes Tools ecosystem.
- levelup.gitconnected.com: Access Kubernetes Objects Data From /Proc Directory' π [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering levelup.gitconnected.com: Access Kubernetes Objects Data From /Proc Directory' π in the Kubernetes Tools ecosystem.
- alexsniffin.medium.com: Debugging Remotely with Go in Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering alexsniffin.medium.com: Debugging Remotely with Go in Kubernetes in the Kubernetes Tools ecosystem.
- vik-y.medium.com: An easier way to auto-remediate memory leaks on Kubernetes! [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering vik-y.medium.com: An easier way to auto-remediate memory leaks on Kubernetes! in the Kubernetes Tools ecosystem.
- medium.com/@yusufkaratoprak: Advanced Troubleshooting Techniques in Kubernetes' Pods [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@yusufkaratoprak: Advanced Troubleshooting Techniques in Kubernetes' Pods in the Kubernetes Tools ecosystem.
- Understanding Kubernetes cluster events [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering Understanding Kubernetes cluster events in the Kubernetes Tools ecosystem.
- hwchiu.medium.com: Kubernetes Network Troubleshooting Approach π [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering hwchiu.medium.com: Kubernetes Network Troubleshooting Approach π in the Kubernetes Tools ecosystem.
- blog.ediri.io: Kubernetes: ImagePullBackOff! [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering blog.ediri.io: Kubernetes: ImagePullBackOff! in the Kubernetes Tools ecosystem.
- medium.com: Kubernetes Tip: How To Disambiguate A Pod Crash To Application' Or To Kubernetes Platform? (CrashLoopBackOff) [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com: Kubernetes Tip: How To Disambiguate A Pod Crash To Application' Or To Kubernetes Platform? (CrashLoopBackOff) in the Kubernetes Tools ecosystem.
- pauldally.medium.com: Why Leaving Pods in CrashLoopBackOff Can Have a Bigger' Impact Than You Might Think [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering pauldally.medium.com: Why Leaving Pods in CrashLoopBackOff Can Have a Bigger' Impact Than You Might Think in the Kubernetes Tools ecosystem.
- tonylixu.medium.com: K8s Troubleshooting β Pod in Terminating or Unknown' Status [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering tonylixu.medium.com: K8s Troubleshooting β Pod in Terminating or Unknown' Status in the Kubernetes Tools ecosystem.
- blog.devgenius.io: K8s Troubleshooting β Pod in Terminating or Unknown Status [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering blog.devgenius.io: K8s Troubleshooting β Pod in Terminating or Unknown Status in the Kubernetes Tools ecosystem.
- medium.com/@reefland: Tracking Down βInvisibleβ OOM Kills in Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@reefland: Tracking Down βInvisibleβ OOM Kills in Kubernetes in the Kubernetes Tools ecosystem.
- baykara.medium.com: A Gentle Inspection of OOMKilled in Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering baykara.medium.com: A Gentle Inspection of OOMKilled in Kubernetes in the Kubernetes Tools ecosystem.
- medium.com/@bm54cloud: Stressing a Kubernetes Pod to Induce an OOMKilled' Error [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@bm54cloud: Stressing a Kubernetes Pod to Induce an OOMKilled' Error in the Kubernetes Tools ecosystem.
- blog.devgenius.io: K8s β pause container [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering blog.devgenius.io: K8s β pause container in the Kubernetes Tools ecosystem.
- blog.kumomind.com: What You Need To Know To Debug A Preempted Pod On Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering blog.kumomind.com: What You Need To Know To Debug A Preempted Pod On Kubernetes in the Kubernetes Tools ecosystem.
- blog.ediri.io: How to remove a stuck namespace [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering blog.ediri.io: How to remove a stuck namespace in the Kubernetes Tools ecosystem.
- medium.com/@it-craftsman: How to fix Kubernetes namespaces stuck in terminating' state [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@it-craftsman: How to fix Kubernetes namespaces stuck in terminating' state in the Kubernetes Tools ecosystem.
- medium.com/@reefland: Access PVC Data without the POD; troubleshooting Kubernetes. [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@reefland: Access PVC Data without the POD; troubleshooting Kubernetes. in the Kubernetes Tools ecosystem.
- medium.com/geekculture: K8s Troubleshooting β How to Debug CoreDNS Issues [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/geekculture: K8s Troubleshooting β How to Debug CoreDNS Issues in the Kubernetes Tools ecosystem.
- Kubernetes Troubleshooting: A Step-by-Step Guide [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering Kubernetes Troubleshooting: A Step-by-Step Guide in the Kubernetes Tools ecosystem.
- How to quarantine pods [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering How to quarantine pods in the Kubernetes Tools ecosystem.
- tetrate.io: How to debug microservices in Kubernetes with proxy, sidecar' or service mesh? [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering tetrate.io: How to debug microservices in Kubernetes with proxy, sidecar' or service mesh? in the Kubernetes Tools ecosystem.
- sumanthkumarc.medium.com: Debugging namespace deletion issue in Kubernetes [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering sumanthkumarc.medium.com: Debugging namespace deletion issue in Kubernetes in the Kubernetes Tools ecosystem.
- medium.com/linux-shots: Debug Kubernetes Pods Using Ephemeral Container [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/linux-shots: Debug Kubernetes Pods Using Ephemeral Container in the Kubernetes Tools ecosystem.
- medium.com/@blgreco72: Debugging Kubernetes Services Locally π [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@blgreco72: Debugging Kubernetes Services Locally π in the Kubernetes Tools ecosystem.
- zendesk.engineering: Debugging containerd [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering zendesk.engineering: Debugging containerd in the Kubernetes Tools ecosystem.
- heka-ai.medium.com: Introduction to Debugging: locally and live on Kubernetes' with VSCode π [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering heka-ai.medium.com: Introduction to Debugging: locally and live on Kubernetes' with VSCode π in the Kubernetes Tools ecosystem.
- eminaktas.medium.com: Debug Containerd in Production [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering eminaktas.medium.com: Debug Containerd in Production in the Kubernetes Tools ecosystem.
- medium.com/@alex.ivenin: Exploring ephemeral containers in kubernetes π [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@alex.ivenin: Exploring ephemeral containers in kubernetes π in the Kubernetes Tools ecosystem.
- medium.com/@danielepolencic: Isolating kubernetes pods for debugging [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@danielepolencic: Isolating kubernetes pods for debugging in the Kubernetes Tools ecosystem.
- medium.com/adaltas: Kubernetes: debugging with ephemeral containers [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/adaltas: Kubernetes: debugging with ephemeral containers in the Kubernetes Tools ecosystem.
- medium.com/@ospalaemon: Introducing Palaemon, the Savior of Kubernetes Pods! [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering medium.com/@ospalaemon: Introducing Palaemon, the Savior of Kubernetes Pods! in the Kubernetes Tools ecosystem.
- Debugging Kubernetes Systems: Practical Advice with Quality Telemetry [COMMUNITY-TOOL] β A curated technical resource and architectural guide covering Debugging Kubernetes Systems: Practical Advice with Quality Telemetry in the Kubernetes Tools ecosystem.
Architecture¶
Pod Lifecycle¶
Ephemeral Containers¶
- (2023) linkedin.com: Kubernetes Ephemeral Containers | Bibin Wilson [COMMUNITY-TOOL] β Practical implementation breakdown demonstrating how engineers can leverage kubectl debug to attach ephemeral containers to scratch or distroless pods. Provides real-world configuration examples and use cases.
- (2023) iximiuz.com: Kubernetes Ephemeral Containers and kubectl debug Command π [COMMUNITY-TOOL] [GUIDE] β A comprehensive deep-dive into the architectural mechanics of kubectl debug and ephemeral containers. Explains container namespaces sharing, security contexts, and step-by-step diagnostic workflows on live clusters.
- (2022) opensource.googleblog.com: Introducing Ephemeral Containers [ADVANCED LEVEL] [COMMUNITY-TOOL] β The foundational announcement from Google detailing the integration of ephemeral containers in Kubernetes. Outlines the architectural mechanics of injecting debugging tooling directly into running distroless or minimal containers via the API.
Chaos Engineering¶
Curated Playbooks¶
Awesome Lists¶
- (2023) Awesome Chaos Engineering β 6589 πππππ [DE FACTO STANDARD] β The premier curated directory of resources, tools, and papers dedicated to the practice of Chaos Engineering. It indexes tools for simulating network latency, injecting resource stress, and terminating instances across various platforms, with a strong focus on cloud-native environments. This is a must-have reference for engineering teams building self-healing, fault-tolerant distributed systems.
Cloud-Native Platforms¶
Operating Systems¶
Flatcar Linux¶
- (2024) kinvolk.io [ADVANCED LEVEL] [COMMUNITY-TOOL] β The official portal of Kinvolk (acquired by Microsoft), pioneers of Flatcar Container Linux, Lokomotive Kubernetes, and Inspektor Gadget. The platform represents an essential pillar in the development of minimal, immutable operating systems and eBPF-based Kubernetes tooling.
Container Runtimes¶
containerd¶
ctr CLI¶
- (2023) labs.iximiuz.com: How to work with container images using ctr [ADVANCED LEVEL] [COMMUNITY-TOOL] [GUIDE] β Deep technical laboratory exercise focused on managing low-level container images using the containerd 'ctr' CLI. Vital for operations engineers debugging nodes directly where high-level runtimes like docker are not installed.
Development Workflow¶
IDE Extensions¶
Bridge to Kubernetes¶
- (2023) marketplace.visualstudio.com: Bridge to Kubernetes (VSCode) [TYPESCRIPT CONTENT] [COMMUNITY-TOOL] β Official VS Code extension that implements the Bridge to Kubernetes local-to-remote cluster redirection framework. Allows developers to step through breakpoints in their local environment while acting as an integrated cluster endpoint.
Local Development¶
Bridge to Kubernetes (1)¶
- (2021) thorsten-hans.com: Debugging apps in Kubernetes with Bridge [COMMUNITY-TOOL] [GUIDE] β Hands-on guide focusing on Microsoft's 'Bridge to Kubernetes' tool for direct local debugging within an active cluster context. Evaluates how the mechanism redirects traffic to a local machine without modifying production routing topologies.
Infrastructure¶
Compute¶
CPU Performance¶
- (2023) The Hidden CPU Throttling Crisis in Kubernetes Clusters [NONE CONTENT] [ADVANCED LEVEL] [COMMUNITY-TOOL] β Investigates latency spikes caused by the Linux CFS (Completely Fair Scheduler) quota enforcement mechanism in Kubernetes. Highlights how kernel bugs throttle containerized workloads even when usage is far below limits, providing remediation strategies such as adjusting quotas or using CPU pins.
Kubernetes¶
Core Mechanics¶
Pod Lifecycle (1)¶
- (2020) erkanerol.github.io: I wish pods were fully restartable [ADVANCED LEVEL] [COMMUNITY-TOOL] β This reflective post explores the architectural limitations of the Kubernetes Pod lifecycle, specifically the inability to easily restart individual containers without recreating the entire Pod. It analyzes proposed community designs and current workarounds, such as mutating controllers or rollout restarts. It provides deep architectural insights into the design decisions governing Kubernetes resource controllers.
Networking¶
Ingress Troubleshooting¶
- (2019) managedkube.com: Troubleshooting a Kubernetes ingress [COMMUNITY-TOOL] β Kubernetes Ingress troubleshooting requires systematically tracing path matching, port definitions, and target service selectors. This guide isolates a common failure pattern where configuration mismatches between Ingress specifications and Pod ports disrupt traffic flow. It provides actionable diagnostic commands to verify endpoint connectivity and route mapping.
Packet Tracing¶
- (2021) itnext.io: Tracing Pod2Pod Network Traffic in Kubernetes | Daniele Polencic [ADVANCED LEVEL] [COMMUNITY-TOOL] β When service-to-service communication fails silently, engineers must dive deep into packet inspection. This tutorial guides readers through capture techniques, using tools like tcpdump and Wireshark inside container network namespaces to trace pod-to-pod traffic. It addresses the complexity of tracing traffic within overlay networks and virtual ethernet pairs.
Observability¶
Events¶
- (2023) groundcover.com: Failure Is an Option: How to Stay on Top of K8s Container Events [COMMUNITY-TOOL] β Kubernetes events are transient cluster logs that provide critical context for system state changes and errors, but their short lifespan makes persistent logging essential. This article analyzes strategies for collecting, storing, and visualizing container events to catch intermittent failures. Utilizing event streaming helps platform engineers build early-warning systems before issues escalate to outages.
Events Logging¶
- (2022) decisivedevops.com: Kubernetes Events β News feed of your cluster [COMMUNITY-TOOL] β This guide explains the utility of Kubernetes events as a real-time diagnostic feed of cluster activities. It demonstrates how to capture, filter, and stream these resource events to external observability backends using open-source collectors. Properly configured event logging ensures that platform teams maintain a historical record of pod crashes and scheduling failures.
Resource Management¶
CPU Scheduling¶
- (2021) andydote.co.uk: The Problem with CPUs and Kubernetes [ADVANCED LEVEL] [COMMUNITY-TOOL] β A low-level analysis of how the Linux kernel scheduler interacting with CFS quota limits can severely throttle Kubernetes pods, even when total CPU utilization is low. The author exposes the architectural friction between container resource limits and multi-threaded application runtimes like Java and Go. Platform engineers will find valuable tuning techniques to mitigate silent latency spikes caused by CPU throttling.
Metrics Collection¶
- (2023) learnitguide.net: How to Check Memory Usage of a Pod in Kubernetes? [COMMUNITY-TOOL] β Monitoring memory consumption is critical to preventing out-of-memory (OOM) evictions and performance degradation. This guide explains how to leverage the Kubernetes Metrics Server via kubectl top, alongside advanced Prometheus queries, to track real-time memory footprints. It contrasts raw memory utilization against requested resource boundaries to help right-size workloads.
OOM and Throttling¶
- (2021) sysdig.com: Kubernetes OOM and CPU Throttling [COMMUNITY-TOOL] β A highly analytical post from Sysdig comparing the operational impacts of memory exhaustion (OOM) versus CPU exhaustion (Throttling). While OOM errors lead to sudden container termination and service disruption, CPU throttling leads to slow response times and latency degradation. The guide explains how to monitor these metrics to balance cluster cost against application performance.
QoS and OOM Score¶
- (2020) cloudyuga.guru: How does Kubernetes assign QoS class to pods through OOM score? [ADVANCED LEVEL] [COMMUNITY-TOOL] β This article provides a rigorous, technical exploration of how Kubernetes maps container memory/CPU requests and limits to Quality of Service (QoS) classes: Guaranteed, Burstable, and BestEffort. It details how the orchestrator translates these classes into Linux OOM scores, which the kernel uses to select victim processes during memory pressure. Mastering this mapping is vital for designing reliable, high-density cluster workloads.
Runtime Security¶
Sandboxed Containers¶
- (2022) cloud.redhat.com: Troubleshooting Sandboxed Containers Operator [ADVANCED LEVEL] [COMMUNITY-TOOL] β Deploying secure, isolated runtimes within Kubernetes demands deep knowledge of sandboxed container configurations, such as Kata Containers. This guide explores troubleshooting the Sandboxed Containers Operator on OpenShift, focusing on kernel module loading and hardware virtualization flags. It is an indispensable resource for platform engineers managing untrusted multi-tenant workloads.
Security¶
Container Images¶
- (2020) youtube: 3 Ways to Detect Evil "Latest" Image Tags in Kubernetes - Kubevious [COMMUNITY-TOOL] β Using the 'latest' container image tag in production is a critical antipattern that breaks deployment reproducibility and introduces security risks. This guide uses Kubevious to outline three automated detection techniques to intercept and block these non-deterministic tags in deployment pipelines. It emphasizes implementing admission controllers to enforce strict version tagging.
Troubleshooting¶
Architectural Slides¶
- (2019) speakerdeck.com/mhausenblas (redhat): Troubleshooting Kubernetes apps [COMMUNITY-TOOL] β This presentation slide deck by Michael Hausenblas maps out the complete lifecycle of Kubernetes application troubleshooting, from log collection to packet analysis. It offers a structured classification of failure domains, targeting scheduling, networking, and application-level errors. Highly recommended for visual learners seeking a comprehensive overview of cloud-native debugging patterns.
Case Study¶
- (2021) thenewstack.io: What David Flanagan Learned Fixing Kubernetes Clusters [COMMUNITY-TOOL] β This synthesis highlights critical real-world lessons from fixing broken Kubernetes clusters under pressure, illustrating how misconfigurations lead to spectacular outages. It emphasizes that basic oversights in DNS configuration, RBAC permissions, and resource constraints cause the majority of cluster-wide failures. This post-mortem review serves as a warning and a guide to preventative cluster maintenance.
Command Line¶
- (2021) thenewstack.io: Living with Kubernetes: Debug Clusters in 8 Commands π [COMMUNITY-TOOL] β A highly practical command-line cheat sheet that isolates eight crucial kubectl commands for diagnosing cluster and application failures. By leveraging advanced output formatting and resource filtering, engineers can quickly drill down from high-level service failures to underlying node exhaustion. This primer is designed for rapid incident response and system validation.
CrashLoopBackOff¶
- (2023) devtron.ai: Troubleshoot: Pod Crashloopbackoff [COMMUNITY-TOOL] β A targeted diagnostic manual focusing exclusively on resolving Pod CrashLoopBackOff errors within a Kubernetes cluster. It details the precise troubleshooting path, examining application configuration errors, environment variable omissions, and permission issues. It also includes strategies for extracting logs from previously terminated container instances.
- (2022) komodor.com: Kubernetes CrashLoopBackOff Error: What It Is and How to Fix It [COMMUNITY-TOOL] β A comprehensive playbook from Komodor designed to help engineers systematically isolate the root causes of CrashLoopBackOff errors. It covers common trigger events such as database connection timeouts, file access permissions, and mismatched configuration maps. By mapping these triggers, it helps engineers transition from reactive troubleshooting to proactive reliability planning.
- (2021) sysdig.com: What is Kubernetes CrashLoopBackOff? And how to fix it π [COMMUNITY-TOOL] β A definitive guide from Sysdig addressing the causes, mechanics, and remediation steps for the ubiquitous CrashLoopBackOff state. The article unpacks how the kubelet manages backoff loops when an application continuously crashes on startup. It outlines diagnostic strategies combining container telemetry with log output to pinpoint runtime exceptions.
Developer Enablement¶
- (2021) thenewstack.io: 6 Kubernetes Best Practices to Empower Devs to Troubleshoot [COMMUNITY-TOOL] β DevOps maturity relies on empowering application developers to diagnose and resolve their own Kubernetes errors without administrative intervention. By providing accessible log aggregation, standardized event alerts, and ephemeral debugging shells, platform teams can eliminate operational bottlenecks. This framework details how to design developer-friendly observability environments.
Development Environments¶
- (2023) devzero.io: Kubernetes Debugging Tips [COMMUNITY-TOOL] β Optimizing the inner-loop developer experience in Kubernetes requires special tooling to bypass slow build-and-deploy cycles. This guide reviews practical tips for configuring local debugging environments, port forwarding, and syncing code directly into remote dev clusters. Implementing these techniques allows developers to debug code in real time within a cluster-like context.
Diagnostic Manuals¶
- (2022) github.com/metaleapca: metaleap-k8s-troubleshooting.pdf β 40 [ADVANCED LEVEL] πππππ [DE FACTO STANDARD] β An exhaustive community-driven reference PDF that serves as a diagnostic field manual for complex Kubernetes cluster issues. It covers everything from low-level CNI networking issues to control plane degradation and API server latency. This document is a valuable offline resource for operations teams handling large-scale production incidents.
Distroless Debugging¶
- (2021) itnext.io: Distroless Container Debugging on K8s/OpenShift [ADVANCED LEVEL] [COMMUNITY-TOOL] β While distroless images enhance security by stripping out non-essential utilities and shells, they present unique challenges when live troubleshooting is required. This article explores strategies to bridge this gap, utilizing Kubernetes ephemeral containers and volume-mounting debug utilities on OpenShift. It demonstrates how to maintain a minimal attack surface without sacrificing operational diagnostics.
Ephemeral Containers (1)¶
- (2022) loft.sh: Using Kubernetes Ephemeral Containers for Troubleshooting [ADVANCED LEVEL] [COMMUNITY-TOOL] β This guide explains how to leverage Kubernetes Ephemeral Containersβthe native, standard way to troubleshoot running pods without pre-packaging debug tools in production images. By launching a temporary diagnostic container within the target pod's namespace, engineers can safely run tools like curl, tcpdump, or strace. This represents the modern, secure gold standard for debugging production distroless containers.
Exit Codes¶
- (2022) komodor.com: Exit Codes In Containers & Kubernetes β The Complete Guide π [COMMUNITY-TOOL] β An industry-standard reference guide demystifying exit codes returned by failing containers, from common application errors (Exit Code 1) to system-level kills (Exit Code 137). It details how the container runtime and Linux kernel interact to generate these codes, providing a crucial diagnostic map for troubleshooting CrashLoopBackOffs. Understanding these codes is essential for automated root-cause analysis.
Fundamentals¶
- (2021) thenewstack.io: Kubernetes Troubleshooting Primer [COMMUNITY-TOOL] β A foundational text that introduces the core mechanics of the Kubernetes control plane and how components interact during deployment failures. It explains the relationship between the API server, controller manager, scheduler, and kubelet, illustrating where common failures arise. This primer establishes the necessary context for executing advanced diagnostics.
General Guide¶
- (2021) blog.alexellis.io: How to Troubleshoot Applications on Kubernetes π [COMMUNITY-TOOL] β A classic, hands-on troubleshooting guide tailored for engineers deploying applications to Kubernetes clusters. It provides a logical flow of commands from checking container logs and inspecting events to executing commands inside a container. This resource is excellent for developing the primary instincts required to resolve containerized application failures.
Methodologies¶
- (2022) freecodecamp.org: How to Simplify Kubernetes Troubleshooting [COMMUNITY-TOOL] β This comprehensive troubleshooting guide reframes cluster diagnostics into a logical, step-by-step methodology rather than chaotic guesswork. It focuses on isolating issues systematically across the application layer, the container runtime, and the underlying orchestrator networking. Utilizing this structured approach significantly lowers the cognitive load of on-call engineers.
OOMKilled¶
- (2021) itnext.io: Kubernetes Silent Pod Killer [ADVANCED LEVEL] [COMMUNITY-TOOL] β This article investigates the elusive scenarios where the Linux kernel OOM killer terminates container processes silently without registering a clean Kubernetes OOMKilled status. This discrepancy occurs when sub-processes within a container are targeted, leaving the main container entrypoint running but degraded. The author shares advanced debugging tips utilizing system logs and container runtimes to detect these hidden failures.
Pod Diagnostics¶
- (2023) learnitguide.net: How To Troubleshoot Kubernetes Pods [COMMUNITY-TOOL] β A focused tutorial describing standard commands and methodologies to diagnose Pod failures such as ImagePullBackOff, CrashLoopBackOff, and Evicted states. It systematically walks through the utilization of kubectl describe, logs, and events to construct an accurate failure timeline. It is perfect for junior engineers looking to build baseline debugging proficiency.
Pod Eviction¶
- (2021) sysdig.com: Understanding Kubernetes Evicted Pods [COMMUNITY-TOOL] β Pod eviction is a proactive defense mechanism executed by the kubelet when a node faces critical resource pressure, such as low disk space or memory exhaustion. This guide details the eviction lifecycle, explaining the difference between soft and hard eviction thresholds. It provides practical strategies for configuring tolerations and cluster autoscaling to prevent widespread application downtime.
Pod Scheduling¶
- (2021) sysdig.com: Understanding Kubernetes pod pending problems [COMMUNITY-TOOL] β Pods trapped in a 'Pending' state point directly to scheduling failures caused by resource exhaustion, node selectors, or volume binding delays. This deep dive from Sysdig explains how to read scheduler events and analyze resource request limits to unlock stuck deployments. Engineers will learn how to identify node-affinity conflicts and taint/toleration mismatches.
Real-World Examples¶
- (2022) tennexas.com: Kubernetes Troubleshooting Examples [COMMUNITY-TOOL] β An essential compilation of practical Kubernetes troubleshooting scenarios focusing on common network, storage, and configuration failures. It provides hands-on diagnostic scripts and commands to dissect failing pods and unresolvable services. This playbook bridges the gap between theoretical knowledge and day-to-day cluster administration.
Troubleshooting Tooling¶
kubectl plugins¶
- (2020) kubectl-debug β 2305 [ADVANCED LEVEL] πππππ [DE FACTO STANDARD] [LEGACY] β Originally a popular community-built plugin to launch debugging containers within target pods,
kubectl-debughas largely been superseded by native Kubernetes Ephemeral Containers (kubectl debugcommand) in modern releases. This project remains a valuable reference for historical context and legacy cluster compatibility. For modern clusters, engineers are strongly advised to transition to built-in Kubernetes diagnostic commands.
Kubernetes Platform Engine¶
Cluster Operations¶
Memory Management¶
- (2025) OOMKilled in Kubernetes: Understanding and Preventing Hidden Memory Leaks [N/A CONTENT] [ADVANCED LEVEL] [COMMUNITY-TOOL] β Diagnoses Kubernetes
OOMKilled(Exit Code 137) events caused by memory leaks, misconfigured resource limits, and JVM heap management issues. Explains how to set appropriate limits/requests while implementing profiling tools to prevent container churn.
Observability (1)¶
Debugging¶
Automation¶
- (2022) github.com/airwallex: k8s-pod-restart-info-collector [GO CONTENT] [COMMUNITY-TOOL] β An automated utility designed to capture and log state configuration, events, and logs instantly when a pod restarts, offering immediate post-mortem insights for ephemeral microservices.
CLI Extensions¶
- (2021) github.com/JamesTGrant/kubectl-debug β 373 [GO CONTENT] ππ [COMMUNITY-TOOL] β A kubectl plugin designed to launch temporary debugging containers within a target pod namespace. Streamlines manual container introspection prior to the widespread adoption of native ephemeral containers.
CLI Operations¶
- (2021) thenewstack.io: Living with Kubernetes: 12 Commands to Debug Your Workloads π [COMMUNITY-TOOL] β Curated list of 12 essential kubectl commands designed to streamline low-level container and network diagnostics. Targets common day-2 operational challenges, addressing resource pressure, storage attachments, and system event inspection.
Container Debugging¶
- (2025) iximiuz/cdebug β 1652 [GO CONTENT] [ADVANCED LEVEL] πππ [COMMUNITY-TOOL] β A specialized CLI tool for debugging running containers. Allows attaching ephemeral tooling environments into running, stripped-down containers (even without Kubernetes, working directly with Docker/containerd).
- (2022) felipecruz91/debug-ctr β 52 [GO CONTENT] π [COMMUNITY-TOOL] β Lightweight helper utility targeting runtimes at the node level. Allows developers to run custom debug tools directly inside active containerd namespaces without impacting root node security models.
Containers¶
- (2023) KDBG: Small Kubernetes debugging container β 36 [SHELL CONTENT] π [COMMUNITY-TOOL] β A lightweight Kubernetes debugging container designed to simplify troubleshooting inside active pods. It packages key diagnostic tools (curl, dig, iproute2, etc.) for direct execution inside the cluster networking space, serving as an effective sidecar or ephemeral debugging agent.
Legacy Tooling¶
- (2021) palaemon.io [LEGACY] β Legacy troubleshooting tool aimed at optimizing container workloads and analyzing scheduling configurations. Now largely archived but remains historically relevant for declarative scheduling architectures.
Pre-flight Checks¶
- (2025) github.com/replicatedhq/troubleshoot β 582 [GO CONTENT] [ADVANCED LEVEL] ππ [COMMUNITY-TOOL] β A framework providing preflight checks and support-bundle collection capabilities for Kubernetes applications. Crucial for enterprise environments deploying applications onto heterogeneous on-premises or customer-managed clusters.
Troubleshooting Guide¶
- (2023) learnk8s.io: A visual guide on troubleshooting Kubernetes deployments [COMMUNITY-TOOL] [GUIDE] β A high-density visual guide detailing a deterministic flowchart for troubleshooting Kubernetes deployment failures. It systematically walks engineers through checking ingress, service routing, selector matching, and pod-level failures (e.g., CrashLoopBackOff).
- (2023) komodor.com: Kubernetes Troubleshooting: The Complete Guide π [COMMUNITY-TOOL] [GUIDE] β An exhaustive architectural manual dissecting everyday Kubernetes failure patterns including OOMKilled, ImagePullBackOff, CrashLoopBackOff, and CPU throttling.
Workloads¶
- (2022) towardsdatascience.com: The Easiest Way to Debug Kubernetes Workloads [COMMUNITY-TOOL] β Practical overview detailing baseline approaches to troubleshooting Kubernetes workloads, focusing on common command-line routines. It bridges the gap for data engineers and developers needing rapid context on kubectl logs, describes, and port-forwards.
Deployments¶
Telemetry¶
- (2021) StatusBay β 387 [GO CONTENT] ππ [COMMUNITY-TOOL] β An active deployment monitoring tool that provides real-time visibility into Kubernetes deployment sequences. By aggregating event logs and state transitions, StatusBay offers clean diagnostic traces for failed releases, improving post-mortem analysis.
Networking (1)¶
API Traffic Analyzer¶
- (2024) kubetools.io: Kubeshark β API Traffic Analyzer for Kubernetes [ADVANCED LEVEL] [COMMUNITY-TOOL] β An advanced network analyzer for Kubernetes (formerly API Tap) that captures and decrypts cluster traffic in real-time. Leverages modern eBPF and packet capture technologies to trace service-to-service communication.
Troubleshooting Platforms¶
Enterprise Monitoring¶
- (2024) komodor.com [COMMUNITY-TOOL] β Commercial troubleshooting and observability platform offering absolute end-to-end lineage visualization for Kubernetes resources. Pinpoints root causes of failures by cross-referencing changes, logs, and metrics.
UI Clients¶
Multi-Cluster¶
- (2024) KubeUI: A Desktop Kubernetes Client β 311 [C# CONTENT] ππ [COMMUNITY-TOOL] β A high-performance, desktop-optimized UI designed to stream, monitor, and interact with live cluster metrics and objects. It enhances developer agility through dynamic views of multi-cluster namespaces and active workload metrics.
π‘ Explore Related: Kubernetes Storage | Kubernetes Alternatives | Kubernetes Client Libraries