Kubernetes is excellent at keeping workloads running, which can hide problems. Pods restart silently, nodes are replaced automatically and resources are shuffled around. Without Kubernetes observability, you see the symptoms only when users do.
This guide covers what to monitor, which tools to use and how to alert without drowning your team.
Kubernetes is excellent at keeping workloads running, which can hide problems.
What makes Kubernetes observability different?
- Ephemeral resources. Pods and containers come and go; their names and IPs change constantly.
- Many layers. A slow request could be caused by the application, the pod's limits, the node, the network or the control plane.
- Shared infrastructure. Many teams run on the same clusters.
- Self-healing hides failure. A crash-looping container restarts automatically, so availability may look fine while performance suffers.
What should you monitor in Kubernetes?
Kubernetes cluster health dashboard in Grafana
Illustration in progress
| Layer | Key signals | Common sources |
|---|---|---|
| Control plane | API server latency and errors, etcd health, scheduler queue | Managed service metrics (EKS, AKS, GKE) |
| Nodes | CPU, memory, disk, network, node conditions | node-exporter, kubelet |
| Workloads | Pod restarts, OOMKills, pending pods, CPU throttling, requests vs usage | kube-state-metrics, cAdvisor |
| Applications | Request rate, errors, latency (RED), business metrics | OpenTelemetry, app metrics |
| Events and logs | Scheduling failures, image pull errors, app logs | Kubernetes events, container logs |
Which Kubernetes metrics matter most?
- Container restarts and OOMKilled events: often the first sign of memory limits set too low.
- Pending pods: the cluster cannot schedule workloads, usually due to capacity or constraints.
- CPU throttling: CPU limits are slowing applications, raising latency.
- Requests vs actual usage: over-provisioning wastes money; under-provisioning risks instability. See Kubernetes cost optimization.
- Node pressure conditions: memory, disk or PID pressure leading to evictions.
- Application RED metrics: rate, errors and duration for each service.
Which tools are used for Kubernetes observability?
- Prometheus (or an OpenTelemetry Collector) to scrape metrics, with kube-state-metrics and node-exporter
- Long-term metrics storage such as Grafana Mimir
- Log collection with Grafana Alloy, Fluent Bit or the OpenTelemetry Collector into Loki or another backend
- Tracing with OpenTelemetry and a backend like Tempo
- Dashboards and alerting in Grafana
Grafana Cloud's Kubernetes Monitoring provides much of this pre-configured. See our LGTM stack guide.
How should you alert on Kubernetes?
Avoid paging on every pod restart. Instead:
Kubernetes alerting: pages vs tickets
Illustration in progress
- Page on user-facing SLO burn rates for each service.
- Page on platform failures that will affect many services: nodes not ready, API server errors, persistent pending pods.
- Ticket on workload hygiene: frequent restarts, OOMKills, throttling, requests far above usage.
- Group by namespace and team so alerts go to owners.
This approach keeps on-call focused, as covered in our alert fatigue playbook.
How do you label Kubernetes telemetry?
Consistent labels make everything else work. Use standard labels such as app.kubernetes.io/name, plus team and environment. Ensure OpenTelemetry resource attributes (k8s.namespace.name, k8s.pod.name, service.name) are attached to every signal, so you can move from a metric to traces and logs for the same pod.
How Crozaint approaches Kubernetes observability
Crozaint runs Kubernetes and OpenShift workloads across AWS, Azure and Google Cloud. Our observability engagement instruments clusters and applications with OpenTelemetry, consolidates telemetry on Grafana Cloud, defines SLOs per service and builds alerting that separates pages from tickets.
For teams that want round-the-clock cover, Crozaint Ops · Live adds 24/7 operations on top of the platform, and our managed services team can run the clusters themselves.
Common mistakes to avoid
- Paging on every pod restart
- Monitoring nodes but not applications
- No resource requests or limits, making metrics hard to interpret
- Inconsistent labels across teams
- Keeping all container logs at full volume forever
Conclusion
Kubernetes observability means seeing every layer and connecting them. Collect platform and application telemetry, label it consistently and alert on what users feel.
Running Kubernetes in production? Book a 30-minute discovery call with Crozaint.
