0%

Preparing the page

Kubernetes Observability: What to Monitor, Which Tools to Use and How to Alert

A practical guide to Kubernetes observability: what to monitor at cluster, node, pod and application level, key metrics, tools and alerting best practices.

Girish

· 4 min read

Share
Placeholder illustration

Kubernetes is excellent at keeping workloads running, which can hide problems. Pods restart silently, nodes are replaced automatically and resources are shuffled around. Without Kubernetes observability, you see the symptoms only when users do.

This guide covers what to monitor, which tools to use and how to alert without drowning your team.

Kubernetes is excellent at keeping workloads running, which can hide problems.

What makes Kubernetes observability different?

  • Ephemeral resources. Pods and containers come and go; their names and IPs change constantly.
  • Many layers. A slow request could be caused by the application, the pod's limits, the node, the network or the control plane.
  • Shared infrastructure. Many teams run on the same clusters.
  • Self-healing hides failure. A crash-looping container restarts automatically, so availability may look fine while performance suffers.

What should you monitor in Kubernetes?

Cluster health
LayerKey signalsCommon sources
Control planeAPI server latency and errors, etcd health, scheduler queueManaged service metrics (EKS, AKS, GKE)
NodesCPU, memory, disk, network, node conditionsnode-exporter, kubelet
WorkloadsPod restarts, OOMKills, pending pods, CPU throttling, requests vs usagekube-state-metrics, cAdvisor
ApplicationsRequest rate, errors, latency (RED), business metricsOpenTelemetry, app metrics
Events and logsScheduling failures, image pull errors, app logsKubernetes events, container logs

Which Kubernetes metrics matter most?

  • Container restarts and OOMKilled events: often the first sign of memory limits set too low.
  • Pending pods: the cluster cannot schedule workloads, usually due to capacity or constraints.
  • CPU throttling: CPU limits are slowing applications, raising latency.
  • Requests vs actual usage: over-provisioning wastes money; under-provisioning risks instability. See Kubernetes cost optimization.
  • Node pressure conditions: memory, disk or PID pressure leading to evictions.
  • Application RED metrics: rate, errors and duration for each service.

Which tools are used for Kubernetes observability?

  • Prometheus (or an OpenTelemetry Collector) to scrape metrics, with kube-state-metrics and node-exporter
  • Long-term metrics storage such as Grafana Mimir
  • Log collection with Grafana Alloy, Fluent Bit or the OpenTelemetry Collector into Loki or another backend
  • Tracing with OpenTelemetry and a backend like Tempo
  • Dashboards and alerting in Grafana

Grafana Cloud's Kubernetes Monitoring provides much of this pre-configured. See our LGTM stack guide.

How should you alert on Kubernetes?

Avoid paging on every pod restart. Instead:

Page vs ticket alert routing
  1. Page on user-facing SLO burn rates for each service.
  2. Page on platform failures that will affect many services: nodes not ready, API server errors, persistent pending pods.
  3. Ticket on workload hygiene: frequent restarts, OOMKills, throttling, requests far above usage.
  4. Group by namespace and team so alerts go to owners.

This approach keeps on-call focused, as covered in our alert fatigue playbook.

How do you label Kubernetes telemetry?

Consistent labels make everything else work. Use standard labels such as app.kubernetes.io/name, plus team and environment. Ensure OpenTelemetry resource attributes (k8s.namespace.name, k8s.pod.name, service.name) are attached to every signal, so you can move from a metric to traces and logs for the same pod.

How Crozaint approaches Kubernetes observability

Crozaint runs Kubernetes and OpenShift workloads across AWS, Azure and Google Cloud. Our observability engagement instruments clusters and applications with OpenTelemetry, consolidates telemetry on Grafana Cloud, defines SLOs per service and builds alerting that separates pages from tickets.

For teams that want round-the-clock cover, Crozaint Ops · Live adds 24/7 operations on top of the platform, and our managed services team can run the clusters themselves.

Common mistakes to avoid

  • Paging on every pod restart
  • Monitoring nodes but not applications
  • No resource requests or limits, making metrics hard to interpret
  • Inconsistent labels across teams
  • Keeping all container logs at full volume forever

Conclusion

Kubernetes observability means seeing every layer and connecting them. Collect platform and application telemetry, label it consistently and alert on what users feel.

Running Kubernetes in production? Book a 30-minute discovery call with Crozaint.

Frequently Asked Questions

What is the best tool for Kubernetes monitoring?

There is no single best tool. Most teams combine Prometheus-compatible metrics, kube-state-metrics and node-exporter with a log backend, tracing via OpenTelemetry and Grafana dashboards. Managed platforms such as Grafana Cloud package these together and reduce the operational effort of running them.

What is kube-state-metrics?

kube-state-metrics is a service that listens to the Kubernetes API and generates metrics about the state of objects such as deployments, pods, nodes and jobs. It shows desired versus actual replicas, pod phases and restart counts, which are essential for workload health monitoring.

How do I monitor EKS, AKS or GKE?

Managed Kubernetes services expose control plane metrics and logs through their cloud providers. Combine those with in-cluster metrics from kube-state-metrics, node-exporter and cAdvisor, application telemetry from OpenTelemetry, and container logs, all sent to a unified observability platform.

Why do my pods keep getting OOMKilled?

Pods are OOMKilled when a container exceeds its memory limit. Common causes are limits set too low, memory leaks or traffic spikes. Compare actual memory usage with limits over time, fix leaks, and set limits based on observed peaks plus a safety margin.

Should I alert on pod restarts?

Usually not as a page. Kubernetes restarts failing containers automatically, and occasional restarts may not affect users. Page on user-facing SLO impact and platform-wide failures; send frequent restarts and OOMKills to a ticket queue for investigation during working hours.

Written by

Girish

Crozaint · 15 articles

Full profile coming soon.

JosephReviewed for technical accuracy by Joseph, Cloud Consulting.

After the reading

Reading About Observability Is the Easy Part.Doing It in Your Estate Is Ours.

Thirty minutes with the people who wrote this. We look at your setup, say what we would fix first and leave you with a plan, whether or not you go further with us.

  • A look at your estate, not a demo
  • What we would fix first, and why
  • A plan you keep, whether or not you hire us
Joseph

Talk to Joseph

Cloud Consulting

Thirty minutes on your estate. Joseph looks at what you have and tells you what we would do first.

Book 30 Minutes

No deck, no pitch, no commitment.