Starting with Dynatrace version 1.348, Dynatrace is updating Kubernetes alert severity classifications and introducing new recommended alert defaults. Learn what changes automatically and what requires your action.
Kubernetes alerting controls which conditions open a Problem, which appear as Warning signals, and which are monitored. The release brings two changes: a severity reclassification that applies automatically to a set of existing alerts, and new recommended defaults you can optionally apply.
Kubernetes alerts that previously raised a Problem now change to Warning signals. This reclassification applies automatically to all environments.
Affected alerts appear as Warning signals in
Kubernetes and contribute to the health status of your Kubernetes objects. Warning signals do not open a Problem investigation or trigger alert notifications. This reduces noise from conditions that indicate a potential concern rather than an active incident.
See Alert reference for affected alerts.
With the Dynatrace version 1.348 release, new alerting defaults are available for all environments. The new defaults define which alerts are enabled and their detection thresholds. Your current configuration is not affected until you explicitly apply them under
Settings > Analyze and alert > Alerts > Kubernetes.
Each alert has two dimensions:
Applying the recommended settings enables additional alerts and updates detection thresholds. Before applying, be aware that some newly enabled alerts have Critical severity and open a Problem when triggered.
The following alerts are disabled in the current defaults and become enabled with Critical severity when you apply the recommended settings. Review them before applying so you know which conditions raise Problems.
| Alert | Trigger condition |
|---|---|
Detect stuck deployments | Stops progressing ≥3 min; within 5 min |
Detect workloads without ready pods | No ready pods ≥10 min; within 15 min |
Detect pod backoff events | Event-based |
Detect pod eviction events | Event-based |
The following tables list all Kubernetes alerts with their recommended settings.
¹ Newly enabled in recommended settings
² Reclassified from health alert to warning signal
| Alert | Enabled by default | Threshold | Severity |
|---|---|---|---|
Detect cluster readiness issues | Yes | Not ready ≥3 min; within 5 min | Critical |
Detect cluster CPU-request saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect cluster memory-request saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect cluster pod-saturation | No | >90% pods / ≥3 min / within 5 min | Warning signal ² |
Detect monitoring issues | Yes ¹ | Not available ≥15 min; within 30 min | Warning signal ¹ ² |
| Alert | Enabled by default | Threshold | Severity |
|---|---|---|---|
Detect node readiness issues | Yes ¹ | Not ready ≥3 min; within 5 min | Warning signal ¹ ² |
Detect problematic node conditions | Yes ¹ | Problematic ≥3 min; within 5 min | Warning signal ¹ ² |
Detect node CPU-request saturation | No | >90% / ≥3 min / within 5 min | Warning signal ² |
Detect node memory-request saturation | No | >90% / ≥3 min / within 5 min | Warning signal ² |
Detect node pod-saturation | No | >90% pods / ≥3 min / within 5 min | Warning signal ² |
| Alert | Enabled by default | Threshold | Severity |
|---|---|---|---|
Detect namespace CPU-request quota saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect namespace CPU-limit quota saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect namespace memory-request quota saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect namespace memory-limit quota saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect namespace pod quota saturation | No | >90% pods / ≥10 min / within 15 min | Warning signal ² |
| Alert | Enabled by default | Threshold | Severity |
|---|---|---|---|
Detect container restarts | Yes ¹ | ≥1 restart per 3 min; within 5 min | Warning signal ¹ ² |
Detect stuck deployments | Yes ¹ | Stops progressing ≥3 min; within 5 min | Critical ¹ |
Detect pods stuck in pending | Yes ¹ | ≥1 pod stuck ≥10 min; within 15 min | Warning signal ¹ ² |
Detect pods stuck in terminating | Yes ¹ | Termination stopped ≥10 min; within 15 min | Warning signal ¹ ² |
Detect workloads without ready pods | Yes ¹ | No ready pods ≥10 min; within 15 min | Critical ¹ |
Detect workloads with non-ready pods | No | >30% / within 60 min | Warning signal ² |
Detect memory usage saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect CPU usage saturation | No | >90% / ≥10 min / within 15 min | Warning signal ² |
Detect high CPU throttling | No | >50% throttled / ≥10 min / within 15 min | Warning signal ² |
Detect out-of-memory kills | Yes ¹ | Event-based | Warning signal ¹ ² |
Detect job failure events | Yes ¹ | Event-based | Warning signal ¹ ² |
Detect pod backoff events | Yes ¹ | Event-based | Critical ¹ |
Detect pod eviction events | Yes ¹ | Event-based | Critical ¹ |
Detect pod preemption events | Yes ¹ | Event-based | Warning signal ¹ ² |
| Alert | Enabled by default | Threshold | Severity |
|---|---|---|---|
Detect low disk space (MiB) | No | ≤100 MiB remaining / ≥3 min / within 5 min | Warning signal ² |
Detect low disk space (%) | Yes ¹ | <3% available / ≥3 min; within 5 min | Warning signal ¹ ² |
To review or adjust your Kubernetes alert configuration
Kubernetes