Browse Library
BASIC RESOURCE MONITORING
Prometheus self-monitoring
36 RULESPrometheus target missing, Prometheus all targets missing, Prometheus target missing with warmup time, Prometheus not connected to alertmanager, Prometheus job missing
Host and hardware
41 RULESHost high CPU load, Host CPU steal noisy neighbor, Host CPU high iowait, Host CPU load saturation, Host CPU is underutilized
S.M.A.R.T Device Monitoring
9 RULESSMART critical warning, SMART device temperature critical, SMART device temperature over trip value, SMART device temperature warning, SMART device temperature nearing trip value
IPMI
17 RULESIPMI chassis power off, IPMI collector down, IPMI SEL almost full, IPMI temperature sensor critical, IPMI fan speed sensor critical
Docker containers
10 RULESContainer killed, Container absent, Container High CPU utilization, Container high throttle rate, Container high low change CPU usage
Blackbox
10 RULESBlackbox probe failed, Blackbox probe HTTP failure, Blackbox probe low uptime, Blackbox slow probe, Blackbox probe slow HTTP
Windows Server
8 RULESWindows Server CPU Usage, Windows Server memory Usage, Windows Server disk Space Usage, Windows Server disk drive Status, Windows Server NTP client delay
VMware
4 RULESVirtual Machine Memory Critical, Virtual Machine Memory Warning, High Number of Snapshots, Outdated Snapshots
Proxmox VE
9 RULESPVE node down, PVE cluster not quorate, PVE VM/CT down, PVE high CPU usage, PVE high memory usage
Netdata
9 RULESNetdata high cpu usage, Netdata CPU steal noisy neighbor, Netdata high memory usage, Netdata low disk space, Netdata predicted disk full
eBPF
3 RULESeBPF exporter no enabled configs, eBPF exporter program not attached, eBPF exporter decoder errors
Process Exporter
10 RULESProcess exporter group down, Process exporter high CPU usage, Process exporter high memory usage, Process exporter high disk write IO, Process exporter file descriptors exhausted
Systemd
7 RULESSystemd unit tasks near limit, Systemd socket high connections, Systemd service crash looping, Systemd unit failed, Systemd unit inactive
DATABASES
MySQL
18 RULESMySQL down, MySQL Slave IO thread not running, MySQL Slave SQL thread not running, MySQL too many connections (> 80%), MySQL high prepared statements utilization (> 80%)
PostgreSQL
25 RULESPostgresql down, Postgresql replication role changed, Postgresql bloat index high (> 80%), Postgresql bloat table high (> 80%), Postgresql long-running transaction
SQL Server
2 RULESSQL Server down, SQL Server deadlock
Oracle Database
8 RULESOracle DB down, Oracle DB tablespace full (> 95%), Oracle DB tablespace reaching capacity (> 85%), Oracle DB sessions reaching limit (> 85%), Oracle DB processes reaching limit (> 85%)
Patroni
1 RULESPatroni has no Leader
PGBouncer
6 RULESPGBouncer max connections, PGBouncer server connection saturation critical, PGBouncer active connections, PGBouncer clients waiting for connection, PGBouncer server connection saturation warning
Redis
16 RULESRedis down, Redis missing master, Redis too many masters, Redis disconnected slaves, Redis replication broken
Memcached
9 RULESMemcached down, Memcached memory usage high (> 90%), Memcached connection limit approaching (> 95%), Memcached connection limit approaching (> 80%), Memcached out of memory errors
MongoDB
22 RULESMongoDB Down, Mongodb replica member unhealthy, MongoDB replica set has no primary, MongoDB replica set primary changed, MongoDB too many connections (Percona)
Elasticsearch
22 RULESElasticsearch Cluster Red, Elasticsearch Healthy Nodes, Elasticsearch Healthy Data Nodes, Elasticsearch High CPU Usage, Elasticsearch Heap Usage Too High
OpenSearch
12 RULESOpenSearch host CPU usage high, OpenSearch process CPU usage high, OpenSearch high heap usage, OpenSearch disk high watermark reached, OpenSearch disk low watermark reached
Meilisearch
2 RULESMeilisearch index is empty, Meilisearch http response time
Cassandra
30 RULESCassandra Node is unavailable, Cassandra connection timeouts total (Instaclustr), Cassandra storage exceptions (Instaclustr), Cassandra tombstone dump (Instaclustr), Cassandra client request unavailable write (Instaclustr)
Clickhouse
23 RULESClickHouse node down, ClickHouse No Available Replicas, ClickHouse No Live Replicas, ClickHouse Interserver Connection Issues, ClickHouse ZooKeeper Connection Issues
CouchDB
24 RULESCouchDB node down, CouchDB atom memory usage critical, CouchDB open OS files critical, CouchDB file descriptors high, CouchDB open databases critical
Solr
10 RULESSolr shard has no leader, Solr high JVM heap usage, Solr disk space very low, Solr disk space low, Solr update errors
MESSAGE BROKERS
RabbitMQ
25 RULESRabbitMQ node down, RabbitMQ memory alarm active, RabbitMQ memory high, RabbitMQ disk space alarm active, RabbitMQ predicted low disk watermark
Zookeeper
4 RULESNo description
Kafka
5 RULESKafka topics replicas, Kafka consumer group lag, Kafka consumer group lag increasing
Pulsar
10 RULESPulsar high ledger disk usage, Pulsar subscription very high number of backlog entries, Pulsar topic very large backlog storage size, Pulsar high write latency, Pulsar read only bookies
Nats
13 RULESNats server down, Nats high CPU usage, Nats high memory usage, Nats high JetStream memory usage, Nats high JetStream store usage
PROXIES, LOAD BALANCERS AND SERVICE MESHES
Nginx
4 RULESNginx high HTTP 4xx error rate, Nginx high HTTP 5xx error rate, Nginx latency high, Nginx internal metric errors
Apache
4 RULESApache down, Apache workers load, Apache response time too high, Apache restart
HaProxy
33 RULESHAProxy HTTP slowing down, HAProxy dropping logs, HAProxy backend healthcheck flapping, HAProxy server healthcheck flapping, HAProxy high HTTP 4xx error rate backend
Traefik
9 RULESTraefik service down, Traefik high HTTP 4xx error rate service, Traefik high HTTP 5xx error rate service, Traefik certificate expiring imminently, Traefik config reload failure
Caddy
3 RULESCaddy Reverse Proxy Down, Caddy high HTTP 4xx error rate service, Caddy high HTTP 5xx error rate service
Envoy
20 RULESEnvoy server not live, Envoy high memory usage, Envoy global downstream connections overflowing, Envoy downstream connections overflowing, Envoy high downstream HTTP 5xx error rate
Linkerd
1 RULESLinkerd high error rate
Istio
11 RULESIstio control plane component down, Istio Pilot Duplicate Entry, Istio Kubernetes gateway availability drop, Istio Pilot high push error rate, Istio Mixer Prometheus dispatches low
RUNTIMES
PHP-FPM
4 RULESPHP-FPM max-children reached, PHP-FPM listen queue saturation, PHP-FPM busy process ratio high, PHP-FPM slow requests ratio
JVM
12 RULESJVM compilation time spike, JVM memory filling up, JVM non-heap memory filling up, JVM old gen GC frequency, JVM file descriptors exhaustion
Golang
10 RULESGo GC CPU fraction high, Go memory usage high, Go heap objects count high, Go heap in-use growing, Go memory leak
Ruby
5 RULESRuby heap live slots high, Ruby heap free slots high, Ruby major GC rate high, Ruby RSS high, Ruby allocated objects spike
Python
5 RULESPython GC objects uncollectable, Python GC generation 2 collections high, Python virtual memory high, Python file descriptors exhaustion, Python GC collections high
Sidekiq
2 RULESSidekiq scheduling latency too high, Sidekiq queue size
DATA ENGINEERING
Apache Flink
12 RULESFlink job is not running, Flink TaskManager heap memory high, Flink JobManager heap memory high, Flink no TaskManagers registered, Flink all task slots used
Apache Spark
8 RULESSpark worker memory exhausted, Spark executor high disk spill, Spark no alive workers, Spark executor all tasks failing, Spark too many waiting apps
Hadoop
15 RULESHadoop Name Node Down, Hadoop Resource Manager Down, Hadoop HDFS Missing Blocks, Hadoop Node Manager CPU High, Hadoop Resource Manager vCore Usage High
ORCHESTRATORS
Kubernetes
46 RULESKubernetes StatefulSet down, Kubernetes Container oom killer, Kubernetes CronJob too long, Kubernetes Node memory pressure, Kubernetes Node disk pressure
Nomad
4 RULESNomad job failed, Nomad job lost, Nomad job queued, Nomad blocked evaluation
Consul
5 RULESConsul missing master node, Consul agent unhealthy, Consul service healthcheck failed, Consul has no raft leader, Consul exporter query failure
Etcd
19 RULESEtcd no Leader, Etcd insufficient members for quorum, Etcd high fsync durations critical, Etcd high fsync durations, Etcd high commit durations
OpenStack
20 RULESOpenStack exporter down, OpenStack Nova agent down, OpenStack Neutron agent down, OpenStack Cinder agent down, OpenStack hypervisor high vCPU usage
CI/CD
Jenkins
8 RULESJenkins node offline, Jenkins no node online, Jenkins healthcheck, Jenkins builds health score, Jenkins outdated plugins
ArgoCD
9 RULESArgoCD service not synced, ArgoCD service unhealthy, ArgoCD application sync failed, ArgoCD cluster connection error, ArgoCD git fetch failures
FluxCD
4 RULESFlux Kustomization Failure, Flux HelmRelease Failure, Flux Source Issue, Flux Image Issue
GitLab CI
32 RULESGitLab Puma workers not running, GitLab high memory usage, GitLab Ruby heap fragmentation, GitLab Puma high queued connections, GitLab database connection pool saturation
Spinnaker
12 RULESSpinnaker thread pool exhaustion, Spinnaker dead messages, Spinnaker polling monitor items over threshold, Spinnaker circuit breaker open, Spinnaker Orca queue backing up
NETWORK AND SECURITY
SpeedTest
2 RULESSpeedTest Slow Internet Download, SpeedTest Slow Internet Upload
SSL/TLS
4 RULESSSL certificate probe failed, SSL certificate revoked, SSL certificate OCSP status unknown, SSL certificate expiry (< 7 days)
cert-manager
4 RULESCert-Manager absent, Cert-Manager certificate not ready, Cert-Manager hitting ACME rate limits, Cert-Manager certificate expiring soon
Juniper
3 RULESJuniper switch down, Juniper critical Bandwidth Usage 1GiB, Juniper warning Bandwidth Usage 1GiB
CoreDNS
9 RULESCoreDNS Forward Plugin SERVFAIL Rate Critical, CoreDNS Panic Count, CoreDNS High Query Latency, CoreDNS SERVFAIL Error Rate Critical, CoreDNS Forward Plugin High Latency
Freeswitch
3 RULESFreeswitch down, Freeswitch Sessions Critical, Freeswitch Sessions Warning
sip-exporter
6 RULESSIP exporter processing channel saturated, SIP exporter kernel socket drops high, SIP exporter session establishment ratio degraded, SIP exporter RTP packet loss high, SIP exporter dialogs missing RTP high
Hashicorp Vault
7 RULESVault autopilot unhealthy, Vault sealed, Vault cluster health, Vault too many infinity tokens, Vault audit log write failures
Keycloak
6 RULESKeycloak no successful logins, Keycloak high login failure rate, Keycloak high token refresh error rate, Keycloak high code-to-token exchange error rate, Keycloak high registration failure rate
Cloudflare
5 RULESCloudflare load balancer pool unhealthy, Cloudflare http 5xx error rate, Cloudflare http 4xx error rate, Cloudflare high threat count, Cloudflare high request rate
SNMP
7 RULESSNMP target down, SNMP interface down, SNMP interface high bandwidth usage inbound, SNMP interface high bandwidth usage outbound, SNMP interface high inbound error rate
Cilium
31 RULESCilium agent unreachable nodes, Cilium agent unreachable health endpoints, Cilium operator exhausted IPAM IPs, Cilium agent high denied rate, Cilium agent conntrack table full
WireGuard
3 RULESWireGuard peer handshake never established, WireGuard peer handshake too old, WireGuard no traffic on peer
STORAGE
Ceph
26 RULESCeph OSD down (>= 10%), Ceph PG down, Ceph monitor quorum at risk, Ceph OSD Down, Ceph monitor down
ZFS
4 RULESZFS offline pool
OpenEBS
4 RULESOpenEBS LVM volume group missing physical volume, OpenEBS used pool capacity, OpenEBS LVM volume group capacity high, OpenEBS LVM thin pool capacity high
Minio
5 RULESMinio cluster erasure set quorum lost, Minio cluster disk offline, Minio node disk offline, Minio disk space usage, Minio KMS unavailable
CLOUD PROVIDERS
AWS CloudWatch
13 RULESAWS EC2 high CPU utilization, AWS RDS high CPU utilization, AWS RDS low free storage space, AWS RDS high database connections, AWS ALB unhealthy targets
Google Cloud Stackdriver
5 RULESStackdriver exporter scrape error, Stackdriver exporter slow scrape, Stackdriver exporter scrape errors increasing, Stackdriver exporter scrape stale, Stackdriver exporter high API calls
DigitalOcean
11 RULESDigitalOcean droplet down, DigitalOcean database down, DigitalOcean Kubernetes cluster down, DigitalOcean load balancer down, DigitalOcean account not active
Azure
3 RULESAzure API read rate limit approaching, Azure API write rate limit approaching, Azure exporter slow collection
OBSERVABILITY
Thanos
49 RULESThanos Compactor Multiple Running, Thanos Compactor Halted, Thanos Compactor High Compaction Failures, Thanos Compact Bucket High Operation Failures, Thanos Compact Has Not Run
Loki
6 RULESLoki compactor not running compaction, Loki request errors, Loki request panic, Loki request latency, Loki process too many restarts
Promtail
3 RULESPromtail request errors, Promtail request latency, Promtail file not being tailed
Cortex
23 RULESCortex not connected to Alertmanager, Cortex Alertmanager state replication failing, Cortex ruler missed evaluations, Cortex compactor has not uploaded blocks, Cortex notifications are being dropped
Grafana Tempo
28 RULESTempo live store unhealthy, Tempo metrics generator unhealthy, Tempo compactions failing, Tempo polls failing, Tempo tenant index failures
Grafana Mimir
55 RULESMimir compactor not running compaction, Mimir ruler missed evaluations, Mimir memory map areas too high, Mimir ingester reaching series limit critical, Mimir ingester reaching tenants limit critical
Grafana Alloy
7 RULESGrafana Alloy service down, Grafana Alloy OpenTelemetry exporter failing to send data, Grafana Alloy OpenTelemetry receiver refusing data, Grafana Alloy unhealthy components, Grafana Alloy slow component evaluations
OpenTelemetry Collector
15 RULESOpenTelemetry Collector down, OpenTelemetry Collector high memory usage, OpenTelemetry Collector receiver refused spans, OpenTelemetry Collector receiver refused metric points, OpenTelemetry Collector receiver refused log records
Jaeger
21 RULESJaeger storage completely unavailable, Jaeger no storage reads succeeding, Jaeger high storage error rate, Jaeger slow storage operations, Jaeger query service high error rate
OTHER
APC UPS
6 RULESAPC UPS low battery voltage, APC UPS high temperature, APC UPS Battery nearly empty, APC UPS Less than 15 Minutes of battery time remaining, APC UPS AC input outage
Graph Node
8 RULESProvider failed because net_version failed, Provider failed because get genesis failed, Provider failed because net_version timeout, Provider failed because get genesis timeout, Store connection very slow
LiteLLM
3 RULESLiteLLM provider spend over budget, LiteLLM proxy failed requests rate high, LiteLLM request latency p95 high