Browse Library

BASIC RESOURCE MONITORING

Prometheus

Prometheus self-monitoring

36 RULESPrometheus target missing, Prometheus all targets missing, Prometheus target missing with warmup time, Prometheus not connected to alertmanager, Prometheus job missing

Placeholder icon

Host and hardware

41 RULESHost high CPU load, Host CPU steal noisy neighbor, Host CPU high iowait, Host CPU load saturation, Host CPU is underutilized

Placeholder icon

S.M.A.R.T Device Monitoring

9 RULESSMART critical warning, SMART device temperature critical, SMART device temperature over trip value, SMART device temperature warning, SMART device temperature nearing trip value

Placeholder icon

IPMI

17 RULESIPMI chassis power off, IPMI collector down, IPMI SEL almost full, IPMI temperature sensor critical, IPMI fan speed sensor critical

Docker

Docker containers

10 RULESContainer killed, Container absent, Container High CPU utilization, Container high throttle rate, Container high low change CPU usage

Placeholder icon

Blackbox

10 RULESBlackbox probe failed, Blackbox probe HTTP failure, Blackbox probe low uptime, Blackbox slow probe, Blackbox probe slow HTTP

Placeholder icon

Windows Server

8 RULESWindows Server CPU Usage, Windows Server memory Usage, Windows Server disk Space Usage, Windows Server disk drive Status, Windows Server NTP client delay

VMware

VMware

4 RULESVirtual Machine Memory Critical, Virtual Machine Memory Warning, High Number of Snapshots, Outdated Snapshots

Proxmox

Proxmox VE

9 RULESPVE node down, PVE cluster not quorate, PVE VM/CT down, PVE high CPU usage, PVE high memory usage

Netdata

Netdata

9 RULESNetdata high cpu usage, Netdata CPU steal noisy neighbor, Netdata high memory usage, Netdata low disk space, Netdata predicted disk full

Placeholder icon

eBPF

3 RULESeBPF exporter no enabled configs, eBPF exporter program not attached, eBPF exporter decoder errors

Placeholder icon

Process Exporter

10 RULESProcess exporter group down, Process exporter high CPU usage, Process exporter high memory usage, Process exporter high disk write IO, Process exporter file descriptors exhausted

Placeholder icon

Systemd

7 RULESSystemd unit tasks near limit, Systemd socket high connections, Systemd service crash looping, Systemd unit failed, Systemd unit inactive

DATABASES

MySQL

MySQL

18 RULESMySQL down, MySQL Slave IO thread not running, MySQL Slave SQL thread not running, MySQL too many connections (> 80%), MySQL high prepared statements utilization (> 80%)

PostgreSQL

PostgreSQL

25 RULESPostgresql down, Postgresql replication role changed, Postgresql bloat index high (> 80%), Postgresql bloat table high (> 80%), Postgresql long-running transaction

Placeholder icon

SQL Server

2 RULESSQL Server down, SQL Server deadlock

Oracle

Oracle Database

8 RULESOracle DB down, Oracle DB tablespace full (> 95%), Oracle DB tablespace reaching capacity (> 85%), Oracle DB sessions reaching limit (> 85%), Oracle DB processes reaching limit (> 85%)

Placeholder icon

Patroni

1 RULESPatroni has no Leader

Placeholder icon

PGBouncer

6 RULESPGBouncer max connections, PGBouncer server connection saturation critical, PGBouncer active connections, PGBouncer clients waiting for connection, PGBouncer server connection saturation warning

Redis

Redis

16 RULESRedis down, Redis missing master, Redis too many masters, Redis disconnected slaves, Redis replication broken

Placeholder icon

Memcached

9 RULESMemcached down, Memcached memory usage high (> 90%), Memcached connection limit approaching (> 95%), Memcached connection limit approaching (> 80%), Memcached out of memory errors

MongoDB

MongoDB

22 RULESMongoDB Down, Mongodb replica member unhealthy, MongoDB replica set has no primary, MongoDB replica set primary changed, MongoDB too many connections (Percona)

Elasticsearch

Elasticsearch

22 RULESElasticsearch Cluster Red, Elasticsearch Healthy Nodes, Elasticsearch Healthy Data Nodes, Elasticsearch High CPU Usage, Elasticsearch Heap Usage Too High

OpenSearch

OpenSearch

12 RULESOpenSearch host CPU usage high, OpenSearch process CPU usage high, OpenSearch high heap usage, OpenSearch disk high watermark reached, OpenSearch disk low watermark reached

Meilisearch

Meilisearch

2 RULESMeilisearch index is empty, Meilisearch http response time

Placeholder icon

Cassandra

30 RULESCassandra Node is unavailable, Cassandra connection timeouts total (Instaclustr), Cassandra storage exceptions (Instaclustr), Cassandra tombstone dump (Instaclustr), Cassandra client request unavailable write (Instaclustr)

ClickHouse

Clickhouse

23 RULESClickHouse node down, ClickHouse No Available Replicas, ClickHouse No Live Replicas, ClickHouse Interserver Connection Issues, ClickHouse ZooKeeper Connection Issues

Placeholder icon

CouchDB

24 RULESCouchDB node down, CouchDB atom memory usage critical, CouchDB open OS files critical, CouchDB file descriptors high, CouchDB open databases critical

Placeholder icon

Solr

10 RULESSolr shard has no leader, Solr high JVM heap usage, Solr disk space very low, Solr disk space low, Solr update errors

MESSAGE BROKERS

RabbitMQ

RabbitMQ

25 RULESRabbitMQ node down, RabbitMQ memory alarm active, RabbitMQ memory high, RabbitMQ disk space alarm active, RabbitMQ predicted low disk watermark

Placeholder icon

Zookeeper

4 RULESNo description

Placeholder icon

Kafka

5 RULESKafka topics replicas, Kafka consumer group lag, Kafka consumer group lag increasing

Placeholder icon

Pulsar

10 RULESPulsar high ledger disk usage, Pulsar subscription very high number of backlog entries, Pulsar topic very large backlog storage size, Pulsar high write latency, Pulsar read only bookies

Placeholder icon

Nats

13 RULESNats server down, Nats high CPU usage, Nats high memory usage, Nats high JetStream memory usage, Nats high JetStream store usage

PROXIES, LOAD BALANCERS AND SERVICE MESHES

NGINX

Nginx

4 RULESNginx high HTTP 4xx error rate, Nginx high HTTP 5xx error rate, Nginx latency high, Nginx internal metric errors

Apache

Apache

4 RULESApache down, Apache workers load, Apache response time too high, Apache restart

Placeholder icon

HaProxy

33 RULESHAProxy HTTP slowing down, HAProxy dropping logs, HAProxy backend healthcheck flapping, HAProxy server healthcheck flapping, HAProxy high HTTP 4xx error rate backend

Placeholder icon

Traefik

9 RULESTraefik service down, Traefik high HTTP 4xx error rate service, Traefik high HTTP 5xx error rate service, Traefik certificate expiring imminently, Traefik config reload failure

Caddy

Caddy

3 RULESCaddy Reverse Proxy Down, Caddy high HTTP 4xx error rate service, Caddy high HTTP 5xx error rate service

Placeholder icon

Envoy

20 RULESEnvoy server not live, Envoy high memory usage, Envoy global downstream connections overflowing, Envoy downstream connections overflowing, Envoy high downstream HTTP 5xx error rate

Linkerd

Linkerd

1 RULESLinkerd high error rate

Istio

Istio

11 RULESIstio control plane component down, Istio Pilot Duplicate Entry, Istio Kubernetes gateway availability drop, Istio Pilot high push error rate, Istio Mixer Prometheus dispatches low

RUNTIMES

Placeholder icon

PHP-FPM

4 RULESPHP-FPM max-children reached, PHP-FPM listen queue saturation, PHP-FPM busy process ratio high, PHP-FPM slow requests ratio

Placeholder icon

JVM

12 RULESJVM compilation time spike, JVM memory filling up, JVM non-heap memory filling up, JVM old gen GC frequency, JVM file descriptors exhaustion

Placeholder icon

Golang

10 RULESGo GC CPU fraction high, Go memory usage high, Go heap objects count high, Go heap in-use growing, Go memory leak

Ruby

Ruby

5 RULESRuby heap live slots high, Ruby heap free slots high, Ruby major GC rate high, Ruby RSS high, Ruby allocated objects spike

Python

Python

5 RULESPython GC objects uncollectable, Python GC generation 2 collections high, Python virtual memory high, Python file descriptors exhaustion, Python GC collections high

Sidekiq

Sidekiq

2 RULESSidekiq scheduling latency too high, Sidekiq queue size

DATA ENGINEERING

Apache

Apache Flink

12 RULESFlink job is not running, Flink TaskManager heap memory high, Flink JobManager heap memory high, Flink no TaskManagers registered, Flink all task slots used

Apache

Apache Spark

8 RULESSpark worker memory exhausted, Spark executor high disk spill, Spark no alive workers, Spark executor all tasks failing, Spark too many waiting apps

Placeholder icon

Hadoop

15 RULESHadoop Name Node Down, Hadoop Resource Manager Down, Hadoop HDFS Missing Blocks, Hadoop Node Manager CPU High, Hadoop Resource Manager vCore Usage High

ORCHESTRATORS

Kubernetes

Kubernetes

46 RULESKubernetes StatefulSet down, Kubernetes Container oom killer, Kubernetes CronJob too long, Kubernetes Node memory pressure, Kubernetes Node disk pressure

Nomad

Nomad

4 RULESNomad job failed, Nomad job lost, Nomad job queued, Nomad blocked evaluation

Consul

Consul

5 RULESConsul missing master node, Consul agent unhealthy, Consul service healthcheck failed, Consul has no raft leader, Consul exporter query failure

etcd

Etcd

19 RULESEtcd no Leader, Etcd insufficient members for quorum, Etcd high fsync durations critical, Etcd high fsync durations, Etcd high commit durations

OpenStack

OpenStack

20 RULESOpenStack exporter down, OpenStack Nova agent down, OpenStack Neutron agent down, OpenStack Cinder agent down, OpenStack hypervisor high vCPU usage

CI/CD

Jenkins

Jenkins

8 RULESJenkins node offline, Jenkins no node online, Jenkins healthcheck, Jenkins builds health score, Jenkins outdated plugins

Placeholder icon

ArgoCD

9 RULESArgoCD service not synced, ArgoCD service unhealthy, ArgoCD application sync failed, ArgoCD cluster connection error, ArgoCD git fetch failures

Placeholder icon

FluxCD

4 RULESFlux Kustomization Failure, Flux HelmRelease Failure, Flux Source Issue, Flux Image Issue

GitLab

GitLab CI

32 RULESGitLab Puma workers not running, GitLab high memory usage, GitLab Ruby heap fragmentation, GitLab Puma high queued connections, GitLab database connection pool saturation

Spinnaker

Spinnaker

12 RULESSpinnaker thread pool exhaustion, Spinnaker dead messages, Spinnaker polling monitor items over threshold, Spinnaker circuit breaker open, Spinnaker Orca queue backing up

NETWORK AND SECURITY

Speedtest

SpeedTest

2 RULESSpeedTest Slow Internet Download, SpeedTest Slow Internet Upload

Placeholder icon

SSL/TLS

4 RULESSSL certificate probe failed, SSL certificate revoked, SSL certificate OCSP status unknown, SSL certificate expiry (< 7 days)

Placeholder icon

cert-manager

4 RULESCert-Manager absent, Cert-Manager certificate not ready, Cert-Manager hitting ACME rate limits, Cert-Manager certificate expiring soon

Placeholder icon

Juniper

3 RULESJuniper switch down, Juniper critical Bandwidth Usage 1GiB, Juniper warning Bandwidth Usage 1GiB

Placeholder icon

CoreDNS

9 RULESCoreDNS Forward Plugin SERVFAIL Rate Critical, CoreDNS Panic Count, CoreDNS High Query Latency, CoreDNS SERVFAIL Error Rate Critical, CoreDNS Forward Plugin High Latency

Placeholder icon

Freeswitch

3 RULESFreeswitch down, Freeswitch Sessions Critical, Freeswitch Sessions Warning

Placeholder icon

sip-exporter

6 RULESSIP exporter processing channel saturated, SIP exporter kernel socket drops high, SIP exporter session establishment ratio degraded, SIP exporter RTP packet loss high, SIP exporter dialogs missing RTP high

HashiCorp

Hashicorp Vault

7 RULESVault autopilot unhealthy, Vault sealed, Vault cluster health, Vault too many infinity tokens, Vault audit log write failures

Keycloak

Keycloak

6 RULESKeycloak no successful logins, Keycloak high login failure rate, Keycloak high token refresh error rate, Keycloak high code-to-token exchange error rate, Keycloak high registration failure rate

Cloudflare

Cloudflare

5 RULESCloudflare load balancer pool unhealthy, Cloudflare http 5xx error rate, Cloudflare http 4xx error rate, Cloudflare high threat count, Cloudflare high request rate

Placeholder icon

SNMP

7 RULESSNMP target down, SNMP interface down, SNMP interface high bandwidth usage inbound, SNMP interface high bandwidth usage outbound, SNMP interface high inbound error rate

Cilium

Cilium

31 RULESCilium agent unreachable nodes, Cilium agent unreachable health endpoints, Cilium operator exhausted IPAM IPs, Cilium agent high denied rate, Cilium agent conntrack table full

WireGuard

WireGuard

3 RULESWireGuard peer handshake never established, WireGuard peer handshake too old, WireGuard no traffic on peer

STORAGE

Ceph

Ceph

26 RULESCeph OSD down (>= 10%), Ceph PG down, Ceph monitor quorum at risk, Ceph OSD Down, Ceph monitor down

Placeholder icon

ZFS

4 RULESZFS offline pool

Placeholder icon

OpenEBS

4 RULESOpenEBS LVM volume group missing physical volume, OpenEBS used pool capacity, OpenEBS LVM volume group capacity high, OpenEBS LVM thin pool capacity high

MinIO

Minio

5 RULESMinio cluster erasure set quorum lost, Minio cluster disk offline, Minio node disk offline, Minio disk space usage, Minio KMS unavailable

CLOUD PROVIDERS

Placeholder icon

AWS CloudWatch

13 RULESAWS EC2 high CPU utilization, AWS RDS high CPU utilization, AWS RDS low free storage space, AWS RDS high database connections, AWS ALB unhealthy targets

Google

Google Cloud Stackdriver

5 RULESStackdriver exporter scrape error, Stackdriver exporter slow scrape, Stackdriver exporter scrape errors increasing, Stackdriver exporter scrape stale, Stackdriver exporter high API calls

DigitalOcean

DigitalOcean

11 RULESDigitalOcean droplet down, DigitalOcean database down, DigitalOcean Kubernetes cluster down, DigitalOcean load balancer down, DigitalOcean account not active

Placeholder icon

Azure

3 RULESAzure API read rate limit approaching, Azure API write rate limit approaching, Azure exporter slow collection

OBSERVABILITY

Thanos

Thanos

49 RULESThanos Compactor Multiple Running, Thanos Compactor Halted, Thanos Compactor High Compaction Failures, Thanos Compact Bucket High Operation Failures, Thanos Compact Has Not Run

Placeholder icon

Loki

6 RULESLoki compactor not running compaction, Loki request errors, Loki request panic, Loki request latency, Loki process too many restarts

Placeholder icon

Promtail

3 RULESPromtail request errors, Promtail request latency, Promtail file not being tailed

Placeholder icon

Cortex

23 RULESCortex not connected to Alertmanager, Cortex Alertmanager state replication failing, Cortex ruler missed evaluations, Cortex compactor has not uploaded blocks, Cortex notifications are being dropped

Grafana

Grafana Tempo

28 RULESTempo live store unhealthy, Tempo metrics generator unhealthy, Tempo compactions failing, Tempo polls failing, Tempo tenant index failures

Grafana

Grafana Mimir

55 RULESMimir compactor not running compaction, Mimir ruler missed evaluations, Mimir memory map areas too high, Mimir ingester reaching series limit critical, Mimir ingester reaching tenants limit critical

Grafana

Grafana Alloy

7 RULESGrafana Alloy service down, Grafana Alloy OpenTelemetry exporter failing to send data, Grafana Alloy OpenTelemetry receiver refusing data, Grafana Alloy unhealthy components, Grafana Alloy slow component evaluations

OpenTelemetry

OpenTelemetry Collector

15 RULESOpenTelemetry Collector down, OpenTelemetry Collector high memory usage, OpenTelemetry Collector receiver refused spans, OpenTelemetry Collector receiver refused metric points, OpenTelemetry Collector receiver refused log records

Jaeger

Jaeger

21 RULESJaeger storage completely unavailable, Jaeger no storage reads succeeding, Jaeger high storage error rate, Jaeger slow storage operations, Jaeger query service high error rate

OTHER

Placeholder icon

APC UPS

6 RULESAPC UPS low battery voltage, APC UPS high temperature, APC UPS Battery nearly empty, APC UPS Less than 15 Minutes of battery time remaining, APC UPS AC input outage

Placeholder icon

Graph Node

8 RULESProvider failed because net_version failed, Provider failed because get genesis failed, Provider failed because net_version timeout, Provider failed because get genesis timeout, Store connection very slow

Placeholder icon

LiteLLM

3 RULESLiteLLM provider spend over budget, LiteLLM proxy failed requests rate high, LiteLLM request latency p95 high