Your whole estate in Prometheus and Grafana
One scrape job turns every figure on the KLYRN VM dashboard into a time series: machines, node capacity, address pools, tasks and alerts.
An operator who already runs Prometheus should not have to write a script to find out that a node is filling up. KLYRN VM serves its estate at /api/v1/metrics in the Prometheus text format.
Connecting it
Create an API token for an administrator with the Read everything scope, and give it to Prometheus as a bearer token:
scrape_configs:
- job_name: klyrn-vm
scheme: https
metrics_path: /api/v1/metrics
authorization:
type: Bearer
credentials_file: /etc/prometheus/klyrn-vm.token
static_configs:
- targets: ['panel.example.com:8443']
The endpoint is for administrators only, because it names every node and counts every machine. That is enforced by the same guard as the rest of the provider's side of the panel, and a test fails if a customer can read it.
What is in it
klyrn_vms{state}: running, stopped, needs attention, creating, unknown.klyrn_node_upper node, and memory, storage and vCPU capacity for every managed node: installed, promised to guests and in use.klyrn_ipv4_freeandklyrn_ipv4_totalacross every pool.- The task queue: running, failed in the last day, waiting for a node.
klyrn_alerts_open{severity}.
Every value is read by the same code as the Overview page, so a Grafana panel and the dashboard cannot disagree.
Alerts worth starting with
- alert: KlyrnNodeDown
expr: klyrn_node_up == 0
for: 5m
- alert: KlyrnIPv4Low
expr: klyrn_ipv4_free / klyrn_ipv4_total < 0.1
for: 30m
- alert: KlyrnStorageNearlyFull
expr: klyrn_node_storage_used_bytes / klyrn_node_storage_bytes > 0.85
for: 30m