KLYRN VM / Blog

Your whole estate in Prometheus and Grafana

One scrape job turns every figure on the KLYRN VM dashboard into a time series: machines, node capacity, address pools, tasks and alerts.

· 1 min read · The KLYRN team

IntegrationsMonitoring

An operator who already runs Prometheus should not have to write a script to find out that a node is filling up. KLYRN VM serves its estate at /api/v1/metrics in the Prometheus text format.

Connecting it

Create an API token for an administrator with the Read everything scope, and give it to Prometheus as a bearer token:

scrape_configs:
            - job_name: klyrn-vm
              scheme: https
              metrics_path: /api/v1/metrics
              authorization:
                type: Bearer
                credentials_file: /etc/prometheus/klyrn-vm.token
              static_configs:
                - targets: ['panel.example.com:8443']

The endpoint is for administrators only, because it names every node and counts every machine. That is enforced by the same guard as the rest of the provider's side of the panel, and a test fails if a customer can read it.

What is in it

  • klyrn_vms{state}: running, stopped, needs attention, creating, unknown.
  • klyrn_node_up per node, and memory, storage and vCPU capacity for every managed node: installed, promised to guests and in use.
  • klyrn_ipv4_free and klyrn_ipv4_total across every pool.
  • The task queue: running, failed in the last day, waiting for a node.
  • klyrn_alerts_open{severity}.

Every value is read by the same code as the Overview page, so a Grafana panel and the dashboard cannot disagree.

Alerts worth starting with

- alert: KlyrnNodeDown
            expr: klyrn_node_up == 0
            for: 5m
          - alert: KlyrnIPv4Low
            expr: klyrn_ipv4_free / klyrn_ipv4_total < 0.1
            for: 30m
          - alert: KlyrnStorageNearlyFull
            expr: klyrn_node_storage_used_bytes / klyrn_node_storage_bytes > 0.85
            for: 30m