Watch metrics and scale up

The API reports a VM's power and nothing from inside the guest: no CPU, memory or disk figures. Run an exporter on each VM, read it on a schedule, and resize before a resource runs out. RAM and the root disk grow only while the VM is stopped, so set thresholds early enough that you choose when it stops.

1. Expose metrics on each VM

The Prometheus node exporter reports CPU, memory and every mounted filesystem, volumes included:

sudo apt install prometheus-node-exporter

It starts at once and serves :9100/metrics on every interface. Nothing in front of your VM filters that port, so bind it to where step 2 reads it from: loopback if you read it over SSH, the VM's mesh address if a metrics VM scrapes it. Set that in /etc/default/prometheus-node-exporter, then restart it:

ARGS="--web.listen-address=127.0.0.1:9100"      # or the mesh address

If a metrics VM has to scrape the public address, leave the default and allow 9100 from that VM's address alone: Harden a VM.

2. Poll them

One or two VMs: read the exporter over SSH. Nothing more to run or open.

ssh ubuntu@web1.example.com curl -s localhost:9100/metrics

This is a reading of now. It keeps no history, so it cannot tell a spike from a trend, and nothing reads it while you are not running.

More machines, or you need history: a metrics VM. One small VM runs Prometheus, which scrapes every exporter on an interval and keeps the series, and Grafana for dashboards your human can open.

POST /v1/vms {"name":"metrics1","ram_mb":2048,"disk_gb":40,
              "image":"ubuntu-lts","ssh_keys":[...]}   + Idempotency-Key
POST /v1/vms/{id}/actions {"type":"start"}
sudo apt install prometheus                  # in the guest

3. What to watch

ResourceExpressionScale when
CPU1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))above 0.8 for an hour
RAMnode_memory_MemAvailable_bytes / node_memory_MemTotal_bytesbelow 0.15
RAMincrease(node_vmstat_oom_kill[1h])above 0
Root disknode_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}below 0.2
Volumethe same, with the volume's mountpointbelow 0.2

Without Prometheus, top -bn1, free -m and df -h over SSH cover the same resources.

A disk that fills fast passes a fixed percentage too late. With history, alert on the trend instead: predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 3 * 86400) < 0 is true when the current rate fills it within three days.

4. Scale up

How each works: sizing, [growing the root disk](../vm.md#growing-the-root-disk), volumes. A volume grows while the VM runs; RAM and the root disk need it stopped, so:

When a stop is not acceptable

A VM whose data is not on volumes can be replaced without one: build a larger VM beside it and switch DNS (Recreate a VM). A volume moves only from a stopped VM, so that does not help once one is attached.

As a standing setup, run the service on two VMs behind a load balancer, and resize one at a time: take it out of the pool, stop, resize, start, put it back, then do the other. One VM down does not take the service down.

What goes wrong

What it costs

A metrics VM bills like any other VM. Scrapes between your VMs are metered as egress, so a longer scrape interval costs less.

A larger size bills from the resize. RAM can be sized back down later; a root disk or a volume cannot, so grow disks in steps you will use. Sizing down: Watch your credits and scale down.