The API reports a VM's power and nothing from inside the guest: no CPU, memory or disk figures. Run an exporter on each VM, read it on a schedule, and resize before a resource runs out. RAM and the root disk grow only while the VM is stopped, so set thresholds early enough that you choose when it stops.
The Prometheus node exporter reports CPU, memory and every mounted filesystem, volumes included:
sudo apt install prometheus-node-exporter
It starts at once and serves :9100/metrics on every interface. Nothing in
front of your VM filters that port, so bind it to where step 2 reads it from:
loopback if you read it over SSH, the VM's mesh address
if a metrics VM scrapes it. Set that in
/etc/default/prometheus-node-exporter, then restart it:
ARGS="--web.listen-address=127.0.0.1:9100" # or the mesh address
If a metrics VM has to scrape the public address, leave the default and allow 9100 from that VM's address alone: Harden a VM.
One or two VMs: read the exporter over SSH. Nothing more to run or open.
ssh ubuntu@web1.example.com curl -s localhost:9100/metrics
This is a reading of now. It keeps no history, so it cannot tell a spike from a trend, and nothing reads it while you are not running.
More machines, or you need history: a metrics VM. One small VM runs Prometheus, which scrapes every exporter on an interval and keeps the series, and Grafana for dashboards your human can open.
POST /v1/vms {"name":"metrics1","ram_mb":2048,"disk_gb":40,
"image":"ubuntu-lts","ssh_keys":[...]} + Idempotency-Key
POST /v1/vms/{id}/actions {"type":"start"}
sudo apt install prometheus # in the guest
--storage.tsdb.retention.time so it cannot fill the disk
(which disk).GET /api/v1/query?query=<expr>
answers one expression for every machine at once.| Resource | Expression | Scale when |
|---|---|---|
| CPU | 1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) | above 0.8 for an hour |
| RAM | node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes | below 0.15 |
| RAM | increase(node_vmstat_oom_kill[1h]) | above 0 |
| Root disk | node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"} | below 0.2 |
| Volume | the same, with the volume's mountpoint | below 0.2 |
Without Prometheus, top -bn1, free -m and df -h over SSH cover the same
resources.
A disk that fills fast passes a fixed percentage too late. With history,
alert on the trend instead:
predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 3 * 86400) < 0
is true when the current rate fills it within three days.
ram_mb. vCPUs follow it.disk_gb, or move the data that keeps growing to a
volume.size_gb.How each works: sizing, [growing the root disk](../vm.md#growing-the-root-disk), volumes. A volume grows while the VM runs; RAM and the root disk need it stopped, so:
GET /v1/quota before you stop the VM. A resize over a cap is
refused, and the stop would have bought nothing.ram_mb and disk_gb in one resize, so the VM stops once.A VM whose data is not on volumes can be replaced without one: build a larger VM beside it and switch DNS (Recreate a VM). A volume moves only from a stopped VM, so that does not help once one is attached.
As a standing setup, run the service on two VMs behind a load balancer, and resize one at a time: take it out of the pool, stop, resize, start, put it back, then do the other. One VM down does not take the service down.
quota_exceeded (403) โ the new size is over a project cap. Nothing was
applied. Ask for a raise with a ticket
(quotas).no_capacity_available (409) โ no room for that much RAM right now. Nothing
was applied; start the VM at its old size and try again later. Raising the
quota does not help.A metrics VM bills like any other VM. Scrapes between your VMs are metered as egress, so a longer scrape interval costs less.
A larger size bills from the resize. RAM can be sized back down later; a root disk or a volume cannot, so grow disks in steps you will use. Sizing down: Watch your credits and scale down.