# Watch metrics and scale up

The API reports a VM's power and nothing from inside the guest: no CPU,
memory or disk figures. Run an exporter on each VM, read it on a schedule, and
resize before a resource runs out. RAM and the root disk grow only while the
VM is stopped, so set thresholds early enough that you choose when it stops.

## 1. Expose metrics on each VM

The Prometheus node exporter reports CPU, memory and every mounted
filesystem, volumes included:

```
sudo apt install prometheus-node-exporter
```

It starts at once and serves `:9100/metrics` on every interface. Nothing in
front of your VM filters that port, so bind it to where step 2 reads it from:
loopback if you read it over SSH, the VM's [mesh](wireguard-mesh.md) address
if a metrics VM scrapes it. Set that in
`/etc/default/prometheus-node-exporter`, then restart it:

```
ARGS="--web.listen-address=127.0.0.1:9100"      # or the mesh address
```

If a metrics VM has to scrape the public address, leave the default and allow
9100 from that VM's address alone: [Harden a VM](harden-your-vm.md).

## 2. Poll them

**One or two VMs: read the exporter over SSH.** Nothing more to run or open.

```
ssh ubuntu@web1.example.com curl -s localhost:9100/metrics
```

This is a reading of now. It keeps no history, so it cannot tell a spike from
a trend, and nothing reads it while you are not running.

**More machines, or you need history: a metrics VM.** One small VM runs
Prometheus, which scrapes every exporter on an interval and keeps the series,
and Grafana for dashboards your human can open.

```
POST /v1/vms {"name":"metrics1","ram_mb":2048,"disk_gb":40,
              "image":"ubuntu-lts","ssh_keys":[...]}   + Idempotency-Key
POST /v1/vms/{id}/actions {"type":"start"}
sudo apt install prometheus                  # in the guest
```

- Keep Prometheus's data on the root disk, and cap it with
  `--storage.tsdb.retention.time` so it cannot fill the disk
  ([which disk](../vm.md#two-kinds-of-disk)).
- Prometheus's port, 9090, has no login. Bind it to loopback or the mesh.
  Grafana comes from its own apt repository; give it TLS and a password before
  your human opens it from the Internet.
- Poll Prometheus rather than each exporter: `GET /api/v1/query?query=<expr>`
  answers one expression for every machine at once.

## 3. What to watch

| Resource | Expression | Scale when |
|---|---|---|
| CPU | `1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))` | above 0.8 for an hour |
| RAM | `node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes` | below 0.15 |
| RAM | `increase(node_vmstat_oom_kill[1h])` | above 0 |
| Root disk | `node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}` | below 0.2 |
| Volume | the same, with the volume's mountpoint | below 0.2 |

Without Prometheus, `top -bn1`, `free -m` and `df -h` over SSH cover the same
resources.

A disk that fills fast passes a fixed percentage too late. With history,
alert on the trend instead:
`predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[6h], 3 * 86400) < 0`
is true when the current rate fills it within three days.

## 4. Scale up

- **CPU or RAM**: a larger `ram_mb`. vCPUs follow it.
- **Root disk**: a larger `disk_gb`, or move the data that keeps growing to a
  volume.
- **Volume**: a larger `size_gb`.

How each works: [sizing](../vm.md#sizing), [growing the root
disk](../vm.md#growing-the-root-disk), [volumes](../vm.md#volumes). A volume
grows while the VM runs; RAM and the root disk need it stopped, so:

- Read `GET /v1/quota` before you stop the VM. A resize over a cap is
  refused, and the stop would have bought nothing.
- Put `ram_mb` and `disk_gb` in one resize, so the VM stops once.

### When a stop is not acceptable

A VM whose data is not on volumes can be replaced without one: build a
larger VM beside it and switch DNS ([Recreate a VM](recreate-vm.md)). A volume
moves only from a stopped VM, so that does not help once one is attached.

As a standing setup, run the service on two VMs behind a load balancer, and
resize one at a time: take it out of the pool, stop, resize, start, put it
back, then do the other. One VM down does not take the service down.

- The load balancer is yours to run: HAProxy or nginx on a small VM of its
  own, or your CDN's load balancing if you [use one](use-a-cdn.md). Point the
  DNS name at it, not at either VM.
- Size each VM to carry the full load alone, since it does while the other is
  stopped.
- This fits a stateless service. A database on both needs its own
  replication, and files the service writes have to live where both reach
  them.
- Both VMs bill all the time: RAM, root disk and address each.

## What goes wrong

- `quota_exceeded` (403) — the new size is over a project cap. Nothing was
  applied. Ask for a raise with a ticket
  ([quotas](../account.md#default-quotas)).
- `no_capacity_available` (409) — no room for that much RAM right now. Nothing
  was applied; start the VM at its old size and try again later. Raising the
  quota does not help.
- **The metric still reads full after a resize.** The filesystem has not
  grown yet: a volume's grows when you grow it in the guest, and the root's at
  the next start ([growing the root disk](../vm.md#growing-the-root-disk)).

## What it costs

A metrics VM bills like any other VM. Scrapes between your VMs are metered as
egress, so a longer scrape interval costs less.

A larger size bills from the resize. RAM can be sized back down later; a root
disk or a volume cannot, so grow disks in steps you will use. Sizing down:
[Watch your credits and scale down](watch-your-credits.md#scale-down).
