On this page

Alerting rules

A Prometheus alerting rules file for Pulse, pulse-alerts.yml. Eleven rules across tick health, engine warnings, log errors, endpoint availability, the worldgen queue and per-mod attribution, each carrying a severity label (warning or critical) and an annotation that says what to check, not just what happened. The thresholds are calibrated against the engine's own numbers: 30 TPS is the game's nominal tick rate, 500 ms is the engine's own overload cutoff, 90%/100% of DieAboveMemoryUsageMb are the engine's own memory thresholds. The comments in the file say where each one comes from.

It is not wired in automatically. Add it to rule_files in your prometheus.yml:

yaml
rule_files:
  - /etc/pulse/pulse-alerts.yml

scrape_configs:
  - job_name: vintagestory
    static_configs:
      - targets: ["127.0.0.1:9464"]

If you're running the container from contrib/grafana, mount this directory alongside it and point rule_files at the mounted path:

sh
docker run -d --rm --name pulse-prom --network host \
  -v "$PWD/contrib/grafana:/etc/pulse" \
  -v "$PWD/contrib/alerts:/etc/pulse-alerts" \
  prom/prometheus --config.file=/etc/pulse/prometheus.yml --web.listen-address=127.0.0.1:9090

with rule_files: [/etc/pulse-alerts/pulse-alerts.yml] added to contrib/grafana/prometheus.yml (left out of that file by default, so the grafana kit stays alerting-free until you ask for it).

That bind is loopback only, same reasoning as contrib/grafana/README.md: --network host puts the container directly on the machine's own network, and Prometheus has no login of its own, so without it, its own UI and API, including /alerts and every metric it holds, would be reachable from anywhere that can reach this machine at all, not just from it. To check /alerts from another computer, tunnel over SSH instead of widening that bind:

sh
ssh -L 9090:127.0.0.1:9090 user@your-server

then open http://localhost:9090/alerts on your own computer.

Prometheus only reads rule_files at startup or on a reload: send it SIGHUP, hit /-/reload if it was started with --web.enable-lifecycle, or restart the container.

A firing alert only gets you as far as Prometheus's own /alerts page. To have it actually notify anyone, point the alerting: block in prometheus.yml at an Alertmanager and configure routing there; that setup is entirely yours; nothing here assumes a particular chat tool or paging service.

Three things worth knowing before you rely on these:

  • PulseEndpointDown matches up{job="vintagestory"}, the job name contrib/grafana's own prometheus.yml uses. If your scrape job is named differently, change that one label.
  • PulseTickSaturationHigh reads pulse_server_tick_busy_seconds, one of the families that only exists when Pulse's engine probe resolved successfully (see the main README's "Degraded mode" section). If that cast ever fails on a game update, the family disappears from /metrics and this rule simply has no data to evaluate; it goes quiet, not green. PulseTickOverrunsHigh looks like it belongs in the same boat but does not: it only reads the tick histogram and pulse_server_tick_budget_seconds, both public API metrics, so it keeps working in degraded mode the same as the tick rate and log/worldgen rules.
  • PulseModHoggingTick reads pulse_server_tick_busy_seconds too, so it is quiet in degraded mode for the same reason. It also needs attribution switched on (see the main README's "Attribution" section) for pulse_mod_tick_share. That family is an observable gauge tied to the same duty cycle, so it disappears from /metrics the moment attribution stops, whether that is /pulse attribution off, a reload with Enabled false, or the duty cycle giving up on its own; it no longer sits there at the last burst's values looking like a real number. The rule keeps its own guard regardless, increase(pulse_attribution_ticks_total[5m]) > 0: that counter only moves when a burst actually completes, a more direct signal that attribution is doing real work than the family merely being present, and one that still protects the rule if a future change ever lets the gauge report something without a completed burst behind it.

Validate the file after editing it with the same promtool container used to write it:

sh
docker run --rm -v "$PWD/contrib/alerts:/a" --entrypoint promtool prom/prometheus check rules /a/pulse-alerts.yml

That only parses the PromQL; it does not run it, so a rule can pass check rules and still never fire, for instance an and or an arithmetic operator whose two sides carry different labels and so never match anything. pulse-alerts.test.yml catches that class of mistake by running every rule against synthetic data, one case where it should fire and one where it should stay quiet:

sh
docker run --rm -v "$PWD/contrib/alerts:/a" --entrypoint promtool prom/prometheus test rules /a/pulse-alerts.test.yml

Run both after touching this file. Adding a rule without adding its two cases here is how the next silent one gets through.

The rules #

download pulse-alerts.yml

pulse-tick #

PulseTickRateLow warning for 5m

Tick rate below 27 TPS for 5m (target 30).

expr
rate(pulse_server_ticks_total[2m]) < 27

Tick rate has been under 27 TPS for 5 minutes; check pulse_server_tick_busy_seconds and the worldgen queue for what's eating the budget before it gets worse.

Reads: pulse_server_ticks_total

Why this rule, from the comments of the file
rate() over the tick counter is TPS. Two-tier: 27 is "visibly behind", 20 is "players are
feeling this". Both use the same alert name so Alertmanager groups them; a very sick
server fires both at once, which is expected, not a bug.

PulseTickRateLow CRITICAL for 5m

Tick rate below 20 TPS for 5m.

expr
rate(pulse_server_ticks_total[2m]) < 20

Tick rate has been under 20 TPS for 5 minutes, deep into player-visible lag; this needs attention now, not a wait-and-see.

Reads: pulse_server_ticks_total

Why this rule, from the comments of the file
rate() over the tick counter is TPS. Two-tier: 27 is "visibly behind", 20 is "players are
feeling this". Both use the same alert name so Alertmanager groups them; a very sick
server fires both at once, which is expected, not a bug.

PulseTickSaturationHigh warning for 5m

Tick busy time over 80% of budget for 5m.

expr
pulse_server_tick_busy_seconds / pulse_server_tick_budget_seconds > 0.8

Average tick busy time has been over 80% of the configured budget for 5 minutes; find what's using the headroom before ticks start missing budget outright.

Reads: pulse_server_tick_busy_seconds engine, pulse_server_tick_budget_seconds

Why this rule, from the comments of the file
Busy time over budget is the leading indicator, before rate() over the tick counter even
moves: a server can hold 30 TPS while spending most of its budget getting there. Only
meaningful when the engine-probe families are being served (the guarded cast in
Pulse/EngineProbe.cs succeeded) -- if it failed, pulse_server_tick_busy_seconds is simply
absent from /metrics and this rule has no data to evaluate, silently not firing either
way. Absence of this alert is not proof of headroom; check /metrics for the family before
trusting silence here.

PulseTickOverrunsHigh warning for 5m

p99 tick time over 3x budget for 5m.

expr
histogram_quantile(0.99, rate(pulse_server_tick_seconds_bucket[5m])) > 3 * pulse_server_tick_budget_seconds

The slowest 1% of ticks have been running past three times the tick budget for 5 minutes; look for worldgen, an autosave, or a mod spiking before this turns into a sustained TPS drop.

Reads: pulse_server_tick_seconds, pulse_server_tick_budget_seconds

Why this rule, from the comments of the file
Two ways to phrase "ticks are overrunning": count buckets past a fixed 50ms line, or take
a quantile against the live budget gauge. Went with the quantile: it tracks
pulse_server_tick_budget_seconds, so it stays correct if an operator retunes the tick rate
with /serverconfig, where a hardcoded 50ms cutoff would quietly go stale. It also doesn't
depend on 0.05 staying a bucket boundary in Pulse/PulseModSystem.cs's TickBuckets array.
The tradeoff is bucket-interpolation error from only ten finite buckets, which is fine
at warning granularity -- this is "is the tail bad", not a latency SLO.

pulse-engine #

PulseEngineWarning warning

Engine {{ $labels.kind }} warning firing.

expr
increase(pulse_engine_warnings_total{kind!="memory"}[10m]) > 0

The engine logged a {{ $labels.kind }} warning in the last 10 minutes; check the server log around that time for what triggered it.

Reads: pulse_engine_warnings_total

Why this rule, from the comments of the file
increase() > 0 over a 10m window already is the debounce; a single stray warning ten
minutes ago still counts, which is the point. Same alert name for both severities, kind
excluded from the warning leg so memory doesn't double-fire under its own critical rule
below.

PulseEngineWarning CRITICAL

Engine memory warning firing.

expr
increase(pulse_engine_warnings_total{kind="memory"}[10m]) > 0

The engine crossed 90% of its memory ceiling in the last 10 minutes and self-terminates at 100%; raise DieAboveMemoryUsageMb or free heap now, not after it restarts itself.

Reads: pulse_engine_warnings_total

Why this rule, from the comments of the file
DieAboveMemoryUsageMb: the "memory" kind is the engine crossing 90% of its ceiling, and
the engine kills itself outright at 100%. That gap is small enough to treat as critical
from the first warning rather than waiting for a repeat.

pulse-log #

PulseLogErrorBurst warning

More than 5 error log entries in 10m.

expr
increase(pulse_log_entries_total{level="error"}[10m]) > 5

More than 5 error-level log entries landed in the last 10 minutes; tail the server log for the recurring one rather than a one-off.

Reads: pulse_log_entries_total

Why this rule, from the comments of the file
A single stray error is normal noise (a malformed packet, a one-off mod hiccup); a burst
is one thing failing repeatedly. 5 in 10 minutes is a judgement call, not a measured
value like the tick numbers above -- tune it if your server's baseline error rate differs.

PulseLogFatal CRITICAL

Fatal log entry.

expr
increase(pulse_log_entries_total{level="fatal"}[10m]) > 0

A fatal log entry landed in the last 10 minutes; read the server log now. Error and fatal entries count the same way toward the engine's own DieAboveErrorCount self-shutdown threshold.

Reads: pulse_log_entries_total

Why this rule, from the comments of the file
Fatal is the engine's own top log severity, so even one is worth a page, not a threshold
like the burst rule above. It is not the only level that can shut the server down though:
VintagestoryLib.dll's ServerSystemMonitor.OnEntryAdded counts error and fatal entries the
same way toward the engine's own DieAboveErrorCount threshold, so a sustained stream of
errors is also a path to an unplanned restart, covered above by the burst rule.

pulse-availability #

PulseEndpointDown CRITICAL for 2m

Pulse scrape target down for 2m.

expr
up{job="vintagestory"} == 0

Prometheus has not been able to scrape the vintagestory job for 2 minutes; confirm the game server process is up and Pulse's endpoint is still bound before assuming it's just network flakiness.

Why this rule, from the comments of the file
`up` is Prometheus's own synthetic series, one per scrape target; job name here matches
contrib/grafana/prometheus.yml's `job_name: vintagestory`. Change the job label if yours
scrapes Pulse under a different name.

pulse-worldgen #

PulseWorldgenQueueStuck warning for 15m

Worldgen queue over 500 columns for 15m.

expr
pulse_worldgen_queue_columns > 500

The worldgen queue has held more than 500 pending columns for 15 minutes without draining; check the worldgen thread isn't stalled and that disk I/O for chunk writes isn't backed up.

Reads: pulse_worldgen_queue_columns

Why this rule, from the comments of the file
A queue that fills during a player exploration burst and drains afterwards is the queue
doing its job. 15m of staying above 500 is long enough that draining stopped, not that it
is merely busy.

pulse-attribution #

PulseModHoggingTick warning for 10m

Mod {{ $labels.modid }} holding over half the tick for 10m while the server is loaded.

expr
(pulse_mod_tick_share{modid!~"engine|unattributed"} > 0.5) and ignoring(modid) (pulse_server_tick_busy_seconds / pulse_server_tick_budget_seconds > 0.8) and ignoring(modid) (increase(pulse_attribution_ticks_total[5m]) > 0)

{{ $labels.modid }} has held more than 50% of the profiled main-thread tick for 10 minutes while tick busy time stayed over 80% of budget; open the dashboard's attribution row for the full breakdown and look at that mod first. Attribution is sampled (about one tick in thirty at the default duty cycle) and only sees listeners and behaviours, not broadcast event handlers, so treat this as where to look first, not a full accounting.

Reads: pulse_mod_tick_share attribution, pulse_server_tick_busy_seconds engine, pulse_server_tick_budget_seconds, pulse_attribution_ticks_total attribution

Why this rule, from the comments of the file
Needs attribution switched on (README.md's Attribution section: the config block or
`/pulse attribution on`). pulse_mod_tick_share is an observable gauge that disappears from
/metrics the moment attribution stops, whether that is switching it off, a reload with
Enabled false, or the duty cycle giving up on its own, so its absence alone already means
nothing is being measured. The increase(pulse_attribution_ticks_total) term below stays as
a second, independent guard: it only moves when a burst actually completes, which is a
more direct proof of live measurement than the share's mere presence, and it is what would
still catch a future change that let the share report again with nothing behind it.

Two conditions on top of that, both required: pulse_mod_tick_share sums to 1 across every
modid including engine and unattributed, so one mod clearing 0.5 already outweighs
everything else on the server combined, and the load gate reuses PulseTickSaturationHigh's
own 80% busy-over-budget threshold, on the same reasoning, a mod owning half the tick is not
worth paging on while the server is coasting at 20% of budget. 10m is several burst
refreshes at the default duty cycle (10 BurstTicks / 10s IntervalSeconds samples roughly
once every 10 seconds), long enough that this is a sustained hog and not one unlucky sample
landing mid-spike, which the README's "What it cannot see" warns can happen on a spiky
server. The excluded modid is not always a third-party mod either: the base game itself
ships as several mods (survival, essentials, game among them), so on an entity-heavy
vanilla server with nothing else installed this can still name one of those.

"and" defaults to matching on the full label set, and the left side carries modid while
neither the load gate nor the ticks guard does, so without "ignoring(modid)" on both no
side ever matches anything and this alert can never fire. Ignoring modid rather than naming
"on(instance, job)" keeps whatever else a target carries part of the match too, for example
a label a scrape job's relabel_configs adds, instead of silently discarding it the way
naming just instance and job would.

Source: contrib/alerts/README.md, contrib/alerts/pulse-alerts.yml