On this page
Alerting rules
A Prometheus alerting rules file for Pulse, pulse-alerts.yml. Eleven rules across tick health,
engine warnings, log errors, endpoint availability, the worldgen queue and per-mod attribution,
each carrying a severity label (warning or critical) and an annotation that says what to
check, not just what happened. The thresholds are calibrated against the engine's own numbers: 30
TPS is the game's nominal tick rate, 500 ms is the engine's own overload cutoff, 90%/100% of
DieAboveMemoryUsageMb are the engine's own memory thresholds. The comments in the file say
where each one comes from.
It is not wired in automatically. Add it to rule_files in your prometheus.yml:
rule_files:
- /etc/pulse/pulse-alerts.yml
scrape_configs:
- job_name: vintagestory
static_configs:
- targets: ["127.0.0.1:9464"]If you're running the container from contrib/grafana, mount this directory alongside it and
point rule_files at the mounted path:
docker run -d --rm --name pulse-prom --network host \
-v "$PWD/contrib/grafana:/etc/pulse" \
-v "$PWD/contrib/alerts:/etc/pulse-alerts" \
prom/prometheus --config.file=/etc/pulse/prometheus.yml --web.listen-address=127.0.0.1:9090with rule_files: [/etc/pulse-alerts/pulse-alerts.yml] added to contrib/grafana/prometheus.yml
(left out of that file by default, so the grafana kit stays alerting-free until you ask for it).
That bind is loopback only, same reasoning as contrib/grafana/README.md: --network host puts
the container directly on the machine's own network, and Prometheus has no login of its own, so
without it, its own UI and API, including /alerts and every metric it holds, would be
reachable from anywhere that can reach this machine at all, not just from it. To check /alerts
from another computer, tunnel over SSH instead of widening that bind:
ssh -L 9090:127.0.0.1:9090 user@your-serverthen open http://localhost:9090/alerts on your own computer.
Prometheus only reads rule_files at startup or on a reload: send it SIGHUP, hit
/-/reload if it was started with --web.enable-lifecycle, or restart the container.
A firing alert only gets you as far as Prometheus's own /alerts page. To have it actually
notify anyone, point the alerting: block in prometheus.yml at an Alertmanager and configure
routing there; that setup is entirely yours; nothing here assumes a particular chat tool or
paging service.
Three things worth knowing before you rely on these:
PulseEndpointDownmatchesup{job="vintagestory"}, the job namecontrib/grafana's ownprometheus.ymluses. If your scrape job is named differently, change that one label.PulseTickSaturationHighreadspulse_server_tick_busy_seconds, one of the families that only exists when Pulse's engine probe resolved successfully (see the main README's "Degraded mode" section). If that cast ever fails on a game update, the family disappears from/metricsand this rule simply has no data to evaluate; it goes quiet, not green.PulseTickOverrunsHighlooks like it belongs in the same boat but does not: it only reads the tick histogram andpulse_server_tick_budget_seconds, both public API metrics, so it keeps working in degraded mode the same as the tick rate and log/worldgen rules.PulseModHoggingTickreadspulse_server_tick_busy_secondstoo, so it is quiet in degraded mode for the same reason. It also needs attribution switched on (see the main README's "Attribution" section) forpulse_mod_tick_share. That family is an observable gauge tied to the same duty cycle, so it disappears from/metricsthe moment attribution stops, whether that is/pulse attribution off, a reload withEnabledfalse, or the duty cycle giving up on its own; it no longer sits there at the last burst's values looking like a real number. The rule keeps its own guard regardless,increase(pulse_attribution_ticks_total[5m]) > 0: that counter only moves when a burst actually completes, a more direct signal that attribution is doing real work than the family merely being present, and one that still protects the rule if a future change ever lets the gauge report something without a completed burst behind it.
Validate the file after editing it with the same promtool container used to write it:
docker run --rm -v "$PWD/contrib/alerts:/a" --entrypoint promtool prom/prometheus check rules /a/pulse-alerts.ymlThat only parses the PromQL; it does not run it, so a rule can pass check rules and still never
fire, for instance an and or an arithmetic operator whose two sides carry different labels and
so never match anything. pulse-alerts.test.yml catches that class of mistake by running every
rule against synthetic data, one case where it should fire and one where it should stay quiet:
docker run --rm -v "$PWD/contrib/alerts:/a" --entrypoint promtool prom/prometheus test rules /a/pulse-alerts.test.ymlRun both after touching this file. Adding a rule without adding its two cases here is how the next silent one gets through.
The rules #
pulse-tick #
PulseTickRateLow warning for 5m
Tick rate below 27 TPS for 5m (target 30).
rate(pulse_server_ticks_total[2m]) < 27Tick rate has been under 27 TPS for 5 minutes; check pulse_server_tick_busy_seconds and the worldgen queue for what's eating the budget before it gets worse.
Reads: pulse_server_ticks_total
Why this rule, from the comments of the file
rate() over the tick counter is TPS. Two-tier: 27 is "visibly behind", 20 is "players are feeling this". Both use the same alert name so Alertmanager groups them; a very sick server fires both at once, which is expected, not a bug.
PulseTickRateLow CRITICAL for 5m
Tick rate below 20 TPS for 5m.
rate(pulse_server_ticks_total[2m]) < 20Tick rate has been under 20 TPS for 5 minutes, deep into player-visible lag; this needs attention now, not a wait-and-see.
Reads: pulse_server_ticks_total
Why this rule, from the comments of the file
rate() over the tick counter is TPS. Two-tier: 27 is "visibly behind", 20 is "players are feeling this". Both use the same alert name so Alertmanager groups them; a very sick server fires both at once, which is expected, not a bug.
PulseTickSaturationHigh warning for 5m
Tick busy time over 80% of budget for 5m.
pulse_server_tick_busy_seconds / pulse_server_tick_budget_seconds > 0.8Average tick busy time has been over 80% of the configured budget for 5 minutes; find what's using the headroom before ticks start missing budget outright.
Reads: pulse_server_tick_busy_seconds engine, pulse_server_tick_budget_seconds
Why this rule, from the comments of the file
Busy time over budget is the leading indicator, before rate() over the tick counter even moves: a server can hold 30 TPS while spending most of its budget getting there. Only meaningful when the engine-probe families are being served (the guarded cast in Pulse/EngineProbe.cs succeeded) -- if it failed, pulse_server_tick_busy_seconds is simply absent from /metrics and this rule has no data to evaluate, silently not firing either way. Absence of this alert is not proof of headroom; check /metrics for the family before trusting silence here.
PulseTickOverrunsHigh warning for 5m
p99 tick time over 3x budget for 5m.
histogram_quantile(0.99, rate(pulse_server_tick_seconds_bucket[5m])) > 3 * pulse_server_tick_budget_secondsThe slowest 1% of ticks have been running past three times the tick budget for 5 minutes; look for worldgen, an autosave, or a mod spiking before this turns into a sustained TPS drop.
Reads: pulse_server_tick_seconds, pulse_server_tick_budget_seconds
Why this rule, from the comments of the file
Two ways to phrase "ticks are overrunning": count buckets past a fixed 50ms line, or take a quantile against the live budget gauge. Went with the quantile: it tracks pulse_server_tick_budget_seconds, so it stays correct if an operator retunes the tick rate with /serverconfig, where a hardcoded 50ms cutoff would quietly go stale. It also doesn't depend on 0.05 staying a bucket boundary in Pulse/PulseModSystem.cs's TickBuckets array. The tradeoff is bucket-interpolation error from only ten finite buckets, which is fine at warning granularity -- this is "is the tail bad", not a latency SLO.
pulse-engine #
PulseEngineWarning warning
Engine {{ $labels.kind }} warning firing.
increase(pulse_engine_warnings_total{kind!="memory"}[10m]) > 0The engine logged a {{ $labels.kind }} warning in the last 10 minutes; check the server log around that time for what triggered it.
Reads: pulse_engine_warnings_total
Why this rule, from the comments of the file
increase() > 0 over a 10m window already is the debounce; a single stray warning ten minutes ago still counts, which is the point. Same alert name for both severities, kind excluded from the warning leg so memory doesn't double-fire under its own critical rule below.
PulseEngineWarning CRITICAL
Engine memory warning firing.
increase(pulse_engine_warnings_total{kind="memory"}[10m]) > 0The engine crossed 90% of its memory ceiling in the last 10 minutes and self-terminates at 100%; raise DieAboveMemoryUsageMb or free heap now, not after it restarts itself.
Reads: pulse_engine_warnings_total
Why this rule, from the comments of the file
DieAboveMemoryUsageMb: the "memory" kind is the engine crossing 90% of its ceiling, and the engine kills itself outright at 100%. That gap is small enough to treat as critical from the first warning rather than waiting for a repeat.
pulse-log #
PulseLogErrorBurst warning
More than 5 error log entries in 10m.
increase(pulse_log_entries_total{level="error"}[10m]) > 5More than 5 error-level log entries landed in the last 10 minutes; tail the server log for the recurring one rather than a one-off.
Reads: pulse_log_entries_total
Why this rule, from the comments of the file
A single stray error is normal noise (a malformed packet, a one-off mod hiccup); a burst is one thing failing repeatedly. 5 in 10 minutes is a judgement call, not a measured value like the tick numbers above -- tune it if your server's baseline error rate differs.
PulseLogFatal CRITICAL
Fatal log entry.
increase(pulse_log_entries_total{level="fatal"}[10m]) > 0A fatal log entry landed in the last 10 minutes; read the server log now. Error and fatal entries count the same way toward the engine's own DieAboveErrorCount self-shutdown threshold.
Reads: pulse_log_entries_total
Why this rule, from the comments of the file
Fatal is the engine's own top log severity, so even one is worth a page, not a threshold like the burst rule above. It is not the only level that can shut the server down though: VintagestoryLib.dll's ServerSystemMonitor.OnEntryAdded counts error and fatal entries the same way toward the engine's own DieAboveErrorCount threshold, so a sustained stream of errors is also a path to an unplanned restart, covered above by the burst rule.
pulse-availability #
PulseEndpointDown CRITICAL for 2m
Pulse scrape target down for 2m.
up{job="vintagestory"} == 0Prometheus has not been able to scrape the vintagestory job for 2 minutes; confirm the game server process is up and Pulse's endpoint is still bound before assuming it's just network flakiness.
Why this rule, from the comments of the file
`up` is Prometheus's own synthetic series, one per scrape target; job name here matches contrib/grafana/prometheus.yml's `job_name: vintagestory`. Change the job label if yours scrapes Pulse under a different name.
pulse-worldgen #
PulseWorldgenQueueStuck warning for 15m
Worldgen queue over 500 columns for 15m.
pulse_worldgen_queue_columns > 500The worldgen queue has held more than 500 pending columns for 15 minutes without draining; check the worldgen thread isn't stalled and that disk I/O for chunk writes isn't backed up.
Reads: pulse_worldgen_queue_columns
Why this rule, from the comments of the file
A queue that fills during a player exploration burst and drains afterwards is the queue doing its job. 15m of staying above 500 is long enough that draining stopped, not that it is merely busy.
pulse-attribution #
PulseModHoggingTick warning for 10m
Mod {{ $labels.modid }} holding over half the tick for 10m while the server is loaded.
(pulse_mod_tick_share{modid!~"engine|unattributed"} > 0.5) and ignoring(modid) (pulse_server_tick_busy_seconds / pulse_server_tick_budget_seconds > 0.8) and ignoring(modid) (increase(pulse_attribution_ticks_total[5m]) > 0){{ $labels.modid }} has held more than 50% of the profiled main-thread tick for 10 minutes while tick busy time stayed over 80% of budget; open the dashboard's attribution row for the full breakdown and look at that mod first. Attribution is sampled (about one tick in thirty at the default duty cycle) and only sees listeners and behaviours, not broadcast event handlers, so treat this as where to look first, not a full accounting.
Reads: pulse_mod_tick_share attribution, pulse_server_tick_busy_seconds engine, pulse_server_tick_budget_seconds, pulse_attribution_ticks_total attribution
Why this rule, from the comments of the file
Needs attribution switched on (README.md's Attribution section: the config block or `/pulse attribution on`). pulse_mod_tick_share is an observable gauge that disappears from /metrics the moment attribution stops, whether that is switching it off, a reload with Enabled false, or the duty cycle giving up on its own, so its absence alone already means nothing is being measured. The increase(pulse_attribution_ticks_total) term below stays as a second, independent guard: it only moves when a burst actually completes, which is a more direct proof of live measurement than the share's mere presence, and it is what would still catch a future change that let the share report again with nothing behind it. Two conditions on top of that, both required: pulse_mod_tick_share sums to 1 across every modid including engine and unattributed, so one mod clearing 0.5 already outweighs everything else on the server combined, and the load gate reuses PulseTickSaturationHigh's own 80% busy-over-budget threshold, on the same reasoning, a mod owning half the tick is not worth paging on while the server is coasting at 20% of budget. 10m is several burst refreshes at the default duty cycle (10 BurstTicks / 10s IntervalSeconds samples roughly once every 10 seconds), long enough that this is a sustained hog and not one unlucky sample landing mid-spike, which the README's "What it cannot see" warns can happen on a spiky server. The excluded modid is not always a third-party mod either: the base game itself ships as several mods (survival, essentials, game among them), so on an entity-heavy vanilla server with nothing else installed this can still name one of those. "and" defaults to matching on the full label set, and the left side carries modid while neither the load gate nor the ticks guard does, so without "ignoring(modid)" on both no side ever matches anything and this alert can never fire. Ignoring modid rather than naming "on(instance, job)" keeps whatever else a target carries part of the match too, for example a label a scrape job's relabel_configs adds, instead of silently discarding it the way naming just instance and job would.
Source: contrib/alerts/README.md, contrib/alerts/pulse-alerts.yml