On this page

Metrics

Metric families #

Queries are the bundled dashboard's own. $__rate_interval is a Grafana variable: in Prometheus itself, write a window of at least four scrape intervals, such as [1m] at 15 seconds. The dotnet_ runtime families are described under Runtime metrics below and charted in the dashboard's Runtime row. 24 entries for 28 families: four entries document two families each.

The metric families it serves:

  • pulse_server_ticks_total (counter): server ticks processed since startup. Prometheus rate() over it is your TPS.
    Used by 2 panels and 2 alert rules

    Panels: At a glance: Tick rate, Tick health: Tick rate

    Alert rules: PulseTickRateLow (warning), PulseTickRateLow (critical)

    Tick rate: ticks per second
    rate(pulse_server_ticks_total[$__rate_interval])
  • pulse_server_tick_seconds (histogram): wall-clock seconds between consecutive ticks, measured with the mod's own stopwatch, in buckets from 25 ms to 1 s.
    Used by 2 panels and 1 alert rule

    Panels: At a glance: Tick time p99, Tick health: Tick time

    Alert rules: PulseTickOverrunsHigh (warning)

    Tick time p99: p99
    histogram_quantile(0.99, sum by (le) (rate(pulse_server_tick_seconds_bucket[$__rate_interval])))
    Tick time: p50
    histogram_quantile(0.50, sum by (le) (rate(pulse_server_tick_seconds_bucket[$__rate_interval])))
    Tick time: p95
    histogram_quantile(0.95, sum by (le) (rate(pulse_server_tick_seconds_bucket[$__rate_interval])))
  • pulse_players_online (gauge): players currently connected.
    Used by 2 panels

    Panels: At a glance: Players online, Players: Players online

    Players online: players
    pulse_players_online
  • pulse_entities_loaded (gauge): entities loaded in the world.
    Used by 2 panels

    Panels: At a glance: Entities loaded, World: Entities loaded

    Entities loaded: entities
    pulse_entities_loaded
  • pulse_server_tick_budget_seconds (gauge): the configured tick budget, so tick_seconds / tick_budget_seconds reads as saturation.
    Used by 3 panels and 3 alert rules

    Panels: Tick health: Tick rate, Tick health: Tick time, Tick health: Busy time against the budget

    Alert rules: PulseTickSaturationHigh (warning), PulseTickOverrunsHigh (warning), PulseModHoggingTick (warning)

    Tick rate: budget allows
    1 / pulse_server_tick_budget_seconds
    Tick time: budget
    pulse_server_tick_budget_seconds
  • pulse_worldgen_queue_columns (gauge): chunk columns waiting in the generation queue.
    Used by 1 panel and 1 alert rule

    Panels: Worldgen: Worldgen queue

    Alert rules: PulseWorldgenQueueStuck (warning)

    Worldgen queue: columns queued
    pulse_worldgen_queue_columns
  • pulse_worldgen_columns_generated_total (counter): chunk columns generated since startup.
    Used by 1 panel

    Panels: Worldgen: Columns generated

    Columns generated: columns per minute
    60 * rate(pulse_worldgen_columns_generated_total[$__rate_interval])
  • pulse_chunks_loaded (gauge): chunks loaded in the world, on a slow cadence of its own.
    Used by 2 panels

    Panels: At a glance: Chunks loaded, World: Chunks loaded

    Chunks loaded: chunks
    pulse_chunks_loaded
  • pulse_log_entries_total{level} (counter): log entries by severity, one series each for warning, error and fatal.
    Used by 1 panel and 2 alert rules

    Panels: Pauses, warnings and logs: Log entries, 5 minute window

    Alert rules: PulseLogErrorBurst (warning), PulseLogFatal (critical)

    Log entries, 5 minute window: {{level}}
    increase(pulse_log_entries_total[5m])
  • pulse_engine_warnings_total{kind} (counter): the engine's own health warnings, recognised by the text it logs. overload is a tick past 500 ms, memory is crossing 90% of DieAboveMemoryUsageMb, suspend_timeout is a server suspend that gave up waiting for a thread, and autosave_io is an autosave arriving while the previous one is still writing.
    Used by 1 panel and 2 alert rules

    Panels: Pauses, warnings and logs: Engine warnings, 5 minute window

    Alert rules: PulseEngineWarning (warning), PulseEngineWarning (critical)

    Engine warnings, 5 minute window: {{kind}}
    increase(pulse_engine_warnings_total[5m])
  • pulse_server_uptime_seconds (gauge): seconds the server has been ticking. This is the engine's own unpaused clock, so it stops during a save and is not process uptime.
    Used by 1 panel

    Panels: At a glance: Uptime

    Uptime: uptime
    pulse_server_uptime_seconds
  • pulse_entities_by_code{code} (gauge): loaded entities by entity code, the ten most numerous plus an other bucket for the rest. Refreshed on the same slow cadence as the chunk count.
    Used by 1 panel

    Panels: World: Entities by code

    Entities by code: {{code}}
    pulse_entities_by_code
  • pulse_player_ping_seconds{stat} (gauge): round trip time to the players online, as avg and max. Both read 0 with nobody connected.
    Used by 1 panel

    Panels: Players: Ping

    Ping: {{stat}}
    pulse_player_ping_seconds
  • pulse_player_deaths_total (counter): player deaths since startup.
    Used by 1 panel

    Panels: Players: Player deaths, 5 minute window

    Player deaths, 5 minute window: deaths
    increase(pulse_player_deaths_total[5m])
  • pulse_server_suspends_total and pulse_server_suspend_seconds_total (counters): how often the server suspended ticking and how long it spent suspended. Every autosave is one of these, and the seconds are the pause players actually feel.
    Used by 2 panels

    Panels: Pauses, warnings and logs: Suspends, 5 minute window, Pauses, warnings and logs: Seconds suspended, 5 minute window

    Suspends, 5 minute window: suspends
    increase(pulse_server_suspends_total[5m])
    Seconds suspended, 5 minute window: seconds suspended
    increase(pulse_server_suspend_seconds_total[5m])
  • pulse_network_sent_bytes_total and pulse_network_received_bytes_total (counters): bytes over the main TCP channel. UDP is not in these two; the public server API does not report it.
    Used by 1 panel

    Panels: Network: Traffic from the byte totals

    Traffic from the byte totals: tcp sent
    rate(pulse_network_sent_bytes_total[$__rate_interval])
    Traffic from the byte totals: tcp received
    rate(pulse_network_received_bytes_total[$__rate_interval])

Six more come from the engine's own accounting, which no public API exposes. See the note on degraded mode below for what happens when they are unavailable.

  • pulse_server_tick_busy_seconds (gauge): the average time one tick spent working, sleep excluded, over the engine's last completed two-second window. This is the number /stats prints, and the only view of headroom below the tick budget that exists at all. The engine measures in whole milliseconds, so an idle server legitimately averages zero.
    Used by 1 panel and 2 alert rules

    Panels: Tick health: Busy time against the budget

    Alert rules: PulseTickSaturationHigh (warning), PulseModHoggingTick (warning)

    Busy time against the budget: busy
    pulse_server_tick_busy_seconds
  • pulse_network_packets_per_second{channel} and pulse_network_bytes_per_second{channel} (gauges): traffic rates over that same window, split tcp and udp. Gauges rather than counters because the engine zeroes its window rather than accumulating it; a window cut short by a suspend reads low for one sample.
    Used by 2 panels

    Panels: Network: Bytes per second by channel, Network: Packets per second by channel

    Bytes per second by channel: {{channel}}
    pulse_network_bytes_per_second
    Packets per second by channel: {{channel}}
    pulse_network_packets_per_second
  • pulse_connection_queue_clients (gauge): clients waiting because the server is full.
    Used by 1 panel

    Panels: Network: Connection queue

    Connection queue: clients waiting
    pulse_connection_queue_clients
  • pulse_network_udp_sent_bytes_total and pulse_network_udp_received_bytes_total (counters): the UDP totals missing from the two public byte counters above.
    Used by 1 panel

    Panels: Network: Traffic from the byte totals

    Traffic from the byte totals: udp sent
    rate(pulse_network_udp_sent_bytes_total[$__rate_interval])
    Traffic from the byte totals: udp received
    rate(pulse_network_udp_received_bytes_total[$__rate_interval])

Four more answer "which mod is eating the tick", and only when you turn them on. They have a section of their own further down.

  • pulse_mod_tick_share{modid} (gauge): the fraction of profiled main-thread busy time that went to one mod over the last completed burst. The shares add up to 1 across every modid, including the two Pulse adds: engine for the server's own systems and for the time no marker named, and unattributed for work that was marked but that no loaded mod claims.
    Used by 2 panels and 1 alert rule
  • pulse_mod_tick_seconds_total{modid} (counter): main-thread seconds attributed to one mod. Sampled, not total: this is time measured inside the bursts, not time since startup. Divide by the tick counter below to compare two servers, or take rate() of it against rate(pulse_attribution_ticks_total) for seconds per profiled tick.
    Used by 1 panel

    Panels: Attribution, only when turned on: Attributed tick time by mod

    Attributed tick time by mod: {{modid}}
    rate(pulse_mod_tick_seconds_total[$__rate_interval]) / ignoring(modid) group_left() rate(pulse_attribution_ticks_total[$__rate_interval])
  • pulse_attribution_ticks_total (counter): ticks actually profiled, which is what makes the sampled seconds mean anything.
    Used by 2 panels and 1 alert rule

    Panels: Attribution, only when turned on: Attributed tick time by mod, Attribution, only when turned on: Profiling health

    Alert rules: PulseModHoggingTick (warning)

    Attributed tick time by mod: {{modid}}
    rate(pulse_mod_tick_seconds_total[$__rate_interval]) / ignoring(modid) group_left() rate(pulse_attribution_ticks_total[$__rate_interval])
    Profiling health: profiled ticks/s
    rate(pulse_attribution_ticks_total[$__rate_interval])
  • pulse_attribution_dropped_samples_total (counter): profiler readings thrown away because they overflowed. The engine accumulates each marker's time into a 32 bit counter of stopwatch ticks, which wraps negative somewhere past two seconds inside a single tick. A wrapped reading is not a large number, it is garbage, so it is dropped and counted here instead of being published as data. Anything but a flat zero means the server had a tick so bad that a single marker ran for over two seconds.
    Used by 1 panel

    Panels: Attribution, only when turned on: Profiling health

    Profiling health: dropped samples/s
    rate(pulse_attribution_dropped_samples_total[$__rate_interval])

The tick period is measured rather than taken from the value the engine hands tick listeners, because that one is rounded to whole milliseconds. Overruns still land exactly: once a tick's work exceeds the budget the engine's throttle sleep is zero, and the period is the busy time. Most gauges come from a snapshot the tick listener refreshes about once a second on the server's main thread, so a scrape never touches live world state.

Loaded chunks are the exception. The engine offers no cheap count, and the one accessor that exists clones the entire loaded-chunk dictionary under the chunk lock, so that gauge gets its own listener at ChunksRefreshSeconds and reads 0 until the first refresh. The entity breakdown rides that same slow listener, but it is not the same kind of gauge: pulse_entities_by_code is recorded only from the listener callback, with no observable default behind it, so it is absent from the exposition entirely until that first refresh, not present at 0. The event-driven counters do not ride the tick listener either: columns generated is incremented from MapChunkGeneration, which fires on the worldgen thread, and the log counters from Logger.EntryAdded, which fires on whichever thread wrote the line. Both handlers classify and increment, and nothing else.

Degraded mode #

Six of those families come from a place the modding API does not reach. Tick busy time, the per-window packet and byte counts, the connection queue depth and the UDP byte totals are all measured by the engine on concrete types inside VintagestoryLib.dll, which the game's authors change freely between versions and make no promises about. Pulse casts the server world to Vintagestory.Server.ServerMain once, at startup, in a try/catch, and every read of those types lives in one small class.

If that cast ever stops working, Pulse logs one warning naming what went wrong and carries on. The six families are simply absent from /metrics rather than present and lying, and every other metric on this page keeps being served exactly as before, including the tick period histogram, which still catches every overrun: once a tick's work exceeds the budget the engine's throttle sleep is zero and the period is the busy time. What you lose is the view of headroom on a healthy server. One warning, no retry loop, no log spam, and nothing that can stop a game server.

The compile-time reference to VintagestoryLib.dll is deliberate for the same reason. A game version that renames or moves ServerMain breaks the Pulse build, loudly, before a release goes out, instead of shipping a mod that quietly serves six families fewer.

Runtime metrics #

With RuntimeMetrics left on, the .NET runtime's own System.Runtime meter is served alongside Pulse's, as dotnet_* families: GC collections and pause time, heap size and fragmentation by generation, working set, CPU time by mode, JIT, thread pool, lock contention, loaded assemblies. None of it is instrumented here. The runtime publishes the meter, and the writer renames each instrument the way Prometheus's otlptranslator does, the library Prometheus's own OTLP receiver, Mimir and Grafana Cloud use to turn an OTLP instrument into a Prometheus name: dots become underscores, the instrument's unit becomes a trailing word unless the name already contains it as a word, and a monotonic counter's name ends in _total, moved there rather than duplicated if the name already spells "total" somewhere. dotnet.gc.collections is served as dotnet_gc_collections_total; dotnet.process.memory.working_set, a gauge in bytes, is served as dotnet_process_memory_working_set_bytes; dotnet.gc.heap.total_allocated, a counter also in bytes, is served as dotnet_gc_heap_allocated_bytes_total rather than the doubled ..._total_allocated_bytes_total. This is the same name Grafana derives when it translates the OTLP export, so a dashboard or alert built against a server scraped over OTLP through Grafana Cloud reads Pulse's own /metrics without translation too.

Nine of these families moved to this spelling in 0.2, to line up with that translation; see the changelog for the full old to new list if you have a dashboard or alert built against the earlier names. pulse_* families are unaffected.

Pulse renders the shape each instrument declares, including where that is arguable. dotnet_thread_pool_thread_count_total is typed as a counter because the runtime publishes it as an ObservableCounter, even though the number goes down as often as up. Second-guessing the framework here would only make the series harder to correlate with any other .NET exporter.

Source: README.md