monitord ... know how happy your systemd is! π
/proc) β available on all standard Linux systemsunix:path=/run/dbus/system_bus_socket)io.systemd.Metrics on /run/systemd/report/io.systemd.Manager,
used today for unit counts/state): v260+; v261+ for JobsQueued, the
exact UnitsTotal, UnitsByLoadStateTotal, per-unit
StateChangeTimestamp, and per-service StatusErrno (on v260,
total_units is approximated by summing mapped per-type counts, load-state
totals are counted from per-unit load states, while jobs_queued,
time_in_state_usecs and status_errno are unavailable)io.systemd.Network.Describe, used today for interface states):
v257+io.systemd.Manager/io.systemd.Unit on
/run/systemd/io.systemd.Manager): v258+ for Manager.Describe, used
today for the systemd version and system state. Unit.List needs v261+
for anything monitord reads from it: the method exists in v258, but the
per-type context sections it returns are added by varlink-service.c and
varlink-timer.c, which first appear in v261 β so per-service stats and
service types are unavailable before then. v261+ also for job listingio.systemd.Machine.List on
/run/systemd/machine/io.systemd.Machine, used today for machine
enumeration): v257+monitord collects systemd health metrics via D-Bus (and optionally Varlink) and outputs them as JSON. It provides visibility into:
systemd-analyze blamesystemd-machined (e.g. systemd-nspawn; VMs are skipped)systemd-analyze verify and reports failing unit counts by typeWe offer the following run modes:
daemon: mode options are set in monitord.confdaemon_stats_refresh_secsOpen to more formats / run methods ... Open an issue to discuss. Depends on the dependencies basically.
monitord is a config driven binary. We plan to keep CLI arguments to a minimum.
INFO level logging is enabled to stderr by default. Use -l LEVEL to increase or decrease logging.
Install monitord:
bash
cargo install monitord
Create a minimal config at /etc/monitord.conf:
```ini
[monitord]
output_format = json-pretty
[units] enabled = true
[pid1] enabled = true ```
bash
monitordThis will collect unit counts and PID 1 stats, then print JSON to stdout and exit. Enable additional collectors in the config as needed (see Configuration below).
Download pre-built binaries from GitHub Releases:
monitord-linux-amd64 β x86_64monitord-linux-aarch64 β ARM64# Example: download and install the latest release (x86_64)
curl -L -o /usr/local/bin/monitord \
https://github.com/cooperlees/monitord/releases/latest/download/monitord-linux-amd64
chmod +x /usr/local/bin/monitord
Install via cargo or use as a dependency in your Cargo.toml.
cargo install monitordmonitord.confmonitord --helpMONITORD_CONFIG env var to set config pathcrl-linux:monitord cooper$ monitord --help
monitord: Know how happy your systemd is! π
Usage: monitord [OPTIONS]
Options:
-c, --config <CONFIG>
Location of your monitord config
[default: /etc/monitord.conf]
-l, --log-level <LOG_LEVEL>
Adjust the console log-level
[default: Info]
[possible values: error, warn, info, debug, trace]
-h, --help
Print help (see a summary with '-h')
-V, --version
Print version
monitord can have the different components monitored. To enable / disabled set the following in our monitord.conf. This file is ini format to match systemd unit files.
# Pure ini - no yes/no for bools
[monitord]
# Set a custom dbus address to connect to
# OPTIONAL: If not set, we default to the Unix socket below
dbus_address = unix:path=/run/dbus/system_bus_socket
# Timeout in seconds for dbus connection/collections
# OPTIONAL: default is 30 seconds
dbus_timeout = 30
# Run as a daemon or 1 time
daemon = false
# Time to refresh systemd stats in seconds
# Daemon mode only
daemon_stats_refresh_secs = 60
# Prefix flat-json key with this value
# The value automatically gets a '.' appended (so don't put here)
key_prefix = monitord
# cron/systemd timer output format
# Supported: json, json-flat, json-pretty
output_format = json
# Grab as much stats from DBus GetStats call
# we can from running dbus daemon
# More tested on dbus-broker daemon
[dbus]
# Summary counters - both dbus-broker + dbus-daemon
enabled = false
# Count stale pidfds held by the system dbus-broker from procfs
stale_fd_stats = true
# dbus.user.* metrics: user stats as reported by dbus-broker
user_stats = false
# dbus.peer.* metrics: peer stats as reported by dbus-broker
peer_stats = false
# Only report peers that own a well-known bus name when peer_stats is on
peer_well_known_names_only = false
# dbus.cgroup.* stats is an aggregation of peer_stats by cgroup
# by dbus-broker
cgroup_stats = false
# Grab networkd stats from files + networkctl
[networkd]
enabled = true
link_state_dir = /run/systemd/netif/links
# Enable grabbing PID 1 stats via procfs
[pid1]
enabled = true
# Services to grab extra stats for
# .service is important as that's what DBus returns from `list_units`
[services]
foo.service
# systemd version and system state (e.g. running, degraded)
[system-state]
enabled = true
[timers]
enabled = true
[timers.allowlist]
foo.timer
[timers.blocklist]
bar.timer
# Grab unit status counts via dbus
[units]
enabled = true
state_stats = true
# Also record how long each unit has been in its current state
state_stats_time_in_state = true
ignore_inactive_oneshot_services = true
# Filter what services you want collect state stats for
# If both lists are configured blocklist is preferred
# If neither exist all units state will generate counters
[units.state_stats.allowlist]
foo.service
[units.state_stats.blocklist]
bar.service
# machines config
[machines]
enabled = true
# Same rules apply as state_stats lists above
[machines.allowlist]
foo
[machines.blocklist]
bar
# Collect via systemd varlink APIs instead of D-Bus where available
# (see the Varlink section below). Falls back to D-Bus unless no_fallback = true
[varlink]
enabled = false
no_fallback = false
# Boot blame metrics - shows the N slowest units at boot
# Similar to `systemd-analyze blame`
# Disabled by default
[boot]
enabled = false
# Cache boot blame stats in <cache_dir>/<boot_id>.boot_blame.bin
# Enabled by default; set false to force recalculation every run
cache_enabled = true
# Directory where boot blame cache files are stored (optional)
# Defaults to /run/monitord if not set; avoid world-writable dirs like /tmp when running privileged
# cache_dir = /tmp
# Number of slowest units to report
num_slowest_units = 5
# Optional: only include specific units in boot blame (if empty, all units are checked)
# Same rules apply as state_stats lists above
[boot.allowlist]
# slow-startup.service
# Optional: exclude specific units from boot blame
[boot.blocklist]
# noisy-but-expected.service
# Unit verification using systemd-analyze verify
# Disabled by default as it can be slow on large systems
[verify]
enabled = false
# Optional: only verify specific units (if empty, all units are checked)
[verify.allowlist]
# example.service
# example.timer
# Optional: skip verification for specific units
[verify.blocklist]
# noisy.service
# broken.timer
When using the provided monitord.service, systemd creates /run/monitord via
RuntimeDirectory=monitord and assigns ownership to the configured service User/Group.
If you run monitord another way and use the default cache_dir (/run/monitord), ensure the
directory exists and is writable by the monitord process user so boot cache files can be created.
Alternatively, set cache_dir to a location like /tmp that is always writable.
From version >=0.11 monitord supports obtaining the same set of key from
systemd 'machines' (i.e. machinectl --list). Only machines of class container
are collected. They are enumerated via machined's io.systemd.Machine.List varlink
API (systemd v257+) when varlink is enabled, falling back to machined's D-Bus API.
The keys are the same format as below in json_flat output but are prefixed with
the machines keyword and machine name. For example:
# $KEY_PREFIX.machines.$MACHINE_NAME
{
...
"monitord.machines.foo.pid1.fd_count": 69,
...
}
Machine collection is the only part of monitord that needs extra privileges, so skip this
if you don't monitor machines. When running as a non-root user (like the shipped
monitord.service), what it needs depends on how machines are collected:
| Machines collected over | Capabilities needed |
|---|---|
D-Bus (the default, or [machines] varlink = false) |
CAP_SYS_PTRACE |
varlink ([varlink] enabled = true and [machines] varlink = true) |
CAP_SYS_PTRACE + CAP_SYS_ADMIN |
# /etc/systemd/system/monitord.service.d/machines.conf
# Drop CAP_SYS_ADMIN if you collect machines over D-Bus.
[Service]
AmbientCapabilities=CAP_SYS_PTRACE CAP_SYS_ADMIN
CapabilityBoundingSet=CAP_SYS_PTRACE CAP_SYS_ADMIN
CAP_SYS_PTRACE lets monitord reach each machine's sockets and files through
/proc/<leader_pid>/root/β¦: the kernel only allows that for processes that may ptrace
the machine's leader. Without it every machine fails with Permission denied, while
host collection keeps working. CAP_DAC_READ_SEARCH is not enough.CAP_SYS_ADMIN is needed for varlink only. systemd's varlink servers only accept
callers inside the machine's PID namespace (see
#211 and
systemd/systemd#43807), so monitord
spawns one small helper per machine inside that namespace, which needs CAP_SYS_ADMIN.
The helper is monitord itself re-executed; it shows up inside the machine as
monitord __monitord-machine-connector. Right after starting it opens the machine's root,
drops to an unprivileged user (nobody if monitord runs as root) with no capabilities
at all, and from then on only connects to a fixed set of systemd sockets and hands the
connection back to monitord. It is spawned once per machine and reused for every
collection until the machine restarts, so there is no fork per collection cycle.
Without CAP_SYS_ADMIN monitord logs one warning and collects machines over D-Bus
(with [varlink] no_fallback = true this is an error instead).Both capabilities are broad: CAP_SYS_PTRACE also allows reading other processes'
memory, and CAP_SYS_ADMIN is close to root. If you'd rather not grant CAP_SYS_ADMIN,
set [machines] varlink = false to collect machines over D-Bus with CAP_SYS_PTRACE
only, or run monitord inside each machine instead of collecting from the host.
Machines started with user namespacing (e.g. machinectl start defaults to
PrivateUsers=pick) also reject monitord's D-Bus connection, as host users are not mapped
inside the machine. Set PrivateUsers=no in /etc/systemd/nspawn/<machine>.nspawn to
collect from them.
Normal serde_json non pretty JSON. All on one line. Most compact format.
Move all key value pairs to the top level and . notate components + sub values. Is semi pretty too + custom. All unittested ...
stat_collection_run_time_ms is emitted in milliseconds (with _ms suffix) to follow
Prometheus metric naming conventions for duration units, which keeps unit semantics
clear and consistent when these keys are transformed into Prometheus metric names.
This sample is from a host with no containers, so it deliberately omits the
machines.<name>.* key space. Those keys duplicate the top level ones documented
here, just prefixed per machine β see Machines support β so
they are not repeated below.
{
"boot.blame.cpe_chef.service": 103.05,
"boot.blame.dev-ttyS0.device": 15.809,
"boot.blame.dnf5-automatic.service": 204.159,
"boot.blame.sys-module-fuse.device": 16.21,
"boot.blame.systemd-networkd-wait-online.service": 1.674,
"collection_timings.list_units_ms": 5.26,
"collection_timings.per_unit_loop_ms": 42.99,
"collection_timings.service_dbus_fetches": 0,
"collection_timings.slowest_units.0.chronyd.service": 5.248485,
"collection_timings.slowest_units.1.fstrim.timer": 5.173884,
"collection_timings.slowest_units.2.systemd-logind.service": 0.042048,
"collection_timings.slowest_units.3.dev-nvme0n1.device": 0.035938,
"collection_timings.slowest_units.4.sys-module-fuse.device": 0.035106,
"collection_timings.state_dbus_fetches": 0,
"collection_timings.timer_dbus_fetches": 24,
"collector_timings.boot_blame.elapsed_ms": 53.36,
"collector_timings.boot_blame.start_offset_ms": 0.08,
"collector_timings.boot_blame.success": 1,
"collector_timings.dbus_stats.elapsed_ms": 1.84,
"collector_timings.dbus_stats.start_offset_ms": 0.11,
"collector_timings.dbus_stats.success": 1,
"collector_timings.machines.elapsed_ms": 0.51,
"collector_timings.machines.start_offset_ms": 0.14,
"collector_timings.machines.success": 0,
"collector_timings.networkd.elapsed_ms": 8.48,
"collector_timings.networkd.start_offset_ms": 0.1,
"collector_timings.networkd.success": 1,
"collector_timings.pid1.elapsed_ms": 0.94,
"collector_timings.pid1.start_offset_ms": 0.1,
"collector_timings.pid1.success": 1,
"collector_timings.system_state.elapsed_ms": 2.38,
"collector_timings.system_state.start_offset_ms": 0.12,
"collector_timings.system_state.success": 1,
"collector_timings.units.elapsed_ms": 53.24,
"collector_timings.units.start_offset_ms": 0.06,
"collector_timings.units.success": 1,
"collector_timings.verify.elapsed_ms": 31.07,
"collector_timings.verify.start_offset_ms": 0.13,
"collector_timings.verify.success": 1,
"collector_timings.version.elapsed_ms": 2.25,
"collector_timings.version.start_offset_ms": 0.09,
"collector_timings.version.success": 1,
"dbus.active_connections": 10,
"dbus.bus_names": 16,
"dbus.cgroup.system.slice-systemd-logind.service.activation_request_bytes": 0,
"dbus.cgroup.system.slice-systemd-logind.service.activation_request_fds": 0,
"dbus.cgroup.system.slice-systemd-logind.service.incoming_bytes": 16,
"dbus.cgroup.system.slice-systemd-logind.service.incoming_fds": 0,
"dbus.cgroup.system.slice-systemd-logind.service.match_bytes": 6942,
"dbus.cgroup.system.slice-systemd-logind.service.matches": 5,
"dbus.cgroup.system.slice-systemd-logind.service.name_objects": 1,
"dbus.cgroup.system.slice-systemd-logind.service.outgoing_bytes": 0,
"dbus.cgroup.system.slice-systemd-logind.service.outgoing_fds": 0,
"dbus.cgroup.system.slice-systemd-logind.service.reply_objects": 0,
"dbus.incomplete_connections": 0,
"dbus.match_rules": 26,
"dbus.peak_bus_names": 33,
"dbus.peak_bus_names_per_connection": 2,
"dbus.peak_match_rules": 33,
"dbus.peak_match_rules_per_connection": 13,
"dbus.peer.org.freedesktop.systemd1.activation_request_bytes": 0,
"dbus.peer.org.freedesktop.systemd1.activation_request_fds": 0,
"dbus.peer.org.freedesktop.systemd1.incoming_bytes": 16,
"dbus.peer.org.freedesktop.systemd1.incoming_fds": 0,
"dbus.peer.org.freedesktop.systemd1.match_bytes": 46533,
"dbus.peer.org.freedesktop.systemd1.matches": 33,
"dbus.peer.org.freedesktop.systemd1.name_objects": 1,
"dbus.peer.org.freedesktop.systemd1.outgoing_bytes": 0,
"dbus.peer.org.freedesktop.systemd1.outgoing_fds": 0,
"dbus.peer.org.freedesktop.systemd1.reply_objects": 0,
"dbus.stale_fds": 3,
"dbus.user.cooper.bytes": 919236,
"dbus.user.cooper.fds": 78,
"dbus.user.cooper.matches": 510,
"dbus.user.cooper.objects": 80,
"dbus.user.root.stale_fds": 3,
"networkd.eno4.address_state": 3,
"networkd.eno4.admin_state": 4,
"networkd.eno4.carrier_state": 5,
"networkd.eno4.ipv4_address_state": 3,
"networkd.eno4.ipv6_address_state": 2,
"networkd.eno4.oper_state": 9,
"networkd.eno4.required_for_online": 1,
"networkd.managed_interfaces": 2,
"networkd.wg0.address_state": 3,
"networkd.wg0.admin_state": 4,
"networkd.wg0.carrier_state": 5,
"networkd.wg0.ipv4_address_state": 3,
"networkd.wg0.ipv6_address_state": 3,
"networkd.wg0.oper_state": 9,
"networkd.wg0.required_for_online": 1,
"pid1.cpu_time_kernel": 48,
"pid1.cpu_user_kernel": 41,
"pid1.fd_count": 245,
"pid1.memory_usage_bytes": 19165184,
"pid1.tasks": 1,
"services.chronyd.service.active_enter_timestamp": 1683556542382710,
"services.chronyd.service.active_exit_timestamp": 0,
"services.chronyd.service.cpuusage_nsec": 328951000,
"services.chronyd.service.inactive_exit_timestamp": 1683556541360626,
"services.chronyd.service.ioread_bytes": 18446744073709551615,
"services.chronyd.service.ioread_operations": 18446744073709551615,
"services.chronyd.service.memory_available": 18446744073709551615,
"services.chronyd.service.memory_current": 5214208,
"services.chronyd.service.nrestarts": 0,
"services.chronyd.service.processes": 1,
"services.chronyd.service.restart_usec": 100000,
"services.chronyd.service.state_change_timestamp": 1683556542382710,
"services.chronyd.service.status_errno": 0,
"services.chronyd.service.tasks_current": 1,
"services.chronyd.service.timeout_clean_usec": 18446744073709551615,
"services.chronyd.service.watchdog_usec": 0,
"stat_collection_run_time_ms": 87.4013,
"system-state": 3,
"timers.fstrim.timer.accuracy_usec": 3600000000,
"timers.fstrim.timer.fixed_random_delay": 0,
"timers.fstrim.timer.last_trigger_usec": 1743397269608978,
"timers.fstrim.timer.last_trigger_usec_monotonic": 0,
"timers.fstrim.timer.next_elapse_usec_monotonic": 0,
"timers.fstrim.timer.next_elapse_usec_realtime": 1744007133996149,
"timers.fstrim.timer.persistent": 1,
"timers.fstrim.timer.randomized_delay_usec": 6000000000,
"timers.fstrim.timer.remain_after_elapse": 1,
"timers.fstrim.timer.service_unit_last_state_change_usec": 1743517244700135,
"timers.fstrim.timer.service_unit_last_state_change_usec_monotonic": 639312703,
"unit_files.root.generated.mount_units": 6,
"unit_files.root.generated.service_units": 1,
"unit_files.root.generated.socket_units": 1,
"unit_files.root.generated.swap_units": 1,
"unit_files.root.transient.scope_units": 1,
"unit_files.root.transient.service_units": 1,
"unit_files.user.transient.scope_units": 19,
"unit_files.user.transient.service_units": 15,
"unit_states.chronyd.service.active_state": 1,
"unit_states.chronyd.service.load_state": 1,
"unit_states.chronyd.service.time_in_state_usecs": 2176062087185,
"unit_states.chronyd.service.unhealthy": 0,
"units.activating_units": 0,
"units.active_units": 403,
"units.automount_units": 1,
"units.device_units": 150,
"units.failed_units": 0,
"units.inactive_units": 159,
"units.jobs_queued": 0,
"units.loaded_units": 497,
"units.masked_units": 25,
"units.mount_units": 52,
"units.not_found_units": 38,
"units.path_units": 4,
"units.scope_units": 17,
"units.service_units": 199,
"units.slice_units": 7,
"units.socket_units": 28,
"units.target_units": 54,
"units.timer_persistent_units": 1,
"units.timer_remain_after_elapse": 1,
"units.timer_units": 20,
"units.total_units": 562,
"varlink_usage.boot_blame": 1,
"varlink_usage.machines": 1,
"varlink_usage.networkd": 1,
"varlink_usage.system_state": 1,
"varlink_usage.units": 1,
"varlink_usage.verify": 1,
"varlink_usage.version": 1,
"verify.failing.device": 43,
"verify.failing.mount": 15,
"verify.failing.service": 31,
"verify.failing.slice": 1,
"verify.failing.total": 97,
"version": "255.7-1.fc40"
}
Normal serde_json pretty representations of each components structs.
monitord records the wall time each collector future spends inside a single
stat_collector cycle and exposes the result on MonitordStats::collector_timings,
plus an inner phase breakdown for the units collector
(SystemdUnitStats::collection_timings).
Each collector that can use varlink reports which transport served it on the
last run: varlink_usage.<collector> is 1 when the varlink attempt succeeded
and 0 when the collector fell back to D-Bus (or file-based collection for
networkd). Collectors with no varlink path (pid1, dbus_stats) and disabled
collectors emit no gauge, so the present gauges are exactly the enabled set β
no separate enabled-collectors counter is needed.
As the systemd under monitord upgrades past each endpoint's minimum version
(networkd v257+, system state/version v258+, units v260+, unit details v261+),
collectors silently flip from 0 to 1 with no config change, so watching these
gauges over time shows varlink adoption climbing across the fleet. Downstream
consumers such as monitord-exporter can aggregate them (share of collectors
reporting 1) from there; monitord itself only makes the gauges available in
its output formats. Per-container gauges are emitted under
machines.<name>.varlink_usage.<collector>. The host machines gauge covers
enumeration only (machined's io.systemd.Machine.List, v257+); per-container
collection reports its own gauges.
| Field | Meaning |
|---|---|
collector_timings.<name>.start_offset_ms |
ms from the top of the cycle until the spawned future was first polled. Should be sub-ms when collectors are running in parallel; a non-trivial value means the spawn loop or runtime is delaying first poll. |
collector_timings.<name>.elapsed_ms |
ms from first poll to completion for that collector. |
collector_timings.<name>.success |
1 if the collector returned Ok, 0 otherwise. |
collection_timings.list_units_ms |
ms for the systemd ListUnits D-Bus call (one batched call). |
collection_timings.per_unit_loop_ms |
ms spent walking each listed unit, including any per-unit D-Bus calls (timer/state/service). |
collection_timings.timer_dbus_fetches |
Count of timer D-Bus property fetches this run. |
collection_timings.state_dbus_fetches |
Count of unit-state D-Bus fetches (only when state_stats_time_in_state is enabled). |
collection_timings.service_dbus_fetches |
Count of per-service D-Bus property fetches. |
collection_timings.slowest_units |
The units.slowest_units_count slowest units this run (unit name, duration ms), descending. Empty when slowest_units_count = 0. D-Bus path only (see parity note below). |
Comparing sum(collector_timings.*.elapsed_ms) against
stat_collection_run_time_ms gives an effective parallelism ratio
(sum / wall β N means N-way parallelism, β 1 means effectively serial).
The per-collector lines are also emitted to logs at debug! level. The end-of-cycle
"stat collection run took {}ms" summary stays at info!.
collection_timings is populated identically by the D-Bus path
(units::parse_unit_state) and the varlink path
(varlink_units::parse_metrics plus the io.systemd.Unit.List detail pass).
In the varlink case, list_units_ms is the bulk varlink List call on
io.systemd.Manager and per_unit_loop_ms covers the local parse loop plus
the per-unit detail pass that fills service and timer stats.
All three *_dbus_fetches counters stay at zero on the varlink path, since
nothing there touches D-Bus any more: per-service stats, timer properties and
service types come from io.systemd.Unit.List, and time-in-state from the
StateChangeTimestamp metric. They only become nonzero if that socket is
unusable and monitord falls back. This makes varlink.enabled = true vs
false directly comparable on the same host.
Convention for new collectors moved to varlink: when porting a collector
from D-Bus to varlink, add the equivalent inner timings so the two
implementations remain comparable. The minimum is wall time of the bulk
fetch (analogous to list_units_ms) and the local parse loop (analogous to
per_unit_loop_ms), recorded onto a struct nested inside the collector's
public stats type. Single-shot varlink calls (e.g. networkd Describe) do
not need an inner split β the outer collector_timings.<name>.elapsed_ms
already covers them.
Many metrics are serialized as integers. Here are the enum mappings:
system-state
| Value | State |
|---|---|
| 0 | unknown |
| 1 | initializing |
| 2 | starting |
| 3 | running |
| 4 | degraded |
| 5 | maintenance |
| 6 | stopping |
| 7 | offline |
active_state (unit_states.*.active_state)
| Value | State |
|---|---|
| 0 | unknown |
| 1 | active |
| 2 | reloading |
| 3 | inactive |
| 4 | failed |
| 5 | activating |
| 6 | deactivating |
load_state (unit_states.*.load_state)
| Value | State |
|---|---|
| 0 | unknown |
| 1 | loaded |
| 2 | error |
| 3 | masked |
| 4 | not-found |
networkd address_state / ipv4_address_state / ipv6_address_state
| Value | State |
|---|---|
| 0 | unknown |
| 1 | off |
| 2 | degraded |
| 3 | routable |
networkd admin_state
| Value | State |
|---|---|
| 0 | unknown |
| 1 | pending |
| 2 | failed |
| 3 | configuring |
| 4 | configured |
| 5 | unmanaged |
| 6 | linger |
networkd carrier_state
| Value | State |
|---|---|
| 0 | unknown |
| 1 | off |
| 2 | no-carrier |
| 3 | dormant |
| 4 | degraded-carrier |
| 5 | carrier |
| 6 | enslaved |
networkd oper_state
| Value | State |
|---|---|
| 0 | unknown |
| 1 | missing |
| 2 | off |
| 3 | no-carrier |
| 4 | dormant |
| 5 | degraded-carrier |
| 6 | carrier |
| 7 | degraded |
| 8 | enslaved |
| 9 | routable |
networkd required_for_online
| Value | State |
|---|---|
| 0 | false |
| 1 | true |
| 255 | unknown |
Note the unknown sentinel is 255 (u8::MAX), not 0 β 0 means an
explicit false. Alerting rules should treat 255 as "no data", not as
"not required for online".
You're going to need to be root or allow permissiong to pull dbus stats.
For dbus-broker here is example config allow a user monitord to query
getStats
When stale_fd_stats = true, monitord also inspects
/proc/<dbus-broker>/fdinfo for stale pidfds reported by the kernel as
Pid: -1. These counters are emitted as dbus.stale_fds and, for metric
compatibility, dbus.user.root.stale_fds; the per-user root value is the
unattributed system-broker count, not proof that root owns those descriptors.
[cooper@l33t ~]# cat /etc/dbus-1/system.d/allow_monitord_stats.conf
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE busconfig PUBLIC
"-//freedesktop//DTD D-BUS Bus Configuration 1.0//EN"
"http://www.freedesktop.org/standards/dbus/1.0/busconfig.dtd">
<busconfig>
<policy user="monitord">
<allow send_destination="org.freedesktop.DBus"
send_interface="org.freedesktop.DBus.Debug.Stats"
send_member="GetStats"
send_path="/org/freedesktop/DBus"
send_type="method_call"/>
</policy>
</busconfig>
To do test runs (requires systemd and systemd-networkd installed)
Pending what you have enabled in your config ...
cargo run -- -c monitord.conf -l debug
Ensure the following pass before submitting a PR (CI checks):
cargo testcargo clippycargo fmtCargo.toml./build_docs.sh to regenerate docsMove to version X.Y.Z for release + update docsgh release create X.Y.Z --title "X.Y.Z" --generate-notescargo install zbus_xmlgenzbus-xmlgen system org.freedesktop.systemd1 /org/freedesktop/systemd1/unit/chronyd_2eserviceThen add the following macros to tell clippy to go away:
#![allow(warnings)]
#![allow(clippy)]
Sometimes I develop from my Mac OS X laptop. So I thought I'd document and add the way I build a Fedora Rawhide container and mount the local repo to /repo in the container to run monitord and test.
docker build -t monitord-dev .docker run --rm --name monitord-dev -it --privileged --tmpfs /run --tmpfs /tmp -v $(pwd):/repo monitord-dev /sbin/init--rm is optional but will remove the container when stoppedYou can now log into the container to build + run tests and run the binary now against systemd.
docker exec -it monitord-dev bashcd /repo ; cargo run -- -c monitordsystemctl start systemd-networkdvarlink_integration_test.py is what the varlink-integration GitHub Action
runs, so a CI failure reproduces locally with the same command:
python3 varlink_integration_test.pyIt rebuilds the image on every run (a no-op when the Dockerfile has not
changed) and recreates its container whenever the image changes, so a local run
uses the same image the Action builds from scratch instead of a stale one. It
reuses the running container otherwise; --fresh rebuilds without the cache.
The image must carry everything the test shells out to, which is why the
Dockerfile names systemd-container, systemd-networkd and shadow-utils
explicitly β see REQUIRED_CONTAINER_TOOLS, which fails the run up front with
one message if any of them is missing.
"Connection refused" or D-Bus connection errors
The system bus connection is created lazily: monitord starts fine without a bus as long as no enabled collector needs D-Bus. Collectors served entirely over varlink, cgroupfs, state files, sysfs or procfs never connect β the networkd file fallback maps ifindexes to names via /sys/class/net/*/ifindex and only reaches for Manager.ListLinks when sysfs yields nothing usable. A collector that does need the bus β including any varlink fallback β reports its own per-collector error and the run continues with the remaining collectors. If a D-Bus collector is failing, ensure the system D-Bus daemon is running and the socket exists at /run/dbus/system_bus_socket. If using a custom address, set dbus_address in [monitord] config. Increase dbus_timeout if running on slow systems.
Empty or missing networkd metrics
systemd-networkd must be installed and running (systemctl start systemd-networkd). If networkd is not in use on your system, disable the collector with enabled = false in [networkd].
Permission denied for machines / containers
A non-root monitord needs CAP_SYS_PTRACE to reach machines via /proc/<leader_pid>/root,
plus CAP_SYS_ADMIN to collect them over varlink. See Permissions under
Machines support.
Permission denied for D-Bus stats
The [dbus] collector requires permission to call org.freedesktop.DBus.Debug.Stats.GetStats. Either run monitord as root or add a D-Bus policy file β see the dbus stats section.
PID 1 stats unavailable
PID 1 stats require Linux with procfs mounted at /proc. This collector is compiled out on non-Linux targets. If /proc is not available (some container runtimes), disable with enabled = false in [pid1].
Collector errors don't crash monitord
When an individual collector fails (e.g., networkd not running, D-Bus timeout), monitord logs a warning and continues with the remaining collectors. Check stderr output or increase the log level (-l debug) to see which collectors had issues.
Large u64 values (18446744073709551615) in output
These represent u64::MAX and mean "not available" or "not tracked" for that metric. This is how systemd reports fields that are unsupported or not configured for the unit. The cgroup-derived service fields report u64::MAX only when neither cgroupfs nor the IPC fallback has data (e.g. ioread_bytes with the io controller disabled); memory_available in particular is computed even without limits, as host MemAvailable.
monitord can be used as a Rust library. See the full API documentation at monitord.xyz.
All monitord's dbus is done via async (tokio) zbus crate.
systemd Dbus APIs are in use in the following modules:
ManagerProxy::list_machines() β fallback only, when varlink enumeration
(io.systemd.Machine.List) is disabled or unavailableManagerProxy::list_links() β last resort only, when sysfs yields no usable ifindex map/run/systemd/netif/links are used by default, with the
ifindexβname map read from /sys/class/net/*/ifindex (no D-Bus); the varlink
io.systemd.Network.Describe API can be enabled instead (see below)ManagerProxy::get_version()ManagerProxy::system_state()TimerProxy::unit() - Find service unit of timerManagerProxy::get_unit()UnitProxy::state_change_timestamp()UnitProxy::state_change_timestamp_monotonic()ManagerProxy::list_units() - Main counting of unit statsServiceProxy::control_group() + ServiceProxy::main_pid() + ServiceProxy::control_pid() - Locate the unit's cgroup and fold PIDs living outside it into the process count, the way GetProcesses doesServiceProxy::nrestarts()ServiceProxy::restart_usec()ServiceProxy::status_errno()ServiceProxy::timeout_clean_usec()ServiceProxy::watchdog_usec()cpuusage_nsec, ioread_bytes, ioread_operations,
memory_current, memory_available, tasks_current, process count) is read from
cgroupfs (cpu.stat, memory.current/max/high, pids.current, io.stat,
cgroup.procs) instead of the ServiceProxy properties / get_processes().
Each field falls back to its D-Bus property when cgroupfs has no data (e.g. cgroup v1 hosts)UnitProxy::active_enter_timestampUnitProxy::active_exit_timestampUnitProxy::inactive_exit_timestamp()UnitProxy::state_change_timestamp() - Used for raw stat + time_in_stateSome of these modules can be disabled via configuration. Due to this, monitord might not always be running / calling all these DBus calls per run.
monitord supports collecting unit statistics via systemd's Varlink metrics API,
available in systemd v260+. When enabled, monitord connects to the io.systemd.Metrics interface
at /run/systemd/report/io.systemd.Manager to collect unit counts, active/load states, and restart counts.
Set enabled = true in the [varlink] section of monitord.conf:
[varlink]
enabled = true
When varlink is enabled, monitord will attempt to collect stats via the varlink APIs first, automatically falling back to D-Bus or file-based collection when a varlink socket is unavailable (e.g., older systemd versions).
Set no_fallback = true alongside it to turn any varlink failure into a loud per-collector
error instead of falling back. That is a verification mode for proving a collector set is
varlink-clean in CI β not a hardening flag: tripped collectors still report success=0 in
collector_timings and the run exits 0 (per-collector failures are deliberately non-fatal,
especially in daemon mode), so CI must assert on the success gauges, not the exit status.
no_fallback only fires inside varlink code paths, so it has no effect while [varlink]
enabled=false (a warning is logged) or on [dbus] stats, which have no varlink path
(they are statistics about the D-Bus daemon itself). Partial varlink data that parses with warnings (e.g. a
skipped metric) is not a fallback either β pair no_fallback with the completeness
assertions, not as a substitute for them.
Each varlink-capable collector ([units], [networkd], [system-state], [boot],
[verify], [machines] for machine enumeration and container collection) also has its own varlink toggle,
defaulting to true. A collector uses varlink only when both the global switch and its section toggle are
true, so collectors can be moved to varlink one at a time by setting a section toggle to
false. Container collection additionally requires [machines] varlink. Timers have no
toggle of their own: they ride the units path and follow [units] varlink.
[system-state] varlink also gates version collection, which shares the
Manager.Describe call β including when [system-state] itself is disabled.
Units (io.systemd.Metrics β systemd v260+, v261+ for jobs queued, exact totals, load-state totals, time-in-state, service errno, and the ActiveTimestamp/InactiveExitTimestamp families boot blame reads):
- Unit counts by type (service, mount, socket, target, device, automount, timer, path, slice, scope)
- Unit counts by state (activating, active, failed, inactive)
- Unit counts by load state (UnitsByLoadStateTotal, v261+: loaded, masked, not-found; counted from per-unit load states on v260)
- Exact total unit count (UnitsTotal, v261+; approximated from mapped per-type counts on v260)
- Queued job count (JobsQueued, v261+; unavailable (0) on v260)
- Per-unit active state and load state (with allowlist/blocklist filtering)
- Per-unit time in state (StateChangeTimestamp, v261+; unavailable on v260)
- Per-unit health status (computed from active + load state)
- Per-service restart counts (nrestarts)
- Per-service errno status (StatusErrno, v261+; unavailable (0) on v260)
- Full per-service stats for units in [services] via io.systemd.Unit.List (systemd v261+): timestamps, restart/timeout/watchdog settings from the reply, and CPU, memory, tasks, IO plus process count read from cgroupfs (same reader as the D-Bus path, so both agree). A field with no cgroupfs data falls back to the reply's runtime.CGroup value, defaulting to u64::MAX when the reply omits it too
- Per-timer stats via io.systemd.Unit.List (systemd v261+): accuracy, delays, next elapse, last trigger, and the triggered unit's state change β no D-Bus backfill
- Falls back to D-Bus collection if the socket is unavailable
System state and version (io.systemd.Manager.Describe β systemd v258+):
- Overall systemd system state (running, degraded, β¦)
- Version of the running systemd manager (which can trail the installed package
until PID 1 re-execs)
- Falls back to the D-Bus SystemState/Version properties if the socket is unavailable
Boot blame (io.systemd.Metrics β systemd v261+):
- Slowest units at boot, from the per-unit ActiveTimestamp/InactiveExitTimestamp metrics rather than a D-Bus property read per unit
- Falls back to D-Bus if the socket is unavailable
Verify (unit names from UnitLoadState metric objects β systemd v260+):
- Unit enumeration for systemd-analyze verify without D-Bus ListUnits; the metric objects are the same set ListUnits returns (verified live, including not-found units)
- The analyze run itself is unchanged and shared by both paths
- Falls back to D-Bus if the socket is unavailable
Networkd interfaces (io.systemd.Network.Describe β systemd v257+):
- Per-interface operational, carrier, admin, and address states
- Falls back to parsing /run/systemd/netif/links state files if the socket is unavailable
Machines (io.systemd.Machine.List on /run/systemd/machine/io.systemd.Machine β systemd v257+):
- Enumeration of containers and their leader PIDs, with the [machines] allowlist/blocklist
- Needs no extra privileges (unlike per-container collection, see below)
- Falls back to machined's D-Bus ListMachines if the socket is unavailable
For systemd-nspawn containers, monitord uses the same varlink endpoints as on the host,
inside the container: units from /run/systemd/report/io.systemd.Manager plus
io.systemd.Unit.List on /run/systemd/io.systemd.Manager, system state and version from
/run/systemd/io.systemd.Manager, and networkd from
/run/systemd/netif/io.systemd.Network (with the same file-based fallback). Unit file
and cgroup data are read below /proc/<leader_pid>/root.
systemd's credential-checking varlink servers (PID 1, networkd) refuse callers outside
their PID namespace (#211,
systemd/systemd#43807), so these
connections are made by a per-machine helper inside the container's PID namespace. That
needs CAP_SYS_ADMIN in addition to CAP_SYS_PTRACE; without it monitord warns once and
collects containers over D-Bus. See Permissions.
varlink might one day replace our DBUS usage. Here are some notes on how to work with systemd varlink
as there isn't really documentation outside man pages.
Here is an example with networkd's interfaces. monitord uses Describe, which
returns per-interface states (GetStates returns the manager-level aggregate
instead):
varlinkctl info unix:/run/systemd/netif/io.systemd.Network
varlinkctl introspect unix:/run/systemd/netif/io.systemd.Network io.systemd.Network
cooper@au:~$ varlinkctl call unix:/run/systemd/netif/io.systemd.Network io.systemd.Network.Describe '{}' | jq '.Interfaces[0]'
{
"Index": 2,
"Name": "eth0",
"AdministrativeState": "configured",
"OperationalState": "routable",
"CarrierState": "carrier",
"AddressState": "routable",
"IPv4AddressState": "routable",
"IPv6AddressState": "routable",
"OnlineState": "online",
"NetworkFile": "/etc/systemd/network/69-eno4.network",
"RequiredForOnline": true
}