Validating Reflector Availability
Two different claims, one check each
Section titled “Two different claims, one check each”“Operational” and “accepting connections” sound like one property. They’re not, and conflating them is exactly how a poller ends up reporting a reflector as up when its D-Star instance has been crash-looped for a week, or as down when a firewall change on the monitor’s own network is the only thing that changed.
- Operational - the reflector process is running and the mode instance’s listener bound successfully. This is a claim about internal state, answered by liveness.
- Accepting connections - the advertised host:port is actually reachable from outside, past NAT, firewall, and routing. This is a claim about the network path, answered by a reachability probe.
The Public Interface guide makes the failure mode explicit: a raw TCP/UDP connect succeeding “proves nothing about the reflector logic behind it,” since a listener can accept a socket while the mode instance behind it is disabled or failed. The inverse gap matters just as much - a reflector can be fully operational internally while sitting behind a broken port forward. This page covers the liveness half in depth, since it’s the one most polling integrations skip, and closes with how to add the reachability half without duplicating what liveness already tells you.
The liveness endpoint
Section titled “The liveness endpoint”Every Vexillum node that opts into the public API exposes a monitoring-focused
snapshot of every configured instance, purpose-built for exactly this
kind of external check. It’s deliberately kept separate from the
dashboard-oriented /overview endpoint, so a polling integration doesn’t have
to track new dashboard fields over time.
| Route | Returns |
|---|---|
GET /api/public/v1/liveness |
Every instance on the node, one call. |
GET /api/public/v1/liveness/:mode/:instance |
One instance, same fields plus found. |
The one call that matters most for a multi-instance node: the bulk route
returns every mode and every instance in a single round trip. A polling
integration doesn’t need to know in advance which modes a listed node runs -
store one base URL per VEX host, fetch /liveness once per poll cycle, and
match every listing on that host against the instances[] array by
mode_value + instance. One host, one request, full coverage, regardless of
whether it’s running one mode or all seven, and regardless of how many instances
of each mode it runs.
flowchart LR
D["poller"] -- "GET /liveness\n(one request)" --> V
subgraph V["one VEX host"]
direction TB
I1["dmr / main"]
I2["dstar / main"]
I3["m17 / relay"]
end
V -- "instances[]\n(all three, one response)" --> D
Fan-in, not fan-out. A three-instance host still costs one request. Requesting per-instance instead means one request per listing, on every poll, forever.
Field reference
Section titled “Field reference”| Field | Present | Meaning |
|---|---|---|
state |
always | One of running, starting, stopping, stopped, failed, disabled. |
state_class |
always | Collapses state to ok, warn, or err - see the state table below. |
reason |
when non-empty | Last transition: started, restarted, stopped, disabled, enabled, or the captured error text when state is failed. |
state_changed_at |
always | RFC3339 timestamp of the last transition. Use it for flap and staleness detection even while stopped. |
uptime_seconds |
only when running |
Derived from state_changed_at. Resets on a Vexillum process restart, since instances run in-process and there’s no independent uptime to preserve. |
active_clients |
always | Current connected-client count for the instance. |
Why running is trustworthy
Section titled “Why running is trustworthy”This is worth grounding, because a monitoring integration is only as good as
what its green light actually verifies. Each Vexillum process tracks every
configured instance through an in-process supervisor with six states.
Critically, the transition into running is gated on the mode’s actual
runtime start call succeeding, which includes binding its listener socket. On
process startup, and again whenever an operator starts or restarts an
instance, if that call fails for any reason (port already in use, invalid
stored config, anything else), the supervisor is set to failed with the
error text captured in reason before the instance is ever exposed as
running.
So state: "running" specifically means this mode’s listener bound and its
runtime accepted its start call - a stronger claim than “an instance is
configured” or “the operator marked it enabled.”
stateDiagram-v2
direction LR
[*] --> stopped: registered
[*] --> disabled: registered, not enabled
stopped --> running: Start(), bind ok
stopped --> failed: Start(), bind fails
failed --> running: Start(), bind ok
running --> stopped: Stop()
running --> failed: restart fails
stopped --> disabled: Disable()
running --> disabled: Disable()
failed --> disabled: Disable()
disabled --> stopped: Enable()
Every path into running passes through a bind attempt, from either
stopped or failed. A successful Restart() isn’t drawn as its own edge:
it re-enters running from running, only resetting uptime_seconds and
updating reason to "restarted". starting and stopping are valid
state_class inputs but aren’t produced by any transition above; handle them
defensively as momentary, not as states you should expect to observe often.
From six states to one verdict
Section titled “From six states to one verdict”A poller doesn’t need to reason about six raw state values. It needs a
decision. Map every response, including the ones where the request itself
fails, into four buckets: ok, warn, err, and unreachable.
| Response | state_class |
Bucket | Read it as |
|---|---|---|---|
state: running |
ok |
Up | Listener bound, instance serving. |
starting / stopping |
warn |
Transitional | Mid-lifecycle. Only worth flagging if it doesn’t resolve within a poll cycle or two. |
stopped |
warn |
Intentionally off | Operator stopped it. Not an error, so don’t alert as one. |
disabled |
warn |
Intentionally off | Operator disabled the instance. Same treatment as stopped. |
failed |
err |
Real fault | Vexillum itself is reachable and diagnosable, since reason carries the actual error. Worth surfacing distinctly, since the operator can act on it immediately. |
404, found: false |
- | unreachable |
Node is up but this mode/instance doesn’t exist there - likely a stale or mistyped listing, not an outage. |
| Timeout, refused, non-2xx, non-JSON | - | unreachable |
Could be the node, the network path, or an operator who never enabled the public API. Indistinguishable from outside, so don’t call it “down.” |
The warn and unreachable buckets are easy to collapse into “not ok” by
mistake. Keep them apart: warn means the reflector answered you and told
you it’s deliberately off; unreachable means you have no idea what’s
happening on the other end. A public-facing status message should say
different things for each.
The check, end to end
Section titled “The check, end to end”- Poll the bulk endpoint once per VEX host. One
GET {base_url}/api/public/v1/livenessper node per cycle covers every mode and instance it runs. Prefer this over per-instance calls whenever a host has more than one listed instance, since it’s strictly fewer requests for the same coverage. - Match listings by
mode_value+instance, case-insensitively. Instance lookups are case-insensitive server-side; the response always reports canonical casing. If you call a per-instance route directly and the name doesn’t match canonical casing, expect anHTTP 307with aLocationheader - follow it rather than treating it as a failure. - Classify into the four buckets above, not a binary up/down. Store the
bucket, the raw
state/reason, andstate_changed_atalongside the listing. The raw fields are what let a human (or the listing owner) understand why, not just that. - Debounce before changing a public-facing status. Use
state_changed_atto avoid flapping a listing’s displayed status on a single transitional or unreachable read. Require the new bucket to hold across a couple of consecutive polls, or for a minimum dwell time, before it’s reflected publicly. - Respect the rate limit. The default budget is 120 requests/minute per
source IP, burst 20, shared across the whole public API (see
Public API Reference). It’s generous for one
bulk call per host per cycle even at a several-minute cadence, but back
off on
HTTP 429and honor itsRetry-Afterheader rather than retrying immediately. - Pick a cadence, not a persistent connection. Vexillum also offers an SSE stream of this snapshot for live dashboards, but a poller checking hundreds of listings doesn’t want hundreds of held-open connections. A one-shot bulk poll every 5-15 minutes is the pragmatic default; tighten it only for listings actively flagged as flapping.
sequenceDiagram
participant P as poller
participant V as VEX host
P->>V: GET /api/public/v1/liveness
alt normal poll
V-->>P: 200 OK, instances[]
P->>P: match by mode_value + instance
else instance renamed since last poll
P->>V: GET /liveness/dmr/Main
V-->>P: 307, Location: /liveness/dmr/main
P->>V: GET /liveness/dmr/main
V-->>P: 200 OK
else over the rate limit
V-->>P: 429, Retry-After
P->>P: wait, retry after header value
end
P->>P: classify into ok / warn / err / unreachable
P->>P: hold new bucket for N polls before publishing
What one poll cycle looks like on the wire, including the two edge cases a naive integration usually gets wrong: treating a 307 as a failure instead of following it, and retrying a 429 immediately instead of backing off.
Adding the reachability half
Section titled “Adding the reachability half”Liveness answers “operational.” Confirming “accepting connections” is a different question about the network path, not the internal state - it means probing the specific host:port the listing advertises to end users, and treating the result as a separate signal rather than folding it into the same verdict.
Why a bare connect proves less than it looks like
Section titled “Why a bare connect proves less than it looks like”Every DV mode listener - D-Star, DMR, M17, NXDN, P25, and YSF - binds a UDP
socket. Only VAFM also opens real TCP listeners, for its control and
WebSocket surfaces. UDP has no handshake: a client-side connect() call on a
UDP socket is a purely local operation that never touches the network, so a
script that “connects” and calls that a reachability check has verified
nothing. It will happily report success against a completely unreachable
host. The only way to learn anything over UDP is to send a datagram and see
whether one comes back, and what comes back (or doesn’t) depends entirely on
what you sent.
flowchart TD
S["send one UDP datagram\nto the advertised host:port"] --> Q{"what was sent?"}
Q -- "garbage / malformed bytes" --> M["parse error\nno reply, ever"]
Q -- "valid protocol frame\n(DExtra poll,\nDMR RPTL login)" --> R["session created\non the reflector"]
R --> A["reply sent\n(poll ack / login ack)"]
M --> T1["timeout"]
A --> T2["reply received"]
T1 --> V1["identical to: firewalled,\ndropped, or wrong port"]
T2 --> V2["port is open, and the mode\ninstance is speaking its protocol"]
Only a real protocol frame produces a signal, and producing it isn’t free. Malformed bytes get silently dropped whether or not the port is reachable, so a timeout there tells you nothing. Getting a real answer means the reflector accepted what looked like the start of a client connecting.
Two examples
Section titled “Two examples”Traced from the mode implementations themselves, since those are what Vexillum actually does on the first packet from an unknown source address. See the DExtra and DMR protocol references for the full wire formats.
| Mode | Send | Frame | Reply on success | Side effect |
|---|---|---|---|---|
| D-Star / DExtra | DEXP poll |
4-byte tag "DEXP" + 10-byte callsign + 1-byte module (15 bytes) |
DEXO, 10-byte reflector name + module |
New source address registers as a session; connect count increments immediately |
| DMR (Homebrew) | RPTL login |
4-byte tag "RPTL" + 4-byte repeater ID (8 bytes) |
RPTACK + a random auth challenge seed |
Same behavior: session registers and connect count increments before any password is checked |
Note what the DMR row does not require: a correct password. RPTL is the
first message of the login handshake, sent before the credential exchange
(RPTK) that a polling integration obviously can’t complete for someone
else’s reflector. It still gets an RPTACK back, and that ack alone is proof
the port is open and speaking the DMR protocol, with no need to actually log
in.
Three practical consequences follow:
- Use a recognizable probe identity. Send a distinctive reserved
placeholder callsign, something like
PROBE1, so an operator glancing at their own client list or last-heard recognizes a monitoring probe instead of wondering whoUNKNOWNwas. - Probe sparingly, not on every poll cycle. Liveness is free to check as often as the rate limit allows, but a reachability probe touches someone else’s session table. Reserve it for listing submission or verification, and for spot-checking a listing that liveness alone has already flagged. It isn’t meant to be a routine companion to every bulk poll.
- Retry a couple of times before calling it unreachable, never once. UDP has no delivery guarantee in either direction, so a single lost datagram looks identical to a closed port. Two or three attempts with a short timeout (a second or two) is enough to separate real unreachability from ordinary packet loss without turning the probe into a load generator.
The other four DV modes (M17, NXDN, P25, YSF) each run their own login/poll sequence over the same kind of unauthenticated-first-packet UDP protocol. The pattern generalizes, but the exact bytes don’t - confirm each mode’s specific handshake (see the Mode Protocol Reference) before scripting a probe for it rather than assuming it matches D-Star or DMR. VAFM is the outlier worth calling out separately: its control and WebSocket surfaces are real TCP, so a standard TCP connect there is a meaningful, side-effect-free reachability signal on its own. The UDP caveats above simply don’t apply to it.
Combining the two signals
Section titled “Combining the two signals”The instance-detail endpoint
includes a public_runtime.rows array with human-readable connect endpoints
per mode. That’s useful for cross-checking, but it isn’t a substitute for the
address a listing already has on file: probe the address actually advertised
to users, not whatever the node reports about itself. With both signals in
hand, resolve the question this page opened with:
| Liveness | Reachability | Reported status |
|---|---|---|
ok |
open | Up. Both halves confirmed. |
ok |
closed / timeout | Operational, unreachable. Vexillum is fine, but the network path to it isn’t. Point the operator at NAT or a firewall rule, not their config. |
warn / err |
not probed | No need to probe, since liveness already explains the state. Report it as-is. |
unreachable |
not probed | Unknown. Skip the probe until the API itself answers again. A probe result here can’t disambiguate a down host from an operator who never enabled [public_api]. |
Cheat sheet
Section titled “Cheat sheet”| Verdict | Rule |
|---|---|
ok |
state_class: "ok" on the matched instance. Show as up; keep uptime_seconds visible if displaying detail. |
warn |
state_class: "warn", or found: false resolved through a redirect. Show as “listed, not currently serving,” not an error. |
err |
state_class: "err" (state: "failed"). Show as a fault, and surface reason if the listing owner can see it. |
unreachable |
Timeout, connection refused, non-JSON, or a real 404. Show as “status unknown,” never as down, until it persists past your debounce window. |
Grounded in Vexillum Public API v1 (Beta). The field set may grow; the four-bucket model above tolerates new fields without changing the verdict.