Skip to content

Validating Reflector Availability

“Operational” and “accepting connections” sound like one property. They’re not, and conflating them is exactly how a poller ends up reporting a reflector as up when its D-Star instance has been crash-looped for a week, or as down when a firewall change on the monitor’s own network is the only thing that changed.

  • Operational - the reflector process is running and the mode instance’s listener bound successfully. This is a claim about internal state, answered by liveness.
  • Accepting connections - the advertised host:port is actually reachable from outside, past NAT, firewall, and routing. This is a claim about the network path, answered by a reachability probe.

The Public Interface guide makes the failure mode explicit: a raw TCP/UDP connect succeeding “proves nothing about the reflector logic behind it,” since a listener can accept a socket while the mode instance behind it is disabled or failed. The inverse gap matters just as much - a reflector can be fully operational internally while sitting behind a broken port forward. This page covers the liveness half in depth, since it’s the one most polling integrations skip, and closes with how to add the reachability half without duplicating what liveness already tells you.

Every Vexillum node that opts into the public API exposes a monitoring-focused snapshot of every configured instance, purpose-built for exactly this kind of external check. It’s deliberately kept separate from the dashboard-oriented /overview endpoint, so a polling integration doesn’t have to track new dashboard fields over time.

Route Returns
GET /api/public/v1/liveness Every instance on the node, one call.
GET /api/public/v1/liveness/:mode/:instance One instance, same fields plus found.
{
  "title": "Public Liveness",
  "runtime_name": "VEX Test Runtime",
  "generated_at": "2026-07-09T14:32:01Z",
  "totals": {
    "instances": 7, "running": 5, "stopped": 2,
    "failed": 0, "disabled": 0, "active_clients": 3
  },
  "instances": [
    {
      "mode": "DMR", "mode_value": "dmr", "instance": "main",
      "detail_path": "/api/public/v1/liveness/dmr/main",
      "state": "running", "state_class": "ok", "reason": "started",
      "state_changed_at": "2026-07-09T14:30:01Z",
      "uptime_seconds": 120, "active_clients": 1
    }
  ]
}

The one call that matters most for a multi-instance node: the bulk route returns every mode and every instance in a single round trip. A polling integration doesn’t need to know in advance which modes a listed node runs - store one base URL per VEX host, fetch /liveness once per poll cycle, and match every listing on that host against the instances[] array by mode_value + instance. One host, one request, full coverage, regardless of whether it’s running one mode or all seven, and regardless of how many instances of each mode it runs.

flowchart LR
    D["poller"] -- "GET /liveness\n(one request)" --> V
    subgraph V["one VEX host"]
        direction TB
        I1["dmr / main"]
        I2["dstar / main"]
        I3["m17 / relay"]
    end
    V -- "instances[]\n(all three, one response)" --> D

Fan-in, not fan-out. A three-instance host still costs one request. Requesting per-instance instead means one request per listing, on every poll, forever.

Field Present Meaning
state always One of running, starting, stopping, stopped, failed, disabled.
state_class always Collapses state to ok, warn, or err - see the state table below.
reason when non-empty Last transition: started, restarted, stopped, disabled, enabled, or the captured error text when state is failed.
state_changed_at always RFC3339 timestamp of the last transition. Use it for flap and staleness detection even while stopped.
uptime_seconds only when running Derived from state_changed_at. Resets on a Vexillum process restart, since instances run in-process and there’s no independent uptime to preserve.
active_clients always Current connected-client count for the instance.

This is worth grounding, because a monitoring integration is only as good as what its green light actually verifies. Each Vexillum process tracks every configured instance through an in-process supervisor with six states. Critically, the transition into running is gated on the mode’s actual runtime start call succeeding, which includes binding its listener socket. On process startup, and again whenever an operator starts or restarts an instance, if that call fails for any reason (port already in use, invalid stored config, anything else), the supervisor is set to failed with the error text captured in reason before the instance is ever exposed as running.

So state: "running" specifically means this mode’s listener bound and its runtime accepted its start call - a stronger claim than “an instance is configured” or “the operator marked it enabled.”

stateDiagram-v2
    direction LR
    [*] --> stopped: registered
    [*] --> disabled: registered, not enabled

    stopped --> running: Start(), bind ok
    stopped --> failed: Start(), bind fails
    failed --> running: Start(), bind ok

    running --> stopped: Stop()
    running --> failed: restart fails

    stopped --> disabled: Disable()
    running --> disabled: Disable()
    failed --> disabled: Disable()
    disabled --> stopped: Enable()

Every path into running passes through a bind attempt, from either stopped or failed. A successful Restart() isn’t drawn as its own edge: it re-enters running from running, only resetting uptime_seconds and updating reason to "restarted". starting and stopping are valid state_class inputs but aren’t produced by any transition above; handle them defensively as momentary, not as states you should expect to observe often.

A poller doesn’t need to reason about six raw state values. It needs a decision. Map every response, including the ones where the request itself fails, into four buckets: ok, warn, err, and unreachable.

Response state_class Bucket Read it as
state: running ok Up Listener bound, instance serving.
starting / stopping warn Transitional Mid-lifecycle. Only worth flagging if it doesn’t resolve within a poll cycle or two.
stopped warn Intentionally off Operator stopped it. Not an error, so don’t alert as one.
disabled warn Intentionally off Operator disabled the instance. Same treatment as stopped.
failed err Real fault Vexillum itself is reachable and diagnosable, since reason carries the actual error. Worth surfacing distinctly, since the operator can act on it immediately.
404, found: false - unreachable Node is up but this mode/instance doesn’t exist there - likely a stale or mistyped listing, not an outage.
Timeout, refused, non-2xx, non-JSON - unreachable Could be the node, the network path, or an operator who never enabled the public API. Indistinguishable from outside, so don’t call it “down.”

The warn and unreachable buckets are easy to collapse into “not ok” by mistake. Keep them apart: warn means the reflector answered you and told you it’s deliberately off; unreachable means you have no idea what’s happening on the other end. A public-facing status message should say different things for each.

  1. Poll the bulk endpoint once per VEX host. One GET {base_url}/api/public/v1/liveness per node per cycle covers every mode and instance it runs. Prefer this over per-instance calls whenever a host has more than one listed instance, since it’s strictly fewer requests for the same coverage.
  2. Match listings by mode_value + instance, case-insensitively. Instance lookups are case-insensitive server-side; the response always reports canonical casing. If you call a per-instance route directly and the name doesn’t match canonical casing, expect an HTTP 307 with a Location header - follow it rather than treating it as a failure.
  3. Classify into the four buckets above, not a binary up/down. Store the bucket, the raw state/reason, and state_changed_at alongside the listing. The raw fields are what let a human (or the listing owner) understand why, not just that.
  4. Debounce before changing a public-facing status. Use state_changed_at to avoid flapping a listing’s displayed status on a single transitional or unreachable read. Require the new bucket to hold across a couple of consecutive polls, or for a minimum dwell time, before it’s reflected publicly.
  5. Respect the rate limit. The default budget is 120 requests/minute per source IP, burst 20, shared across the whole public API (see Public API Reference). It’s generous for one bulk call per host per cycle even at a several-minute cadence, but back off on HTTP 429 and honor its Retry-After header rather than retrying immediately.
  6. Pick a cadence, not a persistent connection. Vexillum also offers an SSE stream of this snapshot for live dashboards, but a poller checking hundreds of listings doesn’t want hundreds of held-open connections. A one-shot bulk poll every 5-15 minutes is the pragmatic default; tighten it only for listings actively flagged as flapping.
sequenceDiagram
    participant P as poller
    participant V as VEX host

    P->>V: GET /api/public/v1/liveness
    alt normal poll
        V-->>P: 200 OK, instances[]
        P->>P: match by mode_value + instance
    else instance renamed since last poll
        P->>V: GET /liveness/dmr/Main
        V-->>P: 307, Location: /liveness/dmr/main
        P->>V: GET /liveness/dmr/main
        V-->>P: 200 OK
    else over the rate limit
        V-->>P: 429, Retry-After
        P->>P: wait, retry after header value
    end
    P->>P: classify into ok / warn / err / unreachable
    P->>P: hold new bucket for N polls before publishing

What one poll cycle looks like on the wire, including the two edge cases a naive integration usually gets wrong: treating a 307 as a failure instead of following it, and retrying a 429 immediately instead of backing off.

Liveness answers “operational.” Confirming “accepting connections” is a different question about the network path, not the internal state - it means probing the specific host:port the listing advertises to end users, and treating the result as a separate signal rather than folding it into the same verdict.

Why a bare connect proves less than it looks like

Section titled “Why a bare connect proves less than it looks like”

Every DV mode listener - D-Star, DMR, M17, NXDN, P25, and YSF - binds a UDP socket. Only VAFM also opens real TCP listeners, for its control and WebSocket surfaces. UDP has no handshake: a client-side connect() call on a UDP socket is a purely local operation that never touches the network, so a script that “connects” and calls that a reachability check has verified nothing. It will happily report success against a completely unreachable host. The only way to learn anything over UDP is to send a datagram and see whether one comes back, and what comes back (or doesn’t) depends entirely on what you sent.

flowchart TD
    S["send one UDP datagram\nto the advertised host:port"] --> Q{"what was sent?"}
    Q -- "garbage / malformed bytes" --> M["parse error\nno reply, ever"]
    Q -- "valid protocol frame\n(DExtra poll,\nDMR RPTL login)" --> R["session created\non the reflector"]
    R --> A["reply sent\n(poll ack / login ack)"]
    M --> T1["timeout"]
    A --> T2["reply received"]
    T1 --> V1["identical to: firewalled,\ndropped, or wrong port"]
    T2 --> V2["port is open, and the mode\ninstance is speaking its protocol"]

Only a real protocol frame produces a signal, and producing it isn’t free. Malformed bytes get silently dropped whether or not the port is reachable, so a timeout there tells you nothing. Getting a real answer means the reflector accepted what looked like the start of a client connecting.

Traced from the mode implementations themselves, since those are what Vexillum actually does on the first packet from an unknown source address. See the DExtra and DMR protocol references for the full wire formats.

Mode Send Frame Reply on success Side effect
D-Star / DExtra DEXP poll 4-byte tag "DEXP" + 10-byte callsign + 1-byte module (15 bytes) DEXO, 10-byte reflector name + module New source address registers as a session; connect count increments immediately
DMR (Homebrew) RPTL login 4-byte tag "RPTL" + 4-byte repeater ID (8 bytes) RPTACK + a random auth challenge seed Same behavior: session registers and connect count increments before any password is checked

Note what the DMR row does not require: a correct password. RPTL is the first message of the login handshake, sent before the credential exchange (RPTK) that a polling integration obviously can’t complete for someone else’s reflector. It still gets an RPTACK back, and that ack alone is proof the port is open and speaking the DMR protocol, with no need to actually log in.

Three practical consequences follow:

  • Use a recognizable probe identity. Send a distinctive reserved placeholder callsign, something like PROBE1, so an operator glancing at their own client list or last-heard recognizes a monitoring probe instead of wondering who UNKNOWN was.
  • Probe sparingly, not on every poll cycle. Liveness is free to check as often as the rate limit allows, but a reachability probe touches someone else’s session table. Reserve it for listing submission or verification, and for spot-checking a listing that liveness alone has already flagged. It isn’t meant to be a routine companion to every bulk poll.
  • Retry a couple of times before calling it unreachable, never once. UDP has no delivery guarantee in either direction, so a single lost datagram looks identical to a closed port. Two or three attempts with a short timeout (a second or two) is enough to separate real unreachability from ordinary packet loss without turning the probe into a load generator.

The other four DV modes (M17, NXDN, P25, YSF) each run their own login/poll sequence over the same kind of unauthenticated-first-packet UDP protocol. The pattern generalizes, but the exact bytes don’t - confirm each mode’s specific handshake (see the Mode Protocol Reference) before scripting a probe for it rather than assuming it matches D-Star or DMR. VAFM is the outlier worth calling out separately: its control and WebSocket surfaces are real TCP, so a standard TCP connect there is a meaningful, side-effect-free reachability signal on its own. The UDP caveats above simply don’t apply to it.

The instance-detail endpoint includes a public_runtime.rows array with human-readable connect endpoints per mode. That’s useful for cross-checking, but it isn’t a substitute for the address a listing already has on file: probe the address actually advertised to users, not whatever the node reports about itself. With both signals in hand, resolve the question this page opened with:

Liveness Reachability Reported status
ok open Up. Both halves confirmed.
ok closed / timeout Operational, unreachable. Vexillum is fine, but the network path to it isn’t. Point the operator at NAT or a firewall rule, not their config.
warn / err not probed No need to probe, since liveness already explains the state. Report it as-is.
unreachable not probed Unknown. Skip the probe until the API itself answers again. A probe result here can’t disambiguate a down host from an operator who never enabled [public_api].
Verdict Rule
ok state_class: "ok" on the matched instance. Show as up; keep uptime_seconds visible if displaying detail.
warn state_class: "warn", or found: false resolved through a redirect. Show as “listed, not currently serving,” not an error.
err state_class: "err" (state: "failed"). Show as a fault, and surface reason if the listing owner can see it.
unreachable Timeout, connection refused, non-JSON, or a real 404. Show as “status unknown,” never as down, until it persists past your debounce window.

Grounded in Vexillum Public API v1 (Beta). The field set may grow; the four-bucket model above tolerates new fields without changing the verdict.