Field · Remote Radio Infrastructure

The borrowed light.

An amateur radio station usually has no public address: carrier NAT, sometimes twice, no port forwarding, no UPnP on a router nobody administers. Yet a radio meant to be used remotely needs an entry on the internet. This is the design that gives it one without asking anyone to touch their network, and the four defaults that made a working install look dead.

1 outbound port: 8989 entry = https://<callsign>.mrrc.vlsc.net/ per-instance certificate pinning routes generated, never hand-edited

By BG1SB  ·   ·  ~12 min read

The interesting part of this system is not the tunnel. Tunnels are a solved shape. The interesting part is what has to be true for a stranger's radio to be reachable at a name derived from their callsign, for that name to be verifiable rather than merely forwarded, and for adding the next radio to be a one-line change instead of an editing session on a production web server.

TL;DR — the design in five claims

  • FactThe instance dials out; the hub never dials in. A tunnel client connects to one port, 8989, and maps the local radio server onto a private loopback port on the hub. The customer's network needs zero inbound access.
  • FactEach instance is verified against its own certificate. Instances self-sign; the hub's trust bundle is generated from the registry, so only registered instances are trusted and renewals need no cross-machine synchronisation.
  • InferenceGenerated routes are what make growth cheap. The registry is label, port, upstream name per line; one command emits the nginx map and the wildcard vhost. Adding an instance is one line.
  • ThesisThe self-service portal performs no privileged action. It can record an application and sign an instance certificate. It cannot reload the web server or touch a firewall, and the two commands that need root are printed for an operator instead.
  • ExcludedNot covered here: radio behaviour. Everything below was verified on machines without a transceiver attached. CAT, audio and transmit paths need physical hardware and are outside this record.

No door, and no permit to cut one

Every design decision below follows from this table, not from preference.

ConstraintConsequence
No public IP, no port forwarding, no UPnP at the instance The entry cannot live at the client, so the tunnel must be outbound
Some networks permit only 80 and 443; others permit almost anything Exactly one control port, and it must be configurable rather than assumed
The user operates a radio, not a server Enrolment has to be a few clicks inside the application, with the operator's approval happening somewhere else entirely
One device belongs to one callsign The entry name is the callsign itself: https://bg9aaa.mrrc.vlsc.net/, with no port, path or product suffix
Instances may run on Windows, macOS, Linux or a Pi Nothing in the path may assume a shell, a package manager or a writable program directory
Fact · measured, not assumed

The port choice was settled by measurement rather than by taste. A second user-facing port that had existed for weeks turned out never to have been reachable from the public internet at all: three different networks refused it while the same host answered 443 without complaint. Applications written before that discovery had been quietly failing on every request except the one that happened to use a fallback path.

Four that held, each paid for once

Each one exists because removing it produced an incident.

① Outbound only

A tunnel client on the instance connects to tunnel.mrrc.vlsc.net:8989 and maps the radio's local port onto a loopback port on the hub, such as 127.0.0.1:18803. The hub's web server proxies the callsign hostname to that loopback port. The instance needs one outbound TCP connection and nothing else. If the tunnel drops it reconnects on its own, every few seconds, with no user action.

② Each is verified against its own certificate

The instance self-signs at first run, with its entry name as the subject. The hub's upstream verification points at a bundle generated from the registry: system roots plus every registered instance's public key. Two properties follow. Renewals need no cross-machine synchronisation, and an instance that is not in the registry cannot be served even if it somehow opens a tunnel.

③ Routes are generated

The registry holds one line per instance: label, loopback port, and the name the upstream certificate is expected to carry. A single command renders the nginx map and the wildcard vhost. Adding an instance is one line plus a regeneration. Nobody edits the web server's configuration by hand, which is the only reason a non-expert can operate this at all.

④ The portal holds no privileges

The self-service portal runs on loopback as an unprivileged service account. It records applications, verifies callsigns against a callsign database, and signs instance certificates. It cannot reload nginx and cannot change a firewall. The two steps that need root are printed as commands for an operator to run. A compromised portal therefore gains no machine.

upstream SSL certificate verify error: (18:self-signed certificate) while SSL handshaking to upstream

上游 SSL 证书校验失败:(18:自签名证书)发生在与上游进行 SSL 握手时。

nginx error log, hub host, 2026-10-01

That single line is invariant ② failing in the only way it can fail visibly. The tunnel was up, the application was serving, the browser reached the hub, and the hub refused its own upstream because the instance's certificate was not in the bundle. Fixing it is a regeneration, not a debugging session, provided the certificate file is named after the registry label.

Inference · why pinning beats a public CA here

A public certificate for each instance would be simpler to reason about and impossible to issue: the names are per-instance hostnames under a wildcard the operator controls, and every renewal would become a coordinated event between two machines. Pinning the self-signed certificate trades that coordination for a generation step that already has to exist, because the registry is the source of truth for routes anyway.

Four silences: no certificate, no logs, no port, no password

Reported as "installed it and it will not start". It was running the whole time.

On the first day the hub went live, a customer installed the application and saw a black screen. The service was up, the port was listening, and the log said so. What the log also said, in a warning that reads like a footnote:

SSL disabled or cert/key not found (cert=C:\Program Files\MRRC Modern\_internal\certs\fullchain.pem key=...\radio.vlsc.net.key)

SSL 已停用,或找不到证书与私钥。证书路径在安装目录内部,私钥文件名是开发机上那一张。

MRRC Modern server log, customer machine, 2026-10-02
DefaultWhat happened on a customer machine
Certificate path inside the packaged directory The program directory is read-only for a normal user, so no certificate was found; the server silently fell back to plain HTTP while the launcher opened an https:// address. The browser reported a protocol error. The screen stayed black.
Log directory inside the packaged directory A permissions error on every start, and therefore no logs at all, which removed the only way to diagnose the previous row.
Serial port defaulting to a macOS device name On Windows the radio was sitting on COM5 while the application looked for /dev/cu.SLAB_USBtoUART, a path that cannot exist there.
Login password defaulting to a constant in the source Anyone on the local network could authenticate until somebody changed it. Only reachable when the launcher, which generates a random one, was bypassed.

Three structural changes closed the class, rather than the four instances of it:

  • Every path the program writes to at runtime defaults into the user's own data directory, resolved per platform and never allowed to raise.
  • A missing certificate is signed on the spot. Serving plain HTTP now requires an explicit flag, and says so loudly when used.
  • A gate test fails the suite if any writable default resolves inside the program directory. It caught a fifth case, a memory-channels file, on its first run.
Thesis · the silent fallback is the actual defect

A wrong default is recoverable if it announces itself. What turned a bad path into an unusable product was the decision to degrade quietly: no certificate became no TLS, no writable directory became no logs, and the user was left with a black rectangle and no evidence. Every fallback in this codebase now either fixes itself or reports itself.

Two defects that exist only in an order

Neither shows up in a unit test, because each is a property of an ordering.

A new certificate, an old process

The TLS context is built once, when the server starts. Enrolling writes a fresh certificate to the same path. The running process keeps serving the previous one while the hub verifies against the new one, so the entry answers 502 with everything individually healthy. The nastiest ordering is restart first, enrol second: the file changes while the process is already up, and any check based on a snapshot taken at import time reports that nothing needs doing.

The criterion that survives every ordering is a comparison of two timestamps: if the certificate file was written after this process started, a restart is required. The interface now says so and offers the restart, and the restart refuses to exit unless it has already started something that will bring the server back.

A tunnel that waits to be asked for

The tunnel used to start from exactly one place: the endpoint the settings dialog calls to render itself. The consequence is that a machine which reboots overnight serves nothing until somebody opens that dialog in the morning. Found during an end-to-end acceptance run, not during development, because the developer had just clicked through the dialog.

An instance that is already enrolled now starts its tunnel at boot. Failure to start is logged and does not prevent the server from coming up, and the dialog still retries, so there are two paths to the same state rather than one.

Inference · why sequence bugs survive review

Both defects are invisible in any single state. Each component behaved correctly: the server served the certificate it had been given, the tunnel started when asked. The defect lived in the gap between one component's write and another's read. Reviewing components cannot find it; running the whole flow from a clean machine, in an inconvenient order, can.

Four rules, each paid for in wasted cycles

Each one cost at least one wasted cycle.

  • Never accept an exit code as proof. A build reported success, the log contained the line saying compilation succeeded, and the version file named the new release. The artifact was the previous one. Proof is a version, a timestamp that includes the year, and a hash different from the last release.
  • Never rebuild under a published version number. The upgrade channel compares version strings, so a machine that already installed that number will never see the rebuild. Its update button replays the same version forever.
  • One driver per release train. Two writers on one tree produced a manifest carrying a new version label with the previous version's size and hash. The fix is to generate the manifest on the server from the files actually published.
  • Every field failure becomes a test the same day. Otherwise it returns under another name. Of the gates written during this work, one found a new bug immediately and another has prevented a whole category from shipping since.
LayerWhere its truth is writtenRead this first when an entry fails
Tunnel The tunnel client's own log on the instance Whether it logged in and registered its proxy
Hub listener Loopback sockets owned by the tunnel server Whether the instance's port is listening at all
Upstream TLS The web server's error log error 18 means the trust bundle; a name mismatch means the registry's third column
Application The instance's own log, in the user's data directory Whether it is serving TLS at all, and which certificate
Network path Nothing; only a probe Whether the port is reachable from outside, tested from two networks

Three steps, and not one shell

No shell, no configuration file, no port forwarding.

  1. Install the application for the platform in use.
  2. Open the settings dialog, choose the cloud entry, type the callsign and submit the application. If the operator has already approved it and handed over a one-time enrolment secret, pasting that secret into the same dialog claims the approved entry directly and skips the wait.
  3. Restart when asked. The application enrols its own certificate, writes the tunnel configuration, starts the tunnel and prints the entry.

The result is a name derived from the callsign and nothing else:

https://bg9aaa.mrrc.vlsc.net/          # 443, no port, no path, no product suffix
https://bg9aaa.mrrc.vlsc.net/login     # the login page
https://bg9aaa.mrrc.vlsc.net/api/health # 401 while healthy: alive, and asking who you are

When the launcher will not start

The server runs without it. Start the server entry from the start menu, open http://127.0.0.1:8888/, and read the password the window prints. The launcher only seeds configuration, starts the server and opens a browser.

When the environment is a mess

A cleanup helper backs up configuration, certificates and logs to the desktop, then removes the application, its settings, its environment variables and its scheduled tasks. It touches nothing outside a fixed list of paths and leaves unrelated products alone unless explicitly asked.

When only the interface needs fixing

Interface and asset fixes ship as a few-hundred-kilobyte patch instead of a full installer. The patch verifies its own hash and the version it applies to, backs up every file it replaces, and restarts the application. Compiled logic cannot be patched and still needs a full release.

When none of this explains it

One button in the interface builds a diagnosis bundle: log tails, a whitelist of environment keys, and a summary that has already been read once by a machine before a human sees it. It uploads over the same 443 the site uses.

Fact · what was verified, and on what

The full sequence was run end to end on real machines: revoke an instance, apply from the application, approve in the portal, refresh, enrol, restart, and reach the entry with a 401 from the health endpoint. It was run on a Windows virtual machine and on a clean macOS machine, including a fresh install from the public download. What was not verified anywhere in this article is radio behaviour: no transceiver was attached to any of those machines.