Post-mortemWorking with AI2026-10-02
Highbrow and Lowbrow
— On Muscle and Mind
Forty-eight hours, two models, six releases, one Windows VM, one Mac, one Hong Kong VPS, and a tunnel that would not come up. This is not about how capable the models were. It is about the one thing those hours made clear: the thing that works and the thing that thinks are different organs.
1. Two kinds of "it works"
Two in the morning, staring at [INFO] mrrc: Server ready!.
The port is listening, the processes are up, the build script exited 0,
and the log says Successful compile. The only problem:
the user opens a browser and sees black.
A default value was dug out of the package:
cert=C:\Program Files\MRRC Modern\_internal\certs\fullchain.pem
key =C:\Program Files\MRRC Modern\_internal\certs\radio.vlsc.net.key
The build machine's paths, baked into a certificate default shipped to
customers. On anyone else's machine neither file exists, so the server
quietly fell back to plain HTTP - while the launcher opened an
https:// URL. The browser reported a protocol error; the
screen stayed black.
Three more of the same that night: logs defaulting into
C:\Program Files\ ([WinError 5], no logs at
all), a serial port defaulting to macOS's
/dev/cu.SLAB_USBtoUART while the radio sat on
COM5, and a login password defaulting to the one written in
the source.
Four bugs, one disease: defaults that only hold on the build machine. And each one cost a full release - build, install, reproduce on real hardware - to find.
2. Lowbrow: what muscle does
The actions that actually pinned things down were none of them clever:
| Action | What it bought |
|---|---|
| Pull the exe back from the VM and hash it locally | The "successful" build was the previous release - same size, same hash, older timestamp |
| Read timestamps with the year | Nearly shipped a two-week-old artifact; now a hard rule |
| Read one line of TOML byte by byte |
A backslash inside a quoted string is an escape;
\U became a Unicode escape and frpc refused the file.
Three hours of reasoning that "the config looks fine" lost to one
line of bytes
|
| Run the server bare in a Windows console and screenshot it | Four warnings in one image; four bugs at once |
ssh to the VPS and run ufw status |
"The cloud panel is not blocking 8989" was wrong - the machine's own firewall was |
Strip every non-ASCII character out of a .ps1 |
PowerShell 5.1 reads BOM-less files as GBK; Chinese text breaks string terminators. Three times. |
They share one property: they all touch the real world - real kernels, real encodings, real firewalls, real antivirus, real desktops. They produce no insight, only facts.
And the real world's favourite forgery is exactly "success":
- exit code 0, because the last statement succeeded;
-
Successful compile, because the compile did succeed - nobody wanted the result; -
version.txtnaming the new version, because it is written before the build that then failed; -
"fetch complete", because the missing
frpc.exewas announced on stderr.
Muscle's value is refusing that evidence: counting bytes, comparing hashes, checking timestamps, pointing a real browser at the address, and reporting back: 502.
3. Highbrow: what mind does
Other work that night changed the nature of the problem while appearing to fix nothing.
The first move was to name it: to call four scattered
bugs one class - "defaults that only hold on the build
machine". Once named, the shape changed from "one more fix" to "where
else can this hide?", and a sweep followed: LOG_DIR,
MEM_FILE, RECORDINGS_DIR,
CERT_DIR, SSL_CERTFILE - each asked "can a
customer's machine write here?"
The second move was to build a gate:
def test_runtime_paths_are_outside_the_program_directory(self):
"""Nothing this program writes to may default inside its own program directory."""
It immediately caught MEM_FILE, a case nobody had noticed.
More importantly, this class can no longer ship: it is not "be careful
next time", it is "red next time". One test is worth a stack of
post-mortems.
The third move was to change the criterion. Whether a restart is needed to serve a new certificate had been decided by a snapshot of the file's identity taken at import - which missed the ordering that matters (the app restarts, then the enrolment completes). It now follows the fact: was the certificate written after this process started? Not cleverer; just closer to what the question meant.
The fourth move was quieter: writing down that the upgrade channel compares version strings. Rebuilding under the same number means anyone who already installed it will never see the new package - their update button replays the same version forever. That is not code; it is an understanding of the mechanism, and it decided that the night's fix had to ship as 1.24.6 rather than "docs only".
4. Two models: one writing design, one crawling through consoles
| The highbrow side | The lowbrow side | |
|---|---|---|
| Output | Decision records, tidy commits, systematic test gates, a launcher that asks the server which address to use | Installing packages, killing processes, reading GBK, opening firewalls, counting hashes, turning four console warnings into four bugs |
| How it failed | Defaults and mechanisms - invisible without real hardware |
Slips of the hand: Chinese in a .ps1,
pkill -f killing its own session,
& blocking forever on a process that never exits
|
| Contribution | Stops it recurring | Makes it visible the first time |
Neither is sufficient. Muscle alone: six releases in 48 hours, each
"fixed", the next bug already queued - a perpetual release machine. Mind
alone: elegant documents, complete tests, and a default still pointing
at C:\Program Files while the customer stares at a black
screen.
There is one direction that closes the loop: muscle produces facts; mind turns facts into structure. Every grubby, lowbrow discovery has to become a test, an invariant, a sentence in the runbook that day - otherwise it was just luck.
One purely social lesson, too: two writers on one repository means one
latest.json overwriting the other's ("1.24.7" label, 1.24.6
numbers - an updater comparing mismatched hashes). Intelligence does not
solve that. Discipline does: a release train has one driver.
5. So which is highbrow?
The old Chinese phrase - yangchun baixue, "spring snow", for the refined; xiali baren, "the rustic songs of the village", for the coarse - is usually read as a ranking. In engineering it is not. They are two strokes of the same hand.
Spring snow is the comment, the decision, the guard test that spares the
next person. The village song is the VM reinstalled at two in the
morning, the quote eaten by GBK, the
ping that returns nothing - the opposite evidence that
makes "it should be fine" falsifiable.
The two sentences I am keeping out of those hours:
1. Do not prove success with exit codes; prove it with facts - a timestamp with a year, a hash different from the last one, a real request coming back 401.
2. Every failure on real hardware must become a test or an invariant that same day, or it will return under another name.
As for the models: over these hours one was the dirty hand and one wrote the gate - and occasionally the same hand broke something (committed once with the suite red, killed its own session three times, treated a half-written hash as final). The old line holds: tools amplify method, not character. With the right method, muscle and mind complete each other; with the wrong one they conspire to build a perpetual release machine.
Written 2026-10-02, the day the MRRC Cloud Hub went up. Every number
here comes from that day's logs, commits or artifact checks: 1.24.2 to
1.24.7, six releases, one Windows VM, one Mac, one Hong Kong VPS.
中文版:阳春白雪与下里巴人
Also: Two Models, Forty-Eight Hours, Seven Releases · The Cloud's First Mile · All posts