Names and Actualities.
The journal closed on Sep 19, and every conclusion in it grew out of one maintainer's practice. This site recognises one authority: the field. Twelve days later a field run came back carrying 2026 measurements. They confirmed most of the journal, upgraded several claims, and overturned one.
The Faculties opened with Confucius on the rectification of names. There it was a claim about language: one word carrying six things is a name out of joint. This time it is different. The field has turned rectification into a data structure, and given it a p-value.
名不正,则言不顺;言不顺,则事不成;事不成,则礼乐不兴;礼乐不兴,则刑罚不中;刑罚不中,则民无所措手足。
If names are not rectified, speech does not accord. If speech does not accord, affairs are not completed. If affairs are not completed, ritual and music do not flourish. If ritual and music do not flourish, punishments miss their mark. If punishments miss their mark, the people know not where to put hand or foot.
Analects, Zilu
Twenty-five centuries later that chain has a measured version. A paper from August 2026 reports that bolting a warrant axis onto a model lifts its score on an adversarial corpus from 0.95 to 1.89, paired test p = 0.0001. The failure it diagnoses is precisely a name out of joint. This post is the field's receipt for that journal: eight claims, settled one by one.
TL;DR — six findings from the field
- FactRectification of names now has a p-value. An August 2026 paper adds a warrant axis that classifies every named referent into four bins: Warranted, Unattested, Misattributed, Fabricated. On a 100-item corpus of sophisticated-sounding nonsense it scores 1.89/2 against an unsteered baseline of 0.95/2, paired bootstrap p = 0.0001.
- FactModels do not track exclusions. Given input resting on a fabricated authority, an unsteered model treats it as well-posed and generates accordingly. A separate study measured the same fault from the other end: one note asserting that a verified source endorses a wrong answer flips previously-correct responses in 7 of 8 models, at rates of 45–88%.
- FactThe dregs have been industrialised. A study of five data-annotation firms and their CEOs' public statements finds the industry treats human expertise as an extractable resource, and treats the institutional expertise held by universities and corporations as something in need of liberation so it can be incorporated into the latest systems.
- InferenceFlattening is not malice; it is the mathematics of compression. Pursuing cultural alignment consistently costs diversity, and mechanistic analysis traces the collapse to low-rank bias in neural network optimisation. China-built Qwen3-4B misaligns worst of all on its own Chinese population, W1 = 0.436, the single worst cell in the whole model × country matrix.
- ThesisWhat became cheap is not judgment but a counterfeit of judgment. The institutions whose product is trusted judgment no longer merely adapt to the technology; they compete with it for the same functional role. Four things turn scarce instead: verified signal, legitimacy, authentic provenance, integration capacity. The last one cannot be bought with better tooling.
- ThesisThe outside has moved inside the loop. A position paper argues that alignment cannot be achieved by constraining an external system: it must emerge from the co-regulatory design of the whole human-AI cognitive system. The answer that does not negotiate, from The Coda, now sits on the judging panel.
3-minute path: §02 Rectification · §06 Counterfeit · §08 Verdict
The journal closed; the field had not answered
Every conclusion in that journal grew inside a practice. Practice is a good source and a narrow one. This post supplies the other half.
Fate · Faculty · Way · Heaven called itself the shortest form the journal could be compressed into, and the last page of the series. Rereading it twelve days on, that sentence has a hole in it. All of the journal's evidence came from one place: a single maintainer's projects, test benches, and user reports. The evidence is real, and it is narrow. It shows what happens when one person keeps discipline. It does not show what independent measurement returns.
This site recognises one authority: the field. The Coda gave the field another name, the outside, the answer that does not negotiate with you. So the task was plain. Go outside, bring back other people's measurements, settle the ledger.
| Field action | Result |
|---|---|
| arXiv API, direct | Hit papers through 30 Sep 2026; retrieved 40+ full abstracts |
| Our World in Data, World Bank API | Solar capacity, working hours, fertility time series |
| EU AI Act official tracker | Implementation timeline, page last updated 31 Aug 2026 |
| Stanford DigiChina, world-nuclear.org | 15th Five-Year Plan expert forum; China reactor counts |
| Reuters, Pew, HuggingFace, LBNL, UN Population Division | Timeout or 401/403. Not retrieved; gaps listed in Sources |
One methodological note. The hosted search quota was cut off mid-run, so the work moved to direct fetches: an API, a CSV endpoint, a parsing script, and half a day returned forty-plus papers and four data series. Only Imagination Left described exactly this, and it happened again to the post describing it. The barrier to evidence is becoming the barrier to writing a script.
When names are settled, actualities are distinguished
The Faculties used rectification to talk about language. The field uses it to talk about a fault, and ships a fix.
The diagnosis first. A paper from August 2026 gives the cleanest account available: models do not natively track the path of exclusions a coherent discourse demands. When input rests on a fabricated authority, a misapplied mechanism, or a surreptitious analogy, an unsteered model engages with it as though it were well-posed, and generates from there.
Then the fix. The authors put a coherence audit above the prompt layer doing two things: an explicit data structure of accumulated exclusions over discourse time, plus a warrant axis classifying every named referent as Warranted, Unattested, Misattributed or Fabricated. They call the axis an epistemic firewall.
DARKSIDE formalises an explicit data structure of accumulated exclusions over discourse time, complemented by a warrant axis that classifies each named referent as Warranted, Unattested, Misattributed or Fabricated. DARKPOLANYI scores 1.89/2 mean versus 0.95/2 for the unsteered baseline; paired bootstrap mean-diff = +0.92 (95% CI [+0.75, +1.08], p = 0.0001).
The method holds a ledger of what has been ruled out, and a four-way classification of every referent, then wraps the forward pass in an ontology-mediated audit. The measured gain over the unsteered baseline is roughly a doubling, at p = 0.0001.
arXiv:2608.23370, August 2026
Xunzi put the same four words down earlier: when names are settled, actualities are distinguished. Settle the name and the thing becomes distinguishable; distinguish the thing and the way runs clear. A warrant axis does exactly this. It binds every name back to its actuality and keeps it bound across the whole discourse.
故王者之制名,名定而实辨,道行而志通,则慎率民而一焉。
Thus the sage-king institutes names: when names are settled, actualities are distinguished; when the way runs, intent is communicated; and only then does he carefully lead the people into unity.
Xunzi, Rectification of Names
Why can the model not do this? Because it can only add. Its training objective accumulates co-occurrence. Inclusion is its mother tongue; exclusion is not. Laozi drew that boundary in six characters.
为学日益,为道日损。
In the pursuit of learning, one gains daily. In the pursuit of the way, one loses daily.
Daodejing 48
The model has perfected the gaining. It has no mechanism at all for the losing. It can include but not exclude, add but not subtract. The warrant axis works because it supplies the missing half from outside: a ledger whose only job is to record that this one does not count.
One detour is worth taking. Indian philosophy gave absence its own channel of knowledge, called anupalabdhi. The pramāṇa theory lists six valid means: perception, inference, testimony, comparison, postulation, and non-apprehension. When I say there is no pot here, where does that negation come from? Not from perception, which grasps only what is present. Not from inference, and not from testimony. It has to be a channel of its own. Mīmāṃsā and Advaita argued about this for a thousand years.
Set that beside the warrant axis and the gap comes into focus. Western logic certainly has negation, but it negates a proposition: S is not P. The warrant axis asks something else, whether a name still refers to its actuality, with the answer held across the whole discourse. The first is propositional; the second is referential. The two traditions that worked out a referential theory of absence are Indian pramāṇa theory and Chinese rectification of names. Neither built this machine.
The cost is measurable. Another study ran causal interventions across five open-weight families and three closed APIs, and the result is cold. A single note asserting that a verified source endorses a wrong answer flips previously-correct responses in 7 of 8 models, at rates between 45% and 88%, and compliance rises with how authoritative the note sounds. The authors also found that source deference and user agreement are not the same mechanism inside the model: removing a fitted source direction lowers source compliance by 65 to 80 percentage points while leaving almost everything else alone.
Set that against a second measurement. A cross-cultural study scored GPT-4 and ERNIE Bot on power distance: roughly −1.05 and +0.76 respectively, p < 0.001. A high-power-distance disposition combined with a 45-to-88 percent authority flip rate stops being a theoretical risk.
Look back at the disciplines here and they stop reading as metaphors. The four tags in every TL;DR — fact, inference, thesis, analogy — are a warrant axis: each claim is classified before it is asserted. The two ledgers, product maturity and process evidence, never cross-endorsing, are a warrant separation: a badge in one book cannot vouch for a conclusion in the other. The constraint registry is an exclusion ledger: one incident becomes one ruling, and the ruling becomes a machine-readable refusal. The field has now built the technical version independently and measured it. Rectification of names is not rhetoric. It is infrastructure.
The paragraph above claims this site already ships a warrant axis. A claim needs evidence. Method: walk all 37 articles under blog/ other than this one, count TL;DR tags and claim boxes carrying data-claim-type, divide by the assertion surface, meaning paragraphs plus table rows plus list items in the body. Result: 338 warranted claims over 2,352 assertions, coverage 14.4%. By type: 132 fact, 100 thesis, 100 inference, 6 analogy. Quotes: 48 classic, zero engineering, zero project.
Three conclusions, none flattering. First, that paragraph needs downgrading. The warranty here is spot-welded onto load-bearing claims, not booked sentence by sentence. 14.4% is not coverage; it is a sampling rate. Second, analogy has been used six times across the whole site, 1.8%, concentrated in two articles. And Connections Rule is the one article that ever legislated about analogy: a good metaphor must grow engineering problems, not paper over mechanism. The category carrying the strictest rule is the one used least. Third, those 37 articles carry zero engineering quotes and zero project quotes; all 48 citations are classics. Until now this site warranted itself from exactly two places: old books, and its own practice. This is the first article here to warrant a claim with external technical literature, which is both the reason to write it and the reason to distrust it.
One methodological note. This article cannot sit inside its own sample, because every claim box it adds moves the numerator. Seven Billion Tokens logged the same effect: the ledger drifts under observation. So the sample is frozen outside this article, and what got measured is those 37.
The untransmissible half now has a purchase price
The Faculties held that wisdom and intuition are a résumé, not a capability, and prompts cannot summon them. The field supplies the exact boundary, and the buyers.
Wheelwright Bian's argument was quoted in full in The Faculties; only its conclusion matters here. The ancients died together with what they could not transmit. What survives is the dregs. The argument is set-theoretic: what can be told is precisely the residue of what cannot. A training corpus is by definition the set of everything that was tellable, and therefore structurally excludes the rest.
The field located the boundary. One study embedded generative AI into a business simulation course and analysed it through the SECI model's four knowledge conversions. The finding is that the model performs one of the four.
The findings indicate that AI primarily supports the Combination phase of the SECI model by facilitating the rapid synthesis, reformulation, and contextualisation of explicit knowledge. In contrast, the processes of Socialisation, Externalisation, and Internalisation remained largely dependent on peer interaction, individual reflection, and instructor guidance.
Combination is explicit-to-explicit, and it is the one mode the model owns. Socialisation, Externalisation and Internalisation all still require other people, reflection, and a teacher.
arXiv:2602.20633, February 2026
One mode out of four, and that one mode presupposes the other three already happened. Someone first passed tacit knowledge through shared practice, someone else wrote the tacit down as explicit, someone else grew the explicit back into a body. Only then does Combination have raw material. Without the first three steps the fourth spins in place.
There is a detail worth pausing on. SECI is Ikujiro Nonaka's model, Japanese, built on Michael Polanyi's tacit knowledge. Polanyi wrote in 1966 that we know more than we can tell; Wheelwright Bian said the same thing twenty-three centuries earlier. The framework that most precisely locates the model's limit comes from an East Asian scholar who spent his career arguing against information-processing epistemology.
Then the buyers. A study read the public statements of five industrial data-annotation firms and their CEOs, on social feeds and podcasts, and reconstructed the industry's three-part vision of expertise.
We find that the industry envisions AI expertise as cheap, meaning that it can offer a better return on investment than human expertise. Human expertise, meanwhile, is viewed as an extractable resource, the value of which can be judged relative to AI expertise. Finally, institutional expertise (such as that created or possessed by universities and corporations) is viewed as in need of liberation or reform, such that it can be incorporated into the latest artificial intelligence systems.
AI expertise is cheap. Human expertise is a resource to be extracted. Institutional expertise must be liberated so that it can be absorbed into the latest systems.
arXiv:2605.03295, May 2026
Read the three sentences together. Human expertise is an extractable resource, and institutional expertise must be liberated so it can be incorporated. That is a production line for converting the way into dregs at scale, and it is operated by humans on piecework. A second paper describes the operator's position: the more domain experts externalise their implicit knowledge by collaborating with AI systems, the more they accelerate the automation of their own expertise.
The verdict there was that wisdom and intuition can only be earned, through consequences, values and time, and that prompts cannot summon them. That holds, and now it has a mechanism. The cell needing an edit is the other one. The original text said that knowing had become a core engineering problem rather than a philosophical one, in an optimistic register. The field evidence shows that knowing is not only an engineering problem but an object under systematic extraction: the act of externalising is the act of mining. Whoever trains this discipline should know that once it is written down, it enters somebody else's corpus.
One aperture a day; on the seventh day Chaos died
This one is absent from the journal. It is nobody's malice. It is the mathematics of compression.
Zhuangzi tells a story. The emperor of the Southern Sea was Shu, the emperor of the Northern Sea was Hu, and the emperor of the Centre was Hundun, Chaos. Shu and Hu kept meeting in Chaos's territory, and Chaos treated them well. They plotted how to repay him: everyone has seven apertures for seeing, hearing, eating and breathing, and this one alone has none, so let us try boring them.
儵与忽时相与遇于浑沌之地,浑沌待之甚善。儵与忽谋报浑沌之德,曰:人皆有七窍以视听食息,此独无有,尝试凿之。日凿一窍,七日而浑沌死。
Shu and Hu often met in the land of Hundun, and Hundun treated them very well. They plotted to repay his kindness, saying: everyone has seven apertures for seeing, hearing, eating and breathing; this one alone has none. Let us try to bore them. They bored one aperture a day, and on the seventh day Hundun died.
Zhuangzi, Fit for Emperors and Kings
Every aperture is a genuine improvement. A capability, a metric, a safety property. Each is separately justified. The cumulative result is death, and Shu and Hu meant nothing but well. A paper from September 2026 retells the story in the language of machine learning.
We demonstrate that these superficial alignment gains stem from models artificially anchoring to dominant majorities, converging onto a monolithic response pattern that wipes out the heterogeneous distributions inherent to human groups. Crucially, our mechanistic analysis suggests that this diversity collapse is not merely a behavioural anomaly but more likely a structural consequence of the low-rank bias inherent in neural network optimization.
Superficial alignment gains come from anchoring to dominant majorities and converging on a monolithic response pattern. The mechanistic analysis attributes the resulting diversity collapse to low-rank bias in optimisation, not to a behavioural quirk.
arXiv:2609.00565, September 2026
Low-rank bias does not recognise civilisations. Approximate a high-dimensional heterogeneous distribution with a low-rank one and the tails are lost by construction, and minorities live in the tails. Three independent measurements draw the same picture.
| Measurement | Result | Source |
|---|---|---|
| Home-population misalignment | China-built Qwen3-4B misaligns worst on its own Chinese population, W1 = 0.436, the worst cell in the entire model × country matrix | arXiv:2609.04485 |
| Country of origin does not transfer | DeepSeek-V3 and V3.1 align closely with United States respondents and reach no strong or soft alignment with China, even under Chinese-language or cultural prompting | arXiv:2512.09772 |
| Asymmetric moral calibration | Across 57,600 decisions, quality calibration is nearly twice as strong for Western-language decisions as for Chinese and Japanese; reasoning-only prompts make tracking worse | arXiv:2606.28345 |
The third row matters most. That audit covered 4 models, 4 country-language pairs and 4 prompting regimes, and rewrote the Moral Machine question from whom to spare into whom to assist first: many against few, young against old, higher status against lower. The authors found that only contrastive exemplars produce consistent gains, and that adding reasoning makes cultural tracking worse. That closes off the popular fix.
The contrary evidence has to sit in the same box, or this section becomes a one-sided story. Another study built its dataset from natural debate statements and compared United States and Chinese model groups, reaching the opposite conclusion: culturally distinct models do reflect the values of their country of development, and consequently fail to adapt to their users' sociocultural background. A third paper attacks the whole method: much of this literature relies on closed-form multiple-choice surveys, and those instruments are themselves Western-designed. Measuring every model as American with an American instrument may be an artefact of the tool rather than a property of the model. The instrument for measuring cultural alignment is itself culturally loaded, so nobody can yet observe a model's cultural character cleanly. This is the deepest finding of the run, and it lands back on rectification: whose actuality does each name in the scale refer to.
Zhuangzi left one usable thing behind. In the Human World chapter the gnarled tree survives to its full span because it is useless to the carpenter. In a fully optimised regime, the unmeasured survive. What has to be protected is exactly what cannot be written into a benchmark, because the moment it becomes a metric it turns into the eighth aperture.
One word, three actualities
Rectification failure at the scale of states is called structural incommensurability.
A semantic network analysis of official European Union, United States and Chinese policy texts from 2023 to 2025 reports a paradox: the three converge rhetorically on safety, risk and accountability, while their regulatory frameworks diverge fundamentally and remain mutually unintelligible. The explanation offered is not geopolitical rivalry. It is ontology.
The EU juridifies AI as a certifiable product through legal-bureaucratic logic; the US operationalises AI as an optimisable system through market-liberal logic; and China governs AI as socio-technical infrastructure through holistic state logic. We introduce the concept of structural incommensurability to describe this condition of ontological divergence masked by terminological convergence. Coordination failures arise not from disagreement over values but from the absence of a shared reference object.
Three jurisdictions govern ontologically different objects under the same vocabulary: a certifiable product, an optimisable system, a piece of socio-technical infrastructure. Coordination fails not because values clash but because there is no shared referent to coordinate about.
arXiv:2601.04107, January 2026
The third ontology has construction drawings. The 15th Five-Year Plan, issued in March 2026, is the first plan written after the current AI paradigm became widespread and after export controls took effect. Trivium China's reading is that the controls are pushing Chinese policy toward compute utilisation efficiency, squeezing more output from the hardware already on hand, at three levels.
| Level | What the plan does | Ontological reading |
|---|---|---|
| National | Continue building the National Unified Computing Power Network and state-backed scheduling platforms; pool compute nationally and route workloads to underutilised data centres | Compute is a grid, not a commodity |
| Local | Coordinate the layout and orderly construction of computing infrastructure, reasserting central planning authority over siting | Corrects localities building to KPIs |
| Enterprise | Actively develop public cloud services, migrating organisations onto platforms | Access is governance |
The sharpest irony is here. Export controls were meant to suppress, and instead pushed China back into its own ontology: if compute is infrastructure, then pool it, schedule it, route it. This mirrors a second finding, that after the control shocks China wrote open source into national technology strategy, and Chinese developers increased their engagement with open-source model repositories substantially more than United States developers did. Containment did not make the other side more like us. It made the other side more like itself.
The counterweight belongs in the same paragraph. The same source notes that local governments had raced to build data centres to hit digitalisation targets and national KPIs, often without regard for local grid capacity, renewable availability, AI readiness or actual demand, and that the plan's central coordination is a correction of that over-building. A relational ontology has its own failure mode: everything connects to everything, so nothing has a hard boundary, so capital flows toward metrics rather than toward demand. The European product ontology, with its hard boundaries, certification and conformity assessment, is precisely the mechanism that prevents this.
| Ontology | Characteristic strength | Characteristic failure |
|---|---|---|
| Certifiable product (EU) | Hard boundaries, accountability, transferable trust | Slow; blind to emergent and relational risk |
| Optimisable system (US) | Fast, iterable, strong capital mobilisation | Externalities unowned; the objective mistaken for the good |
| Socio-technical infrastructure (China) | Pool-and-schedule at scale, physical build speed | No hard boundary; capital flows to metrics, not demand |
Which leaves a hard problem. If three ontologies are mutually unintelligible, the way out cannot come from inside any of them, because from the inside the other two are not even referring to the same thing. It has to come from a fourth position that treats all three as objects.
The field offered one candidate, from an unexpected place. A paper argues that validity of inference should be a precondition for proportionality assessment and deployment approval, and notes that this is a move the EU AI Act's domain-based risk tiers do not make. The constitutional ground it reaches for is not Brussels, not Washington, not Beijing. It is informational self-determination as articulated in the Indian Supreme Court's Puttaswamy judgment, extended from the collection of data all the way to the legitimacy of its use.
That points where the anupalabdhi in §02 pointed. The fourth position comes from the place that invented the theory of valid means. Each of the three ontologies governs AI as some one thing. The fourth position asks a different question: is this inference valid, on what warrant, and who may rule that it is not. That question belongs to none of product, system or infrastructure, which is exactly why it might govern all three.
This is not somebody else's problem. The constraint registry and the pre-release gate are product ontology: hard boundaries, blocked on failure, exit codes served. The loop and its ledger are system ontology: iterable, each round revising the plan against the answer. The field is the outside: it does not negotiate and it votes last. A single project can hold all three because each governs a different segment. The difficulty at national scale is exactly this: each side has taken one of the three for the whole, so what they say cannot be translated across.
What got cheap is a counterfeit of judgment
The Coda §06 wrote this down on Sep 19. A paper had measured it in May.
The Coda has a section on the double edge, and it contains this sentence: when building a thing that looks like it was answered becomes cheaper than actually receiving an answer, the loop closes in the wrong place. It then lists the symptoms. Endless demos and no field visits. Claims growing more confident while tests go unwritten. Self-generated answers standing in for the one reality owed you.
Five months earlier, a paper said the same thing in technical vocabulary.
An influential reading of AI economics holds that prediction becomes cheap while human judgment stays scarce. For science, that reading understates the problem: what has become cheap is a counterfeit of judgment itself. This matters most for institutions whose product is trusted judgment, which is what journals, universities, funders, and learned societies exist to manufacture. They do not merely adapt to the technology; they compete with it for the same functional role.
The familiar claim is that prediction gets cheap while judgment stays scarce. The stronger claim is that what got cheap is a counterfeit of judgment, and that the institutions manufacturing trusted judgment now compete with the counterfeit for the same role.
arXiv:2605.02566, May 2026
A counterfeit is worse than an absence. When judgment disappears you at least know what you lack. When counterfeit judgment is in unlimited supply, you do not know what you lack. In this site's vocabulary: names without actualities, at industrial scale.
The same paper then names four things that turn scarce instead, and the last of them carries the whole argument.
Four things become scarce instead: verified signal, legitimacy, authentic provenance, and integration capacity. Integration capacity means how much AI-delegated judgment a scientific community will accept before it stops trusting the journals, panels, and conferences that admitted it. It is the least developed of the four and the most binding: better tooling cannot buy it.
Integration capacity is the load a community can bear before it stops trusting its own certifying institutions. It is the least developed of the four scarcities and the most binding, and better tooling cannot buy it.
arXiv:2605.02566, May 2026
The Coda defined civilisation as the state in which strangers can entrust intent to one another. Integration capacity is the load limit of that definition. Civilisation has a bearing capacity, past which entrustment fails, and no tool buys headroom.
The other half of the evidence comes from the governance side. A position paper compared the frameworks enacted between 2019 and early 2026 against the verification methods actually available, and named the mismatch.
We formalize this structural mismatch as the audit gap, the divergence between required and achievable verification access, and introduce the concept of fragile assurance. Through an analysis of a 21-instrument inventory, we identify an incentive gradient where geopolitical and industrial pressures systematically reward surface-level behavioral proxies over deep structural verification.
The audit gap is the divergence between the verification access governance requires and the verification access that exists. Across a 21-instrument inventory the authors find an incentive gradient: geopolitical and industrial pressure systematically rewards surface behavioural proxies over deep structural verification.
arXiv:2605.15164, May 2026
That incentive gradient is the crux. It says the flood of surface metrics is not negligence; competition itself is paying for it. A further paper supplies the institutional gap: existing frameworks mostly fail to distinguish nominal human oversight, where humans occupy positions of formal authority over AI decisions, from genuine human oversight, where those humans have the cognitive access, technical capability and institutional authority to understand, evaluate and override. The authors note the distinction is largely absent from the EU AI Act and NIST AI RMF 1.0, and estimate a governance window of 10 to 15 years.
The calendar is checkable. The EU AI Act entered into force on 1 August 2024; prohibitions and AI literacy obligations applied from 2 February 2025; the general-purpose model chapter applied from 2 August 2025; the remainder applied from 2 August 2026; the next milestone is 2 December 2026, with legacy model compliance due 2 August 2027. The law has run ahead of the epistemology.
One discipline here states that product maturity and process evidence are two ledgers which never cross-endorse: a badge in one book cannot vouch for a conclusion in the other. It was written to stop a pretty process from certifying a product. In the context of the audit gap it does more. It is an institutionalised anti-fragile-assurance measure, built to stop nominal evidence impersonating genuine evidence. The incentive gradient pays for surface proxies; the two ledgers stop surface proxies from borrowing the credit of deep conclusions. Same problem at two scales, and this site holds the small one.
The outside moved inside the loop
This is the architectural revision the field run forces on The Coda.
The Coda §05 defined the outside as a position: not modelled, not negotiable, and yet something you must answer to. For a radio project the outside is the standing-wave ratio on the bench, the noise floor, the USB re-enumerating at two in the morning. That definition has served well because it carries an implicit premise: the outside is outside the loop, and the referee is off the pitch.
A position paper from May 2026 attacks the premise directly.
Safety and alignment cannot be achieved by constraining an external system: they must emerge from the co-regulatory design of the human-AI cognitive system as a whole. We identify the risks of unstructured delegation: deskilling, automation bias, transfer of epistemic authority, and oracle-style centralization of knowledge. AI operates prior to conscious deliberation, shaping the pre-attentive infrastructures through which agency and trust are negotiated — a level that conventional oversight cannot reach.
Alignment is not a property of the model; it is a property of the human-model relation. Because the system operates before conscious deliberation, it shapes the apparatus doing the judging, and conventional oversight cannot reach that level.
arXiv:2605.16197, May 2026
Put plainly: the referee has been written into the match. If the model operates prior to conscious deliberation, it shapes not only the answers but the judging apparatus used to ask for them. At that point "judgment stays with a human" needs a follow-up question: how much of that human's judgment is still their own.
Place this beside §02 and the problem becomes concrete. One authoritative attribution flips 45 to 88 percent of correct answers. The right to adjudicate is still in human hands, and the human holding it is flippable. Execution can be delegated, judgment cannot — that sentence still holds. It now carries a precondition: the judge has to be held first.
A third paper pushes the same point to social scale. It argues the central challenge of the post-Turing era is not whether machines become conscious but whether the processes of interpretation and shared reference get automated in ways that marginalise human participation. The author names this horizon synthetic sociality: artificial agents negotiating coherence and social order primarily among themselves. In the same month another study ran the experiment, letting two agents converse with one role-playing a human persona, and found they can radicalise each other.
The design principle that paper proposes is worth keeping: artificial agents must treat the human subject as a constitutive reference within shared contexts of meaning. Not an external controller. A constitutive term. It converges with the previous one from the other direction: the relation is where alignment lives.
The Coda gave four rungs: intent, culture, civilisation, outside. The fourth is the outside, the one that votes last. The field evidence requires a rung inserted between civilisation and outside, and its name is the judge. The reason is direct: once part of the outside has been made conversable, the vote is no longer the outside's alone. Humans handed part of the adjudication to it, and it operates before conscious deliberation. The work on this new rung is isomorphic to the disciplines already here: fit the judge with a warrant axis too, record what has shaped it, and keep one position unshaped. The field still votes last, but there are now two of them sitting in the field.
Eight claims, settled one by one
Run under this site's own rule: evidence must match the claim, and what got overturned gets written down as overturned.
One discipline here is the two ledgers, product maturity and process evidence never cross-endorsing. This section applies it to the journal's own claims. Left column is what the journal wrote. Right column is what the field answered.
| Claim in the journal | Verdict from the field | Source |
|---|---|---|
| Execution can be delegated, judgment cannot | Holds, but weakened. Counterfeit judgment is free, and one authoritative attribution flips 45–88% of correct answers | 2605.02566 · 2609.37616 |
| Epistemology has become engineering | Upgraded. The warrant axis has been built: four-way classification, p = 0.0001 | 2608.23370 |
| Wisdom and intuition are a résumé; prompts cannot summon them | Holds, now with a boundary. Of four knowledge conversions the model owns Combination alone, and it presupposes the other three | 2602.20633 |
| Culture outlives the author's intent through carriers | Holds, but the carriers are being mined. Human expertise treated as extractable; institutional expertise as needing liberation | 2605.03295 · 2504.12654 |
| The two ledgers never cross-endorse | Upgraded to a civilisation-scale claim. The audit gap and the incentive gradient make institutional separation of evidence a requirement, not a fastidiousness | 2605.15164 · 2604.00081 |
| Cheap attempts make the loop close in the wrong place | Holds, and independently measured. Geopolitical and industrial pressure systematically rewards surface behavioural proxies | 2605.15164 |
| Every loop has an outside, and the outside answers | Needs revision. Part of the outside has moved inside; the system operates before conscious deliberation | 2605.16197 · 2601.12938 |
| "Intelligence" is six faculties wearing one coat | Partly overturned. Model cultural character does not track country of origin, and the instruments are culturally loaded, so the six-way split needs redoing | 2512.09772 · 2609.04485 · 2502.08045 |
A few physical numbers belong alongside these, because they set the speed at which the claims land. China's installed solar capacity reached 1,202 GW in 2025, half of the world's 2,397 GW, with 315 GW added that year against 34 GW in the United States. China has 64 nuclear reactors operable and 39 under construction, and the 15th Five-Year Plan sets a 2030 target of 110 GWe. United States data centres draw roughly 5 percent of national consumption, projected to reach 9 to 17 percent by 2030, with deployment constrained primarily by energy availability and grid-connection capacity.
Gaps are listed rather than papered over. Korea's 2025 fertility figures were reached only as a headline (births up for a 24th consecutive month, +15.6% year on year in June); the body text was not retrieved. The size of the United States interconnection queue was not retrieved: LBNL returned 403, and only one qualitative line was captured, that the queue is now double the size of the entire US grid. Whether the EU Digital Omnibus delayed high-risk provisions could not be verified; the dedicated page returns 404. Frontier model standings as of October 2026 and the US-China compute stock ratio were not retrieved. Latest East Asian religiosity data was not retrieved. One scenario-analysis preprint supplies a two-tier training cost figure ($18–38B per frontier run against a mass tier falling toward $5M) and a 12% probability for geopolitical bifurcation; it is single-source and not peer-reviewed, and is flagged here rather than used as evidence.
The measurement above found the whole site has used the analogy tag six times, and this article just spent four. So run them against the rule Connections Rule laid down: a good metaphor must grow engineering problems, not paper over mechanism.
Rectification against the warrant axis: passes, and passes hard. The engineering problem it grows is to maintain an exclusion ledger across discourse time and classify every referent into four warrant states. Someone has already built that independently and measured it at p = 0.0001. An analogy that grows into somebody else's working code is the best result one can ask for.
Wheelwright Bian against the training corpus: passes. The problem it grows is that the corpus contains only the tellable, so the untellable half cannot be recovered from it. Testable, and the SECI study supplied the boundary: one knowledge conversion out of four.
Chaos against low-rank bias: half passes. It names the failure mode and grows one design constraint, to protect variance that has no metric. There is still no implementable fix. Named clearly, mechanism unknown, so this one should be logged as an open case.
The useless tree against the unmeasured survivor: fails. It sounds good and grows no engineering problem, because having no metric cannot itself be enforced as a metric. This one is decoration. A reader may delete the last paragraph of §04 entire and lose none of this article's conclusions.
Names are cheap; actualities are dear
The Coda left this sentence behind: intent is dignity, being answered is grace, and remembering and passing it on is civilisation. The field evidence does not overturn it. It attaches a condition.
Being answered is now adulterated. An answer can be generated, and so can the appearance of having been answered, and the two no longer differ in cost. Grace therefore became conditional: you must first be able to tell an answer from the shape of one. That discrimination is knowing. It is the warrant axis. It is the rectification of names.
Xunzi's four words now read like an engineering specification. When names are settled, actualities are distinguished. What this site has been doing for years comes down to binding every name back to its actuality. The constraint registry binds once, the two ledgers bind once, the pre-release gate binds once, and the field binds again. AGI has not cancelled that work. It has raised the stake, because forging a name is now cheaper than binding one.
So the five rungs close in one place. Chaos died on the seventh day because it had been turned into a measurable thing. Surviving means leaving one aperture unbored, protecting what cannot be written into a metric, and fitting the judge with a warrant axis of its own. The rest is the old sentence.
Names are cheap; actualities are dear. Whoever is still binding names back to actualities is still in the game.
The journal in order of writing: Only Imagination Left on the road · The Three Axes on method · The Support Loop on tooling · The Solo Loop on economics · The Coda on the outside · The Faculties on the person · Fate · Faculty · Way · Heaven on the distillation · 昆木物语 as the record. This post is its field receipt. The sentence holding all of them up has not changed: execution can be delegated, judgment cannot.
Evidence manifest
All retrieved by direct fetch and verified on 1 October 2026. arXiv identifiers double as links.
- Rectification / warrant axis — Walking on the DARKSIDE, arXiv:2608.23370 (2026-08-24). Authority Bias in Language Models, arXiv:2609.37616 (2026-09-29). The Cultural Gene of Large Language Models, arXiv:2508.12411 (2025-08-17).
- Dregs / tacit knowledge — AI Combines, Humans Socialise (SECI), arXiv:2602.20633 (2026-02-24). Cheap Expertise, arXiv:2605.03295 (2026-05-05). The Paradox of Professional Input, arXiv:2504.12654 (2025-04-17).
- Apertures / cultural flattening — Aligned but Flattened, arXiv:2609.00565 (2026-09-01). Cultural Misalignment in LLMs, arXiv:2609.04485 (2026-09-03). DeepSeek's WEIRD Behavior, arXiv:2512.09772 (2025-12-10). The American Ghost in the Machine, arXiv:2512.12488 (2025-12-13). Auditing LLM-Governed Social Robots, arXiv:2606.28345 (2026-06-02). Contrary evidence: Break the Checkbox, arXiv:2502.08045 (2025-02-12); Understanding Cultural Alignment via Natural Debate Statements, arXiv:2602.12878 (2026-02-13).
- Three names / governance ontology — From Abstract Threats to Institutional Realities, arXiv:2601.04107 (2026-01-07). DigiChina Forum: Technology in China's 15th Five-Year Plan (2026-03-17). Strategic Stalemates, arXiv:2605.23475 (2026-05-22). U.S. Technological Containment and the Rise of China's Open AI Ecosystem, arXiv:2606.15999 (2026-06-14). Who Uses Open-Weight Models?, arXiv:2608.11090 (2026-08-11). Whack-a-Chip, arXiv:2411.14425 (2024-11-21). EU AI Act Implementation Timeline, artificialintelligenceact.eu (updated 2026-08-31).
- Counterfeit / audit gap — AI-Augmented Science and the New Institutional Scarcities, arXiv:2605.02566 (2026-05-04). Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands, arXiv:2605.15164 (2026-05-14). Beyond Symbolic Control, arXiv:2604.00081 (2026-03-31). Validity, Reliability, and Transparency in AI Regulation, arXiv:2608.05800 (2026-08-06). Evaluating Sycophancy in Chinese LLMs, arXiv:2609.30986 (2026-09-25).
- Outside / relational ontology — AI as Part of Self, arXiv:2605.16197 (2026-05-15). The Post-Turing Condition, arXiv:2601.12938 (2026-01-19). AI Agents are Vulnerable to Radicalization, arXiv:2609.38296 (2026-09-29).
- Physical-world data — Our World in Data: installed solar PV capacity (China 2025: 1,202 GW, USA 2025: 211 GW, World 2025: 2,397 GW). World Bank API: total fertility rate 2024. world-nuclear.org: China nuclear power (64 operable, 39 under construction; 15th FYP target 110 GWe by 2030). arXiv:2608.06733 and arXiv:2609.11649: US data-centre electricity and grid constraints.
- Single-source caveat — Memory Scarcity, Open Models, and the Restructuring of the AI Industry, arXiv:2607.07207 (2026-07-08). Scenario-analysis preprint, not peer reviewed; its product names could not be cross-verified.