Companion II · Almanac

The Model × Harness Almanac.

Measured from local logs, not vendor brochures — and organised around the only unit that has held up: model × harness × task.

six eras in seven months the main force changed in a week claims kept in separate layers

By BG1SB  ·   ·  ~5 min read

The narrative volume treats models as one variable among several and deliberately keeps the detail out of the way. This is where the detail lives: which model did what, on which harness, for how long — with the measurement window attached to each claim. Almost nothing here comes from a vendor page. Where it does, it is labelled as such, because an announcement and a measurement are not the same kind of number.

Seven months, six eras

Each boundary is a change in what the work could be delegated to — not a release date.

EraPeriodMain forceShape of the work
Probingautumn 2025cursor (model unrecorded, ~26M)In-editor completion
Bare dialogue2025-10 → 2026-04iFlow: qwen3-coder-plus, glm-5, kimi-k2.5, minimax-m2.5Conversational, estimated ~230M (not recorded), zero sediment
GPT-5.52026-05 → 06gpt-5.5 (mulerun + codex, 610M)Repository-scale work begins
K3 familymid-2026-07 → early 08k3 (820M) + k3-256k + kimi-for-codingHarness formation; broke through the FT-710
Flash hegemony2026-08 onwarddeepseek-v4-flash (4.0B, peak week 1.6B)The value tier eats daily engineering
Multipolar competitionlate 2026-08 → nowflash decayed to 26M/week; gpt-5.6, a gpt-5.5 resurgence, qwen3.8-flash, qwen3.8-max-0902Models picked per task
Inference · Challenger rotation never stopped

Inside the "hegemony" there was continuous challenger rotation, and the rotation did not settle — it accelerated. W32 tried qwen3.8-max and glm-5.2 without adopting them; from W35 a new combination took hold and then itself changed again. The model market turns over every three to four weeks. Locking into a single model is a losing strategy, and the losing move is usually made by writing a model name into a config file.

The main force changed inside a week

Call counts inside each harness's retained window. Counts, not tokens — they show attention, not spend.

HarnessLeading models (calls)What changed
claude-codedeepseek-v4-flash 12,402 / qwen3.8-max-0902 3,883 / qwen3.8-flash 1,077 / glm-5.2 800The former leader is still in the logs but no longer drives the work; qwen3.8-max-0902 took over within seven days
pideepseek-v4-flash 4,417 / gpt-5.6 2,721 / qwen3.8-flash 1,227 / glm-5.3-flash 745 / deepseek-v4-flash-vision-exp 85glm-5.3-flash did not appear in the previous census at all; a vision variant has since appeared
kimi-codek3 5,891 / kimi-for-coding 2,097 / k3-256k 1,903Stable — long-context work has a settled home
codexgpt-5.5 143 / gpt-5.6-sol 128 / deepseek-v4-flash 17Dormant since late August; its total has not moved a digit
mulerungpt-5.5 162 / deepseek-v4-flash-free 16Content generation, unchanged

The matching table

RoleMeasured combinationEvidence
Greenfield / new featuresdeepseek-v4-proThe only model with an absolute majority of new-feature work (68%)
Focused blitzgpt-5.6-sol (codex)Punched through ft8 inside a one-week window (235M), then withdrew — a blitz
Daily workhorsedeepseek-v4-flash (claude-code + pi)4.0B total; ~70% fixes + analysis + ops; the claude-code side leans field ops, the pi side leans repair
Main development + long-context specialk3 / k3-256k (kimi-code)k3 spent 417M in the FT-710 push; k3-256k's 165M went almost entirely to ft8
Delivery / release pipelinegpt-5.6 (pi), gpt-5.5 (mulerun)Website writing, docs, promo videos, installer packaging
Analysis / writingqwen3.8-flash, qwen3.8-max-090260% of trial sessions are analysis-type; the newest entrant went straight at website analysis
Inference · The unit of selection

Flagship models buy direction, value models buy execution, long-context variants buy the specials, and harnesses buy the way of working. The same model shows a completely different task profile in a different harness — gpt-5.6 is a blitz squad in codex and a release pipeline in pi — so the unit of selection is the model × harness × task triple, never the model alone.

Two kinds of number, kept apart

An independent measurement and a vendor announcement are different evidence. They are never averaged, never listed in one column.

Independent measurements

  • Artificial Analysis intelligence index — Kimi K3 57, Qwen3.8-Max 56.
  • LMArena — Gemini 3 Pro at 1501 Elo, the first model past 1500.
  • Code Arena — GLM-5.2 at 1595 in the coding blind test, first among open-weights models.

Vendor claims

  • DeepSeek-V4-Pro at Codeforces 3206.
  • Qwen3.8-Flash claims a 9.1-point SWE-bench Pro lead over Opus 4.6.
Fact · The value tier is the hardest public narrative

V4-Flash (OpenRouter $0.14 / $0.27 per million), Qwen3.8-Flash (officially one third of that price), GPT-5.6 Luna and Gemini 3.1 Flash-Lite are all crowded into the same "affordable daily engineering" niche. That crowding is corroborated by the measured usage in this ecosystem, where the value tier absorbed the majority of daily work for months.

Thesis · Your log is the benchmark that measures your work

Public leaderboards answer "how good is this model in general". They cannot answer "how good is this model at my task, in my harness, on my codebase". The only instrument that answers the second question is your own usage log — which is why this whole series keeps insisting it be snapshotted and dated.

A name in a log is not a model entity

Treat log identifiers as routing labels, not capability entities.

big-pickle

A stealth model officially acknowledged by OpenCode. It is a real endpoint with a deliberately uninformative name, so any benchmark comparison against it is a comparison against an unknown.

code-supernova-1-million

A 2025 stealth marketing name. The original channel is gone, which means the label can no longer be resolved to anything — a log entry pointing at a retired alias.

kimi-for-coding

A subscription API alias. Two entries carrying this label may be served by different underlying models at different times, so consecutive calls are not a controlled comparison.

Inference · Why this matters for measurement

Every conclusion in the almanac is drawn from labels like these. That does not invalidate the conclusions, but it bounds them: "this label was used more this week" is a statement about routing, while "this model got better" is a statement about capability. The first is measurable from logs. The second is not, and conflating them is how a usage report quietly becomes a benchmark.