T27.AI

Blog

A multiplicity correction changed the deployment reading

2026-08-17 · 6 min read

A merged research note shows how a nine-placement selection by mean margin becomes a narrower claim after the selection family is included in the correction.

MeasurementReproducibilityStatisticsQuantization

A merged research note changes how one codebook selection should be read. The campaign selected a deployed arm by its mean margin, then applied a multiplicity correction to the nine placements from which that selection was made.

At model level, the comparison has n = 5 checkpoints. NEAR0 is the deployed arm: its mean margin is −4.03 %, with a 95 % interval of [−7.32, −0.63], p = 0.031 before the ×9 correction and 0.279 after it. The corrected table marks it TIE.

The table after the correction

PlacementMean95 % intervalp ×9Verdict
NEAR0 (deployed)−4.03 %[−7.32, −0.63]0.279TIE
MID−2.12 %[−3.07, −1.16]0.036BEATS

MID is marked BEATS in the nine-placement table: its corrected p value is 0.036. That label is narrower than a deployment recommendation. In the paired head-to-head comparison against NEAR0, MID differs by +1.99 % with an interval of [−1.17, +5.25] and p = 0.157, so the two arms are a tie at n = 5.

Why the mean was unstable

NEAR0 varies from −1.06 to −8.07 % across the five checkpoints, a 7.01 percentage-point range. MID varies from −1.21 to −3.09 %, a 1.88 percentage-point range. Removing the Pythia checkpoint moves NEAR0 to −2.98 % but leaves MID at −2.01 %.

A leave-one-checkpoint-out selection by mean margin chooses NEAR0 5 times out of 5. The note’s interpretation is methodological: selecting on a raw mean and reporting after a multiplicity correction can reward a high-variance arm, because the selection statistic and the reported claim answer different questions.

The substrate was part of the result

The same merged change records that the weights directory lived in another session’s /tmp scratchpad and disappeared mid-campaign. The corpus and model were recovered, then a ruler gate reproduced the reference values before post-loss measurements were treated as comparable. This is a reproducibility receipt, not a performance claim.

What this does not establish

A corrected label is not a deployment decision. It is a narrower claim with its selection family made visible.

The useful engineering lesson is to register the selection family and the statistic before treating a measured margin as a result. The receipt is the merged PR and its research notes.

What this does not settle

Receipts

Every figure above is measured, and the limits are named with it.