Blog
A merged research note shows how a nine-placement selection by mean margin becomes a narrower claim after the selection family is included in the correction.
A merged research note changes how one codebook selection should be read. The campaign selected a deployed arm by its mean margin, then applied a multiplicity correction to the nine placements from which that selection was made.
At model level, the comparison has n = 5 checkpoints. NEAR0 is the deployed arm: its mean margin is −4.03 %, with a 95 % interval of [−7.32, −0.63], p = 0.031 before the ×9 correction and 0.279 after it. The corrected table marks it TIE.
| Placement | Mean | 95 % interval | p ×9 | Verdict |
|---|---|---|---|---|
| NEAR0 (deployed) | −4.03 % | [−7.32, −0.63] | 0.279 | TIE |
| MID | −2.12 % | [−3.07, −1.16] | 0.036 | BEATS |
MID is marked BEATS in the nine-placement table: its corrected p value is 0.036. That label is narrower than a deployment recommendation. In the paired head-to-head comparison against NEAR0, MID differs by +1.99 % with an interval of [−1.17, +5.25] and p = 0.157, so the two arms are a tie at n = 5.
NEAR0 varies from −1.06 to −8.07 % across the five checkpoints, a 7.01 percentage-point range. MID varies from −1.21 to −3.09 %, a 1.88 percentage-point range. Removing the Pythia checkpoint moves NEAR0 to −2.98 % but leaves MID at −2.01 %.
A leave-one-checkpoint-out selection by mean margin chooses NEAR0 5 times out of 5. The note’s interpretation is methodological: selecting on a raw mean and reporting after a multiplicity correction can reward a high-variance arm, because the selection statistic and the reported claim answer different questions.
The same merged change records that the weights directory lived in another session’s /tmp scratchpad and disappeared mid-campaign. The corpus and model were recovered, then a ruler gate reproduced the reference values before post-loss measurements were treated as comparable. This is a reproducibility receipt, not a performance claim.
A corrected label is not a deployment decision. It is a narrower claim with its selection family made visible.
The useful engineering lesson is to register the selection family and the statistic before treating a measured margin as a result. The receipt is the merged PR and its research notes.
Every figure above is measured, and the limits are named with it.