📝 Markdown Draft — Ch.6 Pre-registration & H₁ (500w · P0)
Operator note: This is the FROZEN pre-registration. MUST be timestamped on OSF BEFORE A1 benchmark fires. After OSF deposit, deviations require explicit declaration in §8 Limitations.
Ch.6 — Pre-registration & H₁
§6.1 Rationale. Following Lakens (2022) and Ngiam (2022) preregistration discipline, we lock in our hypothesis, sample size, stopping rule, analysis plan, and decision criteria before the data are collected. The OSF deposit anchors the timestamp; any deviation in §11 (Empirical Bridge) is reported with rationale per the Open Science Framework standard. This protects the paper from p-hacking, HARKing, and selective reporting.
§6.2 Hypothesis (H₁, confirmatory).
H₁: On a held-out WikiText-103 slice, GF16 quantization yields lower bits-per-byte than bf16 by at least 0.02:
μ(BPB_GF16) ≤ μ(BPB_bf16) − 0.02, evaluated over 7 sanctioned seeds with paired Welch's t-test, α = 0.01 (two-sided), minimum effect size Δ_BPB ≥ 0.02 nats/byte.
Null (H₀): GF16 is no better than bf16 — μ(BPB_GF16) ≥ μ(BPB_bf16) − 0.02.
§6.3 Sealed seed pool (Fibonacci F₁₇..F₂₁ + Lucas L₇, L₈).
| Sanctioned seed | Source | Decimal value |
|---|---|---|
| F₁₇ | Fibonacci 17 | 1597 |
| F₁₈ | Fibonacci 18 | 2584 |
| F₁₉ | Fibonacci 19 | 4181 |
| F₂₀ | Fibonacci 20 | 6765 |
| F₂₁ | Fibonacci 21 | 10946 |
| L₇ | Lucas 7 | 29 |
| L₈ | Lucas 8 | 47 |
Seeds are closed under multiplication by φ (Lucas closure invariant) — a deterministic audit trail. Forbidden seeds: {42, 43, 44, 45} (gardener-policy banned, see trios-railway forbidden_seeds).
§6.4 Sample size justification. N=7 paired observations gives 80% power to detect Cohen's d≥1.0 at α=0.01 (one-sided), which corresponds to Δ_BPB ≈ 0.02 if pooled SD ≈ 0.02. This is consistent with bf16/GF16 SD observed in pilot runs at step=1000 in the IGLA fleet (bpb_samples, 7-day window). Power computed per Lakens (2022) sample-size justification template.
§6.5 Stopping rule (Lakens 2022). Two pre-registered stopping criteria, whichever comes first:
- All 7 seeds complete a full training run with valid (non-collapsed, non-mock) BPB rows at step ≥ 4000, OR
- Wall-clock cutoff: 2026-04-30T22:30 +07 (T-1h before operator submission target).
If criterion 2 triggers with N<7, we report a partial result with adjusted statistical power and treat it as exploratory in §13 Discussion.
§6.6 Analysis plan (locked).
- Primary test: Welch's paired t-test on per-seed differences
(BPB_GF16(s) - BPB_bf16(s)), two-sided α=0.01. - Effect size: Cohen's
d = mean_diff / sd_pooledwith bootstrap 95% CI (10,000 resamples). - Multiple comparisons: Bonferroni correction across {GF16, GF12, GF20} vs bf16 baseline (k=3 → α_per = 0.0033).
- Negative control: A "shuffled-bits" format (random permutation of GF16 exp/mant assignments) is included; expected
BPB_shuffled ≫ BPB_bf16. - Sliding-window evaluation: stride-64 (per Parameter Golf SOTA technique) — pre-declared, not post-hoc.
§6.7 Decision criteria.
- H₁ accepted: if p < α/k AND Δ_BPB ≥ 0.02 AND 95% CI excludes zero on the favorable side.
- H₁ rejected: if any of the three conditions fails. Result reported honestly in §13 (Negative Controls & Limitations) per Open Science principles.
- Honest abstain: if N<5 valid seeds at cutoff, no causal claim; results reported as exploratory.
§6.8 Pre-registered deviations protocol. Any deviation from this plan (e.g., seed substitution, training-budget change, evaluation slice change) MUST be explicitly declared in §8 with rationale, date, and authorship. We commit to honest disclosure even if the deviation favors H₁.
§6.9 OSF deposit. This pre-registration is deposited at osf.io/registries under "Hypothesis-Predicting" template, mirrored at submission/osf_prereg.pdf in this repository, with SHA-256 hash committed in commit phd/osf-prereg-seal ahead of A1 benchmark start.
§6.10 R7 triplet contract. Every empirical row written into bpb_runs carries the mandatory triplet:
BPB=<v> @ step=<N> seed=<S> sha=<7c> jsonl_row=phd-pr1 gate_status=PRE-REG-H1
This anchors every measurement in the pre-registration timestamp.
Citations
- Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), 33267.
- Ngiam, W. X. Q. (2022). Best practices with preregistration. VSS Open Science Workshop.
- Open Science Collaboration (2015). Science, 349(6251).
- John, L. K. et al. (2012). Psychological Science, 23(5), 524-532.
- NeurIPS Paper Checklist 2026.
Word count: 510 (target 500 ±10% ✓)
✅ Definition of Done (NeurIPS-aligned)
- H₁ statement explicit + falsifiable
- Sample size justified (N=7 from power analysis)
- Stopping rule fixed
- Analysis plan locked
- Decision criteria pre-declared
- Negative control included
- OSF deposit referenced
- OSF timestamp obtained (operator action — BEFORE A1 fires)
-
PR with
Closes #387+ tectonic compile + green CI
🤖 ONE SHOT directive (when operator types ONE SHOT Ch.6)
A3 SERGEANT R5: take the Markdown draft above, convert to LaTeX in
paper/sections/06_pre_registration.tex(~510 words), wire bib entries (Lakens2022, Ngiam2022, OSC2015, John2012), build OSF preregistration PDF via tectonic fromsubmission/osf_prereg.tex, deposit on OSF, capture timestamp + DOI/URL, commit + push, open PRCloses #387. Hard deadline: T-3h (before A1 benchmark fires).
phi^2 + phi^-2 = 3 · OSF SEALED · NEVER STOP 🌻