Locked Prospective Protocol
Tag: p02.5-preregistered
Commit 97eabba (2026-08-23)
Prospective Preregistration Protocol
Phase P02.5 Gate Screening & Phase P04 Production Matrix Execution Protocol under Research Planning Framework (RPF v2.0).
Cryptographic Git Commit Provenance
In accordance with open science standards, this protocol was formulated, timestamped, and cryptographically frozen in Git prior to the execution of any Phase P06 production training runs.
• Initial Preregistration Tag: p02.5-preregistered (Commit 97eabba, 2026-08-23 18:24:10 UTC)
• Parameter Calibration Anchor: ADR-006 (Subgate Calibration, 2026-08-25)
• Production Release Anchor: v2.0-paper02 (Commit f9ba574)
• Authoritative Web Record: https://basyirin-dev.github.io/sigma-model/preregistration.html
• Parameter Calibration Anchor: ADR-006 (Subgate Calibration, 2026-08-25)
• Production Release Anchor: v2.0-paper02 (Commit f9ba574)
• Authoritative Web Record: https://basyirin-dev.github.io/sigma-model/preregistration.html
0. Experimental Matrix Taxonomy & Grid Hierarchy
The experimental plan comprises a mutually exclusive and collectively exhaustive (MECE) 960-run production hierarchy:
- Locked Tier 1 Primary Falsification Matrix (720 runs): 4 benchmarks (ℏ, SCAN, COGS, PCFG-SET) × 1 canonical architecture (Transformer 2L, 0.93M params) × 6 standardized discrete levels (λ ∈ {0.000, 0.015, 0.020, 0.025, 0.030, 0.500}) × n=30 independent seeds = 720 runs.
- Tier 2 Architecture Scaling Matrix (240 runs): 4 benchmarks × 2 scaled architectures (Deep Transformer 4L with 3.80M params; GRU Seq2Seq with 0.41M params) × 3 discrete regimes (λ ∈ {0.000, 0.025, 0.500}) × n=10 independent seeds = 240 runs.
- Derived 11-Point Dense Grid on ℏ (330 evaluated conditions): Formed by the 180 ℏ Tier 1 runs plus 150 targeted boundary-refinement runs (λ ∈ {0.010, 0.018, 0.022, 0.028, 0.050}, n=30 seeds each) used for continuous inflection point estimation.
- Derived Interventional & Grokking Arms: Evaluates 180 branched trajectory checkpoints (30 seeds × 6 timepoints) from the subcritical baseline and 30 extended-horizon 20k-step anti-grokking control runs.
1. Primary Outcome (Criterion 1)
| Metric Name |
Criterion 1 — Sharp escape step-function (Fitted 2-Parameter Logistic Change-Point Estimand & Standardized Transition Scale):Pfit(escape | λ) = 1 / (1 + exp(-k(λ - λcrit))), escape := AccOOD ≥ 80%, and standardized transition scale P*(λ) = (Pfit(λ) - Pbaseline) / (1 - Pbaseline).
|
| Threshold |
Steepness k ≥ 15.0, macroscopic jump ΔPfit ≥ 0.50, and standardized threshold separation P*(λ ≤ 0.010) < 0.50 (with P*(λ ≤ 0.000) < 0.05) and P*(λ ≥ 0.500) > 0.95.
Empirical Outcome: k = 79.48 >> 15.0 (PASS); ΔPfit = 0.7496 >> 0.50 (PASS); P*(λ = 0.010) = 0.1968 < 0.50 (PASS); P*(λ ≥ 0.500) = 1.000 > 0.95 (PASS).
|
| Direction | Higher k = sharper boundary = PASS; separation inequalities must hold. |
| Justification | Distinguishes a macroscopic phase boundary (trap → escape) from a smooth dose-response regularizer curve. In the continuous ODE (Level 1), subcritical escape is identically 0%. In finite-sample discrete mini-batch training (Level 3), small baseline escapes represent stochastic fluctuations from initialization and early mini-batch gradient variance. |
2. Secondary Outcomes (Confirmatory)
| Metric Name | Preregistered Threshold | Tag | Observed Result |
|---|---|---|---|
| Criterion 2 — Late-Onset Recovery | For tint = 1000 runs (trapped at step 1000, AccOOD ≤ 50%), switching on λ = 0.050 reaches final AccOOD ≥ 90% in ≥ 90% of seeds. | Confirmatory | 30/30 (100.0%) |
| Criterion 3 — Supercritical Asymptotic Parity | Pairwise TOST equivalence among λ ∈ {0.025, 0.030, 0.050, 0.100, 0.500} within margin ±2.5%, p < 0.05 (Bonferroni across 10 pairs). | Confirmatory | All pairs p < 0.005 (Qualified for 2L post-separation λ ≥ 0.030) |
3. Exploratory Outcomes
- λ̂crit Point Estimate: Logistic inflection point point estimate and 95% bootstrap CI within [0.015, 0.030], matching the analytical formula λcrit = bC / aC = 0.025.
- Transition-Latency Ordering: τ̂fixed < τ̂mult < τ̂add corresponding to compositional operator complexity.
- CKA/RGA Lead-Lag Dynamics: Granger causality test demonstrating internal representation alignment precedes behavioral generalization jumps (ΔRGAt → ΔOODt+1).
- Whitened GCA Step-0 Orthogonality: P⊥emb verification that embedding layer initializations exhibit exact step-0 orthogonality without spurious alignment artifacts.
4. Objective Exclusion Criteria
| Exclusion Criterion | Operational Rationale | Observed Incurrence (960 Runs) |
|---|---|---|
| NaN or Inf loss/metric at any step | Numerical instability / floating-point overflow | 0 runs (0.0%) |
| Training non-convergence (loss non-decreasing ≥ 500 steps) | Optimization breakdown | 0 runs (0.0%) |
| Checkpoint / file artifact SHA-256 hash mismatch | Data corruption | 0 runs (0.0%) |
| Hardware / OS crash (OOM, killed process) | Compute infrastructure fault | 0 runs (0.0%) |
| Seed mismatch (≠ cell_seed_idx * 42 + 7) | Determinism failure | 0 runs (0.0%) |
5. Statistical Plan & Power Calibration
Sample sizes are calibrated to ensure statistical power > 0.90 for detecting macroscopic phase-boundary jumps:
- Sample Size & Power: n = 30 independent seeds per cell. For binomial escape fraction, standard error is bounded by SE ≤ √(0.5 × 0.5 / 30) ≈ 0.091, resolving the required separation (ΔP ≥ 0.50) with > 5 SE margin.
- Model Selection Likelihood Formulation: Individual seed Bernoulli MLE (AIC = 2k - 2 ln L, saturated ceiling ln Lsat = -165.234) and aggregate proportions (AIC = n ln(RSS/n) + 2k).
- Family-Wise Error Rate Control: Bonferroni correction within each hypothesis family (e.g., α = 0.05 / 10 = 0.005 for TOST equivalence pairs), with pre-ordered gate hierarchy (Primary → Secondary).
- Decision Rule: PASS = Primary meets all criteria (k ≥ 15.0 ∧ ΔPfit ≥ 0.50 ∧ P*(≤0.010) < 0.50 ∧ P*(≥0.500) > 0.95) ∧ Secondary consistent.