Developmental validation and training dynamics
A complete record of how the primary configuration and checkpoints were selected—kept explicitly separate from the sealed outcomes that decide the claims.
Validation selected models; it did not decide the claims.
The study used observed-cell validation data to choose one pilot configuration and one checkpoint for each primary run. These data were available during development, so they are analytically distinct from the sealed test outcomes reported in the Results chapter.
Each of the 20 primary runs recorded macro-validation accuracy every 500 steps through step 50,000. The resulting 2,000 checkpoints provide a complete view of convergence, seed dispersion, and checkpoint selection. They are useful for diagnosing training behavior, but they cannot independently support H1 or the composite mechanism claim.
6.2 · Developmental evidence
From pilot selection to frozen checkpoints.
Pilot and validation metrics selected hyperparameters and checkpoints. Training loss is retained only as a supporting optimization diagnostic.
Six equal-budget configurations.
The tie set contained both 0.0006 learning-rate candidates. The frozen tie rule selected the lower dropout: learning rate 0.0006, dropout 0.0.
Twenty completed runs, twenty frozen checkpoints.
Values are developmental macro-validation accuracy at the selected step. Final training loss is shown for diagnostic context.
OPM_SHAREDDOMAIN_GENERALISTPROC_UNTIEDPROC_CLONEObserved-cell validation across every checkpoint.
All 20 primary runs recorded macro-validation accuracy every 500 steps. These 2,000 developmental checkpoints drove model selection; sealed-test data remained excluded.
Accuracy at six points in training.
Darker cells indicate higher accuracy. Values are five-seed means with the across-seed standard deviation.
Loading validation milestones…
First checkpoint above each threshold.
Earlier is faster. Thresholds are descriptive developmental diagnostics, not preregistered claims.
Loading validation thresholds…
Cross-entropy from every primary run.
Loss plots remain available as supporting training diagnostics. Solid colors, direct labels, and explicit range whiskers replace transparency-based encoding.
Mean loss at six points in training.
Darker cells indicate lower cross-entropy. Values are means across five primary seeds.
Loading loss milestones…
Final-window cross-entropy by seed.
Each dot is a 500-step mean, positioned on a shared log scale from 1 to 1e−9.
Loading seed dispersion…
First window below each loss threshold.
Earlier is faster. A dash means the condition did not cross that threshold by step 50,000.
Loading threshold crossings…
Orders of magnitude removed.
Ratio of the first 500-step mean to the final 500-step mean for each condition.
Loading reduction metrics…
Checkpoint and validation graphs for all 20 runs.
Every run completed its fixed 50,000-step budget. These panels cover the remaining training-time fields recorded in the frozen run summaries.
Selected checkpoint step by seed.
Observed-cell validation selected one frozen checkpoint per run; sealed tests were not used.
Loading checkpoint selections…
Accuracy at the selected checkpoint.
The axis is intentionally restricted to 99.5–100.0% so small seed differences remain legible.
Loading validation results…
Selected checkpoint versus step 50,000.
Condition means across five seeds. Circles mark the frozen selected checkpoint; diamonds mark the last recorded validation checkpoint.
Loading selection comparison…
Final step versus final 500-step window.
A single terminal batch can be noisier than a window mean. This graph shows both five-seed condition averages on the same log scale without transparency.
Loading terminal-loss comparison…
Cross-entropy traces are developmental diagnostics, not sealed-test outcomes. Download the complete loss-analysis JSON or all 2,000 seed-window rows.
How to read the developmental curves
All four conditions learned the observed training distribution to very high validation accuracy. That convergence is important because it shows that the sealed recombination gap was not caused by broad optimization failure. The conditions diverged only when evaluated on combinations that were absent from training.
The selected-versus-final comparison also documents why the frozen checkpoint rule matters. The highest observed-cell validation checkpoint was selected independently for each run rather than assuming the final update was optimal. This selection procedure was fixed before sealed evaluation.