Skip to content
Shadow Theory

Paper 3 · Section 6Learning

What changes when fitting effort increases

Six starts and a longer update budget repair five of the 11 original selected failures on unchanged observations. The retrospective analysis preserves the original results and locates the remaining selection and prediction defects.

Section 7 of 18

6 Post-hoc optimization-budget sensitivity

6.1 Why another analysis was warranted

All 48 original 16-state fits reached the 60-update limit, with one initialization at each size. This records failure to trigger the implemented early-stop rule, not proof of mathematical nonconvergence. It nevertheless leaves a direct finite-budget explanation for part of the produced-bank gap. A subsequent Claude review identified this concern after the original results and audit were known. We therefore froze a separate retrospective protocol and retained the original OII-4 results as the primary study.

The target set is the 11 original R/S all-core failures, including the two selector misses, eight obstructed banks and one mixture-unresolved case. Four originally passing fixtures were chosen by a fixed hash ordering, not by mechanism or error magnitude. This is an outcome-selected diagnostic subset, not a random 15-system replication. No additional interaction data were collected.

6.2 Fixed fitting and public selection

For each fixture and each k∈{2,4,8,16}k\in\{2,4,8,16\}, six starts were run to at most 600 updates. Start zero repeats the original initialization; five additional seeds are fixed hashes of fixture, size and restart. The instrument family, 0.010.01 pseudocounts, fitting and calibration tapes, probability floors and relative stopping tolerance are unchanged. Each run retains the old-cap endpoint as well as its extended endpoint. In total, 360 fits produce 720 saved endpoint records, not 720 independent fits. All 60 original-seed short endpoints reproduce the original parameters exactly.

The principal expanded portfolio S+S^+ contains the original S bank and the 24 extended endpoints for its fixture. Average public calibration NLL selects the model; ties prefer original candidates and then the fixed size/restart order. All 15 batches and selections were committed together before any new core score was computed. The research process had already seen the original failures and source revelations; this local sequencing protects the new selection step, not author-level blindness. The protocol, worker access checks, seeds and commitments are supplied in Appendix F and the supplement.

Three predeclared auxiliary banks separate aspects of the sensitivity: original S plus the four extended original starts; original S plus the 24 old-cap endpoints from all starts; and the 24 extended HMMs without the old automata. The last comparison changes the bank and is not substituted for S+S^+ when it performs better.

Table 10. Post-hoc optimization sensitivity on the 11 outcome-selected failures. Each row uses calibration-only selection; produced counts are evaluated afterward.

ProcedureSelectedProducedSlotsMean worst
Frozen S0/112/11175/2640.652077
Longer, original starts3/115/11201/2640.429132
Six starts, old cap3/114/11193/2640.583670
S+S^+: six starts, long cap5/117/11212/2640.318750
Extended HMMs only6/117/11214/2640.309245

All but the last row retain the original eight-candidate S bank. Old and long caps are 60 and 600 updates; the unchanged tolerance can stop a fit earlier. The last row is an auxiliary bank restriction, not the principal sensitivity result. Original 48-system scores are unchanged.

6.3 Five selected repairs; two remaining selection misses

On the 11 failures, S+S^+ selects an all-core-passing model for five and its expanded bank contains at least one for seven. The frozen S figures on that same subset were zero selected and two produced successes. The selected pass count rises from 175 to 212 of 264 retained slots; mean fixture-worst TV falls from 0.6520770.652077 to 0.3187500.318750. These are post-hoc subset results, not a replacement 48-system headline. The four passing controls all retain their original selected models and all remain passing.

The five selected repairs are checksum protocol, alternating channel, shared budget, delayed receiver and finite budget. Extending only the original starts repairs three; multiple starts at the old cap also repair three, with different membership. The combined expansion repairs five (Table 10, Table 11). The original candidate-production result is therefore materially sensitive to the fitting/search budget. The experiment does not isolate a universal compute effect or certify globally optimal fitting.

Table 11. Individual outcomes on the 11 original failures. Scores are worst core-slot TV; original labels identify the historical diagnosis, not the cause of fitting failure.

Fixture familyOld typeFrozen SS+S^+SelectedClass
leaky occupancySel.0.1527810.246786k=08, r=1B
checksum protocolBank0.3097210.053293k=16, r=2A
alternating channelMix.0.5062130.080255k=08, r=2A
shared budgetBank0.4677880.040041k=16, r=2A
receiverBank0.8044790.016800k=08, r=5A
lock 0Bank0.9999911.000000k=02, r=2C
budgetBank0.8660140.066692k=16, r=5A
challenge 0Bank0.9391870.490052k=16, r=3C
challenge 1Bank0.9301910.315852k=16, r=5C
coinSel.0.1964760.196476old 03B
lock 1Bank1.0000001.000000old 00D

A: selected model passes; B: an expanded-bank individual passes but selection fails; C: none passes and best bank worst-error improves by at least 0.01; D: remaining smaller/no improvement. The two locks and challenges are labeled by their original variant. Complete hashed identifiers, fitting iterations and oracle-best candidates are in the supplement.

The two original selection-miss fixtures remain misses under S+S^+. On leaky occupancy, a new selected eight-state model has core-worst TV 0.2467860.246786, worse than frozen S's 0.1527810.152781, although a produced four-state model reaches 0.0655830.065583. Its public calibration NLL is slightly worse than the selected model's. The old-cap multi-start arm succeeds on this fixture, so additional iterations are not uniformly beneficial after selection. On the coin fixture, the original automaton retains the best observed calibration score and still fails the core panel. The HMM-only auxiliary selector passes it, giving six rather than five repaired fixtures, but discarding old candidates is a distinct bank intervention and not the principal result.

Both locks and both challenges still lack an individually all-core-passing model in the expanded bank. Best worst-slot error improves by at least 0.01 on one lock and both challenges; the second lock changes by less than that descriptive cutoff. The resulting classification is five selected repairs, two selection misses, three improved-but-failing banks, and one smaller/no-improvement bank. The thresholds define reporting categories, not significance or convergence claims.

6.4 Old certificates and new banks

The original eight obstruction certificates remain true about their frozen banks. Four of those fixtures now have a passing extended model, showing why the old certificates cannot be carried over automatically. Re-evaluating the original event against every member of each expanded bank gives three rigorous surviving lower bounds: 0.7936350.793635 and 0.9999980.999998 for the two locks, and 0.4151800.415180 for the single-match challenge. Each exceeds 0.15. The same event no longer excludes the offset-match challenge bank, although no individual candidate passes there. Its expanded-mixture adequacy remains unresolved by this nonexhaustive event check.

Of the 360 fits, 219 meet the unchanged stopping tolerance and 141 reach the 600-update cap. Among the targeted failures, 47 of 66 sixteen-state fits still reach the cap. These counts do not support a claim that the enlarged search exhausts optimization. Across all fits, endpoint fitting NLL improves relative to the old-cap snapshot on 286, calibration NLL on 221 and core-worst TV on 182; remaining cases include ties and deteriorations. The different counts, and the explicit selector misses, rule out treating likelihood improvement as guaranteed downstream adequacy.

6.5 What fails in the adaptive challenge

The revealed challenge generator alternates a random challenge and an immediate response. At phase zero it emits j∈{0,1,2,3}j\in\{0,1,2,3\} uniformly and retains (j,b)(j,b), where bb is a prepared bit. At the next call it emits (2+b)1{a=(j+d) mod 4}(2+b)\mathbf1\{a=(j+d)\bmod4\} and returns to the challenge phase; d∈{0,1}d\in\{0,1\} is the variant. Public actions and outputs use the saved permutations and XOR encoding. Thus the challenge value requires one-call retention, while the prepared bit persists through the episode. This is not a long-delay challenge mechanism.

Both selected tree controls use exactly the preparation code, time modulo two and previous output as split features, together with action-conditioned rows. Those features expose the relevant distinctions directly. Each original fitting tape contains 3,840 response events, covers all 64 preparation/challenge/action rows, and has at least 39 or 45 observations per row, respectively. The observed failure is not absence of every instance of a required response combination.

The frozen selected models both have eight latent states. After the first challenge, their posterior vectors for different challenge values at a fixed preparation have TV distances above 0.997. They therefore do not simply collapse or forget the observed challenge. Their predictions instead mix the two prepared-bit-dependent matching responses: the correct matching-response probability averages 0.50370.5037 and 0.49880.4988, versus above 0.99970.9997 for the tree controls (Table 12). Direct bit-contrasted traces are given in Table 25.

Table 12. Source-revealed challenge diagnosis after selection. The first-response probabilities average equally over four preparations and four challenges, choosing the matching next action.

FixtureFrozen STree PS+S^+S+S^+ core worst
Challenge 00.50370.99980.86670.490052
Challenge 10.49880.99980.98950.315852

The first three numeric columns are probability assigned to the deterministic correct matching response; the source value is one. They are diagnostic conditional probabilities, not extra primary tests. Both new selected models have 16 latent states; the original selected models have eight.

The selected sixteen-state extensions improve those matching-response averages to 0.86670.8667 and 0.98950.9895. Nevertheless, core-worst errors remain 0.4900520.490052 and 0.3158520.315852, and their selected core-pass counts are one and zero. On original-core prefixes, response-row TV decreases, while the error in the nominally uniform challenge-output rows increases. This localizes observed prediction defects more precisely than “insufficient memory,” without proving a unique statistical cause. The receiver preserves a prepared bit behind a native delay and the budget fixture carries a remaining-token count; both selected failures disappear under the same optimization expansion. Their common repair does not establish that they shared the challenge's response-mixing mechanism.