Paper 3 · Appendix CLearning
Exact examples that discipline the inference
A genuine refinement can worsen a different prediction objective; a valid safe replacement can be limited to its measured panel; an event can refute a whole candidate bank. The proofs show precisely why.
C Analytical controls on interpretation
The following finite arguments support specific diagnoses. They are not convergence theorems for the implemented learners and do not replace the interface results in the companion theory paper.
C.1 A true refinement can worsen the selected worst-context error
Take four equally weighted histories with deterministic next-output probabilities . A one-state maximum-likelihood fit is , with worst-history TV . A valid counterexample requires and to be separated. The refined partition fits probabilities and . It obeys that constraint and improves average negative log likelihood from to
At history , however, TV becomes . The statement concerns the result of likelihood refitting, not the optimum under a fixed worst-case loss. If and the same exact risk is minimized globally, then is immediate. A richer class does not itself worsen the best available minimax risk.
This counterexample was analyzed after OII-2 scoring. It motivated OII-3; it was not a preregistered prediction proving that objective mismatch caused every OII-2 failure.
C.2 What the safe-replacement rule certifies
Let be worst-context TV on one fixed measured panel . Suppose a simultaneous confidence event gives for every compared candidate, including the incumbent . Then
The proof is the chain . Comparing two upper bounds alone does not give this implication. Uniformity must cover the selected candidates; using adaptively chosen models with merely pointwise intervals would require another justification.
OII-3 obtained a candidate-uniform bound by bounding the true record law. For a pilot-selected atom set , independent calibration supplied simultaneous lower mass bounds . Every candidate then satisfies
The omitted mass remains uncertain rather than being declared impossible. Independent resets, stationarity and the stated multiple-comparison allocation are hypotheses of the probability guarantee. Whether the bound is useful depends on its resolution. It does not constrain contexts outside without a coverage or structural argument.
For a convex bank with vertex laws , suppose a convex upper objective has value on every vertex and a valid lower certificate proves . Convexity gives the reverse inequality . Thus the objective is identically on that bank. This is the conditional reasoning behind the 74 post-reveal flat-objective diagnoses. It does not make all uncertainty-aware objectives flat.
C.3 Exact operational twins and a finite certificate's reach
Two finite generators with matched initial laws and identical admitted complete record laws cannot be separated by a test using only that contract. Hidden coordinates can be split, relabeled or added without changing those laws. A learner's finite nondetection is weaker: two distinct laws can look identical on a small observation sample. Conversely, state labels from two numerical fits can differ even when their induced public laws agree.
An event witness of Equation 9 is sufficient to exclude the entire frozen convex bank. Failure to find one does not establish that any member is adequate. Similarly, the common-history inequality in Equation 11 is a sufficient refutation of a proposed shared prediction, while Equation 17 shows that it is not a complete criterion for a whole merged class.