Skip to content
Shadow Theory

Paper 3 · Section 2Learning

What an interface must preserve

Preparations, joint outputs, adaptive actions and retained resources belong to one complete interaction law. This chapter defines the contract and the measurements used throughout the study.

Section 3 of 18

2 Operational contract and measures

2.1 Finite public interaction

A source fixture has a finite hidden carrier XX, four actions a∈{0,1,2,3}a\in\{0,1,2,3\}, four joint output symbols o∈{0,1,2,3}o\in\{0,1,2,3\} encoding two bits, and four public preparation codes p∈{0,1}2p\in\{0,1\}^2. Preparation selects an initial distribution μp\mu_p; it does not reveal the resulting hidden state. Its joint output/successor instrument is

Ka,o(x,x′)≥0,∑o,x′Ka,o(x,x′)=1.K_{a,o}(x,x')\geq0,\qquad \sum_{o,x'}K_{a,o}(x,x')=1. (1)

The fixed public clock permits at most eight calls per episode. Resource use, blocked service, mode changes and delayed return are represented by the hidden dynamics and observed outputs. A reset starts another charged preparation episode; it is not an uncharged mid-episode operation.

An experiment e=(p,H,π)e=(p,H,\pi) specifies a preparation, horizon H≤8H\leq8, and causal policy πt(a∣ht−1)\pi_t(a\mid h_{t-1}). The retained history is hH=((a1,o1),…,(aH,oH))h_H=((a_1,o_1),\ldots,(a_H,o_H)), with its timed position and the common stopping convention. Its law is

Pe(hH)=[∏t=1Hπt(at∣ht−1)]μpKa1,o1⋯KaH,oH1.P^e(h_H)=\left[\prod_{t=1}^H\pi_t(a_t\mid h_{t-1})\right] \mu_p K_{a_1,o_1}\cdots K_{a_H,o_H}\one . (2)

For deterministic policies the policy factor is zero or one. The same policy is used for source and candidate evaluation. Joint outputs are not factorized. Matching a singleton marginal is weaker than matching Equation 2.

A candidate may be a finite unifilar transducer, a tree with explicitly update-closed history registers, or a latent-state model whose public state is a posterior vector. Each is executable, but executable syntax alone does not certify source-faithful prediction. Zero-probability branches need an explicit fallback convention; no equality of conditional laws is inferred at a branch that is impossible in the source.

2.2 Adequacy beyond a finite score

For a fixed admitted family E\E, define the target-relative law discrepancy

dE(P,Q)=sup⁡e∈ETV⁡(Pe,Qe),TV⁡(P,Q)=12∑h∣P(h)−Q(h)∣.d_{\E}(P,Q)=\sup_{e\in\E}\TV(P^e,Q^e),\qquad \TV(P,Q)=\tfrac12\sum_h|P(h)-Q(h)|. (3)

The companion paper supplies a sufficient actual-state interface construction: a quotient q:X→Zq:X\to Z must preserve protected marks and make

∑x′:q(x′)=z′Ka,o(x,x′)=K‾a,o(q(x),z′)\sum_{x':q(x')=z'}K_{a,o}(x,x') =\overline K_{a,o}(q(x),z') (4)

independent of the discarded representative. Under the specified matching causal context, retained joint reference variables, resource ownership and prepared-law pushforward, such quotients admit substitution. Uniform conditional joint-row errors ϵi\epsilon_i for at most nin_i calls and initial discrepancy δ0\delta_0 yield the imported finite-history bound

dE(P,P‾)≤1−(1−δ0)∏i(1−ϵi)ni≤δ0+∑iniϵi.d_{\E}(P,\overline P)\leq 1-(1-\delta_0)\prod_i(1-\epsilon_i)^{n_i} \leq\delta_0+\sum_i n_i\epsilon_i. (5)

These are conditional results from [16], not new theorems of this paper. Transferring distance from a resource-restricted simulator family additionally requires transporting that family; target-law fidelity alone is insufficient. We do not repeat those proofs.

The OII learners do not observe XX or receive a map qq satisfying Equation 4. Neither a small fitted model nor success on finitely many tests verifies its universal row hypotheses. The benchmark asks how well candidate interfaces predict nominated legal histories, while leaving the stronger source-quotient and native-resource certification obligations open.

2.3 Scoring units

For method mm, fixture ff and registered core slot jj, let vfmjv_{fmj} be the recorded law error. OII-1–3 use conservative upper values including their stated evaluation allowances; OII-4 uses the full-support floating law evaluation, with separate rigorous event enclosures for selected failure witnesses. The benchmark pass allowance is τ=0.15\tau=0.15 throughout, an engineering tolerance rather than a physical boundary.

For FF fixtures with JJ core slots each, the main summaries are

Npass(m)=∑f,j1{vfmj≤τ},v‾m=1FJ∑f,jvfmj,Nall(m)=∑f1{max⁡jvfmj≤τ},v‾max⁡,m=1F∑fmax⁡jvfmj.\begin{align}N_{\mathrm{pass}}(m)&=\sum_{f,j}\one\{v_{fmj}\leq\tau\},& \overline v_m&=\frac1{FJ}\sum_{f,j}v_{fmj},\tag{6}\\ N_{\mathrm{all}}(m)&=\sum_f\one\{\max_jv_{fmj}\leq\tau\},& \overline v_{\max,m}&=\frac1F\sum_f\max_jv_{fmj}. \tag{7}\end{align}

The mean fixture-worst value is not the overall maximum or a bound on Equation 3. An all-core pass means that every slot in that fixture's finite panel passes, not that every legal future context does. Duplicated slots retain their originally assigned weight.

A red-team failure is a fixture for which the privileged postcommit search finds a legal event witnessing error above τ\tau. An unsafe-merge count instead concerns pairs of equal-clock histories whose true continuation laws differ by more than 2τ2\tau but share one retained model-state key. The panels and denominators are reported separately. No-witness outcomes are not adequacy certificates, and a belief vector can avoid exact key collisions while still predicting poorly.

Table 1. Four obligations that must not be collapsed.

ObjectWhat is establishedWhat is not established
Model syntaxA normalized executable prediction/update ruleAgreement with the unknown source
Core-panel scoreError on registered complete-law slotsUniform adequacy on all legal policies
Source interface theoremSubstitution under marked joint-row and causal-contract premisesThose premises for an opaque learned model
Post-reveal diagnosticA source-relative capacity, obstruction or selection factKnowledge or oracle access available to the learner