Skip to content
Shadow Theory

Paper 3 · Appendix DLearning

What the models store and what fitting costs

Automaton nodes, posterior vectors, streaming registers and mixture states have different meanings. The recorded sizes and fitting times keep those units and shared costs explicit.

Section 15 of 18

D Complexity and recorded computation

Table 21. OII-4 median size within each selected model type. Different size units are not directly comparable.

MethodSize unitFixturesMedian sizeParameter upperObserved keys
BLatent dimension486636521.5
EAutomaton states487844
GAutomaton states485602
HAutomaton states484676.556118526.5
PStreaming registers48411443
RLatent dimension1881020533.5
RAutomaton states306723
SLatent dimension1981020535
SAutomaton states296723

Automaton sizes include the reset-only type and unknown sink. Latent dimension is not a public predictive-state count. Observed keys depend on the finite audit history panel and floating belief representation. Parameter upper counts are storage-table counts, not statistical degrees of freedom.

An automaton's live state is a discrete current node, while a hidden-state predictor retains a posterior vector. For kk latent coordinates the latter can visit arbitrarily many distinct beliefs over unbounded histories, even when the hidden carrier is finite. The stored parameter array and the precision of that vector are separate costs. Observed floating state keys are a finite implementation statistic, not a canonical quotient cardinality.

The passive tree's live state contains all registers needed to update the selected features, not only its current prediction leaf. A mixture additionally stores its constituent predictor states and episode-level posterior weights. Counting a mixture as one state would conceal those costs. Parameter upper counts in the table describe the saved array structures; normalization constraints mean they are not numbers of free parameters.

Table 22. Recorded component-fitting times (seconds), summed over the components entering each method.

MethodMedian per fixtureMaximumTotal
E0.0915420.4822666.150401
G3.59279121.781511289.814325
B1.4242001.71547668.120519
P1.2264021.65500560.662023
O base bank1.6199682.21896381.513990
H0.0423400.2757814.446267
R5.10937623.979253364.085244
S1.5618102.19774274.270920

These are preserved original timings, not new controlled runtime experiments. R and S reuse E/G/B fits, and P/O share tree fits; totals overlap and must not be added. Mixture optimization and other selector overhead are not all included.

All timings were recovered from original per-component fitting logs. They share hardware and software context but were not collected as a compute-matched performance study. Portfolio sums overlap: R and S reuse the same fitted automata and hidden-state candidates, while P and O share tree fits. Repeating these totals as if each portfolio paid an independent experimental budget would overcount both evidence and component computation. The reported sums do not uniformly include selector, serialization or scorer overhead.