Skip to content
Shadow Theory

Extended inquiryExtended inquiry 1

Bounded perspective, real control

Why incomplete knowledge can coexist with exact action, and the precise point at which an aperture obstructs control.

Reading position 12 of 17

An agent never needs to possess the whole world in order to act within it. What it needs depends on the task. A missing distinction matters when different possibilities demand incompatible responses. Other hidden distinctions can remain hidden without preventing success.

This is the first extension of the published account of reflective freedom. The published paper establishes how retained commitments can govern assessment, rule installation and fresh decisions. The supplementary control models developed here ask a different supporting question: what can an embedded system accomplish through a limited interface? Their results belong to this extended website treatment. They are not additional claims of the published paper.

What a perspective preserves

Write the observational relation as

Ω→ΔR. \Omega \xrightarrow{\Delta} R.

The source state is in Ω\Omega, the aperture Δ\Delta supplies a readout, and RR is what the installed interface makes available. If Δ\Delta is many-to-one, several source states have the same readout. That is an exact limitation on discrimination. It does not yet specify a limitation on every possible action.

For a deterministic aperture, the three relevant tests are different:

Question Exact requirement
Can the source state be reconstructed? Δ\Delta must be injective.
Can a particular fact q(ω)q(\omega) be recovered? qq must be constant on each fibre of Δ\Delta.
Can the task be completed without regret? The source states in each fibre must share an optimal executable action.

A fibre is the set of source states giving one readout. Knowing which fibre contains the world can be enough to answer one question and insufficient to answer another. A standpoint is therefore a structured access relation, not a scalar amount of ignorance.

Consider a target register holding an unknown value xx and a receiver prepared at a desired value gg. The reversible swap

(x,g)⟼(g,x) (x,g)\longmapsto(g,x)

sets the target exactly without identifying xx. The unresolved value survives in the receiver. Conversely, even complete observation does not help a controller whose only permitted actuation is the identity. Adding an inaccessible spectator register changes the completeness of its world description while leaving this target task unchanged.

These examples block an unrestricted inference from partial knowledge to absent agency. They also identify the right questions: which actions are compatible with the remaining uncertainty, where can unresolved information go, and which transformations can the apparatus execute?

The exact decision obstruction

Let Θ\Theta be a finite set of source conditions, H\mathcal H a finite record alphabet, and Q(h∣θ)Q(h\mid\theta) the observation channel. Let A\mathcal A be the admitted finite action set. A policy π(a∣h)\pi(a\mid h) can use the record but cannot read the hidden condition directly.

An outcome kernel K(y∣θ,a)K(y\mid\theta,a) and a bounded task score u(θ,y)u(\theta,y) define

vθ(a)=∑yK(y∣θ,a)u(θ,y),vθ∗=max⁡avθ(a),ℓθa=vθ∗−vθ(a). v_\theta(a)=\sum_y K(y\mid\theta,a)u(\theta,y), \qquad v^*_\theta=\max_a v_\theta(a), \qquad \ell_{\theta a}=v^*_\theta-v_\theta(a).

The regret ℓθa\ell_{\theta a} measures the task value lost by selecting aa when θ\theta is the source condition. Every row has at least one zero. Define worst-source regret by

R(Q,ℓ)=min⁡πmax⁡θ∑h,aQ(h∣θ)π(a∣h)ℓθa. R(Q,\ell)=\min_\pi\max_\theta \sum_{h,a}Q(h\mid\theta)\pi(a\mid h)\ell_{\theta a}.

Theorem: exact finite task compatibility. The same value has the dual expression

R(Q,ℓ)=max⁡λ∈Δ(Θ)∑hmin⁡a∑θλθQ(h∣θ)ℓθa. R(Q,\ell)=\max_{\lambda\in\Delta(\Theta)} \sum_h\min_a\sum_\theta\lambda_\theta Q(h\mid\theta)\ell_{\theta a}.

Moreover, R(Q,ℓ)=0R(Q,\ell)=0 exactly when every possible record hh satisfies

⋂θ: Q(h∣θ)>0arg min⁡aℓθa≠∅. \bigcap_{\theta:\,Q(h\mid\theta)>0} \operatorname*{arg\,min}_a\ell_{\theta a}\ne\varnothing.

Proof. The policy set is a product of finite probability simplexes. Replace maximization over the source condition by maximization over source distributions λ\lambda. Finite minimax exchanges that maximization with policy minimization. With λ\lambda fixed, the policy optimization separates over records, and a minimizing action at each record gives the dual formula.

If each intersection is nonempty, choose an action in it. Every supported source-record pair then has zero loss. Conversely, zero worst-source regret means that every nonnegative term with positive Q(h∣θ)π(a∣h)Q(h\mid\theta)\pi(a\mid h) has zero loss. An action given positive probability at hh must be optimal in every source condition compatible with that record. At least one such action exists because the policy probabilities sum to one. This proves both directions.

The obstruction is incompatible required action under the same available record. Randomization does not remove that obstruction at zero error. Every action in a randomized policy's support must still satisfy the common requirement.

Pairwise agreement is weaker than joint agreement. With a single record and loss matrix ℓ=I3\ell=I_3, any two source rows share a zero-loss action, but all three do not. The minimax value is 1/31/3: assigning each action probability 1/31/3 attains it, and some action must receive at least that probability. Testing only pairs would miss the obstruction.

How indistinguishability becomes a quantitative limit

Total variation is

TV⁡(P,Q)=12∑h∣P(h)−Q(h)∣. \operatorname{TV}(P,Q)=\frac12\sum_h|P(h)-Q(h)|.

It measures how well the available records can distinguish two hypotheses. Suppose two source conditions have disjoint optimal action sets and every nonoptimal action loses at least γ>0\gamma>0. Then

R(Q,ℓ)≥γ2[1−TV⁡(Qθ0,Qθ1)]. R(Q,\ell)\ge \frac\gamma2\left[1-\operatorname{TV}(Q_{\theta_0},Q_{\theta_1})\right].

Proof. Their common record mass is c(h)=min⁡{Q(h∣θ0),Q(h∣θ1)}c(h)=\min\{Q(h\mid\theta_0),Q(h\mid\theta_1)\}, with total mass 1−TV⁡(Qθ0,Qθ1)1-\operatorname{TV}(Q_{\theta_0},Q_{\theta_1}). At any action, the two losses sum to at least γ\gamma because the action cannot be optimal in both conditions. Sum over the common mass and take half the resulting total. Maximum risk is at least average risk.

For several source conditions, put

α=∑hmin⁡θQ(h∣θ),rblind=min⁡pmax⁡θ∑ap(a)ℓθa. \alpha=\sum_h\min_\theta Q(h\mid\theta), \qquad r_{\rm blind}=\min_p\max_\theta\sum_a p(a)\ell_{\theta a}.

Then R(Q,ℓ)≥αrblindR(Q,\ell)\ge\alpha r_{\rm blind}. For α>0\alpha>0, the common record mass induces the same action distribution in every condition; discard the nonnegative loss on the remaining mass. For α=0\alpha=0, the statement is just nonnegativity.

Garbling an observation cannot improve the optimum when the actuator and resources stay fixed. If Q′=QBQ'=QB, every policy using Q′Q' can be simulated from QQ by first drawing the garbled record through BB. Its policy class is contained in the original one. This finite argument is the relevant instance of Blackwell's comparison of experiments.

The lesson for reflective agency is direct. An acceptable amendment can exist without being identifiable through the installed read ports. The published paper's common-amendment criterion applies this same structure to endorsement and continued auditability. More information can remove a conflict between compatible contexts. It cannot supply an unavailable actuator or make an inaccessible amendment executable before a deadline.

Capability is a set of achievable consequences

For a specified interface and policy-resource class Π\Pi, define

K(Π)={Kπ(⋅∣θ):π∈Π}. \mathcal K(\Pi)=\{K^\pi(\cdot\mid\theta):\pi\in\Pi\}.

This collects achievable outcome or transcript laws. It permits exact comparisons between a controller before and after learning, between two memory budgets, or between two operation sets. There is no need to force these comparisons into one universal number called freedom.

Theorem: a task witnesses a capability gap. Let C\mathcal C be a compact convex set of finite kernels, let μ\mu have full support, and let KK be a target kernel. Then

min⁡L∈C∑θμθTV⁡(Kθ,Lθ)=max⁡0≤u≤1[∑θ,yμθK(y∣θ)u(θ,y)−max⁡L∈C∑θ,yμθL(y∣θ)u(θ,y)]. \begin{aligned} \min_{L\in\mathcal C}\sum_\theta\mu_\theta\operatorname{TV}(K_\theta,L_\theta) =\max_{0\le u\le1}\Bigg[&\sum_{\theta,y}\mu_\theta K(y\mid\theta)u(\theta,y)\\ &-\max_{L\in\mathcal C}\sum_{\theta,y}\mu_\theta L(y\mid\theta)u(\theta,y)\Bigg]. \end{aligned}

Proof. For probability distributions of equal mass, TV⁡(p,q)=max⁡0≤u≤1∑yuy(py−qy)\operatorname{TV}(p,q)=\max_{0\le u\le1}\sum_yu_y(p_y-q_y). Apply this separately to each source row and exchange the minimum over the compact convex kernel class with the maximum over the score cube by finite minimax. The inner minimum subtracts the largest comparator score.

A new achievable law outside the old convex capability class therefore has a bounded task that reveals the difference. This is a precise sense in which learning can create an effective capability. It does not establish that the learner has escaped its underlying physical laws or acquired ultimate authorship of them.

A world model can be partial and adequate

If estimated action values obey ∣v^θ(a)−vθ(a)∣≤ε|\widehat v_\theta(a)-v_\theta(a)|\le\varepsilon for every value actually compared, the greedy estimated choice loses at most 2ε2\varepsilon relative to that comparison's true optimum. Add and subtract the two estimated values; the greedy inequality cancels the middle difference. The premise must cover the comparison being claimed. An error guarantee on accessible comparisons does not automatically reach an inaccessible source-oracle optimum.

Nor does a lossy observation automatically destroy Markov structure. For a controlled process on XX and a map ϕ:X→Z\phi:X\to Z, an exact controlled quotient exists when

ϕ∗Ka(⋅∣x)=ϕ∗Ka(⋅∣x′)whenever ϕ(x)=ϕ(x′),for every a. \phi_*K^a(\cdot\mid x)=\phi_*K^a(\cdot\mid x') \quad\text{whenever }\phi(x)=\phi(x'),\quad\text{for every }a.

Necessity follows by starting from either state in the same fibre. Sufficiency defines the quotient kernel by the common pushed-forward law and then iterates it. Memory is required when relevant distinctions fail this test, not merely because the representation is incomplete.

Broad competence can demand much richer models than a single reset task. General agents need world models supplies such necessity results under its own task and environment assumptions. A bounded perspective can support strong local competence without becoming a complete representation of the source. That is the setting in which reflective freedom has to be built and tested.