Skip to content
Shadow Theory

Sealed or Leaky Section 3

Experimental anchoring

Section 4 of 17

3 Experimental anchoring

3.1 A finite-sample likelihood statement, not a central-value entropy claim

Status: Statistical model.

Let Ci=(Ai,Bi)C_i=(A_i,B_i) and Wi=(Xi,Yi)W_i=(X_i,Y_i), with history Hi=(C1,W1,…,Ci−1,Wi−1)\mathcal H_i=(C_1,W_1,\ldots,C_{i-1},W_{i-1}). For a fixed classical initial TT, assume

P(Wi=xy∣T,Hi)=qxy>0 \Prb(W_i=xy\mid T,\mathcal H_i)=q_{xy}>0

and that the box of CiC_i conditional on T,HiT,\mathcal H_i is no-signalling. These are sequential premises, not consequences of the one-round assumptions. Their chain factorization is

P(cn,wn∣t)=∏i=1nqwiPi(ci∣wi,Hi,t). P(c^n,w^n\mid t)=\prod_{i=1}^n q_{w_i} P_i(c_i\mid w_i,\mathcal H_i,t). (5)

Summation over the outcomes gives P(wn∣t)=∏iqwiP(w^n\mid t)=\prod_iq_{w_i}. Thus conditioning on the whole setting word does not introduce a future-input factor into the outcome likelihood. This causal factorization is indispensable for the following likelihood calculation.

Put q=min⁡xyqxyq=\min_{xy}q_{xy}, σxy=(−1)xy\sigma_{xy}=(-1)^{xy},

Di=σXiYiAiBiqXiYi,S^=1n∑iDi,Si=E[Di∣T,Hi]. D_i=\frac{\sigma_{X_iY_i}A_iB_i}{q_{X_iY_i}},\qquad \widehat S=\frac1n\sum_iD_i,\qquad S_i=\E[D_i\mid T,\mathcal H_i].

This defines the estimator; it is not automatically the same as a published setting-frequency-normalized central value.

Lemma 3.1 (Conditional-range CHSH concentration)

Status: Proved conditional on the sequential setting law above.

For 0<α<10<\alpha<1, define

rn=1q2ln⁡(1/α)n. r_n=\frac1q\sqrt{\frac{2\ln(1/\alpha)}n}.

For almost every tt, Pt{n−1∑iSi<S^−rn}≤α\Prb_t\{n^{-1}\sum_i S_i<\widehat S-r_n\}\le\alpha. The constant uses the conditional range width 2/q2/q, not the larger bound 2(1/q+4)2(1/q+4) on centered increments.

Proof

Conditionally on the history and tt, Di∈[−1/q,1/q]D_i\in[-1/q,1/q]. The log moment-generating function of Di−SiD_i-S_i has value and first derivative zero at the origin; its second derivative is the variance under an exponentially tilted law, at most (2/q)2/4=1/q2(2/q)^2/4=1/q^2. Integrating twice gives the bounded-range inequality

Et[eθ(Di−Si)∣Hi]≤eθ2/(2q2). \E_t[e^{\theta(D_i-S_i)}\mid\mathcal H_i] \le e^{\theta^2/(2q^2)}.

Iteration and Markov's inequality give Pt(∑i(Di−Si)≥nr)≤inf⁡θ>0e−θnr+nθ2/(2q2)=e−nq2r2/2\Prb_t(\sum_i(D_i-S_i)\ge nr)\le \inf_{\theta>0}e^{-\theta nr+n\theta^2/(2q^2)} =e^{-nq^2r^2/2}. Substitution proves the assertion. Independence of device outputs between rounds was not assumed.

□
Proposition 3.2 (Finite-sample outcome-likelihood certificate)

Status: Proved conditional on (5) and conditional no-signalling; quantum alternative conditional on per-history (QT).

Define FNS(s)=0F_{\rm NS}(s)=0 for s≤2s\le2, FNS(s)=fNS(s)F_{\rm NS}(s)=\fNS(s) for 2<s<42<s<4, and FNS(s)=1F_{\rm NS}(s)=1 for s≥4s\ge4. Then for almost every tt,

Pt ⁣{−log⁡2Pt(Cn∣Wn)<nFNS(S^−rn)}≤α. \Prb_t\!\left\{-\log_2 P_t(C^n\mid W^n) <nF_{\rm NS}(\widehat S-r_n)\right\}\le\alpha. (6)

Under per-history (QT), replace FNSF_{\rm NS} by FQF_Q, obtained by the analogous zero extension below 22 and constant extension above 222\sqrt2 of fQ\fQ. The same rnr_n suffices. For a threshold s∗s_* fixed before the analyzed data, put δ∗=2−nFNS(s∗−rn)\delta_*=2^{-nF_{\rm NS}(s_*-r_n)} and J∗=1{S^≥s∗}J_* =\one_{\{\widehat S\ge s_*\}}. Then

Pt{Pt(Cn∣Wn)>δ∗, J∗=1}≤α. \Prb_t\{P_t(C^n\mid W^n)>\delta_*,\ J_*=1\}\le\alpha. (7)

Neither display alone is a bound on unsmoothed min-entropy of a postselected distribution.

Proof

Each joint outcome probability is no larger than an Alice marginal, so Lemma 2.3 bounds it by 2−FNS(Si)2^{-F_{\rm NS}(S_i)}. By (5),

−log⁡2Pt(cn∣wn)≥∑iFNS(Si)≥nFNS ⁣(n−1∑iSi). -\log_2P_t(c^n\mid w^n) \ge\sum_iF_{\rm NS}(S_i) \ge nF_{\rm NS}\!\left(n^{-1}\sum_i S_i\right).

The function is convex and nondecreasing on the physical domain [−4,4][-4,4]; its extension above that domain is not used in Jensen's inequality. Outside an event of probability α\alpha, Lemma 3.1 makes the last expression at least the right side of (6). If the empirical lower endpoint exceeds 44, that event is necessarily in the exceptional set. The quantum proof uses the imported one-round bound of [22], convexity of FQF_Q on [−22,22][-2\sqrt2,2\sqrt2], and the same concentration inequality. The fixed-threshold consequence follows by monotonicity. Postselection changes both the conditional law and the exceptional probability; Lemma 3.9 supplies that conversion explicitly.

□

3.2 Published CHSH values: illustrations only

Status: Imported central values; derived illustrative evaluations.

The reported values are 2.42±0.202.42\pm0.20 in Hensen et al. [23], 2.221±0.0332.221\pm0.033 in Rosenfeld et al. [26], and 2.0747±0.00332.0747\pm0.0033 in Storz et al. [27]. The corresponding reported trial counts and local-realist significance bounds are 245245, P≤0.039P\le0.039; 10410^4 in the cited run, P<2.57×10−9P<2.57\times10^{-9}; and more than 10610^6, P<10−108P<10^{-108}, respectively. These significance bounds test the authors' stated nulls; they are not source-entropy estimates.

ExperimentCentral SSfNS(S)\fNS(S)fQ(S)\fQ(S)gNS(S)g_{\rm NS}(S)
Hensen et al.2.420.16000.20750.2100
Rosenfeld et al.2.2210.08200.09260.1105
Storz et al.2.07470.02720.02830.03735
Table 1. Evaluations of ideal-box bounds, in bits per round, not finite-sample certifications. fNS\fNS and fQ\fQ bound the appropriate conditional min-entropy and also Shannon entropy; gNSg_{\rm NS} is the sharper no-signalling Shannon bound.

The fNS\fNS evaluations at one reported error bar below and above the central values are, respectively, [0.0816,0.2430][0.0816,0.2430], [0.0695,0.0946][0.0695,0.0946], and [0.0260,0.0284][0.0260,0.0284]; for fQ\fQ they are [0.0921,0.3838][0.0921,0.3838], [0.0769,0.1091][0.0769,0.1091], and [0.0270,0.0296][0.0270,0.0296]. These transformed error bars are not confidence-certified entropy intervals. An ideal identical-round interpretation is needed to read the central entries as constant per-round rates.

Status: Not established from the supplied summaries.

No finite-sample source-entropy number for these three data sets is certified here. Applying Proposition 3.2 needs the exact score or equivalent sufficient statistics, the actual setting law, applicable sequential assumptions and a specified confidence rule; a central estimate and its standard error do not supply those items. Trial lists have not been reconstructed. The photonic analyses in [24, 25] use different Bell statistics; the CHSH expressions are not substituted for those statistics.

3.3 Bierhorst-admissible information and the actual protocol word

Definition 3.3 (Bierhorst-admissible Tier-1 information)

Status: Imported model class; notation translated from [28, Eqs. (2)–(3)].

A classical random variable TT, fixed before the protocol, with no access to data produced during it, is B-admissible when, for each trial and almost every supported history and tt,

P(Xi=x,Yi=y∣T=t,Hi)=14,P(Ai=a∣XiYi,T=t,Hi)=P(Ai=a∣Xi,T=t,Hi),P(Bi=b∣XiYi,T=t,Hi)=P(Bi=b∣Yi,T=t,Hi).\begin{align}\Prb(X_i=x,Y_i=y\mid T=t,\mathcal H_i)&=\tfrac14,\tag{8}\\ \Prb(A_i=a\mid X_iY_i,T=t,\mathcal H_i) &=\Prb(A_i=a\mid X_i,T=t,\mathcal H_i),\tag{9}\\ \Prb(B_i=b\mid X_iY_i,T=t,\mathcal H_i) &=\Prb(B_i=b\mid Y_i,T=t,\mathcal H_i). \tag{10}\end{align}

Here Hi\mathcal H_i includes the settings and outcomes of earlier trials. All public training choices and the fixed Bell function are part of the conditioning context. This is a classical pre-existing isolated-adversary condition; it is not a theorem for quantum side information, later-acquired protocol data, or the internal state of a predictable settings generator. The published bounded-bias variants require their own revised parameters and are not silently invoked in the numerical application below.

Lemma 3.4 (Finite-seed obstruction to exact sequential setting independence)

Status: Proved.

Suppose the full setting word WW is a deterministic function of a finite-valued generator resource GsetG_{\rm set}, including any genuinely fresh randomness used during the protocol. The exact setting condition (8) implies

H(W∣T)=2n≤H(Gset∣T)≤log⁡2∣Gset∣. H(W\mid T)=2n\le H(G_{\rm set}\mid T)\le\log_2|\mathcal G_{\rm set}|.

In particular, a generator resource with less than 2n2n conditional entropy cannot satisfy that condition, even when TT does not reveal the generator state.

Proof

Average (8) over the earlier outcomes, keeping TT and earlier settings fixed. Each next setting pair is still conditionally uniform. The entropy chain rule gives H(W∣T)=∑iH(Wi∣T,W<i)=2nH(W\mid T)=\sum_i H(W_i\mid T,W_{<i})=2n. Deterministic data processing through GsetG_{\rm set} gives the remaining inequalities.

□

Status: Scope.

The preceding obstruction concerns the stronger sequential condition (8), not the original one-round (MI). Correlation between settings at different times need not itself imply correlation between a setting and the measured source.

Alphabet and symbol translation.

Status: Definitions.

Bierhorst et al. label detection by ++ and nondetection by 00. The bijection +↦+1+\mapsto+1, 0↦−10\mapsto-1 puts both into this article's alphabet; it changes no probability or entropy and discards no nondetections. Their Bell function is denoted here by B\mathcal B, not TT; their extractor seed by RR, not CHSH SS; and their extracted string by KK. Write W=XnYnW=X^nY^n, let CC be the protocol's processed outcome word, and let JJ be its pass indicator.

Causal padding is part of the record map.

Status: Imported implementation detail; not reconstructed data.

For Data Set 5, SI S.6 of [28] reports first crossing the fixed threshold at trial 41,243,97641{,}243{,}976, followed by relabelling every later outcome to nondetection. Thus CC is the raw word through that crossing, followed by deterministic (−1,−1)(-1,-1) pairs. Its remaining Bell factors equal one. This history-dependent change is covered by the paper's adaptive theorem: it is selected from the completed past, not from a current distant outcome. The reported raw final product 2.018×10412.018\times10^{41} is not the product of this padded word. We use the reported threshold crossing, not that raw product as a replacement threshold or as a recomputation of CC.

Source and numerical quantities.

Status: Imported published quantities.

Data Set 5 of [28] (arXiv:1803.06219v1 and SI) uses n=55,110,210n=55{,}110{,}210 trials after 5×1065\times10^6 training trials; the reported Bell-function parameter is m=0.0100425m=0.0100425. The fixed threshold, errors, seed length and output length are

v=1.5×1032,p=9.025×10−25,κ=9.5×10−13,p=κ2,ϵext=5×10−14,d=315,844,ℓ=1024,ϵ=10−12. \begin{gathered} v=1.5\times10^{32},\quad p=9.025\times10^{-25},\quad \kappa=9.5\times10^{-13},\quad p=\kappa^2,\\ \epsilon_{\rm ext}=5\times10^{-14},\quad d=315{,}844,\quad \ell=1024,\quad \epsilon=10^{-12}. \end{gathered} (11)

The probability π=P(J=1)\pi=\Prb(J=1) is a property of the repeated-protocol law, not of the particular observed word. The hypothesis π≥κ\pi\ge\kappa remains explicit; observing J=1J=1 does not prove it. The training-based i.i.d. success estimate in the paper is not substituted for an adversarial lower bound on π\pi.

Lemma 3.5 (Rounding-controlled entropy-production parameter)

Status: Proved from the printed Bell table and its stated downward-rounding rule in [28].

The printed coefficients bxyabb_{xy}^{ab}, in the original {+,0}\{+,0\} column order, are

xy+++00+00001.02435563530.97046478040.97355076581011.02561274090.94919512430.99607753341101.02272749880.99627827540.94610913831110.92730405631.00372172251.00392246451 \begin{array}{c|rrrr} xy & ++ & +0 & 0+ & 00\\\hline 00&1.0243556353&0.9704647804&0.9735507658&1\\ 01&1.0256127409&0.9491951243&0.9960775334&1\\ 10&1.0227274988&0.9962782754&0.9461091383&1\\ 11&0.9273040563&1.0037217225&1.0039224645&1 \end{array}

Under uniform settings, their maximum local expectation is 11 and maximum no-signalling expectation is 1.010042507751.01004250775. If the original coefficients are rounded down at the tenth decimal place, their actual no-signalling excess mactm_{\rm act} satisfies

0.01004250775≤mact≤m+:=0.01004250785. 0.01004250775\le m_{\rm act}\le m_+:=0.01004250785. (12)

Put

δ+=[1−exp⁡(ln⁡(pv)/n)−12m+]n,b+=−log⁡2δ+. \delta_+=\left[1-\frac{\exp(\ln(pv)/n)-1}{2m_+}\right]^n, \qquad b_+=-\log_2\delta_+. (13)

For the parameters (11),

1344.91401726<b+<1344.91401728,b++log⁡2κ+5log⁡2ϵext−11>1073.05>1064. 1344.91401726<b_+<1344.91401728, \qquad b_++\log_2\kappa+5\log_2\epsilon_{\rm ext}-11>1073.05>1064. (14)

Consequently δ+\delta_+ is a conservative entropy-production threshold, and the extraction constraint for ℓ=1024\ell=1024 remains satisfied. These checks do not reconstruct the experimental product from the rounded table.

Proof

Enumerate the sixteen local assignments (a0,a1,b0,b1)∈{0,1}4(a_0,a_1,b_0,b_1)\in\{0,1\}^4 and their means 14∑xybxyaxby\frac14\sum_{xy}b_{xy}^{a_xb_y}. The maximum is one, attained by all nondetections. The other eight vertices of the binary no-signalling polytope are the PR boxes [22, App. A.3]; their means are

18∑x,y,a∈{0,1}bxya, a⊕xy⊕rx⊕sy⊕t,(r,s,t)∈{0,1}3. \frac18\sum_{x,y,a\in\{0,1\}} b_{xy}^{a,\ a\oplus xy\oplus rx\oplus sy\oplus t}, \qquad (r,s,t)\in\{0,1\}^3.

Their maximum occurs at (r,s,t)=(0,0,0)(r,s,t)=(0,0,0) and equals 4040170031/40000000004040170031/4000000000; each of the other seven means is below one. These finite sums can be checked directly with the displayed decimal rationals. Increasing every coefficient by less than 10−1010^{-10} changes the expectation under any normalized distribution by less than 10−1010^{-10}. Monotonicity of the maximum then proves (12). The largest printed local mean other than the all-nondetection strategy is 1−5.25×10−101-5.25\times10^{-10}. Raising every uncertain coefficient by 10−1010^{-10} still leaves each such mean below one. The all-nondetection coefficients are fixed to exactly one by the published training constraint. Thus the local bound for the original function also follows from this rounding envelope; no unchecked numerical optimization tolerance is needed.

For pv>1pv>1, the entropy-production expression for δ\delta is increasing in mm wherever its bracket is positive. The actual theorem therefore also holds with the larger δ+\delta_+. Its admissible range pv≤(1+3mact/2)npv\le(1+3m_{\rm act}/2)^n holds already with the lower endpoint in (12). Numerically stable evaluation is

b+=−nln⁡2ln⁡ ⁣(1−exp⁡(ln⁡(pv)/n)−12m+). b_+=-\frac{n}{\ln2}\ln\!\left(1- \frac{\exp(\ln(pv)/n)-1}{2m_+}\right).

The intervals in (14) follow by bounding logarithms with ln⁡x=2∑j≥0((x−1)/(x+1))2j+1/(2j+1)\ln x=2\sum_{j\ge0}((x-1)/(x+1))^{2j+1}/(2j+1) after binary range reduction, and bounding the positive exponential and −ln⁡(1−u)-\ln(1-u) series by their geometric tail bounds. Twenty terms in the range-reduced logarithms and four terms in each of the latter small-argument series suffice for the displayed intervals. Finally ℓ+4log⁡2ℓ=1064\ell+4\log_2\ell=1064. The conservative value agrees with the fragment's rounded 1344.91344.9, but the rounded printed mm is not treated as an exact extremal value.

□

The exact imports used.

Status: Imported conditional theorems from [28, Eq. (4), Eqs. (5)–(6), SI S.2 and S.5].

With TT B-admissible, the entropy-production theorem gives, almost everywhere in tt,

Pt{Pt(C∣W)>δ+, J=1}≤p. \Prb_t\{P_t(C\mid W)>\delta_+,\ J=1\}\le p. (15)

Here Pt(C∣W)P_t(C\mid W) is the probability assigned to the realized word and settings, not a maximum taken after observing the data. For the soundness theorem, additionally take RR uniform on {0,1}d\{0,1\}^d and independent of (C,W,T,J)(C,W,T,J), use the specified extractor K=Ext(C,R)K=\mathrm{Ext}(C,R), and suppose π≥κ\pi\ge\kappa. Then

TV ⁣(PKWRT∣J=1, Uℓ⊗PWRT∣J=1)≤p/π+ϵext≤10−12. \TV\!\left(P_{KWRT\mid J=1},\, U_\ell\otimes P_{WRT\mid J=1}\right) \le p/\pi+\epsilon_{\rm ext}\le10^{-12}. (16)

Here UℓU_\ell and UdU_d are uniform laws and PWRT∣J=1P_{WRT\mid J=1} is the product of PWT∣J=1P_{WT\mid J=1} and UdU_d, with coordinates put in the displayed order. The independence of RR makes the real and ideal (W,R,T)(W,R,T) marginals identical after passing. Equations (15)–(16) are not assumed for a larger class of side information than Definition 3.3.

3.4 Extracted randomness translated into source-response entropy

Lemma 3.6 (Uniformity-to-entropy continuity with the correct dimension)

Status: Proved; imports the entropy-continuity inequality of [29].

Let KK take M=2ℓM=2^\ell values and let EE be classical. If TV(PKE,Uℓ⊗PE)≤ϵ≤1−1/M\TV(P_{KE},U_\ell\otimes P_E)\le\epsilon\le1-1/M, then

H(K∣E)≥ℓ−h2(ϵ)−ϵlog⁡2(M−1)≥ℓ(1−ϵ)−h2(ϵ),pguess(K∣E)≤M−1+ϵ. H(K\mid E)\ge\ell-h_2(\epsilon)-\epsilon\log_2(M-1) \ge\ell(1-\epsilon)-h_2(\epsilon), \quad p_{\rm guess}(K\mid E)\le M^{-1}+\epsilon. (17)
Proof

Write δe=TV(PK∣e,Uℓ)\delta_e=\TV(P_{K\mid e},U_\ell). Equality of the EE marginals gives Eδe≤ϵ\E\delta_e\le\epsilon. Always 0≤δe≤1−1/M0\le\delta_e\le1-1/M, since a point mass is farthest from the uniform law on MM points. Apply Audenaert's bound to the diagonal MM-dimensional density matrices:

H(K∣E=e)≥log⁡2M−h2(δe)−δelog⁡2(M−1). H(K\mid E=e)\ge\log_2M-h_2(\delta_e)-\delta_e\log_2(M-1).

The loss function is concave and nondecreasing on [0,1−1/M][0,1-1/M], since its derivative is log⁡2((M−1)(1−u)/u)≥0\log_2((M-1)(1-u)/u)\ge0 there. Jensen and monotonicity give the first inequality of (17); log⁡2(M−1)≤ℓ\log_2(M-1)\le\ell gives the second. No uncontrolled δe>1/2\delta_e>1/2 exception or auxiliary majorant is needed. Finally, under the ideal law any guess measurable in EE succeeds with probability 1/M1/M. Total variation changes that success probability by at most ϵ\epsilon, including for the optimal real-law guess.

□
Lemma 3.7 (Translation of the certified extracted string)

Status: Proved conditional on B-admissibility, the published protocol and seed hypotheses, π≥κ\pi\ge\kappa, single outcomes and runwise outcome determinism.

Suppose the processed word is a deterministic function C=F(λ,W)C=F(\lambda,W) of an initial source state and the full setting word; let T=τ(λ)T=\tau(\lambda) be B-admissible. Let Zrun=F(λ,⋅)Z_{\rm run}=F(\lambda,\cdot) be the finite response function on the finite set of setting words. With the extraction hypotheses of (16),

H(Zrun∣T,J=1)≥H(Zrun∣T,W,R,J=1)≥H(K∣T,W,R,J=1)≥1024(1−10−12)−h2(10−12)>1024−1.1×10−9,pguess(Zrun∣T,W,R,J=1)≤2−1024+10−12.\begin{align}H(Z_{\rm run}\mid T,J=1) &\ge H(Z_{\rm run}\mid T,W,R,J=1)\notag\\ &\ge H(K\mid T,W,R,J=1)\notag\\ &\ge1024(1-10^{-12})-h_2(10^{-12}) >1024-1.1\times10^{-9},\tag{18}\\ p_{\rm guess}(Z_{\rm run}\mid T,W,R,J=1) &\le2^{-1024}+10^{-12}. \tag{19}\end{align}

Thus the unsmoothed conditional min-entropy is at least −log⁡2(2−1024+10−12)>39.86-\log_2(2^{-1024}+10^{-12})>39.86 bits, also after dropping W,RW,R from the conditioning. No independence of λ\lambda from WW is needed beyond the stated B-admissibility of TT for these implications.

Proof

Apply Lemma 3.6 to the law conditioned on J=1J=1, with E=(T,W,R)E=(T,W,R). The numerical Shannon loss is less than 1.06531×10−91.06531\times10^{-9} bits. The output satisfies K=Ext(Zrun(W),R)K=\mathrm{Ext}(Z_{\rm run}(W),R), so conditioning and deterministic data processing give the entropy chain in (18). They do not require the generally false identity H(Zrun∣T)=H(Zrun∣T,W,R,J=1)H(Z_{\rm run}\mid T)=H(Z_{\rm run}\mid T,W,R,J=1). Any guess of ZrunZ_{\rm run} from (T,W,R)(T,W,R) induces a guess of KK correct whenever the response guess is correct. Its ideal-law success probability is 2−10242^{-1024}, so (16) gives (19). Dropping conditioning cannot improve guessing. All distributions in this argument are conditional passing-ensemble laws, not entropies assigned to one realized string.

□
Corollary 3.8 (The correctly scoped 1024-bit Shannon consequence)

Status: Proved conditional on all hypotheses of Lemma 3.7.

The Data Set 5 soundness certificate implies the lower bound in (18) for every separately specified B-admissible classical readout TT. It is a Shannon bound on the finite response function conditional on passing. The number 10241024 is the extracted output length, not its certified unsmoothed min-entropy and not unsmoothed min-entropy of an unrestricted λ\lambda.

Proof

This is Lemma 3.7 with the published parameters. Keeping its probability space, conditioning and side-information class is essential to the statement.

□

3.5 A stronger finite-sample consequence without an extractor or source–setting independence

Lemma 3.9 (Atom-tail transfer through determinism and postselection)

Status: Proved.

Let Z,CZ,C be finite, EE classical standard Borel, C=f(Z,E)C=f(Z,E), and JJ a pass event of probability π>0\pi>0. Suppose 0<p<π0<p<\pi, 0<δ<10<\delta<1 and

P{J=1, P(C∣E)>δ}≤p. \Prb\{J=1,\ P(C\mid E)>\delta\}\le p. (20)

Then

pguess(Z∣E,J=1)≤min⁡{1,(δ+p)/π},H(Z∣E,J=1)≥(1−p/π)[log⁡2π−pδ]+.\begin{align}p_{\rm guess}(Z\mid E,J=1)&\le\min\{1,(\delta+p)/\pi\},\tag{21}\\ H(Z\mid E,J=1)&\ge (1-p/\pi)\left[\log_2\frac{\pi-p}{\delta}\right]_+. \tag{22}\end{align}

For clarity, define deletion-smoothed conditional min-entropy of a normalized law PP by optimizing over subprobability measures Q≤PQ\le P of mass at least 1−η1-\eta:

Hmin⁡del,η(Z∣E)P=sup⁡0≤Q≤PQ(all)≥1−η[−log⁡2∫max⁡zdQ(z,⋅)dμ(e) dμ(e)], H_{\min}^{\mathrm{del},\eta}(Z\mid E)_P =\sup_{\substack{0\le Q\le P\\Q(\mathrm{all})\ge1-\eta}} \left[-\log_2\int\max_z\frac{dQ(z,\cdot)}{d\mu}(e)\,d\mu(e)\right],

where μ\mu dominates the EE measures. The value is independent of that dominating choice. For the conditional passing law,

Hmin⁡del,p/π(Z∣E,J=1)≥−log⁡2δ+log⁡2π. H_{\min}^{\mathrm{del},p/\pi}(Z\mid E,J=1) \ge-\log_2\delta+\log_2\pi. (23)

The smoothing parameter is discarded probability mass; no claim that it equals a purified-distance or normalized-TV smoothing convention is made.

Proof

Let BB indicate the good event P(C∣E)≤δP(C\mid E)\le\delta, and restrict the conditional passing law to B=1B=1:

Q(z,de)=P(Z=z,E∈de,B=1∣J=1). Q(z,de)=\Prb(Z=z,E\in de,B=1\mid J=1).

This deletes mass at most p/πp/\pi. For every (z,e)(z,e) that survives, determinism and the good-event definition give P(Z=z∣E=e)≤P(C=f(z,e)∣E=e)≤δP(Z=z\mid E=e)\le P(C=f(z,e)\mid E=e)\le\delta. Taking μ=PE\mu=P_E therefore gives

dQ(z,⋅)dPE(e)≤δπ,pguess(Q;Z∣E)≤δπ. \frac{dQ(z,\cdot)}{dP_E}(e)\le\frac{\delta}{\pi},\qquad p_{\rm guess}(Q;Z\mid E)\le\frac{\delta}{\pi}.

This proves (23). The deleted submeasure has guessing probability at most its mass, proving (21).

Let g=P(B=1∣J=1)≥1−p/πg=\Prb(B=1\mid J=1)\ge1-p/\pi. The normalized good law has guessing probability at most δ/(πg)\delta/(\pi g) and hence Shannon entropy at least [log⁡2(πg/δ)]+[\log_2(\pi g/\delta)]_+. Conditioning on BB can only reduce average Shannon entropy, and the bad-law entropy is nonnegative. Therefore

H(Z∣E,J=1)≥g[log⁡2(πg/δ)]+. H(Z\mid E,J=1)\ge g[\log_2(\pi g/\delta)]_+.

The right side is nondecreasing in g≥0g\ge0, giving (22). No independence between ZZ and EE was used. In particular the lemma is stronger, for postselected posterior-response entropy, than a transfer argument requiring an independent prior response law.

□
Corollary 3.10 (Published-parameter finite-sample source-response certificate)

Status: Proved conditional on B-admissibility, the published entropy-production certificate, single outcomes, runwise outcome determinism and π≥κ\pi\ge\kappa.

For Data Set 5, take E=(T,W)E=(T,W), Z=ZrunZ=Z_{\rm run} and δ=δ+\delta=\delta_+ of (13). Without using the extractor or assuming source–setting independence, the published quantities imply

H(Zrun∣T,W,J=1)>1304.97 bits,Hmin⁡del, 9.5×10−13(Zrun∣T,W,J=1)>1304.97 bits,Hmin⁡(Zrun∣T,W,J=1)>39.93 bits.\begin{align}H(Z_{\rm run}\mid T,W,J=1)&>1304.97\ \text{bits},\tag{24}\\ H_{\min}^{\mathrm{del},\,9.5\times10^{-13}} (Z_{\rm run}\mid T,W,J=1)&>1304.97\ \text{bits},\tag{25}\\ H_{\min}(Z_{\rm run}\mid T,W,J=1)&>39.93\ \text{bits}. \tag{26}\end{align}

The Shannon and ordinary guessing conclusions remain true when WW is dropped from the conditioning. These are not obtained by substituting a central CHSH value into fNS\fNS.

Proof

Average (15) over tt and apply Lemma 3.9 to C=Zrun(W)C=Z_{\rm run}(W). Since π≥κ\pi\ge\kappa, deletion mass is at most p/κ=κp/\kappa=\kappa, and the smooth bound is at least b++log⁡2κb_++\log_2\kappa. The Shannon bound is at least

(1−p/κ)[b++log⁡2(κ−p)]>1304.9768>1304.97. (1-p/\kappa)[b_++\log_2(\kappa-p)] >1304.9768>1304.97.

Here monotonicity in π\pi follows directly because both nonnegative factors in (22) increase with π\pi. The unsmoothed bound is

−log⁡2((δ++p)/κ)>39.93. -\log_2\bigl((\delta_++p)/\kappa\bigr)>39.93.

Numerically b++log⁡2κb_++\log_2\kappa lies between 1304.976871304.97687 and 1304.976891304.97689; the Shannon correction to it is less than 1.3×10−91.3\times10^{-9}. These conservative intervals follow from Lemma 3.5; they do not identify 1344.91344.9 with an unconditional or unsmoothed entropy. Dropping public settings can only reduce guessing and increase Shannon entropy.

□
Lemma 3.11 (The fragment's prior-response tail and its additional premise)

Status: Proved conditional on B-admissibility, the entropy-production certificate, runwise outcome determinism and Zrun⊥W∣TZ_{\rm run}\perp W\mid T.

With the additional conditional independence just stated, for almost every tt,

Pt{Pt(Zrun=zobs)>δ+, J=1}≤p, \Prb_t\{P_t(Z_{\rm run}=z_{\rm obs})>\delta_+,\ J=1\}\le p, (27)

where zobsz_{\rm obs} is the realized response function. Independence of the full λ\lambda and WW conditional on TT is sufficient but not necessary: response-level conditional independence suffices. In particular, runwise (MI) and T=τ(λ)T=\tau(\lambda) imply that sufficient condition. This does not redefine the (MI) premise of Section 2.1.

Proof

At the realized setting word ww,

Pt(C=zobs(w)∣W=w)≥Pt(Zrun=zobs∣W=w)=Pt(Zrun=zobs). P_t(C=z_{\rm obs}(w)\mid W=w) \ge P_t(Z_{\rm run}=z_{\rm obs}\mid W=w) =P_t(Z_{\rm run}=z_{\rm obs}).

The event in (27) is therefore contained in that of (15). This is an exceptional-event bound on a prior atom, not itself a conditional entropy after passing. Corollary 3.10 instead proves a posterior entropy statement without this independence premise.

□
Lemma 3.12 (Threshold test of a complete deterministic admissible readout)

Status: Proved conditional on B-admissibility and the published entropy-production theorem; reported threshold crossing imported.

For a predetermined threshold v>1v>1 and the same Bell-function protocol, any source-complete model with single outcomes and C=F(λ,W)C=F(\lambda,W) satisfies

P(J=1)≤v−1. \Prb(J=1)\le v^{-1}.

At the published v=1.5×1032v=1.5\times10^{32} this is 2/(3×1032)2/(3\times10^{32}). Unlike Corollaries 3.8 and 3.10, this test bound does not require π≥κ\pi\ge\kappa. Its rejection event is the reported crossing of the fixed threshold, not a threshold fitted to the final raw product.

Proof

Source completeness makes CC a function of (T,W)(T,W), so the realized conditional likelihood Pt(C∣W)P_t(C\mid W) equals one almost surely. In the entropy-production theorem choose an error parameter p′>v−1p'>v^{-1} arbitrarily close to v−1v^{-1}. Then p′v>1p'v>1, the corresponding δ(p′)\delta(p') is less than one, and the admissible-range condition holds for all sufficiently close choices. Its event inequality consequently bounds the probability of passing by p′p', for almost every tt and hence unconditionally. Let p′↓v−1p'\downarrow v^{-1}. This is a frequentist bound under the joint model hypotheses, not a posterior probability that a source ontology is false. The physical applicability of B-admissibility is not proved by the test.

□

3.6 Who certifies what, and what remains a gap

Remark 3.13 (Certification and source interpretation)

Status: Scope; imported certificate and proved translation distinguished.

Published certificate. Bierhorst et al.'s theorem concerns pre-existing classical information isolated from the protocol. Their passing-ensemble soundness inequality, seed requirement and statistical assumptions are (16) and Definition 3.3, not security against every physically imaginable no-signalling observer. The Methods section reports pseudorandom settings generated by the Mersenne Twister and explains the additional device-trust basis for effective independence [28]. Possession of its predictive internal state is outside the present admissible class. Neither the article nor this translation certifies independence from such a state. Merely hiding that state is also insufficient to prove the exact sequential law: Lemma 3.4 requires 2n=110,220,4202n=110{,}220{,}420 bits of conditional settings-resource entropy. Counting this many pseudorandom output bits does not certify that entropy. The source does not supply an independent verification of this exact generator-entropy condition, and no quantified replacement error for it is imported here. Thus the numerical implications remain conditional model statements, not unconditional information-theoretic certification of the pseudorandom implementation.

Proved here under determinism. Determinism turns recorded-word unpredictability into lower bounds on an initial finite response function, which is a function of λ\lambda. The extracted-string route gives nearly 10241024 conditional Shannon bits but only 39.8639.86 guaranteed unsmoothed min-entropy bits. The entropy-production route gives over 1304.971304.97 Shannon and deletion-smoothed min-entropy bits and over 39.9339.93 unsmoothed bits. The smooth and unsmoothed quantities are not interchangeable. No cardinality, differential entropy, or universal observer-access assertion about an arbitrary λ\lambda is inferred merely from the output length. Without deterministic source responses, the same statistics can instead represent fresh chance.

Passing matters. All near-10241024 and 1304.971304.97 entropy bounds refer to the passing-ensemble law and require the displayed lower bound on its pass probability. Conditioning on passing need not preserve no-signalling; none of the proofs assumes it does. One observed pass is not a measurement of π\pi, and an ensemble conditional entropy is not a property of one known output string. If LL denotes the conditional Shannon lower bound, the unconditioned conclusion is only H(Zrun∣T)≥πL≥κLH(Z_{\rm run}\mid T)\ge\pi L\ge\kappa L, by conditioning additionally on JJ and nonnegativity. A large unconditional entropy would need an independently justified much larger lower bound on π\pi.

Explicit unresolved empirical scope.

Status: Not established.

The supplied summaries do not establish the actual pass probability, exact sequential setting independence for the pseudorandom implementation, independence from every proposed additional variable, security with quantum side information, or an unconditional high-entropy bound for λ\lambda. Nor do they provide the exact scores and setting analyses needed to certify the three central-value CHSH examples. Those gaps are not filled by inventing trial lists. For the Data Set 5 theorem-level translation, the reported threshold crossing and published aggregate parameters do suffice under the stated hypotheses; independent reproduction of the experiment or extractor execution is not claimed. A claim about another error level, a different data set, or biased settings requires its own published or recomputed parameters.