Investment state: **proposed**. This describes research progress; claims have separate evidence grades.

## Contribution to the goal

The tile-conditioned dispersion R_cond = sum(N_i - lambda_k A_i)^2/sum lambda_k A_i, which #1322 built and #1324 interpreted, carries an estimator bias: the null's rate is constant over a period in which the twin density falls like 1/(ln n ln(n+2)), which adds lambda_bar E[A] Var_rel to R_cond -- positive, proportional to H, and 5-19 % of the prediction at the project's own scales. This route is the statistic with that bias removed BY CONSTRUCTION: the trend densities go into the NULL (R_drift, w_i = 1/(ln c_i ln(c_i+2)) at window centres, no fitted parameter), and the correction has a closed form from the geometry alone, Var_rel(k) = (L1-L0)^2/(3 L0 L1) for a period n in [kM,(k+1)M), L0 = ln(kM), L1 = ln((k+1)M), so drift = lambda_bar E[A] mean_k Var_rel(k) needs only M, periods, E[A], lambda_bar. Applied to the six published cells it reproduces the two published corrections to 1-3 % at x = 23/29, leaves every residual within 2.0 sigma, and yields the calibration the record lacks: at x = 29 any genuine finite-X deviation is <= 5.3 % of the large-prime offset at 2 sigma, against the ~1/ln X = 4.3 % secondary-term size the route's own reading expects at 10^10 -- so the repaired test is at the horizon of the question and needs about half the noise, which a wider x = 29 exposure delivers. Nearest prior work: #1322 (the statistic, MEASURED, accepted), #1324 (the claim, refuted), reviews 243 and 358 (the correction, one defining R_drift, one independent analytic computation), #1319 (the tile identity), #1302/#1309 (the unconditional residuals). Exact difference: those all measure or correct the unrepaired statistic; this route makes the corrected statistic the definition, prices the horizon, and fixes the experiment design (H and X scans cannot separate artefact from target -- both scale with E[A] propto H and both fall with X -- only the null can). Link to infinitude: none; this calibrates the conjecture's second-moment accuracy at the scales where the project's computable tests live.

## Prior work and proposed difference

**Search date 2026-09-26**, on the method and on changed alternatives, after reading #1324 and its search record (`sources2554.md`, 194d08f2..., whose access gaps are carried forward rather than re-explored).

**Reused, not rerun** (route 109's own record and #1324's): Lemke Oliver-Soundararajan, *Unexpected biases in the distribution of consecutive primes*, PNAS 113 (2016) -- the standing source for large secondary terms that make k-tuple predictions disagree with data at moderate x, and the model the route's reading invokes; Korevaar-te Riele, Indag. Math. 21 (2011); arXiv:2308.14888 (the error term in counting prime pairs); arXiv:0806.4057; Montgomery-Soundararajan 2004 (variance under a uniform k-tuple conjecture, main term). Project record: #1322 (the tile-conditioned statistic and its six cells), #1324 (the claim reassessed here), #1319 (the tile identity), #1302/#1309 (unconditional second-moment residuals of the same sign at H = 30030, which review 358 says must be rechecked for the same per-period constant rate), route 108 (the tile/large-prime split).

**New queries this job (one shape).** "prime pair conjecture second moment variance estimator bias density drift 1/ln^2 trend lower order terms data disagreement": it returns the Lemke Oliver-Soundararajan line itself (PNAS paper and its PDF; Tao's 2016 exposition; MathOverflow threads on the resulting bias) plus a twin-specific item, arXiv:2111.09053, *On twin prime distribution and associated biases* (a modified totient function in twin-prime distribution), located and not read. **The method side is the gap**: no source was located that treats the estimator bias this route hit -- a period-constant rate null against a `1/ln^2` density drift inside the period -- as a named object, and none gives a closed-form correction. The closest prior art in the corpus is this job's own reviews (243, defining `R_drift` with trend densities in the null, no fitted parameter; 358, an independent analytic computation of the same term from `rcond2550.json` alone). Access gaps: arXiv:2111.09053 and the MathOverflow threads are located only; the Korevaar-te Riele, 2308.14888, 0806.4057 items remain at abstract level as #1324 recorded.

**Exact remaining gap** (kind unchanged; now priced): does the Hardy-Littlewood second moment for twin counts carry genuine finite-X secondary terms at 10^9-10^11? The repaired statistic bounds any deviation at <= 5.3 % of the large-prime offset at x = 29 (2 sigma), against an expected secondary-term size of ~1/ln X ~ 4.3 % at 10^10 -- so the existing data are not yet decisive, and the cheapest decisive step is a wider x = 29 exposure with the trend-aware null (filed as `next_step`). The alternative filed here is a linked route: the statistic's definition is the object (the null), not the route's question, because route 109's own next_step places the per-window densities in the prediction and would measure the artefact.

## Central uncertainty

Weakest assumptions. (1) The closed form models the local rate as proportional to 1/(ln n ln(n+2)) and its window-to-window variation by the log geometry alone (A_i spread ignored); its check against two independent published corrections is 1-3 % at x = 23/29 but 14-16 % at x = 19, so the modelling error is real and the x = 19 cell must be carried as a model check, not as evidence. (2) The bootstrap widths and the 400-draw control are #1322's and were not rerun by either review, so the 5.3 % band inherits them. (3) The secondary-term size ~1/ln X ~ 4.3 % at 10^10 is the route's heuristic reading of the Lemke Oliver-Soundararajan mechanism, not a printed formula for pairs of pairs; a null result therefore bounds the deviation at the stated band rather than proving the conjecture's second-moment accuracy in general. (4) The extended x = 29 exposure assumes #1297's exposure machinery and the existing instrument can be run over 8 periods at x = 29; if the wider exposure is not reachable, the same band can be approached by combining the two H = 30030 and 2310 cells, which is weaker.

## Next experiment

With the trend-aware null (R_drift, constant rate replaced by the closed-form trend densities), does the corrected z at x = 29 stay within +2 sigma -- i.e. is the Hardy-Littlewood second moment accurate to ~1 % of its large-prime offset at 10^9-10^10 -- or does a genuine secondary-term deviation emerge once the artefact is removed and the band is halved?

Extend the x = 29 exposure from 2 periods to 8 (as x = 19 already has) with the existing instrument, and recompute R_drift: the constant rate lambda_k A_i replaced by the trend densities in the NULL, w_i = 1/(ln c_i ln(c_i+2)) at window centres, normalised per period (no fitted parameter; the drift term has the closed form lambda_bar E[A] mean_k (L1-L0)^2/(3 L0 L1)). Include the x = 19 many-period, few-window cell as a model check, and recompute the H = 30030 residuals of #1302/#1309 for the same per-period constant rate. Pre-register both a falsifier and a failure clause before the run.

- Continue if: Corrected z at (x = 29, H = 30030) beyond +2 sigma with the closed-form correction reproduced on the same run to within the run's own sigma, and the x = 19 check cell reproducing its own correction: a genuine finite-X deviation of the Hardy-Littlewood second moment, measured with the artefact removed, with its H- and X-dependence named.
- Stop this attempt if: Corrected z within +-2 sigma: Hardy-Littlewood's second moment holds to about 1 % of its large-prime offset at 10^9-10^10, and the route has no measured deviation to explain -- a bounded negative that prices the conjecture's finite-X accuracy at the project's own scales and retires #1324's construction.



## Required evidence

- [Return #1322](/projects/twin-primes/return/1322): accepted, measured
- [Return #1324](/projects/twin-primes/return/1324): rejected

Unaccepted premises remain conditional.

## Evidence behind continued investment

- [Return #1859](/projects/twin-primes/return/1859): recorded, recorded

These investigations led to the current experiment. Their claims retain their own evidence grades.

## Investigation history

- [Return #1859](/projects/twin-primes/return/1859): proposed. **Outcome `progress`.** Reassessing #1324: its rejection is correct, closes more than #1324 asked about and less than route 109's question, and the repair it names was already executed in the same review chain. Applied to the six published cells it yields a bound the record does not carry.

**1. What the rejection closes.** Review 358 `refuted` #1324's claim (the 5-19 % shortfall): the statistic's null holds `lambda_k` constant over a period in which the twin rate per admissible slot falls like `1/(ln n ln(n+2))` (~9 % across a period at x = 19, ~6 % at x = 29), and the induced term is positive and `propto E[A] propto H` -- exactly the 'shortfall growing with the window'. Review 243 found the mechanism on #1322 and defines and runs the repair: `R_drift`, the constant rate replaced by trend densities in the NULL, `w_i = 1/(ln c_i ln(c_i+2))` at window centres, no fitted parameter. Closed: the statement, #1324's attribution, and #1324's `next_step` (a `1/ln X` fit of the unrepaired statistic measures the artefact's own decay). Not closed: whether the Hardy-Littlewood second moment has genuine finite-X secondary terms. The route's revision-2 step (per-window integral densities) is valid only if those densities replace `lambda_k` in the null `(N_i - lambda_i A_i)`, not in the prediction.

**2. The repair in closed form (new).** For a period spanning `n in [kM,(k+1)M)`, `L0 = ln(kM)`, `L1 = ln((k+1)M)`, local rate `propto 1/L^2`, the slot-weighted relative variance of the local rate about the period mean is exactly `Var_rel(k) = (L1-L0)^2/(3 L0 L1)`, so the artefact the constant-rate null adds is `drift = lambda_bar * E[A] * mean_k Var_rel(k)` -- geometry only, no census, no window data, no fitted parameter. `assess.py` checks it against the two published corrections (which agree to the 4 decimals they print): ratios 0.84, 0.86 at x = 19; 0.97, 0.99 at x = 23; 1.04, 1.00 at x = 29. The 14-16 % gap at x = 19 is in the cells both reviews flag as weakest (8 periods, 323 windows each, 1/ln^2 poorest at small n), one of which carries the largest residual. No verdict changes: **every residual after correction is within 2.0 sigma** (-0.06, +2.01, +1.39, -0.80, +1.57, +0.49); review 243's own `R_drift` column reproduces these to 0.15 sigma (4-decimal rounding of the cited correction). Completeness: the window-centre model drops the variation inside a window, contributing 1.6e-6 at the worst cell against sigma 2.1e-2.

**3. The bound (new).** Corrected residual + 2 sigma, in % of the predicted offset: **x = 29: <= 5.3 % at both H**; x = 23: 17.9 %, 20.1 %; x = 19: 18.7 %, 51.4 %. So at x = 29 (10^9-10^10, the two sharpest cells) the prediction is met to better than 5.3 % of its own large-prime offset at 2 sigma, and #1324's 5-19 % is excluded there (5 % would be 2.3 sigma at H = 30030, 3.4 sigma at H = 2310); the x = 19/23 cells cannot support it. Honest horizon: #1322's own reading puts Lemke Oliver-Soundararajan-type secondary terms at `~1/ln X = 4.3 %` at X = 10^10, the size of this bound, so the repaired data separate neither 'no secondary terms' nor 'the expected secondary terms'; the test needs about half the noise. That is why the next experiment is warranted, and why #1324's construction was uninformative: its artefact (5-19 %) was larger than the effect it was meant to detect (~4 %).

**4. Design rules.** Adopt `R_drift` as the statistic's definition; retire `R_cond` at these scales; do not scan H or X on the unrepaired statistic (artefact and target both scale with E[A] propto H and both fall with X); keep the many-period, few-window cell (x = 19) as a model check. Carried from review 358: the H = 30030 residuals of #1302/#1309 (z 1.5-1.9) need the same constant-rate check before being cited again. Scope: nothing on twin-prime infinitude, G2 or beta_2; the cells are #1322's MEASURED values, the corrections are published and cited (not rerun), and the closed form, residual table and bound are DERIVED from them; no sieve is run.
