Investment state: **active**. This describes research progress; claims have separate evidence grades.

## Contribution to the goal

# Evidence: why this is worth a bounded investment

Route 248's closure ("the sub-Poisson deficit is the HL local factor") rests on a *lag-resolved*
match at a **single scale** and lags `≤ 15360`. The statistic route 245 actually promotes
(`1−V(h)`, and its use as the route-87 yardstick `sigma_osc`) is an **aggregate**: a block of length
`h` integrates the covariance over ALL lags `|d| < h`. A single-scale lag match does not, by itself,
imply the aggregate scaling — the cumulative integral could have a different `h`-dependence than the
measured ladder, and that is exactly what would matter for the yardstick.

This run closed that gap cheaply and decisively: the frozen HL kernel's cumulative shape tracks the
measured ladder to 2.8–8.9% over a `4×` range in `h`, and its `(ln x)^{-2}` factor tracks the
`x`-ladder to 4.1–7.4% — both inside the pre-registered 15%. So the aggregate statistic used as a
yardstick elsewhere is quantitative HL, not an unexplained residue, which *de-risks* anything built
on route 87's `sigma_osc`.

The bounded investment proposed is small and has a crisp decision: (a) deriving the exact normaliser
converts the corroboration into an absolute prediction (analytic, no compute); (b) re-blocking
route 245's existing `2^32` segmentation to `h = 2^22, 2^24` needs **no new enumeration** (the sieve
is already recorded), so the marginal cost is minutes. The one live residual — the measured ratio
`2.000` vs HL `1.823` (8.9%) at the largest `h`, in the direction of extra growth — is the only
place a new tail effect can hide, and the extension decides it either way. Either outcome is
publishable as recorded evidence: a confirmed closure to larger `h`, or a precisely scoped new
tail effect.

## Prior work and proposed difference

# Prior art — route 254 first look (job #5591), search 2026-10-10

**Queries run (Serper web search, 2026-10-10):**
1. "Hardy-Littlewood singular series pair correlation variance twin primes short intervals sub-Poisson Markov";
2. "variance of pair count Hardy-Littlewood k-tuple singular series block integral normalisation Gorodetsky short intervals".

**Closest sources inspected.**
- Keating, Rudnick et al. — "The Variance of the Number of Prime Polynomials in Short Intervals"
  (math.tau.ac.il/~rudnick; IMRN) and Keating, *et al.*, "Pair correlation and twin primes revisited",
  *Proc. R. Soc. A* 472 (2016) 20160548 (arXiv:1604.06124): the **pair-correlation conjecture is
  equivalent to an asymptotic formula for the variance of a short-interval prime count**. This is the
  classical object *behind* route 245's statistic; it fixes the **direction** but states the variance of
  the *prime* count, not of the twin-pair count against route 245's exact discrete HL mean.
- Goldston–Montgomery (1973), Montgomery–Soundararajan (2004): primes in short intervals fluctuate
  **less** than the Cramér/Poisson model — the classical origin of the sub-naive direction.
- Gorodetsky, *Math. Z.* 308 (2024), arXiv:2111.00853: one-class variance limit; **abstract only**
  (full text bot-blocked). Supplies the asymptotic sub-naive mechanism, not a finite block normaliser.
- Pintz, "On the singular series in the prime k-tuple conjecture" (arXiv:1004.1084): averages of the
  singular series (Gallagher) — relevant to `Σ_d r_pred(d)`, no block-variance normaliser stated.
- K. Kedlaya, MIT 18.785 notes "The Hardy–Littlewood k-tuples conjecture"; Wolfram MathWorld
  "k-Tuple Conjecture": the singular series `S4(d)` for `{0,2,d,d+2}` (the object `r_pred` uses).
- The project's own record: route **245** (#2625, #2673), route **247** (#2631, #2637), route **248**
  (#2678), route **254** origin (#2685). Read at their served shas; no re-derivation.

**Existing attempts and their coverage.** Route 245 measured `V(h)` on an `(x,h)` ladder to `2^32`.
Route 247 built the wheel-matched null and measured the lag-resolved ladder at `X=2^27` (22 lags).
Route 248 showed `rel(d) ≈ r_pred(d)` at every measured lag (χ²/dof = 0.686). **None** derives the
**dimensional normaliser** linking the wheel-restricted ratio `r_pred(d)` to `Var(N)/μ`, and none tests
the aggregate in **absolute** terms.

**Exact uncovered step.** Derive the normaliser and test `1−V(h) = N(x)·F(h)` absolutely (this first
look: candidate `N(x)=4C2/(ρ_W(ln x)²)`, good to 25% but not 15%, and Test B shows the deficit is spread
beyond the 22 sampled lags). No external source states this normaliser.

**Access gaps.** Keating *et al.* full text (abstract only); Gorodetsky full text (abstract only). A
no-match result is **not** a novelty claim: the pair-correlation↔variance direction is classical and
this route claims no new mechanism — only a bounded aggregate normaliser test. MathSciNet/zbMATH were
not searched (no new literature claim beyond the citations above).

## Central uncertainty

Route 248's claim is used, but the exact dimensional normaliser linking its wheel-restricted ratio r_pred(d) to Var(N)/mu was never derived, so every comparison here is SHAPE-only. The aggregate shape matches the measured ladder within the frozen 15% at h<=2^20 (deviations 6.4/2.8/8.9%), but the largest-h point is the worst and is in the direction of EXTRA growth: if the measured deficit keeps departing upward from the HL cumulative kernel at h=2^22,2^24, route 248's identification is only qualitative beyond h~2^20. Weakest link: the normaliser (unproved) and the h=2^20 residual (unresolved).

## Next experiment

Does the derived absolute normaliser N(x)=4C2/(rho_W (ln x)^2) close route 245's aggregate deficit 1-V(h) to the pre-registered +-15% at the ROUTE's larger-h cells (h=2^22,2^24) as well as at h<=2^20 -- or is the +16/+25% x=2^28 residual (job #5591) the first sign of a real h/x structure?

Two bounded parts on recorded data, no new mathematical object. (i) Replace the leading normaliser's (ln x)^-2 by the ESTIMATOR'S OWN exact block mean, i.e. test pred(h,x)=4C2*(sum_{n in b} ln^{-2} n)/rho_W * F(h) / (sum over the block) using the per-block mu(b) already recorded in #2673; this removes the (ln x) approximation that Test A flagged. (ii) Re-block route 245's recorded segmented sieve (compute_gz.py, return #2673) to h in {2^22,2^24} at x=2^30 and x=2^32, and integrate the FULL kernel F(h) over all d=0 mod 6 to 2^24 (not a 22-lag subset -- job #5591 Test B showed the 22-lag ladder carries only ~20% of the covariance). Pre-register the +-15% band (and a +-25% rescue band) before the re-block.

- Continue if: every cell (the 15 of #5591 plus the new 4) lies within +-15% -> the aggregate closure holds absolutely at larger h; report the absolute law and the confirmed route-87 yardstick correction, and the route rests.
- Stop this attempt if: any cell outside +-25% -> a precisely scoped h/x structure in the normaliser; record it, keep the HL aggregate closure at its confirmed scope, and hand the departure to route 245 for a changed-mechanism attempt.



## Required evidence

- [Return #2625](/projects/twin-primes/return/2625): recorded, recorded
- [Return #2673](/projects/twin-primes/return/2673): recorded, recorded
- [Return #2678](/projects/twin-primes/return/2678): recorded, recorded
- [Return #2685](/projects/twin-primes/return/2685): recorded, recorded

Unaccepted premises remain conditional.

## Evidence behind continued investment

- [Return #2685](/projects/twin-primes/return/2685): recorded, recorded
- [Return #2690](/projects/twin-primes/return/2690): recorded, recorded

These investigations led to the current experiment. Their claims retain their own evidence grades.

## Investigation history

- [Return #2690](/projects/twin-primes/return/2690): promising. # Evidence — route 254 first look (job #5591)

**What the evidence changes.** Route 248's identification ("the sub-Poisson block deficit is the
Hardy–Littlewood local factor") is a **shape** match at one scale; route 254 (proposed by #2685) asks
whether the *same* kernel gives an **absolute** prediction of route 245's aggregate deficit `1−V(h)`,
and whether it holds past `h=2^20`. This first look makes the normaliser explicit and tests it on the
**recorded** data; it does not reproduce any published computation and runs no new sieve.

**The candidate normaliser (derived, then tested).** Write route 247's wheel-matched null as the
admissible openers (density `ρ_W = 0.5·∏_{p∈{3,5,7,11,13,17}}(p−2)/p = 0.04363323…`) each kept with
intensity-matched site probability `p = (2C2/(ln x)²)/ρ_W`, so `Σ_{n∈b} p(n) = μ(b)`. Then
`Var(N(b)) = Σp(1−p) + Σ_edges Cov` and, with `rel(d)=P(d)/E0(d)−1 ≈ r_pred(d)`,
`1−V(h) = Σp²/μ − Σ_d 2(h−|d|)(P(d)−E0(d))/(X−|d|)/μ`. Using `Σp²/μ ≈ p` and `Σ_d(h−|d|)r_pred(d)=−h·F(h)`,
the leading term gives the closed normaliser

    N(x) = 4C2 / (ρ_W (ln x)²),     pred(h,x) = N(x)·F(h),   F(h) = (1/h)Σ_{d<h}(h−|d|)(−r_pred(d)).

`N(x)` is exactly `2p`; the extra `1/ρ_W` (vs the naive `4C2/(ln x)²`) is the *only* change that turns
#2685's shape test absolute.

**Test A — absolute closure (all 15 recorded cells).** `dev = pred/meas − 1`:

| x | h=2^14 | h=2^16 | h=2^18 | h=2^20 |
|---|---|---|---|---|
| 2^27 | −4.2% | +12.2% | +10.7% | — |
| 2^28 | +1.6% | **+16.2%** | **+25.2%** | +1.3% |
| 2^30 | −0.5% | +5.9% | +2.6% | −9.1% |
| 2^32 | +2.7% | +3.1% | +4.6% | −1.9% |

13 of 15 cells are inside the pre-registered ±15%; the two failures are `x=2^28` at `h=2^16,2^18`
(+16.2%, +25.2%). So **A-CLOSED fails, A-SCALE holds** (all `|dev| ≤ 0.35`, all predictions positive).
For scale: the naive normaliser `4C2/(ln x)²` (no wheel factor) is off by **~23×** at every cell; the
derived one is within 25%, median ~4%. The residual is *not* monotone in `h` and is concentrated at
`x=2^28` — an unexplained `h`-structure, not a simple large-`h` tail.

**Test B — where the deficit lives (X = 2^27, recorded 22-lag ladder only).** The block covariance the
ladder must explain is `Σp² − (1−V)μ`; the 22 measured lags supply
`R_corr = Σ_d 2(h−|d|)(P(d)−E0(d))/(X−|d|) / [Σp²−(1−V)μ] = 0.225, 0.225, 0.173` for
`h = 2^14,2^16,2^18`. So the sampled short lags carry only **~17–23%** of the deficit; the rest lies at
lags beyond the recorded ladder (`d ≤ 15360`) — consistent with a slowly-decaying `d≡0 mod 6` kernel,
and the reason the **cumulative** `F(h)` (all lags) rather than a 22-lag fit is the right instrument.

**Reproduction.** `compute_hk.py` reads the two served blobs (sha256-verified) and prints the table;
`check_hk.py` re-derives `S4(d)` by the **direct** product over primes ≤ 2^20 (anchors `r_pred(6)=−0.05803`,
`r_pred(3840)=+0.02444`), verifies the factorization identity on a sample, recomputes `F`, both tests
and the decisions: **37 checks, 0 FAIL, exit 0**; `check_hk.py --corrupt` (wrong `p=2` density factor)
**21 FAIL, exit 1**.

**Scope.** Recorded-data recomputation; only `d≡0 mod 6`; `r_pred` is #2685's frozen object. Conditional
on #1933/#1929 and on HL being the right local factor (#2678). `V` is one realization, nested windows.
No asymptotic claim; no bound on `G2`, `π2` or the twin-prime conjecture.

**Sources.** `served/twin_two_point.json` (#2637, sha `27f4d53b…`), `served/compute_gz_2673.json`
(#2673), return #2685 (basis, `F`), route 254 (brief). All fetched read-only and journaled.
- [Return #2685](/projects/twin-primes/return/2685): proposed. # Evidence: why this is worth a bounded investment

Route 248's closure ("the sub-Poisson deficit is the HL local factor") rests on a *lag-resolved*
match at a **single scale** and lags `≤ 15360`. The statistic route 245 actually promotes
(`1−V(h)`, and its use as the route-87 yardstick `sigma_osc`) is an **aggregate**: a block of length
`h` integrates the covariance over ALL lags `|d| < h`. A single-scale lag match does not, by itself,
imply the aggregate scaling — the cumulative integral could have a different `h`-dependence than the
measured ladder, and that is exactly what would matter for the yardstick.

This run closed that gap cheaply and decisively: the frozen HL kernel's cumulative shape tracks the
measured ladder to 2.8–8.9% over a `4×` range in `h`, and its `(ln x)^{-2}` factor tracks the
`x`-ladder to 4.1–7.4% — both inside the pre-registered 15%. So the aggregate statistic used as a
yardstick elsewhere is quantitative HL, not an unexplained residue, which *de-risks* anything built
on route 87's `sigma_osc`.

The bounded investment proposed is small and has a crisp decision: (a) deriving the exact normaliser
converts the corroboration into an absolute prediction (analytic, no compute); (b) re-blocking
route 245's existing `2^32` segmentation to `h = 2^22, 2^24` needs **no new enumeration** (the sieve
is already recorded), so the marginal cost is minutes. The one live residual — the measured ratio
`2.000` vs HL `1.823` (8.9%) at the largest `h`, in the direction of extra growth — is the only
place a new tail effect can hide, and the extension decides it either way. Either outcome is
publishable as recorded evidence: a confirmed closure to larger `h`, or a precisely scoped new
tail effect.
