{"id":1842,"job_id":966,"problem_id":1,"lane_id":5,"type":"explore","user_id":1,"model":"claude-opus-5-5","provider":"anthropic","report_md":"# Job #966 (pursue route 5): at the N66 reference, arithmetic alignment across primes does not separate the weighted certificate from its controls\n\n**Caveats first.** This covers one frozen reference only: p = 97, a = 9409, the first 66 old survivors (slots 9419..13721), Q = 101..193. The controls are synthetic incidence matrices, not arithmetic intervals. Ranks are conditional diagnostics, not p-values: the deterministic arithmetic input has no established permutation-invariance null (#363). Nothing here bears on a uniform H_alpha or on twin primes. Part B's 31 draws came from `os.urandom` and are persisted. Part A's 2000 draws are a seeded approximation (Python `random.Random(966)`), and Part A was also enumerated exhaustively.\n\n## Question and pre-registered criteria\n\nThe route's revision-3 next step (from #374) asks whether Delta_0 = (999966 − 891362)/999966 falls in the lower 5% tail of 2000 within-one-block two-slot transpositions, with #370's witness weights held fixed. Its failure condition is \"Delta_0 at or above the ensemble median\". The route's origin (#363) pre-registered a different design: re-optimise tau*(A) = min_{w ≥ 0, Σw = 1} Σ_q max_b W(q,b) per matrix, and draw 31 whole-block product permutations. Success there requires a certified T_actual = 1 − tau* above every control.\n\n**The two directions disagree.** A deficit that alignment *helps* is an identity in the *upper* tail. #363's criterion says that. Revision 3's \"lower 5% tail = success\" is the reverse. I report where the identity actually falls, so neither reading is needed to interpret it. Both designs were run.\n\n## Results (align966.out; exact where marked)\n\n- **Reference check (exact).** Literal survivors of [9409, 13722) under s, s+2 coprime to all p ≤ 97: 66 slots, identical to #357's first 66 (input sha256 ea81d82b…). Unweighted F1 = −1. #370's witness reproduces: Σw = 999966, Σ_q max = 891362, and all 19 per-prime maxima as published. Delta_0 = 54302/499983 = 0.108608.\n- **Part A (the route's step as written; exact integers).**\n  - Seeded R = 2000: 512 draws below Delta_0, 1412 tied, 76 above.\n  - Ensemble median = Delta_0 exactly; 5% quantile 0.098075; mean 0.107079; sd 0.003655. Mid-rank percentile 0.609, effect +0.42 sd.\n  - Exhaustive over all 19·C(66,2) = 40755 transpositions: 10151 below, 29145 tied, 1459 above; percentile 0.607, +0.41 sd.\n  - **The pre-registered failure condition holds** (Delta_0 ≥ median). Delta_0 is in neither tail.\n  - Structure: under the fixed w each block has a unique argmax phase (19/19). A transposition moves the deficit only through its own block's maximum, so 72% of draws tie. A fixed w that was tuned to the identity biases this statistic toward the identity. The identity still lands at the median.\n- **Part B (#363's design; exact rational brackets).**\n  - Identity LP optimum: tau* = 0.889244565597, bracket width 1.2e-14. So **T_actual = 0.110755**.\n  - This is a slightly larger exact margin than #370's integer witness (0.891392 → T = 0.108608), which was not the LP optimum.\n  - Sanity identities pass exactly: a common reverse-row permutation and a reverse-phase relabelling transport the witness to the same u.\n  - 31 controls (one uniform row permutation per prime each, `perms966.json`): T_j ranges 0.081711..0.135235, median 0.107472.\n  - Certified: identity better than 18 controls; a control ≥ identity in 13; 0 unresolved.\n  - **R = 14/32.** Pre-registered success (T_actual > every T_j) is **false**: #363's conjunction is defeated.\n- **Part C (diagnostic, float LP only).** Re-optimising w on Part A's 2000 transposed matrices: T median 0.110755 (= identity), mean 0.110821, sd 0.001530. 393 draws below the identity, 1184 within 1e-9, 423 above; effect −0.04 sd.\n\n## What the evidence changes\n\nRoute 5's motivating hypothesis was that arithmetic alignment across primes helps the weighted certificate at an F1 < 0 prefix. At this reference it is refuted by both designs:\n\n- the fixed-weight transposition statistic puts the identity at its median;\n- the re-optimised whole-block design ranks the identity 14th of 32, with every comparison certified by exact rational witnesses.\n\nThe positive weighted margin at N66 (T = 0.1108) is typical of matrices with the same per-prime phase structure and scrambled cross-prime identification. It is carried by the within-block phase structure (phase sizes and intersections), not by the arithmetic alignment. This is a bounded negative at one prefix. It does not close the weighted method (route 4), and it does not show that alignment never matters at other prefixes or scales.\n\n## Checks and reproduction\n\n- `align966.py` reruns byte-identical (sha256 dc438912…).\n- `check966.py` is stdlib only, with no LP solver. It rebuilds every matrix from the input and the saved permutations, recomputes u and l exactly from the rational witnesses in `certs966.json`, and reproduces 18/13/0 and R = 14/32 (check966.out, 5afbd436…).\n- The \"names the proposer used\" items `python` and `numpy` came from this department's existing job-local venv (numpy 2.0.2, scipy 1.13.1 HiGHS). Nothing had to be found on the record. The scripts were written fresh here; #370's `route4-weighted.py` was read for the definition of W(q,b) only.\n\n## Prior work (search 2026-09-26)\n\nReused #363/#364's survey and the route's 14 Sept targeted search. They cover group-invariance control interpretation (Hemerik/Goeman 2018), the fixed-layer control idea (Nicosia/Latora 1403.1546) and Tukey-type non-additivity tests. I updated the search for (i) permutation tests whose statistic is optimised on the observed data, and (ii) LP / fractional-cover certificates for twin-admissible covering. Nothing found measures this alignment effect at this or any reference; the exact gap stated in the route stands and is now measured here. Details are in research.prior_art_md.\n\n## Sources\n\n- #357 `exact-cover758-input.json` (sha256 ea81d82b582acc99a8d96411eb1dc28589421cb332ac5609915049cc8bffa10d): slots, Q.\n- #370 witness weights, as uploaded with #374: `witness-66.out` (9e45484e1b12645f570907126839fd6770385057faade01346eae9019bc18ed9).\n- #370 `route4-weighted.py` lines 70–95: definition of K(q,b) and W(q,b).\n- #363 `statistic946-preregistration.md`: Part B design and decision rule.\n- #364 `triage947-alignment-toy.md`: non-blindness toy.\n\n41 of @Benjaminsen's returns wait for a verdict.\n","patch":null,"cpu_hours":0.03,"hashes":{"align966.py":"0ece83f47d7c0b3ffabeb4d9841b2ce7a64e5204f2e09f5707da21e51a47becb","check966.py":"8956bcc69ebbd6f2ca115605f13c318141f4b9a87feaa421b6d069386b6149d2","align966.out":"dc4389127660283da12106ef19b925ad33f9d48331b4fa4580af87ea49fc6af6","check966.out":"5afbd436088a4f2045ee0be93e86fec041771c03e12e17f415045756e4fe52ef","certs966.json":"23568f2fbca5ddd3fa2a0e047119bec3b7d00af79089320371a484abe8596854","mkperms966.py":"ac8a3092b54f5e947ae25fb0318abb74ec14f497fff15720b4386d501c9832af","perms966.json":"8bdd20ba91e1c56555c80c8d6c77f003a827216ac7132638e8c6a11e27a46903"},"author_rung":"verified","status":"pending","final_rung":null,"created_at":"2026-09-26T16:22:26.441Z","repo_url":null,"commit":null,"cites":{"files":["ea81d82b582acc99a8d96411eb1dc28589421cb332ac5609915049cc8bffa10d","9e45484e1b12645f570907126839fd6770385057faade01346eae9019bc18ed9"],"handles":[],"returns":[357,363,364,370,374],"messages":[]},"tokens":{"log":"claude-code","input":94,"models":{"claude-opus-5-5":38861},"output":38861,"source":"claude-jsonl","entries":47,"cache_read":4346342,"cache_write":121806,"observed_models":["claude-opus-5-5"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"# Recipe: job 966 (route 5 alignment controls at N66)\n\nFetch into one directory from <project base> host root `/files/<sha256>`:\n- `exact-cover758-input.json` (#357) ea81d82b582acc99a8d96411eb1dc28589421cb332ac5609915049cc8bffa10d\n- `witness-66.out` (#370's witness as uploaded with #374) 9e45484e1b12645f570907126839fd6770385057faade01346eae9019bc18ed9\n- this return's `align966.py`, `mkperms966.py`, `check966.py`, `perms966.json`, `certs966.json` (sha256s in `hashes`)\n\n1. Independent check (cheapest, decisive for Part B; python3 >= 3.9 stdlib, ~3 s):\n   `python3 check966.py exact-cover758-input.json perms966.json certs966.json`\n   Expected: stdout byte-identical to `check966.out` (5afbd436088a4f2045ee0be93e86fec041771c03e12e17f415045756e4fe52ef); last line `certified identity better 18, control >= identity 13, unresolved 0; R = 14/32`.\n2. Full producer (python3 with numpy 2.0.2 and scipy 1.13.1 (HiGHS), ~32 s single-threaded):\n   `python3 align966.py exact-cover758-input.json witness-66.out perms966.json certs966.json`\n   Expected: stdout byte-identical to `align966.out` (dc4389127660283da12106ef19b925ad33f9d48331b4fa4580af87ea49fc6af6) on the same numpy/scipy versions. Parts A and B's decisions are exact integer/rational; Part C's float LP digits may differ in the last places on other solver builds. `certs966.json` is rewritten; compare its sha (23568f2f…) on the same versions.\n3. Do not rerun `mkperms966.py` for replay: it draws fresh os.urandom bytes. The published `perms966.json` (8bdd20ba…) contains the 31 permutations and every accepted/rejected byte draw.\nCost: under 1 CPU minute total; RAM < 200 MB.","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":"high","also_fix":null,"transcript_omitted":{"share":0.04081632653061224,"omitted":2,"outputs":49},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":"2026-09-26T16:24:02.178Z","file_notes":null,"research":{"outcome":"result","route_id":5,"depends_on":[357,370],"evidence_md":"At the frozen N66 reference (p = 97, a = 9409, slots 9419..13721, Q = 101..193; literal survivor list re-derived, equal to #357's first 66; F1 = -1), both alignment designs put the arithmetic identity inside its control ensemble, not in a tail. (A) Route revision-3 step, #370's integer witness w fixed, exact integers: Delta_0 = 54302/499983 = 0.108608 equals the median of 2000 seeded single-block two-slot transpositions (512 below / 1412 tied / 76 above; 5% quantile 0.098075; +0.42 sd) and of the exhaustive 40755 (10151/29145/1459; +0.41 sd). The pre-registered failure condition (Delta_0 >= median) holds. Note: that step's success direction (lower tail) is the reverse of #363's (identity above every control), and a w tuned to the identity biases the fixed-w statistic toward the identity. (B) #363's pre-registered design, w re-optimised per matrix by LP with exact rational primal/dual brackets: identity tau* = 0.889244565597 (T = 0.110755, larger than #370's witness margin 0.108608); 31 os.urandom whole-block product permutations give T_j in 0.0817..0.1352, median 0.1075; certified identity better in 18, control >= identity in 13, 0 unresolved; R = 14/32. #363's success conjunction is defeated. Sanity identities (common reverse-row, reverse-phase) transport the witness exactly. (C) Float diagnostic: re-optimising w on (A)'s 2000 matrices puts the identity at the median (-0.04 sd). So the positive weighted margin at N66 is typical of matrices with the same per-prime phase structure; arithmetic alignment across primes adds nothing measurable here. Scope: one prefix; synthetic controls; ranks are conditional diagnostics. The weighted method (route 4) is untouched. check966.py (stdlib) re-derives (B) from the saved rational witnesses.","prior_art_md":"Search updated 2026-09-26, online, before computing. Reused the recorded surveys of #363/#364 and route 5's 14 Sept targeted search (Hemerik & Goeman, TEST 27 (2018), Def. 1/2, Thm 2, sec. 3.4, https://link.springer.com/article/10.1007/s11749-017-0571-1; Nicosia & Latora arXiv:1403.1546v2 secs IV, VII.1; Tukey 1949 one-df non-additivity, additivityTests package, Wang et al. 2015 PMC4459566) rather than repeating them. New queries for the changed ingredients: (1) 'permutation test statistic optimized on observed data selection bias re-optimize each permutation null distribution': results were general permutation-test material (https://en.wikipedia.org/wiki/Permutation_test; permutation null calibration of an upward-biased optimised statistic, arXiv:2608.09863; model-selection significance by permutation, arXiv:2004.07583). They support the design point used here: a statistic tuned on the observed labelling must be re-tuned inside each permutation (Part B/C), or a fixed-parameter null is biased toward the observed labelling (Part A). Inspected at abstract/snippet level only. (2) 'twin prime admissible slots covering residue classes weighted fractional cover certificate Jacobsthal interval sieve linear programming': Jacobsthal-function LP formulations (Hagedorn et al., arXiv:1611.03310), Ford's sieve notes (https://ford126.web.illinois.edu/sieve2023.pdf), Zenodo 'Atlas of Maximal Gaps' (record 22865056) and 'Counting Survivor Sets' (arXiv:2609.08528), titles/snippets only. None measures a cross-prime alignment control for a weighted non-covering certificate at a fixed admissible prefix. Exact gap before this job: the alignment effect at the N66 reference was unmeasured (route text). This job measures it at that one reference. Uncovered: other prefixes, other p, and any mechanism by which the within-block phase structure alone carries the margin. Access gaps: full texts of arXiv:2608.09863 and 2609.08528 not read."},"research_route_id":5,"verification_plan":{"cost":{"ram_gb":1,"disk_gb":1,"minutes":1,"cpu_hours":0.01,"judgment_minutes":10},"claim":"At the N66 reference (first 66 slots of #357's input, Q = 101..193), the LP optimum tau* of min_w sum_q max_b W(q,b) is bracketed within 1.2e-14 at 0.889244565597 (T = 0.110755), and against the 31 saved whole-block product permutations the identity is certified better than 18 controls and no better than 13 (R = 14/32), so #363's pre-registered success (identity above every control) fails.","scope":"The identity and the 31 controls in perms966.json only. Part A (fixed-w transpositions) and Part C (float diagnostic) are in align966.out and are not covered by this checker. Nothing about other prefixes, p or Q.","tools":["python3"],"inputs":["ea81d82b582acc99a8d96411eb1dc28589421cb332ac5609915049cc8bffa10d","8bdd20ba91e1c56555c80c8d6c77f003a827216ac7132638e8c6a11e27a46903","23568f2fbca5ddd3fa2a0e047119bec3b7d00af79089320371a484abe8596854"],"checker":"8956bcc69ebbd6f2ca115605f13c318141f4b9a87feaa421b6d069386b6149d2","command":"python3 check966.py exact-cover758-input.json perms966.json certs966.json","targets":["check966.out"],"coverage":"decisive","expected":"stdout byte-identical to check966.out (sha 5afbd436088a4f2045ee0be93e86fec041771c03e12e17f415045756e4fe52ef); last line 'certified identity better 18, control >= identity 13, unresolved 0; R = 14/32'","manifest":[{"path":"check966.py","role":"checker","sha256":"8956bcc69ebbd6f2ca115605f13c318141f4b9a87feaa421b6d069386b6149d2"},{"path":"exact-cover758-input.json","role":"input","sha256":"ea81d82b582acc99a8d96411eb1dc28589421cb332ac5609915049cc8bffa10d"},{"path":"perms966.json","role":"input","sha256":"8bdd20ba91e1c56555c80c8d6c77f003a827216ac7132638e8c6a11e27a46903"},{"path":"certs966.json","role":"certificate","sha256":"23568f2fbca5ddd3fa2a0e047119bec3b7d00af79089320371a484abe8596854"},{"path":"check966.out","role":"target","sha256":"5afbd436088a4f2045ee0be93e86fec041771c03e12e17f415045756e4fe52ef"}],"supports":"PASS certifies every identity-vs-control comparison in Part B exactly. It does not validate the uniformity of the random draws or say anything about other references.","comparison":"Exact byte equality of stdout.","assumptions":"Upper bounds from the saved rational w (renormalised to sum 1) and lower bounds from the saved rational lambda (each block renormalised to 1) are valid by weak LP duality; the checker rebuilds every incidence from the slot list and permutations.","coverage_md":"All 32 matrices of Part B (identity + 31 controls), every comparison.","environment":"python3 >= 3.9, standard library only.","availability":{"status":"complete","details":"Checker, inputs and certificates are in the manifest; no network.","network":false,"required_sources":[]},"schema_version":1},"verification_fingerprint":"13003653a717e02d9aab01d41e07168ebc117d3e97c16165f281c22dae640b6c","review_admitted_at":"2026-09-26T16:22:26.441Z","department_id":"dept_cc0a0b6ba2bdfadd5f9c50be","run_id":"run_f22306a36b3bc6198a3528ee","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"handle":"Benjaminsen","job_brief":"First update the online prior-work search for this experiment. If existing work covers it, record that and stop; otherwise run this bounded sprint on the uncovered uncertainty. Use cited published numbers during pursuit; their reproduction belongs in later validation. Build on the supplied findings; do not reconstruct earlier research. Return concrete progress and its cheapest credible check, a useful result for review, or a precisely scoped obstacle. Continued investment requires a distinct experiment.\n\nRead GET <project base>/research-routes/5 and return #374. Return the ordinary report and transcript plus research: {route_id: 5, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes\", prior_art_md: \"updated online search record, sources and exact remaining gap\", next_step: <only for continued pursuit>, obstacle: <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.","review_deferred":false,"in_triage":false,"triage":[],"verification_runs":[],"verification_state":{"execution":"not_attempted","conflict":false,"unresolved_conflict":false,"latest_receipt_id":0,"receipt_count":0,"resolution":null},"verification_summary":{"execution":"not_attempted","headline":"No independent execution recorded yet; a check assignment is queued for a worker on another model.","lines":["Claim: At the N66 reference (first 66 slots of #357's input, Q = 101..193), the LP optimum tau* of min_w sum_q max_b W(q,b) is bracketed within 1.2e-14 at 0.889244565597 (T = 0.110755), and against the 31 saved whole-block product permutations the identity is certified better than 18 controls and no bette… (shortened; full text on the return) Scope: The identity and the 31 controls in perms966.json only. Part A (fixed-w transpositions) and Part C (float diagnostic) are in align966.out and are not covered by this checker. Nothing about other pref… (shortened; full text on the return)","Assumptions declared by the author: Upper bounds from the saved rational w (renormalised to sum 1) and lower bounds from the saved rational lambda (each block renormalised to 1) are valid by weak LP duality; the checker rebuilds every incidence from the slot list and permutations.","Why the check supports the claim, as the author argues it: PASS certifies every identity-vs-control comparison in Part B exactly. It does not validate the uniformity of the random draws or say anything about other references.","Coverage declared by the author: decisive for this scope (a claim for review). All 32 matrices of Part B (identity + 31 controls), every comparison.","Awaiting trusted judgment."],"coverage":"decisive","method":null,"controls":{"reported":false,"itemised":false,"detected":null,"total":null,"missed":[]},"receipts":{"total":0,"independent":0,"pass":0,"fail":0,"unable":0,"reused":0,"excluded":0},"pending_check":"queued","unresolved_conflict":false,"latest_receipt_id":null,"basis":{"claim":"At the N66 reference (first 66 slots of #357's input, Q = 101..193), the LP optimum tau* of min_w sum_q max_b W(q,b) is bracketed within 1.2e-14 at 0.889244565597 (T = 0.110755), and against the 31 saved whole-block product permutations the identity is certified better than 18 controls and no better than 13 (R = 14/32), so #363's pre-registered success (identity above every control) fails.","scope":"The identity and the 31 controls in perms966.json only. Part A (fixed-w transpositions) and Part C (float diagnostic) are in align966.out and are not covered by this checker. Nothing about other prefixes, p or Q.","assumptions":"Upper bounds from the saved rational w (renormalised to sum 1) and lower bounds from the saved rational lambda (each block renormalised to 1) are valid by weak LP duality; the checker rebuilds every incidence from the slot list and permutations.","supports":"PASS certifies every identity-vs-control comparison in Part B exactly. It does not validate the uniformity of the random draws or say anything about other references.","coverage_md":"All 32 matrices of Part B (identity + 31 controls), every comparison.","comparison":"Exact byte equality of stdout."},"coverages":[],"caveats":[],"judgment":{"status":"pending","provisional":false,"by":null,"rung":null,"trusted_reviews":0,"advisory_reviews":0,"receipt_id":null,"sufficiency_md":null}},"canonical_return":null,"review_history":[],"dependencies":[{"id":"357","status":"accepted","final_rung":"measured","canonical_return_id":null},{"id":"370","status":"accepted","final_rung":"verified","canonical_return_id":null}],"research_url":"/projects/twin-primes/research-routes/5","transcript_url":"/projects/twin-primes/return/1842/transcript","files":[{"sha256":"0ece83f47d7c0b3ffabeb4d9841b2ce7a64e5204f2e09f5707da21e51a47becb","name":"align966.py","bytes":11256},{"sha256":"ac8a3092b54f5e947ae25fb0318abb74ec14f497fff15720b4386d501c9832af","name":"mkperms966.py","bytes":1178},{"sha256":"8956bcc69ebbd6f2ca115605f13c318141f4b9a87feaa421b6d069386b6149d2","name":"check966.py","bytes":2342},{"sha256":"8bdd20ba91e1c56555c80c8d6c77f003a827216ac7132638e8c6a11e27a46903","name":"perms966.json","bytes":669052},{"sha256":"23568f2fbca5ddd3fa2a0e047119bec3b7d00af79089320371a484abe8596854","name":"certs966.json","bytes":849695},{"sha256":"dc4389127660283da12106ef19b925ad33f9d48331b4fa4580af87ea49fc6af6","name":"align966.out","bytes":7445},{"sha256":"5afbd436088a4f2045ee0be93e86fec041771c03e12e17f415045756e4fe52ef","name":"check966.out","bytes":2954},{"sha256":"8360e4c33019a23136c6500c3ce5c1af7d6ee1b8750e0cc209ce92da2ef5f971","name":"report966.md","bytes":6407},{"sha256":"454daa56a4a46890ff471d05a4dbba88b56eac1c85be39cfe3c34f453c4e7d3e","name":"recipe966.md","bytes":1638},{"sha256":"ea81d82b582acc99a8d96411eb1dc28589421cb332ac5609915049cc8bffa10d","name":"exact-cover758-input.json","bytes":34281}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}