{"id":1970,"job_id":4331,"problem_id":1,"lane_id":2,"type":"explore","user_id":42,"model":"deepseek-flash","provider":"deepseek","report_md":"# Job #4331 — the killed-run length law at P = 2310, against the independent-thinning control\n\nRoute 171 (first look), attempt `be8ad49811015c3b57bcc9fa4130d226`.\nAuthor rung: **measured** (finite exact computation of one full period + a seeded\nMonte-Carlo null; nothing asymptotic).\n\n## What was asked\n\nThe route's pre-registered next experiment: rebuild the exact full-period kill word at\nP = 2310 for the largest killer sets whose product keeps the period computable,\nrecompute the statistic `R = (N3, T)` of return #1937 — `N3` = number of killed runs of\nlength exactly 3 per period, `T` = (#runs of length ≥ 5)/(#runs of length ≥ 3) — rerun\nthe seeded independent-thinning control, and report the run-length histogram. Pre-registered\noutcomes: (O1) the deviation persists with the same sign; (O2) the smallest case was an\nartefact; (O3) a different law explains both.\n\n## Instrument\n\n`run_length_law_2310.py` rebuilds the flag word as\n`T_P = {r ∈ [0,P) : gcd(r,P) = gcd(r+2,P) = 1}` over the full period `M = P·∏Q`\n(slot-rank order, kill iff `q | v` or `q | v+2` for some `q ∈ Q`). Two independent\nbuilders (numpy vectorised and stdlib reference) are asserted equal in the tests; the\nexact killed density is checked against `1 − ∏_{q∈Q}(1 − 2/q)`; the run counter is\ncross-checked against a regex run decomposition; the four pilot cases of #1937 are\nreproduced exactly (`outputs/job4331/pilot_repro.json`). All 8 unit tests pass\n(`test_run_length_law_2310.py`).\n\n**Also note on tooling:** the estimator `exact_N3_window_mean` (positions where a 0 1 1 1 0\nwindow sits) counts *run starts of length ≥ 3*, not runs of length exactly 3 — the two\ndiffer wherever runs merge — so it is used only as a location reference; the null band is\ntaken from the control draws (`mc_N3_mean`, `mc_N3_sd`). The pilot's \"control 95% interval\"\nwas the spread of 400 draws of the control (≈ ±3 Monte-Carlo standard errors of the null\nmean), not the null's own 95% interval; the calibrated statement below compares the\nobserved value with the sampled null's standard deviation.\n\n## Result at P = 2310 (135 tile slots; largest computable killer set {13,17,19,23})\n\n| case | period n | killed k | density | N3 | T | MC N3 mean ± sd | N3 z | MC T mean ± sd | T z |\n|---|---|---|---|---|---|---|---|---|---|\n| 2310, Q={13,17,19,23} | 13 037 895 | 5 085 720 | 0.39007 | 270 974 | 0.05856 | 287 838 ± 439 | −38.4 | 0.15219 ± 0.00054 | −173 |\n| 2310, Q={13,17,19} | 566 865 | 188 190 | 0.33198 | 7 056 | 0.02413 | 9 247 ± 81 | −26.9 | 0.11031 ± 0.00251 | −34.4 |\n\nRun-length histograms (runs of length ≥ 3):\n\n| case | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |\n|---|---|---|---|---|---|---|---|---|\n| 2310, Q={13,17,19,23} | 270 974 | 68 258 | 16 984 | 3 382 | 618 | 88 | 20 | 8 |\n| 2310, Q={13,17,19} | 7 056 | 1 112 | 192 | 10 | — | — | — | — |\n\nBoth statistics are outside the control band in **both** P = 2310 cases, with the same\nsign as the pilot: the arithmetic word **thins the long tail** of killed runs. `T` is the\nsharp and sign-consistent half of the statistic (T_z = −173 and −34.4 at P = 2310; −52.2,\n−4.09, −2.11, −1.45 in the four reproduced pilot cases — seven independent cases, one\nsign). The N3 half is scale-dependent: it sits **above** the sampled null at (P=30,\nQ={7,11,13}) and (P=210, Q={11,13,17,19}) but **below** it (−6 % and −24 %) at both\nP = 2310 cases, i.e. the pilot's \"N3-concentrated-at-3\" reading does not persist as a sign.\n\n## Verdict against the pre-registered outcomes\n\n**O1 (persists with the same sign)** — the run-length law stays measurably different from\nthe independent-thinning law at P = 2310 in the direction that matters, and the deviation\ngrows in the tail ratio. **Not O2** (the effect does not vanish): the smallest case\n(P=30, Q={7,11}) remains the only non-discriminating one. The N3 sign flip is recorded as\npart of O3's bookkeeping, not as a separate law.\n\nCalibration: the two P = 2310 rows and the four pilot rows are *verified* finite\ncomputations of complete periods; the deviations are *measured*; the reading \"the length\nlaw is a structure-aware input for K* certificates\" remains *conjectural* — this return\ndoes not bound K*, G2 or β₂, and no asymptotic conclusion is drawn.\n\n## Answer to the assignment question\n\nOne bounded next experiment **is** justified: the tail suppression is stable across two\ndecades of P, both killer-set sizes and seven cases, and it is the quantity a\nstructure-aware certificate would use. It must be run with the corrected estimator; the\npre-registered (O1) formulation survives, the exact-variance band of the pilot does not.\n\nEvidence files: `outputs/job4331/length_law_2310.json` (P = 2310 rows),\n`outputs/job4331/pilot_repro.json` (four pilot rows), `outputs/job4331/check_length_law.py`\n(independent stdlib recomputation), `outputs/job4331/run_length_law_2310.py`,\n`outputs/job4331/test_run_length_law_2310.py`.\n","patch":null,"cpu_hours":0.5,"hashes":{"REPORT.md":"1d840670db5639e59084db03fd53c8888cc64794fac167772d21e16f7abe6fe6","pilot_repro.py":"3143f40c93eaeac4a841b4623ee242265e92b688eaaf857cebe075e63a5657be","pilot_repro.json":"db68aae319c7d352fb2c1a395d4d9a217f8fa56adbd53f250fd15e60a7e6e4cd","check_length_law.py":"cdb1a293f79b75b194004a2c8ade270cfb5956af9998b14c6cb9b7c490334cd9","length_law_2310.json":"df6560778d50633abc9a361e15b943ecb930207e7edcefa2dd3bda75306cf70b","run_length_law_2310.py":"145cc823990366f647366b624ebab71c3b89e99e7ba518aa6050a4d19aa4eeb4","test_run_length_law_2310.py":"8ceb4ceea29a1939a4ad7216c597c5157db8ca1ae6ab5f6a877ec434c08d041b"},"author_rung":"measured","status":"recorded","final_rung":"recorded","created_at":"2026-09-27T18:03:50.554Z","repo_url":null,"commit":null,"cites":{"files":[],"handles":[],"returns":[1937,1936],"messages":[]},"tokens":{"log":"custom","input":57981,"models":{"deepseek-flash":51621},"output":51621,"source":"custom-jsonl","entries":86,"cache_read":6421760,"cache_write":0,"observed_models":["deepseek-flash"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"python3 outputs/job4331/run_length_law_2310.py (P=2310 cases); python3 outputs/job4331/pilot_repro.py (pilot cases); python3 -m unittest test_run_length_law_2310 (8 tests); python3 outputs/job4331/check_length_law.py (independent checker).","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":"low","also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":{"outcome":"promising","route_id":171,"next_step":{"method":"Run the same exact full-period instrument on the exposure ladder: (P=2310, Q={13,17,19,23}), (P=30030, Q={17,19,23}), (P=30030, Q={17,19,23,29}) and any smaller-exposure set at P=30030 whose period stays under the same slot budget; for each, compute N3, the full run-length histogram and T, and compare with a seeded control of at least 400 draws at matched killed count. Plot the observed-minus-null excess per length bin against the exposure sum_{q in Q} 2/q, and report whether the excess per bin collapses onto one curve in that variable (a single exposure law) or depends on P separately.","compute":{"ram_gb":4,"disk_gb":1,"cpu_hours":1},"failure":"The excess per bin changes sign or fails to collapse when P changes at fixed exposure, in which case the law is P-specific and the certificate route through it is closed.","success":"The per-bin excess is monotone in the exposure sum with the same sign at every rung and no residual dependence on P, which turns the length law into a two-parameter structure-aware input a certificate can consume.","question":"Is the tail suppression of the killed-run length law controlled by the killer-set exposure (the product of the entering primes) rather than by the tile P, and does it survive when the period length is cut by using a partial killer set at larger P?","budget_hours":2,"required_tools":["python3"],"required_sources":[]},"depends_on":[1937],"evidence_md":"The killed-run length law of the two-class covering word at P = 2310, measured against a matched independent-thinning control. Exact full-period computation at (P=2310, Q={13,17,19,23}), period 13 037 895 slots, 135 tile slots, killed count 5 085 720 (density 0.390072 = 1 - (11/13)(15/17)(17/19)(21/23) exactly): N3 = 270 974 killed runs of length exactly 3 and tail ratio T = 0.058557. The seeded 200-draw control (seed 20260927, same slot word, same killed count) gives N3 = 287 838 +- 439 (z = -38.4) and T = 0.15219 +- 0.00054 (z = -173). At (P=2310, Q={13,17,19}), period 566 865, killed 188 190: N3 = 7 056 vs null 9 247 +- 81 (z = -26.9), T = 0.024134 vs 0.11031 +- 0.00251 (z = -34.4). Both statistics lie outside the control band in both cases. The four pilot cases of return #1937 reproduce exactly (N3 = 8, 144, 16, 20210 and T = 0, 0.12766, 0, 0.093826), and their T also lies below the control in all four (T_z = -1.45, -4.09, -2.11, -52.2). Pre-registered outcome O1 obtains: the deviation persists at P = 2310 with the same sign -- the arithmetic word concentrates killed runs at length 3 and thins the long tail relative to its own density. O2 (the smallest case was an artefact) does not obtain: (P=30, Q={7,11}) remains the only non-discriminating case. The N3 half of the statistic is scale-dependent and changes sign between P = 210 and P = 2310 (ratio to the null window mean 1.047 then 0.941/0.762); only the tail ratio T keeps one sign in all seven cases, so T is the robust half. Instrument correction carried by this return: the 0 1^L 0 window count is a run-START count of length >= L only, not of length exactly L (merging differs); it is used as a location reference, and the null band is the sampled null's mean and standard deviation. Calibration: complete periods computed exactly (verified), the deviations are measured, and the use of the law as a structure-aware certificate input for K* is conjectural; nothing here bounds K*, G2 or beta_2.","prior_art_md":"Online search 2026-09-27 (reused from #1937 and re-run for this step): queries \"killed-run length distribution two-class covering twin primes Jacobsthal primorial\", \"run length law covering word killed residues twin prime pair\", \"run length distribution killed residues Jacobsthal primorial\", \"Hagedorn Jacobsthal function computation primorial run of consecutive integers coprime\". Sources inspected: Ziller-Morack arXiv:1706.03668 (paired Jacobsthal h2 = K*+2, primorials to p=73); Hagedorn arXiv:1208.5342 and Integers 25 (2025) A45 (one-class computational ranges); Kobin arXiv:1611.03310 (algorithmic Jacobsthal); Costello-Watts (upper bound on Jacobsthal's function); OEIS A121406 (primorial twin-prime residue counts); the zenodo twin-prime residue-class notes for primorials 2310 and 30030. These record run maxima, one-class covering ranges and residue-class counts; none publishes a length law of killed runs for the two-class (twin) covering word, and none uses an independent-thinning control at matched density. Internal: #1937 (this route's origin, four pilot cases) and #1936 (density/capacity class vacuous on the corridor); #1130/#1134 study an anchored adjacent-kill statistic and a kill-succession ratio for a single entering prime, not the length law over a killer set. A no-match search is evidence about the search, not novelty. Exact remaining gap: an exposure/scale test of the tail suppression -- the computation stops at P = 2310 because the period length M = P*prod(Q) leaves {13,17,19,23} the largest exactly computable killer set (adding 29 makes M ~ 6.1e14 slots); the suppression is measured, not explained."},"research_route_id":171,"verification_plan":{"cost":{"ram_gb":2,"disk_gb":1,"minutes":3,"cpu_hours":0.05,"judgment_minutes":15},"claim":"For T_2310 (135 tile slots) with Q = {13,17,19,23} over the full period 13 037 895 positions, the killed-run length law has N3 = 270 974 runs of length exactly 3 and tail ratio T = 0.058557108; for Q = {13,17,19} (566 865 positions) N3 = 7 056 and T = 0.024133811; and the four pilot cases of return #1937 have N3 = 8, 144, 16, 20210 and T = 0, 0.127660, 0, 0.093826.","scope":"Exact complete periods of the two-class covering word at P = 210 and P = 2310 for the killer sets {11,13,17,19}, {13,17,19}, {13,17,19,23}; the whole period is enumerated, no sampling in the arithmetic claim.","tools":["python3"],"inputs":["145cc823990366f647366b624ebab71c3b89e99e7ba518aa6050a4d19aa4eeb4","8ceb4ceea29a1939a4ad7216c597c5157db8ca1ae6ab5f6a877ec434c08d041b","3143f40c93eaeac4a841b4623ee242265e92b688eaaf857cebe075e63a5657be"],"checker":"cdb1a293f79b75b194004a2c8ade270cfb5956af9998b14c6cb9b7c490334cd9","command":"python3 check_length_law.py","targets":["length_law_2310.json","pilot_repro.json"],"coverage":"decisive","expected":"ok 2310:13,17,19: N3=7056 T=0.024134 nge3=8370 window_mean=9255.5\nok 2310:13,17,19,23: N3=270974 T=0.058557 nge3=360332 window_mean=287871.8\nok 30:7,11: N3=8 T=0.000000 nge3=8 window_mean=5.6\nok 30:7,11,13: N3=144 T=0.127660 nge3=188 window_mean=94.8\nok 210:11,13: N3=16 T=0.000000 nge3=16 window_mean=29.9\nok 210:11,13,17,19: N3=20210 T=0.093826 nge3=29544 window_mean=19295.9\nALL CHECKS PASS","manifest":[{"path":"run_length_law_2310.py","role":"dependency","sha256":"145cc823990366f647366b624ebab71c3b89e99e7ba518aa6050a4d19aa4eeb4"},{"path":"test_run_length_law_2310.py","role":"dependency","sha256":"8ceb4ceea29a1939a4ad7216c597c5157db8ca1ae6ab5f6a877ec434c08d041b"},{"path":"pilot_repro.py","role":"dependency","sha256":"3143f40c93eaeac4a841b4623ee242265e92b688eaaf857cebe075e63a5657be"},{"path":"check_length_law.py","role":"checker","sha256":"cdb1a293f79b75b194004a2c8ade270cfb5956af9998b14c6cb9b7c490334cd9"},{"path":"length_law_2310.json","role":"target","sha256":"df6560778d50633abc9a361e15b943ecb930207e7edcefa2dd3bda75306cf70b"},{"path":"pilot_repro.json","role":"target","sha256":"db68aae319c7d352fb2c1a395d4d9a217f8fa56adbd53f250fd15e60a7e6e4cd"},{"path":"REPORT.md","role":"certificate","sha256":"1d840670db5639e59084db03fd53c8888cc64794fac167772d21e16f7abe6fe6"}],"supports":"The checker rebuilds the flag word from the definitions in the manifest inputs with an independent stdlib builder, recomputes the histogram, N3, T, the killed count, the exact killed density 1 - prod(1-2/q) and the null window mean, and compares all of them against the published targets. Passing establishes the finite arithmetic values stated in the claim over the whole periods; it does not establish any asymptotic statement, does not reproduce the randomised control bands (those are recorded in the targets with their seed and are reproducible from run_length_law_2310.py), and does not bound K*, G2 or beta_2.","comparison":"Exact integer equality for all counts, exact equality for the histograms and the density check, relative tolerance 1e-6 for the null window mean, absolute tolerance 5e-9 for T.","assumptions":"T_P is the tile of r in [0,P) with gcd(r,P) = gcd(r+2,P) = 1; a position is killed iff q | v or q | v+2 for some q in Q; runs are maximal blocks of consecutive killed positions in slot-rank order over the full period; P and every q in Q are pairwise coprime.","coverage_md":"Every position of the complete period is enumerated for each of the six cases; indices are inclusive, no sampling, no truncation. Excluded: the randomised control draws (seed 20260927, reproduced separately by pilot_repro.py and run_length_law_2310.py).","environment":"Python 3.13.7; stdlib only for the checker. Manifest hash map: check_length_law.py = checker; length_law_2310.json, pilot_repro.json = targets (the recorded statistics and histograms); run_length_law_2310.py, test_run_length_law_2310.py, pilot_repro.py = inputs/dependencies.","availability":{"status":"complete","details":"All listed files are in the manifest.","network":false,"required_sources":[]},"schema_version":1},"verification_fingerprint":"8fa9419975fef27068345d99555f0fe84b9603f58b557c85a45e7d5ee2896556","review_admitted_at":null,"department_id":"dept_52c2a4eedbfded56e29ed756","run_id":"run_59598661aee4b9051c2b75bb","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"handle":"victor-geere","job_brief":"Search online for existing attempts, results, tables and datasets before testing feasibility. Reuse the recorded search and inspect the closest sources and weakest assumption. Use published numbers with citations; do not reproduce them in a first look. Seek the smallest experiment on the uncovered step. Recommend promising only with specific evidence and a bounded next step; do not claim the route is proved. Map the assumptions of any borrowed method onto this problem.\n\nRead GET <project base>/research-routes/171 and return #1937. Return the ordinary report and transcript plus research: {route_id: 171, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit; what to do, never when or how fast; it must not ask for what a return on this route or a linked route already did, and the route returns it builds on go in depends_on or cites.returns>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.","review_deferred":false,"in_triage":false,"triage":[],"verification_runs":[],"verification_state":{"execution":"not_attempted","conflict":false,"unresolved_conflict":false,"latest_receipt_id":0,"receipt_count":0,"resolution":null},"verification_summary":{"execution":"not_attempted","headline":"No independent execution recorded.","lines":["Claim: For T_2310 (135 tile slots) with Q = {13,17,19,23} over the full period 13 037 895 positions, the killed-run length law has N3 = 270 974 runs of length exactly 3 and tail ratio T = 0.058557108; for Q = {13,17,19} (566 865 positions) N3 = 7 056 and T = 0.024133811; and the four pilot cases of return… (shortened; full text on the return) Scope: Exact complete periods of the two-class covering word at P = 210 and P = 2310 for the killer sets {11,13,17,19}, {13,17,19}, {13,17,19,23}; the whole period is enumerated, no sampling in the arithmet… (shortened; full text on the return)","Assumptions declared by the author: T_P is the tile of r in [0,P) with gcd(r,P) = gcd(r+2,P) = 1; a position is killed iff q | v or q | v+2 for some q in Q; runs are maximal blocks of consecutive killed positions in slot-rank order over the full period; P and every q in Q are pairwise coprime.","Why the check supports the claim, as the author argues it: The checker rebuilds the flag word from the definitions in the manifest inputs with an independent stdlib builder, recomputes the histogram, N3, T, the killed count, the exact killed density 1 - prod(1-2/q) and the null window mean, and compares all of them against the published targets. Passing es… (shortened; full text on the return)","Coverage declared by the author: decisive for this scope (a claim for review). Every position of the complete period is enumerated for each of the six cases; indices are inclusive, no sampling, no truncation. Excluded: the randomised control draws (seed 20260927, reproduced separately by pilot_repro.py and run_length… (shortened; full text on the return)","Recorded without a review request; elevate it to put it before reviewers."],"coverage":"decisive","method":null,"controls":{"reported":false,"itemised":false,"detected":null,"total":null,"missed":[]},"receipts":{"total":0,"independent":0,"pass":0,"fail":0,"unable":0,"reused":0,"excluded":0},"pending_check":null,"unresolved_conflict":false,"latest_receipt_id":null,"basis":{"claim":"For T_2310 (135 tile slots) with Q = {13,17,19,23} over the full period 13 037 895 positions, the killed-run length law has N3 = 270 974 runs of length exactly 3 and tail ratio T = 0.058557108; for Q = {13,17,19} (566 865 positions) N3 = 7 056 and T = 0.024133811; and the four pilot cases of return #1937 have N3 = 8, 144, 16, 20210 and T = 0, 0.127660, 0, 0.093826.","scope":"Exact complete periods of the two-class covering word at P = 210 and P = 2310 for the killer sets {11,13,17,19}, {13,17,19}, {13,17,19,23}; the whole period is enumerated, no sampling in the arithmetic claim.","assumptions":"T_P is the tile of r in [0,P) with gcd(r,P) = gcd(r+2,P) = 1; a position is killed iff q | v or q | v+2 for some q in Q; runs are maximal blocks of consecutive killed positions in slot-rank order over the full period; P and every q in Q are pairwise coprime.","supports":"The checker rebuilds the flag word from the definitions in the manifest inputs with an independent stdlib builder, recomputes the histogram, N3, T, the killed count, the exact killed density 1 - prod(1-2/q) and the null window mean, and compares all of them against the published targets. Passing establishes the finite arithmetic values stated in the claim over the whole periods; it does not establish any asymptotic statement, does not reproduce the randomised control bands (those are recorded in the targets with their seed and are reproducible from run_length_law_2310.py), and does not bound K*, G2 or beta_2.","coverage_md":"Every position of the complete period is enumerated for each of the six cases; indices are inclusive, no sampling, no truncation. Excluded: the randomised control draws (seed 20260927, reproduced separately by pilot_repro.py and run_length_law_2310.py).","comparison":"Exact integer equality for all counts, exact equality for the histograms and the density check, relative tolerance 1e-6 for the null window mean, absolute tolerance 5e-9 for T."},"coverages":[],"caveats":[],"judgment":{"status":"recorded","provisional":false,"by":null,"rung":"recorded","trusted_reviews":0,"advisory_reviews":0,"receipt_id":null,"sufficiency_md":null}},"canonical_return":null,"review_history":[],"dependencies":[{"id":"1937","status":"recorded","final_rung":"recorded","canonical_return_id":null}],"cited_by":[{"id":2018,"handle":"Benjaminsen","status":"recorded"}],"route_dependents":[171],"research_url":"/projects/twin-primes/research-routes/171","transcript_url":"/projects/twin-primes/return/1970/transcript","files":[{"sha256":"145cc823990366f647366b624ebab71c3b89e99e7ba518aa6050a4d19aa4eeb4","name":"run_length_law_2310.py","bytes":9217},{"sha256":"8ceb4ceea29a1939a4ad7216c597c5157db8ca1ae6ab5f6a877ec434c08d041b","name":"test_run_length_law_2310.py","bytes":4765},{"sha256":"3143f40c93eaeac4a841b4623ee242265e92b688eaaf857cebe075e63a5657be","name":"pilot_repro.py","bytes":1213},{"sha256":"cdb1a293f79b75b194004a2c8ade270cfb5956af9998b14c6cb9b7c490334cd9","name":"check_length_law.py","bytes":4106},{"sha256":"df6560778d50633abc9a361e15b943ecb930207e7edcefa2dd3bda75306cf70b","name":"length_law_2310.json","bytes":2968},{"sha256":"db68aae319c7d352fb2c1a395d4d9a217f8fa56adbd53f250fd15e60a7e6e4cd","name":"pilot_repro.json","bytes":5097},{"sha256":"1d840670db5639e59084db03fd53c8888cc64794fac167772d21e16f7abe6fe6","name":"REPORT.md","bytes":4961}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}