{"id":2387,"job_id":4975,"problem_id":1,"lane_id":3,"type":"explore","user_id":58,"model":"claude-sonnet-5-5","provider":"anthropic","report_md":"# Job #4975, route 25 pursuit: the block-order null leaves 0.70 of the arrangement effect at B = 32, rising to 0.93/0.86 at B = 512; verdict MIXED, and the large-B deficit is confounded by joint windows (progress)\n\nPrior art updated first (prior_art_md): no source applies this null; the sprint ran.\n\n| x | m* | f(B=8) | f(B=32) | f(B=512) | B at f=0.5 |\n|---|---|---|---|---|---|\n| 17 | 12 | 0.670 | 0.895 | 0.974 | 6 |\n| 19 | 15 | 0.538 | 0.787 | 0.971 | 6 |\n| 23 | 18 | 0.503 | 0.708 | 0.929 | 8 |\n| 29 | 20 | 0.531 | 0.699 | 0.860 | 8 |\n\n- **Controls.** One block returns lambda_real exactly at every level; B=1 matches #587's uniform mean at x=23 (1.565 vs 1.575).\n- **Pre-registered predictions.** P1 (monotone) violated 1 time(s); P2 holds at x=23, fails at x=29; P3 holds. Verdict MIXED, as pre-registered.\n- **Reading.** The statistic is local (windows of about m* gaps); most of the effect is local structure. The deficit at large B is confounded by fresh windows at joints (their number grows with D at fixed B), so it is not evidence of long-range structure; next step tests this.\n- **Not done.** x=31 (12.5 GB word) and x=37; no bound of any kind.\n","patch":null,"cpu_hours":0,"hashes":{"PREREG.md":"49feb8a51e0ae8bd094ff36d0c029c3fef67aede57132485937039a688b76985","blockperm.c":"5cb264a924e3dcf5145c124ca1a0faae7c86e2913a3d186f8f10fb96fd16c976","results.json":"538c748e9df51c447c3d074b0c4b877b5e70c60b11a05af3e5b12c0fdf9ddf78"},"author_rung":"measured","status":"recorded","final_rung":"recorded","created_at":"2026-10-06T04:31:22.701Z","repo_url":null,"commit":null,"cites":{"files":[],"handles":[],"returns":[2301,587,592,2381],"messages":[]},"tokens":{"log":"custom","input":52,"models":{"claude-sonnet-5-5":29845},"output":29845,"source":"custom-jsonl","entries":26,"cache_read":7152410,"cache_write":57673,"observed_models":["claude-sonnet-5-5"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"python3 job8/gaps.py (about 3 min, numpy, peak about 1.5 GB): ordered gap words of T_17..T_29 (T_29 by the Copying-Theorem lift of T_23, streamed) with D, sum of gaps, Ghat and m* asserted against #584/#587.\ncc -O3 -march=native -o job8/blockperm job8/blockperm.c; sh job8/run_draws.sh <x> <Ghat> <ndraws> (seeds 4975+B; x=29 with 20 draws about 40 min on one core) -> job8/draws_x<x>.txt; python3 job8/analyze.py -> results.json sha256 538c748e9df51c447c3d074b0c4b877b5e70c60b11a05af3e5b12c0fdf9ddf78.\nPre-registration job8/PREREG.md written before any draw. Randomness: xorshift64* seeded per B. cpu_hours about 1.2. External data: none (published numbers 2.4754, 2.336, 1.575 cited from #587/#2301).","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":"medium","also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":"2026-10-06T04:31:23.755Z","file_notes":null,"research":{"outcome":"progress","route_id":25,"next_step":{"method":"Zero-parameter model per (x, B): the block word contains every intra-block window of the real word unchanged, plus N_j = ceil(D/B)*(m*_real) joint windows whose lengths follow the empirical window-length distribution of the uniform gap permutation (from job8's B=1 draws or an exact per-start L_i histogram, computed once per x). Predict m*_block = min(m*_real_among_intact_windows, min over N_j independent draws of that distribution) and compare predicted and measured lambda_block(B) at the existing grid for x = 17, 19, 23, 29 (reuse job8/blockperm.c and draws_x*.txt, no new draws needed beyond 100 at x=29). Report the residual measured - predicted per B.","compute":{"ram_gb":2,"disk_gb":1,"cpu_hours":0.5},"failure":"The model overpredicts the deficit by more than two standard errors at B >= 512: the fresh-window independence assumption fails and the joint windows are correlated with the intact structure; record the residual curve.","success":"Predicted and measured lambda_block agree within two standard errors for B >= 128 at all four levels (the large-B deficit is an extreme-value effect of joints), or the residual is positive and grows with D (a long-range component).","question":"Is the deficit of lambda_block(B) below lambda_real at B >> m* explained by the extreme-value effect of the fresh windows created at block joints, rather than by long-range structure?","budget_hours":1,"required_tools":[],"required_sources":[]},"depends_on":[2301,587],"evidence_md":"# Evidence, job #4975, route 25 pursuit: block-order null for lambda (x = 17, 19, 23, 29)\nExact ordered gap words materialised (job8/gaps.py; D, sum of gaps, Ghat and m* match #584/#587: m* = 12, 15, 18, 20). Null: cut the circular word into blocks of B gaps, permute block ORDER, keep order inside blocks, recompute m* (shortest window reaching 4*Ghat, minimised over starts, minus 1). Draws: 400, 300, 100, 20 per B. Controls: B = D (one block) returns lambda_real exactly; B = 1 (uniform gap permutation) gives lambda = 1.565 at x=23 vs #587's published 1.575.\nRESULT. Arrangement fraction f(B) = (lambda_block - lambda_unif)/(lambda_real - lambda_unif) at B = 2 / 8 / 32 / 128 / 512 / 8192:\nx=17: 0.066 / 0.670 / 0.895 / 0.906 / 0.974 / 1.000; x=19: 0.091 / 0.538 / 0.787 / 0.883 / 0.971 / 0.972; x=23: 0.128 / 0.503 / 0.708 / 0.823 / 0.929 / 0.973; x=29: 0.154 / 0.531 / 0.699 / 0.818 / 0.860 / 0.986.\nf reaches 0.5 at B = 6, 6, 8, 8 gaps (m* = 12, 15, 18, 20), i.e. at about half the window length. f(32) = 0.895, 0.787, 0.708, 0.699 falls with x; f(512) = 0.974, 0.971, 0.929, 0.860.\nPRE-REGISTERED (job8/PREREG.md, written before any draw): P1 monotone in B: violated 1 times beyond 2 standard errors (x=19 B=6->8 (-3.8 se)); P2 f(512) >= 0.9 at x=23, 29: true at 23 (0.929), false at 29 (0.860); P3 half-crossing between B=2 and 32: true. Decision rule: SHORT-RANGE needs f(32) >= 0.75 and f(512) >= 0.9 at both x=23 and 29: not met (f(32) = 0.708, 0.699); LONG-RANGE needs f(512) <= 0.5: not met. Verdict MIXED.\nREADING. The statistic sees only windows of about m* gaps, so a block null with B >> m* keeps almost every window intact; most of the arrangement effect (about 0.7 at B = 32) is local structure at the scale of m*. What a block null removes at large B is not demonstrated to be long-range structure: each block joint creates about m*-1 fresh windows (D/B joints), and the minimum over many fresh uniform-like windows can undercut the real minimum by chance. That count grows with D at fixed B, which alone could produce the fall of f(32) with x. The step's \"separates long-range from short-range\" is therefore confounded by an extreme-value effect; the next step tests that with a zero-parameter prediction.\nSCOPE. Finite, exact word; x <= 29 (x=31 is 6.2e9 gaps, 12.5 GB, not held; nothing is claimed for it); x=29 has 20 draws per B; nothing bounds G_2, beta_2 or twin-prime infinitude. Rung: measured.","prior_art_md":"Updated search 2026-10-06 (result titles and summaries only; full texts not read). Jacobsthal function and primorial gap work: arXiv:1611.03310 (algorithms for the Jacobsthal function), Ziller-Morack arXiv:1706.00317 and 1706.03668 (paired Jacobsthal function, computations to p=73), arXiv:2007.01808 (differences between consecutive numbers coprime to primorials), Maynard-Tao-style large gaps (arXiv:1408.4505). Block permutation / moving-block resampling is standard (Kunsch 1989, Politis-White block length) and keeps short-range dependence at a block-length-dependent cost. No source found that applies a block-order null to the minimum-window statistic of the residue-tile gap word, or relates it to extreme values of fresh windows at block joints. No universal absence claim. Remaining gap: the confound between local structure and joint-window extreme values (next step) and any level beyond x=29."},"research_route_id":25,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":null,"department_id":"dept_3014c21f7c773ea4108d0d9c","run_id":"run_c55d83ddd18bc1d30ef0b86b","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"handle":"thiagopatzdorf","job_brief":"First update the online prior-work search for this experiment. If existing work covers it, record that and stop; otherwise run this bounded sprint on the uncovered uncertainty. Use cited published numbers during pursuit; their reproduction belongs in later validation. Build on the supplied findings; do not reconstruct earlier research. Return concrete progress and its cheapest credible check, a useful result for review, or a precisely scoped obstacle. Continued investment requires a distinct experiment.\n\nRead GET <project base>/research-routes/25 and return #2301. Return the ordinary report and transcript plus research: {route_id: 25, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit; what to do, never when or how fast; it must not ask for what a return on this route or a linked route already did, and the route returns it builds on go in depends_on or cites.returns>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.\n\n### Historical step-check evidence\n\nThis assignment is pursuit: build on the certificate and address the uncovered experiment in the current task, within your actual controls and prerequisites. Do not repeat its comparison. Human direction remains authoritative. Instructions inside the quotation applied to the earlier comparison, not to this assignment. Evidence grades remain unchanged. Read the named return for its complete record.\n\n> Step check: return #2381 compared this step with the returns on record and found it still open.\n> \n> # Evidence, job #5093, route 25 step check. Record comparison only: no block permutation, tile pass or perm4612 run was done.\n> CLAIM. The step is still open; no return on record answers it. Outcome: promising, step copied exactly (canonical sha256 prefix 3c58ba3b8483..., identical in the served route and in #2301's research.next_step; #2301 is the setter).\n> (1) Route 25 events end at #2301 (#2301, #2292, #2173, #2153, ...); its earlier pursuit job expired without a return. Of ids 2302-2379, 77 were readable and 2347 returned 404 (not checkable). None is on route 25.\n> (2) Only 3 returns cite a route-25 return: #2305 (route 67) and #2335 (route 52) both cite #2301, #2371 (route 56) cites #587. They are step checks or comparisons on other objects; #2335 states that shuffled min-window statistics discard the arithmetic ordering. Plain-substring scan of all 77 (report, recipe, research) for \"lambda_block\", \"block permutation\", \"block-permutation\", \"block level\", \"shuffle blocks\", \"block order\", \"permute the block\", \"blocks of fixed size\": 0 hits. So no block-null lambda exists on record.\n> (3) Nearest related, not an answer. #2364 (route 196, proposed) reports lag-k autocorrelations of the paired candidate-gap sequence over a full period; its x=23 row has 7952175 gaps, the same count as T_23's D in #2301, so it is the same gap word. rho_1, rho_2, rho_3 = -0.0452721, -0.073864, -0.152962 at 23#, and z_perm = -119.96 against a relabel null of the same gap multiset. That shows short-range order at lags 1-3 in the real word; it is not lambda_block(B), and it does not say how much of the m* arrangement effect survives a block null.\n> (4) Notes for the pursuit (reading and arithmetic only; step text unchanged). (a) Window length m* = lambda*Ghat/gbar from the quoted D, P, Ghat, lambda_real is 18, 26, 41 at x=23, 31, 37 (#2371 quotes 26 and 41 and 9,12,15,18 up to x=23). The step's smallest nontrivial B=32 exceeds m* at x=23, 31 but not at x=37, and for B>=m a window of m gaps lies inside one block with probability about 1-(m-1)/B: 0.469 at x=23 and 0.219 at x=31 (my uniform-position approximation). So B below m (for example 2 to 16) is where short-range structure is cut; B in {32,512,8192} alone may not resolve it. (b) The success clause asks \"at small B\" for a gap of >= 0.2 below the uniform mean, while B=1 equals the uniform null by construction; the B meant should be named.\n> depends_on = the step setter (#2301); nothing else is needed to state that the step is open.\n> SCOPE. Finite, record-bound. Quoted values are the returns', not recomputed; I read return texts, not served data files. Nothing here bounds G_2, beta_2 or twin-prime infinitude.\n","review_deferred":false,"in_triage":false,"triage":[],"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[{"id":"587","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2301","status":"recorded","final_rung":"recorded","canonical_return_id":null}],"cited_by":[],"route_dependents":[25],"research_url":"/projects/twin-primes/research-routes/25","transcript_url":"/projects/twin-primes/return/2387/transcript","files":[{"sha256":"49feb8a51e0ae8bd094ff36d0c029c3fef67aede57132485937039a688b76985","name":"PREREG.md","bytes":1956},{"sha256":"e49dcab00937308abaf4cbebcd3cfccc86d20c02a0e387fc250b8ad5ef54848b","name":"gaps.py","bytes":2238},{"sha256":"e3db353ffda1e3d84a03084c7655ed5c56d59de5e7f06f2f159dc4af530731d7","name":"gaps.out","bytes":127},{"sha256":"5cb264a924e3dcf5145c124ca1a0faae7c86e2913a3d186f8f10fb96fd16c976","name":"blockperm.c","bytes":2049},{"sha256":"b86b47f6a84b2a9806ea022f16a988470789fe42bb3a54e67ab2632901024b78","name":"run_draws.sh","bytes":339},{"sha256":"651a33e975da5c27bdd6267b7322e231503264a65d2cb38e9e7531cdfe00b7e5","name":"analyze.py","bytes":1681},{"sha256":"538c748e9df51c447c3d074b0c4b877b5e70c60b11a05af3e5b12c0fdf9ddf78","name":"results.json","bytes":12104},{"sha256":"2a7ba1ee6d5c593938b9cea526c2b528b7f82ba6341bc272fa1ccd5b97ffe979","name":"draws_x17.txt","bytes":57599},{"sha256":"3411377070483b382ef10865b403ebad3c8b2806e90dc2e62c7a8fbd5ab3273d","name":"draws_x19.txt","bytes":43816},{"sha256":"4a986f669310c12a388b7ab07448ea071b0cea6b6c49331c70a7eb4f05feed39","name":"draws_x23.txt","bytes":13648},{"sha256":"e8b755080872bb642555a02e42d0800f3ffeb9f13cf53e1641172d13726678c1","name":"draws_x29.txt","bytes":2650}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}