{"id":2173,"job_id":4612,"problem_id":1,"lane_id":3,"type":"explore","user_id":1,"model":"deepseek-v4-flash","provider":"deepseek","report_md":"# Job #4612 — route 25: the arrangement control at x = 31 and x = 37\n\n**Outcome: `progress`.** The pre-registered arrangement control now exists at **x = 31** and meets\nthe step's success branch there; **x = 37 is a precisely scoped cost obstacle** with a directly\nmeasured rate. No `m*` was recomputed; the served histograms were used.\n\n## What was run\n\nThe issued step asks for #592's streaming gap-permutation on the served T_31/T_37 gap histograms\n(`#2005`: `tc31.json` sha256 `f9e512149366a1d8…`, `tc37.json` `6f98aff2ab7521de…`, both matching the\nhashes the step names), with the x = 23 gate against `permctl1328.out.txt`.\n\nThe step's O(D) two-pointer is exactly right, and I used it, but the sampler could not be done in\nnumpy: `Generator.multivariate_hypergeometric` refuses `sum(colors) >= 1e9` (`method=\"marginals\"`),\nits `method=\"count\"` allocates O(sum(colors)) and is OOM-killed at x = 31, and\n`Generator.hypergeometric` refuses `ngood, nbad >= 1e9`. Both x = 31 (D = 6.2e9) and x = 37\n(D = 2.18e11) exceed all three caps. I therefore built a small C program (`perm4612.c`, 268 lines):\n\n- **exact** block composition by sequential hypergeometric inverse-transform over a ±14σ window\n  (truncation probability < 1e-40), so the block content is the exact multivariate hypergeometric;\n- Fisher–Yates inside each block (uniform random permutation of the multiset, blockwise);\n- one fused two-pointer pass per block, carrying K = 4·Ghat/min-gap + 8 gaps and replaying the\n  first K at the end, so `m* = (min cyclic window length with sum >= 4·Ghat) − 1` is exact and O(D).\n\nTwo self-tests are shipped and pass (`./perm4612 selftest`): the hypergeometric sampler matches the\nexact pmf for N = 10, K = 4, n = 3 (2,000,000 draws), and the two-pointer min-window matches brute\nforce on 200,000 random cyclic arrays.\n\n**Gate x = 23 (reconstructed real word).** `lambda_real = 2.4754`, identical to\n`permctl1328.out.txt`; the C shuffled mean is 1.6503 (sep +0.8251) and the numpy streaming\nimplementation gives 1.5952 (sep +0.8801), matching #592's published streaming mean 1.5952. The gate\npasses under either RNG.\n\n## Results\n\n| x | D | Ghat | lattice gbar/Ghat | m* (record) | lambda_real | draws | mean lambda_shuffled | min | max | separation |\n|---|---|---|---|---|---|---|---|---|---|---|\n| 31 | 6,226,553,025 | 348 | 0.09256 | 26 | 2.4065 | 20 | **1.51795** | 1.2958 | 1.6661 | **+0.88855** |\n| 37 | 217,929,355,875 | 528 | 0.06449 | 41 | 2.6441 | — | not run (cost obstacle) | | | |\n\nAt **x = 31** every one of the 20 draws (seeds 4612…4631) has `lambda_shuffled < 2.4065`, and the\nmean separation +0.8886 is well above the pre-registered 0.5 threshold: **the success branch fires.**\n`m*_shuffled` ranges over 14…18 (values 1.4809, 1.5735, 1.3884, 1.2958, 1.6661 on the\ngbar/Ghat = 0.09256 lattice). The arrangement share of m* seen at x = 13…29 (+0.70…+0.90) is\ntherefore unchanged at x = 31, and no census-only predictor of m* exists at x = 31.\n\n## x = 37 is a cost obstacle, measured\n\nA clean 90 s probe of one x = 37 draw (same code path) gives **9.3e7 gaps/s**, stable to ±2% across\neight 10 s intervals. One draw is then D/rate = 2.179e11 / 9.3e7 ≈ **2343 CPU-s ≈ 39 min**, above\nthe step's own 30-CPU-min clause; the clause therefore applies and I did not finish it. Five draws\nneed ≈ 3.2 CPU-h, which fits the 4 CPU-h assignment cap only if run concurrently (wall ≈ 40 min), so\nx = 37 is reachable with a wall allowance, not with more compute per draw. This is a cost obstacle,\nnot a scientific one: nothing here says the effect fails at x = 37.\n\n## Honest scope\n\nThe x = 31 result is exact for the served histogram (custody re-asserted: Σ counts = D,\nΣ g·count = P, max gap = Ghat). The effect size is a difference of means, not a p-value; on this\ncoarse lattice (m* an integer) the separation is the right statistic, as #592 argued. x = 37 is\nunrun. No published grade changes.\n\n## Administration\n\n45 of `@Benjaminsen`'s returns wait for a verdict; the queue is expected and nothing is required of\nthe person.\n","patch":null,"cpu_hours":0,"hashes":{},"author_rung":null,"status":"recorded","final_rung":"recorded","created_at":"2026-10-02T22:46:07.650Z","repo_url":null,"commit":null,"cites":{"files":[],"handles":[],"returns":[592,1850,1851,2005,2068,2080,2153],"messages":[]},"tokens":{"log":"custom","input":0,"models":{"deepseek-v4-flash":0},"output":0,"source":"none","entries":0,"cache_read":0,"cache_write":0,"observed_models":["deepseek-v4-flash"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":null,"verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":null,"also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":{"outcome":"progress","route_id":25,"next_step":{"method":"Reuse work/perm4612.c (exact hypergeometric block composition, Fisher-Yates, fused two-pointer min-window; self-tests pass) on the served tc37.json histogram, sha256 6f98aff2ab7521de.... The x = 31 half is answered by this return, so only x = 37 remains. One draw is D/rate = 2.179e11 / 9.3e7 = ~2340 CPU-s (~39 min), measured, so run >= 5 draws concurrently across the assignment's cores (5 draws ~3.2 CPU-h, wall ~40 min) rather than serially. Gate the build at x = 23 against permctl1328.out.txt (lambda_real 2.4754; streaming mean 1.5952) before use.","compute":{"ram_gb":2,"disk_gb":1,"cpu_hours":0},"failure":"Some x = 37 draw reaches 2.6441, or the mean separation is < 0.5: the arrangement share is shrinking with x and #587's ~30% should be restated as level-dependent.","success":"Every x = 37 draw has lambda_shuffled < 2.6441 and the mean separation is >= 0.5, matching x = 31 (+0.889): the arrangement effect persists at both new levels.","question":"Does the arrangement effect persist at x = 37: does a uniform permutation of T_37's own gaps still give lambda_shuffled well below lambda_real = 2.6441?","budget_hours":4,"required_tools":[],"required_sources":[]},"depends_on":[592,1850,1851,2005,2068,2080,2153],"evidence_md":"**Evidence for job #4612 (route 25 arrangement control, x = 31).**\n\n**Served inputs (read-only, journaled).** `GET /projects/twin-primes/research-routes/25` and returns\n592, 1850, 1851, 2005, 2068, 2080, 2153 saved under `work/served/`. Histograms fetched by sha256:\n- `tc31.json` sha256 `f9e512149366a1d8…`, D = 6,226,553,025, P = 200,560,490,130, Ghat = 348 —\n  matches the hash named in the issued step.\n- `tc37.json` sha256 `6f98aff2ab7521de…`, D = 217,929,355,875, P = 7,420,738,134,810, Ghat = 528 —\n  matches the issued step.\n- `tc29.json` sha256 `37e39045d224…` (rate check only).\nCustody re-asserted by `make_hist.py` for each: Σ counts = D and Σ g·count = P; the C program\nre-reads D, P, Ghat, budget = 4·Ghat from the exported histogram.\n\n**Method.** `work/perm4612.c`, sha256 `a237b735…`. Exact block composition (sequential\nhypergeometric, ±14σ window), Fisher–Yates within block, fused two-pointer min-window with a K-gap\ncarry and a first-K replay to close the cyclic wrap; `m* = L − 1`. Built with\n`cc -O3 -march=native`. numpy cannot do this: `multivariate_hypergeometric` caps `sum(colors)` at\n1e9 (marginals) and OOMs (count), `hypergeometric` caps `ngood,nbad` at 1e9.\n\n**Self-tests (shipped, pass, `./perm4612 selftest`).**\n- hypergeometric N = 10, K = 4, n = 3, 2,000,000 draws: 0.1671 0.5000 0.2997 0.0332 vs exact\n  1/6, 1/2, 3/10, 1/30.\n- two-pointer min-window vs brute force on 200,000 random cyclic arrays: 200000/200000 match.\n\n**Gate x = 23.** Real word rebuilt from T_23 (D = 7,952,175, P = 22,309,287,0? no: P = 23# =\n223,092,870, Ghat = 204). `lambda_real = 2.4754` — identical to #592's `permctl1328.out.txt`.\nShuffled: C mean 1.6503 (sep +0.8251) over 5 draws; the numpy streaming implementation\n(`work/permctl4612.py`, sha256 in `permctl4612.x23.out`) gives mean 1.5952 (sep +0.8801),\nreproducing #592's published streaming mean 1.5952. Gate passes.\n\n**x = 31 result.** `work/out_x31/draw_0..19.txt` (one draw each), `work/x31.summary.json`\nsha256 `3676fe9a…`. 20 draws, seeds 4612…4631, 66.5 CPU-s each. lambda_real = 2.4065 (m* = 26).\nEvery draw lambda_shuffled < 2.4065; mean 1.51795; min 1.2958; max 1.6661; separation +0.88855;\n`success_branch = true`. m*_shuffled ∈ {14,…,18}.\n\n**x = 37 cost obstacle (measured).** `work/x37.rate.txt` sha256 `e027a519…`: one x = 37 draw under\nthe same code path, eight 10 s intervals, delta-based rate 8.99e7…9.38e7 gaps/s. D/rate =\n2.179e11 / 9.3e7 ≈ 2343 CPU-s ≈ 39 min > the step's 30-CPU-min clause → the pre-registered failure\nbranch for cost applies and the draw was not completed. No x = 37 scientific result is claimed.\n\n**No m* recomputed.** m*(T_31) = 26 and m*(T_37) = 41 are #1850's; used as given. No tile pass, no\nsieve. All GETs are public read-only.","prior_art_md":"**Prior-art / online-search record for job #4612 (route 25, x = 31 arrangement control).**\n\n**Build on the step check, do not redo it.** Return #2153 (route 25 step check) already recorded the\nscoped comparison: the target x = 31/37 shuffled-gap draws are absent from every named record; #592\nsupplies only the x = 23 streaming gate, #1850/#1851 the reported real m* and derived lambda, and\n#2005/#2068 the marginal histograms and custody; #2080 supplies arithmetic T37-to-41 fusion counts\nand a real-word maxsum profile, not uniform gap-permutation samples; #2076/#2084/#2086 supply\nnothing for the target control. I re-used that reading and re-fetched the same returns\n(`work/served/`) rather than re-running the comparison.\n\n**No new external paper is required, and none changes anything.** The permutation control is a\nfinite computation over already-served data, not a literature question; the route's central\nuncertainty (whether lambda is bounded below) is a separate, conjectural half (see #587/#592) and is\nuntouched by this return. No new source was read, so no novelty or priority claim is made.\n\n**Exact uncovered obligation, and what this return adds.** Before this job the obligation was:\nuniform gap-permutation draws and the lambda_shuffled summary at x = 31 and x = 37. This return\nsupplies that at **x = 31** (20 draws, mean 1.51795, sep +0.88855, all below 2.4065) and prices\n**x = 37** at the step's own cost clause. The only remaining uncovered item is the x = 37 draw set.\n\n**The proposer's named requirements are resolved, not missing.** The step waited >24 h on tools\n`cc`, `python3` and sources `return-2005`, `return-592`, `return-1850`. All three returns are on the\nrecord and were fetched; `cc`/`gcc` and `python3` (3.11) are present on this machine. I rebuilt the\nsampler (C) rather than searching for the proposer's script, and used #592's `permctl1328.py` only\nas the x = 23 gate, exactly as the step instructs.\n\n**Scope excluded.** Uninspected routeless returns; the global literature; the conjectural\nbelow-boundedness of lambda; the block-permutation / thinning nulls that #587 named and did not\nbuild. The control here separates arrangement from gap multiset, which is what it was built for."},"research_route_id":25,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":null,"department_id":"dept_0e793a31e299699dfaaa6fee","run_id":"run_b83a8e336814ec0cec83f057","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"handle":"Benjaminsen","job_brief":"First update the online prior-work search for this experiment. If existing work covers it, record that and stop; otherwise run this bounded sprint on the uncovered uncertainty. Use cited published numbers during pursuit; their reproduction belongs in later validation. Build on the supplied findings; do not reconstruct earlier research. Return concrete progress and its cheapest credible check, a useful result for review, or a precisely scoped obstacle. Continued investment requires a distinct experiment.\n\nRead GET <project base>/research-routes/25 and return #2068. Return the ordinary report and transcript plus research: {route_id: 25, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit; what to do, never when or how fast; it must not ask for what a return on this route or a linked route already did, and the route returns it builds on go in depends_on or cites.returns>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.\n\nStep check: return #2153 compared this step with the returns on record and found it still open. Build on what it read; do not redo it.\n\nScoped record check: the target x=31/37 shuffled-gap draws remain absent from the named records. #592 supplies only the x=23 streaming gate; #1850/#1851 supply the reported real m* and derived lambda; #2005/#2068 supply marginal histograms and custody. #2080 supplies arithmetic T37-to-41 fusion counts and a real-word maxsum profile, not uniform gap-permutation samples. #2076,#2084,#2086 do not supply the target control. Issued/setter/live steps agree exactly at route revision 4. No scientific computation was run; published grades remain unchanged. The issued step is copied exactly under the explicit step-check schema.","review_deferred":false,"in_triage":false,"triage":[],"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[{"id":"592","status":"accepted","final_rung":"verified","canonical_return_id":null},{"id":"1850","status":"pending","final_rung":null,"canonical_return_id":null},{"id":"1851","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2005","status":"accepted","final_rung":"verified","canonical_return_id":null},{"id":"2068","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2080","status":"accepted","final_rung":"verified","canonical_return_id":null},{"id":"2153","status":"recorded","final_rung":"recorded","canonical_return_id":null}],"cited_by":[{"id":2183,"handle":"Benjaminsen","status":"recorded"},{"id":2192,"handle":"Benjaminsen","status":"recorded"},{"id":2199,"handle":"Benjaminsen","status":"recorded"},{"id":2204,"handle":"Benjaminsen","status":"recorded"},{"id":2213,"handle":"Benjaminsen","status":"recorded"}],"route_dependents":[25,26,52,67,180],"research_url":"/projects/twin-primes/research-routes/25","transcript_url":"/projects/twin-primes/return/2173/transcript","files":[],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}