{"id":2683,"job_id":5334,"problem_id":1,"lane_id":4,"type":"explore","user_id":1,"model":"deepseek-v4-flash","provider":"deepseek","report_md":"# Report — job #5334 (route 31 pursue): the matched-normaliser step at a raised draw count\n\nAttempt `[id]…`, run `[private]`, session `6e4e8604…`. Device: #2459's matched-normaliser\nthinning (`prereg_co.md`), adopted unchanged; falsifier pre-registered in `PREREGISTRATION_hf.md`\n(**DRAWS = 2000**) and executed only after the anchoring gate passed.\n\n## 1. What the held step asked (route 31, set by #2546, still open per #2612)\n\n#2546 ran the matched-normaliser device over the whole anchored set at `D = 200` and its failure\nclause fired (6 of 9 cell-cases outside the leave-one-out pseudo-observed max, flipping with the\nseed). The step asks for the **same** device at the **same** three anchored cases and the same seeds,\nwith the draw count raised (e.g. `D = 2000-5000`) **or** the existing draws reused under a\nMonte-Carlo-resolution-corrected p-value, the decision pre-registered on the resolved statistic:\n*does the mixed verdict persist once the range is resolved, or does 200 draws merely under-resolve a\nstable ~5 signal?*\n\n## 2. Reconstruction and anchoring gate (passed before any raised-D number)\n\nThe served estimator was **not** re-implemented: `harness.Case` (#2348), `n4_run.perm_within` (#2265),\n`nullnull_5047.pool_like_newstat / leave_one_out` (#2348) and the served producer/statistic/snull\n(#636 / #654 / #2002) are imported byte-identically (SHA-verified, 45/45 served files). Only the draw\nloop count is parameterised.\n\n* **D = 200 reconstruction is exact**: my `d200/` run reproduces #2546's own `results_es.json`\n  **cell-for-cell, worst relative difference 0.000e+00** over all 9 cell-cases × 4 streams.\n* **#2348's six published seed-4164 `|z_level|`** reproduce to `< 5e-3` (measured 3.005764 /\n  4.791444 / 5.163895 / 5.512534 / 3.313359 / 5.994103; rel. diff ≤ 2.4e-14).\n* The gate is a **fixed-draw-count** check. At the raised count the same-named `own` reading moves\n  (e.g. 2^16 N4u `3.0058 → 3.1332`, rel 4.2%), because `own` is computed against the `D`-draw cloud's\n  mean/sd/PC1 frame. This is disclosed as Preregistration Amendment A; the branch and case set are\n  unchanged.\n\n## 3. Result at D = 2000: the step's **failure** clause fires\n\nStatistic = the matched thinning cloud `T_match`; branch = `|z_m| > A_null` (PLACEMENT-FREE) vs\n`|z_m| <= A_null` (NORMALISER-BOUND), `A_null` = leave-one-out pseudo-observed `abs_max`.\n\n| x | seed | z_m (D=2000) | A_null (D=2000) | p_res | rank | branch | D=200 branch |\n|---|---|---|---|---|---|---|---|\n| 2^16 | 4164 | −4.7765 | 5.4772 | 0.00200 | 4 | NORMALISER-BOUND | PLACEMENT-FREE |\n| 2^16 | 4165 | −4.7800 | 4.8987 | 0.00100 | 2 | NORMALISER-BOUND | NORMALISER-BOUND |\n| 2^16 | 4166 | −4.7211 | 5.7870 | 0.00100 | 2 | NORMALISER-BOUND | PLACEMENT-FREE |\n| 2^17 | 4164 | −5.2518 | 6.2356 | 0.00200 | 4 | NORMALISER-BOUND | NORMALISER-BOUND |\n| 2^17 | 4165 | −5.4335 | 6.0036 | 0.00100 | 2 | NORMALISER-BOUND | PLACEMENT-FREE |\n| 2^17 | 4166 | −5.3497 | 6.1741 | 0.00250 | 5 | NORMALISER-BOUND | PLACEMENT-FREE |\n| 2^20 | 4164 | −5.5916 | 5.1658 | 0.00050 | 1 | **PLACEMENT-FREE** | PLACEMENT-FREE |\n| 2^20 | 4165 | −5.1571 | 6.4438 | 0.00300 | 6 | NORMALISER-BOUND | NORMALISER-BOUND |\n| 2^20 | 4166 | −5.3245 | 6.4990 | 0.00150 | 3 | NORMALISER-BOUND | PLACEMENT-FREE |\n\nTally **8 inside / 1 outside** → **not seed-stable**, so the success clause (`all outside` or\n`all inside`) is **not** reached and the step's prescribed failure clause fires: *the\nmatched-normaliser device is not the right instrument for route 31's amplitude clause; the route\nneeds a different statistic rather than more draws.*\n\n## 4. Why — the range is draw-count-governed, not signal-governed (the substantive finding)\n\nAcross the nine cells, `A_null` is a **maximum over `D` draws** and therefore **increases with `D`**,\nwhile `|z_m|` is essentially `D`-stable:\n\n* `A_null`: **3.1841–6.0780** at `D = 200` → **4.8987–6.4990** at `D = 2000` (8 of 9 cells up);\n* `|z_m|`: 4.7982–6.1443 at `D = 200` → **4.7211–5.5916** at `D = 2000` (unchanged to ~4%);\n* the verdict moves from 6 outside / 3 inside to 8 inside / 1 outside **with no change to the\n  mathematics of the case** — only the draw count.\n\nSo the \"resolved range\" does not converge to a stable bar that more draws estimate better: it\n**inflates** with the draw count, and the observed displacement sits just under/over it. The\ncomparison `|z_m| vs max_i|z_i|` is a `1/(D+1)`-level \"is the observed the single most extreme of\n`D+1` exchangeable values\" test, so at `D = 2000` the bar is strictly higher than at `D = 200`.\nThis is exactly the step's failure mode and its stated response: a different statistic, not more\ndraws.\n\n## 5. The step's own alternative — the resolution-corrected p-value — behaves differently\n\n`p_res = (1 + #{|z_i| >= |z_m|})/(1 + D)` (Phipson–Smyth / Knijnenburg), reported (not branched on),\nranges **0.0005–0.0030** (observed ranks 1–6 of 2001; two-sided ≈ 0.001–0.006). Unlike the raw max,\nthis does **not** inflate with `D`: it is a genuine p-value that sharpens as draws are added. It says\nthe matched-thinning level displacement is small but not zero — a partial rejection, not the strict\n\"observed is the most extreme\" the raw branch demands. That is the honest replacement statistic.\n\n## 6. Contrast that fixes the reading\n\n`N4u_match` (within-class placement **preserved**, normaliser matched) is outside the range in\n**9/9 cells at both `D = 200` and `D = 2000`** (`z = −5.41 … −11.92`, `A_null = 2.78 … 4.11`), i.e.\nseed-stable. So the level displacement is robust against the within-class cloud but **not** against\nthe placement-destroying thinning cloud. This corroborates the recorded conclusion that the thinning\nrejection was not placement-specific; it is consistent with the normaliser/range explanation above\nand does **not** by itself establish local anti-correlation of the sign field.\n\n## 7. Normaliser confound (reproduced, larger at D=2000)\n\nThe observed residual `mean_sq` (11.424174 / 17.309923 / 24.899543) lies **below the entire range** of\nthe draws' own `mean_sq` in all three cases (draw ranges 949–11433 / 1306–12601 / 2131–24658; rel.\nsd 289.1 / 203.0 / 265.8). The thinning cloud is not matched to the observed grid on the estimator's\nnormaliser at all without the device; that confound is what the matched device removes.\n\n## 8. Scope and limits\n\nThree finite cases, three seeds, one control family, level statistic only; `D = 2000`, anchors at\nseed 4164. The leave-one-out pseudo-observations are exchangeable with each other **by construction**,\nso a range verdict concerns the **estimator**, not the observed data. No asymptotic, exponent or\n`E_>(x)` claim; nothing about `mu`'s sign correlations, `G2`, `beta_2` or twin-prime infinitude. The\nregistered occupancy-preserving control (#2348) is **not** built; no claim is made that it is\nunnecessary — this run's result says the opposite. Checker `check_hf.py` **85/85 exit 0**;\n`--corrupt` **8 FAIL exit 1**.\n\n## 9. What this changes and the cheapest next experiment\n\nIt closes the held step: at the resolved draw count the verdict is still mixed, so the device cannot\ncarry route 31's amplitude clause, and the fix is a `D`-stable statistic — the resolution-corrected\np-value, pre-registered and pooled across the anchored set — not more draws. That is `next_step`.\n","patch":null,"cpu_hours":1.4,"hashes":{"sah.py":"21a1d3556191bf54458b13fa0ebe41b4550fb92a33ab9bee6518d82ef222c843","fetch.log":"4d2469e3c6692a0b7adc678a9c65423d2ccd9b1a7832137b02a2350dc750f037","n4_run.py":"664f82b737e2ea948d195d0fd2ed5adf37ba68cdce9c67156d354028dec62477","prereg.md":"d3d558fee856b1ea43bd8d849e58e59722f09ad8e7444b0eb24f2b13e8d0c403","recipe.md":"517ef82aa1d36a4a665d1669a9bb14e04c5048e1ed34ad940ada8b047732cf8d","report.md":"830c986275f88c89dbd285091976dac2d1ca0131dc99c9114a71a5c7c0edb329","board.json":"2f050782a32c1bd1de37f68518bb02c18a5c19b6040a8608fdce39c17943294f","driver.out":"b13da910efbcb9f922dd0b8ad4e39afbfa3680ea817057ab81394568d9ef2ef9","harness.py":"66554b5fec573f51ac765794e2c3c8281fe96bb1831827023a28c01a537e71ac","check_hf.py":"5ce44e2e5604426f12eb9e26b054ca50f704bf1c7736508c29030cfedfd5189e","evidence.md":"d570dea20fe1f31de6080841b7d0af75c5eee69f44ba261c5c5a46a409a6c65f","check_hf.out":"12bd4f768cc3a95231e1beb3636d253337972b0cb29e827e27e5704d320b138a","newstat_p.py":"8a6aa937415fc88cc560422f9bfc74beccf668f7dfae41c4bd24722b87075373","prior-art.md":"a796e530b5e3408ee1e1eebc2e0d7cfd3348e786dd9a73d5656a1a5e1f3e0009","run_cases.sh":"4462baae98f52ca25f826008b42f4de2df02ceec853fe8ed069e3cb59624d74f","combine_hf.py":"eb5610fca20965641f8229b85ee49aa5d004645a46c28e8e3a826a72e3fa1104","compute_hf.py":"345880d7805b9242bb5ee57456bc1008c200a30082f25347c17a2b8612d243d0","route-31.json":"c116da57ccaacb3afe3cb2fc99a1aa21de0286cd68abf6f9dfcc27d58621df8b","sieve-null.py":"a53c50266370a779c75d3e2d7b3d93603425ad2d9dcdc68801ea353a2b62f1c8","next-step.json":"fbb9d96f866e10bb37b2b6d9db3fe2cdf8237cd4b8daa8993e653f3bd94beeff","questions.json":"b64e565197938ecf58e11ece6644d713c5b30c4113cd12799a6b85c201a8335d","fetch-files.log":"75990c745f157b94a5a249a43deac7ca784947c8d0134adaa6330f148f4e732e","results_hf.json":"b92614906180f3fca31167b1e85fd7d39654404296a09597c0977a601eeb665e","followup_2681.py":"6f5b3ca2f12cf7e45f0a24a71b017bcec9c7e8ac3b083e529fc81c930b176d33","nullnull_5047.py":"ce9e52a05eba211d06dbda3c64da3994bb28c2945d1ef7be17d4143514699204","return-2348.json":"7062a3e4e8e86611c2923ae51adf7657114c918076f0abe6e153663b0f990c66","return-2459.json":"a73f9327e35ca2bdd987e61c74146d7c2eea685b3ba1601ee0e29e45308809e6","return-2542.json":"af0ed6752db765fd102f7d904988eac8ce72f906e5cec4e3b6626f8de84fdad2","return-2546.json":"f26e752cc9ce3181efd694b0424756b5789979a8e7fd8839ca83d9f0323e5300","return-2612.json":"e9298129302e739d0e4c905fdcfbaadf1076358621072f305076355ef4393e00","fibre-sign-lag.py":"a74825d84e5421eb330d6b54f93029a0aebdc2fd5120fce02ab5d6d857545b56","followup_2681.out":"6af816b67c7c340a050143e780db0564ccfa6ae406b945f6168e90ae606cd0a3","baseline-d200.json":"01f1aa01f641325585686606009a757313f96200d9646c8d14133d970760170b","check_hf.control.out":"94e3615e9917293f5477da8075e7ff32d235efd16338dcfdc756bed2a563c5b1","export_transcript.py":"029efc05e4b791b297f3cb254a24887e3d23b98b1ab4a6639d1f6dc7b69cc82f","research-routes.json":"e15d2f54a51107120cf8fa5197076ea99e85ad690086768b26752b953292060a","research-protocol.json":"bcca2fe0645d5663739db0d0c51c65a50ebc7fd362e243cb591a42faa532059e","job1438-cls-resid-offset.py":"3907b1518b8bc0133b09f1cd2f28a39f3fff62fc5ded2d522c4016b67cbffce4","note-route31-raised-draws.md":"76d2502c17146864536456fd412dea97d1a6d84f6a7dc9a2f4f35121e45c4a6a"},"author_rung":null,"status":"recorded","final_rung":"recorded","created_at":"2026-10-10T06:15:32.050Z","repo_url":null,"commit":null,"cites":{"returns":[2546]},"tokens":{"log":"custom","input":0,"models":{"deepseek-v4-flash":0},"output":0,"source":"none","entries":0,"cache_read":0,"cache_write":0,"observed_models":["deepseek-v4-flash"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"# Recipe — route 31 held step at a raised draw count (job #5334)\n\nReproduce offline with numpy + the served stack (no network; stdlib + numpy only).\n\n## 1. Lay out the served estimator (byte-identical, hash-verified)\n\nAttached served files (fetch each by its served sha; `_` in a served name is a copy target, not a\nsecond stored file):\n\n```\nserved file `n4_run.py`        sha256 664f82b737e2…   -> save as  work/served/n4_run.py\nserved file `newstat_p.py`     sha256 c3b03d2fcacf…   -> save as  work/served/newstat_p.py\nserved file `fibre-sign-lag.py` sha256 a74825d84e54…  -> save as  work/served/f636_fibre-sign-lag.py\nserved file `job1438-cls-resid-offset.py` sha256 3907b1518b8b… -> save as work/served/f654_job1438-cls-resid-offset.py\nserved file `sieve-null.py`    sha256 a53c50266370…   -> save as  work/served/f2002_sieve-null.py\nattached harness.py            -> work/harness.py        (#2348 shared harness: Case, modules)\nattached nullnull_5047.py      -> work/nullnull_5047.py  (#2348 pool_like_newstat / leave_one_out)\n```\n`harness.modules()` hard-codes the `f636_`/`f654_`/`f2002_` names under `work/served/`; the copy step\nabove provides them from the served files (the producer copy must keep sha `a74825d84e54…`, which\n`n4_run.PRODUCER_SHA` checks). `compute_hf.py` then imports the estimator unchanged.\n\n## 2. Anchoring gate (fixed-D check) — must pass first\n\n```bash\n# whole anchored set at the canonical D=200 (the published draw count)\nbash work/run_cases.sh 200 work/d200 \"65536 10,10,1,1\" \"131072 14,14,1,1\" \"1048576 14,14,1,1\"\n# expect: each d200/x<X>_s<S>.json `anchor.<stream>.ok == true`, and result identical to #2546's\n# results_es.json cell-for-cell (worst relative difference 0.000e+00 over the 4 streams x 9 cells)\n```\nThe `own` z_level is computed against the `D`-draw cloud, so this gate is a **fixed-D** check; do not\nrun it at a raised D (Preregistration Amendment A).\n\n## 3. Raised-draw run (the experiment)\n\n```bash\nbash work/run_cases.sh 2000 work/d2000 \"65536 10,10,1,1\" \"131072 14,14,1,1\" \"1048576 14,14,1,1\"\npython3 work/combine_hf.py --raised work/d2000 --low work/d200 --out work/results_hf.json\npython3 work/check_hf.py                 # expect: 85/85 PASS, exit 0\npython3 work/check_hf.py --corrupt       # expect: 8 FAIL, exit 1\n```\n`run_cases.sh` runs the three seeds (4164/4165/4166) in parallel per case; each writes its own\n`x<X>_s<S>.json`. One (case, seed) can be reproduced alone:\n```bash\npython3 work/compute_hf.py --x 1048576 --cfg 14,14,1,1 --draws 2000 --seed 4164 --out out.json\n```\nRuntimes on this machine (wall): 2^16 134 s, 2^17 202 s, 2^20 1288 s per seed; D=200 gate ≈ 110 s.\n\n## 4. Field-by-field anchors\n\n- Observed `mean_sq`: 11.424174460408176 / 17.30992258229879 / 24.899543479159398 (2^16/2^17/2^20).\n- `M/NR/P_U`: 32768/32771/210; 65536/65539/30030; 524288/524291/30030.\n- Branch: `|z_m| > A_null` → PLACEMENT-FREE, else NORMALISER-BOUND; step statistic = `T_match`.\n- `p_res = (1+#{|z_i|>=|z_m|})/(1+D)`; rank_obs = `1+#{|z_i|>=|z_m|}`.\n- D=2000 tally: 8 inside / 1 outside (the one outside is 2^20 seed 4164, rank 1, p_res = 1/2001).\n\n## 5. Not run / not claimed\n\nNo sieve, no new coefficient field, no reconstruction of any earlier research; no asymptotic,\nexponent or `E_>(x)` claim; the registered occupancy-preserving control (#2348) is not built. The\ncanonical-D (200) numbers are a reconstruction of #2546, not new evidence. `cpu_hours ≈ 1.4`\n(measured wall 1624 s × 3 seeds, parallel 3-way on 10 cores).","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":null,"also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":{"outcome":"progress","route_id":31,"next_step":{"method":"Reuse the draws already computed here (D=2000, 3 cases x 3 seeds, matched-normaliser device) and replace the per-cell raw max branch with the resolution-corrected p-value p_res = (1 + #{i: |z_i| >= |z_m|})/(1 + D) of the matched thinning cloud T_match. Pre-register ONE pooled decision statistic over the nine cell-cases (e.g. Fisher combination of the nine two-sided p_res, or the minimum p_res with a pre-registered multiplicity rule) and a single alpha, written before the number is read; report each cell's p_res, the pooled statistic and the anchoring gate from the canonical-D=200 reconstruction first. Do not add draws (the max-based range inflates with D; adding draws does not resolve it) and do not rebuild the registered occupancy-preserving control (#2348).","compute":{"ram_gb":1,"disk_gb":0.1,"cpu_hours":0.2},"failure":"The pooled p_res is above the pre-registered alpha: the matched-thinning displacement is within the estimator's own resolution-corrected range, the thinning rejection was a normaliser/dispersion effect, and the registered occupancy-preserving control (#2348) stays necessary; the route then needs an instrument that is not this estimator at all.","success":"The pooled p_res is below the pre-registered alpha across the anchored set: the matched-thinning level displacement is real and placement-free, so route 31 may drop the occupancy-preserving control for the amplitude clause.","question":"Once the decision statistic is made D-stable (the resolution-corrected p-value, not the leave-one-out maximum), does route 31's amplitude clause read on within-class placement at all — i.e. is the matched-thinning level displacement below a pre-registered alpha across the whole anchored set, or is it indistinguishable from the placement-destroying cloud?","budget_hours":0.2,"required_tools":["python3","numpy"],"required_sources":["served_return_records","served_pipeline_files"]},"depends_on":[2348,2459,2546,2612],"evidence_md":"# Evidence — route 31 held step executed at a raised draw count (job #5334)\n\n**Device (adopted unchanged from #2459's `prereg_co.md`).** Matched-normaliser thinning: each draw's\nresidual grid of `stat.rl` is evaluated with a common `mean_sq` equal to the observed value\n(`match`), compared with the served reading (`own`). Streams: `N4u` (within-class, `seed+9000`) and\n`T` thinning (placement-destroyed, `seed+9002`). Statistic: `z_level` of the orthogonal level\ndirection; branch on `|z_m|` vs the leave-one-out pseudo-observed `abs_max` `A_null`.\n\n**Nothing re-implemented.** Served modules imported byte-identically (SHA-verified 45/45):\n`harness.py`, `nullnull_5047.py`, `n4_run.py` `664f82b7…`, `newstat_p.py` `c3b03d2f…`,\n`fibre-sign-lag.py` `a74825d8…` (#636 producer, matches n4_run's PRODUCER_SHA), `job1438-…py`\n`3907b151…`, `sieve-null.py` `a53c5026…`. Only the draw count is parameterised.\n\n**Anchoring gate (fixed-D check, D=200) — PASSED, and the reconstruction is exact.**\n1. My `D=200` run reproduces #2546's `results_es.json` **cell-for-cell: worst relative difference\n   0.000e+00** over all 9 cell-cases × 4 streams (`z_level_observed`, `loo_abs_max`, `level_null_sd`).\n2. #2348's six published seed-4164 `|z_level|` reproduce to rel. diff ≤ 2.4e-14\n   (3.0057637 / 4.7914445 / 5.1638949 / 5.5125344 / 3.3133590 / 5.9941028).\nDisclosed (Prereg Amendment A): at the raised D the `own` reading moves with the cloud\n(2^16 N4u 3.0058→3.1332, rel 4.2%; 2^20 T 5.9941→7.0626, rel 17.8%), so the raised-D `own` is\nreported but is not the gate.\n\n**Raised-draw result (D=2000, matched thinning `T_match`, the step's statistic).**\n2^16: s4164 z=−4.7765/A=5.4772/p=0.00200(D200: OUT); s4165 −4.7800/4.8987/0.00100(in);\ns4166 −4.7211/5.7870/0.00100(D200 OUT).\n2^17: s4164 −5.2518/6.2356/0.00200(in); s4165 −5.4335/6.0036/0.00100(D200 OUT);\ns4166 −5.3497/6.1741/0.00250(D200 OUT).\n2^20: s4164 −5.5916/5.1658/0.00050(**OUT**, rank 1); s4165 −5.1571/6.4438/0.00300(in);\ns4166 −5.3245/6.4990/0.00150(D200 OUT).\n**Tally 8 inside / 1 outside → not seed-stable → the step's failure clause fires.**\n\n**Mechanism (the finding).** `A_null` is a max over D draws and therefore **grows with D**:\n3.1841–6.0780 (D=200) → 4.8987–6.4990 (D=2000), 8/9 cells up; `|z_m|` is D-stable\n(4.798–6.144 → 4.721–5.592). The verdict is draw-count-governed, not placement- or\nnormaliser-governed. The resolution-corrected p-value\n`p_res = (1+#{|z_i|≥|z_m|})/(1+D)` (recorded, not branched on) is 0.0005–0.0030 (ranks 1–6 of 2001)\nand does **not** inflate with D.\n\n**Contrast.** `N4u_match` is outside the range **9/9 at both D** (z = −5.41…−11.92, A_null\n2.78–4.11) → seed-stable; the level displacement is robust under the within-class cloud but not\nunder thinning.\n\n**Confound (reproduced).** Observed `mean_sq` 11.424174 / 17.309923 / 24.899543 lies below the whole\ndraw range (949–11433 / 1306–12601 / 2131–24658); rel. sd 289.1 / 203.0 / 265.8.\n\n**Controls.** Checker `check_hf.py` 85/85 exit 0 — re-derives every branch, the p_res arithmetic, the\nconfound statement, the D=200 reconstruction, and the served-estimator equivalence\n(`nullnull_5047.pool_like_newstat` vs served `newstat_p.pool_stats_ext`, two code paths, agree to\n<1e-9); `--corrupt` 8 FAIL exit 1. Producer `bounded --limit 3600`: exit 0, `survivors_seen: []`.\n\n**Scope.** 3 finite cases, 3 seeds, one control family, level statistic only, D=2000; pseudo-obs are\nexchangeable by construction so a range verdict concerns the estimator, not the data. No\nasymptotic/exponent/`E_>(x)` claim. Registered occupancy-preserving control (#2348) NOT built.","prior_art_md":"# Prior-art record — route 31 held step, raised draw count (job #5334)\n\n**This run's online search (2026-10-10, before the raised-D numbers were read).** Re-ran the external\nfamily the route's own records name and checked for anything new bearing on a **finite-draw\nleave-one-out maximum** used as a decision range.\n\n* **Monte-Carlo permutation p-value resolution** — unchanged and still the only direct external\n  signal. Phipson & Smyth, \"Permutation P-values should never be zero: calculating exact P-values when\n  permutations are randomly drawn\" (Stat. Appl. Genet. Mol. Biol. 2010; arXiv:1603.05766): report\n  `(1+#{null ≥ obs})/(1+D)`, never 0. Knijnenburg et al., \"Fewer permutations, more accurate\n  P-values\" (Bioinformatics 2009, PMC2687965): the minimal obtainable p-value and its resolution are\n  both coupled to the number of permutations. Both **agree** with the mechanism measured here: a\n  raw `#{null ≥ obs}` of 0 gives a resolution floor `1/(D+1)`, so a raw max-based branch tightens as\n  `D` grows and a corrected p-value is the D-stable substitute.\n* **Restricted / structured permutation schemes and validity** — Besag & Clifford cyclic-shift\n  schemes (`permute`; gavinsimpson.github.io/permute), \"Restricted Block Permutation for Two-Sample\n  Testing\" (arXiv:2512.00668, exact validity for any fixed restricted permutation set). Unchanged; no\n  match to this project's object.\n* **Max-based / adaptive permutation** — search returned only generic multivariate-permutation-null\n  and adaptive-permutation material (partitioning permutations for small p-values, Segal et al.;\n  maximal-null constructions, Chau et al. ISMRM 2004). None states this project's specific object.\n* **Null-of-null / leave-one-out calibration** — search returned only standard permutation-test\n  expositions (Holt 2023; standard textbook null-distribution constructions). No template for an\n  orthogonal **level** statistic whose null range is a leave-one-out maximum.\n\n**Exact remaining gap.** No external work states this project's object: an orthogonal level statistic\nwhose null range is a finite-draw leave-one-out **maximum**, with a branch read against that range.\nThe external family supplies only the vocabulary (resolution, D-coupling, exact validity) for the\nobstruction this run measured directly. **No match found is not a novelty claim**; this is a scoped\nprior-work assessment, not a refutation. Nothing here reproduces an external published number, and no\nexternal template decides the branch.\n\n**Served-return prior art examined (read-only).** #2459 (the device), #2546 (the setter whose step is\nexecuted; its own `results_es.json` is reproduced exactly here), #2542 and #2612 (the two step checks\nthat found the step open), #2348 (the anchor source and the registered occupancy-preserving control),\n#2374/#2265/#2021/#2002 (the served stack's upstream). #2612 scanned return ids 2547..2607 and found\n**no route-31 executor** of this step; this run is the executor, and it adds the raised-D result the\nroute's `next_step` asked for.\n\n**What is genuinely new here (internal, not external).** The measured fact that the leave-one-out\n**max** range `A_null` **increases with D** (3.18–6.08 at D=200 → 4.90–6.50 at D=2000) while `|z_m|`\nbarely moves, so the branch is draw-count-governed — and therefore that the D-stable decision\nstatistic is the resolution-corrected p-value, which the step already named as its alternative. This\nis a property of the estimator, not a claim about the sign field."},"research_route_id":31,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":null,"department_id":"dept_0e793a31e299699dfaaa6fee","run_id":"run_6a66423f347c0f9b0734365d","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"paper_exposition":null,"handle":"Benjaminsen","job_brief":"First update the online prior-work search for this experiment. If existing work covers it, record that and stop; otherwise run this bounded sprint on the uncovered uncertainty. Use cited published numbers during pursuit; their reproduction belongs in later validation. Build on the supplied findings; do not reconstruct earlier research. Return concrete progress and its cheapest credible check, a useful result for review, or a precisely scoped obstacle. Continued investment requires a distinct experiment.\n\nRead GET <project base>/research-routes/31 and return #2546. Return the ordinary report and transcript plus research: {route_id: 31, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit; what to do, never when or how fast; it must not ask for what a return on this route or a linked route already did, and the route returns it builds on go in depends_on or cites.returns>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.\n\n### Historical step-check evidence\n\nThis assignment is pursuit: build on the certificate and address the uncovered experiment in the current task, within your actual controls and prerequisites. Do not repeat its comparison. Human direction remains authoritative. Instructions inside the quotation applied to the earlier comparison, not to this assignment. Evidence grades remain unchanged. Read the named return for its complete record.\n\n> Step check: return #2612 compared this step with the returns on record and found it still open.\n> \n> # Evidence — route 31 step check (job #5435, first look)\n> \n> Record comparison only. All facts re-derived from this run's own served snapshots\n> (`work/served/**`) by the independent offline checker `work/check_gf.py` (26 checks, 0 FAIL, exit 0;\n> `--corrupt` 1 FAIL, exit 1). No experiment, no sieve, no published number recomputed.\n> \n> ## 1. The route record and the held step\n> \n> - `GET /research-routes/31`: `state = active`, `revision = 23`, `origin_return_id = 636`,\n>   `last_return_id = 2546`; event return ids max **2546** (23 events) — no route-31 event after the\n>   setter.\n> - The served `next_step` is **canonically equal to `#2546.research.next_step`** (full-object JSON\n>   equality asserted in both directions). Canonical sha256 =\n>   `320dbb1a9a193613c7b886bd9a9e2583dc2b4ac603006e602351eb2da8c3a51e`. So the held step is the one\n>   #2546 set, and `#2546` is a route-31 return with `outcome = progress`.\n> - The **superseded** step (`#2459`'s, the one `#2542` found open) has canonical sha256\n>   `71bba79c274e58d27006e2242e190e225ee367f691214141255099bb1223519b`; `#2459` and `#2542` publish it\n>   byte-equal, and it is **not** the held step.\n> \n> ## 2. Nothing recorded after #2546 executes the step\n> \n> Full scan of every return with an id in **2547..2607** (the maximum existing id is 2607): 69 probed,\n> 57 carry route metadata.\n> \n> - **No return has `route_id = 31`.**\n> - **No return carries the step's instrument** (`mean_sq`, `grid_of`, `leave-one-out`,\n>   `matched-normaliser`, `pseudo-observed`, `resolved range`, `resolution-corrected`, `seed-stability`).\n> - **The only post-setter return that names #2546 is #2548** (route 227, `promising`, created\n>   2026-10-08T11:24Z). It is a **route-227 step check**, not a route-31 executor:\n>   - `#2548.research.route_id = 227`; its `depends_on` includes `2546` only as one of the three\n>     post-setter returns it inspected (2545/2546/2547);\n>   - its only textual `227` reference is a substring collision — \"a `227` inside a numeric table\n>     (`3.2276` region of the `A_null` table)\";\n>   - its own `next_step` is the route-227 **bridge test** (`M_2k(h)`, gap-law tail), a different\n>     object with different required sources.\n> \n> So the step has **no executor**: neither route 31 nor any linked route runs the matched-normaliser\n> device over the anchored set with a raised draw count or a resolution-corrected p-value.\n> \n> ## 3. What is still required to answer it (the step, unchanged)\n> \n> `#2546` executed the device over the whole set at D = 200 draws and its failure clause fired: 6 of 9\n> cell-cases outside the leave-one-out pseudo-observed max, flipping with the seed. Its set step asks\n> for the **same** device at the **same** three anchored cases (`x = 2^16` beta `U = 10`; `x = 2^17`,\n> `2^20` alpha `U = 14`), the same seeds 4164/4165/4166, with the draw count raised (e.g. D =\n> 2000-5000) **or** the existing draws reused under a Monte-Carlo-resolution-corrected p-value, and the\n> decision pre-registered on that resolved statistic. The anchoring gate against #2348's published\n> seed-4164 `|z_level|` readings must be passed first — the six readings (`3.0058, 4.7914, 5.1639,\n> 5.5125, 3.3134, 5.9941`) are on record in `#2459`.\n> \n> ## 4. Verdict\n> \n> `outcome: promising` — the held step is still open and is copied **exactly** as `next_step`; the\n> pursuit goes back out with this note. Rung: **recorded** (record comparison only; no computation).\n> \n> **Scope and limits.** Three finite cases, 200 draws historically, one control family, level statistic\n> only; no asymptotic, exponent or `E_>(x)` claim; nothing about `G2`, `beta_2` or twin-prime\n> infinitude. The scan covers all return ids 2547..2607; 12 of the 69 fetched ids had no route metadata\n> (types `paper`/`check`/`direction`/`explore` on other routes) and none carried the instrument. An\n> unchanged-step comparison on another route is not new evidence, and none is claimed here.\n","review_deferred":false,"in_triage":false,"triage":[],"lean_statement_binding":null,"lean_execution_binding":null,"lean_scientific_identity":null,"lean_execution_identity":null,"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[{"id":"2348","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2459","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2546","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2612","status":"recorded","final_rung":"recorded","canonical_return_id":null}],"cited_by":[],"route_dependents":[31],"research_url":"/projects/twin-primes/research-routes/31","transcript_url":"/projects/twin-primes/return/2683/transcript","files":[{"sha256":"830c986275f88c89dbd285091976dac2d1ca0131dc99c9114a71a5c7c0edb329","name":"report.md","bytes":7384},{"sha256":"d570dea20fe1f31de6080841b7d0af75c5eee69f44ba261c5c5a46a409a6c65f","name":"evidence.md","bytes":3680},{"sha256":"a796e530b5e3408ee1e1eebc2e0d7cfd3348e786dd9a73d5656a1a5e1f3e0009","name":"prior-art.md","bytes":3532},{"sha256":"517ef82aa1d36a4a665d1669a9bb14e04c5048e1ed34ad940ada8b047732cf8d","name":"recipe.md","bytes":3509},{"sha256":"d3d558fee856b1ea43bd8d849e58e59722f09ad8e7444b0eb24f2b13e8d0c403","name":"prereg.md","bytes":6172},{"sha256":"fbb9d96f866e10bb37b2b6d9db3fe2cdf8237cd4b8daa8993e653f3bd94beeff","name":"next-step.json","bytes":1967},{"sha256":"76d2502c17146864536456fd412dea97d1a6d84f6a7dc9a2f4f35121e45c4a6a","name":"note-route31-raised-draws.md","bytes":3569},{"sha256":"345880d7805b9242bb5ee57456bc1008c200a30082f25347c17a2b8612d243d0","name":"compute_hf.py","bytes":7545},{"sha256":"4462baae98f52ca25f826008b42f4de2df02ceec853fe8ed069e3cb59624d74f","name":"run_cases.sh","bytes":705},{"sha256":"eb5610fca20965641f8229b85ee49aa5d004645a46c28e8e3a826a72e3fa1104","name":"combine_hf.py","bytes":7286},{"sha256":"5ce44e2e5604426f12eb9e26b054ca50f704bf1c7736508c29030cfedfd5189e","name":"check_hf.py","bytes":10515},{"sha256":"12bd4f768cc3a95231e1beb3636d253337972b0cb29e827e27e5704d320b138a","name":"check_hf.out","bytes":11},{"sha256":"94e3615e9917293f5477da8075e7ff32d235efd16338dcfdc756bed2a563c5b1","name":"check_hf.control.out","bytes":608},{"sha256":"b92614906180f3fca31167b1e85fd7d39654404296a09597c0977a601eeb665e","name":"results_hf.json","bytes":39855},{"sha256":"01f1aa01f641325585686606009a757313f96200d9646c8d14133d970760170b","name":"baseline-d200.json","bytes":37235},{"sha256":"66554b5fec573f51ac765794e2c3c8281fe96bb1831827023a28c01a537e71ac","name":"harness.py","bytes":6132},{"sha256":"ce9e52a05eba211d06dbda3c64da3994bb28c2945d1ef7be17d4143514699204","name":"nullnull_5047.py","bytes":10833},{"sha256":"6f5b3ca2f12cf7e45f0a24a71b017bcec9c7e8ac3b083e529fc81c930b176d33","name":"followup_2681.py","bytes":1682},{"sha256":"6af816b67c7c340a050143e780db0564ccfa6ae406b945f6168e90ae606cd0a3","name":"followup_2681.out","bytes":1587},{"sha256":"b13da910efbcb9f922dd0b8ad4e39afbfa3680ea817057ab81394568d9ef2ef9","name":"driver.out","bytes":279},{"sha256":"4d2469e3c6692a0b7adc678a9c65423d2ccd9b1a7832137b02a2350dc750f037","name":"fetch.log","bytes":3676},{"sha256":"75990c745f157b94a5a249a43deac7ca784947c8d0134adaa6330f148f4e732e","name":"fetch-files.log","bytes":2303},{"sha256":"664f82b737e2ea948d195d0fd2ed5adf37ba68cdce9c67156d354028dec62477","name":"n4_run.py","bytes":12502},{"sha256":"8a6aa937415fc88cc560422f9bfc74beccf668f7dfae41c4bd24722b87075373","name":"newstat_p.py","bytes":8733},{"sha256":"a74825d84e5421eb330d6b54f93029a0aebdc2fd5120fce02ab5d6d857545b56","name":"fibre-sign-lag.py","bytes":11979},{"sha256":"3907b1518b8bc0133b09f1cd2f28a39f3fff62fc5ded2d522c4016b67cbffce4","name":"job1438-cls-resid-offset.py","bytes":16697},{"sha256":"a53c50266370a779c75d3e2d7b3d93603425ad2d9dcdc68801ea353a2b62f1c8","name":"sieve-null.py","bytes":13888},{"sha256":"c116da57ccaacb3afe3cb2fc99a1aa21de0286cd68abf6f9dfcc27d58621df8b","name":"route-31.json","bytes":276850},{"sha256":"f26e752cc9ce3181efd694b0424756b5789979a8e7fd8839ca83d9f0323e5300","name":"return-2546.json","bytes":34424},{"sha256":"a73f9327e35ca2bdd987e61c74146d7c2eea685b3ba1601ee0e29e45308809e6","name":"return-2459.json","bytes":29192},{"sha256":"7062a3e4e8e86611c2923ae51adf7657114c918076f0abe6e153663b0f990c66","name":"return-2348.json","bytes":38438},{"sha256":"af0ed6752db765fd102f7d904988eac8ce72f906e5cec4e3b6626f8de84fdad2","name":"return-2542.json","bytes":28351},{"sha256":"e9298129302e739d0e4c905fdcfbaadf1076358621072f305076355ef4393e00","name":"return-2612.json","bytes":25133},{"sha256":"bcca2fe0645d5663739db0d0c51c65a50ebc7fd362e243cb591a42faa532059e","name":"research-protocol.json","bytes":66698},{"sha256":"e15d2f54a51107120cf8fa5197076ea99e85ad690086768b26752b953292060a","name":"research-routes.json","bytes":470391},{"sha256":"b64e565197938ecf58e11ece6644d713c5b30c4113cd12799a6b85c201a8335d","name":"questions.json","bytes":27653},{"sha256":"2f050782a32c1bd1de37f68518bb02c18a5c19b6040a8608fdce39c17943294f","name":"board.json","bytes":132310},{"sha256":"21a1d3556191bf54458b13fa0ebe41b4550fb92a33ab9bee6518d82ef222c843","name":"sah.py","bytes":56280},{"sha256":"029efc05e4b791b297f3cb254a24887e3d23b98b1ab4a6639d1f6dc7b69cc82f","name":"export_transcript.py","bytes":10230},{"sha256":"2437974e3a6f4bb6d5f3b2d4a6edabfd29d9d4668afdf5f35990ad246cee8d34","name":"f636_fibre-sign-lag.py","bytes":625},{"sha256":"b43138b04de4cd408b7fd72285f9011a060c0a8763149d39d8653c44fd51aa43","name":"f654_job1438-cls-resid-offset.py","bytes":666},{"sha256":"f9bb5129fcde982401b1560e2f8997a22036a234539c3cc87ee13cbec775b8a1","name":"f2002_sieve-null.py","bytes":623}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}