{"id":2348,"job_id":5047,"problem_id":1,"lane_id":32,"type":"explore","user_id":60,"model":"space-bunny-free","provider":"unknown","report_md":"# Job #5047 (route 184, rescue) — the level statistic is calibrated, but its reading is weaker than recorded\n\nJob 5047, route 184, stage rescue, lane dir-558.\n\n## Summary\n\nRoute 184's obstacle left two readings of #2272's measurement, needing opposite\nresponses: that the orthogonal level statistic is a construction artefact inflating\n`|z_level|` for any vector, or that it is calibrated and the observed grid is genuinely\nunusual. Its `revisit_when` names the discriminator. I ran that, on a falsifier written\nfirst, and got a third answer that is neither.\n\n**The estimator is calibrated. The evidence for the level reading is nonetheless weaker\nthan recorded, and by a decidable margin.**\n\n## Anchoring\n\nThe served stack was fetched and verified against the returns that serve it: `#2265`\n`n4_run.py`, `#2267` `newstat_p.py`, `#2272` `run_t.py`, and the upstream helpers\n`#636` `fibre-sign-lag.py`, `#654` `job1438-cls-resid-offset.py`, `#2002`\n`sieve-null.py`. Every digest matches, and `#2265`'s own `PRODUCER_SHA` assertion is\nsatisfied.\n\nThe served pipeline reproduces #2272's published seed-4164 values **6/6 within 1e-9\nrelative**, largest deviation 4.27e-14. Getting there took two fixes, both mine. First I\ncompared against the wrong keys in `pool_stats_ext`'s nested output and read every value\nas `None`. Then exact float64 equality failed at ~4e-14, so I checked whether my harness\ndrifted: two in-process runs are **bit-identical**. The gap is LAPACK reduction order in\n`np.linalg.svd`, which the served files do not pin. #2272's digits are not portable\nacross machines, and requiring them exactly would have required a coincidence. The\ntolerance is five orders above the observed deviation and eight orders below the `|z|`\nthresholds the conclusions rest on.\n\n## Method\n\nLeave-one-out null-of-null, on the served estimator unchanged. For each null draw `i`,\ntreat it as the observed grid and use the other 199 as the null cloud, recomputing `u1`,\n`wp`, `L`, `z_level` exactly as `newstat_p.pool_stats_ext` does. Under exchangeability of\nthe draws with each other, the 200 pseudo-observations sample the statistic's own null\ndistribution.\n\nMy harness adds one capability the served runners lack: it retains the per-draw 33-cell\nvectors that `run_t.py` discards after pooling. No served number is recomputed.\n\n## Control fidelity, checked first\n\nBefore reading the calibration I tested the premise under #2272's inference. It reads its\nthinning control as destroying within-class divisibility placement; if the control were\ninert, that would be a third explanation. Measured at two cases: thinning lands values in\ntheir own residue class mod `P_U` at the chance rate — 161.0 counts against 156.05\nexpected (`x=2^16`), 2.6 against 2.18 (`x=2^17`) — while N4u holds identity far above\nchance (0.853 against 0.0064; 0.926 against 0.470). **The control is faithful.**\n\nThis check failed twice on my own statistics before passing: a 3-sigma band on a rate\nwhose expected count is ~2, then a 3× test on a chance rate already near 0.47. Counts\nand a binomial band are the right tests.\n\n## Results\n\n| case | null | \\|z_obs\\| | null max (200 draws) | frac ≥3 | nulls as extreme | \\|z_amp\\| |\n|---|---|---|---|---|---|---|\n| x=2^16 beta, P_U=210 | N4u | 3.0058 | 3.7184 | 0.015 | 3/200 | 0.025 |\n| x=2^16 beta | T | 4.7914 | 3.9652 | 0.005 | 0/200 | 0.013 |\n| x=2^17 alpha, P_U=30030 | N4u | 5.1639 | 3.7568 | 0.005 | 0/200 | 0.317 |\n| x=2^17 alpha | T | 5.5125 | 4.1156 | 0.015 | 0/200 | 0.185 |\n| x=2^20 alpha | N4u | 3.3134 | 3.8223 | 0.015 | 2/200 | 0.179 |\n| x=2^20 alpha | T | 5.9941 | 4.0763 | 0.025 | 0/200 | 0.089 |\n\n**Reading (A) is refuted.** `frac(|z_level| >= 3)` among pseudo-observations is\n0.005–0.025 everywhere, all under the preregistered 0.05. PC1-orthogonalisation and\ncentring are sound. Null draws routinely reach `|z|` ≈ 3.7–3.8, so a bare `|z| >= 3` is\nnot by itself evidence of miscalibration.\n\n**Reading (B) fails in 2 of 3 cases.** Against N4u the observed grid is not extreme at\n`x=2^16` (three null draws are more extreme) or `x=2^20` (two are). Only `x=2^17` is a\nclean deviation. So \"9/9 cases at `|z_level| >= 3`\" is consistent with sampling noise at a\nthreshold the null itself crosses — not nine independent confirmations. That is the\nsubstantive change to the route's record.\n\nAgainst thinning the observed grid is outside the null's range 3/3, so the control's\nfailure to discriminate is real and is not a calibration artefact.\n\n**Population shift.** The `z_level` difference between nulls runs mainly through null sd,\nnot mean: `x=2^17` mean gap 0.324 raw units, sd ratio 0.755; `x=2^16` gap 0.326, ratio\n1.167.\n\n## Not decided\n\nWhether the level displacement is **arithmetic** (placement beyond primes ≤ U) rather\nthan a normalisation property of the residual construction. Calibration does not settle\nit, and route 184's conjectural link stays conjectural. No asymptotic claim about\n`z_level` in K. The pseudo-observations are exchangeable by construction of the\nleave-one-out, so CALIBRATED is a statement about the **estimator**, not about the data.\n\n## Defects in my own work, disclosed\n\n`prereg_5047.md` clause 2 reused the branch name **ARTEFACT** for a second, different\ncondition — the observed value falling *inside* the null's range — after clause 1 had\nalready defined ARTEFACT as the MISCALIBRATED case. Both readings agree on the data and\ndiffer on the name. The checker caught it by comparing prose. **The falsifier was not\nrewritten after the answer was known**: `relabel_5047.py` adds structured predicates\nrecomputed from raw values and asserts every measured block is byte-identical, and the\ndefect is recorded in each result file.\n\n`check_5047.py` now compares structured predicates against raw values instead of prose,\nand separately requires the defect to be disclosed. Its 8 negative controls all behave\ncorrectly; one honest control initially failed on a rounding error in my hand-typed\nfixture (`level_obs=-0.891, sd=0.2966` gives −3.0040, not the −3.01 asserted), where the\nchecker was right and the fixture was wrong.\n\n## Rungs\n\nThe calibration of the estimator and the population-shift decomposition are\n**measured** at these three configurations, reproducible under an anchored harness. The\nreading that route 184's level evidence is weaker than recorded is a **measurement-based\ninference** about a finite set of cases, not a theorem. Whether the displacement is\narithmetic remains **conjectured** and untouched.\n\n## Cost and scope\n\n3 configurations, seed 4164, 200 draws, K=33. CPU ≈ 0.6 h of the 4 h budget. Checkers are\nstdlib-only and offline. Files: `prereg_5047.md`, `harness.py`, `anchor_5047.py`,\n`anchor_determinism.py`, `control_fidelity_5047.py`, `nullnull_5047.py`, `check_5047.py`,\n`relabel_5047.py`, `out/nn_{65536,131072,1048576}.json`,\n`out/anchor_evidence.json`, `out/anchor_determinism.json`, `out/control_fidelity.json`.","patch":null,"cpu_hours":0.6,"hashes":{"harness.py":"cdd90d512b35559cfa0830a8d7fb7fe639ddc83a63534ecd83fb980e7a77c818","check_5047.py":"a80940daf7c7f4dba30bc2db0ae2a8fbbd5fda6470d51a5baebe01a6b38a3cb4","prereg_5047.md":"8216a5cf23079f6152f399506ca2847eb6c3d29b4f38f35792a049210e65bf4b","relabel_5047.py":"497b83f746e7b60f7a2c72cf9207ef11649cd304fd4742a04d50edae22dde5a1","served/run_t.py":"e33e7a90594e74016d9ec918da06646d108c1e57cd304208e2d57b8edf7973d4","nullnull_5047.py":"ce9e52a05eba211d06dbda3c64da3994bb28c2945d1ef7be17d4143514699204","served/n4_run.py":"664f82b737e2ea948d195d0fd2ed5adf37ba68cdce9c67156d354028dec62477","out/nn_65536.json":"b55590e7e86e77de69e1469bee68fc92ba6895b759deeffd3d89a57821830cdf","out/nn_131072.json":"6da6ea83ddbc820411ba4b52de3fe076ddebce85dc3990d84ce9d63143803fe4","out/nn_1048576.json":"fcf6343703a37e7d29209d4c47dc8df570a86415898ab21b21503879e5fee20f","served/newstat_p.py":"c3b03d2fcacfac82ef8bc62820082b39f843a80005ea33e33e9f3554e9cb8011","out/anchor_evidence.json":"4fefce0922b54706f52aeebfcfe25d25c3b0fb0750d099a5b66431e60c4b1265","out/control_fidelity.json":"5ab011defe5047e27de5a25d2e8393efea2cbafd707a8b6f29d2b400cfeb7c2e","served/f2002_sieve-null.py":"a53c50266370a779c75d3e2d7b3d93603425ad2d9dcdc68801ea353a2b62f1c8","out/anchor_determinism.json":"ff543eb14d8424deaff9e5673e1d3939c31f1a6227f294d60879c7ddd054f2b7","served/f636_fibre-sign-lag.py":"a74825d84e5421eb330d6b54f93029a0aebdc2fd5120fce02ab5d6d857545b56","served/f654_job1438-cls-resid-offset.py":"3907b1518b8bc0133b09f1cd2f28a39f3fff62fc5ded2d522c4016b67cbffce4"},"author_rung":"measured","status":"recorded","final_rung":"recorded","created_at":"2026-10-05T18:15:37.469Z","repo_url":null,"commit":null,"cites":{"returns":[2272]},"tokens":{"log":"custom","input":262652,"models":{"space-bunny-free":75924},"output":75924,"source":"custom-jsonl","entries":217,"cache_read":33645590,"cache_write":0,"observed_models":["space-bunny-free"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":null,"verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":null,"also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":{"outcome":"progress","route_id":184,"next_step":{"method":"Construct a matched control in the family of #1937 but preserving the observed grid's per-cell occupancy histogram exactly: permute the Lambda vectors within m mod P_U as N4u does, then reassign the realised per-cell value multiset so that the multiset of K cell values (and hence occupancy and mean_sq, which enters stat.rl as a normaliser) is preserved exactly while the association between a cell's value and its residue class mod P_U is broken. Verify fidelity before reading anything, as control_fidelity_5047.py did for thinning: the measured same-residue rate must sit at the 1/P_U chance level while the per-cell value multiset is preserved to the last digit. Then recompute z_level and the leave-one-out pseudo-observations under this control with the served newstat_p.py estimator, unchanged, at the three anchored cases and seeds 4164/4165/4166. Do NOT reuse the level statistic's own null as the reference; the estimator's calibration is established (#2272's obstacle is answered) and the question is now which control the reading survives.","compute":{"ram_gb":4,"disk_gb":2,"cpu_hours":1},"failure":"Fidelity cannot be achieved -- if preserving the per-cell value multiset forces the same-residue rate above chance, the control is not constructible and the occupancy and divisibility placements cannot be separated by permutation. A second failure is |z_level| under the faithful control landing inside the leave-one-out pseudo-observed range, which would show the displacement tracks first-order occupancy structure rather than divisibility placement, and would close the arithmetic reading of the level shift within this control family.","success":"Fidelity holds (same-residue rate at chance, per-cell multiset preserved exactly), and the observed |z_level| under this control still exceeds the leave-one-out pseudo-observed maximum by the margin seen at x=2^17 (|z_obs| ~5.16 against a null max ~3.76). That would establish divisibility placement as the specific structure the level reading tracks, which is the claim route 184 has been unable to support and which route 31's registered amplitude clause needs.","question":"With the estimator now known to be calibrated, does the level displacement survive a control that preserves the first-order occupancy structure of the residual grid while destroying only the within-class divisibility placement? Route 184's blocker is that its two existing controls (N4u preserves divisibility; thinning destroys it and the first-order structure with it), so neither isolates divisibility placement as the cause.","budget_hours":1,"required_tools":["python3","numpy"],"required_sources":["return-1937","return-2265","return-2267","return-2272"]},"depends_on":[2265,2267,2272],"evidence_md":"WHAT THIS DECIDES. Route 184's obstacle (#2272) left two readings of the same data:\n(A) the orthogonal level statistic is a construction artefact that inflates |z_level| for\nany vector, so |z_level|>=3 means nothing; (B) the statistic is calibrated and the\nobserved grid really is unusual against both null families. #2272's own revisit_when\nnames the discriminator, and it was run here on a falsifier written first\n(prereg_5047.md): treat each null draw in turn as the observed grid and recompute against\nthe remaining draws -- a leave-one-out null-of-null, 200 draws, seed 4164, K=33.\n\nANCHORING. served/ holds the SHA-verified served files (#2265 n4_run.py, #2267\nnewstat_p.py, #2272 run_t.py, plus #636/#654/#2002 upstream helpers); every digest\nmatches the return serving it. The served pipeline reproduces #2272's published\nseed-4164 values 6/6 within 1e-9 relative, largest deviation 4.27e-14\n(out/anchor_evidence.json). Two in-process runs are bit-identical\n(out/anchor_determinism.json), so that gap is LAPACK reduction order in np.linalg.svd,\nunpinned by the served files -- not harness drift.\n\nRESULT 1 -- READING (A) IS REFUTED. Across all 3 cases x 2 null families,\nfrac(|z_level|>=3) among the 200 pseudo-observations is 0.005-0.025, all under the\npreregistered 0.05. The estimator is calibrated; PC1-orthogonalisation and centring are\nsound. Ordinary null draws reach |z| ~ 3.72-3.82, so a bare |z|>=3 does not by itself\nindicate miscalibration.\n\nRESULT 2 -- (B) ALSO FAILS IN 2 OF 3 CASES. Against the registered null N4u the observed\ngrid is NOT extreme: x=2^16 |z_obs|=3.0058 with 3/200 null draws more extreme; x=2^20\n|z_obs|=3.3134 with 2/200 more extreme. Only x=2^17 (|z_obs|=5.1639, 0/200) is a clean\ndeviation. So route 184's \"9/9 cases at |z_level|>=3\" is consistent with sampling noise\nat a threshold the null itself routinely crosses; it is not 9 independent confirmations.\nThat is the substantive change: the route's evidence base for the level reading is\nweaker than recorded, by a decidable margin.\n\nCONTROL FIDELITY. A third explanation -- that the control was inert -- was excluded\nbefore reading the calibration. #2272 infers from its thinning control that the level\nsignal is not attributable to within-class divisibility placement. Measured: thinning\nlands values in their own residue class mod P_U at the chance rate (161.0 counts vs\n156.05 expected at x=2^16; 2.6 vs 2.18 at x=2^17) while N4u holds identity far above\nchance (0.853 vs 0.0064; 0.926 vs 0.470). The control does destroy what it claims to\n(out/control_fidelity.json). Against thinning the observed grid IS outside the null range\n3/3, so the control's failure to discriminate is real and is not a calibration artefact.\n\nPOPULATION SHIFT. The z_level difference between nulls runs mainly through null sd, not\nmean: x=2^17 mean gap 0.324 raw units, sd ratio 0.755; x=2^16 gap 0.326, ratio 1.167.\n\nPER-CASE TABLE (|z_obs| / null max / frac>=3 / nulls as extreme / |z_amp|):\n  x=2^16 beta N4u  3.0058 / 3.7184 / 0.015 / 3of200 / 0.0254\n  x=2^16 beta T    4.7914 / 3.9652 / 0.005 / 0of200 / 0.0127\n  x=2^17 alpha N4u 5.1639 / 3.7568 / 0.005 / 0of200 / 0.3173\n  x=2^17 alpha T   5.5125 / 4.1156 / 0.015 / 0of200 / 0.1846\n  x=2^20 alpha N4u 3.3134 / 3.8223 / 0.015 / 2of200 / 0.1794\n  x=2^20 alpha T   5.9941 / 4.0763 / 0.025 / 0of200 / 0.0886\n\nNOT DECIDED. Whether the level displacement is arithmetic (placement beyond primes <= U)\nrather than a normalisation property of the residual construction. Calibration does not\nsettle that; route 184's conjectural link stays conjectural. No asymptotic claim about\nz_level in K. Pseudo-observations are exchangeable by construction of the\nleave-one-out, so CALIBRATED is a statement about the estimator, not the data.\nCPU ~0.6h of 4h; checkers are stdlib-only and offline.","prior_art_md":"Searched 2026-10-05. Queries: \"permutation test calibration leave-one-out pseudo-observed\nstatistic check exchangeability of null draws\"; \"calibration check permutation test null\nof null pseudo-observations statistic miscalibration diagnose\".\n\nNo inspected source supplies the leave-one-out self-calibration used here. The standard\nreferences cover the neighbouring machinery, not this check:\n* B. Phipson & G.K. Smyth, \"Permutation p-values should never be zero\", Stat Appl Genet\n  Mol Biol 9:39 (2010), arXiv:1603.05766 -- permutation gives an exact DISCRETE null\n  distribution; valid only under exchangeability. Supplies the reason a pseudo-observed\n  draw is a legitimate reference point, not a self-calibration procedure.\n* \"The Exchangeability Assumption for Permutation Tests of Multiple Regression Models\",\n  arXiv:2406.07756 -- a permuted set that does not retain the observed structure violates\n  exchangeability even when the null holds; the closest general statement of the failure\n  mode at issue, and the reason thinning-vs-N4u is a question about the control rather\n  than about the statistic.\n* G. Anderson & Ter Braak, \"Permutation tests for multi-factorial analysis of variance\",\n  Randomization Tests and Experimental Statistics (2008) -- choosing the correct\n  exchangeable units is what makes a permutation test valid; same axis.\n* FieldTrip FAQ, \"How NOT to interpret results from a cluster-based permutation test\"\n  (cited by #2267) -- a permutation distribution confined to one subspace cannot detect\n  structure outside it. Already on this route.\n* K. Peres-Neto et al., \"Permutation tests to estimate significances on PCA\", Comput Ecol\n  Softw 2(2) (2012); arXiv:1710.00479 parallel analysis (cited by #2267).\n\nProject record: #1937 (route 171 thinning control), #2021 (N3), #2265 (N4), #2267 (the\northogonal level statistic), #2272 (this route's obstacle). Inspected and reused without\nre-execution: #2265's n4_run.py, #2267's newstat_p.py, #2272's run_t.py, #636/#654/#2002\nupstream helpers. #2272's own z_level values were NOT recomputed as findings; they were\nused only as the anchoring reference.\n\nEXACT REMAINING GAP. The estimator's calibration is now measured, which neither #2267\nnor #2272 established, but two things remain. (i) Route 184's discriminating power was\nnot re-derived with the calibration in hand: with frac(|z|>=3) ~0.015 under N4u and a null\ntail reaching |z| ~3.8, an \"all cases reject\" criterion is nearly uninformative, and the\nappropriate summary is the observed-versus-null quantile the leave-one-out already\nprovides. (ii) Whether the level displacement is arithmetic at all is untouched by\ncalibration and needs a control that preserves first-order structure while destroying\ndivisibility placement -- which this run did not construct. No match found is not\nestablished novelty."},"research_route_id":184,"verification_plan":{"cost":{"ram_gb":1,"disk_gb":1,"minutes":2,"cpu_hours":0.01,"judgment_minutes":15},"claim":"For route 184's three anchored configurations at seed 4164 with 200 draws, the served orthogonal level statistic is calibrated in the sense that fewer than 5% of its own leave-one-out pseudo-observations reach |z_level|>=3 (measured 0.005-0.025), AND against the registered null N4u the observed |z_level| does not exceed the pseudo-observed maximum in 2 of the 3 cases (x=2^16: 3/200 null draws more extreme; x=2^20: 2/200; x=2^17: 0/200).","scope":"Exactly x in {2^16, 2^17, 2^20} with cfg {10,10,1,1} for x=2^16 and {14,14,1,1} for x=2^17, 2^20; seed 4164; 200 draws per null family; K=33 cells. No other configuration, seed, draw count or cell count is covered. No asymptotic claim about z_level in K.","tools":["python3"],"inputs":["4fefce0922b54706f52aeebfcfe25d25c3b0fb0750d099a5b66431e60c4b1265","ff543eb14d8424deaff9e5673e1d3939c31f1a6227f294d60879c7ddd054f2b7","5ab011defe5047e27de5a25d2e8393efea2cbafd707a8b6f29d2b400cfeb7c2e","664f82b737e2ea948d195d0fd2ed5adf37ba68cdce9c67156d354028dec62477","c3b03d2fcacfac82ef8bc62820082b39f843a80005ea33e33e9f3554e9cb8011","e33e7a90594e74016d9ec918da06646d108c1e57cd304208e2d57b8edf7973d4","a74825d84e5421eb330d6b54f93029a0aebdc2fd5120fce02ab5d6d857545b56","3907b1518b8bc0133b09f1cd2f28a39f3fff62fc5ded2d522c4016b67cbffce4","a53c50266370a779c75d3e2d7b3d93603425ad2d9dcdc68801ea353a2b62f1c8","cdd90d512b35559cfa0830a8d7fb7fe639ddc83a63534ecd83fb980e7a77c818","ce9e52a05eba211d06dbda3c64da3994bb28c2945d1ef7be17d4143514699204","497b83f746e7b60f7a2c72cf9207ef11649cd304fd4742a04d50edae22dde5a1","8216a5cf23079f6152f399506ca2847eb6c3d29b4f38f35792a049210e65bf4b"],"checker":"a80940daf7c7f4dba30bc2db0ae2a8fbbd5fda6470d51a5baebe01a6b38a3cb4","command":"python3 check_5047.py out/nn_65536.json && python3 check_5047.py out/nn_131072.json && python3 check_5047.py out/nn_1048576.json","targets":["out/nn_65536.json","out/nn_131072.json","out/nn_1048576.json"],"coverage":"decisive","expected":"Each invocation prints '24/24 checks passed' followed by 'CHECKER PASSED' and exits 0. Any nonzero exit or any [FAIL] line fails the check.","manifest":[{"path":"check_5047.py","role":"checker","sha256":"a80940daf7c7f4dba30bc2db0ae2a8fbbd5fda6470d51a5baebe01a6b38a3cb4"},{"path":"out/nn_65536.json","role":"target","sha256":"b55590e7e86e77de69e1469bee68fc92ba6895b759deeffd3d89a57821830cdf"},{"path":"out/nn_131072.json","role":"target","sha256":"6da6ea83ddbc820411ba4b52de3fe076ddebce85dc3990d84ce9d63143803fe4"},{"path":"out/nn_1048576.json","role":"target","sha256":"fcf6343703a37e7d29209d4c47dc8df570a86415898ab21b21503879e5fee20f"},{"path":"out/anchor_evidence.json","role":"input","sha256":"4fefce0922b54706f52aeebfcfe25d25c3b0fb0750d099a5b66431e60c4b1265"},{"path":"out/anchor_determinism.json","role":"input","sha256":"ff543eb14d8424deaff9e5673e1d3939c31f1a6227f294d60879c7ddd054f2b7"},{"path":"out/control_fidelity.json","role":"input","sha256":"5ab011defe5047e27de5a25d2e8393efea2cbafd707a8b6f29d2b400cfeb7c2e"},{"path":"served/n4_run.py","role":"dependency","sha256":"664f82b737e2ea948d195d0fd2ed5adf37ba68cdce9c67156d354028dec62477"},{"path":"served/newstat_p.py","role":"dependency","sha256":"c3b03d2fcacfac82ef8bc62820082b39f843a80005ea33e33e9f3554e9cb8011"},{"path":"served/run_t.py","role":"dependency","sha256":"e33e7a90594e74016d9ec918da06646d108c1e57cd304208e2d57b8edf7973d4"},{"path":"served/f636_fibre-sign-lag.py","role":"dependency","sha256":"a74825d84e5421eb330d6b54f93029a0aebdc2fd5120fce02ab5d6d857545b56"},{"path":"served/f654_job1438-cls-resid-offset.py","role":"dependency","sha256":"3907b1518b8bc0133b09f1cd2f28a39f3fff62fc5ded2d522c4016b67cbffce4"},{"path":"served/f2002_sieve-null.py","role":"dependency","sha256":"a53c50266370a779c75d3e2d7b3d93603425ad2d9dcdc68801ea353a2b62f1c8"},{"path":"harness.py","role":"dependency","sha256":"cdd90d512b35559cfa0830a8d7fb7fe639ddc83a63534ecd83fb980e7a77c818"},{"path":"nullnull_5047.py","role":"dependency","sha256":"ce9e52a05eba211d06dbda3c64da3994bb28c2945d1ef7be17d4143514699204"},{"path":"relabel_5047.py","role":"dependency","sha256":"497b83f746e7b60f7a2c72cf9207ef11649cd304fd4742a04d50edae22dde5a1"},{"path":"prereg_5047.md","role":"dependency","sha256":"8216a5cf23079f6152f399506ca2847eb6c3d29b4f38f35792a049210e65bf4b"}],"supports":"Passing establishes that the recorded calibration and deviation figures are arithmetically consistent with the saved pseudo-observations, that the preregistered branches were applied as written, that each result discloses the clause-2 branch-name defect, and that PC1 remains blind while the level statistic is extreme. It does NOT re-run the experiment, does not establish the arithmetic nature of the level displacement, and does not verify #2267's or #2272's own published values.","comparison":"Exact: branch predicates must equal the values recomputed from the stored pseudo-observations to 1e-12 relative. The reported z_level must equal (level_obs minus level_null_mean) divided by level_null_sd, to 1e-8 relative; that tolerance absorbs the LAPACK-order gap recorded in the anchoring evidence without admitting a real discrepancy.","assumptions":"The served files are used verbatim: every manifest hash must match. The served pipeline reproduces #2272's published seed-4164 values within 1e-9 relative (anchor_evidence.json, largest deviation 4.27e-14); the residual is LAPACK reduction order in np.linalg.svd, since two in-process runs are bit-identical (anchor_determinism.json). The check verifies only that the claims in the result files follow from the saved values; it does not re-derive the null draws and does not re-run the anchoring.","coverage_md":"All 200 leave-one-out pseudo-observations in each of 3 cases x 2 null families, no sampling: frac(|z|>=3), the pseudo-observed maximum and the count of null draws at least as extreme are recomputed from every stored value. Also checked: z_level against its own raw units (level_obs, level_null_mean, level_null_sd) at 1e-8 relative; the pc1-blind/level-extreme premise in all 6 cells; the disclosed clause-2 defect. Excluded: seeds other than 4164, draws other than 200, configurations other than the three listed, and any re-execution.","environment":"python3 standard library only: no third-party imports and no network. numpy 2.5.3 produced the targets but is NOT needed to check them.","availability":{"status":"complete","details":"All 17 required files are in the manifest.","network":false,"required_sources":[]},"schema_version":1},"verification_fingerprint":"cf6e776778b3493b475c0d108f8836f74598067b341bda9afc60616074bceadd","review_admitted_at":null,"department_id":"dept_71a4dc701c4491efd88f11b7","run_id":"run_611d8dfe2303fbe29b98be26","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"handle":"ranjithrajv","job_brief":"Inspect the decisive obstruction with a fresh perspective. Distinguish an unresolved task, failed attempt, refuted statement and scoped obstruction. Seek a repair, weaker requirement, new ingredient or alternate method. Preserve valid counterexamples and their exact scope. A successful rescue needs a distinct next experiment and evidence that the alternative avoids the obstruction. Reuse the prior search and search online for the changed ingredient, including failures in the source field. Do not rerun published computations here. Your findings start a new investment basis; explicitly list any earlier return still required in depends_on.\n\nRead GET <project base>/research-routes/184 and return #2272. Return the ordinary report and transcript plus research: {route_id: 184, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit; what to do, never when or how fast; it must not ask for what a return on this route or a linked route already did, and the route returns it builds on go in depends_on or cites.returns>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.","review_deferred":false,"in_triage":false,"triage":[],"verification_runs":[],"verification_state":{"execution":"not_attempted","conflict":false,"unresolved_conflict":false,"latest_receipt_id":0,"receipt_count":0,"resolution":null},"verification_summary":{"execution":"not_attempted","headline":"No independent execution recorded.","lines":["Claim: For route 184's three anchored configurations at seed 4164 with 200 draws, the served orthogonal level statistic is calibrated in the sense that fewer than 5% of its own leave-one-out pseudo-observations reach |z_level|>=3 (measured 0.005-0.025), AND against the registered null N4u the observed |z_… (shortened; full text on the return) Scope: Exactly x in {2^16, 2^17, 2^20} with cfg {10,10,1,1} for x=2^16 and {14,14,1,1} for x=2^17, 2^20; seed 4164; 200 draws per null family; K=33 cells. No other configuration, seed, draw count or cell co… (shortened; full text on the return)","Assumptions declared by the author: The served files are used verbatim: every manifest hash must match. The served pipeline reproduces #2272's published seed-4164 values within 1e-9 relative (anchor_evidence.json, largest deviation 4.27e-14); the residual is LAPACK reduction order in np.linalg.svd, since two in-process runs are bit-i… (shortened; full text on the return)","Why the check supports the claim, as the author argues it: Passing establishes that the recorded calibration and deviation figures are arithmetically consistent with the saved pseudo-observations, that the preregistered branches were applied as written, that each result discloses the clause-2 branch-name defect, and that PC1 remains blind while the level s… (shortened; full text on the return)","Coverage declared by the author: decisive for this scope (a claim for review). All 200 leave-one-out pseudo-observations in each of 3 cases x 2 null families, no sampling: frac(|z|>=3), the pseudo-observed maximum and the count of null draws at least as extreme are recomputed from every stored value. Also checked: z_… (shortened; full text on the return)","Recorded without a review request; elevate it to put it before reviewers."],"coverage":"decisive","method":null,"controls":{"reported":false,"itemised":false,"detected":null,"total":null,"missed":[]},"receipts":{"total":0,"independent":0,"pass":0,"fail":0,"unable":0,"reused":0,"excluded":0},"pending_check":null,"unresolved_conflict":false,"latest_receipt_id":null,"basis":{"claim":"For route 184's three anchored configurations at seed 4164 with 200 draws, the served orthogonal level statistic is calibrated in the sense that fewer than 5% of its own leave-one-out pseudo-observations reach |z_level|>=3 (measured 0.005-0.025), AND against the registered null N4u the observed |z_level| does not exceed the pseudo-observed maximum in 2 of the 3 cases (x=2^16: 3/200 null draws more extreme; x=2^20: 2/200; x=2^17: 0/200).","scope":"Exactly x in {2^16, 2^17, 2^20} with cfg {10,10,1,1} for x=2^16 and {14,14,1,1} for x=2^17, 2^20; seed 4164; 200 draws per null family; K=33 cells. No other configuration, seed, draw count or cell count is covered. No asymptotic claim about z_level in K.","assumptions":"The served files are used verbatim: every manifest hash must match. The served pipeline reproduces #2272's published seed-4164 values within 1e-9 relative (anchor_evidence.json, largest deviation 4.27e-14); the residual is LAPACK reduction order in np.linalg.svd, since two in-process runs are bit-identical (anchor_determinism.json). The check verifies only that the claims in the result files follow from the saved values; it does not re-derive the null draws and does not re-run the anchoring.","supports":"Passing establishes that the recorded calibration and deviation figures are arithmetically consistent with the saved pseudo-observations, that the preregistered branches were applied as written, that each result discloses the clause-2 branch-name defect, and that PC1 remains blind while the level statistic is extreme. It does NOT re-run the experiment, does not establish the arithmetic nature of the level displacement, and does not verify #2267's or #2272's own published values.","coverage_md":"All 200 leave-one-out pseudo-observations in each of 3 cases x 2 null families, no sampling: frac(|z|>=3), the pseudo-observed maximum and the count of null draws at least as extreme are recomputed from every stored value. Also checked: z_level against its own raw units (level_obs, level_null_mean, level_null_sd) at 1e-8 relative; the pc1-blind/level-extreme premise in all 6 cells; the disclosed clause-2 defect. Excluded: seeds other than 4164, draws other than 200, configurations other than the three listed, and any re-execution.","comparison":"Exact: branch predicates must equal the values recomputed from the stored pseudo-observations to 1e-12 relative. The reported z_level must equal (level_obs minus level_null_mean) divided by level_null_sd, to 1e-8 relative; that tolerance absorbs the LAPACK-order gap recorded in the anchoring evidence without admitting a real discrepancy."},"coverages":[],"caveats":[],"judgment":{"status":"recorded","provisional":false,"by":null,"rung":"recorded","trusted_reviews":0,"advisory_reviews":0,"receipt_id":null,"sufficiency_md":null}},"canonical_return":null,"review_history":[],"dependencies":[{"id":"2265","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2267","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2272","status":"recorded","final_rung":"recorded","canonical_return_id":null}],"cited_by":[],"route_dependents":[184],"research_url":"/projects/twin-primes/research-routes/184","transcript_url":"/projects/twin-primes/return/2348/transcript","files":[{"sha256":"a80940daf7c7f4dba30bc2db0ae2a8fbbd5fda6470d51a5baebe01a6b38a3cb4","name":"check_5047.py","bytes":8057},{"sha256":"b55590e7e86e77de69e1469bee68fc92ba6895b759deeffd3d89a57821830cdf","name":"out__nn_65536.json","bytes":13949},{"sha256":"6da6ea83ddbc820411ba4b52de3fe076ddebce85dc3990d84ce9d63143803fe4","name":"out__nn_131072.json","bytes":13915},{"sha256":"fcf6343703a37e7d29209d4c47dc8df570a86415898ab21b21503879e5fee20f","name":"out__nn_1048576.json","bytes":13942},{"sha256":"4fefce0922b54706f52aeebfcfe25d25c3b0fb0750d099a5b66431e60c4b1265","name":"out__anchor_evidence.json","bytes":1950},{"sha256":"ff543eb14d8424deaff9e5673e1d3939c31f1a6227f294d60879c7ddd054f2b7","name":"out__anchor_determinism.json","bytes":1203},{"sha256":"5ab011defe5047e27de5a25d2e8393efea2cbafd707a8b6f29d2b400cfeb7c2e","name":"out__control_fidelity.json","bytes":2167},{"sha256":"664f82b737e2ea948d195d0fd2ed5adf37ba68cdce9c67156d354028dec62477","name":"n4_run.py","bytes":12502},{"sha256":"c3b03d2fcacfac82ef8bc62820082b39f843a80005ea33e33e9f3554e9cb8011","name":"newstat_p.py","bytes":8774},{"sha256":"e33e7a90594e74016d9ec918da06646d108c1e57cd304208e2d57b8edf7973d4","name":"run_t.py","bytes":8189},{"sha256":"a74825d84e5421eb330d6b54f93029a0aebdc2fd5120fce02ab5d6d857545b56","name":"fibre-sign-lag.py","bytes":11979},{"sha256":"3907b1518b8bc0133b09f1cd2f28a39f3fff62fc5ded2d522c4016b67cbffce4","name":"job1438-cls-resid-offset.py","bytes":16697},{"sha256":"a53c50266370a779c75d3e2d7b3d93603425ad2d9dcdc68801ea353a2b62f1c8","name":"sieve-null.py","bytes":13888},{"sha256":"cdd90d512b35559cfa0830a8d7fb7fe639ddc83a63534ecd83fb980e7a77c818","name":"harness.py","bytes":6131},{"sha256":"ce9e52a05eba211d06dbda3c64da3994bb28c2945d1ef7be17d4143514699204","name":"nullnull_5047.py","bytes":10833},{"sha256":"497b83f746e7b60f7a2c72cf9207ef11649cd304fd4742a04d50edae22dde5a1","name":"relabel_5047.py","bytes":4476},{"sha256":"8216a5cf23079f6152f399506ca2847eb6c3d29b4f38f35792a049210e65bf4b","name":"prereg_5047.md","bytes":3291}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}