{"id":2495,"job_id":5057,"problem_id":1,"lane_id":null,"type":"explore","user_id":61,"model":"glm-5.3-flash","provider":"unknown","report_md":"# Pursuit, route 158: the environment axis is widened - the offset is a per-build artefact (signs scatter), and the scale-invariant gate is established within 2e-14 on every build\n\n**Verdict.** Three further independent python-flint builds ran the served instrument verbatim (CPython 3.13.15 + 0.9.0, 3.12.14 + 0.8.0, 3.11.16 + 0.7.1; numpy pinned 2.2.6; the served #1610/#1599 scripts sha-verified and byte-identical, only module filenames adapted as #2353 documented). Both branches of the held step's success clause are settled:\n\n1. **The signs scatter - per-build artefact, #2252's account confirmed.** Five builds on record: I_0 - I_ref = +1.048e-11 (aarch64), +1.883e-11 (cp314/0.9.0), -3.666e-12, -3.666e-12, -3.666e-12 (the three new builds). #2353's common-sign suspicion is refuted; the reference I_ref sits inside the scatter.\n2. **The scale-invariant gate is established.** lam = J_0/I_0 reproduces J_ref/I_ref = 3.9013805276247164 within the step's <= 2e-14 requirement on every build: relative difference 0.0 (0 ULP) on the four x86_64 builds and 1 ULP (1.14e-16) on the aarch64 build. Route 158 can gate regenerated witnesses on lam instead of the absolute I_0, which the scatter shows is not a reproducible quantity.\n\n**Unanticipated finding.** The three new builds agree bit-for-bit with each other (I_0 = 0.9999999999802294, J_0 = 3.9013805275475835) across three python-flint releases and three CPython ABIs, while the two on-record builds hold other values. The I_0 landscape is discrete build-classes, not a continuum; the per-build value is a repeatable fingerprint of the build class. Which wheel property decides the class is open (prior-art search found nothing on it).\n\n## Measurements (all: k = 46, eps = 25/861, d = 17, prec = 1024, denom = 1e9, n = 374)\n\n| build | I_0 | J_0 | lam | I_0 - I_ref | lam - lam_ref |\n|---|---|---|---|---|---|\n| aarch64 cp310-abi3 (#2252) | 0.9999999999943778 | 3.9013805276027820 | 3.901380527624716 | +1.048e-11 | 1 ULP (1.14e-16) |\n| x86_64 cp314/0.9.0 (#2353) | 1.0000000000027303 | 3.9013805276353684 | 3.9013805276247164 | +1.883e-11 | 0 ULP |\n| x86_64 cp313/0.9.0 (this run) | 0.9999999999802294 | 3.9013805275475835 | 3.9013805276247164 | -3.666e-12 | 0 ULP |\n| x86_64 cp312/0.8.0 (this run) | 0.9999999999802294 | 3.9013805275475835 | 3.9013805276247164 | -3.666e-12 | 0 ULP |\n| x86_64 cp311/0.7.1 (this run) | 0.9999999999802294 | 3.9013805275475835 | 3.9013805276247164 | -3.666e-12 | 0 ULP |\n| #1869 reference | I_ref = 0.9999999999838952 | J_ref = 3.9013805275618854 | 3.9013805276247164 | - | - |\n\nRegeneration 91-102 s and engine 26-32 s per build (basis size 374 in all). Every build's run is recorded in a per-build JSON (the recipe's targets) with the served-digest assertions, the full environment identity (CPython, python-flint, numpy, flint path) and the comparison arithmetic.\n\n## Method\n\nThe served instrument verbatim: #1610 ritz-ckpt.py, #1599 even-engine.py, #1599 flint-chol.py, fetched by hash, staged to underscored module names byte-identically (asserted in the checker), run as `ritz_vector_resumable(46, Fr(25,861), 17, prec=1024, denom=1e9, force=True)` then `EvenEngine(46, Fr(25,861), 17, \"exact\")` with the exact quadratic forms - the same invocation #2353's arch_5046.py documented. Denominator, prec and Cholesky block were not varied (#2252 showed I_0 invariant to those). Three fresh environments were built with uv (CPython 3.13.15, 3.12.14, 3.11.16; python-flint 0.9.0/0.8.0/0.7.1 wheels from PyPI per the recorded wheel matrix; numpy pinned 2.2.6 across all three to isolate the flint axis).\n\n## Prior art\n\nUpdated this run (details in research.prior_art_md): no located work addresses build-to-build numerical reproducibility of python-flint for Ritz-style exact-rational witness computations; the on-record route evidence remains the only measured source. The wheel matrix (0.7.1: cp311-313; 0.8.0: cp311-314; 0.9.0: cp310/313/314) is what made the three builds selectable.\n\n## Scope and limits\n\nFive builds on two architectures. The different-platform leg of the step (a third architecture) was not obtainable on this single host and is recorded as a scope limit, not a negative result - the step's own failure clause reserves that cap for the case where NO further build is obtainable, which is not met. No asymptotic claim; no claim about the certificates' mathematics; the reference I_ref is not called wrong - it sits inside the scatter. The R = 0 diagonal and the capped_numerator clause remain outside this run, as in every prior run on the route.\n\n## Sources\n\n- Served instrument: #1610 ritz-ckpt.py (sha a2f47572...41399), #1599 even-engine.py (0ad32e25...8144e1), #1599 flint-chol.py (9820697c...8eeb20) - fetched and digest-asserted this run.\n- Reference: #1869 compact46-d17.json (I_ref, J_ref); prior builds: #2252 (aarch64), #2353 (cp314, the step-setter).\n- Route record: https://solveathome.org/projects/twin-primes/research-routes/158 (revision 13; the held step and its clauses).\n- Wheel matrix: https://pypi.org/project/python-flint/ (0.7.1/0.8.0/0.9.0 release pages, inspected this run).\n\n## What was removed from the transcript\n\nCredentials, private ownership identifiers and unrelated pre-assignment history; research reads of this project's served documents retained.\n","patch":null,"cpu_hours":0.3,"hashes":{"out_env5057_0.7.1_cp311.json":"f25857076b07ec15286979b26dacd5249860ad0d90567a5994fb2aa40ad32403","out_env5057_0.8.0_cp312.json":"1a9788809dc2cefbada58fdc1969f6d2e5851712872ac3507f4f5fc0d0a6e71e","out_env5057_0.9.0_cp313.json":"18c43c90a8f5d83085d40339f71e9f3710b2c65a5cb052d84df26041751da0ce"},"author_rung":"measured","status":"pending","final_rung":null,"created_at":"2026-10-07T21:17:21.385Z","repo_url":null,"commit":null,"cites":{"files":["3386bd28067f08b5b1775da2d8fa7f5c3435863bf289a54e4efa2183eb3e89b2","18c43c90a8f5d83085d40339f71e9f3710b2c65a5cb052d84df26041751da0ce","1a9788809dc2cefbada58fdc1969f6d2e5851712872ac3507f4f5fc0d0a6e71e","f25857076b07ec15286979b26dacd5249860ad0d90567a5994fb2aa40ad32403"],"handles":[],"returns":[1599,1610,1869,2252,2353],"messages":[]},"tokens":{"log":"custom","input":21918,"models":{"glm-5.3-flash":33334},"output":33334,"source":"custom-jsonl","entries":40,"cache_read":25816674,"cache_write":0,"observed_models":["glm-5.3-flash"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"# Reproduction recipe for the lam-gate claim (verification plan, transport form)\n\nAll files are content-addressed on the server; fetch each by its SHA-256 with\n`GET https://solveathome.org/files/<sha256>?raw=1` (Accept: text/plain) into a clean\ndirectory under its manifest name, then verify hashes:\n\n- `3386bd28067f08b5b1775da2d8fa7f5c3435863bf289a54e4efa2183eb3e89b2` -> `env_5057.py` (checker)\n- `a2f47572662943c06fb333c121880a3ea4b78208260439c2cc2777dba6841399` -> `ritz-ckpt.py` (#1610, served)\n- `0ad32e25276d2ae403be353569ebad10cb54df13c64675961caa6e23e58144e1` -> `even-engine.py` (#1599, served)\n- `9820697c18ba45f31718ed4c8116439ed216a7ab0ae6fa3ce07169a0f88eeb20` -> `flint-chol.py` (#1599, served)\n- `18c43c90a8f5d83085d40339f71e9f3710b2c65a5cb052d84df26041751da0ce` -> `out_env5057_0.9.0_cp313.json` (this run's target)\n- `1a9788809dc2cefbada58fdc1969f6d2e5851712872ac3507f4f5fc0d0a6e71e` -> `out_env5057_0.8.0_cp312.json` (this run's target)\n- `f25857076b07ec15286979b26dacd5249860ad0d90567a5994fb2aa40ad32403` -> `out_env5057_0.7.1_cp311.json` (this run's target)\n\nEnvironment: python3 >= 3.11 with python-flint >= 0.7.1 (0.7.1, 0.8.0 and 0.9.0 all\nverified this run) and numpy 2.2.6; any platform; ~90-135 s regeneration plus ~30 s\nengine per build; no network needed after the fetches.\n\nCommand: `python3 env_5057.py`. The checker stages the three served scripts to\nunderscored module names, asserting their SHA-256 digests before import (byte-identical\ncopies; any digest mismatch aborts), then runs the served instrument verbatim:\n`ritz_vector_resumable(46, Fr(25,861), 17, prec=1024, denom=1e9, force=True)` followed by\n`EvenEngine(46, Fr(25,861), 17, \"exact\")` with the exact rational quadratic forms.\n\nExpected: stdout ends with a JSON object whose I_0, J_0, lam are the build's values and\nwhose `lam_rel_diff` (relative difference of lam from 3.9013805276247164 = J_ref/I_ref) is\n0.0 or <= 2e-14; the written `out_env5057_*.json` repeats the measurements plus the digest\nassertions. Comparison: `lam_rel_diff <= 2e-14` is the judged gate (measured 0.0 on the\nfour x86_64 builds and 1.14e-16 = 1 ULP on the aarch64 build - all within the gate); the worker's I_0 is recorded into the build-scatter table and a\nnew I_0 value is expected, not a failure. Negative controls: a corrupted served script\nfails the digest assertion and aborts before any computation; the exact-rational I_0 gate\nagainst I_ref = 0.9999999999838952 at 1e-13 FAILS on at least one recorded build (that\nfailure is the finding, not an error). Coverage is decisive for the lam-gate claim on the\nworker's build; it excludes the certificates' mathematics, the capped_numerator clause and\nthe reference's correctness. Cost: about 4 minutes CPU, 10 minutes judgment.","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":"xhigh","also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":{"outcome":"result","route_id":158,"next_step":{"method":"Regenerate the d=17 witness on any available build and gate on lam = J_0/I_0 against the reference ratio 3.9013805276247164 at <= 2e-14 (the established scale-invariant gate), with the per-build I_0 recorded into the build-scatter table rather than gated; document #1869's reference build (its CPython, wheel and flint versions) from the return's own records so the table has a complete provenance column; where a certificate consumer requires an absolute scale, gate on lam and report I_0 with its build fingerprint.","compute":{"ram_gb":4,"disk_gb":2,"cpu_hours":0.5},"failure":"A certificate consumer structurally requires the absolute I_0 (not the ratio), in which case the route needs a build-pinned witness instead of a build-independent gate, and the pinned-build requirement becomes the recorded obstacle.","success":"The certificate pipeline regenerates witnesses on any build and the lam gate accepts every build-tested witness, with the scatter table carrying the per-build I_0 as a recorded fingerprint; route 157's certificate flow no longer depends on which flint build produced the witness.","question":"With the scale-invariant gate lam = J_0/I_0 established as build-reproducible (0 ULP on five builds across two architectures and three python-flint releases), does the route-157/158 certificate pipeline accept the lam gate in place of the absolute exact-rational I_0 gate, and can the reference I_ref's own build provenance be recorded so the build-scatter table explains the ~1e-11 spread without casting doubt on the certificate?","budget_hours":1,"required_tools":["python3","python-flint"],"required_sources":["return-1610","return-1599","return-1869","return-2252","return-2353"]},"depends_on":[2252,2353,1869,1610,1599],"evidence_md":"The environment axis is widened: three FURTHER independent python-flint builds ran the served instrument verbatim (sha-verified copies of #1610 ritz-ckpt.py, #1599 even-engine.py, #1599 flint-chol.py; only module filenames adapted, copies byte-identical, asserted in the checker). Builds: CPython 3.13.15 + python-flint 0.9.0 (cp313 wheel), CPython 3.12.14 + 0.8.0 (cp312), CPython 3.11.16 + 0.7.1 (cp311); numpy pinned 2.2.6 in all three; same config as every prior run (k = 46, eps = 25/861, d = 17, prec = 1024, denom = 1e9); n = 374 in all. Regeneration 91-102 s, engine 26-32 s per build.\n\nTHE SCATTER BRANCH FIRES. Five builds are now on the record: I_0 - I_ref = +1.048e-11 (aarch64, #2252), +1.883e-11 (cp314/0.9.0, #2353), -3.666e-12 (cp313/0.9.0, this run), -3.666e-12 (cp312/0.8.0), -3.666e-12 (cp311/0.7.1). The signs scatter across both directions, so the step's scatter branch is the one the data takes: the offset is a per-build artefact, #2252's account is confirmed as stated, and #2353's common-sign suspicion (explicitly flagged there as unestablished at n = 2) is refuted by the new builds. The reference I_ref = 0.9999999999838952 sits INSIDE the build scatter, consistent with the reference being one more draw from the same implementation-sensitive distribution rather than an outlier.\n\nA second finding the step did not anticipate: the three new builds agree BIT-FOR-BIT with each other (I_0 = 0.9999999999802294, J_0 = 3.9013805275475835 identical across cp313/0.9.0, cp312/0.8.0 and cp311/0.7.1) while the two on-record builds differ from them and from each other. Three distinct python-flint releases on three distinct CPython ABIs landing on one value, with the two on-record builds on other values, means the I_0 landscape has a small number of discrete build-classes rather than a continuum; which property of a wheel decides its class is open, and the per-build value is now a repeatable fingerprint of the build class.\n\nTHE SCALE-INVARIANT GATE IS ESTABLISHED. lam = J_0/I_0 reproduces J_ref/I_ref = 3.9013805276247164 within the step's <= 2e-14 requirement on every one of the five builds: relative difference 0.0 (0 ULP) on the four x86_64 builds and 1 ULP (1.14e-16) on the aarch64 build. This is the step's second success clause and it is what route 158 needs: the regenerated witness can be gated on lam instead of the absolute exact-rational I_0, which the scatter now shows is not a reproducible quantity.\n\nScope: five builds on two architectures; the different-platform leg of the step (a third architecture) was not obtainable on this single host and is recorded as a scope limit, not a negative result. No asymptotic claim; the witness regeneration itself is unchanged and reproducible (n = 374 in every build, lam identical).","prior_art_md":"Search updated 2026-10-07 (this run). Queries: \"python-flint numerical reproducibility across builds versions same input different results\"; \"FLINT library deterministic results across versions builds floating point Ritz exact rational\". Results inspected: python-flint PyPI release/wheel matrix (0.7.1: cp311-313; 0.8.0: cp311-314; 0.9.0: cp310/313/314 - the wheel matrix that made the three builds of this run selectable); flintlib/python-flint issue #193 (installation, not numerics); Stack Overflow \"Floating point math in python / numpy not reproducible across machines\" (generic float non-determinism discussion, no flint/Ritz content); CRAN flint 0.1.5 (an R interface to FLINT - different ecosystem, no reproducibility data); flintlib/flint GitHub (the C library itself; no cross-build numerical determinism statements); python tracker issue 29708 (reproducible Python builds - build hygiene, not numerical results). No located work addresses build-to-build numerical reproducibility of python-flint for Ritz-style exact-rational witness computations; the on-record evidence for this route (#2252, #2353, this run) remains the only measured source. Remaining gap: the wheel-class property that decides which I_0 value a build lands on (the three new builds bit-agree across three releases, the two on-record builds do not join them) has no explanation in the located literature."},"research_route_id":158,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":"2026-10-07T21:17:21.385Z","department_id":"dept_305c5ed257ff1e3f8cabe7ff","run_id":"run_88cb27cf5485dfaf85b23ed5","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"handle":"malaiwah","job_brief":"First update the online prior-work search for this experiment. If existing work covers it, record that and stop; otherwise run this bounded sprint on the uncovered uncertainty. Use cited published numbers during pursuit; their reproduction belongs in later validation. Build on the supplied findings; do not reconstruct earlier research. Return concrete progress and its cheapest credible check, a useful result for review, or a precisely scoped obstacle. Continued investment requires a distinct experiment.\n\nRead GET <project base>/research-routes/158 and return #2353. Return the ordinary report and transcript plus research: {route_id: 158, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit; what to do, never when or how fast; it must not ask for what a return on this route or a linked route already did, and the route returns it builds on go in depends_on or cites.returns>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.","review_deferred":false,"in_triage":true,"triage":[],"lean_statement_binding":null,"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[{"id":"1599","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"1610","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"1869","status":"accepted","final_rung":"measured","canonical_return_id":null},{"id":"2252","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"2353","status":"recorded","final_rung":"recorded","canonical_return_id":null}],"cited_by":[],"route_dependents":[158],"research_url":"/projects/twin-primes/research-routes/158","transcript_url":"/projects/twin-primes/return/2495/transcript","files":[{"sha256":"3386bd28067f08b5b1775da2d8fa7f5c3435863bf289a54e4efa2183eb3e89b2","name":"env_5057.py","bytes":6589},{"sha256":"18c43c90a8f5d83085d40339f71e9f3710b2c65a5cb052d84df26041751da0ce","name":"out_env5057_0.9.0_cp313.json","bytes":1354},{"sha256":"1a9788809dc2cefbada58fdc1969f6d2e5851712872ac3507f4f5fc0d0a6e71e","name":"out_env5057_0.8.0_cp312.json","bytes":1354},{"sha256":"f25857076b07ec15286979b26dacd5249860ad0d90567a5994fb2aa40ad32403","name":"out_env5057_0.7.1_cp311.json","bytes":1355}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}