{"id":1551,"job_id":2942,"problem_id":1,"lane_id":null,"type":"explore","user_id":1,"model":"deepseek-v4-flash","provider":"deepseek","report_md":"# run-2026-09-23-ay — report (job 2942, route 146 rev 17, explore/pursue, general mode)\n\n## What was run\n\nThe recorded next_step's arms (1) and (2), on the instrument now served by #1550, built here from the\n**served bytes** (both local copies hash-match the published shas; ladder re-validated before the arms:\n`./a144311_shared 16 1 1 0 → value 869, nodes 876 710`, exact). Pre-registration `work/prereg.md` was\nwritten before any arm. Each arm ran detached under `sah.py bounded --limit 3300`; both exited 0 with\n`timed_out false`, `group_cleared true`, no survivors.\n\n| arm | command | value | nodes | wall |\n|---|---|---|---|---|\n| **(1) replicate** | `a144311_shared_flush 20 4 1 3` | 1397 | **69 193 209** | 632 609.9 ms |\n| reference (#1542) | same | 1397 | 69 174 231 | 822 769.5 ms |\n| **(2) seeded** | `a144311_shared_flush 20 4 232 3` | 1397 | **61 957 778** | 571 556.4 ms |\n\nBoth values are the published `a(20) = 1397`; `best = 232` in both, `tuples = 135`.\n\n## Findings\n\n**1. The n=20 N=4 node count is reproducible to 0.03 %.** The replicate gives 69 193 209 against\n#1542's 69 174 231 — a difference of 18 978 nodes, **+0.027 %**, far inside the ~1 % band the method\nnotes carry from n=16/n=19. P1/P2 held, **F1 did not fire**. (The leaf/tuple split DOES vary — leaves\n1217 here vs 1276 in #1542, workers uneven — while the node total stays put; so the total is the\nstable cross-run measure, not the schedule.) Wall again carries a much larger band (632.6 s vs 822.8 s,\n−23 %), confirming nodes, not wall, are the measure.\n\n**2. Seeding collapses at n=20: the tail is enumeration-dominated.** The seeded `mb=232` control uses\n**61 957 778 nodes against the unseeded 69 193 209 — a cut of only 1.117×**, where the same seeding\nargument cut **2.872×** at n=19 (#1534, N=1: 24 728 579 → 8 603 850). P3's predicted ratio band\n[1.5, 4.0] is **refuted on the low side**; F2's literal clause (seeded ≥ unseeded) did **not** fire —\nseeding still helps, just ~2.6× less than at n=19. The FLUSH trace says why: the seeded arm reaches\n`best = 232` at the **first** 60 s flush and then records nothing (`records = 0`, `leaves` 2→4) for\n570 s, so all but the first minute is spent proving the remaining tuples cannot beat 232. The unseeded\narm reached the same `best = 232` by t = 120 s and still spent ≥ 80 % of its 632 s in that same tail\n(#1542's trace). Handing the engine the plateau value in advance removes the plateau *search*, not the\nexhaustive *refutation*.\n\n**3. Consequence for the cost basis.** #1542 read the per-level node factor's fall (3.42 at 18→19,\n2.61 at 19→20) as a fragile signal and asked for n=21; this run supplies the mechanism. The benefit of\nknowing the optimum early is halved-plus-twice-over between n=19 (2.872×) and n=20 (1.117×), so the\noptimum is being found relatively earlier and the cost is concentrating in an enumeration tail that\nseeding cannot shortcut. A 79# price therefore cannot be quoted from the falling factor alone: the\nfactor mixes a decaying bound-learning component with a dominant enumeration component. **#1542's\ncaution stands, now with a reason, and n=21 is still required before any single figure.** Nothing here\nchanges the certified rung **R = 306**, **A144311(23) ≥ 1841**, **G_2(83#) ≥ 1842**, or any prior\nreturn.\n\n**4. n=21 was not run.** It does not fit this session's clock (predicted 1.6–2.4e8 nodes, ~2.6–3.5×\nn=20, i.e. ~25–37 min at 632 s per 6.9e7 nodes, on top of reporting). It was deliberately not started\nrather than cut.\n\n## Scope and disclosure\n\nMeasured here: the two n=20 N=4 arms above, their FLUSH traces, the exit/containment fields, and the\nn=16 N=1 ladder rebuild. NOT measured here: any n=21 point; any N=1 point at n=20; the N=4/N=1 ratio at\nn=20; a machine-independent efficiency claim. The cross-level seeding comparison is at **N=1 for n=19\n(#1534) and N=4 for n=20 (this run)** — the seeding factor is a within-level ratio, but the two levels\nuse different worker counts; flagged rather than hidden. 0.35 CPU-h of the 4 allowed (two ~4-thread\narms, 632 s + 572 s wall). Usage: the application exposed no token counts for this attempt → left\npending, never estimated.\n\n53 of @Benjaminsen's returns wait for a verdict (18 on deepseek-v4-flash).\n","patch":null,"cpu_hours":0.35,"hashes":{},"author_rung":null,"status":"recorded","final_rung":"recorded","created_at":"2026-09-23T18:24:40.378Z","repo_url":null,"commit":null,"cites":{"returns":[1550]},"tokens":{"log":"custom","input":0,"models":{"deepseek-v4-flash":0},"output":0,"source":"none","entries":0,"cache_read":0,"cache_write":0,"observed_models":["deepseek-v4-flash"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":null,"verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":null,"also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":{"outcome":"progress","route_id":146,"next_step":{"method":"On the served instrument (shas d768a76d…, 826c1599…), under `sah.py bounded`: (1) run n=21 N=4 `a144311_shared_flush 21 4 1 3` (predicted 1.6-2.4e8 nodes; budget ~1 h wall on the 4-CPU quota) and compare its node factor against the measured 3.42 (18->19) and 2.61 (19->20); (2) on the same 4-core handle run the n=20 N=1 pair (`20 1 1 0` unseeded and `20 1 232 0` seeded, the seeded one >=2216 s per #1539) to get the seeding factor at N=1, so the n=19-vs-n=20 comparison is made at fixed N. Report nodes, not wall.","compute":{"ram_gb":2,"disk_gb":1,"cpu_hours":0},"failure":"n=21 falls outside 1.6-2.4e8 nodes, or the N=1 n=20 seeding factor is also ~1.1x while n=19's was 2.872x at the same N, in which case the seeding benefit is decaying with level for a structural reason and every 79# price quoted from a node factor must be given as a range with the level attached.","success":"n=21 lands inside 1.6-2.4e8 nodes, fixing the per-level factor at two adjacent levels so the 79# price can be quoted from a measured range, and the N=1 n=20 pair shows whether the seeding collapse is a level effect or an N effect.","question":"Why does the seeding advantage collapse from 2.872x at n=19 to 1.117x at n=20, and does the per-level node factor keep falling at n=21?","budget_hours":3,"required_tools":[],"required_sources":[]},"depends_on":[1542,1531,1530],"evidence_md":"Ran the recorded next_step's arms (1) and (2) on the instrument served by #1550, built from the served\nbytes (flush source sha256 826c1599668464640ff59e3528b7cd2a3e7d9cb2fd3118f7f42acad02ec30433; shared\nsource d768a76d0533d42f82689b73a3f6b0a111350744ad6aab7c393c1e2cbbf09b3b — both local copies hash-match\nthe published shas). Ladder re-validated before the arms: ./a144311_shared 16 1 1 0 -> value 869, nodes\n876710 (exact #1531 match). Pre-registration work/prereg.md written before any arm. Each arm ran\ndetached under `sah.py bounded --limit 3300`; both exited 0, timed_out false, group_cleared true.\n\nARM (1) REPLICATE  a144311_shared_flush 20 4 1 3\n  SHARED n=20 N=4 mb=1 splitk=3 value=1397 best=232 nodes=69193209 tuples=135 records=25 leaves=1217 wall_ms=632609.9\n  workers 17204831/17243125/17057497/17687756 nodes\n  Reference #1542: value 1397, nodes 69174231, wall 822769.5 ms.\n  Delta = +18978 nodes = +0.0274 % -> far inside the ~1 % N=4 band. F1 did not fire.\n  FLUSH: best=206 @60 s, best=232 @120 s, frozen to 600 s (same shape as #1542).\n\nARM (2) SEEDED  a144311_shared_flush 20 4 232 3   (m_bound = (1397-5)/6 = 232)\n  SHARED n=20 N=4 mb=232 splitk=3 value=1397 best=232 nodes=61957778 tuples=135 records=0 leaves=4 wall_ms=571556.4\n  workers 15194188/15897194/15665870/15200526 nodes\n  FLUSH: best=232 already at t=60, records=0, leaves=2 -> 4 by t=540.\n  Seeding factor = 69193209/61957778 = 1.117x, vs 2.872x at n=19 (#1534: 24728579 -> 8603850).\n  P3's band [1.5,4.0] refuted low; F2 (seeded >= unseeded) did NOT fire. Wall cut 632.6 -> 571.6 s (-9.7 %).\n\nInterpretation: the plateau (best=232) is found within 60-120 s of a 632 s unseeded run, and handing the\nengine that plateau value in advance removes only ~11 % of the work at n=20, against ~65 % at n=19. So\nthe n=20 cost is enumeration-dominated: the exhaustive proof that no tuple beats 232 is what costs, and\nseeding cannot shortcut it. This gives #1542's falling per-level factor (3.42 -> 2.61) a mechanism and\nreinforces that a 79# price cannot be quoted from that factor alone; n=21 is required.\n\nNOT measured: any n=21 point, any N=1 n=20 point, any N=4/N=1 ratio at n=20, any machine-independent\nefficiency claim. Certified rung R=306, A144311(23) >= 1841, G_2(83#) >= 1842 and every prior return\nunchanged. 0.35 CPU-h used (two arms, 632 s + 572 s wall on the 4-CPU quota).","prior_art_md":"# Prior-art update — route 146, run-2026-09-23-ay (job 2942), search 2026-09-23T17:57Z\n\nSearched for the quantities this experiment measures — a *parallel* A144311 / covering-system engine,\nits **node counts**, wall times, or a seeding argument for this traversal, and any new OEIS term.\nNothing new since the route's 17:25Z record in #1544 and #1550's 17:43Z record.\n\n* **OEIS A144311** (read 2026-09-23T17:58Z): still `1, 5, 11, 29, 41, 65, 107, 149, 203, 257, 347,\n  527, 545, 617, 707, 869, 965, 1079, 1283, 1397, 1529, 1709`, only `a(1)..a(22)`, last extension\n  `a(17)-a(22)` by Jinyuan Wang, Nov 2024. **No `a(23)` is published**, so the route's `A144311(23) >=\n  1841` remains a lower bound and the first REFUTED `R = 307` (exactness) is untouched. The one linked\n  program is Jinyuan Wang's single-threaded C++; no parallel engine, node count or seeding argument is\n  published there.\n* **Web search** (2026-09-23T17:56Z, \"A144311 parallel engine covering system node count twin primes\"):\n  the only returned items are the 2016 threads (math.stackexchange 1779109 etc., traversal idea only),\n  the Wikipedia/MathWorld twin-prime pages, general Zhang/Maynard/Tao news, and one 2025 HAL preprint\n  \"On generating an infinite number of twin primes\" (hal-05371780) about `a(a+2)` coprime to a primorial\n  — an *existence* question, not this route's cost curve. The HAL full text was bot-gated (Anubis) and\n  not read; its abstract-level subject does not cover the node-count/seeding quantities measured here.\n* **Exact remaining gap (unchanged in kind):** the route's instrument is now served (#1550), closing the\n  provenance/access gap #1544 recorded. What remains open is the arithmetic: no n=21 point, no\n  N=4/N=1 ratio at n=20, no quotable 79# price from a single level factor, and exactness of\n  `A144311(23)`."},"research_route_id":146,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":null,"department_id":"dept_0e793a31e299699dfaaa6fee","run_id":"run_ca8519bc029a971665e9baa1","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"handle":"Benjaminsen","job_brief":"First update the online prior-work search for this experiment. If existing work covers it, record that and stop; otherwise run this bounded sprint on the uncovered uncertainty. Use cited published numbers during pursuit; their reproduction belongs in later validation. Build on the supplied findings; do not reconstruct earlier research. Return concrete progress and its cheapest credible check, a useful result for review, or a precisely scoped obstacle. Continued investment requires a distinct experiment.\n\nRead GET <project base>/research-routes/146 and return #1550. Return the ordinary report and transcript plus research: {route_id: 146, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.","review_deferred":false,"in_triage":false,"triage":[],"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[{"id":"1530","status":"recorded","final_rung":"recorded","canonical_return_id":null},{"id":"1531","status":"accepted","final_rung":"measured","canonical_return_id":null},{"id":"1542","status":"recorded","final_rung":"recorded","canonical_return_id":null}],"research_url":"/projects/twin-primes/research-routes/146","transcript_url":"/projects/twin-primes/return/1551/transcript","files":[],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}