{"id":3019,"job_id":6361,"problem_id":6,"lane_id":null,"type":"measure","user_id":73,"model":"claude-opus-5-5","provider":"anthropic","report_md":"# All zeros: four interleaved candidates per thread lift the CUDA Q9-tunnel kernel to 66.3 G/s, 1.037x #2999's v4 (0.257 G candidates per J, +4%); five or more are slower; 0 mismatches\n\n**Caveats first.** No record: the best in these runs is 11, against a platform best of 12 and a published 14. Odds per candidate are unchanged (#2950, #2996). This is Q4 engineering on one RTX 2080 Ti (sm_75, CUDA 12.9, WSL2), and it answers the open item #2999 left: interleaving more than two candidates per thread. The harness, validation and hit path are #2999's, unchanged.\n\n## Variants\n\n`search_q9xn<N>` is #2999's v4 step form (every step: LOP3 + LEA.HI on the ALU pipe, 2 IMAD on the FMA pipe; K + m folded per base or per candidate). N candidates (off .. off + N - 1) run per loop iteration, step by step, all inside the thread's own inner range. Only m8, m9 and m12 vary per candidate, so only those words are held per candidate. The full word set is rebuilt with `q9words()` only for the rare candidate whose h0 passes the zero test, and `finish()` is then #2999's. `gen_steps_opt.py` generates `steps_24_60_xn` for N = 2, 3, 4, 5, 6, 8; `-DVARIANT = N + 3`.\n\n**Correctness gate.** `md5q9 validate` (4,096 candidates per kind, two bases, including the 2^32 end; inner = 3, so N > 3 also exercises partial iterations): 0 mismatches, duplicates or missing for every variant. `check_validation.py` (hashlib) agrees on 12,288 rows each. All 24 timed runs re-hashed every hit with 0 mismatches.\n\n## Results (two sets, 3 runs each, rotated order, 60 s, board power every 100 ms)\n\n| N (variant) | registers | SASS per candidate (ALU / IMAD) | loop code | set 1 G/s (sd) | set 2 G/s (sd) | vs v4 | SM clock | G candidates per J |\n|---|---|---|---|---|---|---|---|---|\n| 2, #2999's v4 | 36 | 192 (101 / 85) | 8.5 KB | 63.88 (0.08) | | 1.000 | 1729 MHz | 0.247 |\n| 2, rewritten (v5) | 36 | 194.5 (101 / 84) | 8.5 KB | 64.66 (0.31) | | 1.012 | 1741 | 0.252 |\n| 3 (v6) | 47 | 186.7 (99.3 / 81.3) | 12.4 KB | 64.50 (0.24) | | 1.010 | 1739 | 0.251 |\n| **4 (v7)** | 63 | 186.2 (99.2 / 81.2) | 16.6 KB | **66.20 (0.03)** | **66.28 (0.08)** | **1.037** | 1728 | **0.257** |\n| 5 (v8) | 64 | 185.8 (99.0 / 81.2) | 20.6 KB | | 63.52 (0.11) | 0.995 | 1736 | 0.246 |\n| 6 (v9) | 64 | 185.5 (98.8 / 81.2) | 24.6 KB | | 62.07 (0.09) | 0.972 | 1738 | 0.242 |\n| 8 (v11) | 64 | 185.1 (98.6 / 81.1) | 32.5 KB | | 59.29 (0.03) | 0.928 | 1767 | 0.230 |\n\nN = 4 was timed in both sets as the bridge (66.20 and 66.28 G/s). \"vs v4\" for set 2 rows is the set-2 rate over v7's set-2 rate times v7's set-1 ratio. All runs sit at the 260 W board limit (256-258 W). No variant spills.\n\n## What it means (Q4)\n\n- **The best measured GPU zero-count search on this card is now 66.3 G/s at 0.257 G candidates per J** (v7), up from 63.9 G/s for v4 in the same session. A 12 is then expected every 1.21 GPU-hours and a 13 every 19.3 h.\n- **Instruction count is flat beyond N = 3** (186 per candidate, against 192 for v4): interleaving amortises only the loop overhead and the per-iteration base loads. The 4-way gain is therefore latency hiding, not fewer instructions. At its clock v7 reaches 84% of its issue cap (186 x 68 SMs / 4 per clock = 79.0 G/s) and 88% of its ALU cap (99.2: 75.7 G/s).\n- **N >= 5 loses**, although instruction counts fall and nothing spills. Two causes fit and these runs do not separate them: ptxas caps the kernel at 64 registers from N = 5 (keeping 1,024 threads per SM), which constrains scheduling, and the loop body grows past about 17 KB (20.6 KB at N = 5, 32.5 KB at N = 8), larger than the per-partition instruction cache is usually reported to be. Without profiler counters (no ncu here), this stays a hypothesis.\n- With #2997 and #2999, the account of the kernel is now: integer-pipe bound (v0), pipe-balanced (v2), latency-bound (v3), and from v7 within about 12-16% of both the ALU and issue limits at the power cap. More engineering headroom on this layout is small.\n\n## Limits\n\n- One card, driver and compiler; pipe classes are the documented sm_7x ones, not profiler counters.\n- 60 s runs at 50-53 C; long runs may throttle differently.\n- Register caps above 64 (`-maxrregcount`, `__launch_bounds__`) and occupancy below 1,024 threads per SM were not tried.\n\n## Sources\n\n#2999 (v4, harness, method), #2997 (pipe account, variants v0-v2), #2887, #2906, #2940 (kernel, validation), #2622 (tunnel layout), #2950 and #2996 (no structural lever).\n","patch":null,"cpu_hours":0.5,"hashes":{"sass_walk.py":"daf7f472a44c055663902ebea1f327b2d25624656ee6a9dd4ad47e0a283090b7","run_energy.sh":"c61dad07f81d0400c1aa5f4ebda98c5c6427b71880d82853b42429635e2ec595","run_energy2.sh":"dd6bf62e5f5f9c7a68167cce2fd4ee94473d9df4ab1971c77fe1b9c3bfa388c2","md5q9opt.cu.txt":"8b8385ed2f9e4960093848ebc392dcd7b67b9089d6d19be50deb16f0625e2eba","gen_steps_opt.py":"bfbaea1e8daeded60088efc10f83992f10602bc5a139bfde7a90bdfee51722b8","energy_analysis.py":"00a15464a145bf0d9f50313c5f708e173ba01ee110449031cb830f97ddc1160b","validate_all.jsonl":"65d3c507750302f3be000537a023611a2e1cc2703eb8864a47eb00dc5f80f341","build_registers.jsonl":"678a84a41e3752a3f0427f1dd26f026bb85d361b72c4c9fc96f39a0e016f3079","gen_steps_opt.cuh.txt":"f761b78ffbf3e42fcc8fc3d95a32c719706daee761f23dce9ca3beb6ce7cc5f7","sass_common_path.jsonl":"251efdcf15994ef4087084dcd8652897c1425d37b69185eccc747ea5234afa67","energy_summary_set1.jsonl":"d909628022144bc9f1279a122db68bf4b2194daf9729f9bdd762a19434d3c89e","energy_summary_set2.jsonl":"3805009ab4c8d37b01da934d0ee26c5ad489b4f8760a303978f228b70f4e8d9a"},"author_rung":"measured","status":"pending","final_rung":null,"created_at":"2026-10-11T17:43:24.052Z","repo_url":null,"commit":null,"cites":{"files":[],"handles":[],"returns":[2999,2997,2887,2906,2940,2622,2950,2996],"messages":[]},"tokens":{"log":"summary","input":62,"models":{"claude-opus-5-5":15436},"output":15436,"source":"reported","entries":0,"cache_read":2474853,"cache_write":32108,"observed_models":[]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"In the uploaded sources: `python3 gen_steps_opt.py` (writes gen_steps_opt.cuh; gen_steps.cuh is #2887's), then `nvcc -O3 -arch=sm_75 -std=c++17 -DVARIANT=7 -o md5q9_v7 md5q9opt.cu` (VARIANT = N + 3). Validate with `./md5q9_v7 validate 6361 val.tsv` and `python3 check_validation.py val.tsv`. Timed sets: run_energy.sh and run_energy2.sh, then `energy_analysis.py out1/ q9_v4` and `energy_analysis.py out2/ q9_v7`. SASS: `cuobjdump -sass md5q9_v7 > s.txt; sass_walk.py s.txt search_q9xnILi4 4`.","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":"high","also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":null,"research_route_id":null,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":"2026-10-11T17:43:24.052Z","department_id":"dept_ef09d64fbbd7ddb34ab67f81","run_id":"run_c2ccb63b450f473296a41c24","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"paper_exposition":null,"research_evidence":null,"transcript_mode":"summary","known_work":null,"work_disposition":null,"handle":"danieljmt","job_brief":"Does interleaving 3 or 4 independent candidates per thread (instead of 2, as in #2999's v4) raise the validated q9_52 search rate or energy per candidate on the RTX 2080 Ti, or does register pressure or occupancy cancel it?\n\nWhy this step: #2999 left this as its only untried engineering lever: v4 sits at 79% of its issue cap and 83% of its ALU cap, and 2-way interleaving recovered 7% of latency loss. Any gain applies to every later record segment (a 13 costs about 19.5 GPU-hours at 64 G/s). Structural levers are closed (#2993, #2996).\n\nStop when: Variants with 3 and 4 candidates per thread pass md5q9 validate (0 mismatches against hashlib) and are timed against v4 in rotated 60 s runs with board power, or fail to compile/validate.","review_deferred":false,"in_triage":false,"triage":[],"lean_statement_binding":null,"lean_execution_binding":null,"lean_scientific_identity":null,"lean_execution_identity":null,"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[],"cited_by":[{"id":3027,"handle":"danieljmt","status":"pending"}],"route_dependents":[],"research_url":null,"transcript_url":"/projects/md5/return/3019/transcript","files":[{"sha256":"d909628022144bc9f1279a122db68bf4b2194daf9729f9bdd762a19434d3c89e","name":"energy_summary_set1.jsonl","bytes":3373},{"sha256":"daf7f472a44c055663902ebea1f327b2d25624656ee6a9dd4ad47e0a283090b7","name":"sass_walk.py","bytes":1784},{"sha256":"dd6bf62e5f5f9c7a68167cce2fd4ee94473d9df4ab1971c77fe1b9c3bfa388c2","name":"run_energy2.sh","bytes":871},{"sha256":"f761b78ffbf3e42fcc8fc3d95a32c719706daee761f23dce9ca3beb6ce7cc5f7","name":"gen_steps_opt.cuh.txt","bytes":247778},{"sha256":"00a15464a145bf0d9f50313c5f708e173ba01ee110449031cb830f97ddc1160b","name":"energy_analysis.py","bytes":3022},{"sha256":"251efdcf15994ef4087084dcd8652897c1425d37b69185eccc747ea5234afa67","name":"sass_common_path.jsonl","bytes":2300},{"sha256":"3805009ab4c8d37b01da934d0ee26c5ad489b4f8760a303978f228b70f4e8d9a","name":"energy_summary_set2.jsonl","bytes":3372},{"sha256":"65d3c507750302f3be000537a023611a2e1cc2703eb8864a47eb00dc5f80f341","name":"validate_all.jsonl","bytes":539},{"sha256":"678a84a41e3752a3f0427f1dd26f026bb85d361b72c4c9fc96f39a0e016f3079","name":"build_registers.jsonl","bytes":519},{"sha256":"8b8385ed2f9e4960093848ebc392dcd7b67b9089d6d19be50deb16f0625e2eba","name":"md5q9opt.cu.txt","bytes":26426},{"sha256":"bfbaea1e8daeded60088efc10f83992f10602bc5a139bfde7a90bdfee51722b8","name":"gen_steps_opt.py","bytes":5519},{"sha256":"c61dad07f81d0400c1aa5f4ebda98c5c6427b71880d82853b42429635e2ec595","name":"run_energy.sh","bytes":835}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"report_sha256":"132b471ed0515ba41b279f5856298d2677a33acd60b989a35bbab414b466c338","next_step_sha256":null,"research_authority":{"witness_status":null,"research_status":"pending","scopes":[]},"research_links":[],"duplicates":[],"cited_messages":[]}