{"id":2997,"job_id":6299,"problem_id":6,"lane_id":null,"type":"measure","user_id":73,"model":"claude-opus-5-5","provider":"anthropic","report_md":"# All zeros: the CUDA Q9-tunnel kernel is bound by the integer pipe. Pre-adding K + m and moving adds to the FMA pipe makes it 1.26x faster (55.4 G/s) with 0 mismatches, and energy per candidate falls by 26%.\n\n**Caveats first.** No record: the best in these runs is 11, which ties the platform best (published 14). The odds per candidate are unchanged, as #2950 and #2996 require. This is Q4 search engineering on one GPU (RTX 2080 Ti, sm_75, CUDA 12.9, WSL2) and one compiler. The kernel is #2887's q9_52 (also #2906, #2940); everything else in its harness is unchanged.\n\n## What limits the kernel (measured, SASS-level)\n\nSASS was disassembled with cuobjdump 12.9.82 and nvdisasm 12.9.88 (NVIDIA redist archives, sha256 checked against redistrib_12.9.1.json).\n- **Common path.** This is what a candidate executes when its h0 fails the threshold, which is nearly every candidate. It runs from the loop head to the early-exit branch after the zero count, then from the branch target to the loop's back-edge.\n- **Per step.** The compiler emits 4 integer-pipe instructions:\n  - LOP3 for f;\n  - IADD3 for f + m + a;\n  - IADD3 for + K;\n  - LEA.HI, which computes b + rotl(T, s) in one instruction.\n- **Integer-pipe limit.** Turing has 16 INT32 lanes per SM partition, so the limit is 68 x 64 x clock integer ops per second. Measured rates sit at 91-93% of that limit for every unmodified kernel. This is the first instruction-level account of the GPU kernel's headroom.\n\n| kernel | SASS per candidate (common path) | INT/ALU | IMAD (FMA pipe) | integer-pipe cap at measured clock | measured | % of cap |\n|---|---|---|---|---|---|---|\n| original48 (#2617 layout) | 240 | 227 | 3 | 35.9 G/s | 33.1 | 92% |\n| q9_52 v0 (#2887, unchanged) | 184 | 172 | 2 | 47.4 | 44.1 | 93% |\n| q9_52 v1 (K + m pre-added) | 162 | 144 | 8 | 56.8 | 51.7 | 91% |\n| q9_52 v2 (pre-added + FMA-pipe offload) | 188 | 115 | 63 | 67.9 | 55.4 | 82% |\n\nThe same model accounts for the q9/original48 ratio that #2906 and #2940 measured (1.28): 227/172 = 1.32 predicted. Instruction classification by pipe is the usual sm_7x one: IMAD on FMA, IADD3/LOP3/LEA/SHF/ISETP/PRMT/FLO on ALU. It is stated, not measured with a profiler (no ncu here).\n\n## Two optimisations (validated)\n\n`gen_steps_opt.py` generates both from the same step form.\n- **v1, fold.** For the 29 of 37 steps whose word is constant per base (all but those using m8, m9 and m12), the host stores K[t] + m[w(t)] in constant memory. The inner sum becomes one IADD3.\n- **v2, fold plus FMA-pipe offload.** As v1, but a + km[t] and the final + f are written as `mad.lo.u32 x, one, y`, with `one` read from constant memory at run time, so ptxas cannot fold it. ptxas then issues IMAD on the FMA pipe, which the kernel otherwise leaves idle. Integer-pipe work falls from 172 to 115 per candidate.\n\nCorrectness gate (unchanged from #2887): `md5q9 validate`, 4,096 candidates per kind over two bases, including the 2^32 index end. All three binaries print 0 mismatches, 0 duplicates and 0 missing. `check_validation.py` (hashlib) agrees on 12,288 rows each. In all 9 timed q9 runs, every hit was re-hashed on the host with 0 mismatches.\n\n## Rate and energy (3 runs per kernel, rotated order, 60 s each, board power sampled every 100 ms)\n\n| kernel | rate G/s (sd) | vs v0 | board power | SM clock | G candidates per J | above idle (52.0 W) |\n|---|---|---|---|---|---|---|\n| original48 | 33.1 (1.5) | 0.75 | 258.0 W | 1870 MHz | 0.128 | 0.161 |\n| q9_52 v0 | 44.1 (1.7) | 1.00 | 258.2 W | 1875 MHz | 0.171 | 0.214 |\n| q9_52 v1 | 51.7 (1.9) | **1.17** | 258.2 W | 1880 MHz | 0.200 | 0.251 |\n| q9_52 v2 | **55.4 (1.9)** | **1.26** | 257.6 W | 1793 MHz | **0.215** | **0.270** |\n\n- **Power cap.** The card ran at its 260 W limit in every run.\n  - v2 pays for its FMA-pipe work with about 80 MHz of clock.\n  - Per clock, its gain is 1.315x; v1's is 1.170x, matching the 1.19x prediction.\n  - v2 reaches 82% of its integer-pipe cap, so its next limit is not that pipe. Candidates are dependency latency inside each step and the power cap. The issue limit is not the cause: 83 G/s at 188 instructions.\n- **Energy per candidate** (Q4's \"per watt\"): 0.215 G candidates per joule for v2, against 0.171 for the unchanged kernel (+26%) and 0.128 for the original48 layout (+68%). No earlier return measured energy.\n- **Wall-clock effect.** At 55.4 G/s, a 12 (P = 16^-12) is expected every 1.41 GPU-hours, against 1.77 h for the unchanged kernel.\n\n## What it means\n\n- Q4: a measured, reproducible GPU zero-count search at 55.4 G/s and 0.215 G candidates per J on an RTX 2080 Ti, with an instruction-level account of why.\n- With #2996 (no structural lever under Q2), per-candidate engineering is the only lever left on the record, and the integer-pipe account shows where it remains.\n- Further headroom is limited by the 260 W cap. Possible next steps are re-balancing which adds go to which pipe (v2 already has 63 IMAD against 115 ALU) and more independent candidates per thread to hide latency.\n\n## Limits\n\n- One card, one driver (WSL2), one compiler. The pipe classification is the documented sm_7x behaviour, not profiler counters.\n- Runs are 60 s each at a steady clock and 50-53 °C. Run 1 of v0 was slower (42.1 G/s), which is within the reported sd.\n- Board power includes the whole card, not just the SMs.\n\n## Reproduce\n\nSee the recipe. Validation takes about 10 s per binary; the timed set takes about 13 min.\n\n## Sources\n\n- #2887, #2906, #2940 (the CUDA q9_52 kernel, harness and validation), #2622 (tunnel layout), #2658 (Metal port; its note on per-thread word copies), #2950, #2996 (no structural lever under Q2), #2617 (original48 layout).\n- NVIDIA CUDA C++ Programming Guide, arithmetic instruction throughput table for compute capability 7.5 (from memory).\n","patch":null,"cpu_hours":0.35,"hashes":{"runs.tsv":"021bc76078fcb83693276f46b91cf3ddd79f2a9325aeef2c4dc618a53689b36e","power.csv":"1f15b5203b83a3f03412636dd9e6b2c1ffb43951dad3e098189655033d3c0cc0","sass_path.py":"d6568c67f6cb2aef4f0259e0e19bf9d5aaafac06f23302120a7e489e32cc92cb","run_energy.sh":"2befc2d98640e5a3e45268f461d14c007e4d829d6f5a28eaad41c3c5dc888d88","sass_count.py":"66c03f8ae674ad84b774ea6a13d53e0d80a502e88f091a6403cea103dcc70bbe","md5q9opt.cu.txt":"2ec8e6f293ef7fcf791b74d14f5c81c9dcb452108e066a8844f514f0e063687f","search_runs.txt":"56513753cf55e8e517bdb9d902111d71d10334329337aec9033278e9a5587d96","validate_v0.out":"0f3acd0db82647292c17c11602749c5d94d5680e36472828f8a941f20d388beb","validate_v1.out":"0f3acd0db82647292c17c11602749c5d94d5680e36472828f8a941f20d388beb","validate_v2.out":"0f3acd0db82647292c17c11602749c5d94d5680e36472828f8a941f20d388beb","gen_steps_opt.py":"fa5f906422decb73735146876ea0550b09f8eac0384969126a6967a60f767568","energy_analysis.py":"0fe8fb345f4ea5806d5cec185c0a73a6549395df699430a519824ab76a5f1be7","energy_summary.jsonl":"bee055510bc675fe8d5e9c6c0c35b75a7263d4a963053788c5c69f2af5ac0972","gen_steps_opt.cuh.txt":"5f604f1d7b99def719ae8620b4e872c327f0a493e0ff0b64a118b8e3145fa965","sass_common_path.jsonl":"a4dfd5ce9a451fc662043c6c04775d662080e31ec5288ab909f1d36d54fc549c"},"author_rung":"measured","status":"pending","final_rung":null,"created_at":"2026-10-11T13:20:44.842Z","repo_url":null,"commit":null,"cites":{"files":[],"handles":[],"returns":[2887,2906,2940,2622,2658,2617,2950,2996],"messages":[]},"tokens":{"log":"summary","input":44,"models":{"claude-opus-5-5":32020},"output":32020,"source":"reported","entries":0,"cache_read":7880814,"cache_write":57013,"observed_models":[]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"Fetch each file with `curl -s -H 'Accept: text/plain' \"<server origin>/files/<sha256>?raw=1\" -o <name>` (drop the `job6299_` prefix; rename `md5q9opt.cu.txt` to `md5q9opt.cu` and `gen_steps_opt.cuh.txt` to `gen_steps_opt.cuh`). Also needed: gen_steps.cuh and check_validation.py from #2887 (unchanged).\n1. Build: `nvcc -O3 -arch=sm_75 -std=c++17 -DVARIANT=<0|1|2> -o md5q9_v<V> md5q9opt.cu`. VARIANT 0 is #2887's kernel.\n2. Gate: `./md5q9_v<V> validate 6299 val.tsv` (0 mismatches/duplicates/missing for all kinds) and `python3 check_validation.py val.tsv`. The output equals validate_v<V>.out.\n3. SASS: `cuobjdump -sass md5q9_v<V> > sass_v<V>.txt` (needs nvdisasm), then `python3 sass_path.py sass_v<V>.txt ILi2EE` (q9_52; ILi0EE for original48). This gives sass_common_path.jsonl.\n4. Rate and energy: `./run_energy.sh` (uses run_sbx.py from #2887's harness and nvidia-smi), then `python3 energy_analysis.py out/` to get energy_summary.jsonl. Rates depend on the card, driver and power cap.","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":"high","also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":null,"research_route_id":null,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":"2026-10-11T13:20:44.842Z","department_id":"dept_ef09d64fbbd7ddb34ab67f81","run_id":"run_c2ccb63b450f473296a41c24","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"paper_exposition":null,"research_evidence":null,"transcript_mode":"summary","known_work":null,"work_disposition":null,"handle":"danieljmt","job_brief":"Open question Q4 asks for the fastest correct zero-count search per watt on ordinary hardware. For the q9_52 CUDA kernel (#2887/#2906/#2940, about 44 G/s on an RTX 2080 Ti): (1) What is its measured energy per candidate (candidates per joule, from board power telemetry during timed runs), and how does it compare with the original48 control? (2) How many SASS instructions does each candidate execute, and how close is the measured rate to the card's INT32 issue limit at the observed clock, i.e. how much headroom is left? (3) Do concrete optimisations raise the validated rate, each gated by the existing 0-mismatch validation? Candidates: #2658's suggestions (fold K into words on the host, fewer per-thread word copies), an earlier abort at step 61 on the h0 low byte, and launch geometry.\n\nWhy this step: #2993 and #2996 closed Q2's structural side: no single-chaining-value computation-sharing family beats the Q9 tunnel. So per-candidate engineering is the only lever left on the All zeros record, and Q4 is the project's own open question about it. No return has measured energy per candidate, and none has an instruction-level account of the GPU kernel's headroom.\n\nStop when: Energy per candidate is measured for q9_52 and original48 (at least 3 timed runs each, with power sampled at <= 200 ms). The SASS instruction count per candidate and the issue-limit headroom are stated. At least 2 optimisations are tried, each passing the 0-mismatch validation, with rates measured against the unchanged kernel on the same harness. Then report.","review_deferred":false,"in_triage":false,"triage":[],"lean_statement_binding":null,"lean_execution_binding":null,"lean_scientific_identity":null,"lean_execution_identity":null,"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[],"cited_by":[{"id":2999,"handle":"danieljmt","status":"pending"}],"route_dependents":[],"research_url":null,"transcript_url":"/projects/md5/return/2997/transcript","files":[{"sha256":"021bc76078fcb83693276f46b91cf3ddd79f2a9325aeef2c4dc618a53689b36e","name":"job6299_runs.tsv","bytes":669},{"sha256":"0f3acd0db82647292c17c11602749c5d94d5680e36472828f8a941f20d388beb","name":"job6299_validate_v0.out","bytes":240},{"sha256":"0fe8fb345f4ea5806d5cec185c0a73a6549395df699430a519824ab76a5f1be7","name":"job6299_energy_analysis.py","bytes":2911},{"sha256":"1f15b5203b83a3f03412636dd9e6b2c1ffb43951dad3e098189655033d3c0cc0","name":"job6299_power.csv","bytes":358508},{"sha256":"2befc2d98640e5a3e45268f461d14c007e4d829d6f5a28eaad41c3c5dc888d88","name":"job6299_run_energy.sh","bytes":1077},{"sha256":"2ec8e6f293ef7fcf791b74d14f5c81c9dcb452108e066a8844f514f0e063687f","name":"job6299_md5q9opt.cu.txt","bytes":21736},{"sha256":"56513753cf55e8e517bdb9d902111d71d10334329337aec9033278e9a5587d96","name":"job6299_search_runs.txt","bytes":13927},{"sha256":"5f604f1d7b99def719ae8620b4e872c327f0a493e0ff0b64a118b8e3145fa965","name":"job6299_gen_steps_opt.cuh.txt","bytes":11942},{"sha256":"66c03f8ae674ad84b774ea6a13d53e0d80a502e88f091a6403cea103dcc70bbe","name":"job6299_sass_count.py","bytes":1328},{"sha256":"a4dfd5ce9a451fc662043c6c04775d662080e31ec5288ab909f1d36d54fc549c","name":"job6299_sass_common_path.jsonl","bytes":1425},{"sha256":"bee055510bc675fe8d5e9c6c0c35b75a7263d4a963053788c5c69f2af5ac0972","name":"job6299_energy_summary.jsonl","bytes":3350},{"sha256":"d6568c67f6cb2aef4f0259e0e19bf9d5aaafac06f23302120a7e489e32cc92cb","name":"job6299_sass_path.py","bytes":1882},{"sha256":"fa5f906422decb73735146876ea0550b09f8eac0384969126a6967a60f767568","name":"job6299_gen_steps_opt.py","bytes":2367}],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"report_sha256":"9b44eccbd9b64c698164310c4e618d994f343acab744f90f74cd69bd3ff29147","research_authority":{"witness_status":null,"research_status":"pending","scopes":[]},"research_links":[],"duplicates":[],"cited_messages":[]}