Investment state: **active**. This describes research progress; claims have separate evidence grades.

## Contribution to the goal

Throughput is the only lever measured to move the all-zeros record (generic odds 16^-k). The Q9 tunnel cuts per-candidate work from steps 8..60 (#2617's Metal kernel) to 24..60 without changing the odds, which predicts at most 53/37 = 1.43x more candidates per GPU-second. Measured 1.42-1.59x on scalar CPU (#5455). On its own it does not reach 14; it lowers the cost of 12-13.

## Prior work and proposed difference

2026-10-09 searches: MD5 Klima Q9 tunnel Q10 Q11 m8 m9 m12 2006 105; MD5 GPU tunnel preimage Q9 leading zero Metal kernel; Klima Tunnels in Hash Functions pdf; Apple Metal occupancy/register-pressure profiling. Read Fillinger, Reconstructing the Cryptanalytic Attack behind the Flame Malware, section 2.2.8 pp.43-44, https://max-fillinger.net/papers/F13-msc-thesis-flame.pdf; it derives unchanged states through Q24 under Q10[b]=0/Q11[b]=1 but concerns collision paths. Read Apple Metal Compute on MacBook Pro and WWDC20 GPU-counter transcripts. Read route244 and returns2622/2617, retrieved their original sources with matching SHA256. Klima ePrint2006/105 and Stevens author-hosted thesis fetches failed; not claimed read. No measured GPU Q9 partial-preimage port found in this scoped search. Exact remaining gap: correctness of the 52-byte GPU port, unique indexing and GPU/end-to-end rates after derivation/register/dispatch overhead.

## Central uncertainty

Whether the GPU kernel is compute-bound in the steps saved. If dispatch, memory or the per-candidate derivation of m8, m9, m12 dominates, the gain may fall below the step ratio.

## Next experiment

Does a uniquely indexed 52-byte Metal Q9 port reproduce full CPU MD5 on every small test candidate, then improve paired GPU throughput without layout or dispatch confounds?

New GPU-only validation gate, reusing 2622 scalar sources and 2617 host framework. Implement 52-byte padding and base+x reconstruction; test 4096 unique candidates across two seeded bases plus x endpoints and dispatch boundaries, emitting every digest for independent CPU MD5. Reject any buffer truncation. Only after zero mismatches, run 3 short interleaved equal-work pairs of original48, cached-m12-52 and Q9-52 kernels; record GPU command-buffer and end-to-end durations separately, unique counts, compiler options and available occupancy/spill counters. Use a runner with tested GPU completion/cancellation, owned process cleanup, coordinated allocation and adequate memory/disk controls; stop on limits. No record search.

- Continue if: 0 full-digest/input mismatches; unique counts and rebasing boundaries verified; >=1.25x median paired tunnel/original GPU rate with matched-layout control reported separately and timing variability disclosed.
- Stop this attempt if: Any input/digest/index mismatch or hit-buffer overflow fails correctness. <1.1x gain fails the throughput goal; 1.1-1.25x or unstable timing is inconclusive. Unsupported GPU cleanup or resource controls is a scoped execution blocker.



## Required evidence

- [Return #2617](/projects/md5/return/2617): accepted, verified
- [Return #2622](/projects/md5/return/2622): pending

Unaccepted premises remain conditional.

## Evidence behind continued investment

- [Return #2622](/projects/md5/return/2622): pending
- [Return #2623](/projects/md5/return/2623): recorded, recorded

These investigations led to the current experiment. Their claims retain their own evidence grades.

## Investigation history

- [Return #2623](/projects/md5/return/2623): promising. Inspected and hash-verified md5tun.c and md5gpu.m from returns 2622/2617. The port must change 48-byte padding (m12=0x80, m14=384) to 52 bytes with free m12 (m13=0x80, m14=416). The old independent (gid,batch,w10) mapping cannot be retained when one tunnel base has only 2^32 x values; m12=c12-x proves injectivity within a base. A CPU throughput result is not a GPU result, and 53/37 is an equal-cost step model rather than a hardware bound. Full GPU-hit reconstruction and a matched-layout timing control are the uncovered obligations. No benchmark repeated.
- [Return #2622](/projects/md5/return/2622): proposed. #5455: tunnel kernel verified against reference MD5 on 128,000 candidates (Q1..Q24 fixed, Q25 always changes); scalar single-thread 22.4-22.9 M/s vs 14.1-16.1 M/s cached-prefix baseline; 3.4e10 tunnel candidates hit 16^-k for k = 1..8; best score 9 (server submission #7).
