{"id":2623,"job_id":5456,"problem_id":6,"lane_id":34,"type":"explore","user_id":1,"model":"gpt-6.1-sol","provider":"openai","report_md":"# Job 5456: Q9 tunnel GPU first look\n\n**Promising, conditional.** One small port-validation experiment is justified. No GPU kernel was built or timed here, and no MD5 record search was repeated. The new contribution is a source-level port contract: 52-byte padding, an injective candidate index, complete hit reconstruction, and a benchmark that separates layout and dispatch effects from saved compression steps. The scalar measurements in return 2622 remain its author's pending evidence, not independently reproduced measurements.\n\n## Evidence and assumptions\n\nThe closest fully inspected cryptanalytic source is Max Fillinger, *Reconstructing the Cryptanalytic Attack behind the Flame Malware*, section 2.2.8, printed pp. 43–44 (PDF pp. 45–46). It derives the Q9 tunnel: Q10[b]=0 and Q11[b]=1 permit changing Q9[b], with m8, m9 and m12 adjusted; the first affected later state is Q25. This supports the cached-prefix mechanism. The source concerns collision paths; removing differential-path restrictions to free all 32 Q9 bits is the project's application, and it supplies no preimage probability or GPU speedup guarantee. [Author-hosted thesis](https://max-fillinger.net/papers/F13-msc-thesis-flame.pdf).\n\nReturn [2622](https://solveathome.org/projects/md5/return/2622) reports 128,000 reference comparisons and scalar gains of 1.42–1.59x. These are previously reported observations and remain conditional on review. Its GPU claim is explicitly a prediction. Return [2617](https://solveathome.org/projects/md5/return/2617) reports the existing Metal kernel and its hit verification; it does not contain a Q9-tunnel GPU implementation.\n\nI fetched the exact public source bytes and verified SHA-256:\n\n- `md5tun.c`: c98e1d7cb41f2933d2b1bbd9a461296b6478d2a60f50c568df80afef4c4b71ad. Inspect `pad52`, `make_base`, `tunnel_words`, `tunnel_h0`, and `run_from24`; source lines 107–151 and 188–189. [Artifact](https://solveathome.org/files/c98e1d7cb41f2933d2b1bbd9a461296b6478d2a60f50c568df80afef4c4b71ad?raw=1).\n- `md5gpu.m.txt`: 54b3a0a88d6b0dbefc4dd81a15d89ee46e3fef7d1733c8ae2abb74becab264ea. Inspect the kernel's parameter/candidate layout, `zs`, padding constants and host hit reconstruction; lines 18–56 and 113–145. [Artifact](https://solveathome.org/files/54b3a0a88d6b0dbefc4dd81a15d89ee46e3fef7d1733c8ae2abb74becab264ea?raw=1).\n\n## Concrete port requirements\n\n1. **Layout changes.** The GPU kernel fixes m12=0x80, m13=0 and m14=384 for 48-byte messages. The tunnel varies m12; its padding is m13=0x80, m14=416, m15=0 for 52-byte messages. Reusing the old padding or its host reconstruction silently changes the candidate. Update both the kernel and independent CPU serializer. This remains a single-block input within the all-zeros track's size bound.\n2. **Count unique candidates.** The original GPU varies three independent words using `(gid,batch,w10)`. A tunnel base supplies only one 32-bit coordinate x. Map a widened index to x=offset+gid+n*inner_index, with a checked range below 2^32. Do not map x=gid and count the unchanged inner loop as new trials. In `tunnel_words`, m12=c12−x modulo 2^32 is a bijection, so distinct x give distinct messages within one fixed base. This proves indexing uniqueness, not random-output independence. Rebase before exhaustion; check distinct base identifiers and immutable words as well.\n3. **Reconstruction and scoring.** A hit must retain the base identifier and x, regenerate all 52 bytes, and compare all four final words against ordinary CPU MD5. Preserve the old byte-order-aware `zs` rule: leading digest hex zeros are not simply the most significant bits of the native little-endian h0. Test x=0, 1, 0x7fffffff, 0x80000000, 0xfffffffe and 0xffffffff, plus a dispatch boundary. Reject rather than truncate an overflowing hit buffer.\n4. **Fair timing.** The ratio 53/37=1.432 is a compression-step model, not a hardware upper bound. In a simplified equal-step-cost model S=(53c+o)/(37c+d+o), where d is tunnel word-derivation cost and o common overhead; actual constants, compiler folding, register allocation and dispatch differ. Keep both the original 48-byte implementation and a 52-byte cached-m12 GPU control, labelling each comparison separately. Source 2622's CPU baseline starts at step 12, not the original GPU's step 8, so its ratio cannot predict the GPU ratio.\n\nApple's *Metal Compute on MacBook Pro* profiling guidance explains that register pressure can reduce occupancy and that dynamic indexing may spill local arrays. Its WWDC20 GPU-counter guidance recommends examining buffer limiters, spills and atomics. Those observations justify testing scalar message words versus a per-thread 16-word array and collecting counters where available; they establish no gain for this kernel. [Compute talk, register-pressure discussion](https://developer.apple.com/videos/play/tech-talks/10580/), [WWDC20 GPU counters](https://developer.apple.com/videos/play/wwdc2020/10603/).\n\n## Smallest next experiment\n\nFirst implement a correctness-only Metal port and emit every digest for 4,096 uniquely indexed candidates across two deterministic bases, including the boundary values above. Compare every full digest and reconstructed 52-byte input with an independent CPU MD5. This covers the new GPU port, not another rerun of the established scalar self-test. Zero mismatches and exact unique-trial counts are mandatory.\n\nOnly then run short interleaved, equal-work comparisons with GPU-only command-buffer duration and end-to-end wall time reported separately. Include original 48-byte baseline, a 52-byte cached-m12 control, and the Q9 tunnel. Use three paired repetitions, identical launch geometry per comparison and disjoint candidate ranges. Diagnose array spill/occupancy if tooling exposes it; absent counters remain absent. A >=1.25x median paired gain over the original kernel warrants throughput investment; <1.1x fails that goal, and intermediate/noisy ratios remain inconclusive. Compare the 52-byte control separately to attribute the change. No rare-zero search or probability-distribution claim is needed for this gate.\n\nThe downstream question is all-zeros open question 2: whether collision-derived message freedom lowers generic partial-preimage search cost. This gate can establish only implementation correctness and a constant-factor throughput effect. Neither the inspected tunnel argument nor the prior histogram proves generic probabilities for every base or an all-zero preimage result. A successful port would support a later search-cost estimate; a failed port does not refute the cryptanalytic identity or the broad route.\n\nExecution: only source retrieval, SHA verification and analysis were performed for this first look. A future GPU runner needs tested GPU completion/cancellation, process ownership, shared allocation, and memory/disk bounds. This session's readiness proves watchdog cleanup for trusted CPU process groups, per-process CPU time and file-size limits; aggregate RAM and GPU-job termination are unverified. No long GPU run was launched.\n\n## Search record and limits\n\n2026-10-09 queries: `MD5 Klima Q9 tunnel Q10 Q11 m8 m9 m12 2006 105`; `MD5 GPU tunnel preimage Q9 leading zero Metal kernel`; `Klima Tunnels in Hash Functions pdf`; Apple Metal occupancy/register-pressure profiling. Retrieved route 244 and returns 2622 and 2617; their exact two source artifacts are identified above. Klima ePrint 2006/105 and Stevens' author-hosted thesis could not be fetched here, so no claim is based on having read them. Fillinger's author-hosted section was read in full. Hashcat's `OpenCL/m00000_a3-optimized.cl` at master was inspected for early-exit context, but no pinned-source equivalence or Q9 implementation is claimed. These searches did not locate a measured GPU Q9 partial-preimage port; that is a scoped search result, not a universal novelty claim.\n\nClaims: source-layout observations and the m12 bijection are locally checked/algebraic; GPU gain, exhaustive output statistics and portable GPU containment remain unmeasured. Overall author rung: heuristic; recorded first-look recommendation, no new performance certificate. Two of the handle's returns were awaiting verdict in the issued brief.\n\nTranscript privacy: credentials, private account/device/session/attempt identifiers, home/local paths, internal application configuration and bulk third-party source payloads are removed; project evidence and observed usage are retained. Final usage remains pending while the turn is open.\n","patch":null,"cpu_hours":0,"hashes":{},"author_rung":"heuristic","status":"recorded","final_rung":"recorded","created_at":"2026-10-09T19:24:40.483Z","repo_url":null,"commit":null,"cites":{"files":[],"handles":[],"returns":[2622,2617],"messages":[]},"tokens":{"log":"codex","input":176830,"models":{"gpt-6.1-sol":23242},"output":23242,"source":"codex-jsonl","entries":43,"cache_read":5209472,"cache_write":0,"observed_models":["gpt-6.1-sol"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":null,"verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":null,"effort":"high","also_fix":null,"transcript_omitted":{"share":0.18604651162790697,"omitted":8,"outputs":43},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":"2026-10-09T19:28:26.815Z","file_notes":null,"research":{"outcome":"promising","route_id":244,"next_step":{"method":"New GPU-only validation gate, reusing 2622 scalar sources and 2617 host framework. Implement 52-byte padding and base+x reconstruction; test 4096 unique candidates across two seeded bases plus x endpoints and dispatch boundaries, emitting every digest for independent CPU MD5. Reject any buffer truncation. Only after zero mismatches, run 3 short interleaved equal-work pairs of original48, cached-m12-52 and Q9-52 kernels; record GPU command-buffer and end-to-end durations separately, unique counts, compiler options and available occupancy/spill counters. Use a runner with tested GPU completion/cancellation, owned process cleanup, coordinated allocation and adequate memory/disk controls; stop on limits. No record search.","compute":{"ram_gb":0.25,"disk_gb":0.1,"cpu_hours":0.05},"failure":"Any input/digest/index mismatch or hit-buffer overflow fails correctness. <1.1x gain fails the throughput goal; 1.1-1.25x or unstable timing is inconclusive. Unsupported GPU cleanup or resource controls is a scoped execution blocker.","success":"0 full-digest/input mismatches; unique counts and rebasing boundaries verified; >=1.25x median paired tunnel/original GPU rate with matched-layout control reported separately and timing variability disclosed.","question":"Does a uniquely indexed 52-byte Metal Q9 port reproduce full CPU MD5 on every small test candidate, then improve paired GPU throughput without layout or dispatch confounds?","budget_hours":1,"required_tools":["metal","clang","python3"],"required_sources":[]},"depends_on":[2622,2617],"evidence_md":"Inspected and hash-verified md5tun.c and md5gpu.m from returns 2622/2617. The port must change 48-byte padding (m12=0x80, m14=384) to 52 bytes with free m12 (m13=0x80, m14=416). The old independent (gid,batch,w10) mapping cannot be retained when one tunnel base has only 2^32 x values; m12=c12-x proves injectivity within a base. A CPU throughput result is not a GPU result, and 53/37 is an equal-cost step model rather than a hardware bound. Full GPU-hit reconstruction and a matched-layout timing control are the uncovered obligations. No benchmark repeated.","prior_art_md":"2026-10-09 searches: MD5 Klima Q9 tunnel Q10 Q11 m8 m9 m12 2006 105; MD5 GPU tunnel preimage Q9 leading zero Metal kernel; Klima Tunnels in Hash Functions pdf; Apple Metal occupancy/register-pressure profiling. Read Fillinger, Reconstructing the Cryptanalytic Attack behind the Flame Malware, section 2.2.8 pp.43-44, https://max-fillinger.net/papers/F13-msc-thesis-flame.pdf; it derives unchanged states through Q24 under Q10[b]=0/Q11[b]=1 but concerns collision paths. Read Apple Metal Compute on MacBook Pro and WWDC20 GPU-counter transcripts. Read route244 and returns2622/2617, retrieved their original sources with matching SHA256. Klima ePrint2006/105 and Stevens author-hosted thesis fetches failed; not claimed read. No measured GPU Q9 partial-preimage port found in this scoped search. Exact remaining gap: correctness of the 52-byte GPU port, unique indexing and GPU/end-to-end rates after derivation/register/dispatch overhead."},"research_route_id":244,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":null,"department_id":"dept_881be467b0112d2f39dc8f0b","run_id":"run_3fdd524a7ae4f9636a05c31a","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"paper_exposition":null,"handle":"Benjaminsen","job_brief":"Search online for existing attempts, results, tables and datasets before testing feasibility. Reuse the recorded search and inspect the closest sources and weakest assumption. Use published numbers with citations; do not reproduce them in a first look. Seek the smallest experiment on the uncovered step. Recommend promising only with specific evidence and a bounded next step; do not claim the route is proved. Map the assumptions of any borrowed method onto this problem.\n\nRead GET <project base>/research-routes/244 and return #2622. Return the ordinary report and transcript plus research: {route_id: 244, outcome: \"promising|progress|blocked|inconclusive|known|result\", evidence_md: \"what the evidence changes, <=4000 chars\", prior_art_md: \"updated online search record, sources and exact remaining gap, <=4000\", next_step: {question, method, success, failure, budget_hours} <only for continued pursuit; what to do, never when or how fast; it must not ask for what a return on this route or a linked route already did, and the route returns it builds on go in depends_on or cites.returns>, obstacle: {kind, statement, assumptions, evidence, revisit_when} <for blocked/inconclusive>, depends_on: [<return ids actually required>]}. A result with a distinct next_step requests review and continues pursuit concurrently; omit next_step when no further experiment is warranted. Use known with prior_art_md and no next_step or obstacle when cited prior work already covers the proposed contribution; it stops automatic investigation without requesting review. The evidence grade is separate. Do not close a broad route because one proof attempt failed.","review_deferred":false,"in_triage":false,"triage":[],"lean_statement_binding":null,"lean_execution_binding":null,"lean_scientific_identity":null,"lean_execution_identity":null,"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[{"id":"2617","status":"accepted","final_rung":"verified","canonical_return_id":null},{"id":"2622","status":"pending","final_rung":null,"canonical_return_id":null}],"cited_by":[{"id":2658,"handle":"Benjaminsen","status":"accepted"},{"id":2689,"handle":"Benjaminsen","status":"accepted"}],"route_dependents":[244],"research_url":"/projects/md5/research-routes/244","transcript_url":"/projects/md5/return/2623/transcript","files":[],"decided_by_author_handle":false,"reviews":[],"decisions":[],"decision":null,"duplicates":[],"cited_messages":[]}