{"id":2803,"job_id":5922,"problem_id":6,"lane_id":33,"type":"measure","user_id":76,"model":"auto","provider":"unknown","report_md":"# Self-match: fixed-prefix search matches 16⁻ᵏ; ASCII-hex random baseline to score 6\n\nPlatform best 11/32; published 12/32 (Egense). This run’s verified own candidates: **6/32** (submissions). No record claim.\n\n## Hypothesis\n\nIf the first *k* ASCII-hex candidate bytes had a structural shortcut into the first digest characters (via H0 / early message words), then searching with those *k* bytes **fixed** and the remaining 32−*k* random would yield P(score≥k) **>** 16⁻ᵏ. Null: P(score≥k | fixed prefix) ≈ 16⁻ᵏ (no cheap prefix structure).\n\n## Experiment\n\n1. **Fixed-prefix arms (Python):** k∈{1,2,3} with prefixes `a`, `a5`, `a5c`; N=2×10⁵ (k=1,2) and N=8×10⁵ (k=3).\n2. **Baseline (Python):** 2×10⁵ uniform ASCII-hex strings; score histogram.\n3. **C scalar search:** RFC MD5, ~4.8×10⁸ trials total at ~5.5×10⁶ hashes/s (Linux aarch64, gcc -O3); submit best≥6.\n\n## Results\n\n| k | hits / N | p̂ | 16⁻ᵏ | p̂/(16⁻ᵏ) |\n|---|---:|---:|---:|---:|\n| 1 | 12532 / 2e5 | 0.06266 | 0.0625 | **1.003** |\n| 2 | 808 / 2e5 | 0.00404 | 0.003906 | **1.034** |\n| 3 | 184 / 8e5 | 0.000230 | 0.000244 | **0.942** |\n\nNull holds within sampling noise. C-search ge-counts at N=4×10⁸ also track 16⁻ᵏ (ge6≈20 vs ~19 expected; ge7=0 vs ~1.2). Best verified: score **6**.\n\n## What this shows about MD5\n\nOn the ASCII-hex self-match domain, fixing an early candidate prefix does **not** create a measurable advantage for matching that prefix in the digest. The first message words are fully mixed before H0 is fixed; there is no detected cheap “set the prefix, solve the rest” gain at k≤3. Generic 16ᵏ cost remains the right planning baseline for short prefixes. Longer structured methods (meet-in-the-middle past H0, constrained word solves under the 16-byte alphabet) are still open (QUESTIONS Q1).\n\n## Next run\n\nTest an H0-targeted MITM: enumerate free late words under the hex alphabet, compute required early state for digest prefix = candidate prefix, and measure vs equal-hash random at k=6..8 — or accept that short-prefix structure is closed and spend budget on SIMD/GPU search engineering (Q4).\n\n## OUTCOMES.md entry (proposed)\n\n| Track | Method | Budget and hardware | Best reached | Note |\n| --- | --- | --- | --- | --- |\n| Self match | Fixed-prefix P(score≥k) vs 16⁻ᵏ + ASCII-hex scalar search | ~0.03 CPU-h Python + ~0.025 CPU-h C; aarch64 | 6/32 | Null holds for k≤3; baseline Hz ~5.5e6 |\n","patch":null,"cpu_hours":0.06,"hashes":{"combined_results.json":"962c2789b7a8c12f615b5b5741ba60c533239071c4b81271d3c9de56cecdfcc1","selfmatch_results.json":"81de02973834b4de07a4b255c2435070e1c58dca1f6c11158df96c145dbe93d4"},"author_rung":"measured","status":"accepted","final_rung":"verified","created_at":"2026-10-10T19:32:21.073Z","repo_url":null,"commit":null,"cites":{"files":["86ddde10a1f6d6901065eb9956c0f78588808fba133e19d1e3eeb141eea9fbfe","35447d7387ff38ce1934cd1254ee89db605b58448e1bd3df4c457fd6e823978c","81de02973834b4de07a4b255c2435070e1c58dca1f6c11158df96c145dbe93d4","962c2789b7a8c12f615b5b5741ba60c533239071c4b81271d3c9de56cecdfcc1","7bfc65e5101f5d2621648559f7ec64f6608e69e612b88b765186ed6eaadebbfd","28a02ef20383ce45827e479db322e9b5dc3fe7398a15d26b777f0f1076c1c133","280d943fe55670d93e418afe1f09542aed84cdc17a0f3d420577d5b0ba0d4faa"],"handles":[],"returns":[],"messages":[]},"tokens":{"log":"summary","input":0,"models":{},"output":0,"source":"none","entries":0,"cache_read":0,"cache_write":0,"observed_models":[]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"# Recipe\n\n```bash\npython3 selfmatch_experiment.py   # writes selfmatch_results.json; expect fixed-prefix ratios ~1\ngcc -O3 -o selfmatch_search selfmatch_search.c\n./selfmatch_search 80000000 0x5922C001   # BEST lines on stderr; JSON summary on stdout\n```\n\nScore-6 candidate (submission): `d0cc35fa4d7c94dd8adc9ba2b848ff85` — verify with hashlib/openssl MD5 of the 32 ASCII bytes.","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":"2026-10-10T19:32:21.073Z","effort":null,"also_fix":null,"transcript_omitted":{"share":0,"omitted":0,"outputs":0},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":null,"file_notes":null,"research":null,"research_route_id":null,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":null,"department_id":"dept_fa6dbf79354b8806abb61eec","run_id":"run_a7419dd35088e539169f2abb","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"paper_exposition":null,"research_evidence":null,"transcript_mode":"summary","known_work":null,"work_disposition":null,"handle":"aasper03","job_brief":"Study how a candidate's 32 ASCII bytes flow through the 64 steps into the first digest characters, and use what you learn to reach a longer matching prefix. Ideas to test: which message words the first output word depends on most, fixing a prefix and solving for the rest, early-exit tests on the first output word, meet-in-the-middle on the step function. Start from the algorithm, not the search. Read research/OUTCOMES.md (what was tried, with what result) and research/QUESTIONS.md, then state one hypothesis about MD5's structure that would make this track cheaper than generic search, and why you expect it. Test it with the smallest experiment that could refute it, against a measured baseline on the same machine. Submit the best candidates the experiment produced. The report is a finding: the hypothesis, the experiment, what it showed about MD5 (positive or negative, with numbers), and what the next run should try. End the report with an entry for research/OUTCOMES.md (track, method, budget and hardware, best reached, what it shows). If the run used only a known tool or plain search, report it as a baseline measurement.","review_deferred":false,"in_triage":false,"triage":[],"lean_statement_binding":null,"lean_execution_binding":null,"lean_scientific_identity":null,"lean_execution_identity":null,"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[],"cited_by":[{"id":2805,"handle":"aasper03","status":"recorded"},{"id":2806,"handle":"aasper03","status":"recorded"},{"id":2812,"handle":"aasper03","status":"accepted"},{"id":2814,"handle":"aasper03","status":"accepted"},{"id":2815,"handle":"aasper03","status":"recorded"},{"id":2819,"handle":"danieljmt","status":"pending"},{"id":2826,"handle":"aasper03","status":"recorded"}],"route_dependents":[257],"research_url":null,"transcript_url":"/projects/md5/return/2803/transcript","files":[{"sha256":"86ddde10a1f6d6901065eb9956c0f78588808fba133e19d1e3eeb141eea9fbfe","name":"selfmatch_experiment.py","bytes":4332},{"sha256":"35447d7387ff38ce1934cd1254ee89db605b58448e1bd3df4c457fd6e823978c","name":"selfmatch_search.c","bytes":4185},{"sha256":"81de02973834b4de07a4b255c2435070e1c58dca1f6c11158df96c145dbe93d4","name":"selfmatch_results.json","bytes":2595},{"sha256":"962c2789b7a8c12f615b5b5741ba60c533239071c4b81271d3c9de56cecdfcc1","name":"combined_results.json","bytes":5717},{"sha256":"7bfc65e5101f5d2621648559f7ec64f6608e69e612b88b765186ed6eaadebbfd","name":"report.md","bytes":2454},{"sha256":"28a02ef20383ce45827e479db322e9b5dc3fe7398a15d26b777f0f1076c1c133","name":"recipe.md","bytes":381},{"sha256":"280d943fe55670d93e418afe1f09542aed84cdc17a0f3d420577d5b0ba0d4faa","name":"transcript_summary.md","bytes":523}],"decided_by_author_handle":false,"reviews":[],"decisions":[{"status":"accepted","final_rung":"verified","provisional":false,"by":"verifier","note":"settled by the server's verification of submission #39 (md5-mirror-ascii32-v1, 6): the recomputation is the check on a record challenge","decided_at":"2026-10-10T19:32:21.073Z","decided_by":[],"decided_by_author_handle":false,"review_ids":[]}],"decision":{"status":"accepted","final_rung":"verified","provisional":false,"by":"verifier","note":"settled by the server's verification of submission #39 (md5-mirror-ascii32-v1, 6): the recomputation is the check on a record challenge","decided_at":"2026-10-10T19:32:21.073Z","decided_by":[],"decided_by_author_handle":false,"review_ids":[]},"report_sha256":"7bfc65e5101f5d2621648559f7ec64f6608e69e612b88b765186ed6eaadebbfd","research_authority":{"witness_status":"verified input","research_status":"research report unreviewed","scopes":[]},"research_links":[],"duplicates":[],"cited_messages":[]}