{"id":2644,"job_id":5510,"problem_id":6,"lane_id":33,"type":"measure","user_id":1,"model":"gpt-6.1-sol","provider":"openai","report_md":"# Digest-guided prefix overwrite on full MD5 ASCII32\n\nOne bounded prospective measurement found no support for a changed-input advantage at overwrite lengths 1, 2, 4, or 8. The originally preregistered contrast also contained a conditioned-control defect; the corrected contrasts are explicitly post hoc and do not establish equivalence. The best own candidate is score **5/32**, `d387621c5e23d6b0a5437c316b67ca8b` → `d3876c286fc1a26ebe948fc815ecb6ad`, generated by the **fresh baseline**, r=8, index=14507. No candidate was submitted by this worker, and no record improvement is claimed.\n\nThis concerns supplied QUESTIONS question 1, self-match beyond generic search. The score uses the first differing character of the full lowercase MD5 digest of the **32 literal ASCII bytes**. It never hashes hex-decoded candidate bytes. The model assumption is only a reference for interpreting finite measurements, not a proof of MD5 hardness or fixed-point existence.\n\n## Prior comparison and preregistration\n\nThe supplied sanitized public scientific reports 2610/2618/2626 cover forward gates, cached computation and diffusion; 2627 covers Q tunnels; 2639 tests single-character suffix neighbours of exact H0 hits. Report 2633 references 2630's frozen-state final-X4 reinjection, but that report itself was not supplied here. None of the supplied reports documents the present digest-prefix overwrite measurement. This is a scoped comparison to the supplied evidence, not an exhaustive literature claim.\n\nOriginal `preregister.json` was saved before data. Seed `measure-5510-prefix-overwrite-v1`; 32,768 bases for each r in (1,2,4,8). Synthetic base, random prefix and fresh candidate are derived by the exact SHA256 domain-separated recipe in the source. They are deterministic pseudorandom samples, not a guaranteed independent uniform random sample from MD5's input domain.\n\nEach trial hashes the base, replaces its first r characters with its observed digest prefix and hashes that candidate, hashes a random replacement of the same r characters, and hashes a separately generated fresh candidate. All four calls execute, even if an input repeats. The three separately compared arms each receive the base plus one candidate: **65,536 charged MD5 evaluations per arm per r**, 262,144 per arm over all lengths. Actual union trial calls are **524,288**; common base calls are physically shared and are not claimed as extra compute. These are equal hash-count budgets, not equal measured per-arm wall time. Candidate generation, score aggregation and set bookkeeping are included in the experiment's measured runtime; no throughput speedup is inferred.\n\nPrimary rule: among guided inputs changed from their bases, guided versus random replacement score≥1 must have positive discordance difference exceeding 4.5√(total discordances). All four original primary z values (3.183, 0.081, 0.931, 0.244) fall below 4.5. Secondary score distributions and two-evaluation best scores were preregistered. The threshold is a descriptive conservative diagnostic for this finite sample; no exact independence assumption or calibrated p-value is asserted.\n\n## Preserved design correction\n\nAfter original data generation, the parent review identified that `guided != base` implies base score<r. A random replacement can leave that conditioned base unchanged. At r=1 it then deterministically fails the first-character criterion. An ideal independently sampled random-map heuristic would itself predict a guided/random ratio near 16/15 in that original conditional comparison, because about 1/16 of random replacements waste their fresh query; this is **not MD5 structure**. The original preregister/source/evidence were preserved. `control_correction.py` uses the already recorded score/flag bytes, performs **zero MD5 evaluations**, and labels its output post hoc. It compares both-changed pairs and the fresh arm.\n\n| r | Guided unchanged | Random unchanged | Both changed n | Guided score≥1 | Random score≥1 | Both-changed paired z | Guided-changed vs fresh z |\n|---:|---:|---:|---:|---:|---:|---:|---:|\n| 1 | 2040 | 2062 | 28802 | 1776 | 1735 | 0.737 | -1.028 |\n| 2 | 123 | 120 | 32525 | 2050 | 2051 | -0.016 | -0.449 |\n| 4 | 0 | 1 | 32767 | 2102 | 2043 | 0.947 | 0.959 |\n| 8 | 0 | 0 | 32768 | 2047 | 2032 | 0.244 | -0.032 |\n\nAt r=1 the original conditioned guided/random hit counts are 1,916/1,735, z=3.183. Requiring both inputs to change gives 1,776/1,735, z=0.737; comparing changed guided inputs with fresh candidates gives 1,916/1,978, z=-1.028. The apparent directional difference depends on conditioned duplicates. The larger lengths likewise offer no signal in these corrected descriptive contrasts. Both-changed guided-minus-random first-character effects for r=1/2/4/8 are +0.142 percentage points, -0.003 percentage points, +0.180 percentage points, +0.046 percentage points. Corresponding four-standard-error descriptive widths are ±0.772, ±0.760, ±0.761, ±0.752 percentage points; these are finite-sample resolution summaries, not equivalence bounds. Both-changed filtering does not make all input values independent: random replacements can also equal guided replacements, and the pairs share a base. Exact scores, duplicate flags and paired discordances are retained instead of treating repeated hashes as new trials.\n\n## Equal-cost discovery comparison\n\nCounts below are trial budgets whose **best of base and second candidate** reaches k≥1/2/3, rather than counting a reused old hit as a new success.\n\n| r | Guided budget hits k1/k2/k3 | Random-prefix budget hits k1/k2/k3 | Fresh generic budget hits k1/k2/k3 |\n|---:|---:|---:|---:|\n| 1 | 3956/246/13 | 3775/242/14 | 4018/249/14 |\n| 2 | 3969/244/19 | 3942/262/24 | 3971/253/20 |\n| 4 | 4052/251/18 | 4025/239/16 | 4011/243/19 |\n| 8 | 4005/245/16 | 3996/237/16 | 3998/222/21 |\n\nThe best overall score 5 came from fresh generic input, not the repair method. Across all actual trial calls there are **517,868 distinct inputs**, **6,420 repeated calls**, and **2,433 repeated score≥1 calls**. At r=1, all 2,040 unchanged guided calls repeat already observed first-character hits. Random-prefix calls also duplicate a guided candidate, so union reuse is larger than their base-unchanged count (4,027 versus 2,062 at r=1; 229 versus 120 at r=2). No same-input hit is interpreted as a probability gain. Within each individual method/r there are no repeated inputs across its 32,768 rows; duplication occurs across methods within shared trials. Global union membership was measured exactly.\n\nExact marginal score distributions follow (scores 6..32 were all zero). Conditional histograms, base→guided transitions and k1/k2/k3 discordances are in the JSON evidence.\n\n| r | Method | score0 | score1 | score2 | score3 | score4 | score5 |\n|---:|---|---:|---:|---:|---:|---:|---:|\n| 1 | base | 30728 | 1912 | 120 | 7 | 1 | 0 |\n| 1 | guided | 28812 | 3710 | 233 | 12 | 1 | 0 |\n| 1 | randomprefix | 30775 | 1870 | 116 | 6 | 1 | 0 |\n| 1 | fresh | 30658 | 1989 | 115 | 6 | 0 | 0 |\n| 2 | base | 30768 | 1877 | 111 | 10 | 2 | 0 |\n| 2 | guided | 30585 | 1939 | 225 | 17 | 2 | 0 |\n| 2 | randomprefix | 30705 | 1924 | 127 | 12 | 0 | 0 |\n| 2 | fresh | 30667 | 1971 | 122 | 8 | 0 | 0 |\n| 4 | base | 30677 | 1966 | 112 | 13 | 0 | 0 |\n| 4 | guided | 30666 | 1975 | 122 | 5 | 0 | 0 |\n| 4 | randomprefix | 30724 | 1929 | 112 | 3 | 0 | 0 |\n| 4 | fresh | 30726 | 1922 | 114 | 6 | 0 | 0 |\n| 8 | base | 30695 | 1967 | 100 | 4 | 2 | 0 |\n| 8 | guided | 30721 | 1908 | 129 | 10 | 0 | 0 |\n| 8 | randomprefix | 30736 | 1901 | 121 | 9 | 1 | 0 |\n| 8 | fresh | 30719 | 1932 | 102 | 12 | 2 | 1 |\n\n## Correctness, cost and limits\n\nHashlib and the independent mathematical RFC1321 implementation agree on all seven RFC1321 Appendix A.5 vectors, **256 trial candidate digests** (first 16 trials in every r, all four methods), and the full best digest. The independent implementation covers standard IV, padding, all 64 updates, feed-forward, and little-endian serialization. RFC vectors are controls, never candidate submissions. Constants and vectors are identified with RFC1321; this worker did not fetch external sources. The supplied prior reports cite the RFC primary source.\n\nObserved hardware is Darwin 24.6.0, arm64, 10 logical CPUs; the experiment uses one Python worker, no GPU. CPU model was not separately measured and is not inferred. Python 3.14.6. Experiment process wall **1.433834 s**, process CPU **1.432062 s**, peak RSS **117,932,032 bytes**. Post hoc process wall **0.388185 s**, process CPU **0.387328 s**. Total observed bounded child/watchdog CPU is **1.938394 s** (0.000538443 CPU hours); this excludes unmeasured artifact formatting, metadata reads and agent inference.\n\nThe actual pinned adapter `bounded` enforced measurement wall≤60 s, CPU≤30 s/process, file≤8 MiB; post hoc analysis used wall≤20 s, CPU≤10 s/process, file≤8 MiB. Both exited 0, watchdog exited 0, group cleanup succeeded, and separate signal-0 checks observed ESRCH for their owned process groups. No background jobs remain. RAM was small and observed, not supervisor-enforced. Runtime/hardware were emitted separately from deterministic scientific stdout; private original receipts retain PGIDs and FD redirection. Native usage is reconciled separately by the parent from original native events, without estimates.\n\nOnly this one seed, these lengths and 131,072 synthetic bases were tested. Scores≥3 are sparse. Sampling uncertainty, selection and shared-base dependencies limit inference. No general route, rare structure, fixed-point existence question, MD5 hardness claim or worldwide record is closed. No platform candidate verification is claimed; that is the parent's separate responsibility.\n\n## OUTCOMES entry and next step\n\nSelf match — digest-guided prefix overwrite r=1/2/4/8 on 32,768 synthetic bases each, 524,288 actual MD5 trial calls, one arm64 CPU worker, observed bounded-child CPU 1.938394 s. Original primary tests did not meet threshold; post hoc both-changed/fresh controls removed conditioned-duplicate bias and showed no first-character advantage at this resolution. Best score5 originated in the fresh baseline. Full input-reuse accounting, paired score distributions, and independent RFC MD5 controls retained. Questions1/5 remain open.\n\nCheapest useful next step: preregister a new independent seed with random replacement explicitly sampled to differ from its base, and keep the fresh two-evaluation baseline plus exact union reuse accounting. Replicate the corrected contrast before proposing larger compute; do not enlarge this sample post hoc to cross a threshold.\n\n\nParent completion evidence: own fresh-baseline candidate is verified as submission13, score5, duplicate=false, site_record=false, personal_best=false. The full digest agreed in two implementations before submission. This verifies candidate values, not the written experimental inference. Original worker report and original preregistered/posthoc evidence remain at their exact file hashes. Parent best-candidate preflight additionally used 0.058605 observed CPU seconds including watchdog/supervisor; total reported scoped scientific execution is 1.996999 CPU seconds. No record or method speedup is claimed; written claims retain measured scope and requested independent review.\n","patch":null,"cpu_hours":0.0005547219444444444,"hashes":{"recipe.md":"ed735df72598fdef6c69e44272764268984447a7cdfb06e418d38a96c02a53a7","report.md":"4a28c58065788dc6bef1a07604b8163499f56838c64b792c74481efea1843d01","timing.json":"199dce3b4962c1995a5838fca8639ab80fa40caf4f4cfdfc013b498626bac07e","sources.json":"a444864e339e3966460f25eaf4b355774d1b7266eadfa507c87eb6c9131d7c47","evidence.json":"8e676643d82fa0638241fa65492c8a8ff2ff871a522d610b202842f2226b5989","preregister.json":"aec7a7d102b116795543a49d41815d69d5c072f66f5cb5588de679fd25509066","artifact-index.json":"1e8f66abde426c7504e3fa9666072fdd3092003d8ee196c7c480fefb9d1e97c7","prefix_overwrite.py":"f9eb94574e8e7d448cbf8f12acbef0c6e24f46db9e08f3b7de4675a318675ec6","evidence-base64.json":"fbaae6fd9d61b0eff5448d73fcada59ac6834443fa14ca7ebd3388237b9d66a1","control_correction.py":"7ec9b5ea99df92a07bb115d059a0bbcbb9558d07d62f294eb483211ba4d3c374","scientific-result.json":"1fba844e8664ff352160449ff49cfec1a263aa3eef8ce1f4b9278f560291c427","control-correction.json":"e6a107e89a0d7162e60ec7242990c8a903a994f1016712dbdf80b9dd2c165628","parent-preflight-evidence.json":"997a07f4333b913ea8126c950cb190460edf6b9f981482bad50f4efff19ccfed"},"author_rung":"measured","status":"accepted","final_rung":"verified","created_at":"2026-10-09T22:44:42.574Z","repo_url":null,"commit":null,"cites":{"files":[],"handles":[],"returns":[2610,2618,2626,2627,2633,2639],"messages":[]},"tokens":{"log":"codex","input":184892,"models":{"gpt-6.1-sol":56786},"output":56786,"source":"codex-jsonl","entries":59,"cache_read":7082112,"cache_write":0,"observed_models":["gpt-6.1-sol"]},"paper_slug":null,"revision_path":null,"revision_sha":null,"recipe_md":"# Reproduce the preserved measurement\n\nUse Python 3 stdlib and one CPU worker. No dependencies, GPU or network are required. Place prefix_overwrite.py and control_correction.py in a new directory. Retain preregister.json unchanged. Under a reviewed process-group supervisor enforcing wall<=60 s, CPU<=30 s/process and per-file<=8 MiB, run:\n\n```sh\npython3 prefix_overwrite.py > evidence.json 2> measurement.stderr.private.json\npython3 control_correction.py > control-correction.json 2> correction.stderr.private.json\n```\n\nThe original used pinned adapter bounded(argv, seconds=60, cpu_seconds=30, file_bytes=8388608); correction used seconds=20, cpu_seconds=10. Capture real process groups and verify cleanup. The supervisor framework is private and not a portable dependency. Timing output is separate and is not a byte-reproduction target. This is one fixed measurement, not a record search.\n\nTo reconstruct the original binary from the exact UTF-8 envelope, check its artifact SHA256 first, then:\n\n```python\nimport base64, hashlib, json\nfrom pathlib import Path\np = json.loads(Path('evidence-base64.json').read_text(encoding='utf-8'))\nassert p['version'] == 1 and p['encoding'] == 'base64'\ns = p['data'].encode('ascii')\nassert len(s) == p['base64_ascii_bytes'] <= 2_000_000\nassert hashlib.sha256(s).hexdigest() == p['base64_ascii_sha256']\nb = base64.b64decode(s, validate=True)\nassert len(b) == p['decoded_bytes'] == 917504\nassert hashlib.sha256(b).hexdigest() == p['decoded_sha256']\nPath('evidence.bin').write_bytes(b)\n```\n\nBinary format: 7 unsigned bytes per trial, no header, outer order r=1,2,4,8, inner index0..32767. Bytes are base/guided/randomprefix/fresh scores, guided_changed, random_changed, union_duplicate_bitmask (bit0 base, bit1 guided, bit2 randomprefix, bit3 fresh). Source explicitly fixes candidate generation and tie-breaking. Deterministic JSON evidence must agree byte-for-byte; compare artifact-manifest hashes. Evidence binary digest/length are also in evidence.json.\n\nBest candidate can be regenerated without searching: SHA256 of ASCII `measure-5510-prefix-overwrite-v1:fresh:8:14507`, rendered lowercase hex and truncated to32 characters, gives `d387621c5e23d6b0a5437c316b67ca8b`. Hash the literal ASCII candidate with full MD5: `d3876c286fc1a26ebe948fc815ecb6ad`, common prefix5. Independent md5_reference in prefix_overwrite.py gives the same full digest. RFC control vectors are not search candidates.\n\nPost hoc correction is preserved as such; it was added after original data in response to a design review. Do not relabel it preregistered. See report.md for scope, conditioning, equal paid budgets, duplicate accounting and measured costs. Original native identity, provenance, private FD captures, process receipts and framework omissions remain with the parent. No upload or verification is claimed until the parent observes its actual receipt.","verification":null,"target":null,"finding":null,"human_md":null,"provisional":false,"effects_applied_at":"2026-10-09T22:44:42.574Z","effort":"high","also_fix":null,"transcript_omitted":{"share":0.07142857142857142,"omitted":4,"outputs":56},"patch_hash":null,"superseded_by":null,"duplicate_of":null,"transcript_resubmitted_at":"2026-10-09T22:50:26.474Z","file_notes":null,"research":null,"research_route_id":null,"verification_plan":null,"verification_fingerprint":null,"review_admitted_at":null,"department_id":"dept_881be467b0112d2f39dc8f0b","run_id":"run_3fdd524a7ae4f9636a05c31a","triage_lead":null,"revision_base_sha":null,"integration":null,"resolves":null,"paper_exposition":null,"handle":"Benjaminsen","job_brief":"Study how a candidate's 32 ASCII bytes flow through the 64 steps into the first digest characters, and use what you learn to reach a longer matching prefix. Ideas to test: which message words the first output word depends on most, fixing a prefix and solving for the rest, early-exit tests on the first output word, meet-in-the-middle on the step function. Start from the algorithm, not the search. Read research/OUTCOMES.md (what was tried, with what result) and research/QUESTIONS.md, then state one hypothesis about MD5's structure that would make this track cheaper than generic search, and why you expect it. Test it with the smallest experiment that could refute it, against a measured baseline on the same machine. Submit the best candidates the experiment produced. The report is a finding: the hypothesis, the experiment, what it showed about MD5 (positive or negative, with numbers), and what the next run should try. End the report with an entry for research/OUTCOMES.md (track, method, budget and hardware, best reached, what it shows). If the run used only a known tool or plain search, report it as a baseline measurement.","review_deferred":false,"in_triage":false,"triage":[],"lean_statement_binding":null,"lean_execution_binding":null,"lean_scientific_identity":null,"lean_execution_identity":null,"verification_runs":[],"verification_state":null,"verification_summary":null,"canonical_return":null,"review_history":[],"dependencies":[],"cited_by":[{"id":2654,"handle":"Benjaminsen","status":"accepted"},{"id":2663,"handle":"Benjaminsen","status":"accepted"}],"route_dependents":[],"research_url":null,"transcript_url":"/projects/md5/return/2644/transcript","files":[{"sha256":"4a28c58065788dc6bef1a07604b8163499f56838c64b792c74481efea1843d01","name":"measure5510-report.md","bytes":10632},{"sha256":"ed735df72598fdef6c69e44272764268984447a7cdfb06e418d38a96c02a53a7","name":"measure5510-recipe.md","bytes":2883},{"sha256":"aec7a7d102b116795543a49d41815d69d5c072f66f5cb5588de679fd25509066","name":"measure5510-preregister.json","bytes":2423},{"sha256":"f9eb94574e8e7d448cbf8f12acbef0c6e24f46db9e08f3b7de4675a318675ec6","name":"measure5510-prefix_overwrite.py","bytes":8426},{"sha256":"7ec9b5ea99df92a07bb115d059a0bbcbb9558d07d62f294eb483211ba4d3c374","name":"measure5510-control_correction.py","bytes":2463},{"sha256":"8e676643d82fa0638241fa65492c8a8ff2ff871a522d610b202842f2226b5989","name":"measure5510-evidence.json","bytes":18571},{"sha256":"fbaae6fd9d61b0eff5448d73fcada59ac6834443fa14ca7ebd3388237b9d66a1","name":"measure5510-evidence-base64.json","bytes":1223609},{"sha256":"e6a107e89a0d7162e60ec7242990c8a903a994f1016712dbdf80b9dd2c165628","name":"measure5510-control-correction.json","bytes":6562},{"sha256":"199dce3b4962c1995a5838fca8639ab80fa40caf4f4cfdfc013b498626bac07e","name":"measure5510-timing.json","bytes":1219},{"sha256":"a444864e339e3966460f25eaf4b355774d1b7266eadfa507c87eb6c9131d7c47","name":"measure5510-sources.json","bytes":1649},{"sha256":"1fba844e8664ff352160449ff49cfec1a263aa3eef8ce1f4b9278f560291c427","name":"measure5510-scientific-result.json","bytes":1924},{"sha256":"1e8f66abde426c7504e3fa9666072fdd3092003d8ee196c7c480fefb9d1e97c7","name":"measure5510-artifact-index.json","bytes":2882},{"sha256":"997a07f4333b913ea8126c950cb190460edf6b9f981482bad50f4efff19ccfed","name":"measure5510-parent-preflight-evidence.json","bytes":919}],"decided_by_author_handle":false,"reviews":[],"decisions":[{"status":"accepted","final_rung":"verified","provisional":false,"by":"verifier","note":"settled by the server's verification of submission #13 (md5-mirror-ascii32-v1, 5): the recomputation is the check on a record challenge","decided_at":"2026-10-09T22:44:42.574Z","decided_by":[],"decided_by_author_handle":false,"review_ids":[]}],"decision":{"status":"accepted","final_rung":"verified","provisional":false,"by":"verifier","note":"settled by the server's verification of submission #13 (md5-mirror-ascii32-v1, 5): the recomputation is the check on a record challenge","decided_at":"2026-10-09T22:44:42.574Z","decided_by":[],"decided_by_author_handle":false,"review_ids":[]},"duplicates":[],"cited_messages":[]}