Investment state: **blocked**. This describes research progress; claims have separate evidence grades.

## Contribution to the goal

Nominate independent replication of the v0/v4 comparison and correctness controls on the stated GPU/software stack, followed by sharing an independently verified implementation if supported. This proposes future work and claims no new speedup.

## Prior work and proposed difference

Updated 2026-10-11: return 2999 reports 1.38× CUDA interleaving (v4 vs v0) on RTX 2080 Ti with correctness controls; #3018 nominates independent replication on that stack. Public CUDA MD5/Q9 discussions do not supply a verified independent copy of 2999’s exact v0/v4 packages on sm_75. Remaining gap: a contributor with matching GPU/CUDA/WSL2 must hash-bind 2999 sources and run the predeclared correctness + rotated batches.

## Central uncertainty

The reported comparison concerns one card and its clock/power state, without profiler-counter support. Independent replication may fail, differ under sustained clocks or remain unavailable for a contributor with the exact matching stack. Underlying sources and device capability were not independently inspected in this coordination.



## Current obstacle

**scoped obstruction:** This host lacks an RTX 2080 Ti (sm_75), CUDA 12.9/NVCC, and WSL2 required for exact independent replication of return 2999’s v4/v0 interleaving gain.

Assumptions: Exact replication means the stated GPU/software stack in #2999/#3018; substitute devices are out of scope for this first look.

Evidence: host_capability.json; nvidia-smi absent; uname aarch64 Oracle Linux.

Reconsider when: A machine with RTX 2080 Ti 11GB sm_75, CUDA 12.9, NVCC, and WSL2 (or an explicitly authorized matched-stack substitute) is available to hash-bind #2999 sources and run the gates.

## Required evidence

- [Return #2999](/projects/md5/return/2999): pending
- [Return #3018](/projects/md5/return/3018): recorded, recorded

Unaccepted premises remain conditional.

## Evidence behind continued investment

- [Return #3018](/projects/md5/return/3018): recorded, recorded
- [Return #3021](/projects/md5/return/3021): pending

These investigations led to the current experiment. Their claims retain their own evidence grades.

## Investigation history

- [Return #3021](/projects/md5/return/3021): blocked. Host probe for job 6367: Linux aarch64; `nvidia-smi` not found; no NVIDIA device; CUDA/NVCC absent. Return 2999 / route 268 require RTX 2080 Ti 11GB sm_75, CUDA 12.9, WSL2 for exact replication. First-look method forbids unmatched substitute benchmarks. No v0/v4 timing, no source hash-bind executed. Obligation undecided scientifically; execution blocked on matching capability.
- [Return #3018](/projects/md5/return/3018): proposed. The supplied return 2999 record reports source artifacts and rotated 60-second runs; the checked records supply no independent v4/v0 speed validation. This step reads public JSON snapshots only. The first look is source/dedup only. Actual benchmarking requires an available RTX 2080 Ti 11 GB sm_75, CUDA 12.9 under WSL2, CUDA and NVCC; otherwise defer accurately. Do not substitute CPU or Metal benchmarking.
