proof: before/after evidence for port-cuda-kernel-to-triton, plus an int64 pointer-arithmetic correction - #12
Conversation
Four runs across two models, with and without the skill file, porting a CUDA row-max reduction whose critical path is a __shfl_down_sync butterfly. The base task was byte-identical in all four runs; the only variable is whether SKILL.md was prepended. Opus 5 with the skill fails all 11 shapes. The other three arms pass all 11. The skill-guided port calls .to(tl.int64) on a stride argument, and Triton folds integer arguments equal to 1 into compile-time constants, so a specialized argument reaches the kernel body as a plain Python int and the call raises AttributeError. A contiguous row-major tensor has unit stride in its last dimension, so the port fails for every contiguous input while still compiling for strided views. Correct the pointer-arithmetic rule in Kernel design rules to promote the index tensor rather than the scalar argument, add the matching entry to Common failure modes, and add a review checklist item. No throughput claim is made. The collected timings spanned 2.1x across repeats of the same measurement and the largest working set was L2-resident, so they describe cache and clock state rather than kernel throughput and are not reported. Signed-off-by: Ibteshamul Haque <ibteshamulhaque01@gmail.com>
The plan filed on issue tensormux#1 named a Colab T4 and said any change would be stated rather than switched silently. The run was done on a local RTX 3050 and the substitution was never noted. Recording it now, in the proof doc rather than only in the PR thread, since that is where a later reader looks. The hardware note sits under Hardware and setup because it is load-bearing there: both reasons the timing data is discarded are properties of the laptop card. The correctness matrix is unaffected, as the 0/11 is a Triton compile-time AttributeError with no hardware dependence. Also records the three smaller deviations, being four runs across two models rather than two runs on one, the row-max/ path rather than reduction/, and the reference chain running through the CUDA kernel rather than directly to PyTorch. Neither failure mode predicted in that comment appeared. proof/README.md marks this entry as a negative result so the summary table does not read it as a validation of the skill. Signed-off-by: Ibteshamul Haque <ibteshamulhaque01@gmail.com>
|
Correction to the plan I filed on #1. I said I would run on a Colab T4 and would state it rather than quietly switch if that changed. It changed — this ran on a local RTX 3050 — and I did not state it. It is load-bearing for one part of the entry. The timings I discarded were discarded because of the laptop GPU's clock ramp and its 1.5 MB L2, both properties of the card I actually used rather than the one I named; a fixed-clock datacenter GPU with a larger L2 would likely not have produced that particular failure. The correctness matrix is unaffected, since the 0/11 is an
It also notes that neither failure mode I predicted in that comment appeared. No arm produced wrong numbers, and the Last thing: |
Answers the
port-cuda-kernel-to-tritonrow in #1.Adds the first entry under
proof/portability/, forport-cuda-kernel-to-triton.Four runs, two models, one variable: the same CUDA row-max reduction (a
__shfl_down_syncbutterfly plus a shared-memory cross-warp combine) ported to Triton with and without the skill file. The base task was byte-identical in all four runs.Opus 5 with the skill: 0/11 shapes. The other three arms: 11/11.
The skill-guided port calls
.to(tl.int64)on a stride argument. Triton folds integer arguments equal to1into compile-time constants, so a specialized argument reaches the kernel body as a plain Pythonintand the call raisesAttributeError. A contiguous row-major tensor has unit stride in its last dimension, so the port fails for every contiguous input and compiles only for strided views.This also corrects the pointer-arithmetic rule in the skill's Kernel design rules to promote the index tensor rather than the scalar argument, adds the matching entry to Common failure modes, and adds a review checklist item.
npm run validate:skillspasses.No throughput claim is made. The collected timings spanned 2.1x across repeats of the same measurement and the largest working set was L2-resident, so they describe cache and clock state rather than kernel throughput.
RTX 3050 Laptop (sm_86), Triton 3.7.1, PyTorch 2.13.0+cu130, float32. Full reproduction under
repro/, including the CUDA source, the shared base task, the harness, and the four unedited model outputs.Planned on a Colab T4 per my comment on #1; run on a local RTX 3050 instead, which I should have flagged at the time and did not. That substitution is why the timing data is discarded — the clock ramp and the 1.5 MB L2 are both properties of the laptop card. The correctness matrix is unaffected, since the 0/11 is a compile-time
AttributeErrorwith no hardware dependence. Full provenance and the three smaller deviations from that plan are recorded under Hardware and setup and Deviations from the plan filed on issue #1 in the proof doc.