Skip to content

proof: before/after evidence for port-cuda-kernel-to-triton, plus an int64 pointer-arithmetic correction - #12

Open
titoatwork wants to merge 2 commits into
tensormux:masterfrom
titoatwork:proof/portability-port-cuda-kernel-to-triton
Open

titoatwork wants to merge 2 commits into
tensormux:masterfrom
titoatwork:proof/portability-port-cuda-kernel-to-triton

Conversation

@titoatwork

@titoatwork titoatwork commented Aug 26, 2026 •

Copy link
Copy Markdown

Answers the port-cuda-kernel-to-triton row in #1.

Adds the first entry under proof/portability/, for port-cuda-kernel-to-triton.

Four runs, two models, one variable: the same CUDA row-max reduction (a __shfl_down_sync butterfly plus a shared-memory cross-warp combine) ported to Triton with and without the skill file. The base task was byte-identical in all four runs.

Opus 5 with the skill: 0/11 shapes. The other three arms: 11/11.

The skill-guided port calls .to(tl.int64) on a stride argument. Triton folds integer arguments equal to 1 into compile-time constants, so a specialized argument reaches the kernel body as a plain Python int and the call raises AttributeError. A contiguous row-major tensor has unit stride in its last dimension, so the port fails for every contiguous input and compiles only for strided views.

This also corrects the pointer-arithmetic rule in the skill's Kernel design rules to promote the index tensor rather than the scalar argument, adds the matching entry to Common failure modes, and adds a review checklist item. npm run validate:skills passes.

No throughput claim is made. The collected timings spanned 2.1x across repeats of the same measurement and the largest working set was L2-resident, so they describe cache and clock state rather than kernel throughput.

RTX 3050 Laptop (sm_86), Triton 3.7.1, PyTorch 2.13.0+cu130, float32. Full reproduction under repro/, including the CUDA source, the shared base task, the harness, and the four unedited model outputs.


Planned on a Colab T4 per my comment on #1; run on a local RTX 3050 instead, which I should have flagged at the time and did not. That substitution is why the timing data is discarded — the clock ramp and the 1.5 MB L2 are both properties of the laptop card. The correctness matrix is unaffected, since the 0/11 is a compile-time AttributeError with no hardware dependence. Full provenance and the three smaller deviations from that plan are recorded under Hardware and setup and Deviations from the plan filed on issue #1 in the proof doc.

Four runs across two models, with and without the skill file, porting a CUDA
row-max reduction whose critical path is a __shfl_down_sync butterfly. The base
task was byte-identical in all four runs; the only variable is whether SKILL.md
was prepended.

Opus 5 with the skill fails all 11 shapes. The other three arms pass all 11.
The skill-guided port calls .to(tl.int64) on a stride argument, and Triton folds
integer arguments equal to 1 into compile-time constants, so a specialized
argument reaches the kernel body as a plain Python int and the call raises
AttributeError. A contiguous row-major tensor has unit stride in its last
dimension, so the port fails for every contiguous input while still compiling
for strided views.

Correct the pointer-arithmetic rule in Kernel design rules to promote the index
tensor rather than the scalar argument, add the matching entry to Common failure
modes, and add a review checklist item.

No throughput claim is made. The collected timings spanned 2.1x across repeats
of the same measurement and the largest working set was L2-resident, so they
describe cache and clock state rather than kernel throughput and are not
reported.

Signed-off-by: Ibteshamul Haque <ibteshamulhaque01@gmail.com>
The plan filed on issue tensormux#1 named a Colab T4 and said any change would be
stated rather than switched silently. The run was done on a local RTX 3050
and the substitution was never noted. Recording it now, in the proof doc
rather than only in the PR thread, since that is where a later reader looks.

The hardware note sits under Hardware and setup because it is load-bearing
there: both reasons the timing data is discarded are properties of the
laptop card. The correctness matrix is unaffected, as the 0/11 is a Triton
compile-time AttributeError with no hardware dependence.

Also records the three smaller deviations, being four runs across two models
rather than two runs on one, the row-max/ path rather than reduction/, and
the reference chain running through the CUDA kernel rather than directly to
PyTorch. Neither failure mode predicted in that comment appeared.

proof/README.md marks this entry as a negative result so the summary table
does not read it as a validation of the skill.

Signed-off-by: Ibteshamul Haque <ibteshamulhaque01@gmail.com>
@titoatwork

Copy link
Copy Markdown
Author

Correction to the plan I filed on #1. I said I would run on a Colab T4 and would state it rather than quietly switch if that changed. It changed — this ran on a local RTX 3050 — and I did not state it.

It is load-bearing for one part of the entry. The timings I discarded were discarded because of the laptop GPU's clock ramp and its 1.5 MB L2, both properties of the card I actually used rather than the one I named; a fixed-clock datacenter GPU with a larger L2 would likely not have produced that particular failure. The correctness matrix is unaffected, since the 0/11 is an AttributeError raised at Triton compile time before any kernel runs.

b934dc8 puts the provenance in the proof doc rather than only in this thread, since that is where a later reader would look, and records three smaller deviations from the same comment: four runs across two models rather than two runs on one model (this is what shows the failure is specific to Opus rather than to the skill in general), proof/portability/row-max/ rather than the proof/portability/reduction/ I named, and correctness checked against the CUDA kernel — itself validated against torch.max — rather than directly against PyTorch.

It also notes that neither failure mode I predicted in that comment appeared. No arm produced wrong numbers, and the -inf identity trap caught neither model with or without the skill. The finding here is not the one the experiment was designed to find.

Last thing: proof/README.md now marks this entry as a negative result, so the summary table does not read it as a validation of the skill alongside the seven passing entries.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant