fix(topology): a cube's units are contiguous in ABSOLUTE_POS - #1595
Merged
Merged
Conversation
A cube dim of (32, 8, 1), which `CubeDim::new` returns for every 256-unit launch, put the eight planes of a cube `CUBE_COUNT_X * CUBE_DIM_X` elements apart under the old linearization: one cube's stores went to eight streams a megabyte apart rather than to one contiguous run. On an RTX 4070 Ti SUPER that costs the write path 2.4%, measured against the same kernel hand-written in CUDA, and the read path nothing. The bijection onto `0..total_units` holds either way, and a plane stays contiguous in both, so `testgen_topology` and plane operations are unaffected. What changes is that ABSOLUTE_POS no longer linearizes the (ABSOLUTE_POS_X, Y, Z) coordinate: it is now the cube index times the cube size plus the unit index within the cube.
The existing test scatters value i to index i and compares against 0..n, which holds under any bijection, so it cannot see a cube whose units are spread across the grid.
Merged
8 tasks
ThierryCantin-Demers
marked this pull request as ready for review
September 3, 2026 17:18
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

CubeDim::newalways returns a 2D cube, andABSOLUTE_POSis linearized grid-major, so its y term strides the whole grid. For a cube of (32, 8, 1) on a grid of (2112, 1, 1) it comes out asty*(gridDim.x*32) + bx*32 + tx, and the eight planes of one cube land on eight streams 1.03 MiB apart instead of one contiguous run.Each warp is still fully coalesced, so the generated code is unchanged. What changes is how many concurrent streams reach DRAM, and stores are far more sensitive to that than loads.
This linearizes it as
cube_pos() * CUBE_DIM + UNIT_POSin all four backends.Measured
RTX 4070 Ti SUPER, the
throughputexample, three runs each:RTX 2060, same example, three runs each, main against this branch on the same day. This card has a 3 MiB L2 against the 4070's 48 MiB, and pays far more for scattered streams:
The after column is within two points of what the same card reports under Vulkan, which was the only backend on it that reached those rates before.
Copy is the
out[i] = f(in[i])shape most elementwise kernels have, so about 1% is the realistic figure for those. Read is unchanged, which is expected: only the store side is stream sensitive.cubek kernels
cubek main pinned at this branch's base against this branch, RTX 4070 Ti SUPER, same day. Rows are one strategy on one shape; a negative change is faster.
The two large moves are gains: the autotuned gemm entry on the vector-matrix shapes, -13%, and the AlexNet-shaped conv2d, -8%. Unary at vec4 is 1.5% slower in all three rounds, the one row pointing the other way; at vec8 it is 5% faster in two of three.
Compatibility
ABSOLUTE_POSno longer linearizes(ABSOLUTE_POS_X, Y, Z). Both forms are bijections onto the same range, so a kernel computingf(ABSOLUTE_POS)covers identical indices and only which unit gets which index changes. The axes themselves are unchanged, and nothing outside their definition sites pairs them withABSOLUTE_POS.Tests
The existing topology test writes value
ito indexiand compares against0..n, which holds under any bijection, so it cannot see this. Added one that writesCUBE_POSto each unit'sABSOLUTE_POSslot and asserts contiguous runs per cube. It fails on main and passes here, on CUDA and on Vulkan.test_cmma_manualon sm_89)