Skip to content

fix(topology): a cube's units are contiguous in ABSOLUTE_POS - #1595

Merged
ThierryCantin-Demers merged 5 commits into
mainfrom
fix/store-codegen
Sep 3, 2026
Merged

ThierryCantin-Demers merged 5 commits into
mainfrom
fix/store-codegen

Conversation

@ThierryCantin-Demers

@ThierryCantin-Demers ThierryCantin-Demers commented Sep 2, 2026 •

Copy link
Copy Markdown
Member

CubeDim::new always returns a 2D cube, and ABSOLUTE_POS is linearized grid-major, so its y term strides the whole grid. For a cube of (32, 8, 1) on a grid of (2112, 1, 1) it comes out as ty*(gridDim.x*32) + bx*32 + tx, and the eight planes of one cube land on eight streams 1.03 MiB apart instead of one contiguous run.

Each warp is still fully coalesced, so the generated code is unchanged. What changes is how many concurrent streams reach DRAM, and stores are far more sensitive to that than loads.

This linearizes it as cube_pos() * CUBE_DIM + UNIT_POS in all four backends.

Measured

RTX 4070 Ti SUPER, the throughput example, three runs each:

before after
memory write 512 MiB 581.4 GB/s 596.4 +2.6%
memory copy 1 GiB 592.1 GB/s 597.3 +0.9%
memory read 512 MiB 638.9 GB/s 638.4 unchanged

RTX 2060, same example, three runs each, main against this branch on the same day. This card has a 3 MiB L2 against the 4070's 48 MiB, and pays far more for scattered streams:

before after
memory write 512 MiB 255.7 to 256.9 GB/s 284.9 to 285.9 +11%, 76% of bus to 85%
memory copy 1 GiB 251.3 to 252.3 GB/s 269.2 to 269.7 +7%
memory read 512 MiB 305.3 to 307.8 GB/s 319.3 +4%

The after column is within two points of what the same card reports under Vulkan, which was the only backend on it that reached those rates before.

Copy is the out[i] = f(in[i]) shape most elementwise kernels have, so about 1% is the realistic figure for those. Read is unchanged, which is expected: only the store side is stream sensitive.

cubek kernels

cubek main pinned at this branch's base against this branch, RTX 4070 Ti SUPER, same day. Rows are one strategy on one shape; a negative change is faster.

bench rows median range runs per side
reduce 132 0.0% -0.7% to +2.6% 1
gemm, tensor-core strategies on 4096³ f16/f32 and a skinny shape 12 +0.0% -0.1% to +0.4% 2
gemv 22 +0.2% -13.4% to +3.1% 1
conv2d 2 -0.7% -8.1% to -0.7% 1
unary 3 +1.4% -5.3% to +1.8% 3

The two large moves are gains: the autotuned gemm entry on the vector-matrix shapes, -13%, and the AlexNet-shaped conv2d, -8%. Unary at vec4 is 1.5% slower in all three rounds, the one row pointing the other way; at vec8 it is 5% faster in two of three.

Compatibility

ABSOLUTE_POS no longer linearizes (ABSOLUTE_POS_X, Y, Z). Both forms are bijections onto the same range, so a kernel computing f(ABSOLUTE_POS) covers identical indices and only which unit gets which index changes. The axes themselves are unchanged, and nothing outside their definition sites pairs them with ABSOLUTE_POS.

Tests

The existing topology test writes value i to index i and compares against 0..n, which holds under any bijection, so it cannot see this. Added one that writes CUBE_POS to each unit's ABSOLUTE_POS slot and asserts contiguous runs per cube. It fails on main and passes here, on CUDA and on Vulkan.

suite result
cubecl-cuda runtime tests, RTX 4070 Ti SUPER 933 pass, 17 fail: the same 17 as the base commit (16 complex math, test_cmma_manual on sm_89)
cubecl-cpu runtime tests 766 pass
cubecl-wgpu runtime tests, Vulkan 743 pass
burn main against this branch builds
cubek main against this branch builds; kernels benchmarked above

A cube dim of (32, 8, 1), which `CubeDim::new` returns for every 256-unit
launch, put the eight planes of a cube `CUBE_COUNT_X * CUBE_DIM_X` elements
apart under the old linearization: one cube's stores went to eight streams a
megabyte apart rather than to one contiguous run. On an RTX 4070 Ti SUPER that
costs the write path 2.4%, measured against the same kernel hand-written in
CUDA, and the read path nothing.

The bijection onto `0..total_units` holds either way, and a plane stays
contiguous in both, so `testgen_topology` and plane operations are unaffected.
What changes is that ABSOLUTE_POS no longer linearizes the
(ABSOLUTE_POS_X, Y, Z) coordinate: it is now the cube index times the cube
size plus the unit index within the cube.
The existing test scatters value i to index i and compares against 0..n,
which holds under any bijection, so it cannot see a cube whose units are
spread across the grid.
@ThierryCantin-Demers ThierryCantin-Demers changed the title Linearize ABSOLUTE_POS cube-major perf: stop ABSOLUTE_POS scattering a cube across the grid Sep 2, 2026
@ThierryCantin-Demers ThierryCantin-Demers changed the title perf: stop ABSOLUTE_POS scattering a cube across the grid fix(topology): a cube's units are contiguous in ABSOLUTE_POS Sep 2, 2026

@louisfd louisfd left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Image

@ThierryCantin-Demers
ThierryCantin-Demers merged commit 1f8b3dc into main Sep 3, 2026
6 checks passed
@ThierryCantin-Demers
ThierryCantin-Demers deleted the fix/store-codegen branch September 3, 2026 21:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants