Skip to content

Small optimizations to local_join kernels - #2751

Open
sherylll wants to merge 4 commits into
NVIDIA:mainfrom
sherylll:update-nnd-bbq
Open

sherylll wants to merge 4 commits into
NVIDIA:mainfrom
sherylll:update-nnd-bbq

Conversation

@sherylll

@sherylll sherylll commented Oct 6, 2026

Copy link
Copy Markdown
Contributor
  • Since u4 MMA only exists in a limited range of architectures, switch to native u8 MMA for better compatibility. Perf is not negatively affected because the kernel is latency bound.
  • In MMA kernels, add warp guard on MMA instructions for elements past the useful range (new_size/old_size)

@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Skip warps whose output block is out of range;
Remove redundant syncthreads in local_join_kernel_wmma
Simplify BBQ staging function
@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/cuvs/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: a4e4f8fa-fefe-4a14-89e2-11f8c7e9de71
📥 Commits

Reviewing files that changed from the base of the PR and between 3e6a7bb and bd372af.

📒 Files selected for processing (2)
  • cpp/src/neighbors/detail/nn_descent.cuh
  • cpp/src/neighbors/detail/nn_descent_gnnd.hpp
💤 Files with no reviewable changes (1)
  • cpp/src/neighbors/detail/nn_descent_gnnd.hpp

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Performance
    • Improved efficiency in accelerated neighbor search by skipping unnecessary computation for inactive processing tiles, including tiles beyond the active neighbor-list extents. Processing remains synchronized across participating workgroups. These optimizations apply to dense and BBQ local joins, including BBQ joins that process packed 4-bit values.

Walkthrough

NN-descent graph insertion now uses a fixed device-side degree. WMMA joins skip inactive tiles. The BBQ path derives self-join behavior from layouts and uses byte-fragment MMA to accumulate packed low and high nibbles.

Changes

NN-descent kernels

Layer / File(s) Summary
Fixed device-side graph degree
cpp/src/neighbors/detail/nn_descent.cuh, cpp/src/neighbors/detail/nn_descent_gnnd.hpp
Graph insertion uses DEGREE_ON_DEVICE instead of a graph-width argument. Kernel declarations and launches drop that argument. The two constants are removed from nn_descent_gnnd.hpp.
WMMA inactive-tile checks
cpp/src/neighbors/detail/nn_descent.cuh
WMMA joins skip fragment loads and stores for inactive tiles. Block-wide barriers remain unconditional.
BBQ SIMT self-join dispatch
cpp/src/neighbors/detail/nn_descent.cuh
The BBQ SIMT kernel derives self-join behavior from document and query layouts. Kernel calls and layout-pair dispatch no longer pass a self-join tag or graph width.
BBQ packed-nibble WMMA
cpp/src/neighbors/detail/nn_descent.cuh
The BBQ WMMA kernel stages promoted tiles across block threads and uses 16×16×16 byte fragments to accumulate low and high nibbles. It skips MMA work and stores for inactive tiles.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Refactor

Suggested reviewers: divyegala

Merge Risk: ⚪ Minimal · up to bd372

No specific issue has been established that would prevent merging after normal CUDA build and runtime checks.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the local_join kernel optimizations described in the changeset.
Description check ✅ Passed The description explains the switch to native u8 MMA and the addition of warp guards, which are central changes in the pull request.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant