Skip to content

Fix scalar CAGRA graph byte encoding - #2715

Open
nvzm123 wants to merge 6 commits into
NVIDIA:mainfrom
nvzm123:zackm_scalar_cagra_byte_encoding
Open

nvzm123 wants to merge 6 commits into
NVIDIA:mainfrom
nvzm123:zackm_scalar_cagra_byte_encoding

Conversation

@nvzm123

@nvzm123 nvzm123 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fix the scalar bytes supplied to cuVS when building a CAGRA graph for CPU HNSW search. The previous (byte) (value & 0xff) copy did not change the signed byte bits, while the cuVS BYTE matrix interprets them as unsigned. The quantizer now emits 0–127 directly and initializes per-dimension extrema correctly for all-negative values. The redundant full-vector copy is removed.

Two unit regressions and a live GPU-build/HNSW-search test cover negative and mixed-sign vectors, two segments, force merge, GPU-writer selection, self hits, uniqueness, and recall against exact Euclidean neighbors. The affected generated API pages are updated. This PR contains only the standalone scalar fix; it does not include PR #2476 or PR #2653.

Prewarmed search results

The changes in this PR were benchmarked on Jasper-10M in an integrated tree that also contained PR #2476 and PR #2653. Both codecs used the same integrated tree. These are integrated-stack results, not a measurement of the standalone PR head.

The data were 1536-dimensional Euclidean vectors with one segment and no force merge; graph/intermediate degrees were 32/48, with efSearch=1500 and topK=1500. For each of two fresh-JVM search repeats, index files were evicted, ANN vector and graph files were prewarmed, and 210 warmup plus 1,000 measured queries ran. The scalar original-float .vec file was left cold; indiscriminate index prewarming was disabled.

Codec Peak search RSS, repeats (GiB) Mean latency, repeats (ms/query) Recall@1500
Float CAGRA_HNSW 58.605 / 58.637 12.933 / 12.620 98.0034%
Scalar CAGRA_HNSW_SCALAR 15.694 / 15.691 8.326 / 8.312 97.8883%

Across those repeats, scalar had 42.93 GiB (73.2%) lower hot-process RSS and 34.9% lower mean latency. Most of the RSS difference was file-backed mmap data, not Java heap. The raw float .vec remains on disk and can become resident with blanket prewarming or exact-vector access. These are repeats on one built index per codec, not a confidence interval. Scalar also had a larger index (72.550 vs 58.207 GiB) and slower single-build indexing (332.236 vs 268.413 s).

This is a scalar-versus-float codec comparison, not an isolated search-speed gain caused by this PR: the search reader is unchanged. A separate clone-only ablation kept the corrected quantizer on both sides and restored only the old bit-preserving clone. Two cold scalar builds per variant measured peak build RSS of 96.976 / 96.893 GiB without the clone and 107.007 / 107.381 GiB with it, a median 10.26 GiB (9.6%) reduction. These Jasper results do not establish quality-equivalent scalar behavior on Deep1B-100M.

Validation and compatibility

On NVIDIA A10G with Java 22, Maven 3.9.16, and matching cuVS 26.12, the functional six-file patch in this PR passed Spotless and generated-page checks. The focused quantizer/GPU command ran 3 tests with no failures, errors, or skips; full Maven clean verify ran 350 tests with no failures/errors and 30 other skips. The live test now uses the same assumeTrue(isSupported()) prerequisite convention as the existing cuvs-lucene GPU tests; the recorded GPU run executed with no skips and verified GPU-writer selection. Follow-up cleanup names the quantization bound and refactors test constants for readability; full Maven was not rerun after that nonfunctional cleanup.

Fern site validation was unavailable locally because Node.js was absent. Existing native initial-dimension warnings were also observed in an unchanged main-branch CAGRA test.

The public quantizer byte values intentionally change. Codec identifiers, on-disk format version, and reader code are unchanged, so older indexes are expected to remain readable, though mixed-version reads were not separately tested. Newly built graph topology may differ. Non-Euclidean quality has not been validated.

@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/cuvs/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: d90e9b78-a7e2-40df-9ced-a2b3fb9041a2
📥 Commits

Reviewing files that changed from the base of the PR and between d6d9662 and bb2939a.

📒 Files selected for processing (2)
  • java/cuvs-lucene/src/main/java/com/nvidia/cuvs/lucene/AcceleratedHNSWUtils.java
  • java/cuvs-lucene/src/main/java/com/nvidia/cuvs/lucene/LuceneAcceleratedHNSWScalarQuantizedVectorsWriter.java

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Bug Fixes

    • Improved scalar quantization for dimensions containing only negative values, mapping values across the full unsigned 7-bit range (0–127).
    • Preserved quantized values when building and merging GPU-accelerated HNSW graphs.
    • Improved consistency of search results for scalar-quantized graphs, including after segment merges, with accurate ranking and fewer duplicate results.
  • Documentation

    • Clarified that scalar-quantized values use unsigned 7-bit encoding stored in Java bytes.

Walkthrough

Scalar quantization now encodes values across the unsigned 7-bit range. The scalar graph writer passes quantized byte vectors directly to CuVS. Tests check quantization output and graph recall before and after segment merging.

Changes

Scalar Quantization and Graph Construction

Layer / File(s) Summary
Unsigned scalar quantization
java/cuvs-lucene/src/main/java/com/nvidia/cuvs/lucene/AcceleratedHNSWUtils.java, java/cuvs-lucene/src/test/java/com/nvidia/cuvs/lucene/TestScalarQuantization.java, fern/pages/lucene_api/lucene-api-com-nvidia-cuvs-lucene-acceleratedhnswutils.md
The quantizer initializes per-dimension extrema with positive and negative infinity and clamps codes to 0–127. Documentation and tests describe the encoding and check negative-only and mixed-value dimensions.
Scalar vectors passed to graph construction
java/cuvs-lucene/src/main/java/com/nvidia/cuvs/lucene/LuceneAcceleratedHNSWScalarQuantizedVectorsWriter.java, fern/pages/lucene_api/lucene-api-com-nvidia-cuvs-lucene-luceneacceleratedhnswscalarquantizedvectorswriter.md
The writer passes scalar-quantized byte vectors to CuVS without converting their values. The API documentation source references are updated.
GPU graph recall across merge
java/cuvs-lucene/src/test/java/com/nvidia/cuvs/lucene/TestScalarQuantizedCagraGraph.java
The GPU-backed test checks recall before and after merging two segments. It also checks GPU writer openings, result count, uniqueness, ranking, and exact-neighbor overlap.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix

Suggested reviewers: mythrocks, imotov

Merge Risk: ⚪ Minimal · up to bb293

The scalar graph path uses unsigned 0–127 codes, which CuVS consumes as unsigned values. No concrete merge-blocking issue remains; mixed-version compatibility is not guaranteed for this experimental format.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 28.57% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 14 functions across 4 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check Passed The title clearly and concisely identifies the main change: fixing scalar CAGRA graph byte encoding.
Description check Passed The description is directly related to the changeset and explains the byte encoding fix, quantizer corrections, tests, validation, and compatibility considerations.
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@imotov imotov assigned imotov and nvzm123 and unassigned imotov Oct 8, 2026
@imotov imotov added bug Something isn't working non-breaking Introduces a non-breaking change Lucene labels Oct 8, 2026
@imotov

imotov commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

/ok to test d6d9662

@imotov imotov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch! Thanks for fixing it!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working Lucene non-breaking Introduces a non-breaking change

Projects

Status: Approved

Development

Successfully merging this pull request may close these issues.

2 participants