Skip to content

Add PyLucene backend for cuVS-Lucene vector search in cuvs-bench - #2385

Draft
nvzm123 wants to merge 27 commits into
NVIDIA:mainfrom
nvzm123:agent/pylucene-benchmark-backend
Draft

nvzm123 wants to merge 27 commits into
NVIDIA:mainfrom
nvzm123:agent/pylucene-benchmark-backend

Conversation

@nvzm123

@nvzm123 nvzm123 commented Jul 31, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR adds a built-in pylucene backend to cuVS Bench for building and searching local Lucene vector indexes through PyLucene and cuVS-Lucene.

It supports:

  • Lucene101AcceleratedHNSWCodec for cuVS-assisted HNSW construction with Lucene HNSW search;
  • CuVS2510GPUSearchCodec for GPU CAGRA construction and search;
  • packaged configurations, deterministic parameter sweeps, automatic tuning, dry runs, index reuse, forced rebuilds, query batching, and standard cuVS Bench CSV results; and
  • h5py as an explicit dependency for the existing dataset-preparation path.

Selecting --backend pylucene is sufficient to activate the backend. It resolves a matching local cuvs-java/cuvs-lucene JAR pair and native-library path from conventional build or Maven locations, while retaining paired explicit overrides for nonstandard installations. Missing or mismatched prerequisites fail before JVM startup with actionable messages.

This PR also owns the PyLucene end-to-end suite. Pytest owns its scenarios, assertions, parameterization, and reporting under cuvs_bench/tests/pylucene; reusable Python helpers remain non-test modules, and test-only Java adapters live under python/cuvs_bench/tests/java. Those adapters are compiled into a temporary directory before the process-wide JVM starts and are excluded from published packages.

HNSW parameters and topology

The HNSW backend exposes these benchmark parameters:

  • m maps to AcceleratedHNSWParams.maxConn;
  • ef_construction maps to AcceleratedHNSWParams.beamWidth;
  • both inputs use cuVS-Lucene's SAME_GRAPH_FOOTPRINT heuristic to derive the underlying CAGRA build parameters;
  • num_candidates controls Lucene's KnnFloatVectorQuery candidate budget and may be swept independently for each built index; and
  • direct_single_segment requests one direct Lucene segment without a force merge and verifies the committed topology.

Stock PyLucene instantiates codecs through no-argument constructors. The backend therefore compiles a small Java adapter before JVM startup that delegates to the production accelerated codec with the requested heuristic parameters. It does not reimplement the production codec.

Automatic tuning covers m, ef_construction, and num_candidates, with the candidate lower bound resolved from top_k. Build parameters and segment count are recorded in schema-v4 provenance so incompatible indexes cannot be silently reused.

Runtime and integrity behavior

The backend validates the PyLucene 10.2 ABI, configured JAR roles and matching Maven versions, native-library paths, Lucene SPI codecs, dataset shape, and index location before execution. It requires the base cuvs-java JAR and standard thin cuvs-lucene JAR, rejects native-classifier/fat artifacts, and initializes PyLucene and the JVM lazily.

HNSW preserves cuVS-Lucene's production behavior: it uses the accelerated writer when available and intentionally falls back to Lucene's CPU writer otherwise. CAGRA requires GPU support and is validated fail-closed. A generic cuVS-unavailable CAGRA failure is augmented with the relevant PyLucene/JAR/native-library setup checks without changing unrelated Java errors. CAGRA segment metadata, vector dimensions and counts, file coverage, headers, footers, and checksums are verified before results are accepted.

Atomic, commit-bound provenance protects index reuse. Failed new builds remove only their partial output, and preflight failures preserve an existing index. --force removes the prior index before replacement, so a failed forced rebuild leaves no previous index to reuse. In-process backend results use the shared direct CSV exporter and preserve build/search identity, latency percentiles, and failure handling; result filenames and CSV rows retain algorithm, group, and optional scope identity.

Requirements and limits

  • FLOAT32 Euclidean/L2 datasets, at most 4096 dimensions, and at least two indexed vectors
  • latency mode with one search thread; query batching through --batch-size
  • num_candidates >= top_k for HNSW; this is Lucene's candidate budget, not a direct cuVS ef_search setting
  • CAGRA searches with k <= 1024
  • direct_single_segment remains subject to Lucene's per-indexing-thread hard RAM limit and fails if Lucene commits more than one segment
  • process-wide PyLucene JVM settings cannot change after initialization

PyLucene 10.2 remains a source-built external dependency. This PR depends on the production codec and compatibility changes in NVIDIA/cuvs#2475 until that PR is merged. The Bench guide pins the exact tested producer revision rather than a moving PR head.

Test coverage

Non-live pytest coverage includes backend registration, zero-boilerplate runtime discovery, paired override validation, configuration expansion and tuning, parameter validation, adapter compilation and classpath failures, JVM and SPI validation, configured-artifact role validation, index build/reuse/cleanup, single- and multi-segment topology, search candidate propagation, result export, provenance mismatches, CAGRA corruption detection, built-wheel resource ownership, and CLI failure handling.

The opt-in live matrix explicitly distinguishes:

  • CPU HNSW build and HNSW search;
  • GPU CAGRA build followed by one-layer or three-layer HNSW search; and
  • GPU CAGRA build and CAGRA search.

GPU-required cases assert the accelerated writer, reader, and query implementations and fail on unavailable cuVS or CPU fallback. CPU cases explicitly verify and report stock Lucene HNSW execution.

Coverage includes a single live document, one and ten segments, 10-to-1 and 100-to-10 force merges, CAGRA searchWidth values 1, 16, and 32, deletion, vectorless documents, selective filtering, persisted HNSW graph degree/layers, an exact Lucene filtered-search boundary, deterministic brute-force recall, duplicate-hit exclusion, inactive/filter-rejected document exclusion, and rank-one self matches. CAGRA-built HNSW also covers top_k=2000 with num_candidates=2500; direct CAGRA search retains its supported k <= 1024 boundary. The Java-side query bridge verifies the retained production searchWidth value rather than echoing Python configuration.

CAGRA configurations use graphDegree=32, intermediateGraphDegree=64, and enough documents to avoid cuVS clamping warnings, including 24,832 vectors for the three-layer case. Every live path case captures process output and fails on the known cuVS graph-clamping diagnostics before printing a clear CPU HNSW, GPU/CAGRA-built HNSW, or GPU CAGRA-search label.

Artifact gates require the production Java adapter and PyLucene YAML resources in built distributions while rejecting test Java, PyLuceneTestSupport, and .class payloads. The wheel-content test runs in the standard Python test environment with CUDA disabled, so it does not require nvcc. A Bench-owned formatting POM, the pre-commit matcher, and the Spotless wrapper cover both production and test Java roots without adding a Maven artifact to the package tree.

Validation

Final revisions:

Validated on an NVIDIA A10G with JDK 22 and Lucene/PyLucene 10.2. The GPU run used the official RAPIDS 26.12 development nightly libcuvs 26.12.00a8 (cuda12_260910004918_29e1101b), built from the tested main revision 29e1101b merged into both branches. The runtime reported cuVS 26.12.0 and used the RMM 26.12 ABI; the cuvs-java and cuvs-lucene 26.12.0 JARs were rebuilt from the merged producer branch.

After that full run, PR #2475 merged current main (ef29c4cfd53082d31c1fc05ec35251c835120efa) in 57ce4920374dce1cecd09091566b18710fe01c82. That merge adds only the four upstream CMake dependency-discovery changes and leaves the validated Java, cuvs-lucene, PyLucene, and test inputs unchanged; git diff --check passed, and the GPU suite was not redundantly rerun.

Producer validation

  • full GPU-enabled cuvs-java and cuvs-lucene Maven suites: 497 tests, 0 failures, 0 errors, 29 skipped (cuvs-java: 112 with 1 skipped; cuvs-lucene: 385 with 28 skipped, including both passing ThinJarContentsIT cases)

cuVS Bench validation

  • python -m pytest -q -s cuvs_bench/tests/pylucene --run-pylucene: 393 passed in 55.05 seconds
  • relevant shell syntax checks: passed

The Java suites exercised the native GPU paths and emitted no version-mismatch or linkage errors. Their inherited randomized and small-dataset cases emitted native cuVS graph-parameter diagnostics (271 warning lines from cuvs-java and 2,651 from cuvs-lucene). The PyLucene suite emitted no cuVS configuration or CPU-fallback warnings; it emitted the expected JVM notice for the incubating vector module.

Related work

nvzm123 added 5 commits July 31, 2026 13:27
Dataset preparation imports h5py at runtime. Declare it in both the dependency manifest and project metadata so supported environments install it consistently.
Add opt-in atomic JSON persistence for Python-native backends, preserve canonical result fields, and keep derived CSV artifacts synchronized. Propagate sweep failures through the CLI and cover result identity, export, and cleanup behavior.
Register a local PyLucene backend for cuVS-Lucene HNSW and CAGRA codecs. Add deterministic config selection, lazy JVM and codec resolution, GPU writer validation, safe index lifecycle handling, commit-bound provenance, CAGRA integrity checks, and focused unit coverage.
Exercise real JVM and cuVS-Lucene HNSW and CAGRA build/search paths behind an opt-in pytest marker. Cover persisted GPU formats, index reuse, CLI execution, fallback rejection, and integrity failures.
Document the verified dependency build, runtime configuration, supported codecs and limits, manual smoke workflow, index reuse behavior, and benchmark result semantics.
@copy-pr-bot

copy-pr-bot Bot commented Jul 31, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@cjnolet cjnolet added improvement Improves an existing functionality non-breaking Introduces a non-breaking change labels Aug 1, 2026
nvzm123 added 17 commits August 14, 2026 17:00
…ark-backend

Signed-off-by: nvzm123 <zmeeks@nvidia.com>

# Conflicts:
#	python/cuvs_bench/cuvs_bench/orchestrator/orchestrator.py
#	python/cuvs_bench/cuvs_bench/run/__main__.py
#	python/cuvs_bench/cuvs_bench/run/data_export.py
#	python/cuvs_bench/cuvs_bench/tests/test_data_export.py
Use the shared direct CSV exporter for in-process backends and remove the redundant JSON persistence layer. Preserve scoped result identities, latency percentiles, safe artifact paths, stale-result cleanup, and accurate CLI exit behavior.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Target the production HNSW and CAGRA codecs exposed by the current cuVS-Lucene PR. Preserve the intentional HNSW CPU fallback, keep fail-closed CAGRA integrity checks, and verify actual HNSW writer selection through a downstream test-only codec adapter.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Document the current JDK, PyLucene, cuVS Java, and cuVS-Lucene requirements. Describe the two supported production codecs, HNSW fallback behavior, CAGRA validation contract, and the current manual validation workflow.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Validate the generated PyLucene ABI before JVM startup and follow the current cuVS-Lucene codec and writer-selection contracts. Preserve Lucene defaults for HNSW while keeping CAGRA segment files directly verifiable across flushes and merges. Extend provenance and live coverage for GPU selection, CPU fallback, and merged CAGRA indexes.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Document the custom Lucene 10.2 wrapper requirement, temporary PR 174 artifact workflow, codec-specific compound-file policies, fallback behavior, and latency units.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Pass m and ef_construction through a PyLucene-compatible codec adapter, expose num_candidates sweeps, and support verified direct single-segment builds. Persist the complete build identity and extend unit and live coverage for writer selection, fallback, topology, tuning, and reuse.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Describe the validated cuVS-Lucene source combination and adapter requirements. Document supported HNSW parameters, tuning ranges, single-segment constraints, and a manual sweep workflow.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Use cuVS PR NVIDIA#2475 as the pinned source for matching native, cuvs-java, and cuvs-lucene artifacts. Refresh monorepo paths, validation commands, and adapter compatibility guidance.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Group PyLucene unit, runtime, provenance, and integration coverage under a dedicated test subtree. Preserve recursive discovery and update shared helper and project fixture paths.

Signed-off-by: nvzm123 <zmeeks@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

improvement Improves an existing functionality non-breaking Introduces a non-breaking change

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants