Repository navigation
Conversation
|
/ok to test fe21bf4 |
@achirkin, there was an error processing your request: See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/ |
|
/ok to test fe21bf4 |
achirkin
left a comment
There was a problem hiding this comment.
Thanks, the speedup looks promising! A small request below
| _max_iterations = minimum_depth + raft::ceildiv(mc_itopk_size - minimum_depth, num_ctas); | ||
| _max_iterations += raft::ceildiv(static_cast<size_t>(topk), mc_itopk_size) - 1; |
There was a problem hiding this comment.
Could you please add small comments explaining the logic for setting the max_iterations like this on these two lines?
…ed; add missing update of filtering rate in the multi-partition case
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughCAGRA now computes filtering rates from bitsets when no explicit rate is set. Automatic ChangesCAGRA search tuning
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to No search-impacting issue was established; this change is mergeable after normal checks. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
For multi-CTA mode, current
max_iterationsis assigned a large number. However, at lowitopk, extra iterations only add latency and does not do much to recall. So the intuition is to use a lowermax_iterationsfor smalleritopk.minimum_depth = 8 * (search_quality - 1), since we don't provide a knob forsearch_qualityfor the moment, hardcoding this to 16.search_quality = 5would mean the previous default 32.When the problem is harder, or when more results are needed, it is still a good idea to turn up the
max_iterations.Quality = 3 is the new default, and quality = 5 is the old default.
This is a subset of #2502, where the change to width or hashmap size seem to have different effect on different GPU generations, which still remains to be investigated. But the conclusion on
max_iterationsshould hold.