Describe the bug
cuml.model_selection.KFold returns training indices in a different order from the equivalent sklearn.model_selection.KFold call when shuffle=True and the same random_state is used.
For the deterministic three-sample example below, the first split returned by cuML is:
while scikit-learn returns:
The train/test membership is the same in both implementations; the discrepancy is specifically the ordering of the training indices. The remaining two folds match exactly.
This difference can affect code that expects cuML's scikit-learn-compatible splitter API to produce the same deterministic index ordering for the same input, n_splits, shuffle, and random_state.
Steps/Code to reproduce bug
cuML reproducer:
import numpy as np
from cuml.model_selection import KFold
X = np.array([
[1, 2],
[3, 4],
[5, 6],
])
for train, test in KFold(
n_splits=3,
shuffle=True,
random_state=1,
).split(X):
print(train, test)
Output:
[2 1] [0]
[0 1] [2]
[0 2] [1]
For comparison, the equivalent scikit-learn code:
import numpy as np
from sklearn.model_selection import KFold
X = np.array([
[1, 2],
[3, 4],
[5, 6],
])
for train, test in KFold(
n_splits=3,
shuffle=True,
random_state=1,
).split(X):
print(train, test)
Output:
[1 2] [0]
[0 1] [2]
[0 2] [1]
Expected behavior
For the same input and deterministic parameters, cuml.model_selection.KFold should return training and test indices in an ordering compatible with scikit-learn.
For this reproducer, the expected output is:
[1 2] [0]
[0 1] [2]
[0 2] [1]
In particular, the first training split should be returned as [1, 2] rather than [2, 1].
If cuML intentionally does not guarantee the ordering of indices within each train/test split, that difference should be documented clearly because the returned index arrays are observably different even though the fold membership is identical.
Environment details (please complete the following information):
- Environment location: Docker
- Linux Distro/Architecture: Ubuntu 24.04 / x86_64
- GPU Model/Driver: NVIDIA GeForce RTX 4090 / 595.71.05
- CUDA: 13.2
- Method of cuDF & cuML install: conda
conda list:
conda list
# packages in environment at /opt/conda/envs/rapids-26.08:
#
# Name Version Build Channel
# Name Version Build Channel
python 3.14.6 h242f9ac_102_cp314 conda-forge
numpy 2.4.6 py314h2b28147_0 conda-forge
scipy 1.16.3 py314hf07bd8e_2 conda-forge
scikit-learn 1.9.0 np2py314hf09ca88_0 conda-forge
rapids 26.08.00 cuda13_260806_c2656556 rapidsai
cuml 26.08.00 cuda13_cp311_abi3_260805_265b9da6 rapidsai
libcuml 26.08.00 cuda13_260805_265b9da6 rapidsai
cudf 26.08.00 cuda13_cp311_abi3_260805_ff5b362d rapidsai
libraft 26.08.00 cuda13_260805_ebf92684 rapidsai
libraft-headers 26.08.00 cuda13_260805_ebf92684 rapidsai
pylibraft 26.08.00 cuda13_cp311_abi3_260805_ebf92684 rapidsai
cuvs 26.08.01 cuda13_cp311_abi3_260806_25b1be43 rapidsai
libcuvs 26.08.01 cuda13_260806_25b1be43 rapidsai
cupy 14.1.1 py314hdea9c46_0 conda-forge
cupy-core 14.1.1 py314hcd3b49b_0 conda-forge
numba 0.64.0 py314h8169c2f_0 conda-forge
numba-cuda 0.30.4 py314h42812f9_0 conda-forge
rmm 26.08.00 cuda13_cp311_abi3_260805_42d059f1 rapidsai
librmm 26.08.00 cuda13_260805_42d059f1 rapidsai
cuda-version 13.3 hcbadf70_3 conda-forge
cuda-bindings 13.3.1 py314h42812f9_1 conda-forge
cuda-cudart 13.3.29 hecca717_0 conda-forge
cuda-nvrtc 13.3.33 hecca717_0 conda-forge
libcublas 13.6.0.2 h676940d_0 conda-forge
libcusolver 12.2.6.9 h676940d_0 conda-forge
libcusparse 12.8.2.51 hecca717_0 conda-forge
libcurand 10.4.3.29 h676940d_0 conda-forge
Additional context
The discrepancy is limited to the order of indices within the first training split:
cuML: train=[2, 1], test=[0]
scikit-learn: train=[1, 2], test=[0]
As sets, both training splits contain exactly the same samples {1, 2}. Therefore, this reproducer does not show an incorrect fold assignment; it shows an ordering incompatibility in the index arrays returned by split().
With shuffle=True, shuffling should determine which samples are assigned to each fold. Once the test fold is selected, the returned training indices should remain deterministic and compatible with scikit-learn for the same random seed. Here, cuML appears to preserve a shuffled ordering for the first training split rather than the ordering returned by scikit-learn.
This distinction matters for downstream code that directly compares split indices, serializes splits, uses the returned order to index data without an additional sort, or expects drop-in compatibility with scikit-learn's deterministic output.
Describe the bug
cuml.model_selection.KFoldreturns training indices in a different order from the equivalentsklearn.model_selection.KFoldcall whenshuffle=Trueand the samerandom_stateis used.For the deterministic three-sample example below, the first split returned by cuML is:
while scikit-learn returns:
The train/test membership is the same in both implementations; the discrepancy is specifically the ordering of the training indices. The remaining two folds match exactly.
This difference can affect code that expects cuML's scikit-learn-compatible splitter API to produce the same deterministic index ordering for the same input,
n_splits,shuffle, andrandom_state.Steps/Code to reproduce bug
cuML reproducer:
Output:
For comparison, the equivalent scikit-learn code:
Output:
Expected behavior
For the same input and deterministic parameters,
cuml.model_selection.KFoldshould return training and test indices in an ordering compatible with scikit-learn.For this reproducer, the expected output is:
In particular, the first training split should be returned as
[1, 2]rather than[2, 1].If cuML intentionally does not guarantee the ordering of indices within each train/test split, that difference should be documented clearly because the returned index arrays are observably different even though the fold membership is identical.
Environment details (please complete the following information):
conda list:Additional context
The discrepancy is limited to the order of indices within the first training split:
As sets, both training splits contain exactly the same samples
{1, 2}. Therefore, this reproducer does not show an incorrect fold assignment; it shows an ordering incompatibility in the index arrays returned bysplit().With
shuffle=True, shuffling should determine which samples are assigned to each fold. Once the test fold is selected, the returned training indices should remain deterministic and compatible with scikit-learn for the same random seed. Here, cuML appears to preserve a shuffled ordering for the first training split rather than the ordering returned by scikit-learn.This distinction matters for downstream code that directly compares split indices, serializes splits, uses the returned order to index data without an additional sort, or expects drop-in compatibility with scikit-learn's deterministic output.