Skip to content

[BUG] KFold with shuffle=True returns training indices in a different order from scikit-learn #8631

Description

@apiqwe

Describe the bug

cuml.model_selection.KFold returns training indices in a different order from the equivalent sklearn.model_selection.KFold call when shuffle=True and the same random_state is used.

For the deterministic three-sample example below, the first split returned by cuML is:

[2 1] [0]

while scikit-learn returns:

[1 2] [0]

The train/test membership is the same in both implementations; the discrepancy is specifically the ordering of the training indices. The remaining two folds match exactly.

This difference can affect code that expects cuML's scikit-learn-compatible splitter API to produce the same deterministic index ordering for the same input, n_splits, shuffle, and random_state.

Steps/Code to reproduce bug

cuML reproducer:

import numpy as np
from cuml.model_selection import KFold

X = np.array([
    [1, 2],
    [3, 4],
    [5, 6],
])

for train, test in KFold(
    n_splits=3,
    shuffle=True,
    random_state=1,
).split(X):
    print(train, test)

Output:

[2 1] [0]
[0 1] [2]
[0 2] [1]

For comparison, the equivalent scikit-learn code:

import numpy as np
from sklearn.model_selection import KFold

X = np.array([
    [1, 2],
    [3, 4],
    [5, 6],
])

for train, test in KFold(
    n_splits=3,
    shuffle=True,
    random_state=1,
).split(X):
    print(train, test)

Output:

[1 2] [0]
[0 1] [2]
[0 2] [1]

Expected behavior

For the same input and deterministic parameters, cuml.model_selection.KFold should return training and test indices in an ordering compatible with scikit-learn.

For this reproducer, the expected output is:

[1 2] [0]
[0 1] [2]
[0 2] [1]

In particular, the first training split should be returned as [1, 2] rather than [2, 1].

If cuML intentionally does not guarantee the ordering of indices within each train/test split, that difference should be documented clearly because the returned index arrays are observably different even though the fold membership is identical.

Environment details (please complete the following information):

  • Environment location: Docker
  • Linux Distro/Architecture: Ubuntu 24.04 / x86_64
  • GPU Model/Driver: NVIDIA GeForce RTX 4090 / 595.71.05
  • CUDA: 13.2
  • Method of cuDF & cuML install: conda

conda list:

conda list
# packages in environment at /opt/conda/envs/rapids-26.08:
#
# Name                              Version          Build                                         Channel
# Name              Version       Build                                      Channel
python              3.14.6        h242f9ac_102_cp314                         conda-forge
numpy               2.4.6         py314h2b28147_0                            conda-forge
scipy               1.16.3        py314hf07bd8e_2                            conda-forge
scikit-learn        1.9.0         np2py314hf09ca88_0                         conda-forge
rapids              26.08.00      cuda13_260806_c2656556                     rapidsai
cuml                26.08.00      cuda13_cp311_abi3_260805_265b9da6          rapidsai
libcuml             26.08.00      cuda13_260805_265b9da6                     rapidsai
cudf                26.08.00      cuda13_cp311_abi3_260805_ff5b362d          rapidsai
libraft             26.08.00      cuda13_260805_ebf92684                     rapidsai
libraft-headers     26.08.00      cuda13_260805_ebf92684                     rapidsai
pylibraft           26.08.00      cuda13_cp311_abi3_260805_ebf92684          rapidsai
cuvs                26.08.01      cuda13_cp311_abi3_260806_25b1be43          rapidsai
libcuvs             26.08.01      cuda13_260806_25b1be43                     rapidsai
cupy                14.1.1        py314hdea9c46_0                            conda-forge
cupy-core           14.1.1        py314hcd3b49b_0                            conda-forge
numba               0.64.0        py314h8169c2f_0                            conda-forge
numba-cuda          0.30.4        py314h42812f9_0                            conda-forge
rmm                 26.08.00      cuda13_cp311_abi3_260805_42d059f1          rapidsai
librmm              26.08.00      cuda13_260805_42d059f1                     rapidsai
cuda-version        13.3           hcbadf70_3                                 conda-forge
cuda-bindings       13.3.1         py314h42812f9_1                            conda-forge
cuda-cudart         13.3.29        hecca717_0                                 conda-forge
cuda-nvrtc          13.3.33        hecca717_0                                 conda-forge
libcublas           13.6.0.2       h676940d_0                                 conda-forge
libcusolver         12.2.6.9       h676940d_0                                 conda-forge
libcusparse         12.8.2.51      hecca717_0                                 conda-forge
libcurand           10.4.3.29      h676940d_0                                 conda-forge

Additional context

The discrepancy is limited to the order of indices within the first training split:

cuML:         train=[2, 1], test=[0]
scikit-learn: train=[1, 2], test=[0]

As sets, both training splits contain exactly the same samples {1, 2}. Therefore, this reproducer does not show an incorrect fold assignment; it shows an ordering incompatibility in the index arrays returned by split().

With shuffle=True, shuffling should determine which samples are assigned to each fold. Once the test fold is selected, the returned training indices should remain deterministic and compatible with scikit-learn for the same random seed. Here, cuML appears to preserve a shuffled ordering for the first training split rather than the ordering returned by scikit-learn.

This distinction matters for downstream code that directly compares split indices, serializes splits, uses the returned order to index data without an additional sort, or expects drop-in compatibility with scikit-learn's deterministic output.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions