Skip to content

[xpu-alignment] histc(out=) on XPU silently casts instead of rejecting a mismatched out dtype #5237

Description

@laifenxiawucha

Test bug skip issues

Independent local verification

XPU rerun 2026-09-07 (Max 1100, torch 2.15.0.dev20260902+xpu): float64 input with int64 out was accepted and silently cast. This supersedes the skipped behavior in #2234; that issue is linked for deduplication and intentionally unchanged.

The upstream result is still tracked as upstream-dependent; this issue records the concrete XPU defect and does not assert that upstream has finalized cross-backend semantics.

alignment scan 2026-09-05, run 34005971680

[xpu-alignment] histc(out=) on XPU silently casts instead of rejecting a mismatched out dtype

Upstream source

The PR body states explicitly that the XPU entry points at _histc_out_xpu, "a same-shaped
wrapper that lives out of tree in the xpu fork; it likely shares this gap and should get the
same fix there." This review confirms on hardware that it does.

Observed XPU behavior

torch.histc with an out= whose dtype differs from the input is accepted on XPU and the
floating-point histogram is silently truncated into the out dtype. CPU raises for the same
call today, so this is a current CPU-versus-XPU argument-validation divergence and does not
depend on PR 196100 landing.

CPU path: histc.out dispatches to histogram_histc_out ->
histogramdd_prepare_out (aten/src/ATen/native/Histogram.cpp:132), which raises
torch.histogram: input tensor and hist tensor should have the same dtype, ....

Reproducer

import torch

torch.manual_seed(0)
x = torch.randn(8, dtype=torch.float64, device="xpu")
out = torch.empty(4, dtype=torch.int64, device="xpu")

h = torch.histc(x, bins=4, min=-2.0, max=2.0, out=out)   # expected: RuntimeError
torch.xpu.synchronize()
print("returned without raising:", h.dtype, h.cpu().tolist())

matching = torch.empty(4, dtype=torch.float64, device="xpu")
h2 = torch.histc(x, bins=4, min=-2.0, max=2.0, out=matching)
print("matching out dtype:", h2.dtype, h2.cpu().tolist())

Observed output (runner/logs/pr-196100.log):

input dtype: torch.float64, mismatched out dtype: torch.int64
x: device=xpu:0, dtype=torch.float64
out: device=xpu:0, dtype=torch.int64
CALL: torch.histc(x, bins=4, min=-2.0, max=2.0, out=out)
returned without raising: dtype=torch.int64, values=[0, 3, 4, 0]
matching out dtype: returned dtype=torch.float64, device=xpu:0, values=[0.0, 3.0, 4.0, 0.0]

The same call on CPU raises
RuntimeError: torch.histogram: input tensor and hist tensor should have the same dtype, but got input double and hist long int.

Target-path evidence

  • native_functions.yaml (frozen head, line 9792) dispatches histc.out to
    XPU: _histc_out_xpu, implemented in src/ATen/native/xpu/SummaryOps.cpp.
  • _histc_out_xpu computes auto ret = _histc_xpu(...) in the input dtype, then does
    at::native::resize_output(result, ret.sizes()); result.copy_(ret);. copy_ casts
    silently, so any out dtype is accepted. There is no TORCH_CHECK on
    self.dtype() == result.dtype().
  • Runtime proof the XPU path ran: both operands are xpu:0 tensors, the histogram was
    actually computed ([0, 3, 4, 0], matching the float64 result [0.0, 3.0, 4.0, 0.0]
    from the matching-dtype call), and the return dtype is the mismatched int64.

Suggested fix

Add the CPU/CUDA check to _histc_out_xpu in src/ATen/native/xpu/SummaryOps.cpp, with the
same message text upstream uses:

TORCH_CHECK(
    self.dtype() == result.dtype(),
    "torch.histogram: input tensor and hist tensor should",
    " have the same dtype, but got input ", self.dtype(),
    " and hist ", result.dtype());

The fix should also remove "test_out_histc_xpu_float32" from
test/xpu/skip_list_common.py:183, which is currently skipped for exactly this gap.

Prior XPU history (not a live tracker)

#2234 reported the same defect through
TestCommonXPU::test_out_histc_xpu_float32 ("Expected RuntimeError when doing an unsafe cast
from a result of dtype torch.float32 into an out= with dtype torch.long"). It was closed as
completed by #3119, which only added the test to
skip_list_common.py on the stated grounds that "Both tests are expected to fail, same as on
CUDA." PR 196100 removes that premise by adding the check to CUDA, and the upstream test it
adds (test/test_reductions.py::test_histc_out_dtype) is device-generic apart from
@skipIfMPS, so XPU will start failing it once the PR lands. Issue 2234 is therefore not a
canonical tracker for this work: it is closed and its resolution was a skip, not a fix.

Ownership

The code change belongs in intel/torch-xpu-ops: _histc_out_xpu is implemented here and
PR 196100 only touches aten/src/ATen/native/cuda/SummaryOps.cu. No pytorch/pytorch issue
or PR claims the XPU side, and no open intel/torch-xpu-ops tracker covers it (searches for
histc, histc out dtype, and _histc_out_xpu returned only the closed items above and
unrelated histc work).

Environment

torch: 2.15.0.dev20260904+xpu (git_version 3780801c1d295481c4ddeb2d43049cd00f4c5697)
torch file: /opt/conda/lib/python3.13/site-packages/torch/__init__.py
python: 3.13.13 (/opt/conda/bin/python3.13)
xpu available: True, device_count: 8
xpu device 0: Intel(R) Arc(TM) Pro B60 Graphics

Evidence artifacts: scripts/repro_pr-196100.py
(sha256 c75c448132fd81362606622e032d5a6f077992270d996aae3493dbf99d9bdd69),
runner/logs/pr-196100.log
(sha256 3c3d468a291cc469ab31c216b0e2e4681d19ae87575d22aff2c3bc0d6be88024).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ai_generatedContent generated by an AI assistant

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions