Skip to content

Check noise shape and contiguity in rrelu_with_noise_kernel - #5241

Open
cyyever wants to merge 8 commits into
intel:mainfrom
cyyever:fix/rrelu-with-noise-xpu-shape-check
Open

cyyever wants to merge 8 commits into
intel:mainfrom
cyyever:fix/rrelu-with-noise-xpu-shape-check

Conversation

@cyyever

@cyyever cyyever commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

_rrelu_with_noise_xpu_train writes self.numel() distinct values into noise directly. Nothing checked that noise's shape matches self's, and for a non-contiguous noise (e.g. an expanded view) the kernel's .contiguous() temporary was written into instead and discarded, so the caller's noise tensor silently kept its original values with no error.

Adds the same shape check the CPU kernel already has, and writes noise through a contiguous temporary that is copied back into the caller's noise when it isn't contiguous -- mirroring the output handling already in this kernel and the CPU fix in pytorch/pytorch#196087.

Test Plan: no XPU device available where this was written; offered for CI.

This PR was authored with the assistance of an AI coding assistant.

@guangyey guangyey left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall LGTM. Could you please add a test to test/regressions/

@cyyever
cyyever force-pushed the fix/rrelu-with-noise-xpu-shape-check branch from 69ffe97 to 335ec76 Compare September 20, 2026 01:04
@github-actions github-actions Bot added disable_e2e Disable all e2e test jobs for the PR disable_distributed Disable distributed UT test jobs for the PR labels Sep 20, 2026
@cyyever
cyyever force-pushed the fix/rrelu-with-noise-xpu-shape-check branch from 335ec76 to 70030ea Compare September 20, 2026 01:08
@cyyever
cyyever requested a review from guangyey September 23, 2026 01:05
@guangyey
guangyey requested a review from jianyizh September 23, 2026 02:34
Comment thread test/regressions/test_rrelu_with_noise.py Outdated
@guangyey

Copy link
Copy Markdown
Contributor

please ensure the ut pass.

@guangyey

guangyey commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

@torchxpubot ut-check

@torchxpubot

Copy link
Copy Markdown
Contributor

Replying to this comment by @guangyey

UT Result Check: PR #5241

New Failures

36 new failure(s) detected (not in known issues). Listed below with the PR-related ones first; see the truncation note for the remainder.

Test Category Status Related to PR?
TestRreluWithNoise::test_noise_written op_regression failed Related
TestRreluWithNoise::test_noise_non_contiguous op_regression failed Related
TestRreluWithNoise::test_noise_shape_mismatch op_regression failed Related
CppSerdesTestExport::test_flex_attention_export_cpp_serdes op_ut failed Unrelated
StrictExportTestExport::test_flex_attention_export_strict op_ut failed Unrelated
RetraceExportNonStrictTestExport::test_flex_attention_export_retraceability_nonstrict op_ut failed Unrelated
RetraceExportTestExport::test_flex_attention_export_retraceability_strict op_ut failed Unrelated
SerDesExportNonStrictTestExport::test_flex_attention_export_serdes_nonstrict op_ut failed Unrelated
SerDesExportTestExport::test_flex_attention_export_serdes_strict op_ut failed Unrelated
StrictExportV2TestExport::test_flex_attention_export_strict_export_v2 op_ut failed Unrelated
TestAOTExport::test_aot_export_module_joint op_ut failed Unrelated
TestFunctionalizeXPU::test_functionalize_fx_reapply_views_simple_xpu op_ut failed Unrelated
TestFunctionalizeCPU::test_functionalize_fx_reapply_views_simple_cpu op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_op_has_batch_rule_log_normal_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_vmap_exhaustive_cauchy_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_op_has_batch_rule_cauchy_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_op_has_batch_rule_nn_functional_interpolate_nearest-exact_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_vmap_exhaustive_exponential_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_op_has_batch_rule_exponential_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_vmap_exhaustive_geometric_xpu_float32 op_ut failed Unrelated

... and 16 more failure(s). See CI logs for the full list. All 16 are Unrelated: 6 more TestVmapOperatorsOpInfoXPU vmap-randomness / vmap-fallback cases, 8 TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_*, TestExecutionTraceXPU::test_execution_trace_env_enabled_with_pt2_xpu, and TestFX::test_wrap_does_not_keep_caller_frames_alive.

Failure Relevance Analysis

Related (3) -- the PR's own new tests. All three fail at the same point with AttributeError: module 'torch' has no attribute 'rrelu_with_noise'. This is a test-harness bug, not a kernel bug: rrelu_with_noise is not exposed on the torch top-level namespace. It is reachable as torch._C._nn.rrelu_with_noise or torch.ops.aten.rrelu_with_noise (the existing references in test/xpu/test_nn_xpu.py are all to the aten:: op, never to a torch. attribute). Because the tests die on attribute lookup, the modified kernel in src/ATen/native/xpu/sycl/RreluWithNoiseKernels.cpp is exercised by zero passing assertions in this run -- the change is effectively unvalidated.

Unrelated (33), grouped by shared root cause:

  • 7 test_flex_attention_export_* + test_aot_export_module_joint + 2 test_functionalize_fx_reapply_views_simple_*: expected-graph string mismatches from upstream FX/export tracing changes (_remove_batch_dim/alias numbering, copy_ vs view). Pure upstream churn.
  • 12 TestVmapOperatorsOpInfoXPU cases: vmap: called random operation while in randomness error mode and hit the vmap fallback which is currently disabled (_upsample_nearest_exact1d.vec, hash_tensor, nonzero_static). Upstream functorch coverage gaps.
  • 8 test_linalg_batched_lu_stability_large_inputs_*: Cannot set preferred backend to cuSOLVER if PyTorch has not been compiled with cuSOLVER -- a CUDA-only test not properly gated for XPU.
  • test_wrap_does_not_keep_caller_frames_alive: missing libtorchbind_test.so, an environment/packaging issue.
  • test_execution_trace_env_enabled_with_pt2_xpu: profiler scalar mismatch, known flake territory.

None of these touch RReLU, activation kernels, or the files this PR changes.

New Test Coverage

This PR adds/modifies 3 test(s): 0 passed, 3 failed, 0 skipped, 0 not run.

New/Modified Test Status
TestRreluWithNoise::test_noise_written FAILED
TestRreluWithNoise::test_noise_non_contiguous FAILED
TestRreluWithNoise::test_noise_shape_mismatch FAILED

Recommendation

Not safe to merge. All three newly added tests in test/regressions/test_rrelu_with_noise.py fail on AttributeError: module 'torch' has no attribute 'rrelu_with_noise', so the kernel change has no passing verification behind it -- switch the call to torch.ops.aten.rrelu_with_noise (or torch._C._nn.rrelu_with_noise) and re-run. The other 33 new failures are upstream-churn, CUDA-gating, and environment issues with no connection to this PR.

Custom skills applied: ut-check.


Generated by ut-check job.

@guangyey

guangyey commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

@cyyever The CI failure seems to be related.

cyyever and others added 8 commits October 6, 2026 08:36
_rrelu_with_noise_xpu_train writes self.numel() distinct values into
noise directly. Nothing checked that noise's shape matches self's, or
that noise is contiguous; a non-contiguous noise (e.g. an expanded
view) got a noise_.contiguous() temporary written into instead and
discarded, so the caller's noise tensor silently kept its original
values. Add the same shape check the CPU kernel has, and require noise
be contiguous so the kernel can write into it directly. CUDA checks
neither: its train kernel still writes into a discarded
noise_.contiguous() temporary, so a non-contiguous noise silently
keeps its old values there as well.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Exercises the checks added in the previous commit: shape-mismatched
and non-contiguous noise must raise, and a contiguous noise must be
written by the training kernel.

No XPU device locally; verified with ruff:

```
ruff check test/regressions/test_rrelu_with_noise.py
ruff format --check --line-length 120 test/regressions/test_rrelu_with_noise.py
```

Co-Authored-By: Claude Code <noreply@anthropic.com>
Materializing non-contiguous noise introduced an unchecked copy_ back
into the caller's tensor. An expanded noise aliases the same storage
across elements, so the copy silently wrote distinct values through
overlapping memory instead of failing as intended.

Test Plan:
pytest test/regressions/test_rrelu_with_noise.py -v
@cyyever
cyyever force-pushed the fix/rrelu-with-noise-xpu-shape-check branch from a198456 to 0cb8f03 Compare October 6, 2026 00:37
@cyyever

cyyever commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

@guangyey Fixed

@guangyey

guangyey commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@torchxpubot review

@guangyey

guangyey commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

@torchxpubot ut-check

@torchxpubot

Copy link
Copy Markdown
Contributor

Replying to this comment by @guangyey

Confirmed the key fact from upstream. Here is the report.

UT Result Check: PR #5241

New Failures

28 new failure(s) detected (not in known issues): 3 Related, 25 Unrelated.

Test Category Status Related to PR?
TestRreluWithNoise::test_noise_written op_regression failed Related
TestRreluWithNoise::test_noise_non_contiguous op_regression failed Related
TestRreluWithNoise::test_noise_shape_mismatch op_regression failed Related
MiscTests::test_copy_reduce_override op_ut failed Unrelated
DynamicShapesMiscTests::test_copy_reduce_override_dynamic_shapes op_ut failed Unrelated
TestFX::test_wrap_does_not_keep_caller_frames_alive op_ut failed Unrelated
test_meta_xpu (collection error) op_ut error Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_False_xpu_float32 op_ut failed Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_False_xpu_float64 op_ut failed Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_False_xpu_complex64 op_ut failed Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_False_xpu_complex128 op_ut failed Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_True_xpu_float32 op_ut failed Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_True_xpu_float64 op_ut failed Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_True_xpu_complex64 op_ut failed Unrelated
TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_pivot_True_xpu_complex128 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_vmap_exhaustive_cauchy_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_vmap_exhaustive_exponential_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_op_has_batch_rule_hash_tensor_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_op_has_batch_rule_nonzero_static_xpu_float32 op_ut failed Unrelated
TestVmapOperatorsOpInfoXPU::test_op_has_batch_rule_nn_functional_interpolate_nearest-exact_xpu_float32 op_ut failed Unrelated

... and 8 more failure(s), all in TestVmapOperatorsOpInfoXPU with the same vmap-randomness error mode signature. See CI logs for the full list.

Failure Relevance Analysis

Related (3) — the PR's own new tests, all failing with the same error:

AttributeError: module 'torch' has no attribute 'rrelu_with_noise'

This is a test-side API error, not a kernel defect. Verified against upstream aten/src/ATen/native/native_functions.yaml:12290: rrelu_with_noise is declared with python_module: nn, so the generated binding lives at torch._C._nn.rrelu_with_noise, never at torch.rrelu_with_noise. The same holds for the .out and in-place variants (lines 12282, 12305). The new test file needs to call torch._C._nn.rrelu_with_noise (or go through torch.nn.functional.rrelu).

The practical consequence: the three tests fail at attribute lookup before touching the device, so the kernel change in src/ATen/native/xpu/sycl/RreluWithNoiseKernels.cpp is currently unvalidated by CI — there is no evidence either for or against the kernel fix itself.

Unrelated (25) — environment and upstream-test issues, three clusters:

  • 13x TestVmapOperatorsOpInfoXPU — vmap randomness error mode on random ops (cauchy, log_normal, exponential, geometric, normal_in_place) plus missing batch rules (hash_tensor, nonzero_static, _upsample_nearest_exact1d.vec). Upstream functorch coverage gaps.
  • 8x TestLinalgXPU::test_linalg_batched_lu_stability_large_inputs_* — Cannot set preferred backend to cuSOLVER; an upstream test that is not guarded for non-CUDA builds.
  • 4x infrastructure — ModuleNotFoundError: No module named 'test.support' (2, a Python path collision between PyTorch's test/ package and the stdlib test package), missing libtorchbind_test.so (1), and a test_meta_xpu collection error for a nonexistent TestMeta.test_fill_stride (1).

None of these touch rrelu, activation kernels, or the regression-test harness.

Data note: the test_meta_xpu record has shifted fields (status holds the AttributeError text, message is "XML"), and passed_tests_count (222013) disagrees slightly with totals.passed (222026). Both are collection/reporting artifacts and do not change the analysis.

New Test Coverage

This PR adds/modifies 3 test(s): 0 passed, 3 failed, 0 skipped, 0 not run.

New/Modified Test Status
TestRreluWithNoise::test_noise_written FAILED
TestRreluWithNoise::test_noise_non_contiguous FAILED
TestRreluWithNoise::test_noise_shape_mismatch FAILED

Recommendation

Not safe to merge. Every test this PR adds fails, so the RreluWithNoiseKernels.cpp change has zero passing validation; the fix is small and confined to the test file — switch torch.rrelu_with_noise to torch._C._nn.rrelu_with_noise (confirmed from upstream native_functions.yaml, where the op carries python_module: nn) and re-run so the kernel change is actually exercised. The other 25 failures are pre-existing environment and upstream-test problems unrelated to this diff.


Generated by ut-check job.

@guangyey

guangyey commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

op_regression,third_party.torch-xpu-ops.test.regressions.test_rrelu_with_noise.TestRreluWithNoise,test_noise_non_contiguous
op_regression,third_party.torch-xpu-ops.test.regressions.test_rrelu_with_noise.TestRreluWithNoise,test_noise_shape_mismatch
op_regression,third_party.torch-xpu-ops.test.regressions.test_rrelu_with_noise.TestRreluWithNoise,test_noise_written
These failures seem to be related.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

disable_distributed Disable distributed UT test jobs for the PR disable_e2e Disable all e2e test jobs for the PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants