Skip to content

[bot] Automate bot fix on CI DISABLED issues - #5473

Open
ZhaoqiongZ wants to merge 36 commits into
mainfrom
bot/ci-disabled-queue
Open

ZhaoqiongZ wants to merge 36 commits into
mainfrom
bot/ci-disabled-queue

Conversation

@ZhaoqiongZ

@ZhaoqiongZ ZhaoqiongZ commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

1. Automate @torchxpubot fix on CI DISABLED issues. #5272 is the queue SOURCE is #5490,

  • per-open-episode dedup -- deduplicate the issues in case same issue in several comments
  • serialization -- a fix , wait for result, then another fix.
  • one batch per trigger, first in first out
    The cron is on: 0 0,9-23 * * *, hourly over 17:00-08:00 Beijing

2. Tracking issues for skip-adding UT: template and reference format.

3. Fix the patch collision issue when fix in both repo with same name/file.

Authored with an AI assistant.

@ZhaoqiongZ
ZhaoqiongZ marked this pull request as draft September 22, 2026 02:43
@ZhaoqiongZ
ZhaoqiongZ force-pushed the bot/ci-disabled-queue branch from 82c3e33 to 6d9d8d7 Compare September 22, 2026 02:45
Every `@torchxpubot fix` on #5272 is hand-copied from a batch comment on
chuanqi129/pytorch-xpu-ci#364. The copying is mechanical, but two things
about it are not, which is why it stayed manual: #364 re-lists the same
upstream issue on every later failing commit (pytorch#197144 appears eight
times), and a `fix` run costs a GPU and a couple of hours, so two triggers
must not overlap.

Both are decidable from GitHub state, so the queue keeps no database. An
issue is skipped if it is closed upstream, or if it has stayed open since
the last `@torchxpubot fix` comment naming it. Dedup is per open episode
rather than forever: pytorch#194562 was mirrored on 08-24, closed on 09-03,
reopened on 09-16 and re-reported the same day -- that is a fresh breakage
and gets queued again. `closed_at` cannot see this, since reopening clears
it, so the issue's `closed` events are read instead. Serialization is a
check for a `fix` job that is queued or running; the job's own 300min
timeout bounds the wait, so no staleness escape hatch is needed.

bot.yml's `concurrency: bot-<issue>` cannot do the serializing: one group
holds a single pending run, so posting several comments at once would get
all but the last cancelled.

START is where the queue takes over from the humans, and it has to exist
because "a human already has a PR for this" is not recorded anywhere GitHub
can be asked -- these issues carry no assignee and no linked pull request,
their cross-reference events are both incomplete (pytorch#195558 has no xref
to its fix pytorch#195534) and noisy (an unrelated ROCm PR appears on two of
them), and the ownership table that does have it is a spreadsheet. So the
queue owns nothing before START: that backlog was triaged by hand, and the
leftovers with no PR and no owner are being triggered by hand too.

One batch per run, oldest first, so a backlog drains a batch per completed
fix instead of all at once. The cron covers 17:00-08:00 Beijing, when the
xpu-agent runners are idle.

Test Plan

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

The self-test covers the parser, the trigger body, and per-episode dedup. It
runs on every PR touching these two files.

Dry-run against live #364 and #5272 reports an empty queue, which is correct
for START=2026-09-22: #364's last batch comment is 09-18. Moving START back
to 09-18 in a scratch copy selects that batch and nothing else:

```
GH_TOKEN=... python3 .github/scripts/ci_disabled_queue.py --dry-run
-> nothing to queue since 2026-09-22; 9 issues mirrored so far
START=2026-09-18 -> 09-18 [197521 197522 197523 197524]
```

A steady-state run costs 6 API calls.

Not verified here, both need a run on the default branch: that MERGE_TOKEN
can read the private chuanqi129/pytorch-xpu-ci, and that a comment posted
by torchxpubot passes `fix`'s OWNER/MEMBER gate. torchxpubot is an intel
org member, so the webhook's `author_association` should be MEMBER, but the
REST API reports CONTRIBUTOR for the same user and the two disagree.

Authored with an AI assistant.
The queue's one unverifiable-from-a-branch dependency: it reads issue
comments from a private repo owned by a personal account, so whether
MERGE_TOKEN can reach it decides whether the workflow works at all, and
neither the token's kind nor its grants are visible from the outside.

Probe it on pull_request the way bot.yml's token-permission-test probes its
own token needs. The failure message carries the likely diagnosis, since a
404 with torchxpubot already a collaborator points at the token being
fine-grained rather than at a missing grant.
…s reachable

MERGE_TOKEN turns out to be a fine-grained PAT, so discovery cannot work
yet: fine-grained tokens are scoped to one resource owner, and SOURCE lives
in a repo owned by another personal account. Adding torchxpubot as a
collaborator there does not help and did not -- the token-access job gets a
404 with the grant in place. A cron would 404 hourly, so it is held back
rather than shipped broken, and the probe now warns instead of failing: it
turns green by itself when the report issue moves into this repo, which is
the signal to restore it. Nothing but the SOURCE_* constants changes then.

Meanwhile ISSUES queues an operator-supplied list with no SOURCE read at
all, which is what drains the triaged backlog and what proves a comment
authored by torchxpubot clears the fix command's OWNER/MEMBER gate.

Both dispatch inputs travel through the environment rather than the run
string, since an interpolated workflow_dispatch input is an injection
vector, and only the digits of ISSUES are kept. The self-test pins that:
a hostile value reduces to its issue numbers and nothing else.
An unset or misspelled DRY_RUN used to fall through to posting, which is the
wrong way round for a switch whose live side comments on a tracking issue
and starts a multi-hour GPU job. Only the exact string "false" posts now,
so a restored cron has to say so outright rather than inherit an empty
workflow_dispatch input.
A bare issue URL makes GitHub file a cross-referenced event on the target,
so every trigger leaves an `#5272` backlink on the
upstream issue. All nine posted by hand did it, and at one trigger per batch
forever that is a steady drip of noise onto someone else's tracker -- the
same noise that made cross-references useless as a has-a-PR signal earlier
in this work, where an unrelated ROCm PR showed up on two XPU issues.

Wrap the URLs in backticks. Code spans are not scanned for references, and
ISSUE_RE matches inside them either way, so dedup keeps working. Confirmed
on the live issue: pytorch#188953 is cited in #5272 only as a backticked URL
and has zero cross-references from this repo, while the nine bare ones each
have theirs.
The rule said upstream references go in backticks, and said nothing about
this repo's own, so the model generalised it: the fix session for
pytorch#195876 wrote the tracking issue it had just filed as
`#5476`. A code span is not linkified, so #5476 sat
there unreachable -- no cross-reference from #5272, no way back to the issue
that caused it short of reading the whole session comment. The 09-11 session
got this right with #5329, so it is drift the rule allows rather than asks
for.

Say the other half outright in both places the rule lives. The back-link is
exactly what a tracking issue needs; only someone else's tracker needs
protecting from it.
Two things read wrong now that discovery is dormant. The usage block still
advertised DRY_RUN as the opt-in to printing, when printing became the
default and `DRY_RUN=false` the only value that posts. And the dispatch
input offers "empty = discover", which today reaches a 404 and surfaced as
an HTTPError traceback -- the stack burying the one sentence the operator can
act on. Exit with that sentence instead, and leave anything other than a 404
to raise.
`fix-implement` files a tracking issue whenever it adds a skip, and asked for
`--label "agent-added,module: xpu"`. Neither label exists in this repo, so
`gh issue create` failed and each run picked its own: #5476 got
`module: inductor`, #5329 got `skipped,module: sdpa`, and neither is findable by
filtering for UT issues. The title prefix `[skip-added]` was invented here too
and matches nothing else in the tracker.

Both are fixed strings now, taken from what the 47 existing UT tracking issues
use (most recently #5305-#5309 and #5374-#5378): title
`[upstream_ut] <test>: <symptom>`, label `test: ut`.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Prompt-only change; the next `@torchxpubot fix` run that adds a skip exercises
it.
The template's body placeholder read "from <the torch-xpu-ops issue this run was
triggered on, bare so it back-links>", which is an instruction wearing a
placeholder's clothes, wrapped across two lines, using "bare" as jargon. It says
`intel/torch-xpu-ops#<N>` now, and the reason to leave the backticks off moved
into prose below the block, next to the pytorch/pytorch rule it is the mirror of.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
…not show

The template already shows the prefix, the label and the bare issue number;
twelve lines restating them plus the rationale were longer than the code they
explained. What is left is the part a reader cannot infer: a label this repo does
not have fails the call, and backticks kill the cross-reference.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
Five lines and three clauses to say "here the back-link is wanted". The rule
above it already explains what a rendered reference does; this one only has to
say that here we want it.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
A SOURCE batch can list ten DISABLED issues. Mirroring all of them in one
trigger gives the fix job ten unrelated problems to carry inside one 300min
run, one patch artifact holding several fixes, and a review that needs `item=`
markers to say which verdict is about what. So the queue now cuts a batch into
problems and mirrors the oldest one, leaving the rest to later rounds -- and
since `fix_in_flight` already holds the next trigger until the last one has
answered, the tracking issue reads as trigger, result, trigger, result.

Grouping is on the title, which every DISABLED issue shapes the same way:
`DISABLED test_foo_xpu_float32 (__main__.TestBarXPU)`. The key is the class plus
the test name with its parametrisation stripped. The class alone is not enough --
`test_redispatch_scatter_xpu_float8_*` and
`test_redispatch_nn_functional_grid_sample_xpu_*` are both
`TestTorchFunctionRedispatchOpsDeviceXPU` and are unrelated failures -- and the
test name alone would split one test's dtype variants across four runs.

This is batching, not a verdict on what shares a cause. /issue-handler settles
that by evidence (fix the first entry, re-run the rest against the staged fix,
list what passes in `covers`), so grouping on the title only has to put the
entries that mechanism pays off on into the same run. The title also costs
nothing: `pending` already fetches each issue to check it is open, and the title
comes back in the same response.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Self-test covers the real titles of both same-class families (grouped apart),
one test's dtype variants (grouped together), two unrelated tests in one class
(apart), an unparseable title (alone), and that a single-issue trigger does not
claim to be a parametrised family.
`re.sub(r"_(cpu|cuda|xpu)(_.*)?$", "", test)` matches leftmost, so a device token
in the middle of a name takes the rest of the name with it:
`test_copy_xpu_to_cuda` and `test_copy_cpu_to_xpu` both reduce to `test_copy` and
would have been mirrored as one problem. The parametrisation is always the tail,
so cut at the last device token instead. Same length, right boundary.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Self-test now pins that pair apart. Re-checked against the ten real backlog
titles: unchanged (each dtype family collapses, each unrelated test stands
alone).
A human who wants a specific test fixed now types `@torchxpubot fix` on the
tracking issue from their own account. Dispatching a workflow to have the bot
type it instead is more steps for the same comment, and bot.yml's gate takes
either author (OWNER/MEMBER covers the human; a PAT-authored comment was only
ever needed to prove the event fires). So the input, `explicit_body`, the digit
sanitising and its self-test go away, and `main` is just discovery.

This leaves the workflow dormant until the report issue moves into this repo --
a dispatch today prints the 404 explanation and posts nothing. That is what it
was already worth: the ISSUES path existed to drain the pre-START backlog, which
finished this morning (its last batch, 197521-197524, went out at 10:20).

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
Review pass over the script. The one that cost something real: `pending` called
`is_open_upstream(num)` and `one_problem` called `upstream_title(num)`, two
functions hitting the same `GET /issues/<num>` -- ten issues in a batch meant ten
wasted calls. One `@functools.cache`d `upstream()` serves both, and the self-test
needs one stub instead of two.

The other three are dead weight:

- `trigger_body` took a `key` argument, `next_batch` returned a third element and
  `one_problem` returned a tuple, all to print "All of these parametrise X::Y" in
  the comment. The agent reads the linked issues anyway, so the titles tell it
  that without being told.
- `paged` chose `&` or `?` for the query separator; none of its three callers
  passes a path with a query.
- The module docstring listed START, dedup and serialization, each of which is
  also explained by the function that implements it (and by the PR description).
  Three copies drift; the one at the implementation is the one that gets read.

`one_problem`'s docstring was also longer than its body, keeping only the part a
reader cannot infer: the reason not to turn this into a model call.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Behaviour is unchanged: same grouping, same dedup, same trigger body minus that
one sentence. Net -27 lines.
`fix_in_flight` read `runs?per_page=30` and skipped the completed ones. Measured
on the live workflow, 30 runs of bot.yml span about three hours (09:59 to 13:06
today), and a `fix` run takes one to five: a long one slides out of the window,
the gate reports idle, and the queue posts a trigger that starts a second GPU job
on top of the first. `?status=queued` and `?status=in_progress` return every
unfinished run whatever its age, which is the question this function was asking
all along. Two calls instead of one, no window.

The job-level check is unchanged and is still the point: bot.yml serves every
@torchxpubot command, and only `fix` is exclusive, so a 30-second `revert` run
must not hold the queue.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Both filters verified against the live workflow (`total_count` 0 for each with
nothing running). The self-test does not cover this function -- it is pure API
shape, nothing to assert without stubbing the transport.
With only a cron, a finished `fix` leaves the GPU idle until the next hour. A
`fix` takes 1-5h, so that slack is 20-100% of the wall clock, and the fix is
three lines: `workflow_run: {workflows: [bot], types: [completed]}` wakes the
queue as the run that just answered completes, so the tracking issue reads
trigger, result, trigger, result back to back with nothing waiting on a clock.

Not enabled here, for the same reason the cron is not: discovery still 404s, and
bot.yml completes dozens of times a day, so every one of them would become a
failed queue run. Written down where whoever restores the cron will read it.

It is not a replacement for `fix_in_flight`: a human can start a `fix` on any
issue at any time, and that run is not on the chain.

Test Plan:

```
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/ci_disabled_queue.yml'))"
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
Dedup asked whether a comment *contains* `@torchxpubot fix`. bot.yml fires on
`startsWith(comment.body, '@torchxpubot')` plus `/^@torchxpubot\s+(\S+)/i`, so
the two disagreed about the bot's own session comments, which quote the command
they are reporting on. Measured against #5272's history, the old rule read
196748's mirror time off the session comment (08:30:47) instead of its trigger
(06:05:50) -- a two-hour window in which a fresh upstream report of that issue
would have been dropped as already queued.

No task was lost yet, because that session comment happened to link only issues
that had been triggered. The mechanism for losing one was there: any pytorch
issue a session comment links becomes "already mirrored", and a later SOURCE
report of it is suppressed forever.

`FIX_CMD_RE` is start-anchored, whitespace-flexible and case-insensitive, which
is bot.yml's rule spelled in Python: if bot.yml would not have fired, this is not
a trigger.

The remaining ceiling is in the docstring: a comment that reads as a trigger but
never ran one (author rejected by the permission gate) still counts, because
asking "did a fix run for this issue" means matching comments to runs by
timestamp. The comment is the record.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Self-test pins both directions: the body this script posts matches, and a session
comment quoting the command does not. Also replayed both rules over all of
#5272's comments live -- 19 issues either way, one timestamp corrected.
Re-queueing a reopened issue worked, but by inference: an issue that is open now
and has a `closed` event after its last trigger must have been reopened after it
too. That derivation borrows an invariant from another function (`pending` only
ever triggers issues that are open at the time), and it misses a state change
recorded without a `closed` event at all.

`reopened_since` reads the `reopened` events instead, which is the thing the rule
means: the test is blocking CI again, for a reason the last fix did not settle.
Same call, same cost.

Verified on pytorch#194562 -- reported 08-24, `closed` 2026-09-03T03:29:44Z,
`reopened` 2026-09-16T01:00:20Z, and `closed_at` is null today, which is why
neither predicate can come from that field.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
A SOURCE comment's timestamp is when CI actually broke -- the comment carries the
failing commit and the xpu.yml run -- so walking the comments oldest-first put the
staleest breakage at the head of the queue. A fresh disable is a regression from
recent commits, where the cause is still in reach and the skip is hiding a live
bug; a month-old one has already waited and is waiting on something harder.

One word, `reversed`. Nothing starves at the current rates: a night's window fits
several batches and SOURCE adds a few a week, and if inflow ever passed
throughput, the old tail is what should go to humans anyway.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

The self-test covers `pending` and the grouping directly, not the walk order,
which has nothing to assert without stubbing a comment list.
One SOURCE comment is one failing xpu.yml run, so its issues come from the same
commit and are likelier to share a cause than anything a test name can tell you.
Splitting them by class-plus-test name was worse than not grouping at all.
Measured over every multi-issue batch since August, 7 of 10 would have been split,
and the splits cut through single bugs:

| batch | issues | split into | actually |
|---|---|---|---|
| 09-01 | 4 | 4 | `test_index_add_bfloat16_{deterministic,dim0,direct,scalar_index}` -- one bf16 index_add bug, four builds |
| 09-10 | 2 | 2 | `test_effn_attn_uniform_zero_bias` and its `_backward` |
| 08-20 | 2 | 2 | `special_i1` and `special_i1e`, adjacent kernels |
| 09-03 | 3 | 3 | three inductor regional tests from one commit |
| 09-16, 09-18, 09-22 | 4, 4, 2 | 1 each | the only batches grouping left alone |

The concern it was built for -- one run carrying several unrelated problems for
300 minutes -- is not what the data shows either: the largest batch observed is
four. And /issue-handler already handles a mixed batch better than the queue can,
by evidence instead of by name: `batch_kind=heterogeneous` fans out, and `One fix
may cover several sub-items` fixes the first entry, re-runs the rest against the
staged fix and lists what it covers. Grouping in the queue is what stopped that
from saving the builds.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

-82 lines: `problem_key`, `one_problem`, `TITLE_RE` and their self-tests.
Reverts the newest-batch-first walk. Dedup already makes a repeat impossible, so
jumping the queue buys nothing, and newest-first can starve the tail whenever
SOURCE reports faster than the queue drains -- which the commit that introduced
it waved away as "the old tail is what should go to humans anyway". Oldest-first
cannot starve anything: the batch that has waited longest goes next.

The observation that prompted it still holds and still earns its keep on START: a
SOURCE comment's timestamp is when CI actually broke, because the comment carries
the failing commit.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
…eps both halves

`dest_name` used the branch alone, and `issue-handler` names the branch after the
issue in EACH repo it touches. A fix spanning pytorch and torch-xpu-ops therefore
arrives as two units with the same branch name, lands in one directory, and
`format-patch` numbers both series from 0001.

Seen in the artifacts of today's grid_sample batch (pytorch#197521-197524), which
came one commit subject away from losing half a fix:

    fix-issue-5272-patch/fix-pytorch-issue-197521/
      0001-Enable-5D-bicubic-grid_sample-on-XPU-in-the-meta-ker.patch   (pytorch)
      0001-Implement-tricubic-grid_sampler_3d-on-XPU.patch              (torch-xpu-ops)

Equal subjects would have made the second overwrite the first, while `made`
counted two and the job stayed green -- and the filenames do not say which repo a
half applies to anyway. The directory is now `<branch>-<target_repo>`, which
answers both, and a dest that already exists is an error instead of a silent
merge: same branch and same repo means the agent broke its own branch naming, and
that must not be a green run.

Test Plan:

```
python3 -m pytest .github/scripts/test_bot_export_fix_patch.py
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

14 passed. Two new cases: one bug across both repos keeps both halves in separate
directories, and two units claiming one directory report an error rather than
overwriting. The four existing name assertions now spell the suffix.
#5490 ("XPU Periodic Run Auto Skip & Rerun") was opened on
2026-09-22, same title, same author and the same comment format as
chuanqi129/pytorch-xpu-ci#364, which is where this workflow used to have to read
from. MERGE_TOKEN is a fine-grained PAT scoped to this org, so reading an issue in
this repo needs nothing special -- the 404 that made discovery dormant is gone.

- `SOURCE_REPO` / `SOURCE_ISSUE` -> `intel/torch-xpu-ops`, 5490. SOURCE and
  TARGET are now two issues in one repo, two roles, so the constant says which.
- `START` stays 2026-09-22, which is now exactly the day SOURCE moved here: the
  queue owns what SOURCE has reported since, and the older hand-triaged backlog
  stays with the humans.
- The 404 path no longer explains fine-grained PATs. A 404 on our own issue means
  it was deleted or renumbered, so it says that instead.
- The `token-access` job is deleted. It existed to probe the one thing that could
  not be checked locally and to turn green when the read started working; the read
  moved instead, so there is nothing left for it to watch.

Still no `cron`, for the one reason that is left: a `fix` run that never reaches a
working GPU still counts as queued, so an unattended schedule would silently spend
entries against a dead card. The comment in the workflow says so, and turning the
cron on is one line once that is handled.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
GH_TOKEN=<read+write here> python3 .github/scripts/ci_disabled_queue.py
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

The second command is discovery running against the live issues for the first
time, since until today the read 404'd:

```
nothing to queue since 2026-09-22; 19 mirrored so far
```

#5490's only comment so far is a "rerunning" notice with no `Disable issues:`
list, which the parser yields nothing for -- the self-test pins exactly that case.
Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
Hourly over 17:00-08:00 Beijing, the window where the xpu-agent runners are
otherwise idle. Everything it needed arrived with #5490: SOURCE is readable, so a
tick either finds a `Disable issues:` batch and posts one trigger, or prints
`nothing to queue` and exits.

The dead-GPU gap stays open by decision: a `fix` run that never reaches a working
card still counts as queued, so the entry is silently spent. Driver deaths are
worked around by hand for now -- watch #5272 for a trigger whose run died early and
re-post it -- and the preflight-and-requeue gate comes later. The workflow comment
says exactly that, so whoever meets it first does not have to rediscover it.

The `workflow_run` chain stays commented, with its reason rewritten: it would cut
the idle gap between a finished fix and the next trigger, but the cron already
bounds that at an hour against a 1-5h run, and it would fire on every bot.yml
completion, dozens a day, mostly with nothing to queue.

`DRY_RUN` needs no change: the step reads
`github.event_name == 'schedule' && 'false' || inputs.dry_run`, so a tick posts
while a dispatch still defaults to printing.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
python3 -c "import yaml; print(yaml.safe_load(open('.github/workflows/ci_disabled_queue.yml'))[True]['schedule'])"
```

The last one prints `[{'cron': '0 0,9-23 * * *'}]`; a malformed `on:` block would
otherwise only show up as a workflow that never runs.
The comment said "driver deaths". It is #5432: a Level Zero DEVICE_LOST /
OUT_OF_RESOURCES / OUT_OF_DEVICE_MEMORY state, known driver bugs (GSD-13281,
COMPLRLLVM-77770) with no fix due before 2026.3, so this gap stays open for a
while rather than a week.

Two details from that issue change what a future gate has to do, and are worth
carrying next to the cron they qualify:

- `torch.xpu.is_available()` returns True in that state, and `xpu-smi discovery`
  reports every device normal, while `test_copy.py` fails 9/9 on plain copies. So
  bot.yml's existing preflight assert passes and the run proceeds -- a preflight
  has to run a real kernel, not ask whether a device enumerates.
- runs can hang rather than fail, so the requeue side has to treat a 300min
  timeout as "never got a GPU" too, not just an early exit.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Comment-only.
`fix-root-cause` Step 4 sent any fix needing both pytorch and torch-xpu-ops to
NEEDS_HUMAN, with the reason "agent supports only single-repo fixes in this run".
That was true of the pipeline once; today's grid_sample run disproved it by doing
both halves and verifying them together (pytorch#197521-197524: the XPU SYCL kernel
in torch-xpu-ops, and the pytorch-side meta registration plus the OpInfo entries
that stop xfailing it). The pytorch clone carries torch-xpu-ops as its submodule,
so both halves are in one tree and one build covers them.

Because the case was unspecified, that run had to invent how to record it -- a
`fix_result-<slug>-pytorch.json` file name, a top-level `patches` array nothing
reads, and the landing order in a free-text `notes`. The landing order is the part
that matters: the torch-xpu-ops kernel has to merge before the pytorch side, which
waits on the `xpu.txt` pin bump. Whoever opens the PRs cannot get it from the
artifact, where the two patch series look independent.

So this specifies it instead:

- Step 4 returns `target_repo` (the half that lands first) plus a new
  `companion_repo`, and `cross_repo_coordinated` is redefined to what it should
  always have meant -- a repo the agent cannot build, oneDNN or IGC or a driver
  component. Both example uses of the code are annotated so it is not read as
  "pytorch + torch-xpu-ops" again.
- Stage 4 calls `fix-implement` once per repo, `target_repo` first. The skill still
  stages one repo's diff per invocation; a two-repo fix is two invocations.
- Stage 5 verifies the pair in one build and gives both records that verdict, with
  a note saying so, so a lone `PASSED` is not read as "this half stands alone".
- The fix_result schema gains `depends_on`, the companion file is the primary's
  name plus `-<target_repo>`, and `batch_summary`'s sub-item keeps pointing at the
  unit that fixes the test while gaining a `companion` key. No invented top-level
  keys.
- The Summary block gains the ordered instructions for whoever opens the PR, as a
  `> [!IMPORTANT]` list, plus a one-line version OUTSIDE `</details>` -- that block
  renders collapsed, and a notice nobody expands is a notice nobody reads.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```

Prompt-and-contract only; the next `@torchxpubot fix` that spans both repos
exercises it, and `bot_export_fix_patch.py` already keeps the two halves apart
since the `<branch>-<target_repo>` change earlier in this PR.
"both repos are in one checkout, so one build verifies the pair" was written out
in seven places: Step 4, the `cross_repo_coordinated` code definition, the
`unresolvable_statically` note, `fix-implement`'s single-repo paragraph, the
two-unit recording block, Stage 4 and Stage 5. These files are prompts; a rule
repeated with fresh justification each time reads as seven rules.

Now each place says only what it owns. Step 4 decides which repos. The reason code
says what it is for. The outputs section says how the two units are recorded.
Stage 4 says to call the implement step twice, Stage 5 that one build covers both.
The Summary template gives the ordered instructions without arguing for them.

Net -32 lines against the same behaviour.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
Both scripts had grown prose that explained how a rule was arrived at rather than
what it is. The module docstring was 33 lines and two of its claims had gone stale
-- it still listed "grouping" as a rule (removed) and still said a human posts the
triggers by hand (the cron does).

Kept in every case: the fact that stops someone undoing the fix. `fix_in_flight`
still says why it asks by status, `last_mirrored` still says why the match is
start-anchored, `reopened_since` still says `closed_at` is cleared by a reopen,
`dest_name` still says why the repo is in the directory name. Dropped: the
incident each one came from.

-61/+39 lines of prose, no behaviour change.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
python3 -m pytest .github/scripts/test_bot_export_fix_patch.py
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
"nobody has asked about yet" was the dedup rule smuggled into the summary line,
where it also left a dangling question -- asked whom. Dedup has its own paragraph
below. The line now names the mechanism: it posts a comment carrying the command.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
`cross_repo_coordinated` named oneDNN, IGC and the driver. Triton is missing, and
so is whatever comes next, so the code now points at the vocabulary that already
exists for this: the repo's `dependency component: *` labels (oneDNN, Triton, IGC,
Level_Zero, oneAPI, driver, third_party, community...). Both examples say
oneDNN/Triton/IGC/driver rather than picking one.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
`last_mirrored`'s docstring re-explained the start-anchored match, which the
`FIX_CMD_RE` comment already does at the definition, with bot.yml's own source
next to it. And `main`'s two-line note on `DRY_RUN` restated the usage block at
the top of the file; what a reader needs at that line is why the comparison is
against a literal, which is one line.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
18 lines to 12, and it now says one thing more.

Deleted: the paragraph restating what `# print` / `# post` already say in the usage
block, and the two lines defending the absence of per-test grouping -- a
non-feature does not need a defence in the file, and "batches kept whole" carries
it in three words. The list of which rule lives in which function went too; the
useful half of that sentence is that there is no local state.

Added, because it was missing and is headline behaviour: never while a `fix` job is
still running.

Test Plan:

```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
@ZhaoqiongZ
ZhaoqiongZ force-pushed the bot/ci-disabled-queue branch from 4b5f376 to 3a29d1b Compare September 24, 2026 01:38
@ZhaoqiongZ
ZhaoqiongZ marked this pull request as ready for review September 24, 2026 03:17
@ZhaoqiongZ ZhaoqiongZ changed the title [bot] Queue upstream DISABLED tests for the fix bot on a schedule [bot] Automate bot fix on CI DISABLED issues Sep 24, 2026
@github-actions

Copy link
Copy Markdown

Performance outliers, please check!

  • 🟡 [80%, 90%), may be fluctuations
Category Model Target vs. Baseline [Eager] Target vs. Baseline [Inductor]
timm_models_bfloat16_training tf_efficientnet_b0 0.893254 0.836332

Comment thread .claude/skills/fix-implement/SKILL.md
Comment thread .claude/skills/issue-handler/SKILL.md Outdated
Comment thread .claude/skills/issue-handler/SKILL.md
`fix-root-cause` had eleven NEEDS_HUMAN reason codes. Nothing reads them: no
script under `.github/` parses `reason`, and every one of them lands on the same
terminal state -- `NEEDS_HUMAN` status, `agent:needs-human` label
(execution-modes.md's stage table is the whole mapping). The claim in
issue-handler that "each reason maps to a different final `agent:status` value"
was simply false, and it was the only justification the list had.

So the enum keeps the three codes an orchestrator actually branches on --
`no_registered_domain` (the retry policy refuses to loop on it),
`invalid_reproduction` (re-run reproduce first), `security_concern` (stop now) --
and everything else is `other` with a mandatory one-line justification in
`reason_detail`. Step 5's signal-to-code table becomes a list of blockers seen so
far, marked explicitly as examples and not an enum, so the next unfamiliar
failure gets described instead of filed under the nearest label.

This deletes `cross_repo_coordinated` along with the rest, which undoes 2377f36
and 0aafbdb's work on it. That code was the clearest case for the change: after
being redefined away from pytorch + torch-xpu-ops it needed a parenthetical at
every use site to stop readers drawing the obvious wrong conclusion, and a name
that has to be corrected wherever it appears costs more than it carries. The
`dependency component: *` label vocabulary those commits found is kept -- it now
lives in the justification, which is where a maintainer reads it.

Also fixed the matching overclaim on `fix-implement`'s codes in Stage 4.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV -m origin/main
python -m pytest .github/scripts/test_bot_export_fix_patch.py -q
grep -rn "cross_repo_coordinated\|unresolvable_statically\|hardware_specific" --include=*.md --include=*.py --include=*.yml .
```

The grep returns nothing: no dangling references to the deleted codes.
Branch naming claimed one branch per bug so a reader "knows which issue a PR
closes", and stopped there. With a `companion_repo` that name now exists in two
checkouts and answers to two PRs, and the thing that tells them apart -- the
`-<target_repo>` suffix on the patch directory -- lived only in
`bot_export_fix_patch.py`. A reader looking the rule up would conclude one fix is
one branch is one PR, and the obvious repair is to suffix the branch itself, which
buys `fix-pytorch-issue-197521-pytorch-pytorch`. So the section now says the name
is shared on purpose, where the repo does appear, and that a branch plus
target_repo collision fails the export rather than overwriting half the patches.

And the two-PR callout gains the rule that neither PR may close the issue: a
closing keyword on the half that lands first marks the bug fixed while CI is still
red on the other half. A human closes it once both are in.

Test Plan:

```
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV -m origin/main
python -m pytest .github/scripts/test_bot_export_fix_patch.py -q
```
@github-actions

Copy link
Copy Markdown

Performance outliers, please check!

  • 🟡 [80%, 90%), may be fluctuations
Category Model Target vs. Baseline [Eager] Target vs. Baseline [Inductor]
timm_models_bfloat16_training tf_efficientnet_b0 0.865542 0.853292

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants