[bot] Automate bot fix on CI DISABLED issues - #5473
Open
ZhaoqiongZ wants to merge 36 commits into
Open
ZhaoqiongZ wants to merge 36 commits into
ZhaoqiongZ wants to merge 36 commits into
Conversation
ZhaoqiongZ
marked this pull request as draft
September 22, 2026 02:43
ZhaoqiongZ
force-pushed
the
bot/ci-disabled-queue
branch
from
September 22, 2026 02:45
82c3e33 to
6d9d8d7
Compare
Every `@torchxpubot fix` on #5272 is hand-copied from a batch comment on chuanqi129/pytorch-xpu-ci#364. The copying is mechanical, but two things about it are not, which is why it stayed manual: #364 re-lists the same upstream issue on every later failing commit (pytorch#197144 appears eight times), and a `fix` run costs a GPU and a couple of hours, so two triggers must not overlap. Both are decidable from GitHub state, so the queue keeps no database. An issue is skipped if it is closed upstream, or if it has stayed open since the last `@torchxpubot fix` comment naming it. Dedup is per open episode rather than forever: pytorch#194562 was mirrored on 08-24, closed on 09-03, reopened on 09-16 and re-reported the same day -- that is a fresh breakage and gets queued again. `closed_at` cannot see this, since reopening clears it, so the issue's `closed` events are read instead. Serialization is a check for a `fix` job that is queued or running; the job's own 300min timeout bounds the wait, so no staleness escape hatch is needed. bot.yml's `concurrency: bot-<issue>` cannot do the serializing: one group holds a single pending run, so posting several comments at once would get all but the last cancelled. START is where the queue takes over from the humans, and it has to exist because "a human already has a PR for this" is not recorded anywhere GitHub can be asked -- these issues carry no assignee and no linked pull request, their cross-reference events are both incomplete (pytorch#195558 has no xref to its fix pytorch#195534) and noisy (an unrelated ROCm PR appears on two of them), and the ownership table that does have it is a spreadsheet. So the queue owns nothing before START: that backlog was triaged by hand, and the leftovers with no PR and no owner are being triggered by hand too. One batch per run, oldest first, so a backlog drains a batch per completed fix instead of all at once. The cron covers 17:00-08:00 Beijing, when the xpu-agent runners are idle. Test Plan ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` The self-test covers the parser, the trigger body, and per-episode dedup. It runs on every PR touching these two files. Dry-run against live #364 and #5272 reports an empty queue, which is correct for START=2026-09-22: #364's last batch comment is 09-18. Moving START back to 09-18 in a scratch copy selects that batch and nothing else: ``` GH_TOKEN=... python3 .github/scripts/ci_disabled_queue.py --dry-run -> nothing to queue since 2026-09-22; 9 issues mirrored so far START=2026-09-18 -> 09-18 [197521 197522 197523 197524] ``` A steady-state run costs 6 API calls. Not verified here, both need a run on the default branch: that MERGE_TOKEN can read the private chuanqi129/pytorch-xpu-ci, and that a comment posted by torchxpubot passes `fix`'s OWNER/MEMBER gate. torchxpubot is an intel org member, so the webhook's `author_association` should be MEMBER, but the REST API reports CONTRIBUTOR for the same user and the two disagree. Authored with an AI assistant.
The queue's one unverifiable-from-a-branch dependency: it reads issue comments from a private repo owned by a personal account, so whether MERGE_TOKEN can reach it decides whether the workflow works at all, and neither the token's kind nor its grants are visible from the outside. Probe it on pull_request the way bot.yml's token-permission-test probes its own token needs. The failure message carries the likely diagnosis, since a 404 with torchxpubot already a collaborator points at the token being fine-grained rather than at a missing grant.
…s reachable MERGE_TOKEN turns out to be a fine-grained PAT, so discovery cannot work yet: fine-grained tokens are scoped to one resource owner, and SOURCE lives in a repo owned by another personal account. Adding torchxpubot as a collaborator there does not help and did not -- the token-access job gets a 404 with the grant in place. A cron would 404 hourly, so it is held back rather than shipped broken, and the probe now warns instead of failing: it turns green by itself when the report issue moves into this repo, which is the signal to restore it. Nothing but the SOURCE_* constants changes then. Meanwhile ISSUES queues an operator-supplied list with no SOURCE read at all, which is what drains the triaged backlog and what proves a comment authored by torchxpubot clears the fix command's OWNER/MEMBER gate. Both dispatch inputs travel through the environment rather than the run string, since an interpolated workflow_dispatch input is an injection vector, and only the digits of ISSUES are kept. The self-test pins that: a hostile value reduces to its issue numbers and nothing else.
An unset or misspelled DRY_RUN used to fall through to posting, which is the wrong way round for a switch whose live side comments on a tracking issue and starts a multi-hour GPU job. Only the exact string "false" posts now, so a restored cron has to say so outright rather than inherit an empty workflow_dispatch input.
A bare issue URL makes GitHub file a cross-referenced event on the target, so every trigger leaves an `#5272` backlink on the upstream issue. All nine posted by hand did it, and at one trigger per batch forever that is a steady drip of noise onto someone else's tracker -- the same noise that made cross-references useless as a has-a-PR signal earlier in this work, where an unrelated ROCm PR showed up on two XPU issues. Wrap the URLs in backticks. Code spans are not scanned for references, and ISSUE_RE matches inside them either way, so dedup keeps working. Confirmed on the live issue: pytorch#188953 is cited in #5272 only as a backticked URL and has zero cross-references from this repo, while the nine bare ones each have theirs.
The rule said upstream references go in backticks, and said nothing about this repo's own, so the model generalised it: the fix session for pytorch#195876 wrote the tracking issue it had just filed as `#5476`. A code span is not linkified, so #5476 sat there unreachable -- no cross-reference from #5272, no way back to the issue that caused it short of reading the whole session comment. The 09-11 session got this right with #5329, so it is drift the rule allows rather than asks for. Say the other half outright in both places the rule lives. The back-link is exactly what a tracking issue needs; only someone else's tracker needs protecting from it.
Two things read wrong now that discovery is dormant. The usage block still advertised DRY_RUN as the opt-in to printing, when printing became the default and `DRY_RUN=false` the only value that posts. And the dispatch input offers "empty = discover", which today reaches a 404 and surfaced as an HTTPError traceback -- the stack burying the one sentence the operator can act on. Exit with that sentence instead, and leave anything other than a 404 to raise.
`fix-implement` files a tracking issue whenever it adds a skip, and asked for `--label "agent-added,module: xpu"`. Neither label exists in this repo, so `gh issue create` failed and each run picked its own: #5476 got `module: inductor`, #5329 got `skipped,module: sdpa`, and neither is findable by filtering for UT issues. The title prefix `[skip-added]` was invented here too and matches nothing else in the tracker. Both are fixed strings now, taken from what the 47 existing UT tracking issues use (most recently #5305-#5309 and #5374-#5378): title `[upstream_ut] <test>: <symptom>`, label `test: ut`. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Prompt-only change; the next `@torchxpubot fix` run that adds a skip exercises it.
The template's body placeholder read "from <the torch-xpu-ops issue this run was triggered on, bare so it back-links>", which is an instruction wearing a placeholder's clothes, wrapped across two lines, using "bare" as jargon. It says `intel/torch-xpu-ops#<N>` now, and the reason to leave the backticks off moved into prose below the block, next to the pytorch/pytorch rule it is the mirror of. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
…not show The template already shows the prefix, the label and the bare issue number; twelve lines restating them plus the rationale were longer than the code they explained. What is left is the part a reader cannot infer: a label this repo does not have fails the call, and backticks kill the cross-reference. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
Five lines and three clauses to say "here the back-link is wanted". The rule above it already explains what a rendered reference does; this one only has to say that here we want it. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
A SOURCE batch can list ten DISABLED issues. Mirroring all of them in one trigger gives the fix job ten unrelated problems to carry inside one 300min run, one patch artifact holding several fixes, and a review that needs `item=` markers to say which verdict is about what. So the queue now cuts a batch into problems and mirrors the oldest one, leaving the rest to later rounds -- and since `fix_in_flight` already holds the next trigger until the last one has answered, the tracking issue reads as trigger, result, trigger, result. Grouping is on the title, which every DISABLED issue shapes the same way: `DISABLED test_foo_xpu_float32 (__main__.TestBarXPU)`. The key is the class plus the test name with its parametrisation stripped. The class alone is not enough -- `test_redispatch_scatter_xpu_float8_*` and `test_redispatch_nn_functional_grid_sample_xpu_*` are both `TestTorchFunctionRedispatchOpsDeviceXPU` and are unrelated failures -- and the test name alone would split one test's dtype variants across four runs. This is batching, not a verdict on what shares a cause. /issue-handler settles that by evidence (fix the first entry, re-run the rest against the staged fix, list what passes in `covers`), so grouping on the title only has to put the entries that mechanism pays off on into the same run. The title also costs nothing: `pending` already fetches each issue to check it is open, and the title comes back in the same response. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Self-test covers the real titles of both same-class families (grouped apart), one test's dtype variants (grouped together), two unrelated tests in one class (apart), an unparseable title (alone), and that a single-issue trigger does not claim to be a parametrised family.
`re.sub(r"_(cpu|cuda|xpu)(_.*)?$", "", test)` matches leftmost, so a device token in the middle of a name takes the rest of the name with it: `test_copy_xpu_to_cuda` and `test_copy_cpu_to_xpu` both reduce to `test_copy` and would have been mirrored as one problem. The parametrisation is always the tail, so cut at the last device token instead. Same length, right boundary. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Self-test now pins that pair apart. Re-checked against the ten real backlog titles: unchanged (each dtype family collapses, each unrelated test stands alone).
A human who wants a specific test fixed now types `@torchxpubot fix` on the tracking issue from their own account. Dispatching a workflow to have the bot type it instead is more steps for the same comment, and bot.yml's gate takes either author (OWNER/MEMBER covers the human; a PAT-authored comment was only ever needed to prove the event fires). So the input, `explicit_body`, the digit sanitising and its self-test go away, and `main` is just discovery. This leaves the workflow dormant until the report issue moves into this repo -- a dispatch today prints the 404 explanation and posts nothing. That is what it was already worth: the ISSUES path existed to drain the pre-START backlog, which finished this morning (its last batch, 197521-197524, went out at 10:20). Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
Review pass over the script. The one that cost something real: `pending` called `is_open_upstream(num)` and `one_problem` called `upstream_title(num)`, two functions hitting the same `GET /issues/<num>` -- ten issues in a batch meant ten wasted calls. One `@functools.cache`d `upstream()` serves both, and the self-test needs one stub instead of two. The other three are dead weight: - `trigger_body` took a `key` argument, `next_batch` returned a third element and `one_problem` returned a tuple, all to print "All of these parametrise X::Y" in the comment. The agent reads the linked issues anyway, so the titles tell it that without being told. - `paged` chose `&` or `?` for the query separator; none of its three callers passes a path with a query. - The module docstring listed START, dedup and serialization, each of which is also explained by the function that implements it (and by the PR description). Three copies drift; the one at the implementation is the one that gets read. `one_problem`'s docstring was also longer than its body, keeping only the part a reader cannot infer: the reason not to turn this into a model call. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Behaviour is unchanged: same grouping, same dedup, same trigger body minus that one sentence. Net -27 lines.
`fix_in_flight` read `runs?per_page=30` and skipped the completed ones. Measured on the live workflow, 30 runs of bot.yml span about three hours (09:59 to 13:06 today), and a `fix` run takes one to five: a long one slides out of the window, the gate reports idle, and the queue posts a trigger that starts a second GPU job on top of the first. `?status=queued` and `?status=in_progress` return every unfinished run whatever its age, which is the question this function was asking all along. Two calls instead of one, no window. The job-level check is unchanged and is still the point: bot.yml serves every @torchxpubot command, and only `fix` is exclusive, so a 30-second `revert` run must not hold the queue. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Both filters verified against the live workflow (`total_count` 0 for each with nothing running). The self-test does not cover this function -- it is pure API shape, nothing to assert without stubbing the transport.
With only a cron, a finished `fix` leaves the GPU idle until the next hour. A
`fix` takes 1-5h, so that slack is 20-100% of the wall clock, and the fix is
three lines: `workflow_run: {workflows: [bot], types: [completed]}` wakes the
queue as the run that just answered completes, so the tracking issue reads
trigger, result, trigger, result back to back with nothing waiting on a clock.
Not enabled here, for the same reason the cron is not: discovery still 404s, and
bot.yml completes dozens of times a day, so every one of them would become a
failed queue run. Written down where whoever restores the cron will read it.
It is not a replacement for `fix_in_flight`: a human can start a `fix` on any
issue at any time, and that run is not on the chain.
Test Plan:
```
python3 -c "import yaml; yaml.safe_load(open('.github/workflows/ci_disabled_queue.yml'))"
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
Dedup asked whether a comment *contains* `@torchxpubot fix`. bot.yml fires on `startsWith(comment.body, '@torchxpubot')` plus `/^@torchxpubot\s+(\S+)/i`, so the two disagreed about the bot's own session comments, which quote the command they are reporting on. Measured against #5272's history, the old rule read 196748's mirror time off the session comment (08:30:47) instead of its trigger (06:05:50) -- a two-hour window in which a fresh upstream report of that issue would have been dropped as already queued. No task was lost yet, because that session comment happened to link only issues that had been triggered. The mechanism for losing one was there: any pytorch issue a session comment links becomes "already mirrored", and a later SOURCE report of it is suppressed forever. `FIX_CMD_RE` is start-anchored, whitespace-flexible and case-insensitive, which is bot.yml's rule spelled in Python: if bot.yml would not have fired, this is not a trigger. The remaining ceiling is in the docstring: a comment that reads as a trigger but never ran one (author rejected by the permission gate) still counts, because asking "did a fix run for this issue" means matching comments to runs by timestamp. The comment is the record. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Self-test pins both directions: the body this script posts matches, and a session comment quoting the command does not. Also replayed both rules over all of #5272's comments live -- 19 issues either way, one timestamp corrected.
Re-queueing a reopened issue worked, but by inference: an issue that is open now and has a `closed` event after its last trigger must have been reopened after it too. That derivation borrows an invariant from another function (`pending` only ever triggers issues that are open at the time), and it misses a state change recorded without a `closed` event at all. `reopened_since` reads the `reopened` events instead, which is the thing the rule means: the test is blocking CI again, for a reason the last fix did not settle. Same call, same cost. Verified on pytorch#194562 -- reported 08-24, `closed` 2026-09-03T03:29:44Z, `reopened` 2026-09-16T01:00:20Z, and `closed_at` is null today, which is why neither predicate can come from that field. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
A SOURCE comment's timestamp is when CI actually broke -- the comment carries the failing commit and the xpu.yml run -- so walking the comments oldest-first put the staleest breakage at the head of the queue. A fresh disable is a regression from recent commits, where the cause is still in reach and the skip is hiding a live bug; a month-old one has already waited and is waiting on something harder. One word, `reversed`. Nothing starves at the current rates: a night's window fits several batches and SOURCE adds a few a week, and if inflow ever passed throughput, the old tail is what should go to humans anyway. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` The self-test covers `pending` and the grouping directly, not the walk order, which has nothing to assert without stubbing a comment list.
One SOURCE comment is one failing xpu.yml run, so its issues come from the same
commit and are likelier to share a cause than anything a test name can tell you.
Splitting them by class-plus-test name was worse than not grouping at all.
Measured over every multi-issue batch since August, 7 of 10 would have been split,
and the splits cut through single bugs:
| batch | issues | split into | actually |
|---|---|---|---|
| 09-01 | 4 | 4 | `test_index_add_bfloat16_{deterministic,dim0,direct,scalar_index}` -- one bf16 index_add bug, four builds |
| 09-10 | 2 | 2 | `test_effn_attn_uniform_zero_bias` and its `_backward` |
| 08-20 | 2 | 2 | `special_i1` and `special_i1e`, adjacent kernels |
| 09-03 | 3 | 3 | three inductor regional tests from one commit |
| 09-16, 09-18, 09-22 | 4, 4, 2 | 1 each | the only batches grouping left alone |
The concern it was built for -- one run carrying several unrelated problems for
300 minutes -- is not what the data shows either: the largest batch observed is
four. And /issue-handler already handles a mixed batch better than the queue can,
by evidence instead of by name: `batch_kind=heterogeneous` fans out, and `One fix
may cover several sub-items` fixes the first entry, re-runs the rest against the
staged fix and lists what it covers. Grouping in the queue is what stopped that
from saving the builds.
Test Plan:
```
python3 .github/scripts/ci_disabled_queue.py --self-test
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
-82 lines: `problem_key`, `one_problem`, `TITLE_RE` and their self-tests.
Reverts the newest-batch-first walk. Dedup already makes a repeat impossible, so jumping the queue buys nothing, and newest-first can starve the tail whenever SOURCE reports faster than the queue drains -- which the commit that introduced it waved away as "the old tail is what should go to humans anyway". Oldest-first cannot starve anything: the batch that has waited longest goes next. The observation that prompted it still holds and still earns its keep on START: a SOURCE comment's timestamp is when CI actually broke, because the comment carries the failing commit. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
…eps both halves
`dest_name` used the branch alone, and `issue-handler` names the branch after the
issue in EACH repo it touches. A fix spanning pytorch and torch-xpu-ops therefore
arrives as two units with the same branch name, lands in one directory, and
`format-patch` numbers both series from 0001.
Seen in the artifacts of today's grid_sample batch (pytorch#197521-197524), which
came one commit subject away from losing half a fix:
fix-issue-5272-patch/fix-pytorch-issue-197521/
0001-Enable-5D-bicubic-grid_sample-on-XPU-in-the-meta-ker.patch (pytorch)
0001-Implement-tricubic-grid_sampler_3d-on-XPU.patch (torch-xpu-ops)
Equal subjects would have made the second overwrite the first, while `made`
counted two and the job stayed green -- and the filenames do not say which repo a
half applies to anyway. The directory is now `<branch>-<target_repo>`, which
answers both, and a dest that already exists is an error instead of a silent
merge: same branch and same repo means the agent broke its own branch naming, and
that must not be a green run.
Test Plan:
```
python3 -m pytest .github/scripts/test_bot_export_fix_patch.py
lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV
```
14 passed. Two new cases: one bug across both repos keeps both halves in separate
directories, and two units claiming one directory report an error rather than
overwriting. The four existing name assertions now spell the suffix.
#5490 ("XPU Periodic Run Auto Skip & Rerun") was opened on 2026-09-22, same title, same author and the same comment format as chuanqi129/pytorch-xpu-ci#364, which is where this workflow used to have to read from. MERGE_TOKEN is a fine-grained PAT scoped to this org, so reading an issue in this repo needs nothing special -- the 404 that made discovery dormant is gone. - `SOURCE_REPO` / `SOURCE_ISSUE` -> `intel/torch-xpu-ops`, 5490. SOURCE and TARGET are now two issues in one repo, two roles, so the constant says which. - `START` stays 2026-09-22, which is now exactly the day SOURCE moved here: the queue owns what SOURCE has reported since, and the older hand-triaged backlog stays with the humans. - The 404 path no longer explains fine-grained PATs. A 404 on our own issue means it was deleted or renumbered, so it says that instead. - The `token-access` job is deleted. It existed to probe the one thing that could not be checked locally and to turn green when the read started working; the read moved instead, so there is nothing left for it to watch. Still no `cron`, for the one reason that is left: a `fix` run that never reaches a working GPU still counts as queued, so an unattended schedule would silently spend entries against a dead card. The comment in the workflow says so, and turning the cron on is one line once that is handled. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test GH_TOKEN=<read+write here> python3 .github/scripts/ci_disabled_queue.py lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` The second command is discovery running against the live issues for the first time, since until today the read 404'd: ``` nothing to queue since 2026-09-22; 19 mirrored so far ``` #5490's only comment so far is a "rerunning" notice with no `Disable issues:` list, which the parser yields nothing for -- the self-test pins exactly that case.
Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
Hourly over 17:00-08:00 Beijing, the window where the xpu-agent runners are otherwise idle. Everything it needed arrived with #5490: SOURCE is readable, so a tick either finds a `Disable issues:` batch and posts one trigger, or prints `nothing to queue` and exits. The dead-GPU gap stays open by decision: a `fix` run that never reaches a working card still counts as queued, so the entry is silently spent. Driver deaths are worked around by hand for now -- watch #5272 for a trigger whose run died early and re-post it -- and the preflight-and-requeue gate comes later. The workflow comment says exactly that, so whoever meets it first does not have to rediscover it. The `workflow_run` chain stays commented, with its reason rewritten: it would cut the idle gap between a finished fix and the next trigger, but the cron already bounds that at an hour against a 1-5h run, and it would fire on every bot.yml completion, dozens a day, mostly with nothing to queue. `DRY_RUN` needs no change: the step reads `github.event_name == 'schedule' && 'false' || inputs.dry_run`, so a tick posts while a dispatch still defaults to printing. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV python3 -c "import yaml; print(yaml.safe_load(open('.github/workflows/ci_disabled_queue.yml'))[True]['schedule'])" ``` The last one prints `[{'cron': '0 0,9-23 * * *'}]`; a malformed `on:` block would otherwise only show up as a workflow that never runs.
The comment said "driver deaths". It is #5432: a Level Zero DEVICE_LOST / OUT_OF_RESOURCES / OUT_OF_DEVICE_MEMORY state, known driver bugs (GSD-13281, COMPLRLLVM-77770) with no fix due before 2026.3, so this gap stays open for a while rather than a week. Two details from that issue change what a future gate has to do, and are worth carrying next to the cron they qualify: - `torch.xpu.is_available()` returns True in that state, and `xpu-smi discovery` reports every device normal, while `test_copy.py` fails 9/9 on plain copies. So bot.yml's existing preflight assert passes and the run proceeds -- a preflight has to run a real kernel, not ask whether a device enumerates. - runs can hang rather than fail, so the requeue side has to treat a 300min timeout as "never got a GPU" too, not just an early exit. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Comment-only.
`fix-root-cause` Step 4 sent any fix needing both pytorch and torch-xpu-ops to NEEDS_HUMAN, with the reason "agent supports only single-repo fixes in this run". That was true of the pipeline once; today's grid_sample run disproved it by doing both halves and verifying them together (pytorch#197521-197524: the XPU SYCL kernel in torch-xpu-ops, and the pytorch-side meta registration plus the OpInfo entries that stop xfailing it). The pytorch clone carries torch-xpu-ops as its submodule, so both halves are in one tree and one build covers them. Because the case was unspecified, that run had to invent how to record it -- a `fix_result-<slug>-pytorch.json` file name, a top-level `patches` array nothing reads, and the landing order in a free-text `notes`. The landing order is the part that matters: the torch-xpu-ops kernel has to merge before the pytorch side, which waits on the `xpu.txt` pin bump. Whoever opens the PRs cannot get it from the artifact, where the two patch series look independent. So this specifies it instead: - Step 4 returns `target_repo` (the half that lands first) plus a new `companion_repo`, and `cross_repo_coordinated` is redefined to what it should always have meant -- a repo the agent cannot build, oneDNN or IGC or a driver component. Both example uses of the code are annotated so it is not read as "pytorch + torch-xpu-ops" again. - Stage 4 calls `fix-implement` once per repo, `target_repo` first. The skill still stages one repo's diff per invocation; a two-repo fix is two invocations. - Stage 5 verifies the pair in one build and gives both records that verdict, with a note saying so, so a lone `PASSED` is not read as "this half stands alone". - The fix_result schema gains `depends_on`, the companion file is the primary's name plus `-<target_repo>`, and `batch_summary`'s sub-item keeps pointing at the unit that fixes the test while gaining a `companion` key. No invented top-level keys. - The Summary block gains the ordered instructions for whoever opens the PR, as a `> [!IMPORTANT]` list, plus a one-line version OUTSIDE `</details>` -- that block renders collapsed, and a notice nobody expands is a notice nobody reads. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ``` Prompt-and-contract only; the next `@torchxpubot fix` that spans both repos exercises it, and `bot_export_fix_patch.py` already keeps the two halves apart since the `<branch>-<target_repo>` change earlier in this PR.
"both repos are in one checkout, so one build verifies the pair" was written out in seven places: Step 4, the `cross_repo_coordinated` code definition, the `unresolvable_statically` note, `fix-implement`'s single-repo paragraph, the two-unit recording block, Stage 4 and Stage 5. These files are prompts; a rule repeated with fresh justification each time reads as seven rules. Now each place says only what it owns. Step 4 decides which repos. The reason code says what it is for. The outputs section says how the two units are recorded. Stage 4 says to call the implement step twice, Stage 5 that one build covers both. The Summary template gives the ordered instructions without arguing for them. Net -32 lines against the same behaviour. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
Both scripts had grown prose that explained how a rule was arrived at rather than what it is. The module docstring was 33 lines and two of its claims had gone stale -- it still listed "grouping" as a rule (removed) and still said a human posts the triggers by hand (the cron does). Kept in every case: the fact that stops someone undoing the fix. `fix_in_flight` still says why it asks by status, `last_mirrored` still says why the match is start-anchored, `reopened_since` still says `closed_at` is cleared by a reopen, `dest_name` still says why the repo is in the directory name. Dropped: the incident each one came from. -61/+39 lines of prose, no behaviour change. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test python3 -m pytest .github/scripts/test_bot_export_fix_patch.py lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
"nobody has asked about yet" was the dedup rule smuggled into the summary line, where it also left a dangling question -- asked whom. Dedup has its own paragraph below. The line now names the mechanism: it posts a comment carrying the command. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
`cross_repo_coordinated` named oneDNN, IGC and the driver. Triton is missing, and so is whatever comes next, so the code now points at the vocabulary that already exists for this: the repo's `dependency component: *` labels (oneDNN, Triton, IGC, Level_Zero, oneAPI, driver, third_party, community...). Both examples say oneDNN/Triton/IGC/driver rather than picking one. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
`last_mirrored`'s docstring re-explained the start-anchored match, which the `FIX_CMD_RE` comment already does at the definition, with bot.yml's own source next to it. And `main`'s two-line note on `DRY_RUN` restated the usage block at the top of the file; what a reader needs at that line is why the comparison is against a literal, which is one line. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
18 lines to 12, and it now says one thing more. Deleted: the paragraph restating what `# print` / `# post` already say in the usage block, and the two lines defending the absence of per-test grouping -- a non-feature does not need a defence in the file, and "batches kept whole" carries it in three words. The list of which rule lives in which function went too; the useful half of that sentence is that there is no local state. Added, because it was missing and is headline behaviour: never while a `fix` job is still running. Test Plan: ``` python3 .github/scripts/ci_disabled_queue.py --self-test lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV ```
ZhaoqiongZ
force-pushed
the
bot/ci-disabled-queue
branch
from
September 24, 2026 01:38
4b5f376 to
3a29d1b
Compare
ZhaoqiongZ
marked this pull request as ready for review
September 24, 2026 03:17
Performance outliers, please check!
|
CuiYifeng
reviewed
Sep 28, 2026
`fix-root-cause` had eleven NEEDS_HUMAN reason codes. Nothing reads them: no script under `.github/` parses `reason`, and every one of them lands on the same terminal state -- `NEEDS_HUMAN` status, `agent:needs-human` label (execution-modes.md's stage table is the whole mapping). The claim in issue-handler that "each reason maps to a different final `agent:status` value" was simply false, and it was the only justification the list had. So the enum keeps the three codes an orchestrator actually branches on -- `no_registered_domain` (the retry policy refuses to loop on it), `invalid_reproduction` (re-run reproduce first), `security_concern` (stop now) -- and everything else is `other` with a mandatory one-line justification in `reason_detail`. Step 5's signal-to-code table becomes a list of blockers seen so far, marked explicitly as examples and not an enum, so the next unfamiliar failure gets described instead of filed under the nearest label. This deletes `cross_repo_coordinated` along with the rest, which undoes 2377f36 and 0aafbdb's work on it. That code was the clearest case for the change: after being redefined away from pytorch + torch-xpu-ops it needed a parenthetical at every use site to stop readers drawing the obvious wrong conclusion, and a name that has to be corrected wherever it appears costs more than it carries. The `dependency component: *` label vocabulary those commits found is kept -- it now lives in the justification, which is where a maintainer reads it. Also fixed the matching overclaim on `fix-implement`'s codes in Stage 4. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV -m origin/main python -m pytest .github/scripts/test_bot_export_fix_patch.py -q grep -rn "cross_repo_coordinated\|unresolvable_statically\|hardware_specific" --include=*.md --include=*.py --include=*.yml . ``` The grep returns nothing: no dangling references to the deleted codes.
Branch naming claimed one branch per bug so a reader "knows which issue a PR closes", and stopped there. With a `companion_repo` that name now exists in two checkouts and answers to two PRs, and the thing that tells them apart -- the `-<target_repo>` suffix on the patch directory -- lived only in `bot_export_fix_patch.py`. A reader looking the rule up would conclude one fix is one branch is one PR, and the obvious repair is to suffix the branch itself, which buys `fix-pytorch-issue-197521-pytorch-pytorch`. So the section now says the name is shared on purpose, where the repo does appear, and that a branch plus target_repo collision fails the export rather than overwriting half the patches. And the two-PR callout gains the rule that neither PR may close the issue: a closing keyword on the half that lands first marks the bug fixed while CI is still red on the other half. A human closes it once both are in. Test Plan: ``` lintrunner --skip CLANGTIDY,CLANGFORMAT,MERGE_CONFLICTLESS_CSV -m origin/main python -m pytest .github/scripts/test_bot_export_fix_patch.py -q ```
Performance outliers, please check!
|
CuiYifeng
approved these changes
Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
1. Automate
@torchxpubot fixon CI DISABLED issues. #5272 is the queue SOURCE is #5490,fix, wait for result, then another fix.The cron is on:
0 0,9-23 * * *, hourly over 17:00-08:00 Beijing2. Tracking issues for skip-adding UT: template and reference format.
3. Fix the patch collision issue when fix in both repo with same name/file.
Authored with an AI assistant.