Skip to content

fix(ci): preserve full-suite evidence and verify completion - #1369

Merged
apackeer merged 63 commits into
mainfrom
fix/full-suite-evidence-fixes
Sep 25, 2026
Merged

apackeer merged 63 commits into
mainfrom
fix/full-suite-evidence-fixes

Conversation

@apackeer

@apackeer apackeer commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Full Suite failures exposed timeouts below the cost of valid work on loaded runners, incomplete failure evidence, and incorrect workflow/render expectations. This change uses shared execution backstops throughout the test runner, native drivers, fixture subprocesses, tools, sensors, and supported hook registrations. Commands still finish immediately when their work completes; explicit timeout calibrations and ownership checks remain enforced.

The repair also checks completed workflow state, publishes cancellation observations atomically, retains sanitized failed journeys, and supplies the concrete defect requested by the requirements-analysis fixture. Full verification runs every credential-free job against an immutable PR head; live coverage of a candidate uses the existing live verification mode. Both remain distinct from release-purpose qualification.

Changes

  • Replace operational case/startup/process/lock limits with shared backstops and keep original deadlines across nested operations. Reserve cleanup time once and require observed process retirement.
  • Use five-minute ordinary, fifteen-minute compound, and thirty-minute enclosing tool defaults. Keep explicit caller limits, lock ownership/stale rules, automatic doctor responsiveness, and IDE broken-channel behavior.
  • Align sensor subprocesses, manifest caps, and supported native hook registrations. Codex trust hashes include the configured timeout, so changing a hook timeout changes its trust identity; omitted legacy timeouts retain their native default.
  • Keep non-live CI file/run/step/job limits consistent and retain the existing credentialed-live boundary. Full Suite continues to declare excluded providers rather than claiming all-harness coverage.
  • Replace timing proxies with direct readiness, ownership, completion, and cancellation evidence. Preserve deliberate timeout and parser-regression checks; probe timeouts cannot become capability skips.
  • Preserve release metadata. Claude SessionEnd retains its native 60-second cap; Cursor and Kiro IDE outer-timeout overrides remain unverified.
  • Tests that deliberately hold a lock that never releases pass the existing contention override, because the production wait is now as long as their process ceiling. Scratch fixtures copy the new runtime budget module, and inline Bun scripts in Windows PowerShell avoid embedded double quotes.
  • Lock release retries within the lock's own contention budget instead of giving up after half a second of gate contention or Windows rename refusals, so a sensor no longer keeps the audit lock while its subprocess runs.
  • The plan-approval guard lets the human answer a held Code Generation gate. Opening that gate moves the state past the issued directive, so a compliant conductor could not submit the approval, and under a relaxed fence the refusal named no remedy. Only that stage's approved or rejected report is admitted; the engine still requires the human's exact answer.
  • Live shards made only of production-guard journeys run with --production-guards, so the fix: make the guards a fence for the agents and a gate the human holds the key to #1262 live chat-lowering proof executes instead of skipping. The native terminal driver treats an endpoint closed by its own published stop as success only after the same generation retires.
  • A typed guard-policy or ceremony switch (--guard-policy, --change-control, --sensors, --learnings, --summary-confirmation) now ends the turn as the engine instructs. The Stop hook recognized only depth, test-strategy, and review as terminal configuration, so it told the agent to resume a workflow the human had only asked to reconfigure.
  • Before Plan Approval, output discarded to the null device (/dev/null, NUL) is no longer treated as a workspace write, so read-only probes stay available. Real writes are still refused.
  • A Plan Approval prompt recorded for a named --session the project has not seen still records, but its output warns that the human's answer cannot bind and names the AIDLC Runtime Session: line. With feat(plan-approval): auto-resolve --session from the invoking conversation #1379's auto-resolution, an omitted --session resolves from the invoking conversation and gets no warning; the refusal when nothing resolves and the unprompted-receipt refusal carry the same pointer. This removes the double approval prompt seen live, without adding a refusal.
  • The no-follow file reader treats a file replaced between open and fstat (zero links) as changed and retries it, instead of reporting a hardlink the user does not have.
  • Test fixes: revision recovery hands a structured single-select follow-up question to the answer loop by shape (it previously waited out the whole budget), and the native cancellation witness and the Windows lifecycle fixture's handshake records are published by rename.
  • Live waits end on the observed end of the agent's turn, not only on the clock. The TUI driver recognizes a finished turn by the screen (work was seen, then Claude's empty prompt stayed unchanged for 30 seconds; the longest idle-looking pause inside a turn across 517 recorded sessions was 0.4 seconds). wait, answer-gate, and revision recovery then fail or act at once with the last pane, instead of hanging until the file ceiling reports only a timeout. Unrecognized screens keep the backstop.
  • Deadline reporting names the actual cause: an overall answer-gate expiry, a native cleanup that outlived its deadline, and a supervisor error published at the deadline are no longer reported as other failures. The isolated supervisor reads the clock once per pass, so a shared ready and work deadline always expires the work.
  • Live files and runs get a 3,600-second ceiling (70-minute steps, 80-minute jobs, 64-minute Windows task). The one-hour Bedrock session is assumed just before the run step and model work stops at the cleanup reserve, so work always ends while credentials are valid.
  • Every captured CI test step creates its run.log before starting the background runner, so the log poll cannot race the redirect under load.
  • Live audit assertions read every shard; one test had read a stray agent-written file that sorted first. The native capture fixture breaks capture only after earlier output is captured.
  • Full verification of an unmerged head no longer reaches the credentialed live jobs (live_prepare, live_hosted, live_windows), and every Full Suite checkout sets persist-credentials: false; live coverage of a candidate uses live_verification. Manual deterministic dispatch defaults to the smoke tier.
  • The Codex workspace journey attributes a route only to one exact bun .codex/tools/aidlc.ts invocation inside Codex's own shell wrapper, including its Windows PowerShell wrapper, and a newly created in-flight intent must be Running.

Validation

  • Current head: ce548d9b. 0e16bdf7 merges main again (fix(windows): repair the cross-OS test failures on Windows #1393, fix(doctor): compare the audit against the per-stage checkboxes #1272, fix(sensor): drop a superseded detail file when the sensor passes #1266); 91bee555 merged main (feat(plan-approval): auto-resolve --session from the invoking conversation #1379, fix(sensor): route a sensor's path argument from its declared input_schema #1239, chore(ci): run the cross-OS jobs in the merge queue, not on every PR push #1389, chore(ci): supersede Full Suite verification across branch heads #1390, fix: never record an unreadable review findings table as no findings #1163). Release metadata matches main. Conflicts were resolved in the Plan Approval session path (main's auto-resolution kept, with this branch's hint on its refusal), the sensor import and t92 parameters, two reference docs, and the regenerated coverage registry.
  • Live verification on 91bee555 passed: result passed: true for all families, 196 of 196 jobs; it is the live evidence for Claude SDK, Codex, opencode, and the release contract. This branch's later changes are deterministic tests that no live job runs; main's fix(windows): repair the cross-OS test failures on Windows #1393, fix(doctor): compare the audit against the per-stage checkboxes #1272, and fix(sensor): drop a superseded detail file when the sensor passes #1266, merged after it, are not covered by this live run.
  • Full verification on ce548d9b passed: result passed: true, with the three live jobs omitted by design. Claude TUI live verification on ce548d9b passed: result passed: true, 62 of 62 live jobs. ce548d9b changes the native terminal driver's wait deadline, which the live TUI journeys use, so the TUI family was re-verified on this head. Two more Windows integration-tier runs under load passed all 131 files.
  • Full verification on 43e93253 passed: result passed: true, with the three live jobs omitted by design. Three further Windows integration-tier runs under load found two more flakes from this branch's tighter timing, both fixed in ce548d9b: a native terminal wait could outlive the operation deadline its captures budget against and crash with a budget error instead of reporting "timed out" (9405b6a4), and the outer file deadline case gave cleanup 250 ms, too little to confirm worker retirement on a loaded Windows runner (ce548d9b). Both runs are PR-verification evidence and are not consumed by preview or stable publication.
  • Full verification on 0e16bdf7 failed on Windows in three places. t334 and the swarm continuation test still expected the Windows backslashes that c9fcc485 had loosened them to accept; fix(windows): repair the cross-OS test failures on Windows #1393 fixed the guard to record portable paths, so a30d72f9 restores main's expectations. t46 caught one race in which a single process held the audit lock for the whole 300-second budget while four waited (only one BOLT_STARTED landed; the retained lock and a fresh claim both named that live owner). This branch's release retry loop (0e932d02) is the prime suspect; the cause is not yet known, so 43e93253 records each child's outcome and the lock names over the race, and repeated Windows runs gathered evidence before any lock change: 60 further races (30 isolated, 30 under full integration-tier load) all completed normally in at most 3.4 seconds, so the stall has not recurred and its cause is still unknown. The diagnostics stay in place so a recurrence names it.
  • The second merge resolved four test files that fix(windows): repair the cross-OS test failures on Windows #1393 also fixed for Windows: main's t340, t333, and t-guard-plan-continuation-swarm; and this branch's instrumented t-e2e-isolated-runner case with fix(windows): repair the cross-OS test failures on Windows #1393's Windows rule that an exited sibling is judged by its native identity, because signal 0 can still succeed while a handle remains.
  • Full verification on 91bee555 failed two Windows test defects, both fixed. t340 (from fix(doctor): report Kiro IDE ignore sources that hide .kiro/ #1161, merged today) wrote a native Windows path into a git config file, which git rejects as a bad config line, so the doctor correctly reported that it could not evaluate the source; de26b715 writes the path with forward slashes. The native cancellation fixture published its started witness before the target drew, so a prompt cancel archived an empty screen; df9aa671 waits for the target first.
  • Local on this head: bun run check and the coverage registry pass. The full local unit and integration run on the merge (487 files, 11,198 executed cases) failed only the registry freshness check, regenerated in 2ad636ea, and two t345 cases that lost the captured-log race fixed in 91bee555; both files pass on rerun.
  • Before the merge, 728504b6 had complete evidence: full verification passed, Codex live passed on e1328bba, and Claude SDK, Claude TUI, opencode, and release-contract live passed on c04fafe6.
  • Full verification on e1328bba was cancelled after its legacy Windows node-pty job failed: a kill that the persistent non-root CIM case injects to fail spent the whole 15-minute cleanup backstop (30 seconds before this change raised the backstop), and the retry's one-shot liveness check then saw the session's recorded PID alive. 728504b6 bounds only that deliberately refused kill to its former 30 seconds.
  • On c04fafe6, full verification passed (result passed: true, with the three live jobs omitted by design). Its live run had two Windows failures. The Codex workspace journey rejected Codex's non-login wrapper (pwsh.exe -NoProfile -Command; the model chooses a login or non-login shell per call, and a non-login POSIX call is <shell> -c), fixed in e1328bba. A t-tui-custom-harness SDK case passed but its project folder stayed busy for the whole cleanup window: on Windows the SDK's abort kills only the CLI, about 7 seconds later, and the CLI had started a Bash tool just before. That harness race is not a core defect; the job was rerun and passed (5 of 5 cases in 254 seconds), and containing the SDK CLI in a Windows job object is a follow-up.
  • The first pair on d9bf2ea3 was cancelled early, the live run before any live job started: full verification's Windows native lifecycle test read target.json while its fixture was still writing it ("Unexpected EOF"). c04fafe6 publishes that record by rename, like the fixture's other handshake files.
  • The run on a3ca141b was cancelled after 152 passing jobs and no failures, when 4e7fd87d moved the head and removed credentialed jobs from full verification. Its Windows Codex evidence showed the new exact-invocation parser rejecting all 17 of Codex's PowerShell-wrapped commands, which would fail the Windows workspace journey; d9bf2ea3 accepts that wrapper, and the 55 recorded Linux, macOS, and Windows commands all parse.
  • The run on 2e703c94 was cancelled after 119 passing jobs to pick up two fixes for its only failures: a Windows lifecycle fixture read its exit-code file before the test finished writing it, and t29 asserted the engine error's wording although the conductor relays it in its own words (it now asserts the engine error in the native transcript and its facts on screen). The new turn-end wait ended that t29 case 30 seconds after the turn, not at the 40-minute ceiling.
  • The previous complete run passed on d6320483 (236 of 236 jobs). A repeat on the same head passed 233 jobs and failed two, both fixed here: the supervisor ready/work deadline race (Windows unit tests) and the single-shard audit read (t138 on Windows).
  • Live-only verification ran the fix: make the guards a fence for the agents and a gate the human holds the key to #1262 live chat-lowering journey with production guards on Linux, macOS, and Windows: 1 executed, 0 skipped, coverage complete on each.

The two current-head runs are the merge-readiness evidence; focused local passes do not replace them. Full Suite retains its declared provider exclusions.

Checklist

  • I have reviewed the contributing guidelines
  • I have performed a self-review of this change
  • Changes have been tested
  • Changes are documented
  • Hook timeout changes are reflected in the Codex trust identity, as described above

Acknowledgment

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of the project license.

@github-actions github-actions Bot added documentation Improvements or additions to documentation github labels Sep 23, 2026
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

AIDA findings ledger

ID Sev Status Title Decided by
F1 P1 🔴 open Substring matching lets a bypassed command satisfy the Codex journey —
F2 P2 🔴 open Progress-tolerant creation omits the workflow status invariant —
F3 P3 ✅ resolved The test README still documents trace retention as opt-in AIDA · 8f302f1
F4 P3 🔴 open Production collection suppresses its recovery instruction —
F5 P3 🔴 open Supply-chain documentation still describes verification as live-only —
F6 P0 🔴 open Full verification exposes credentials to PR-controlled code —
F7 P3 🔴 open Manual deterministic dispatch defaults to an invalid selection —
F8 P2 🔴 open Failed SDK fixtures are retained outside uploaded evidence —
F9 P2 🔴 open Test-only timeout hides the shipped linter budget failure —

Open blocking findings (P0/P1): 2. Accepted and rejected findings never count toward the next action.

Maintainer commands (repository write access) — put them on the first lines of a comment, one per line, several ids per line allowed:
/aida accept F# [F#…] <reason> · /aida reject F# [F#…] <reason> · /aida reopen F# [F#…] · /aida status · /aida full (next review covers the whole head)
P0 and P1 findings can be accepted (visible, risk owned by the maintainer) but not rejected. A comment is applied all-or-nothing.
Do not edit this comment: AIDA verifies its digest and refuses to run on an edited ledger. To start over, delete it.

ledger.json
{
  "version": 4,
  "pullRequest": 1369,
  "nextId": 10,
  "findings": [
    {
      "id": "F1",
      "priority": "P1",
      "category": "correctness",
      "title": "Substring matching lets a bypassed command satisfy the Codex journey",
      "anchors": [
        {
          "kind": "line",
          "path": "tests/harness/codex-workspace-evidence.ts",
          "side": "RIGHT",
          "sha256": "3175ceaf00f1ff72870f1dfe6e502bd729345b4cacbcf04c3933bd0755a298c0"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
        "at": "2026-09-23T13:09:29.126Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    },
    {
      "id": "F2",
      "priority": "P2",
      "category": "workflow-state",
      "title": "Progress-tolerant creation omits the workflow status invariant",
      "anchors": [
        {
          "kind": "line",
          "path": "tests/harness/codex-workspace-evidence.ts",
          "side": "RIGHT",
          "sha256": "65320124378afc5e354149ef31d7971cac2db3c3f503217edbded259fb38b217"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
        "at": "2026-09-23T13:09:29.126Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    },
    {
      "id": "F3",
      "priority": "P3",
      "category": "contracts",
      "title": "The test README still documents trace retention as opt-in",
      "anchors": [
        {
          "kind": "line",
          "path": "docs/reference/09-testing.md",
          "side": "RIGHT",
          "sha256": "b578cffe1850d94c1bb9389fa27e396abfb52e7e461b3d1640c8c15ecc12c4bf"
        }
      ],
      "status": "resolved",
      "firstSeen": {
        "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
        "at": "2026-09-23T13:09:29.126Z"
      },
      "lastSeen": {
        "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
        "at": "2026-09-23T23:53:49.492Z"
      }
    },
    {
      "id": "F4",
      "priority": "P3",
      "category": "user-experience",
      "title": "Production collection suppresses its recovery instruction",
      "anchors": [
        {
          "kind": "line",
          "path": ".github/scripts/prepare-live-runtime.ps1",
          "side": "RIGHT",
          "sha256": "8b4e82e0b43f1c15d0023eccb223f3a324cceab08b961553fe1806e294c53499"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
        "at": "2026-09-23T13:09:29.126Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    },
    {
      "id": "F5",
      "priority": "P3",
      "category": "contracts",
      "title": "Supply-chain documentation still describes verification as live-only",
      "anchors": [
        {
          "kind": "line",
          "path": ".github/workflows/full-suite.yml",
          "side": "RIGHT",
          "sha256": "063a689273e67be6489ba7eff6888de8d5f95fd5d77e0bc53f1397387573c087"
        },
        {
          "kind": "line",
          "path": ".github/workflows/full-suite.yml",
          "side": "RIGHT",
          "sha256": "311aece3ce4ca5c124e4060d4c265d49a9b537580ee6c9b5735b93c214495f36"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
        "at": "2026-09-23T19:29:37.875Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    },
    {
      "id": "F6",
      "priority": "P0",
      "category": "security",
      "title": "Full verification exposes credentials to PR-controlled code",
      "anchors": [
        {
          "kind": "line",
          "path": ".github/workflows/full-suite.yml",
          "side": "RIGHT",
          "sha256": "063a689273e67be6489ba7eff6888de8d5f95fd5d77e0bc53f1397387573c087"
        },
        {
          "kind": "line",
          "path": ".github/workflows/full-suite.yml",
          "side": "RIGHT",
          "sha256": "247cfadf662d36db893b1cd15c2504562bf296a6bb910afd837662754d56061c"
        },
        {
          "kind": "line",
          "path": ".github/workflows/full-suite.yml",
          "side": "RIGHT",
          "sha256": "e7d162d250978200908408bf975276c531eae3200cb7fc525da667b4473ab697"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
        "at": "2026-09-23T23:53:49.492Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    },
    {
      "id": "F7",
      "priority": "P3",
      "category": "user-experience",
      "title": "Manual deterministic dispatch defaults to an invalid selection",
      "anchors": [
        {
          "kind": "line",
          "path": ".github/workflows/deterministic-tests.yml",
          "side": "RIGHT",
          "sha256": "d50a8f57f8baa0334395bbdbafbc9ca88986c96b9644b7ee03cdd8ead702d865"
        },
        {
          "kind": "line",
          "path": ".github/workflows/deterministic-tests.yml",
          "side": "RIGHT",
          "sha256": "d5b9b63bf7234c3ce987cd7caafc84c34299b6f811183b28e9382171b5436dd6"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
        "at": "2026-09-23T23:53:49.492Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    },
    {
      "id": "F8",
      "priority": "P2",
      "category": "workflow-state",
      "title": "Failed SDK fixtures are retained outside uploaded evidence",
      "anchors": [
        {
          "kind": "line",
          "path": "tests/integration/t183-codekb-placement-reverify.sdk.test.ts",
          "side": "RIGHT",
          "sha256": "a10105e36d38bf077bb49cfa8c7cbb74e38a833e502734de8e3648339b2992a6"
        },
        {
          "kind": "line",
          "path": "tests/integration/t193-compose-report-journey.sdk.test.ts",
          "side": "RIGHT",
          "sha256": "699c9add7ffd5e3dbf68a818c578d800f0e1d059a986f05afbda2d63a2fec376"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "dc201a91743e593c75d720c24007934894b79661",
        "at": "2026-09-24T02:14:21.134Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    },
    {
      "id": "F9",
      "priority": "P2",
      "category": "correctness",
      "title": "Test-only timeout hides the shipped linter budget failure",
      "anchors": [
        {
          "kind": "line",
          "path": "tests/integration/t92.test.ts",
          "side": "RIGHT",
          "sha256": "eae13c11ca2c0cdb9005cc9731b355a351469576418da1e6ca24e264a609232c"
        },
        {
          "kind": "line",
          "path": "tests/integration/t92.test.ts",
          "side": "RIGHT",
          "sha256": "f1ca2c5d0c83a736cdd132d1fc0abf39ff8c1a83d53466bfc9a769091d8823bd"
        }
      ],
      "status": "open",
      "firstSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      },
      "lastSeen": {
        "head": "79f328fe743afc23357283893d7bbde866fd91b1",
        "at": "2026-09-24T03:11:53.553Z"
      }
    }
  ],
  "events": [
    {
      "at": "2026-09-23T13:09:29.126Z",
      "kind": "opened",
      "by": "aida",
      "id": "F1",
      "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
    },
    {
      "at": "2026-09-23T13:09:29.126Z",
      "kind": "opened",
      "by": "aida",
      "id": "F2",
      "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
    },
    {
      "at": "2026-09-23T13:09:29.126Z",
      "kind": "opened",
      "by": "aida",
      "id": "F3",
      "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
    },
    {
      "at": "2026-09-23T13:09:29.126Z",
      "kind": "opened",
      "by": "aida",
      "id": "F4",
      "head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
    },
    {
      "at": "2026-09-23T19:29:37.875Z",
      "kind": "seen",
      "by": "aida",
      "id": "F1",
      "head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T19:29:37.875Z",
      "kind": "seen",
      "by": "aida",
      "id": "F2",
      "head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T19:29:37.875Z",
      "kind": "seen",
      "by": "aida",
      "id": "F3",
      "head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T19:29:37.875Z",
      "kind": "seen",
      "by": "aida",
      "id": "F4",
      "head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T19:29:37.875Z",
      "kind": "opened",
      "by": "aida",
      "id": "F5",
      "head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9"
    },
    {
      "at": "2026-09-23T23:53:49.492Z",
      "kind": "seen",
      "by": "aida",
      "id": "F1",
      "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T23:53:49.492Z",
      "kind": "seen",
      "by": "aida",
      "id": "F2",
      "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T23:53:49.492Z",
      "kind": "resolved",
      "by": "aida",
      "id": "F3",
      "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
      "reason": "declared corrected by the judge; cited files changed since the last review"
    },
    {
      "at": "2026-09-23T23:53:49.492Z",
      "kind": "seen",
      "by": "aida",
      "id": "F4",
      "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T23:53:49.492Z",
      "kind": "seen",
      "by": "aida",
      "id": "F5",
      "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-23T23:53:49.492Z",
      "kind": "opened",
      "by": "aida",
      "id": "F6",
      "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5"
    },
    {
      "at": "2026-09-23T23:53:49.492Z",
      "kind": "opened",
      "by": "aida",
      "id": "F7",
      "head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5"
    },
    {
      "at": "2026-09-24T02:14:21.134Z",
      "kind": "seen",
      "by": "aida",
      "id": "F6",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:14:21.134Z",
      "kind": "seen",
      "by": "aida",
      "id": "F1",
      "head": "dc201a91743e593c75d720c24007934894b79661",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T02:14:21.134Z",
      "kind": "seen",
      "by": "aida",
      "id": "F2",
      "head": "dc201a91743e593c75d720c24007934894b79661",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T02:14:21.134Z",
      "kind": "seen",
      "by": "aida",
      "id": "F4",
      "head": "dc201a91743e593c75d720c24007934894b79661",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T02:14:21.134Z",
      "kind": "seen",
      "by": "aida",
      "id": "F5",
      "head": "dc201a91743e593c75d720c24007934894b79661",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T02:14:21.134Z",
      "kind": "seen",
      "by": "aida",
      "id": "F7",
      "head": "dc201a91743e593c75d720c24007934894b79661",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T02:14:21.134Z",
      "kind": "opened",
      "by": "aida",
      "id": "F8",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:40:42.380Z",
      "kind": "seen",
      "by": "aida",
      "id": "F6",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:40:42.380Z",
      "kind": "seen",
      "by": "aida",
      "id": "F1",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:40:42.380Z",
      "kind": "seen",
      "by": "aida",
      "id": "F2",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:40:42.380Z",
      "kind": "seen",
      "by": "aida",
      "id": "F8",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:40:42.380Z",
      "kind": "seen",
      "by": "aida",
      "id": "F4",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:40:42.380Z",
      "kind": "seen",
      "by": "aida",
      "id": "F5",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T02:40:42.380Z",
      "kind": "seen",
      "by": "aida",
      "id": "F7",
      "head": "dc201a91743e593c75d720c24007934894b79661"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "seen",
      "by": "aida",
      "id": "F6",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "seen",
      "by": "aida",
      "id": "F1",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "seen",
      "by": "aida",
      "id": "F2",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "seen",
      "by": "aida",
      "id": "F4",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "seen",
      "by": "aida",
      "id": "F5",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "seen",
      "by": "aida",
      "id": "F7",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "seen",
      "by": "aida",
      "id": "F8",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1",
      "reason": "retained: still open per the judge, not restated"
    },
    {
      "at": "2026-09-24T03:11:53.553Z",
      "kind": "opened",
      "by": "aida",
      "id": "F9",
      "head": "79f328fe743afc23357283893d7bbde866fd91b1"
    }
  ],
  "review": {
    "head": "79f328fe743afc23357283893d7bbde866fd91b1",
    "readiness": 1,
    "risk": 5,
    "decision": "change"
  }
}

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 94f05c1298ff1e92405d6f1d7db1d8e35c401356 against 79ff00818620014ffcd3749414e2de9c7e62157e and current repository behavior.

Inspection: 12 changed files. Scope: full head (first review of this pull request).

Final Assessment

Human decision aid only: Readiness 5/5 is best; Risk 1/5 is best. These scores inform the maintainer; the next action below follows finding severity (any open P0/P1 → author/change) and does not approve or merge the PR.

Readiness: 2/5 — The evidence-retention changes are substantial, but the new Codex verifier can produce a false green for a release-gating journey because it does not bind success to an actual allowed utility invocation. Additional workflow-state and documentation corrections remain.

Risk: 4/5 — The blocking flaw affects required hosted live evidence used for stable publication. A broken or bypassed Codex command path could be recorded as successful, although the impact is limited to validation integrity rather than production data mutation.

Decision required: Author — make changes before this PR proceeds. The P1 command-evidence flaw must be corrected before this release-gating verification can be trusted.

Validation performed:

  • Inspected all 12 changed-file snapshots, the full diff, PR metadata, review scope, discussion, findings ledger, and specialist outputs.
  • Compared the Codex evidence helper with its callers, regression tests, existing command validator, state-machine contract, and required Full Suite release path.
  • Compared Windows collection changes with their production caller, native fixtures, tests, workflow upload steps, sanitizer behavior, and testing documentation.
  • The ledger contains no open, accepted, rejected, or archived findings requiring disposition.

Findings: 1 blocking, 3 advisory.

Ledger: 4 open, 0 retained blocking, 0 accepted, 0 suppressed as rejected by a maintainer. Maintainers act on findings with /aida commands in the ledger comment.

Correctness & Reliability

P1 [F1]: Substring matching lets a bypassed command satisfy the Codex journey

Evidence: tests/harness/codex-workspace-evidence.ts:33.

Problem: A completed command is selected solely because its agent-controlled shell text contains the tool path and matches the route regex. A compound command, shell comment, or decoy invocation can include those strings while another command directly constructs the expected registry, state, include, and audit files. Because the model has workspace-write access and the resulting command event can exit zero, turn completion and every disk assertion can pass without the required utility path succeeding.

Impact: The required Codex Full Suite journey can produce false release evidence for a broken or bypassed workflow, allowing stable publication to rely on an invalid supported-harness verification.

Required correction: Parse and validate one exact allowed utility invocation, rejecting shell composition, substitutions, redirections, comments, and unrelated commands. Add adversarial tests showing that valid-looking disk state cannot rescue compound or decoy commands.

Workflow, State & Recovery

P2 [F2]: Progress-tolerant creation omits the workflow status invariant

Evidence: tests/harness/codex-workspace-evidence.ts:112.

Problem: When stopAfterCreate is false, the helper skips every assertion inside this branch, including Status: Running. The first skill-driven journey can therefore accept a newly registered in-flight intent whose state is missing Status or incorrectly says Completed or Archived, despite the canonical workflow machine requiring a newly created intent to remain Running.

Impact: The live Codex journey can report successful creation while persisting contradictory or unusable lifecycle state, weakening the release evidence even after command validation is corrected.

Required correction: Always require workflow Status to be Running for a newly created in-flight intent. Keep only the exact handoff stage and last-completed-stage assertions conditional, and add negative coverage for invalid status when stopAfterCreate is false.

User Experience

User experience change: CI maintainers diagnosing Full Suite failures receive retained sanitized traces by default and independently collected Windows launch logs, while completed Codex turns may now be accepted without command output when disk evidence exists.

Before: Eligible traces were dropped unless explicitly enabled, Windows test-log collection failure could prevent launch-log preservation, and Codex verification required completion text.

After: Eligible traces are retained unless explicitly disabled, Windows collection publishes complete sources independently with a report, and Codex verification uses command events plus workspace state.

Example: Set AIDLC_NIGHTLY_UPLOAD_TRACES=0 to opt out; failed Windows collection retains windows-collection-<uuid>.json and any independently valid launch logs.

Assessment: Failure diagnosis improves, but permissive Codex command matching can misleadingly present a bypassed journey as successful, and the production collection error hides its new recovery instruction.

P3 [F4]: Production collection suppresses its actionable recovery message

Evidence: .github/scripts/prepare-live-runtime.ps1:397.

Problem: Collect-RuntimeLogs throws a fixed safe instruction to inspect the collection report and retained independent logs, but the unchanged top-level catch replaces it with a generic stage, exception type, and line-number message. The new native tests call the function directly and therefore do not cover what workflow operators actually see.

Impact: Windows CI maintainers are not told where the newly retained recovery evidence is located when collection fails.

Required correction: Surface a fixed collection-specific recovery message from the production catch without echoing arbitrary exception text, and add a caller-level test for the emitted failure output.

Contracts & Compatibility

P3 [F3]: The test README still documents trace retention as opt-in

Evidence: docs/reference/09-testing.md:1532.

Problem: The workflows and reference guide now make retention the default with AIDLC_NIGHTLY_UPLOAD_TRACES=0 as the opt-out, but tests/README.md still says traces are dropped by default and that setting the variable to 1 opts in.

Impact: Contributors receive contradictory guidance about uploaded artifact contents and residual disclosure risk.

Required correction: Update tests/README.md to document default retention and the value 0 as the repository-level opt-out.

Residual risk: Repository code and tests were not executed, as required by the read-only review contract. Windows collection behavior was assessed statically rather than on a Windows host.

Reviewed by AIDA (AI-DLC Developer Agent).

[AI-PR-REVIEWED] 94f05c1

@github-actions github-actions Bot added action:change AIDA indicates changes are required before the PR proceeds aida:reviewed AIDA successfully reviewed the latest PR state next:author AIDA indicates the PR author needs to act next labels Sep 23, 2026
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 17:57 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 17:57 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 17:57 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:02 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:02 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:02 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:05 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:05 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:05 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:06 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:06 — with GitHub Actions Active
@apackeer
apackeer deployed to ai-pr-review September 23, 2026 18:06 — with GitHub Actions Active
@github-actions github-actions Bot removed aida:reviewed AIDA successfully reviewed the latest PR state next:author AIDA indicates the PR author needs to act next action:change AIDA indicates changes are required before the PR proceeds labels Sep 23, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 60f2833e249612ce1ea0db4c55596d37f8dff5e9 against 79ff00818620014ffcd3749414e2de9c7e62157e and current repository behavior.

Inspection: 15 changed files. Scope: incremental — 6 files with lines changed since the review at 94f05c12; the security lenses reviewed the full head. Findings on lines unchanged since that review are deferred, not decisive.

Final Assessment

Human decision aid only: Readiness 5/5 is best; Risk 1/5 is best. These scores inform the maintainer; the next action below follows finding severity (any open P0/P1 → author/change) and does not approve or merge the PR.

Readiness: 2/5 — The full-verification and diagnostic-retention work is substantial, but the required Codex journey can still produce false release evidence because command success is established by substring matching. Additional state, recovery, and documentation defects remain.

Risk: 4/5 — The P1 flaw affects required hosted Codex evidence consumed by the stable-release workflow. A bypassed utility invocation can appear successful, although the direct impact is validation integrity rather than production data loss.

Decision required: Author — make changes before this PR proceeds. Re-derived from the ledger: 1 open blocking finding (F1) was not restated this run and the cited code is unchanged, so the author still needs to act. Judge's note, superseded by finding severity: Re-derived: 4 findings outside the incremental review scope were deferred and no blocking finding remains. Judge's note, superseded by finding severity: The open P1 Codex command-evidence flaw must be corrected before release-gating verification can be trusted.

Validation performed:

  • Inspected all 15 changed-file snapshots, the complete diff, incremental scope, PR metadata, discussion, ledger, and specialist outputs.
  • Compared workflow changes with result reduction, release-consumer checks, authorization tests, Codex evidence contracts, Windows collection callers, and testing documentation.
  • Confirmed all four open ledger defects remain present; no accepted or rejected decisions apply.

Findings: 0 blocking, 1 advisory.

Ledger: 5 open, 1 retained blocking, 3 retained advisory, 0 accepted, 0 suppressed as rejected by a maintainer. The next decision was re-derived from the ledger. Maintainers act on findings with /aida commands in the ledger comment.

Contracts & Compatibility

P3 [F5]: Supply-chain references still describe candidate verification as live-only

Evidence: .github/workflows/full-suite.yml:21, .github/workflows/full-suite.yml:698.

Problem: The new mode runs every declared job and emits full-suite-verification-result, but docs/reference/19-supply-chain-security.md still says candidate verification intentionally omits native, deterministic, and production-guard jobs. The artifact inventory in docs/reference/09-testing.md also omits the purpose-specific result artifacts.

Impact: Maintainers receive an inaccurate account of privileged candidate execution and may look for the wrong evidence artifact.

Required correction: Document both verification modes in the supply-chain reference, including authorization, coverage, filters, artifact names, purposes, and release ineligibility. Update the artifact inventory accordingly.

User Experience

User experience change: Maintainers can run the complete Full Suite against an unmerged PR and receive more retained diagnostic evidence from failed CI runs.

Before: Candidate verification covered live jobs only, eligible traces were opt-in, and Windows collection failures could lose independent launch evidence.

After: full_verification=true runs every declared job with separate non-release evidence; traces are retained by default and Windows sources are collected independently.

Example: gh workflow run full-suite.yml --ref '<candidate-branch>' -f 'ref=<exact-head-sha>' -f full_verification=true

Assessment: The workflow is more diagnosable and gives maintainers broader pre-merge validation, but stale guidance and suppressed recovery text make the new behavior harder to use safely.

Retained advisory findings

Open P2/P3 findings from the ledger that the judge marked still open without restating them. They do not affect the decision.

P2 [F2]: Progress-tolerant creation omits the workflow status invariant — first reported at 94f05c12.

P3 [F3]: The test README still documents trace retention as opt-in — first reported at 94f05c12.

P3 [F4]: Production collection suppresses its actionable recovery message — first reported at 94f05c12.

Retained blocking findings

Open P0/P1 findings from the ledger that this review did not restate and whose cited code is unchanged. They keep the next action with the author until the code changes or a maintainer accepts them.

P1 [F1]: Substring matching lets a bypassed command satisfy the Codex journey — first reported at 94f05c12; cited: tests/harness/codex-workspace-evidence.ts.

Deferred (outside the review scope)

Reported by the judge but citing only lines unchanged since the review at 94f05c12. They do not affect the decision; comment /aida full to have the next review cover the whole head.

P1 · correctness: Substring matching lets a bypassed command satisfy the Codex journey — tests/harness/codex-workspace-evidence.ts

P2 · workflow-state: Progress-tolerant creation omits the workflow status invariant — tests/harness/codex-workspace-evidence.ts

P3 · contracts: The test README still documents trace retention as opt-in — docs/reference/09-testing.md

P3 · user-experience: Production collection suppresses its actionable recovery message — .github/scripts/prepare-live-runtime.ps1

Residual risk: Repository code and tests were not executed under the read-only review contract. Windows collection behavior was assessed statically rather than on a Windows host.

Reviewed by AIDA (AI-DLC Developer Agent).

[AI-PR-REVIEWED] 60f2833

@github-actions
github-actions Bot dismissed their stale review September 23, 2026 19:29

Superseded by AI review of 60f2833

@github-actions github-actions Bot added the action:change AIDA indicates changes are required before the PR proceeds label Sep 23, 2026
apackeer and others added 14 commits September 24, 2026 15:35
When a file's work tail is shorter than native startup, the ready deadline is the work cutoff. The loop read the clock twice per pass, so a pass that crossed the cutoff between the two reads skipped the work expiry and reported "e2e supervisor did not become ready" instead (t-bun-test-deadlines on Windows, repeat Full Suite). Use one reading per pass.
Three live failures on this PR were a driver waiting for something that could no longer happen: t139 on an unrecognized menu (25 minutes), t29 after the agent had already finished (33 minutes), and the earlier t126 deadlock. Each ended only at the file ceiling, so the report said "exceeded file ceiling" instead of what went wrong.

The TUI driver now recognizes a finished turn by the screen: work was seen (status spinner or its elapsed timer, a background-agent wait, a running subagent row, or a running command's background hint) and Claude's empty prompt then stayed unchanged for 30 seconds. Across 517 recorded sessions the longest idle-looking pause inside a turn was 0.4 seconds. An unrecognized screen counts as working, so the backstop still applies.

- wait fails with the last pane when its pattern never painted (--through-turn-end opts out).
- answer-gate fails when no menu appeared and its terminator is unmet, and an overall-budget expiry is reported as such rather than as a per-gate timeout.
- Revision recovery types free-text feedback only at an idle prompt; the quiet timer it replaces could type while a structured menu was still painting.
- The t27 audit wait and t29 scope wait use the same observation.
- Native cleanup no longer reports "native kill exited 0" when the kill succeeded but the cleanup deadline passed before retirement could be confirmed.
- Native daemon startup and target launch surface a supervisor error published in the last poll interval instead of a timeout.
- An endpoint seen vacant only after the deadline is reported as unconfirmed, not as still active.
auditFilePathFor returns the first .md in the audit folder. t138 on Windows failed with zero STAGE_STARTED events because the agent had appended its own notes to a second file there, and that file sorted first. Audit assertions now read every shard, as the engine does.
The capture fixture replaced its log with a directory and then printed its marker. When Bun's lazily flushed file header reached the pipe as a separate chunk after the swap, that chunk failed capture first and the runner stopped the child before the marker was printed. Print a marker, wait until the runner has written it, then break capture.
…al session

Live journeys normally take 24 to 29 minutes, against about 35 minutes of usable work in the 2,400-second file ceiling, so one detour ended a journey that was still progressing (t51 on macOS). With waits now ending on the observed end of a turn, a stuck journey fails within about 30 seconds, so a generous ceiling only serves progressing work.

Live files and runs get 3,600 seconds. The one-hour Bedrock role session is assumed about 13 seconds before the run step and model work stops at the five-minute cleanup reserve, so it always ends while its credentials are valid; the constant documents that bound. Live steps become 70 minutes, live jobs 80 minutes, and the Windows isolated run task 64 minutes.
The target fixture waits for finish to exist and then reads its exit code; the test wrote it in place, so a read between creation and write got an empty string and exited 0 (t-tui-bun-lifecycle on Windows, competing exit and stop). leaf.json had the same shape, and its reader treats a partial file as a failure. Write each to a sibling temp file and rename it into place.
The case waited for the engine's error text verbatim, but the conductor relays engine errors in its own words more often than not (11 of 14 across Full Suite traces), accurately and with a remedy, so it failed two runs in three. Assert the error where it is deterministic, as the engine's error result in the native transcript, and require the screen to name AWS_AIDLC_DEFAULT_SCOPE and its value; the no-state assertion is unchanged.
… Codex evidence

Review follow-up for AIDA F6 (P0), F1 (P1), F2, F5 and F7.

- F6: full_verification may select an unmerged head, so it no longer runs
  live_prepare, live_hosted or live_windows, the jobs that receive OIDC and
  the AWS role; live coverage of a candidate keeps the separately authorized
  live_verification mode. The result reducer requires exactly those three to
  be skipped. Every Full Suite checkout now sets persist-credentials: false.
  A contract test evaluates every credentialed job's condition under
  full-verification and pins persist-credentials on every checkout.
- F1: the Codex journey evidence parses one exact
  `bun .codex/tools/aidlc.ts ...` invocation inside Codex's own shell wrapper
  and matches the route against its argv. A command that names the tool and
  route but composes, redirects, substitutes, comments, or decoys fails
  outright; adversarial tests show valid disk state cannot rescue it.
- F2: a newly created in-flight intent must be Running even when the beat
  may progress past bootstrap.
- F5: the supply-chain reference documents both verification modes, and
  the artifact inventory names each purpose-specific result.
- F7: manual deterministic dispatch defaults to the executable smoke tier.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y evidence

The exact-invocation parser unwrapped only Codex's POSIX `<shell> -lc` form. On Windows, Codex reports every command as `"C:\\Program Files\\PowerShell\\7\\pwsh.exe" -Command '...'`, so all 17 utility commands in Full Suite 36043454212's Windows workspace journey were rejected, and any that carries a checked route fails the journey. Recognize that wrapper as well; the inner command keeps the same plain-word rules, which already reject PowerShell composition, substitution and redirection. The unit cases use the observed Windows shape and add PowerShell decoys and a foreign executable.
The target fixture wrote target.json in place while the test polled it, and the reader deliberately treats a partial file as a failure, so full verification 36058777252 read an empty file ("Unexpected EOF") in the detached-descendant case. Publish it by rename, as leaf.json and finish already are. The other cross-process files are safe: status.json is published whole by the supervisor, and release, stop, and ctrl-c are only tested for existence.
…evidence

The model chooses per call whether Codex runs a login shell: Codex's shell tool takes a `login` parameter, and a non-login call runs `<shell> -c` or `pwsh.exe -NoProfile -Command`. The parser accepted only the login forms, so live verification 36061390100's Windows workspace journey failed on `pwsh.exe -NoProfile -Command 'bun .codex/tools/aidlc.ts engine intent create ...'` after the intent was created correctly. Accept both forms; the inner command rules are unchanged. Of 106 utility commands across 36 recent Codex artifacts, the only one still rejected is a genuine compound printf.
The persistent non-root CIM case injects a process-kill failure that can never clear, so its first kill must refuse. Since the shared cleanup backstop rose from 30 seconds to 15 minutes, that refusal spent the full 15 minutes, and the retry's one-shot liveness check then found the session's recorded PID alive (full verification 36066982685). A PID that long exited is free for Windows to reuse. Bound only the refused kill to its former 30 seconds; the retry keeps the full backstop.
Brings in #1379 (Plan Approval auto-resolves --session), #1239 (sensor path argument from its input schema), #1389 (cross-OS CI jobs in the merge queue), #1390 (Full Suite supersedes runs across branch heads) and #1163.

Conflict resolutions:
- aidlc-log.ts: main's resolvePlanApprovalSession at every site; its refusal when no session resolves keeps this branch's AIDLC Runtime Session hint. The unknown-session warning now applies only to a named --session, since an auto-resolved one came from the invoking conversation.
- aidlc-sensor.ts and t92: both sides' imports and parameters.
- 06-hooks-and-tools.md: main's session sentence plus this branch's warning sentence, reworded for auto-resolution.
- 09-testing.md: main's pull-request and merge-queue rows plus the manual deterministic dispatch row; the budget hierarchy names merge-queue native-terminal checks.
- Coverage registry regenerated.
- t265 blanks the hook-injected session override for its omitted-session case, as t328 does.
…tions

The registry was regenerated during the merge before the projections were rebuilt, so it enumerated the pre-merge functions and missed those #1163 added (1,134 units instead of 1,140).
Every captured test step starts run-tests.sh in the background with its output redirected to run.log, then immediately polls that file with sed under `set -euo pipefail`. The background job creates the file only when it runs its redirection, so a reader scheduled first exits 2 and ends the step. t345 executes this script and failed that way twice during a full local -P 8 run ("sed: can't read .../tmp/ci-deterministic/run.log"), while passing alone; a loaded CI runner can lose the same race. Create the file after its directory in all four steps (ci.yml native and node-pty, full-suite.yml native, deterministic-tests.yml). A 300-iteration CPU-load stress did not reproduce the failure either way, so the fix rests on the mechanism, not a reproduction.
…t on Windows

The custom core.excludesFile case wrote the fixture's native path into a git config file. Git reads backslashes in a config value as escapes, so on Windows the line is invalid ("fatal: bad config line 2", exit 128) and the doctor correctly warns that it could not evaluate the source instead of reporting the rule, which the case expected (full verification 36094162239, Windows unit-5). Write the path with forward slashes, which git accepts on every OS. The product check is unchanged.
…get is on screen

The runner fixture's inner test started its native session and published the started witness at once, without waiting for the target to draw, unlike the file's own start helper. In cancel mode the parent cancels as soon as the witness exists, so on Windows the runner archived the session before the target printed and the archived screen was empty (full verification 36094162239, Windows integration). Wait for NATIVE READY before publishing the witness; the archive itself records the screen correctly.
Brings in #1393 (Windows cross-OS test repairs and the plan-approval guard's portable blocked path), #1272 and #1266.

Conflict resolutions:
- t340, t333, t-guard-plan-continuation-swarm: main's versions, which carry the same Windows fixes this branch made (t333 keeps its existing utilityError helper for its other callers).
- t-e2e-isolated-runner: this branch's instrumented case, now async, with main's Windows liveness rule in its probe loop: an exited sibling is judged by its native identity, because signal 0 can still succeed while a handle remains.
…now records

#1393 made the guard's stand-aside row record a blocked path with forward slashes and an upper-case drive, as the write-audit hook does, so the ledger reads the same on every platform. c9fcc48 had instead loosened t334 and the swarm continuation test to accept Windows backslashes, which now fails on Windows (full verification 36098150282, unit-5). Restore main's portable expectations.
Full verification 36098150282 caught one Windows race where a single process held the audit lock for the whole 300-second budget while four waited: only one BOLT_STARTED landed, and the retained lock and its fresh claim both named the same live owner. The children's output was discarded, so the cause could not be told apart. Keep each child's pid, exit code or signal, exit time, stdout, and stderr, and sample the temp directory's lock-related names every 50 ms without opening anything inside a lock directory (which is itself what Windows refuses a rename for). The record is written to the run's log directory and inlined in each race assertion. The race and its assertions are unchanged.
…get crash

The wait loop computed its deadline a few milliseconds after the operation deadline that every nested capture budgets against, so on a loaded runner a final capture could start after the operation deadline and throw TestBudgetExhaustedError, crashing the driver instead of reporting "timed out" (Windows integration load probe 36101895203, t-tui-bun-backend). Never let the loop outlive the operation deadline, and treat a capture that exhausted the wait's own budget, or failed after its deadline, as the wait timing out. A file-level exhaustion still propagates.
…erve

A file's cleanup reserve is a quarter of its deadline, so the one-second deadline left 250 ms to retire the hung worker. The runner requires observed retirement, and on a loaded Windows runner it could not confirm it in time and correctly reported ERROR instead of a timed-out FAIL (Windows integration load probe 36101889728). The file hangs, so a 20-second deadline exercises the same path with a 5-second reserve. The nested run's output now accompanies each report assertion.
Brings in #1398 and #1391. #1398 added core/tools/aidlc-inline-context.ts and this branch added core/tools/aidlc-runtime-budget.ts; each side raised the README tool inventory from 75 to 76, so the identical lines merged silently while the combined tree has 77 tools. t239 derives the count from core/tools and failed on the merge queue's commit. The README now says 77.

This branch had an error being deployed

1 failed deployment
ai-pr-review — 29b83927 Deployed Sep 25, 2026 by apackeer via Review pull request #891
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

aida:review-error AIDA could not produce a reliable review for the latest PR state documentation Improvements or additions to documentation github

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants