fix(ci): preserve full-suite evidence and verify completion - #1369
Conversation
AIDA findings ledger
Open blocking findings (P0/P1): 2. Accepted and rejected findings never count toward the next action. Maintainer commands (repository write access) — put them on the first lines of a comment, one per line, several ids per line allowed: ledger.json{
"version": 4,
"pullRequest": 1369,
"nextId": 10,
"findings": [
{
"id": "F1",
"priority": "P1",
"category": "correctness",
"title": "Substring matching lets a bypassed command satisfy the Codex journey",
"anchors": [
{
"kind": "line",
"path": "tests/harness/codex-workspace-evidence.ts",
"side": "RIGHT",
"sha256": "3175ceaf00f1ff72870f1dfe6e502bd729345b4cacbcf04c3933bd0755a298c0"
}
],
"status": "open",
"firstSeen": {
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
"at": "2026-09-23T13:09:29.126Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
},
{
"id": "F2",
"priority": "P2",
"category": "workflow-state",
"title": "Progress-tolerant creation omits the workflow status invariant",
"anchors": [
{
"kind": "line",
"path": "tests/harness/codex-workspace-evidence.ts",
"side": "RIGHT",
"sha256": "65320124378afc5e354149ef31d7971cac2db3c3f503217edbded259fb38b217"
}
],
"status": "open",
"firstSeen": {
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
"at": "2026-09-23T13:09:29.126Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
},
{
"id": "F3",
"priority": "P3",
"category": "contracts",
"title": "The test README still documents trace retention as opt-in",
"anchors": [
{
"kind": "line",
"path": "docs/reference/09-testing.md",
"side": "RIGHT",
"sha256": "b578cffe1850d94c1bb9389fa27e396abfb52e7e461b3d1640c8c15ecc12c4bf"
}
],
"status": "resolved",
"firstSeen": {
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
"at": "2026-09-23T13:09:29.126Z"
},
"lastSeen": {
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"at": "2026-09-23T23:53:49.492Z"
}
},
{
"id": "F4",
"priority": "P3",
"category": "user-experience",
"title": "Production collection suppresses its recovery instruction",
"anchors": [
{
"kind": "line",
"path": ".github/scripts/prepare-live-runtime.ps1",
"side": "RIGHT",
"sha256": "8b4e82e0b43f1c15d0023eccb223f3a324cceab08b961553fe1806e294c53499"
}
],
"status": "open",
"firstSeen": {
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356",
"at": "2026-09-23T13:09:29.126Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
},
{
"id": "F5",
"priority": "P3",
"category": "contracts",
"title": "Supply-chain documentation still describes verification as live-only",
"anchors": [
{
"kind": "line",
"path": ".github/workflows/full-suite.yml",
"side": "RIGHT",
"sha256": "063a689273e67be6489ba7eff6888de8d5f95fd5d77e0bc53f1397387573c087"
},
{
"kind": "line",
"path": ".github/workflows/full-suite.yml",
"side": "RIGHT",
"sha256": "311aece3ce4ca5c124e4060d4c265d49a9b537580ee6c9b5735b93c214495f36"
}
],
"status": "open",
"firstSeen": {
"head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
"at": "2026-09-23T19:29:37.875Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
},
{
"id": "F6",
"priority": "P0",
"category": "security",
"title": "Full verification exposes credentials to PR-controlled code",
"anchors": [
{
"kind": "line",
"path": ".github/workflows/full-suite.yml",
"side": "RIGHT",
"sha256": "063a689273e67be6489ba7eff6888de8d5f95fd5d77e0bc53f1397387573c087"
},
{
"kind": "line",
"path": ".github/workflows/full-suite.yml",
"side": "RIGHT",
"sha256": "247cfadf662d36db893b1cd15c2504562bf296a6bb910afd837662754d56061c"
},
{
"kind": "line",
"path": ".github/workflows/full-suite.yml",
"side": "RIGHT",
"sha256": "e7d162d250978200908408bf975276c531eae3200cb7fc525da667b4473ab697"
}
],
"status": "open",
"firstSeen": {
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"at": "2026-09-23T23:53:49.492Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
},
{
"id": "F7",
"priority": "P3",
"category": "user-experience",
"title": "Manual deterministic dispatch defaults to an invalid selection",
"anchors": [
{
"kind": "line",
"path": ".github/workflows/deterministic-tests.yml",
"side": "RIGHT",
"sha256": "d50a8f57f8baa0334395bbdbafbc9ca88986c96b9644b7ee03cdd8ead702d865"
},
{
"kind": "line",
"path": ".github/workflows/deterministic-tests.yml",
"side": "RIGHT",
"sha256": "d5b9b63bf7234c3ce987cd7caafc84c34299b6f811183b28e9382171b5436dd6"
}
],
"status": "open",
"firstSeen": {
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"at": "2026-09-23T23:53:49.492Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
},
{
"id": "F8",
"priority": "P2",
"category": "workflow-state",
"title": "Failed SDK fixtures are retained outside uploaded evidence",
"anchors": [
{
"kind": "line",
"path": "tests/integration/t183-codekb-placement-reverify.sdk.test.ts",
"side": "RIGHT",
"sha256": "a10105e36d38bf077bb49cfa8c7cbb74e38a833e502734de8e3648339b2992a6"
},
{
"kind": "line",
"path": "tests/integration/t193-compose-report-journey.sdk.test.ts",
"side": "RIGHT",
"sha256": "699c9add7ffd5e3dbf68a818c578d800f0e1d059a986f05afbda2d63a2fec376"
}
],
"status": "open",
"firstSeen": {
"head": "dc201a91743e593c75d720c24007934894b79661",
"at": "2026-09-24T02:14:21.134Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
},
{
"id": "F9",
"priority": "P2",
"category": "correctness",
"title": "Test-only timeout hides the shipped linter budget failure",
"anchors": [
{
"kind": "line",
"path": "tests/integration/t92.test.ts",
"side": "RIGHT",
"sha256": "eae13c11ca2c0cdb9005cc9731b355a351469576418da1e6ca24e264a609232c"
},
{
"kind": "line",
"path": "tests/integration/t92.test.ts",
"side": "RIGHT",
"sha256": "f1ca2c5d0c83a736cdd132d1fc0abf39ff8c1a83d53466bfc9a769091d8823bd"
}
],
"status": "open",
"firstSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
},
"lastSeen": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"at": "2026-09-24T03:11:53.553Z"
}
}
],
"events": [
{
"at": "2026-09-23T13:09:29.126Z",
"kind": "opened",
"by": "aida",
"id": "F1",
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
},
{
"at": "2026-09-23T13:09:29.126Z",
"kind": "opened",
"by": "aida",
"id": "F2",
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
},
{
"at": "2026-09-23T13:09:29.126Z",
"kind": "opened",
"by": "aida",
"id": "F3",
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
},
{
"at": "2026-09-23T13:09:29.126Z",
"kind": "opened",
"by": "aida",
"id": "F4",
"head": "94f05c1298ff1e92405d6f1d7db1d8e35c401356"
},
{
"at": "2026-09-23T19:29:37.875Z",
"kind": "seen",
"by": "aida",
"id": "F1",
"head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T19:29:37.875Z",
"kind": "seen",
"by": "aida",
"id": "F2",
"head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T19:29:37.875Z",
"kind": "seen",
"by": "aida",
"id": "F3",
"head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T19:29:37.875Z",
"kind": "seen",
"by": "aida",
"id": "F4",
"head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T19:29:37.875Z",
"kind": "opened",
"by": "aida",
"id": "F5",
"head": "60f2833e249612ce1ea0db4c55596d37f8dff5e9"
},
{
"at": "2026-09-23T23:53:49.492Z",
"kind": "seen",
"by": "aida",
"id": "F1",
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T23:53:49.492Z",
"kind": "seen",
"by": "aida",
"id": "F2",
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T23:53:49.492Z",
"kind": "resolved",
"by": "aida",
"id": "F3",
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"reason": "declared corrected by the judge; cited files changed since the last review"
},
{
"at": "2026-09-23T23:53:49.492Z",
"kind": "seen",
"by": "aida",
"id": "F4",
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T23:53:49.492Z",
"kind": "seen",
"by": "aida",
"id": "F5",
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-23T23:53:49.492Z",
"kind": "opened",
"by": "aida",
"id": "F6",
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5"
},
{
"at": "2026-09-23T23:53:49.492Z",
"kind": "opened",
"by": "aida",
"id": "F7",
"head": "8f302f15ad4271486954db5b55dcc6bd7b4668c5"
},
{
"at": "2026-09-24T02:14:21.134Z",
"kind": "seen",
"by": "aida",
"id": "F6",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:14:21.134Z",
"kind": "seen",
"by": "aida",
"id": "F1",
"head": "dc201a91743e593c75d720c24007934894b79661",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T02:14:21.134Z",
"kind": "seen",
"by": "aida",
"id": "F2",
"head": "dc201a91743e593c75d720c24007934894b79661",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T02:14:21.134Z",
"kind": "seen",
"by": "aida",
"id": "F4",
"head": "dc201a91743e593c75d720c24007934894b79661",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T02:14:21.134Z",
"kind": "seen",
"by": "aida",
"id": "F5",
"head": "dc201a91743e593c75d720c24007934894b79661",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T02:14:21.134Z",
"kind": "seen",
"by": "aida",
"id": "F7",
"head": "dc201a91743e593c75d720c24007934894b79661",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T02:14:21.134Z",
"kind": "opened",
"by": "aida",
"id": "F8",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:40:42.380Z",
"kind": "seen",
"by": "aida",
"id": "F6",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:40:42.380Z",
"kind": "seen",
"by": "aida",
"id": "F1",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:40:42.380Z",
"kind": "seen",
"by": "aida",
"id": "F2",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:40:42.380Z",
"kind": "seen",
"by": "aida",
"id": "F8",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:40:42.380Z",
"kind": "seen",
"by": "aida",
"id": "F4",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:40:42.380Z",
"kind": "seen",
"by": "aida",
"id": "F5",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T02:40:42.380Z",
"kind": "seen",
"by": "aida",
"id": "F7",
"head": "dc201a91743e593c75d720c24007934894b79661"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "seen",
"by": "aida",
"id": "F6",
"head": "79f328fe743afc23357283893d7bbde866fd91b1"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "seen",
"by": "aida",
"id": "F1",
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "seen",
"by": "aida",
"id": "F2",
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "seen",
"by": "aida",
"id": "F4",
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "seen",
"by": "aida",
"id": "F5",
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "seen",
"by": "aida",
"id": "F7",
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "seen",
"by": "aida",
"id": "F8",
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"reason": "retained: still open per the judge, not restated"
},
{
"at": "2026-09-24T03:11:53.553Z",
"kind": "opened",
"by": "aida",
"id": "F9",
"head": "79f328fe743afc23357283893d7bbde866fd91b1"
}
],
"review": {
"head": "79f328fe743afc23357283893d7bbde866fd91b1",
"readiness": 1,
"risk": 5,
"decision": "change"
}
} |
There was a problem hiding this comment.
Reviewed 94f05c1298ff1e92405d6f1d7db1d8e35c401356 against 79ff00818620014ffcd3749414e2de9c7e62157e and current repository behavior.
Inspection: 12 changed files. Scope: full head (first review of this pull request).
Final Assessment
Human decision aid only: Readiness 5/5 is best; Risk 1/5 is best. These scores inform the maintainer; the next action below follows finding severity (any open P0/P1 → author/change) and does not approve or merge the PR.
Readiness: 2/5 — The evidence-retention changes are substantial, but the new Codex verifier can produce a false green for a release-gating journey because it does not bind success to an actual allowed utility invocation. Additional workflow-state and documentation corrections remain.
Risk: 4/5 — The blocking flaw affects required hosted live evidence used for stable publication. A broken or bypassed Codex command path could be recorded as successful, although the impact is limited to validation integrity rather than production data mutation.
Decision required: Author — make changes before this PR proceeds. The P1 command-evidence flaw must be corrected before this release-gating verification can be trusted.
Validation performed:
- Inspected all 12 changed-file snapshots, the full diff, PR metadata, review scope, discussion, findings ledger, and specialist outputs.
- Compared the Codex evidence helper with its callers, regression tests, existing command validator, state-machine contract, and required Full Suite release path.
- Compared Windows collection changes with their production caller, native fixtures, tests, workflow upload steps, sanitizer behavior, and testing documentation.
- The ledger contains no open, accepted, rejected, or archived findings requiring disposition.
Findings: 1 blocking, 3 advisory.
Ledger: 4 open, 0 retained blocking, 0 accepted, 0 suppressed as rejected by a maintainer. Maintainers act on findings with /aida commands in the ledger comment.
Correctness & Reliability
P1 [F1]: Substring matching lets a bypassed command satisfy the Codex journey
Evidence: tests/harness/codex-workspace-evidence.ts:33.
Problem: A completed command is selected solely because its agent-controlled shell text contains the tool path and matches the route regex. A compound command, shell comment, or decoy invocation can include those strings while another command directly constructs the expected registry, state, include, and audit files. Because the model has workspace-write access and the resulting command event can exit zero, turn completion and every disk assertion can pass without the required utility path succeeding.
Impact: The required Codex Full Suite journey can produce false release evidence for a broken or bypassed workflow, allowing stable publication to rely on an invalid supported-harness verification.
Required correction: Parse and validate one exact allowed utility invocation, rejecting shell composition, substitutions, redirections, comments, and unrelated commands. Add adversarial tests showing that valid-looking disk state cannot rescue compound or decoy commands.
Workflow, State & Recovery
P2 [F2]: Progress-tolerant creation omits the workflow status invariant
Evidence: tests/harness/codex-workspace-evidence.ts:112.
Problem: When stopAfterCreate is false, the helper skips every assertion inside this branch, including Status: Running. The first skill-driven journey can therefore accept a newly registered in-flight intent whose state is missing Status or incorrectly says Completed or Archived, despite the canonical workflow machine requiring a newly created intent to remain Running.
Impact: The live Codex journey can report successful creation while persisting contradictory or unusable lifecycle state, weakening the release evidence even after command validation is corrected.
Required correction: Always require workflow Status to be Running for a newly created in-flight intent. Keep only the exact handoff stage and last-completed-stage assertions conditional, and add negative coverage for invalid status when stopAfterCreate is false.
User Experience
User experience change: CI maintainers diagnosing Full Suite failures receive retained sanitized traces by default and independently collected Windows launch logs, while completed Codex turns may now be accepted without command output when disk evidence exists.
Before: Eligible traces were dropped unless explicitly enabled, Windows test-log collection failure could prevent launch-log preservation, and Codex verification required completion text.
After: Eligible traces are retained unless explicitly disabled, Windows collection publishes complete sources independently with a report, and Codex verification uses command events plus workspace state.
Example: Set AIDLC_NIGHTLY_UPLOAD_TRACES=0 to opt out; failed Windows collection retains windows-collection-<uuid>.json and any independently valid launch logs.
Assessment: Failure diagnosis improves, but permissive Codex command matching can misleadingly present a bypassed journey as successful, and the production collection error hides its new recovery instruction.
P3 [F4]: Production collection suppresses its actionable recovery message
Evidence: .github/scripts/prepare-live-runtime.ps1:397.
Problem: Collect-RuntimeLogs throws a fixed safe instruction to inspect the collection report and retained independent logs, but the unchanged top-level catch replaces it with a generic stage, exception type, and line-number message. The new native tests call the function directly and therefore do not cover what workflow operators actually see.
Impact: Windows CI maintainers are not told where the newly retained recovery evidence is located when collection fails.
Required correction: Surface a fixed collection-specific recovery message from the production catch without echoing arbitrary exception text, and add a caller-level test for the emitted failure output.
Contracts & Compatibility
P3 [F3]: The test README still documents trace retention as opt-in
Evidence: docs/reference/09-testing.md:1532.
Problem: The workflows and reference guide now make retention the default with AIDLC_NIGHTLY_UPLOAD_TRACES=0 as the opt-out, but tests/README.md still says traces are dropped by default and that setting the variable to 1 opts in.
Impact: Contributors receive contradictory guidance about uploaded artifact contents and residual disclosure risk.
Required correction: Update tests/README.md to document default retention and the value 0 as the repository-level opt-out.
Residual risk: Repository code and tests were not executed, as required by the read-only review contract. Windows collection behavior was assessed statically rather than on a Windows host.
Reviewed by AIDA (AI-DLC Developer Agent).
[AI-PR-REVIEWED] 94f05c1
There was a problem hiding this comment.
Reviewed 60f2833e249612ce1ea0db4c55596d37f8dff5e9 against 79ff00818620014ffcd3749414e2de9c7e62157e and current repository behavior.
Inspection: 15 changed files. Scope: incremental — 6 files with lines changed since the review at 94f05c12; the security lenses reviewed the full head. Findings on lines unchanged since that review are deferred, not decisive.
Final Assessment
Human decision aid only: Readiness 5/5 is best; Risk 1/5 is best. These scores inform the maintainer; the next action below follows finding severity (any open P0/P1 → author/change) and does not approve or merge the PR.
Readiness: 2/5 — The full-verification and diagnostic-retention work is substantial, but the required Codex journey can still produce false release evidence because command success is established by substring matching. Additional state, recovery, and documentation defects remain.
Risk: 4/5 — The P1 flaw affects required hosted Codex evidence consumed by the stable-release workflow. A bypassed utility invocation can appear successful, although the direct impact is validation integrity rather than production data loss.
Decision required: Author — make changes before this PR proceeds. Re-derived from the ledger: 1 open blocking finding (F1) was not restated this run and the cited code is unchanged, so the author still needs to act. Judge's note, superseded by finding severity: Re-derived: 4 findings outside the incremental review scope were deferred and no blocking finding remains. Judge's note, superseded by finding severity: The open P1 Codex command-evidence flaw must be corrected before release-gating verification can be trusted.
Validation performed:
- Inspected all 15 changed-file snapshots, the complete diff, incremental scope, PR metadata, discussion, ledger, and specialist outputs.
- Compared workflow changes with result reduction, release-consumer checks, authorization tests, Codex evidence contracts, Windows collection callers, and testing documentation.
- Confirmed all four open ledger defects remain present; no accepted or rejected decisions apply.
Findings: 0 blocking, 1 advisory.
Ledger: 5 open, 1 retained blocking, 3 retained advisory, 0 accepted, 0 suppressed as rejected by a maintainer. The next decision was re-derived from the ledger. Maintainers act on findings with /aida commands in the ledger comment.
Contracts & Compatibility
P3 [F5]: Supply-chain references still describe candidate verification as live-only
Evidence: .github/workflows/full-suite.yml:21, .github/workflows/full-suite.yml:698.
Problem: The new mode runs every declared job and emits full-suite-verification-result, but docs/reference/19-supply-chain-security.md still says candidate verification intentionally omits native, deterministic, and production-guard jobs. The artifact inventory in docs/reference/09-testing.md also omits the purpose-specific result artifacts.
Impact: Maintainers receive an inaccurate account of privileged candidate execution and may look for the wrong evidence artifact.
Required correction: Document both verification modes in the supply-chain reference, including authorization, coverage, filters, artifact names, purposes, and release ineligibility. Update the artifact inventory accordingly.
User Experience
User experience change: Maintainers can run the complete Full Suite against an unmerged PR and receive more retained diagnostic evidence from failed CI runs.
Before: Candidate verification covered live jobs only, eligible traces were opt-in, and Windows collection failures could lose independent launch evidence.
After: full_verification=true runs every declared job with separate non-release evidence; traces are retained by default and Windows sources are collected independently.
Example: gh workflow run full-suite.yml --ref '<candidate-branch>' -f 'ref=<exact-head-sha>' -f full_verification=true
Assessment: The workflow is more diagnosable and gives maintainers broader pre-merge validation, but stale guidance and suppressed recovery text make the new behavior harder to use safely.
Retained advisory findings
Open P2/P3 findings from the ledger that the judge marked still open without restating them. They do not affect the decision.
P2 [F2]: Progress-tolerant creation omits the workflow status invariant — first reported at 94f05c12.
P3 [F3]: The test README still documents trace retention as opt-in — first reported at 94f05c12.
P3 [F4]: Production collection suppresses its actionable recovery message — first reported at 94f05c12.
Retained blocking findings
Open P0/P1 findings from the ledger that this review did not restate and whose cited code is unchanged. They keep the next action with the author until the code changes or a maintainer accepts them.
P1 [F1]: Substring matching lets a bypassed command satisfy the Codex journey — first reported at 94f05c12; cited: tests/harness/codex-workspace-evidence.ts.
Deferred (outside the review scope)
Reported by the judge but citing only lines unchanged since the review at 94f05c12. They do not affect the decision; comment /aida full to have the next review cover the whole head.
P1 · correctness: Substring matching lets a bypassed command satisfy the Codex journey — tests/harness/codex-workspace-evidence.ts
P2 · workflow-state: Progress-tolerant creation omits the workflow status invariant — tests/harness/codex-workspace-evidence.ts
P3 · contracts: The test README still documents trace retention as opt-in — docs/reference/09-testing.md
P3 · user-experience: Production collection suppresses its actionable recovery message — .github/scripts/prepare-live-runtime.ps1
Residual risk: Repository code and tests were not executed under the read-only review contract. Windows collection behavior was assessed statically rather than on a Windows host.
Reviewed by AIDA (AI-DLC Developer Agent).
[AI-PR-REVIEWED] 60f2833
Superseded by AI review of 60f2833
When a file's work tail is shorter than native startup, the ready deadline is the work cutoff. The loop read the clock twice per pass, so a pass that crossed the cutoff between the two reads skipped the work expiry and reported "e2e supervisor did not become ready" instead (t-bun-test-deadlines on Windows, repeat Full Suite). Use one reading per pass.
Three live failures on this PR were a driver waiting for something that could no longer happen: t139 on an unrecognized menu (25 minutes), t29 after the agent had already finished (33 minutes), and the earlier t126 deadlock. Each ended only at the file ceiling, so the report said "exceeded file ceiling" instead of what went wrong. The TUI driver now recognizes a finished turn by the screen: work was seen (status spinner or its elapsed timer, a background-agent wait, a running subagent row, or a running command's background hint) and Claude's empty prompt then stayed unchanged for 30 seconds. Across 517 recorded sessions the longest idle-looking pause inside a turn was 0.4 seconds. An unrecognized screen counts as working, so the backstop still applies. - wait fails with the last pane when its pattern never painted (--through-turn-end opts out). - answer-gate fails when no menu appeared and its terminator is unmet, and an overall-budget expiry is reported as such rather than as a per-gate timeout. - Revision recovery types free-text feedback only at an idle prompt; the quiet timer it replaces could type while a structured menu was still painting. - The t27 audit wait and t29 scope wait use the same observation.
- Native cleanup no longer reports "native kill exited 0" when the kill succeeded but the cleanup deadline passed before retirement could be confirmed. - Native daemon startup and target launch surface a supervisor error published in the last poll interval instead of a timeout. - An endpoint seen vacant only after the deadline is reported as unconfirmed, not as still active.
auditFilePathFor returns the first .md in the audit folder. t138 on Windows failed with zero STAGE_STARTED events because the agent had appended its own notes to a second file there, and that file sorted first. Audit assertions now read every shard, as the engine does.
The capture fixture replaced its log with a directory and then printed its marker. When Bun's lazily flushed file header reached the pipe as a separate chunk after the swap, that chunk failed capture first and the runner stopped the child before the marker was printed. Print a marker, wait until the runner has written it, then break capture.
…al session Live journeys normally take 24 to 29 minutes, against about 35 minutes of usable work in the 2,400-second file ceiling, so one detour ended a journey that was still progressing (t51 on macOS). With waits now ending on the observed end of a turn, a stuck journey fails within about 30 seconds, so a generous ceiling only serves progressing work. Live files and runs get 3,600 seconds. The one-hour Bedrock role session is assumed about 13 seconds before the run step and model work stops at the five-minute cleanup reserve, so it always ends while its credentials are valid; the constant documents that bound. Live steps become 70 minutes, live jobs 80 minutes, and the Windows isolated run task 64 minutes.
The target fixture waits for finish to exist and then reads its exit code; the test wrote it in place, so a read between creation and write got an empty string and exited 0 (t-tui-bun-lifecycle on Windows, competing exit and stop). leaf.json had the same shape, and its reader treats a partial file as a failure. Write each to a sibling temp file and rename it into place.
The case waited for the engine's error text verbatim, but the conductor relays engine errors in its own words more often than not (11 of 14 across Full Suite traces), accurately and with a remedy, so it failed two runs in three. Assert the error where it is deterministic, as the engine's error result in the native transcript, and require the screen to name AWS_AIDLC_DEFAULT_SCOPE and its value; the no-state assertion is unchanged.
… Codex evidence Review follow-up for AIDA F6 (P0), F1 (P1), F2, F5 and F7. - F6: full_verification may select an unmerged head, so it no longer runs live_prepare, live_hosted or live_windows, the jobs that receive OIDC and the AWS role; live coverage of a candidate keeps the separately authorized live_verification mode. The result reducer requires exactly those three to be skipped. Every Full Suite checkout now sets persist-credentials: false. A contract test evaluates every credentialed job's condition under full-verification and pins persist-credentials on every checkout. - F1: the Codex journey evidence parses one exact `bun .codex/tools/aidlc.ts ...` invocation inside Codex's own shell wrapper and matches the route against its argv. A command that names the tool and route but composes, redirects, substitutes, comments, or decoys fails outright; adversarial tests show valid disk state cannot rescue it. - F2: a newly created in-flight intent must be Running even when the beat may progress past bootstrap. - F5: the supply-chain reference documents both verification modes, and the artifact inventory names each purpose-specific result. - F7: manual deterministic dispatch defaults to the executable smoke tier. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…y evidence The exact-invocation parser unwrapped only Codex's POSIX `<shell> -lc` form. On Windows, Codex reports every command as `"C:\\Program Files\\PowerShell\\7\\pwsh.exe" -Command '...'`, so all 17 utility commands in Full Suite 36043454212's Windows workspace journey were rejected, and any that carries a checked route fails the journey. Recognize that wrapper as well; the inner command keeps the same plain-word rules, which already reject PowerShell composition, substitution and redirection. The unit cases use the observed Windows shape and add PowerShell decoys and a foreign executable.
The target fixture wrote target.json in place while the test polled it, and the reader deliberately treats a partial file as a failure, so full verification 36058777252 read an empty file ("Unexpected EOF") in the detached-descendant case. Publish it by rename, as leaf.json and finish already are. The other cross-process files are safe: status.json is published whole by the supervisor, and release, stop, and ctrl-c are only tested for existence.
…evidence The model chooses per call whether Codex runs a login shell: Codex's shell tool takes a `login` parameter, and a non-login call runs `<shell> -c` or `pwsh.exe -NoProfile -Command`. The parser accepted only the login forms, so live verification 36061390100's Windows workspace journey failed on `pwsh.exe -NoProfile -Command 'bun .codex/tools/aidlc.ts engine intent create ...'` after the intent was created correctly. Accept both forms; the inner command rules are unchanged. Of 106 utility commands across 36 recent Codex artifacts, the only one still rejected is a genuine compound printf.
The persistent non-root CIM case injects a process-kill failure that can never clear, so its first kill must refuse. Since the shared cleanup backstop rose from 30 seconds to 15 minutes, that refusal spent the full 15 minutes, and the retry's one-shot liveness check then found the session's recorded PID alive (full verification 36066982685). A PID that long exited is free for Windows to reuse. Bound only the refused kill to its former 30 seconds; the retry keeps the full backstop.
Brings in #1379 (Plan Approval auto-resolves --session), #1239 (sensor path argument from its input schema), #1389 (cross-OS CI jobs in the merge queue), #1390 (Full Suite supersedes runs across branch heads) and #1163. Conflict resolutions: - aidlc-log.ts: main's resolvePlanApprovalSession at every site; its refusal when no session resolves keeps this branch's AIDLC Runtime Session hint. The unknown-session warning now applies only to a named --session, since an auto-resolved one came from the invoking conversation. - aidlc-sensor.ts and t92: both sides' imports and parameters. - 06-hooks-and-tools.md: main's session sentence plus this branch's warning sentence, reworded for auto-resolution. - 09-testing.md: main's pull-request and merge-queue rows plus the manual deterministic dispatch row; the budget hierarchy names merge-queue native-terminal checks. - Coverage registry regenerated. - t265 blanks the hook-injected session override for its omitted-session case, as t328 does.
…tions The registry was regenerated during the merge before the projections were rebuilt, so it enumerated the pre-merge functions and missed those #1163 added (1,134 units instead of 1,140).
Every captured test step starts run-tests.sh in the background with its output redirected to run.log, then immediately polls that file with sed under `set -euo pipefail`. The background job creates the file only when it runs its redirection, so a reader scheduled first exits 2 and ends the step. t345 executes this script and failed that way twice during a full local -P 8 run ("sed: can't read .../tmp/ci-deterministic/run.log"), while passing alone; a loaded CI runner can lose the same race. Create the file after its directory in all four steps (ci.yml native and node-pty, full-suite.yml native, deterministic-tests.yml). A 300-iteration CPU-load stress did not reproduce the failure either way, so the fix rests on the mechanism, not a reproduction.
…t on Windows
The custom core.excludesFile case wrote the fixture's native path into a git config file. Git reads backslashes in a config value as escapes, so on Windows the line is invalid ("fatal: bad config line 2", exit 128) and the doctor correctly warns that it could not evaluate the source instead of reporting the rule, which the case expected (full verification 36094162239, Windows unit-5). Write the path with forward slashes, which git accepts on every OS. The product check is unchanged.
…get is on screen The runner fixture's inner test started its native session and published the started witness at once, without waiting for the target to draw, unlike the file's own start helper. In cancel mode the parent cancels as soon as the witness exists, so on Windows the runner archived the session before the target printed and the archived screen was empty (full verification 36094162239, Windows integration). Wait for NATIVE READY before publishing the witness; the archive itself records the screen correctly.
Brings in #1393 (Windows cross-OS test repairs and the plan-approval guard's portable blocked path), #1272 and #1266. Conflict resolutions: - t340, t333, t-guard-plan-continuation-swarm: main's versions, which carry the same Windows fixes this branch made (t333 keeps its existing utilityError helper for its other callers). - t-e2e-isolated-runner: this branch's instrumented case, now async, with main's Windows liveness rule in its probe loop: an exited sibling is judged by its native identity, because signal 0 can still succeed while a handle remains.
…now records #1393 made the guard's stand-aside row record a blocked path with forward slashes and an upper-case drive, as the write-audit hook does, so the ledger reads the same on every platform. c9fcc48 had instead loosened t334 and the swarm continuation test to accept Windows backslashes, which now fails on Windows (full verification 36098150282, unit-5). Restore main's portable expectations.
Full verification 36098150282 caught one Windows race where a single process held the audit lock for the whole 300-second budget while four waited: only one BOLT_STARTED landed, and the retained lock and its fresh claim both named the same live owner. The children's output was discarded, so the cause could not be told apart. Keep each child's pid, exit code or signal, exit time, stdout, and stderr, and sample the temp directory's lock-related names every 50 ms without opening anything inside a lock directory (which is itself what Windows refuses a rename for). The record is written to the run's log directory and inlined in each race assertion. The race and its assertions are unchanged.
…get crash The wait loop computed its deadline a few milliseconds after the operation deadline that every nested capture budgets against, so on a loaded runner a final capture could start after the operation deadline and throw TestBudgetExhaustedError, crashing the driver instead of reporting "timed out" (Windows integration load probe 36101895203, t-tui-bun-backend). Never let the loop outlive the operation deadline, and treat a capture that exhausted the wait's own budget, or failed after its deadline, as the wait timing out. A file-level exhaustion still propagates.
…erve A file's cleanup reserve is a quarter of its deadline, so the one-second deadline left 250 ms to retire the hung worker. The runner requires observed retirement, and on a loaded Windows runner it could not confirm it in time and correctly reported ERROR instead of a timed-out FAIL (Windows integration load probe 36101889728). The file hangs, so a 20-second deadline exercises the same path with a 5-second reserve. The nested run's output now accompanies each report assertion.
Brings in #1398 and #1391. #1398 added core/tools/aidlc-inline-context.ts and this branch added core/tools/aidlc-runtime-budget.ts; each side raised the README tool inventory from 75 to 76, so the identical lines merged silently while the combined tree has 77 tools. t239 derives the count from core/tools and failed on the merge queue's commit. The README now says 77.
Summary
Full Suite failures exposed timeouts below the cost of valid work on loaded runners, incomplete failure evidence, and incorrect workflow/render expectations. This change uses shared execution backstops throughout the test runner, native drivers, fixture subprocesses, tools, sensors, and supported hook registrations. Commands still finish immediately when their work completes; explicit timeout calibrations and ownership checks remain enforced.
The repair also checks completed workflow state, publishes cancellation observations atomically, retains sanitized failed journeys, and supplies the concrete defect requested by the requirements-analysis fixture. Full verification runs every credential-free job against an immutable PR head; live coverage of a candidate uses the existing live verification mode. Both remain distinct from release-purpose qualification.
Changes
--production-guards, so the fix: make the guards a fence for the agents and a gate the human holds the key to #1262 live chat-lowering proof executes instead of skipping. The native terminal driver treats an endpoint closed by its own published stop as success only after the same generation retires.--guard-policy,--change-control,--sensors,--learnings,--summary-confirmation) now ends the turn as the engine instructs. The Stop hook recognized only depth, test-strategy, and review as terminal configuration, so it told the agent to resume a workflow the human had only asked to reconfigure./dev/null,NUL) is no longer treated as a workspace write, so read-only probes stay available. Real writes are still refused.--sessionthe project has not seen still records, but its output warns that the human's answer cannot bind and names theAIDLC Runtime Session:line. With feat(plan-approval): auto-resolve --session from the invoking conversation #1379's auto-resolution, an omitted--sessionresolves from the invoking conversation and gets no warning; the refusal when nothing resolves and the unprompted-receipt refusal carry the same pointer. This removes the double approval prompt seen live, without adding a refusal.wait,answer-gate, and revision recovery then fail or act at once with the last pane, instead of hanging until the file ceiling reports only a timeout. Unrecognized screens keep the backstop.run.logbefore starting the background runner, so the log poll cannot race the redirect under load.live_prepare,live_hosted,live_windows), and every Full Suite checkout setspersist-credentials: false; live coverage of a candidate useslive_verification. Manual deterministic dispatch defaults to the smoke tier.bun .codex/tools/aidlc.tsinvocation inside Codex's own shell wrapper, including its Windows PowerShell wrapper, and a newly created in-flight intent must be Running.Validation
ce548d9b.0e16bdf7merges main again (fix(windows): repair the cross-OS test failures on Windows #1393, fix(doctor): compare the audit against the per-stage checkboxes #1272, fix(sensor): drop a superseded detail file when the sensor passes #1266);91bee555merged main (feat(plan-approval): auto-resolve --session from the invoking conversation #1379, fix(sensor): route a sensor's path argument from its declared input_schema #1239, chore(ci): run the cross-OS jobs in the merge queue, not on every PR push #1389, chore(ci): supersede Full Suite verification across branch heads #1390, fix: never record an unreadable review findings table as no findings #1163). Release metadata matches main. Conflicts were resolved in the Plan Approval session path (main's auto-resolution kept, with this branch's hint on its refusal), the sensor import and t92 parameters, two reference docs, and the regenerated coverage registry.91bee555passed: resultpassed: truefor all families, 196 of 196 jobs; it is the live evidence for Claude SDK, Codex, opencode, and the release contract. This branch's later changes are deterministic tests that no live job runs; main's fix(windows): repair the cross-OS test failures on Windows #1393, fix(doctor): compare the audit against the per-stage checkboxes #1272, and fix(sensor): drop a superseded detail file when the sensor passes #1266, merged after it, are not covered by this live run.ce548d9bpassed: resultpassed: true, with the three live jobs omitted by design. Claude TUI live verification once548d9bpassed: resultpassed: true, 62 of 62 live jobs.ce548d9bchanges the native terminal driver's wait deadline, which the live TUI journeys use, so the TUI family was re-verified on this head. Two more Windows integration-tier runs under load passed all 131 files.43e93253passed: resultpassed: true, with the three live jobs omitted by design. Three further Windows integration-tier runs under load found two more flakes from this branch's tighter timing, both fixed ince548d9b: a native terminal wait could outlive the operation deadline its captures budget against and crash with a budget error instead of reporting "timed out" (9405b6a4), and the outer file deadline case gave cleanup 250 ms, too little to confirm worker retirement on a loaded Windows runner (ce548d9b). Both runs are PR-verification evidence and are not consumed by preview or stable publication.0e16bdf7failed on Windows in three places. t334 and the swarm continuation test still expected the Windows backslashes thatc9fcc485had loosened them to accept; fix(windows): repair the cross-OS test failures on Windows #1393 fixed the guard to record portable paths, soa30d72f9restores main's expectations. t46 caught one race in which a single process held the audit lock for the whole 300-second budget while four waited (only one BOLT_STARTED landed; the retained lock and a fresh claim both named that live owner). This branch's release retry loop (0e932d02) is the prime suspect; the cause is not yet known, so43e93253records each child's outcome and the lock names over the race, and repeated Windows runs gathered evidence before any lock change: 60 further races (30 isolated, 30 under full integration-tier load) all completed normally in at most 3.4 seconds, so the stall has not recurred and its cause is still unknown. The diagnostics stay in place so a recurrence names it.91bee555failed two Windows test defects, both fixed. t340 (from fix(doctor): report Kiro IDE ignore sources that hide .kiro/ #1161, merged today) wrote a native Windows path into a git config file, which git rejects as a bad config line, so the doctor correctly reported that it could not evaluate the source;de26b715writes the path with forward slashes. The native cancellation fixture published its started witness before the target drew, so a prompt cancel archived an empty screen;df9aa671waits for the target first.bun run checkand the coverage registry pass. The full local unit and integration run on the merge (487 files, 11,198 executed cases) failed only the registry freshness check, regenerated in2ad636ea, and two t345 cases that lost the captured-log race fixed in91bee555; both files pass on rerun.728504b6had complete evidence: full verification passed, Codex live passed one1328bba, and Claude SDK, Claude TUI, opencode, and release-contract live passed onc04fafe6.e1328bbawas cancelled after its legacy Windows node-pty job failed: a kill that the persistent non-root CIM case injects to fail spent the whole 15-minute cleanup backstop (30 seconds before this change raised the backstop), and the retry's one-shot liveness check then saw the session's recorded PID alive.728504b6bounds only that deliberately refused kill to its former 30 seconds.c04fafe6, full verification passed (resultpassed: true, with the three live jobs omitted by design). Its live run had two Windows failures. The Codex workspace journey rejected Codex's non-login wrapper (pwsh.exe -NoProfile -Command; the model chooses a login or non-login shell per call, and a non-login POSIX call is<shell> -c), fixed ine1328bba. At-tui-custom-harnessSDK case passed but its project folder stayed busy for the whole cleanup window: on Windows the SDK's abort kills only the CLI, about 7 seconds later, and the CLI had started a Bash tool just before. That harness race is not a core defect; the job was rerun and passed (5 of 5 cases in 254 seconds), and containing the SDK CLI in a Windows job object is a follow-up.d9bf2ea3was cancelled early, the live run before any live job started: full verification's Windows native lifecycle test readtarget.jsonwhile its fixture was still writing it ("Unexpected EOF").c04fafe6publishes that record by rename, like the fixture's other handshake files.a3ca141bwas cancelled after 152 passing jobs and no failures, when4e7fd87dmoved the head and removed credentialed jobs from full verification. Its Windows Codex evidence showed the new exact-invocation parser rejecting all 17 of Codex's PowerShell-wrapped commands, which would fail the Windows workspace journey;d9bf2ea3accepts that wrapper, and the 55 recorded Linux, macOS, and Windows commands all parse.2e703c94was cancelled after 119 passing jobs to pick up two fixes for its only failures: a Windows lifecycle fixture read its exit-code file before the test finished writing it, and t29 asserted the engine error's wording although the conductor relays it in its own words (it now asserts the engine error in the native transcript and its facts on screen). The new turn-end wait ended that t29 case 30 seconds after the turn, not at the 40-minute ceiling.d6320483(236 of 236 jobs). A repeat on the same head passed 233 jobs and failed two, both fixed here: the supervisor ready/work deadline race (Windows unit tests) and the single-shard audit read (t138 on Windows).The two current-head runs are the merge-readiness evidence; focused local passes do not replace them. Full Suite retains its declared provider exclusions.
Checklist
Acknowledgment
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of the project license.