Skip to content

test(tui): wait for the run to end before a scenario submits - #118

Merged
yusiwen merged 1 commit into
masterfrom
fix/lsp-scenario-run-race
Oct 7, 2026
Merged

yusiwen merged 1 commit into
masterfrom
fix/lsp-scenario-run-race

Conversation

@yusiwen

@yusiwen yusiwen commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

What changed

  • tui/testdata/scenarios/lsp-tools.scenario and tui/testdata/scenarios/lsp-diagnostics.scenario: after the tool-call line, wait for the stub's final answer (stub-final-ok) and a short stable before sending /diagnostics.

Closes #117.

Why

The two LSP scenarios sent /diagnostics as soon as the tool-call line appeared, but that line renders before the tool runs. The TUI refuses a submit while a run is in flight (tui/update.go: m.status != StatusStreaming && !m.runIsActive()), so the Enter was dropped, the command stayed in the input box, and the assertion timed out on a frame showing the finished answer next to the unexecuted text. Waiting for the model's final answer — the signal tool-call.scenario and tool-multi.scenario already use — puts the submit after the run; stable 200ms absorbs the StreamDone that flips the input back on.

This is a scenario defect, not an LSP regression. Run #297 passed and run #298 failed on byte-identical trees:

$ git rev-parse cc1e728^{tree} 6e1a438^{tree}
d3fe4d3b77cca032728c011543d8645d5da5e45b
d3fe4d3b77cca032728c011543d8645d5da5e45b

lsp-diagnostics.scenario has the same shape (wait --text "read_file" then the command) and is exposed to the same race even though it did not lose it in #298; it is fixed here too.

Verification

  • Mechanism probe (pretend the run is slow). A stub whose second response sleeps 8s keeps the run active after the tool-call line is on screen:

    probe result
    /diagnostics sent on the tool-call line fails; final frame is stub-final-ok + ┃ /diagnostics in the input — byte-identical to run #298's failure frame
    same, after stub-final-ok + stable 200ms passes: saw LSP not available, 15 steps, rc=0
  • lsp-tools.scenario under the fix: 6/6 pass locally with gopls v0.23.0. The pre-fix content also passed 6/6 — the window is too narrow to hit by repetition here, which is why the probe above exists.

  • go test ./... -count=1 (redirected HOME/GOPATH/GIT_CONFIG_GLOBAL) → 13 packages ok / 0 failed / 0 ignored. The changed files are scenario data, not Go sources, so this suite does not exercise them; it is recorded for completeness.

  • CI's lsp job (make test-tui-lsp) is the gate; result below.

Honest scope

The TUI scenarios need a PTY. Under the default workspace-write sandbox, macOS Seatbelt denies the character devices (/dev/ptmx, /dev/zero, /dev/tty → EPERM; /dev/null is allowed), so the first attempts reported pty: start /bin/sh: operation not permitted; the earlier note that widening did not help was wrong — that retry had not taken effect. With full access the PTY works, and the runs above were made there. CI remains the only environment that runs them under the normal policy.

Notes

  • The refusal itself — a submit during a run — is intentional (it prevents two concurrent runs); only the scenario's assumption about when the run ends was wrong. The silent part of that refusal is filed as [tui] A submitted command is dropped silently while a run is active #119.
  • The annotation baseline is not zero (one ubuntu-latest notice per job). Reviewed with annotations_count per job.

CI

All 8 checks pass on 11f71df, including LSP integration (gopls + typescript) — the gate for this fix, which runs make test-tui-lsp — and TUI visual. Annotations match the documented baseline (1 per job; 2 for browser and tui-visual), so no new annotations.

The lsp-tools and lsp-diagnostics scenarios sent /diagnostics right after the
tool-call line, which renders before the tool runs. While a run is in flight the
TUI refuses a submit, so the Enter was dropped and the command stayed in the
input box; the assertion then timed out on a frame showing the answered run next
to the unexecuted command. Run #298 lost exactly that race while run #297 passed
on the identical tree (tree d3fe4d3), which is what makes this a scenario defect
and not an LSP regression.

Waiting for the stub's final answer — the signal tool-call.scenario and
tool-multi.scenario already use — puts the submit after the run, and a short
stable absorbs the StreamDone that flips the input back on.

The lsp scenarios need a PTY, which this machine's sandbox denies (/dev/ptmx is
"operation not permitted"), so the scenario run itself is verified by CI's lsp
job; locally only the parse was checked (both files reach the open step).

Refs #117
@yusiwen yusiwen added bug Something isn't working tests Test coverage and test infrastructure labels Oct 7, 2026
@yusiwen
yusiwen merged commit 9b4832d into master Oct 7, 2026
8 checks passed
@yusiwen
yusiwen deleted the fix/lsp-scenario-run-race branch October 7, 2026 07:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working tests Test coverage and test infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[tui] The lsp-tools scenario races the agent run and drops the /diagnostics submit

1 participant