Repository navigation
V2: preserve permission and question waits across server restart #36347
Description
Activity
- addedbugSomething isn't workingSomething isn't workingcoreAnything pertaining to core functionality of the application (opencode server stuff)Anything pertaining to core functionality of the application (opencode server stuff)
on Jul 11, 2026 opencode-agent commented
on Jul 11, 2026 on Jul 11, 2026 – with OpenCode AgentContributorMore actionsFollow-up: permission preflight and persisted forms
The permission path can be made significantly more restartable if authorization is an explicit phase before any physical tool side effect.
A useful durable distinction is:
tool call: authorizing -> executing -> settled permission request: pending -> approved | rejectedRecovery may reconstruct an
authorizingcall, but must never replay an ambiguousexecutingcall. Permission request IDs should be deterministic for the owning tool call (and request ordinal/resource set), with the request payload and terminal answer persisted. After all required approvals are present, transition the tool call toexecutingdurably before entering the body.Most current built-ins already place
permission.assert(...)before their physical side effects, but this is a convention inside each executor rather than an enforceable phase. Several tools perform path resolution/inspection first and some issue multiple permission assertions. A restartable design therefore needs one of:- an explicit tool authorization/preparation seam that computes every request before execution; or
- a built-in-only resumable protocol with a tested invariant that all work before the final authorization barrier is safe to repeat.
Dynamic/plugin tools and MCP calls should remain non-replayable unless they opt into that contract. In particular, an MCP elicitation may occur after a remote call has already started, so the generic permission/form recovery rule cannot safely replay it.
Persisting form CRUD state itself is straightforward: durable form row, payload, owner Session, status, answer, and compare-and-set reply/cancel.
Form.askshould use a deterministic owner ID and rejoin an existing pending or answered form after reconstruction. Shutdown must not publish a user cancellation merely because the Location graph is closing.The continuation still matters:
- The built-in question tool is a good first target because it has no physical work between authorization and
Form.ask; it can safely reconstruct and rejoin the durable form. - API-created/global forms can persist independently without resuming a tool.
- Forms raised mid-side-effect (notably MCP elicitation) need a separate resumable protocol or an explicit interrupted outcome.
Focused recovery coverage should prove: restart while authorizing, approval/rejection before and after the new process attaches, answer/cancel before and after reconstruction, no replay once
executing, deterministic retry IDs, and changed path/resource resolution causing a fresh authorization rather than using an approval for a different target.The concrete
next-15386tonext-15388incident in #36585 confirms this failure mode. #36591 provides the narrow TUI escape hatch: if a stale form reply returnsFormNotFoundError, refresh authoritative form state so the user is no longer trapped.The long-term design has now been worked through in
specs/v2/durable-session-interactions.md. The main conclusion is that persisting the form alone is insufficient. The original JavaScript continuation andDeferredare also gone, so recovery needs to park and later complete the original Step without replaying its provider request or ordinary Tool side effects.Key decisions from the reviewed plan:
- Start with the Core-owned
questionTool only. Durable permission waits need a separate preflight design because permissions can occur inside arbitrary side-effecting Tool execution. - After the provider attempt completes, settle ordinary eager Tools, prepare questions, capture the end snapshot, then publish one replayable atomic
session.step.parkedfact containing all questions and unconsumed Step facts. - Project each linked question call into a durable
waitingstate. Waiting is not process-local running state. - Admit answer/cancellation with a transactional first-response-wins compare-and-set. Current clients send a response ID for exact retry; legacy payloads remain temporarily accepted with weaker retry behavior.
- Treat a terminal question linked to a Waiting Tool as durable eligible Session work. A recovery loop schedules it at startup and on a bounded live cadence, so a lost advisory wake cannot strand it.
- Resume through a versioned Core-owned
question@1continuation adapter. Do not rerun the provider, ordinary Tools, before-hooks, permission, or the original question executor. - Block compaction, prompt promotion, and provider continuation while the Step is waiting.
- Human cancellation terminates the whole Waiting Step and derives sibling question aborts so no form remains live.
- Keep global MCP elicitation on a separately owned ephemeral path for the first slice. Do not let the
"global"sentinel enter Session persistence. - Use reader-before-writer rollout, a disabled writer flag through race testing, and a precursor database floor guard so unsupported rollback binaries refuse the database before serving or mutating it.
The plan has two gates before persistence work may begin:
- An executable question-continuation interface prototype proving codecs, output finalization, hook semantics, version resolution, and Code Mode exclusion without a second Tool representation.
- An atomic parking/response harness proving there is no commit boundary with a Waiting Tool but no replayable parked Step, and that terminal publication consumes pending Step state once.
The existing issue currently groups permissions and questions together. The vetted recommendation is to use the durable question path as the tracer bullet, then design permission durability separately rather than assuming it is another adapter over the same continuation.
- Start with the Core-owned
@kitlangton I think I've fixed this in this PR: #36604 not sure if it's worth a look
- added a commit that references this issue
on Jul 16, 2026 Gate 1 clarification: retryable finalization
One crash cut seems worth making explicit. If
question@1finalizes the answer and the process dies before the terminal Step transaction commits, recovery has no durable result to read and must finalize again. A journal written after finalization cannot remove that window.Should Gate 1 therefore require finalization to be deterministic and side-effect-free, with exactly-once applying to the durable terminal transition—including consuming pending Step state once—not to adapter invocation count?
A focused harness case would be:
finalize → crash before terminal commit → restart → finalize again → one committed terminal resultProvider work, ordinary Tools, hooks, permission checks, and the original question executor remain outside this recovery path.
I reproduced the QuestionV2 restart case on current
devusing two fresh runtime graphs backed by the same SQLite database.The Session survives the restart, while the pending Question disappears from
QuestionV2.list()in the new runtime.I'd like to take a narrow Question-only tracer bullet first. My current direction would be to make the waiting point durable and let a fresh runtime discover and settle the original request identity, without replaying provider/tool work.
Before I implement that boundary: is that still the intended direction for V2, or has the preferred continuation model changed since the earlier discussion?
- added a commit that references this issue
on Sep 19, 2026 - added a commit that references this issue
on Sep 20, 2026 Quick follow-up: I now have a local, tested Question-only prototype that preserves the original request identity across a fresh runtime and continues the Session without replaying the provider turn.
I haven't pushed it because the earlier discussion points to the parked/waiting continuation model, and I don't want to publish against a superseded boundary.
Is the Question-only tracer bullet still useful on current
dev, or should I wait for the continuation model to be clarified first? I'm happy to open it as a draft PR if seeing the concrete implementation would help.Field report: V2 2.0.24 — stale permission prompts; late reply gets 404; agent turn left dead
Reproducing this failure mode in a real long-running session on
@opencode/cli2.0.24 (npm,latestchannel). Windows 10 10.0.19045 (win32 x64), Cygwin bash configured shell, pluginopencode-working-memory.Setup
- Long-running agent sessions (hours per run) that regularly hit
askpermission checks (mostlyshell). - The user is frequently away from the console when the prompt appears and replies hours later, from both the TUI and the Web UI.
Observed behavior
- Agent issues a tool call → permission resolves to
ask→ client shows the prompt. - The prompt stays pending for hours while the user is away.
- The waiting turn is interrupted meanwhile (in my logs: both by service restarts from auto-update/manual restart, and — importantly — within a single uninterrupted service run).
- User returns and clicks "Allow".
- Nothing happens. The server answers the reply with 404, no
permission.repliedor other terminal event is emitted, so the client keeps rendering the prompt. The agent turn is dead, the session sits idle. The only recovery is asking the agent to repeat the last procedure.
Server-side evidence
From
~/.local/share/opencode/log/opencode.log(same sessionses_f6d5d693cffeWAy94MY594C1AZ, project "vm-iredmail"):2026-10-07T08:33:28Z POST /api/session/ses_f6d5d693.../permission/per_113d1fa7.../reply 404 (run 7ca6c751) 2026-10-07T11:48:18Z POST /api/session/ses_f6d5d693.../permission/per_115826a0.../reply 404 (run 7ca6c751) 2026-10-08T02:02:11Z POST /api/session/ses_f6d5d693.../permission/per_1186f09f.../reply 404 (run 87e9ed3b) 2026-10-09T05:10:00Z POST /api/session/ses_f6d5d693.../permission/per_1194e20d.../reply 404 (run 87e9ed3b) 2026-10-09T10:46:44Z POST /api/session/ses_f6d5d693.../permission/per_11f7d067.../reply 404 (run 87e9ed3b) 2026-10-09T19:19:17Z POST /api/session/ses_f6d5d693.../permission/per_1208f39e.../reply 404 (run 87e9ed3b) 2026-10-10T05:45:13Z POST /api/session/ses_f6d5d693.../permission/per_1224612c.../reply 404 (run 87e9ed3b)Notes:
run=...is the service process id. Run7ca6c751started with the 2.0.23 → 2.0.24 auto-update (2026-10-06 08:57); run87e9ed3bstarted with a manualopencode service stop/start (2026-10-07 22:10) and ran continuously through the last four 404s — i.e. the pending ask was dropped without any service restart in between (turn interruption within the run).GET /api/session/{id}/permissionfor the affected session returns empty while the client still renders the prompt.
Expected behavior (from the user's perspective)
Pending permission requests should wait for the user the way they did in V1: the session stays parked at the ask, and the user's reply hours later resumes the turn. In V2, anything that interrupts the waiting turn (service restart, provider stream abort, user action) drops the in-memory request and reconciles the tool as interrupted, so:
- the client has no way to know the prompt is dead (no terminal event) and keeps showing it;
- the user's reply 404s and is silently lost;
- the only recovery is re-prompting the agent to repeat the last procedure — with re-execution risk for tools that already had partial side effects.
Minimal asks that would unblock users on 2.0.x:
- Emit a terminal event (e.g.
permission.reverted/permission.expired) when a pending ask is dropped, so clients can dismiss the stale prompt; - or park the ask durably so a late reply resumes the session (along the durable-interactions direction discussed in this issue).
Related: #29422 (Web/Desktop stale prompt → "Permission request not found", closed not_planned), #28312 (TUI stale dialog), #36604 (TUI detach/reattach loses prompts). The TUI-side dismissal from #40960 is present in 2.0.24 — it hides the symptom in TUI but does not restore the turn.
- Long-running agent sessions (hours per run) that regularly hit
Summary
Pending permission requests and question forms are process-local in V2. Restarting the server destroys the request and its suspended continuation. Managed-service restart continuity may resume the Session, but it cannot restore or accept a reply to the old blocker.
We need a durable model for interactive waits that covers both permissions and questions without blindly replaying the tool call that created them.
Current behavior
Permissions keep pending requests in an in-memory
Mapwith aDeferredowned by the running tool fiber:packages/core/src/permission.tsQuestions use
Form.Service, whose entries andDeferreds live in an in-memory cache:packages/core/src/form.tsforms.ask(...)inside its tool executionOn managed-server shutdown, active Sessions are suspended for restart continuity. On startup, the runner resumes the Session and reconciles stale running tools as interrupted. The old permission/form request no longer exists, so clients cannot display or answer it.
This also gives shutdown the wrong domain meaning for forms: finalization publishes cancellation even though the user did not cancel the question.
Design pressure
Persisting the request payload alone is insufficient. The reply currently settles a process-local
Deferredinside a tool call, and that continuation disappears with the process.Permissions are the harder case because
permission.assert(...)can be called from within tool execution. Re-entering the tool to reach the permission check may repeat work or side effects that happened before the check. Treating every interrupted tool as replayable is therefore unsafe.The design should decide where the restartable boundary lives. Possibilities include:
Questions and permissions should share the same lifecycle concepts where possible, even if permission evaluation and form rendering remain separate APIs.
Required semantics
Coverage
Add an integration test that:
Cover permission approval, permission rejection, question answer, and question cancellation.
Related
specs/v2/session-restart-continuation.mddocuments the current restart boundary.