Skip to content

V2: preserve permission and question waits across server restart #36347

Description

@kitlangton

Summary

Pending permission requests and question forms are process-local in V2. Restarting the server destroys the request and its suspended continuation. Managed-service restart continuity may resume the Session, but it cannot restore or accept a reply to the old blocker.

We need a durable model for interactive waits that covers both permissions and questions without blindly replaying the tool call that created them.

Current behavior

Permissions keep pending requests in an in-memory Map with a Deferred owned by the running tool fiber:

  • packages/core/src/permission.ts
  • layer finalization declines every pending request and clears the map

Questions use Form.Service, whose entries and Deferreds live in an in-memory cache:

  • packages/core/src/form.ts
  • layer finalization cancels every pending form
  • the built-in question tool waits on forms.ask(...) inside its tool execution

On managed-server shutdown, active Sessions are suspended for restart continuity. On startup, the runner resumes the Session and reconciles stale running tools as interrupted. The old permission/form request no longer exists, so clients cannot display or answer it.

This also gives shutdown the wrong domain meaning for forms: finalization publishes cancellation even though the user did not cancel the question.

Design pressure

Persisting the request payload alone is insufficient. The reply currently settles a process-local Deferred inside a tool call, and that continuation disappears with the process.

Permissions are the harder case because permission.assert(...) can be called from within tool execution. Re-entering the tool to reach the permission check may repeat work or side effects that happened before the check. Treating every interrupted tool as replayable is therefore unsafe.

The design should decide where the restartable boundary lives. Possibilities include:

  • Make interactive blockers durable workflow records whose terminal reply can settle the owning tool call without restoring the old JavaScript continuation.
  • Move authorization to an explicit boundary before the physical tool attempt, so an approved attempt can start safely after restart.
  • Define a resumable tool protocol for tools that can suspend on interactive input, while ordinary tools remain non-replayable.

Questions and permissions should share the same lifecycle concepts where possible, even if permission evaluation and form rendering remain separate APIs.

Required semantics

  • A pending permission or question survives a graceful managed-service restart and remains visible through the existing list APIs.
  • The original request ID remains replyable after restart.
  • Server shutdown does not record a user rejection or cancellation.
  • A reply settles the blocker exactly once and allows the owning Session to continue.
  • Recovery does not repeat provider work or tool side effects that may already have executed.
  • Saved permission rules remain distinct from one pending permission decision.
  • Reconnect hydration and live events converge on the same durable blocker state.
  • Location and Session ownership checks continue to prevent cross-location or cross-session replies.
  • The design explicitly states whether unexpected process death is supported or remains outside the first implementation.

Coverage

Add an integration test that:

  1. Starts a Session that blocks on a permission request or question form.
  2. Tears down the server/location graph while preserving the database.
  3. Builds a fresh graph and performs managed restart continuation.
  4. Verifies the same blocker is listed and replyable.
  5. Replies and verifies the Session continues without duplicate tool side effects.

Cover permission approval, permission rejection, question answer, and question cancellation.

Related

Activity

  1. added
    bugSomething isn't working
    coreAnything pertaining to core functionality of the application (opencode server stuff)
    on Jul 11, 2026
  2. opencode-agent commented on Jul 11, 2026

    @opencode-agent
    Contributor

    Follow-up: permission preflight and persisted forms

    The permission path can be made significantly more restartable if authorization is an explicit phase before any physical tool side effect.

    A useful durable distinction is:

    tool call: authorizing -> executing -> settled
    permission request: pending -> approved | rejected
    

    Recovery may reconstruct an authorizing call, but must never replay an ambiguous executing call. Permission request IDs should be deterministic for the owning tool call (and request ordinal/resource set), with the request payload and terminal answer persisted. After all required approvals are present, transition the tool call to executing durably before entering the body.

    Most current built-ins already place permission.assert(...) before their physical side effects, but this is a convention inside each executor rather than an enforceable phase. Several tools perform path resolution/inspection first and some issue multiple permission assertions. A restartable design therefore needs one of:

    • an explicit tool authorization/preparation seam that computes every request before execution; or
    • a built-in-only resumable protocol with a tested invariant that all work before the final authorization barrier is safe to repeat.

    Dynamic/plugin tools and MCP calls should remain non-replayable unless they opt into that contract. In particular, an MCP elicitation may occur after a remote call has already started, so the generic permission/form recovery rule cannot safely replay it.

    Persisting form CRUD state itself is straightforward: durable form row, payload, owner Session, status, answer, and compare-and-set reply/cancel. Form.ask should use a deterministic owner ID and rejoin an existing pending or answered form after reconstruction. Shutdown must not publish a user cancellation merely because the Location graph is closing.

    The continuation still matters:

    • The built-in question tool is a good first target because it has no physical work between authorization and Form.ask; it can safely reconstruct and rejoin the durable form.
    • API-created/global forms can persist independently without resuming a tool.
    • Forms raised mid-side-effect (notably MCP elicitation) need a separate resumable protocol or an explicit interrupted outcome.

    Focused recovery coverage should prove: restart while authorizing, approval/rejection before and after the new process attaches, answer/cancel before and after reconstruction, no replay once executing, deterministic retry IDs, and changed path/resource resolution causing a fresh authorization rather than using an approval for a different target.

  3. kitlangton commented on Jul 13, 2026

    @kitlangton
    ContributorAuthor

    The concrete next-15386 to next-15388 incident in #36585 confirms this failure mode. #36591 provides the narrow TUI escape hatch: if a stale form reply returns FormNotFoundError, refresh authoritative form state so the user is no longer trapped.

    The long-term design has now been worked through in specs/v2/durable-session-interactions.md. The main conclusion is that persisting the form alone is insufficient. The original JavaScript continuation and Deferred are also gone, so recovery needs to park and later complete the original Step without replaying its provider request or ordinary Tool side effects.

    Key decisions from the reviewed plan:

    • Start with the Core-owned question Tool only. Durable permission waits need a separate preflight design because permissions can occur inside arbitrary side-effecting Tool execution.
    • After the provider attempt completes, settle ordinary eager Tools, prepare questions, capture the end snapshot, then publish one replayable atomic session.step.parked fact containing all questions and unconsumed Step facts.
    • Project each linked question call into a durable waiting state. Waiting is not process-local running state.
    • Admit answer/cancellation with a transactional first-response-wins compare-and-set. Current clients send a response ID for exact retry; legacy payloads remain temporarily accepted with weaker retry behavior.
    • Treat a terminal question linked to a Waiting Tool as durable eligible Session work. A recovery loop schedules it at startup and on a bounded live cadence, so a lost advisory wake cannot strand it.
    • Resume through a versioned Core-owned question@1 continuation adapter. Do not rerun the provider, ordinary Tools, before-hooks, permission, or the original question executor.
    • Block compaction, prompt promotion, and provider continuation while the Step is waiting.
    • Human cancellation terminates the whole Waiting Step and derives sibling question aborts so no form remains live.
    • Keep global MCP elicitation on a separately owned ephemeral path for the first slice. Do not let the "global" sentinel enter Session persistence.
    • Use reader-before-writer rollout, a disabled writer flag through race testing, and a precursor database floor guard so unsupported rollback binaries refuse the database before serving or mutating it.

    The plan has two gates before persistence work may begin:

    1. An executable question-continuation interface prototype proving codecs, output finalization, hook semantics, version resolution, and Code Mode exclusion without a second Tool representation.
    2. An atomic parking/response harness proving there is no commit boundary with a Waiting Tool but no replayable parked Step, and that terminal publication consumes pending Step state once.

    The existing issue currently groups permissions and questions together. The vetted recommendation is to use the durable question path as the tracer bullet, then design permission durability separately rather than assuming it is another adapter over the same continuation.

  4. Mykhol commented on Jul 14, 2026

    @Mykhol

    @kitlangton I think I've fixed this in this PR: #36604 not sure if it's worth a look

  5. dev404ai commented on Aug 10, 2026

    @dev404ai

    Gate 1 clarification: retryable finalization

    One crash cut seems worth making explicit. If question@1 finalizes the answer and the process dies before the terminal Step transaction commits, recovery has no durable result to read and must finalize again. A journal written after finalization cannot remove that window.

    Should Gate 1 therefore require finalization to be deterministic and side-effect-free, with exactly-once applying to the durable terminal transition—including consuming pending Step state once—not to adapter invocation count?

    A focused harness case would be:

    finalize → crash before terminal commit → restart → finalize again → one committed terminal result

    Provider work, ordinary Tools, hooks, permission checks, and the original question executor remain outside this recovery path.

  6. gcomneno commented on Sep 17, 2026

    @gcomneno

    I reproduced the QuestionV2 restart case on current dev using two fresh runtime graphs backed by the same SQLite database.

    The Session survives the restart, while the pending Question disappears from QuestionV2.list() in the new runtime.

    I'd like to take a narrow Question-only tracer bullet first. My current direction would be to make the waiting point durable and let a fresh runtime discover and settle the original request identity, without replaying provider/tool work.

    Before I implement that boundary: is that still the intended direction for V2, or has the preferred continuation model changed since the earlier discussion?

  7. gcomneno commented on Sep 24, 2026

    @gcomneno

    Quick follow-up: I now have a local, tested Question-only prototype that preserves the original request identity across a fresh runtime and continues the Session without replaying the provider turn.

    I haven't pushed it because the earlier discussion points to the parked/waiting continuation model, and I don't want to publish against a superseded boundary.

    Is the Question-only tracer bullet still useful on current dev, or should I wait for the continuation model to be clarified first? I'm happy to open it as a draft PR if seeing the concrete implementation would help.

  8. vc commented on Oct 10, 2026

    @vc

    Field report: V2 2.0.24 — stale permission prompts; late reply gets 404; agent turn left dead

    Reproducing this failure mode in a real long-running session on @opencode/cli 2.0.24 (npm, latest channel). Windows 10 10.0.19045 (win32 x64), Cygwin bash configured shell, plugin opencode-working-memory.

    Setup

    • Long-running agent sessions (hours per run) that regularly hit ask permission checks (mostly shell).
    • The user is frequently away from the console when the prompt appears and replies hours later, from both the TUI and the Web UI.

    Observed behavior

    1. Agent issues a tool call → permission resolves to ask → client shows the prompt.
    2. The prompt stays pending for hours while the user is away.
    3. The waiting turn is interrupted meanwhile (in my logs: both by service restarts from auto-update/manual restart, and — importantly — within a single uninterrupted service run).
    4. User returns and clicks "Allow".
    5. Nothing happens. The server answers the reply with 404, no permission.replied or other terminal event is emitted, so the client keeps rendering the prompt. The agent turn is dead, the session sits idle. The only recovery is asking the agent to repeat the last procedure.

    Server-side evidence

    From ~/.local/share/opencode/log/opencode.log (same session ses_f6d5d693cffeWAy94MY594C1AZ, project "vm-iredmail"):

    2026-10-07T08:33:28Z  POST /api/session/ses_f6d5d693.../permission/per_113d1fa7.../reply  404   (run 7ca6c751)
    2026-10-07T11:48:18Z  POST /api/session/ses_f6d5d693.../permission/per_115826a0.../reply  404   (run 7ca6c751)
    2026-10-08T02:02:11Z  POST /api/session/ses_f6d5d693.../permission/per_1186f09f.../reply  404   (run 87e9ed3b)
    2026-10-09T05:10:00Z  POST /api/session/ses_f6d5d693.../permission/per_1194e20d.../reply  404   (run 87e9ed3b)
    2026-10-09T10:46:44Z  POST /api/session/ses_f6d5d693.../permission/per_11f7d067.../reply  404   (run 87e9ed3b)
    2026-10-09T19:19:17Z  POST /api/session/ses_f6d5d693.../permission/per_1208f39e.../reply  404   (run 87e9ed3b)
    2026-10-10T05:45:13Z  POST /api/session/ses_f6d5d693.../permission/per_1224612c.../reply  404   (run 87e9ed3b)
    

    Notes:

    • run=... is the service process id. Run 7ca6c751 started with the 2.0.23 → 2.0.24 auto-update (2026-10-06 08:57); run 87e9ed3b started with a manual opencode service stop/start (2026-10-07 22:10) and ran continuously through the last four 404s — i.e. the pending ask was dropped without any service restart in between (turn interruption within the run).
    • GET /api/session/{id}/permission for the affected session returns empty while the client still renders the prompt.

    Expected behavior (from the user's perspective)

    Pending permission requests should wait for the user the way they did in V1: the session stays parked at the ask, and the user's reply hours later resumes the turn. In V2, anything that interrupts the waiting turn (service restart, provider stream abort, user action) drops the in-memory request and reconciles the tool as interrupted, so:

    1. the client has no way to know the prompt is dead (no terminal event) and keeps showing it;
    2. the user's reply 404s and is silently lost;
    3. the only recovery is re-prompting the agent to repeat the last procedure — with re-execution risk for tools that already had partial side effects.

    Minimal asks that would unblock users on 2.0.x:

    • Emit a terminal event (e.g. permission.reverted / permission.expired) when a pending ask is dropped, so clients can dismiss the stale prompt;
    • or park the ask durably so a late reply resumes the session (along the durable-interactions direction discussed in this issue).

    Related: #29422 (Web/Desktop stale prompt → "Permission request not found", closed not_planned), #28312 (TUI stale dialog), #36604 (TUI detach/reattach loses prompts). The TUI-side dismissal from #40960 is present in 2.0.24 — it hides the symptom in TUI but does not restore the turn.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    2.0bugSomething isn't workingcoreAnything pertaining to core functionality of the application (opencode server stuff)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions