Skip to content

serve: file watcher re-registration storm during bulk skills writes terminates service process #50594

Description

@caimf1991

Summary

During bulk writes to the skills directories (a sync script within ~1 minute backs up / overwrites 7 SKILL.md files from a custom source, rewrites them with translations, and runs an installer that fully rebuilds skills into ~/.claude/skills), the background service enters a file-watcher re-registration storm on ~/.claude/skills, then the opencode-cli serve --service process dies with no ERROR/FATAL log line and no crash dump. A new serve process starts seconds later and in-session calls abort with "The server restarted while you were working".

Environment

  • opencode version: v2.0.12 (channel=latest, curl installer)
  • OS: Windows_NT 10.0.19045 (Windows 10 22H2, x64)
  • Terminal: Unavailable: TERM_PROGRAM not set (desktop app + background service)
  • Shell: C:\WINDOWS\system32\cmd.exe
  • Install: 2.0.12, channel=latest; background service runs opencode serve --service
  • Active plugins: superpowers.js (~/.config/opencode/plugins/superpowers.js, 17.6 KB, V1/V2 dual-compatible) — present and watched by the service; opencode.jsonc has no explicit plugins field (loading mechanism not verified)

Reproduction

  1. Start the background service (watches ~/.config/opencode/skills, ~/.claude/skills, ~/.config/opencode/superpowers/skills/*).
  2. Run a sync script that within ~1 minute: backs up skill dirs, overwrites 7 SKILL.md files (custom source), rewrites them with translations, and runs an installer that fully rebuilds skills/commands into ~/.claude/skills (68 dirs: 67 gsd-* + android-cli).
  3. Watch ~/.local/share/opencode/log/opencode.log during the write phase.
  4. Observe repeated watcher stopped / watcher subscribe / watcher started entries for the .claude/skills directory — dozens of times per second.
  5. Within 20–40 s of the storm peak, the serve process disappears from the log with no shutdown or error; seconds later a new cli starting ... args=["serve","--service"] appears and active sessions abort with "The server restarted while you were working".

Expected Behavior

The watcher should debounce/coalesce file events during bulk writes so the service survives large bursts of skill file changes.

Actual Behavior

  • Storm window 17:26:00–17:26:45 (local): 1137 watcher entries; calm window 17:20:00–17:20:45 (same process, same day): 0.
  • Serve process terminated twice today (restart records 17:07 and 17:26), both during the bulk-write phase — identical pattern: watcher storm → silent death → new serve begins seconds later; no ERROR/FATAL/panic line, no Windows Event 1000, no crash dump.
  • Windows Error Reporting recorded RADAR_PRE_LEAK_64 for opencode-cli.exe (file version 1.4.2.0) the same day — pre-crash virtual-memory growth warning, consistent with memory pressure before termination.
  • Desktop app processes (earlier processes) survived both incidents; only the background serve restarted.
  • Intermittent: 2 crashes, 2 clean runs today; both crashes coincided with the watcher storm during bulk writes.

Additional Context

Log excerpt (sanitized — usernames, run IDs, full paths replaced):

[17:26:31 run=***] message="watcher subscribe" path="C:\Users\<user>\.claude\skills" type=directory
[17:26:31 run=***] message="watcher stopped"   path="C:\Users\<user>\.claude\skills" type=directory
[17:26:31 run=***] message="watcher started"   path="C:\Users\<user>\.claude\skills" type=directory backend=windows
...   repeats dozens of times per second; 1137 entries in the 45 s storm window, 0 in the calm window ...
[17:26:45 run=***] message="cli starting" version=2.0.12 channel=latest args=["serve","--service"] role=cli
  • Storm target is exactly the directory written by the skills rebuild (~/.claude/skills).
  • skills rescanned entries appear for ~/.config/opencode/skills/**/*.md.en.orig during the same phase (second watched directory being written).
  • Desktop app survive, only serve restarted; new serve starts seconds later.

Activity

  1. github-actions commented on Sep 22, 2026

    @github-actions
    Contributor

    This issue might be a duplicate of an existing issue. Please check:

    The two reports share identical symptoms (watcher stop/start cycling dozens of times per second, silent process death, new serve starting seconds later, active sessions aborting), the same Windows platform, and the same affected component (~/.claude/skills or equivalent global skills directory). #47505 also references PR #43522 which mentions one-off Bun watcher crashes without stable reproduction — your detailed log excerpt (1137 entries in 45 s, two crash timestamps, RADAR_PRE_LEAK_64 entry) may help establish the reproduction path.

  2. juliojesus commented on Sep 30, 2026

    @juliojesus

    Another data point on Windows 11 (OpenCode v2.0.19 -> v2.0.20), with a different trigger and a slower failure mode than the bulk-write crash described here.

    Symptom: opencode serve --service climbs to ~1 CPU core while idle and stays there, with the log almost silent (a couple of lines per minute). opencode service restart brings it back to ~0.01 cores, but it degraded again ~2.5 hours after a restart, while about 7 TUIs were attached.

    What the log shows in the 30 minutes before a restart:

    message count
    watcher subscribe 571
    watcher started 37
    watcher stopped 37
    plugin reconciliation started / completed 60 / 60
    configuration normalization diagnostic 117

    Almost all watcher entries target the skills tree and its aliases: %LOCALAPPDATA%\ClaudeCode\skills (156), ~/.config/opencode/skills (40, a directory junction to the former), ~/.config/opencode/skill (40), ~/.opencode/skills (16), ~/.opencode/skill (16). Subscriptions far outnumber stops, so they appear to accumulate.

    Trigger: any CLI invocation against the service. Even opencode --version is followed within ~50 ms by ~8 watcher subscribe lines on the same skills directory and then back-to-back plugin reconciliation started/completed pairs (18 ms each, 14 plugins). Tools and scripts that call the CLI often (status checks, opencode api ...) therefore keep adding watchers during the day.

    Earlier on the same machine: the skills tree had a stray 2,300-file repository copied into one skill folder; trimming it to its SKILL.md (362 files in total now) lowered the load, but the server still drifts to ~1 core over hours, which points at the accumulation rather than the tree size alone. The junction also fits #50463.

    Would help: dedupe/reuse the skills watcher across CLI requests and locations (one watcher per real path, junctions resolved), and debounce reconciliation. Happy to collect more logs or run a debug build.

  3. juliojesus commented on Oct 1, 2026

    @juliojesus

    Follow-up: still happening on v2.0.21 (Windows 11).

    • opencode serve --service was at ~0 CPU right after starting at 07:57 and stayed fine for at least 4.4 h; at 5.8 h it was at ~1.05-1.2 cores while idle, memory normal (working set ~0.55 GB, commit 1.1 GB, 40 threads, 456 handles).
    • All of it is one thread: thread sampling shows a single thread at ~97% of a core, created when the process started (the main JS thread); every other thread idle. So it looks like synchronous work on the event loop, not I/O.
    • Log for the 90 minutes before I captured it:
    message count
    watcher subscribe 2136
    watcher started / watcher stopped 145 / 142
    plugin reconciliation started / completed 258 / 258
    configuration normalization diagnostic 648
    • Subscriptions arrive in bursts that line up with CLI calls against the service: 676 between 12:15 and 12:30 and 906 between 13:00 and 13:15 (local), around opencode -s <session>, opencode reload, opencode service status and opencode --version. Quiet 15-minute windows had 48-194.
    • opencode service restart returns it to ~0.01 cores until the next drift.

    I captured a full ProcDump of the degraded server (1.4 GB) before restarting. It contains credentials held in memory, so I won't post it here, but I can share it privately with a maintainer or run a debug build / extra tracing if that helps.

    For context, the TUI-side freezes from #51761 are gone for me on v2.0.21 (OpenTUI 0.5.13); this server-side drift is the remaining issue.

  4. juliojesus commented on Oct 3, 2026

    @juliojesus

    Update on v2.0.22 (Windows 11, 20 logical cores): the degradation has not shown up so far.

    • opencode serve --service started on 2026-10-02 06:39 and has now been up for ~30 h.
    • It is at ~0.01-0.03 cores while idle. Total CPU time is 49 min over 30 h, about 0.03 cores on average, and the working set is 0.54 GB.
    • On v2.0.19-v2.0.21 the same machine and setup went to ~1-1.2 cores on the main thread after 1.7-5.8 h every time, and only opencode service restart brought it back.
    • Same workload as before: 6 TUIs attached (hosted in Herdr), the same skills tree (363 files), and no restarts in between.

    I don't know whether something in 2.0.22 touched the watcher or I've just been lucky, so I'll report back if it comes back. If a change in 2.0.22 is expected to fix this, a pointer would help, and this can probably be closed after a few more days.

  5. juliojesus commented on Oct 4, 2026

    @juliojesus

    Correction to my previous comment: it did recur on v2.0.22. It just took longer.

    • opencode serve --service started 2026-10-02 06:39. It was fine at ~30 h (0.01-0.03 cores) and at 41.3 h it was at ~1.0-1.3 cores while idle, all of it on one thread (the other 40 threads idle). Working set was 0.65 GB, so memory is not the issue.
    • The log has the same signature: 21,581 watcher subscribe lines for the same directory (%LOCALAPPDATA%\ClaudeCode\skills, 363 files), still arriving in bursts of 40-80 per minute while nothing was writing to that tree:
    timestamp=2026-10-04T05:56:20.556Z level=INFO message="watcher subscribe" path="C:\Users\julio\AppData\Local\ClaudeCode\skills" type=directory ignores=0
    timestamp=2026-10-04T05:56:20.563Z level=INFO message="watcher subscribe" path="C:\Users\julio\AppData\Local\ClaudeCode\skills" type=directory ignores=0
    ... (40 lines with the same path within ~80 ms)
    
    • opencode service restart brought it back to 0.00 cores right away, as before.

    Before restarting I took a full ProcDump (-ma, 1.65 GB) of the degraded server. It contains credentials, so I can't attach it publicly, but I'm happy to run specific WinDbg commands on it or share native stacks of the hot thread if that helps.

    So on Windows 11 the timeline is: 1.7-5.8 h on v2.0.19-2.0.21, and 30-41 h on v2.0.22. That points to subscriptions still accumulating per directory instead of being deduplicated or released.

  6. juliojesus commented on Oct 4, 2026

    @juliojesus

    Related data point, from the TUI side this time (Windows 11, v2.0.22). It may be the same watcher problem showing up in the client.

    Two TUIs (opencode -s <id>, attached to the service) were each burning ~1 core while idle for hours. Memory stayed flat at ~0.25 GB, so this is not the #51761 leak. Restarting the server did not calm them. Both had very large sessions (5.2M and 3.0M tokens).

    Before closing them I took full dumps (procdump -ma). In both, the cost is a single thread, the JS main thread, and both were sitting in the same native frame:

    process main thread CPU top of stack at dump time
    hot TUI #1 18,028 s ntdll!NtNotifyChangeDirectoryFileEx < KERNELBASE!ReadDirectoryChangesW+0xb1 < opencode.exe+0xaf786b
    hot TUI #2 17,822 s same frames, same return address opencode.exe+0xaf786b
    idle TUI (same cwd, for comparison) 1,697 s KERNELBASE!GetQueuedCompletionStatusEx (normal event-loop wait)

    Every other thread in the hot processes has under 85 s of CPU.

    So the busy loop seems to be a directory watch being re-armed synchronously on the main thread (ReadDirectoryChangesW from opencode.exe), rather than rendering. It's one snapshot per process, so not proof, but both landed on the identical call site. I couldn't get the watched path: the dump has no handle names. I also don't have Bun symbols to resolve opencode.exe+0xaf786b. If you can map that offset for the v2.0.22 Windows x64 build, it would name the function.

    The dumps contain credentials, so I can't attach them, but I can run whatever commands you'd like on them.

  7. juliojesus commented on Oct 7, 2026

    @juliojesus

    I think I found a reliable trigger for the TUI side of this: a plugin reload while TUIs are open (Windows 11, v2.0.24).

    What happened today

    • 9 TUIs had been open since 07:22, averaging ~0.04 cores each.
    • At 15:54 I reinstalled a local plugin, and the install script ran opencode reload. The log shows the reload, followed by watcher subscribe lines for each file of that plugin (type=file).
    • From that moment, all 9 TUIs sat at ~0.35-0.5 cores each while idle (~3.5 cores in total). Memory stayed flat (~0.2 GB each).
    • A TUI opened after the reload stays at 0.00-0.02 cores, and so does the server (0.02-0.05). Only the TUIs that were open during the reload are affected.

    Where the time goes: I took three procdump -mm samples of one hot TUI, a few seconds apart. The cost is the main thread (1,374 s of CPU; no other thread has more than 16 s):

    sample 1: ntdll!NtNotifyChangeDirectoryFileEx < KERNELBASE!ReadDirectoryChangesW+0xb1
    sample 2: ntdll!NtQueryInformationFile        < KERNELBASE!GetFileInformationByHandleEx+0x78
    sample 3: ntdll!NtNotifyChangeDirectoryFileEx < KERNELBASE!ReadDirectoryChangesW+0xb1
    

    This is the same ReadDirectoryChangesW frame (opencode.exe+0xaf786b on 2.0.22) as in my previous comment. Looking back, that case fits too: those two TUIs got hot right after two plugin installs that each ran opencode reload.

    Repro (Windows)

    1. Open a few TUIs and let them idle. They sit near 0 cores.
    2. Reinstall (copy over) a local plugin under ~/.config/opencode/plugins/ and/or run opencode reload. In my case both happened at once, so I can't yet say whether the explicit reload or the file-change hot reload (the host watches the plugin files) is the trigger.
    3. Every TUI that was open now uses ~0.4 cores while idle, indefinitely. TUIs started after the reload are fine. Restarting the server does not fix the affected ones; only reopening them does.

    So on reload the TUI seems to re-create its file watches on the main thread without releasing the old ones, or to re-arm them in a tight loop. As a workaround I've stopped running opencode reload from my install script. That may not be enough if the file-watch hot reload alone causes it; I'll report back after the next reinstall.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions