Repository navigation
serve: file watcher re-registration storm during bulk skills writes terminates service process #50594
Description
Activity
github-actions commented
on Sep 22, 2026 on Sep 22, 2026 – with GitHub ActionsContributorMore actionsThis issue might be a duplicate of an existing issue. Please check:
- core: global skill update terminates shared service on Windows #47505: Same root cause — bulk skill file writes trigger a watcher re-registration storm on Windows, causing the
serve --serviceprocess to terminate silently with no error log, followed by automatic restart and session interruption.
The two reports share identical symptoms (watcher stop/start cycling dozens of times per second, silent process death, new serve starting seconds later, active sessions aborting), the same Windows platform, and the same affected component (
~/.claude/skillsor equivalent global skills directory). #47505 also references PR #43522 which mentions one-off Bun watcher crashes without stable reproduction — your detailed log excerpt (1137 entries in 45 s, two crash timestamps, RADAR_PRE_LEAK_64 entry) may help establish the reproduction path.- core: global skill update terminates shared service on Windows #47505: Same root cause — bulk skill file writes trigger a watcher re-registration storm on Windows, causing the
Another data point on Windows 11 (OpenCode v2.0.19 -> v2.0.20), with a different trigger and a slower failure mode than the bulk-write crash described here.
Symptom:
opencode serve --serviceclimbs to ~1 CPU core while idle and stays there, with the log almost silent (a couple of lines per minute).opencode service restartbrings it back to ~0.01 cores, but it degraded again ~2.5 hours after a restart, while about 7 TUIs were attached.What the log shows in the 30 minutes before a restart:
message count watcher subscribe571 watcher started37 watcher stopped37 plugin reconciliation started/completed60 / 60 configuration normalization diagnostic117 Almost all watcher entries target the skills tree and its aliases:
%LOCALAPPDATA%\ClaudeCode\skills(156),~/.config/opencode/skills(40, a directory junction to the former),~/.config/opencode/skill(40),~/.opencode/skills(16),~/.opencode/skill(16). Subscriptions far outnumber stops, so they appear to accumulate.Trigger: any CLI invocation against the service. Even
opencode --versionis followed within ~50 ms by ~8watcher subscribelines on the same skills directory and then back-to-backplugin reconciliation started/completedpairs (18 ms each, 14 plugins). Tools and scripts that call the CLI often (status checks,opencode api ...) therefore keep adding watchers during the day.Earlier on the same machine: the skills tree had a stray 2,300-file repository copied into one skill folder; trimming it to its
SKILL.md(362 files in total now) lowered the load, but the server still drifts to ~1 core over hours, which points at the accumulation rather than the tree size alone. The junction also fits #50463.Would help: dedupe/reuse the skills watcher across CLI requests and locations (one watcher per real path, junctions resolved), and debounce reconciliation. Happy to collect more logs or run a debug build.
Follow-up: still happening on v2.0.21 (Windows 11).
opencode serve --servicewas at ~0 CPU right after starting at 07:57 and stayed fine for at least 4.4 h; at 5.8 h it was at ~1.05-1.2 cores while idle, memory normal (working set ~0.55 GB, commit 1.1 GB, 40 threads, 456 handles).- All of it is one thread: thread sampling shows a single thread at ~97% of a core, created when the process started (the main JS thread); every other thread idle. So it looks like synchronous work on the event loop, not I/O.
- Log for the 90 minutes before I captured it:
message count watcher subscribe2136 watcher started/watcher stopped145 / 142 plugin reconciliation started/completed258 / 258 configuration normalization diagnostic648 - Subscriptions arrive in bursts that line up with CLI calls against the service: 676 between 12:15 and 12:30 and 906 between 13:00 and 13:15 (local), around
opencode -s <session>,opencode reload,opencode service statusandopencode --version. Quiet 15-minute windows had 48-194. opencode service restartreturns it to ~0.01 cores until the next drift.
I captured a full ProcDump of the degraded server (1.4 GB) before restarting. It contains credentials held in memory, so I won't post it here, but I can share it privately with a maintainer or run a debug build / extra tracing if that helps.
For context, the TUI-side freezes from #51761 are gone for me on v2.0.21 (OpenTUI 0.5.13); this server-side drift is the remaining issue.
Update on v2.0.22 (Windows 11, 20 logical cores): the degradation has not shown up so far.
opencode serve --servicestarted on 2026-10-02 06:39 and has now been up for ~30 h.- It is at ~0.01-0.03 cores while idle. Total CPU time is 49 min over 30 h, about 0.03 cores on average, and the working set is 0.54 GB.
- On v2.0.19-v2.0.21 the same machine and setup went to ~1-1.2 cores on the main thread after 1.7-5.8 h every time, and only
opencode service restartbrought it back. - Same workload as before: 6 TUIs attached (hosted in Herdr), the same skills tree (363 files), and no restarts in between.
I don't know whether something in 2.0.22 touched the watcher or I've just been lucky, so I'll report back if it comes back. If a change in 2.0.22 is expected to fix this, a pointer would help, and this can probably be closed after a few more days.
Correction to my previous comment: it did recur on v2.0.22. It just took longer.
opencode serve --servicestarted 2026-10-02 06:39. It was fine at ~30 h (0.01-0.03 cores) and at 41.3 h it was at ~1.0-1.3 cores while idle, all of it on one thread (the other 40 threads idle). Working set was 0.65 GB, so memory is not the issue.- The log has the same signature: 21,581
watcher subscribelines for the same directory (%LOCALAPPDATA%\ClaudeCode\skills, 363 files), still arriving in bursts of 40-80 per minute while nothing was writing to that tree:
timestamp=2026-10-04T05:56:20.556Z level=INFO message="watcher subscribe" path="C:\Users\julio\AppData\Local\ClaudeCode\skills" type=directory ignores=0 timestamp=2026-10-04T05:56:20.563Z level=INFO message="watcher subscribe" path="C:\Users\julio\AppData\Local\ClaudeCode\skills" type=directory ignores=0 ... (40 lines with the same path within ~80 ms)opencode service restartbrought it back to 0.00 cores right away, as before.
Before restarting I took a full ProcDump (
-ma, 1.65 GB) of the degraded server. It contains credentials, so I can't attach it publicly, but I'm happy to run specific WinDbg commands on it or share native stacks of the hot thread if that helps.So on Windows 11 the timeline is: 1.7-5.8 h on v2.0.19-2.0.21, and 30-41 h on v2.0.22. That points to subscriptions still accumulating per directory instead of being deduplicated or released.
Related data point, from the TUI side this time (Windows 11, v2.0.22). It may be the same watcher problem showing up in the client.
Two TUIs (
opencode -s <id>, attached to the service) were each burning ~1 core while idle for hours. Memory stayed flat at ~0.25 GB, so this is not the #51761 leak. Restarting the server did not calm them. Both had very large sessions (5.2M and 3.0M tokens).Before closing them I took full dumps (
procdump -ma). In both, the cost is a single thread, the JS main thread, and both were sitting in the same native frame:process main thread CPU top of stack at dump time hot TUI #1 18,028 s ntdll!NtNotifyChangeDirectoryFileEx<KERNELBASE!ReadDirectoryChangesW+0xb1<opencode.exe+0xaf786bhot TUI #2 17,822 s same frames, same return address opencode.exe+0xaf786bidle TUI (same cwd, for comparison) 1,697 s KERNELBASE!GetQueuedCompletionStatusEx(normal event-loop wait)Every other thread in the hot processes has under 85 s of CPU.
So the busy loop seems to be a directory watch being re-armed synchronously on the main thread (
ReadDirectoryChangesWfromopencode.exe), rather than rendering. It's one snapshot per process, so not proof, but both landed on the identical call site. I couldn't get the watched path: the dump has no handle names. I also don't have Bun symbols to resolveopencode.exe+0xaf786b. If you can map that offset for the v2.0.22 Windows x64 build, it would name the function.The dumps contain credentials, so I can't attach them, but I can run whatever commands you'd like on them.
I think I found a reliable trigger for the TUI side of this: a plugin reload while TUIs are open (Windows 11, v2.0.24).
What happened today
- 9 TUIs had been open since 07:22, averaging ~0.04 cores each.
- At 15:54 I reinstalled a local plugin, and the install script ran
opencode reload. The log shows the reload, followed bywatcher subscribelines for each file of that plugin (type=file). - From that moment, all 9 TUIs sat at ~0.35-0.5 cores each while idle (~3.5 cores in total). Memory stayed flat (~0.2 GB each).
- A TUI opened after the reload stays at 0.00-0.02 cores, and so does the server (0.02-0.05). Only the TUIs that were open during the reload are affected.
Where the time goes: I took three
procdump -mmsamples of one hot TUI, a few seconds apart. The cost is the main thread (1,374 s of CPU; no other thread has more than 16 s):sample 1: ntdll!NtNotifyChangeDirectoryFileEx < KERNELBASE!ReadDirectoryChangesW+0xb1 sample 2: ntdll!NtQueryInformationFile < KERNELBASE!GetFileInformationByHandleEx+0x78 sample 3: ntdll!NtNotifyChangeDirectoryFileEx < KERNELBASE!ReadDirectoryChangesW+0xb1This is the same
ReadDirectoryChangesWframe (opencode.exe+0xaf786bon 2.0.22) as in my previous comment. Looking back, that case fits too: those two TUIs got hot right after two plugin installs that each ranopencode reload.Repro (Windows)
- Open a few TUIs and let them idle. They sit near 0 cores.
- Reinstall (copy over) a local plugin under
~/.config/opencode/plugins/and/or runopencode reload. In my case both happened at once, so I can't yet say whether the explicit reload or the file-change hot reload (the host watches the plugin files) is the trigger. - Every TUI that was open now uses ~0.4 cores while idle, indefinitely. TUIs started after the reload are fine. Restarting the server does not fix the affected ones; only reopening them does.
So on reload the TUI seems to re-create its file watches on the main thread without releasing the old ones, or to re-arm them in a tight loop. As a workaround I've stopped running
opencode reloadfrom my install script. That may not be enough if the file-watch hot reload alone causes it; I'll report back after the next reinstall.
Summary
During bulk writes to the skills directories (a sync script within ~1 minute backs up / overwrites 7
SKILL.mdfiles from a custom source, rewrites them with translations, and runs an installer that fully rebuilds skills into~/.claude/skills), the background service enters a file-watcher re-registration storm on~/.claude/skills, then theopencode-cli serve --serviceprocess dies with no ERROR/FATAL log line and no crash dump. A new serve process starts seconds later and in-session calls abort with "The server restarted while you were working".Environment
TERM_PROGRAMnot set (desktop app + background service)C:\WINDOWS\system32\cmd.exeopencode serve --servicesuperpowers.js(~/.config/opencode/plugins/superpowers.js, 17.6 KB, V1/V2 dual-compatible) — present and watched by the service;opencode.jsonchas no explicitpluginsfield (loading mechanism not verified)Reproduction
~/.config/opencode/skills,~/.claude/skills,~/.config/opencode/superpowers/skills/*).SKILL.mdfiles (custom source), rewrites them with translations, and runs an installer that fully rebuilds skills/commands into~/.claude/skills(68 dirs: 67gsd-*+android-cli).~/.local/share/opencode/log/opencode.logduring the write phase.watcher stopped/watcher subscribe/watcher startedentries for the.claude/skillsdirectory — dozens of times per second.cli starting ... args=["serve","--service"]appears and active sessions abort with "The server restarted while you were working".Expected Behavior
The watcher should debounce/coalesce file events during bulk writes so the service survives large bursts of skill file changes.
Actual Behavior
RADAR_PRE_LEAK_64foropencode-cli.exe(file version 1.4.2.0) the same day — pre-crash virtual-memory growth warning, consistent with memory pressure before termination.Additional Context
Log excerpt (sanitized — usernames, run IDs, full paths replaced):
~/.claude/skills).skills rescannedentries appear for~/.config/opencode/skills/**/*.md.en.origduring the same phase (second watched directory being written).