Symptom
task-done prints:
Skipping zellij close-tab (no tab id captured; close this tab manually if needed).
and the tab does not close. Reported on both kiro and claude, intermittently.
Cause: the capture fails silently at a different time, in a different command
claude-gh/bin/task-work:295-298 (and the six sibling copies, region info-tab-id):
INFO_TAB_ID=$(aw_zellij_tab_id_by_name "$SLUG" || true)
if [[ -n "$INFO_TAB_ID" ]]; then
printf 'TAB_ID=%s\n' "$INFO_TAB_ID" >> "$INFO_FILE"
fi
When the lookup returns empty, no TAB_ID= line is written and nothing is printed. The block is not guarded on $ZELLIJ either — it runs unconditionally and fails quietly.
So task-work succeeds, the worker runs for hours, and the first sign anything went wrong is task-done's message at teardown — by which time the context that caused it is gone. The user cannot act on the message because it reports a consequence in the wrong command.
This is the same failure shape the repo has now fixed three times: #194 (an empty check list reads as green), #182 (a filter that never executed and failed open), #229 (two distinct failures that both wrote no log line). A state that is invisible because its failure mode produces no artifact.
Evidence
Auditing every <worktree-base>/.<slug>.info on this machine — 9 of 36 sidecars across 8 repos have no TAB_ID, and they are not randomly distributed:
gdi-chatbot-worktrees:
.add-readme.info TAB_ID=[]
.add-readme-7qegu.info .add-readme-9wgw1.info TAB_ID=[]
.add-readme-g8sgw.info .add-readme-ldwpy.info TAB_ID=[]
.add-readme-smy9q.info TAB_ID=[]
.bug-model-renders-table-but-the-chat-already-outpu.info TAB_ID=[]
.bug-model-renders-table-but-the-chat-already-outpu-hu1s5.info TAB_ID=[]
.fix-stores-skill.info TAB_ID=[]
every other repo (retail-scraper, recommender-systems, groppo, portai, task-force): TAB_ID populated
Two shapes stand out, both produced by task-work itself:
- Collision suffixes.
task-work:206-207 appends a random 5-char hash on a name collision (SLUG="${SLUG}-${HASH}"). Six add-readme variants exist, so that path ran repeatedly.
- 50-char truncation.
task-work:136 does SLUG="${SLUG:0:50}". bug-model-renders-table-but-the-chat-already-outpu is exactly 50 characters.
Both are cases where the name task-work asks zellij for is one it derived through a transformation — which is where a mismatch between the slug and the actual tab name would live. Not proven; see below.
What the lookup can return empty for
lib/zellij-tab.sh:38-52, aw_zellij_tab_id_by_name, ends in | .[0] // empty, so empty means no tab matched the name. Four ways to get there:
$ZELLIJ unset — task-work run outside zellij, so no tab was created at all. The block does not check for this.
- no
jq on PATH — :42 returns empty.
- the tab's name is not byte-equal to
$SLUG after paint-stripping. Note the paint set (⏸️ /▶️ /❓︎ ) matches radio's exactly, so the documented race the comment addresses is genuinely covered — this would be a different mismatch, e.g. a long name zellij renders differently.
- a genuine race despite (3).
I could not determine which of these fires from artifacts alone. The distribution above points at (3) and the two transformations that produce it, but the failing tabs are long gone and the message carries no diagnostic.
Fix
Two parts, and the first matters more than the second:
1. Make the failure loud, at the point it happens. When the lookup comes back empty, task-work should say so and say why it can tell — zellij absent, jq absent, or a name lookup that found nothing — and name the slug it searched for. Log it too, so the reconstruction is a grep rather than an audit of every sidecar on the machine. Whatever the trigger turns out to be, this is what makes the next report diagnosable in one line instead of this.
2. Then find the trigger, with (1) in place to report it. Worth checking specifically whether zellij's list-tabs --json .name is byte-equal to the requested name for a 50-char slug, and whether anything in the collision-suffix path creates the tab under a different name than it writes to the sidecar.
Do not fall back to the unscoped zellij action close-tab in task-done — #107/#108 established that closes whichever tab is focused, and task-done:201 already explains why that is worse than skipping.
Verification
- With the lookup forced to return empty,
task-work reports it (message + log line) and still completes — a failed tab-id capture must not abort a worker launch.
- The message distinguishes the causes it can distinguish: no zellij, no jq, no match for
<slug>.
- A successful capture is byte-unchanged in output (regression guard).
task-done's existing "no tab id captured" message gains a pointer to the task-work log line, so the two ends of the failure are linkable.
- All seven loadouts: this lives in the drift-guarded
info-tab-id region, so the fix lands in every copy or the guard fails.
References
Symptom
task-doneprints:and the tab does not close. Reported on both kiro and claude, intermittently.
Cause: the capture fails silently at a different time, in a different command
claude-gh/bin/task-work:295-298(and the six sibling copies, regioninfo-tab-id):When the lookup returns empty, no
TAB_ID=line is written and nothing is printed. The block is not guarded on$ZELLIJeither — it runs unconditionally and fails quietly.So
task-worksucceeds, the worker runs for hours, and the first sign anything went wrong istask-done's message at teardown — by which time the context that caused it is gone. The user cannot act on the message because it reports a consequence in the wrong command.This is the same failure shape the repo has now fixed three times: #194 (an empty check list reads as green), #182 (a filter that never executed and failed open), #229 (two distinct failures that both wrote no log line). A state that is invisible because its failure mode produces no artifact.
Evidence
Auditing every
<worktree-base>/.<slug>.infoon this machine — 9 of 36 sidecars across 8 repos have noTAB_ID, and they are not randomly distributed:Two shapes stand out, both produced by
task-workitself:task-work:206-207appends a random 5-char hash on a name collision (SLUG="${SLUG}-${HASH}"). Sixadd-readmevariants exist, so that path ran repeatedly.task-work:136doesSLUG="${SLUG:0:50}".bug-model-renders-table-but-the-chat-already-outpuis exactly 50 characters.Both are cases where the name
task-workasks zellij for is one it derived through a transformation — which is where a mismatch between the slug and the actual tab name would live. Not proven; see below.What the lookup can return empty for
lib/zellij-tab.sh:38-52,aw_zellij_tab_id_by_name, ends in| .[0] // empty, so empty means no tab matched the name. Four ways to get there:$ZELLIJunset —task-workrun outside zellij, so no tab was created at all. The block does not check for this.jqonPATH—:42returns empty.$SLUGafter paint-stripping. Note the paint set (⏸️/▶️/❓︎) matches radio's exactly, so the documented race the comment addresses is genuinely covered — this would be a different mismatch, e.g. a long name zellij renders differently.I could not determine which of these fires from artifacts alone. The distribution above points at (3) and the two transformations that produce it, but the failing tabs are long gone and the message carries no diagnostic.
Fix
Two parts, and the first matters more than the second:
1. Make the failure loud, at the point it happens. When the lookup comes back empty,
task-workshould say so and say why it can tell — zellij absent, jq absent, or a name lookup that found nothing — and name the slug it searched for. Log it too, so the reconstruction is a grep rather than an audit of every sidecar on the machine. Whatever the trigger turns out to be, this is what makes the next report diagnosable in one line instead of this.2. Then find the trigger, with (1) in place to report it. Worth checking specifically whether zellij's
list-tabs --json.nameis byte-equal to the requested name for a 50-char slug, and whether anything in the collision-suffix path creates the tab under a different name than it writes to the sidecar.Do not fall back to the unscoped
zellij action close-tabintask-done— #107/#108 established that closes whichever tab is focused, andtask-done:201already explains why that is worse than skipping.Verification
task-workreports it (message + log line) and still completes — a failed tab-id capture must not abort a worker launch.<slug>.task-done's existing "no tab id captured" message gains a pointer to thetask-worklog line, so the two ends of the failure are linkable.info-tab-idregion, so the fix lands in every copy or the guard fails.References
TAB_IDis persisted in$INFO_FILEat all, and the paint-race the current comment addressestask-donemust not fall back to unscopedclose-tabTAB_IDcase (tab not closed, or wrong tab closed); this is the absent case, same observabletask-donewiped, which is what makes a failed close look like a failed unregister