Describe the bug
When the compaction summarizer's LLM call fails, the deterministic fallback
merges the previous rolling checkpoint and the new fallback projection into
one text blob. The merge is purely additive and only flattens the outermost
marker, so every subsequent fallback compaction nests the entire previous
checkpoint one level deeper (matryoshka effect).
That nested summary is later capped on injection by
_COMPACTION_SUMMARY_CONTEXT_MAX_CHARS = 16_000
(src/opensquilla/session/context_view.py:25), which retains the head.
The head of a nested summary is the oldest content, so after a handful of
fallback compactions the model sees only the earliest session topic — recent
state, tasks, and decisions are pushed past the cap and lost. The session is
not just compacted; it is permanently amnesiac with no visible error.
Evidence from a real session
Session agent:main:webchat:8db531b2, rows in session_summaries:
| # |
trigger |
summary_source |
tokens_before |
tokens_after |
removed/kept |
summary_text chars |
| 0 |
preflight |
llm |
3,593,093 |
211,508 |
73/5 |
913 |
| 1 |
t3_upgrade |
fallback |
803,518 |
153 |
21/1 |
132,186 |
| 2 |
preflight |
fallback |
478,082 |
220,020 |
12/15 |
238,392 |
| 3 |
preflight |
fallback |
638,021 |
352,223 |
25/7 |
482,796 |
| 4 |
preflight |
fallback |
681,146 |
328,999 |
6/19 |
564,671 |
| 5 |
preflight |
fallback |
604,941 |
151 |
33/1 |
247,753 |
| 6 |
preflight |
fallback |
149,923 |
56 |
9/1 |
266,123 |
| 7 |
preflight |
fallback |
272,138 |
52 |
4/1 |
296,332 |
Summarizer failure modes observed in the gateway debug log (one day, same
session, 11 failed calls):
provider output exceeded compaction token budget (compaction.py:1540)
provider returned an empty summary (compaction.py:1582)
- HTTP 429 upstream rate limit on the shared chat provider
Any flaky or rate-limited summarization model can put a long session on this
path, and each subsequent compaction makes it strictly worse.
Root cause pointers
src/opensquilla/session/compaction.py:1076 _merge_rolling_fallback —
strips one [Deterministic rolling context] header from previous_summary
then concatenates previous + new. Inner layers are never unwrapped:
N fallbacks ⇒ N layers.
src/opensquilla/session/compaction.py:~1000 — the budget-replan path
deliberately refuses to shrink the previous checkpoint ("must either
receive the previous checkpoint in full or decline this candidate").
Combined with additive merging, the checkpoint can only grow.
src/opensquilla/session/context_view.py:25 — the 16k injection cap with
head retention makes the oldest bytes win against the newest state.
Expected behavior
- A fallback compaction should produce a roughly bounded checkpoint that
replaces the previous one — the previous summary is input to be
re-digested, not an immutable payload to be prepended.
- Repeated fallbacks must not monotonically grow
summary_text.
- When a checkpoint must be truncated for injection, retention should favor
the newest state (or structured section priority such as
current_plan_or_next_action), not the oldest head.
coverage_status should not report pass when checkpoint nesting/growth
indicates content was effectively dropped.
Suggested direction
- Make
_merge_rolling_fallback idempotent: recursively unwrap
[Deterministic rolling context] layers before merging, and allow
re-projection of the previous checkpoint under budget pressure instead of
declining the candidate outright.
- On injection, keep tail/section priority so the freshest state survives a
truncation.
- Consider a dedicated, pinned auxiliary summarization model for the
compaction lane: a small model with a fixed, template-constrained output
schema, decoupled from main chat routing and its shared quota. Most
failure modes above (oversized output, empty output, 429 under a chat
provider's rate limit) exist because compaction piggybacks on
general-purpose chat deployments. A fixed-output auxiliary model would
make the fallback path rare — and therefore survivable.
Happy to prototype (1) + (2) if maintainers agree on the direction.
Environment
- OpenSquilla
main (observed across 2026-08-19 … 2026-09-03), Windows
- Compaction config:
[compaction] model = "deepseek-v4-flash",
provider = "tokenrhythm" (third-party, shared with chat traffic)
Describe the bug
When the compaction summarizer's LLM call fails, the deterministic fallback
merges the previous rolling checkpoint and the new fallback projection into
one text blob. The merge is purely additive and only flattens the outermost
marker, so every subsequent fallback compaction nests the entire previous
checkpoint one level deeper (matryoshka effect).
That nested summary is later capped on injection by
_COMPACTION_SUMMARY_CONTEXT_MAX_CHARS = 16_000(
src/opensquilla/session/context_view.py:25), which retains the head.The head of a nested summary is the oldest content, so after a handful of
fallback compactions the model sees only the earliest session topic — recent
state, tasks, and decisions are pushed past the cap and lost. The session is
not just compacted; it is permanently amnesiac with no visible error.
Evidence from a real session
Session
agent:main:webchat:8db531b2, rows insession_summaries:summary_textgrows monotonically to ~300k charswhile
tokens_aftercollapses to 151 / 56 / 52 andkept=1.[Deterministic rolling context]/[Structured Compaction Summary]headers followed by the oldest content from Aug 19. The newest progress
lives at the tail — beyond the 16k injection cap, i.e. invisible to the model.
coverage_statusstayedpassfor all eight compactions. The failure is silent.Summarizer failure modes observed in the gateway debug log (one day, same
session, 11 failed calls):
provider output exceeded compaction token budget(compaction.py:1540)provider returned an empty summary(compaction.py:1582)Any flaky or rate-limited summarization model can put a long session on this
path, and each subsequent compaction makes it strictly worse.
Root cause pointers
src/opensquilla/session/compaction.py:1076_merge_rolling_fallback—strips one
[Deterministic rolling context]header fromprevious_summarythen concatenates previous + new. Inner layers are never unwrapped:
N fallbacks ⇒ N layers.
src/opensquilla/session/compaction.py:~1000— the budget-replan pathdeliberately refuses to shrink the previous checkpoint ("must either
receive the previous checkpoint in full or decline this candidate").
Combined with additive merging, the checkpoint can only grow.
src/opensquilla/session/context_view.py:25— the 16k injection cap withhead retention makes the oldest bytes win against the newest state.
Expected behavior
replaces the previous one — the previous summary is input to be
re-digested, not an immutable payload to be prepended.
summary_text.the newest state (or structured section priority such as
current_plan_or_next_action), not the oldest head.coverage_statusshould not reportpasswhen checkpoint nesting/growthindicates content was effectively dropped.
Suggested direction
_merge_rolling_fallbackidempotent: recursively unwrap[Deterministic rolling context]layers before merging, and allowre-projection of the previous checkpoint under budget pressure instead of
declining the candidate outright.
truncation.
compaction lane: a small model with a fixed, template-constrained output
schema, decoupled from main chat routing and its shared quota. Most
failure modes above (oversized output, empty output, 429 under a chat
provider's rate limit) exist because compaction piggybacks on
general-purpose chat deployments. A fixed-output auxiliary model would
make the fallback path rare — and therefore survivable.
Happy to prototype (1) + (2) if maintainers agree on the direction.
Environment
main(observed across 2026-08-19 … 2026-09-03), Windows[compaction] model = "deepseek-v4-flash",provider = "tokenrhythm"(third-party, shared with chat traffic)