Skip to content

[Bug]: compaction fallback nests the previous checkpoint indefinitely; 16k head-truncation turns it into silent permanent amnesia #1563

Description

@QinLuza

Describe the bug

When the compaction summarizer's LLM call fails, the deterministic fallback
merges the previous rolling checkpoint and the new fallback projection into
one text blob. The merge is purely additive and only flattens the outermost
marker, so every subsequent fallback compaction nests the entire previous
checkpoint one level deeper
(matryoshka effect).

That nested summary is later capped on injection by
_COMPACTION_SUMMARY_CONTEXT_MAX_CHARS = 16_000
(src/opensquilla/session/context_view.py:25), which retains the head.
The head of a nested summary is the oldest content, so after a handful of
fallback compactions the model sees only the earliest session topic — recent
state, tasks, and decisions are pushed past the cap and lost. The session is
not just compacted; it is permanently amnesiac with no visible error.

Evidence from a real session

Session agent:main:webchat:8db531b2, rows in session_summaries:

# trigger summary_source tokens_before tokens_after removed/kept summary_text chars
0 preflight llm 3,593,093 211,508 73/5 913
1 t3_upgrade fallback 803,518 153 21/1 132,186
2 preflight fallback 478,082 220,020 12/15 238,392
3 preflight fallback 638,021 352,223 25/7 482,796
4 preflight fallback 681,146 328,999 6/19 564,671
5 preflight fallback 604,941 151 33/1 247,753
6 preflight fallback 149,923 56 9/1 266,123
7 preflight fallback 272,138 52 4/1 296,332

Summarizer failure modes observed in the gateway debug log (one day, same
session, 11 failed calls):

  • provider output exceeded compaction token budget (compaction.py:1540)
  • provider returned an empty summary (compaction.py:1582)
  • HTTP 429 upstream rate limit on the shared chat provider

Any flaky or rate-limited summarization model can put a long session on this
path, and each subsequent compaction makes it strictly worse.

Root cause pointers

  1. src/opensquilla/session/compaction.py:1076 _merge_rolling_fallback —
    strips one [Deterministic rolling context] header from previous_summary
    then concatenates previous + new. Inner layers are never unwrapped:
    N fallbacks ⇒ N layers.
  2. src/opensquilla/session/compaction.py:~1000 — the budget-replan path
    deliberately refuses to shrink the previous checkpoint ("must either
    receive the previous checkpoint in full or decline this candidate").
    Combined with additive merging, the checkpoint can only grow.
  3. src/opensquilla/session/context_view.py:25 — the 16k injection cap with
    head retention makes the oldest bytes win against the newest state.

Expected behavior

  • A fallback compaction should produce a roughly bounded checkpoint that
    replaces the previous one — the previous summary is input to be
    re-digested, not an immutable payload to be prepended.
  • Repeated fallbacks must not monotonically grow summary_text.
  • When a checkpoint must be truncated for injection, retention should favor
    the newest state (or structured section priority such as
    current_plan_or_next_action), not the oldest head.
  • coverage_status should not report pass when checkpoint nesting/growth
    indicates content was effectively dropped.

Suggested direction

  1. Make _merge_rolling_fallback idempotent: recursively unwrap
    [Deterministic rolling context] layers before merging, and allow
    re-projection of the previous checkpoint under budget pressure instead of
    declining the candidate outright.
  2. On injection, keep tail/section priority so the freshest state survives a
    truncation.
  3. Consider a dedicated, pinned auxiliary summarization model for the
    compaction lane: a small model with a fixed, template-constrained output
    schema, decoupled from main chat routing and its shared quota. Most
    failure modes above (oversized output, empty output, 429 under a chat
    provider's rate limit) exist because compaction piggybacks on
    general-purpose chat deployments. A fixed-output auxiliary model would
    make the fallback path rare — and therefore survivable.

Happy to prototype (1) + (2) if maintainers agree on the direction.

Environment

  • OpenSquilla main (observed across 2026-08-19 … 2026-09-03), Windows
  • Compaction config: [compaction] model = "deepseek-v4-flash",
    provider = "tokenrhythm" (third-party, shared with chat traffic)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    has-linked-prAn open pull request is linked to this issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions