- Replaced legacy `compaction.strategy` and `remoteEnabled` settings with `compaction.methodOrder` across session maintenance, schema, and tests.
- Added automatic fallback mechanism to try subsequent compaction methods upon failure or unsupported model capabilities.
- Added mouse drag-and-drop reordering support and click handlers to multi-select settings submenus.
- Updated documentation and test suites to reflect ordered compaction strategy preferences and fallback chains.
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Updated eager compaction and plan reference tests to track call indices and task delegation markers instead of text strings.
- Removed obsolete context message marker checks, vibe mode assertions, and prompt gating test cases.
- Simplified prewalk, workflow, and Gemini instruction test expectations across agent modules.
- Removed the system prompt personality test suite entirely.
Post-compaction auto-continuation is gated on remaining work since
#5721; an active goal keeps the continuation vehicle these #1246
regressions ride on.
- Normalized completed image_generation_call results into assistant image blocks.
- Persisted image bytes through the session blob store and rendered them in live, replay, ACP, proxy, telemetry, and HTML paths.
- Added response normalization, persistence, and TUI rendering regressions.
Fixes#4768
claude-sonnet-4-5's bundled context window grew to 1M in the catalog regen, so the fixed 191k high-usage turn no longer crossed the ~85% auto-compaction threshold. The auto-continuation never fired and the three re-injection tests hung to their 5s timeout, failing the coding-agent runtime/session CI bucket and blocking the release.
Pin the harness model to a 200k window (mirrors agent-session-eager-compaction / -auto-compaction-queue), keeping the trigger stable across future catalog regenerations.
- Unify `previousText` resolution to correctly concatenate `textHead` and `textTail` during re-compaction.
- Ensure summary fallback logic correctly handles non-text legacy archives.
- Add test coverage for cross-compaction text retention and legacy archive continuity.
- Enable snapcompact strategy in agent plan reference re-injection tests.
- Privatized the legacy `nextToolChoice` method to `#nextHardToolChoice` to ensure all tool-choice directives flow through the unified `nextToolChoiceDirective` entry point.
- Eliminated redundant dual entry points for fetching tool choices, which previously bypassed the soft pending-preview lifecycle.
- Updated test suites to consume `nextToolChoiceDirective` where appropriate to maintain consistency with internal agent-loop logic.
After plan approval the executor delivers the plan-mode-reference exactly
once and sets `#planReferenceSent = true`. Both compaction paths — `compact()`
and `#runAutoCompaction()` — replace the conversation history that carried
that reference but never cleared the flag, so `#buildPlanReferenceMessage()`
short-circuited to null on every subsequent turn and the executor permanently
lost the plan it was working on (exactly the long-session failure reported).
Clear `#planReferenceSent` right after `replaceMessages()` in both paths so the
next turn re-reads the plan from disk and re-injects it. The reset is a no-op
for ordinary sessions: the default plan path (PLAN.md in session-local scratch)
has no file on disk, so `#buildPlanReferenceMessage()` still returns null there.
Adds a deterministic regression test (short-circuited compaction, mock stream)
that fails before this change and passes after, plus a guard proving normal
sessions get no spurious plan injection.
Fixes#1246