Commit Graph

3 Commits

Author SHA1 Message Date
can1357 d6acef8b76 perf(status-line): replaced index-based token cache with message sidecar cache
- Added Symbol-keyed sidecar on each AgentMessage to memoize estimateTokens, with a cheap content fingerprint to detect in-place mutations.
- Fixed stale cache on same-length replaceMessages, post-hoc error attachment, and branch rebuild edge cases.
- Fixed usage fetch error backoff: stamped fetchedAt on failure so the 5-min TTL also gates retries during outages.
- Extracted computeNonMessageBreakdown as shared helper to prevent drift between status-line and context panel token counts.
2026-05-25 14:06:05 +02:00
can1357 8e74996513 chore: adjust tests 2026-05-25 12:22:43 +02:00
Leo Kim 4eb2919674 fix(coding-agent): incremental per-message token cache for status-line context% (avoid 1.1s freeze on long sessions)
Root cause (verified on user's environment):

- User commit `296641213` swapped status-line's context% computation from cheap `calculatePromptTokens(lastAssistantMessage.usage)` to `computeContextBreakdown(session)`, which walks EVERY message and runs native `countTokens` (~0.5 ms per message).

- The 2-second TTL cache helps for steady-state idle but every cache MISS is a full sweep.

- `updateEditorTopBorder()` is invoked on EVERY agent event (event-controller.ts:163 — `agent_start`, `delta`, `agent_end`, `tool_*`). Each delta during streaming can trigger a cache miss.

- User session has 2,312 messages → each full sweep is ~1,120 ms blocking.

- During streaming the UI freezes for ~1.1 s every ~2 s, producing the user-visible 'jittery rendering' ("버벅거림") and 'status bar disappearing' symptoms.

Fix:

`StatusLineComponent.getCachedContextBreakdown()` (renamed from `#getCachedContextBreakdown` so unit tests can exercise it directly) now uses an incremental per-message token cache that exploits the append-only nature of `session.messages`:

  1. Message tokens (the dominant cost): cached per-index. New messages are tokenized as they arrive; previously-cached messages are reused. The LAST message is always recomputed because its content may still be growing during streaming. Compaction (messages.length shrinks) resets the cache.

  2. Non-message tokens (system prompt + tools + skills): cached separately, invalidated only when a cheap inputs-identity fingerprint changes (model swap, skill toggle, tool registration). These rarely change during a session.

Required exposing three helpers from `modes/utils/context-usage.ts` (`estimateSkillsTokens`, `estimateToolSchemaTokens`, `computeNonMessageTokens`) so the status-line cache can call them directly.

Performance (2,300-message synthetic session, measured on user's M-series Mac):

  - COLD warm-up call: ~75 ms (one-time, runs at OMP startup before any streaming)

  - WARM refresh, no new message: ~0.04 ms (20 calls = 0.7 ms total)

  - WARM refresh, 1 new message: ~0.02 ms

vs. prior implementation:

  - Per cache-miss call: ~1,120 ms blocking

  - 28,000× speedup on warm-state refresh

`computeContextBreakdown` itself is untouched — `/context` slash command continues to use it, and its output matches the status-line context% for the same session state (parity preserved).

Tests: 6 new cases in `packages/coding-agent/test/status-line-context-cache.test.ts` covering cold/warm/append/compaction/non-message-invalidation/zero-messages and a perf smoke test asserting 20 warm refreshes on a 200-message session complete in <100 ms.

Full suite: 3,199 tests, 26 pre-existing failures (status-line accent / log_experiment timing-flaky / skills / github tool / workspace-tree / tool path — all unrelated and baseline-confirmed). Lint: 1 pre-existing import-order issue in `event-controller-plan-ready.test.ts` unchanged.
2026-05-25 12:11:15 +02:00