Commit Graph

16 Commits

Author SHA1 Message Date
can1357 2a6d551063 revert(status-line): restored single-row status bar with priority drop
- Reverted PR #5751 (issue #5749): continuation rows wrapped the editor
  top border onto extra lines, which is unacceptable for the input frame.
- EditorTopBorder is back to a single content/width pair; narrow widths
  drop right segments, shrink the path, then drop left segments.
2026-07-17 05:49:13 +02:00
roboomp 8ca790cb6f fix(status-line): wrapped overflow segments
Preserved configured segment priority by packing overflow into continuation rows.

Extended EditorTopBorder to render ordered rows inside the editor frame.

Fixes #5749
2026-07-16 20:19:54 +00:00
roboomp a52ed682c7 fix(ai): separated codex orchestration usage
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.

Fixes #4469
2026-07-03 16:44:12 +00:00
can1357 4e54e557bc feat: improved context usage display and update model configurations
- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
2026-06-22 08:04:42 +02:00
can1357 409196bf2e fix(coding-agent/session): fixed context usage breakdown to prefer completed in-turn anchors
- Fixed context breakdown to anchor estimates on the latest completed assistant usage message after compaction.
- Adjusted pending-context usage selection to prefer an in-turn provider anchor when available at/after cutoff.
- Added a contextUsageRevision cache token so status-line context memo invalidates after snapshot clear.
2026-06-17 12:24:20 +02:00
can1357 48decd15d7 fix(coding-agent): fixed context usage tracking to keep status and selector totals in sync
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
2026-06-17 12:24:20 +02:00
can1357 6d534a3c54 fix(agent): corrected agent event listener errors to structured warnings
- Replaced `Agent` listener error reporting with structured `logger.warn` calls.
2026-06-13 18:03:39 +02:00
can1357 495d49f588 refactor(coding-agent): reorganized status-line context cache usage flow
- Refactored status-line context caching to be keyed by message, tail, and window.
- Updated context usage flow to pass breakdown tokens and expose null usage when unknown.
- Adjusted context percentage handling so zero or unknown windows yield nullable values.
- Expanded status-line cache tests for usage provenance, memoization, and invalidation.
2026-06-13 18:02:08 +02:00
ben af98154b9e fix(coding-agent): cache status line context usage 2026-06-10 08:26:00 +02:00
can1357 9d457f73d9 test: migrated test imports to package subpath exports
- Replaced relative `../src` imports with `@oh-my-pi/pi-ai` and `@oh-my-pi/pi-agent-core` subpaths.
2026-06-08 19:03:55 +02:00
can1357 cc283cf50f feat(packages/coding-agent): added streaming append-only preview behavior
- Added `isStreamingPreviewAppendOnly` to `ToolRenderer` for per-tool streaming mode selection.
- Updated `ToolExecutionComponent` to query append-only predicates only while a call preview is streaming.
- Marked expanded write previews as append-only so over-tall streaming output can commit head rows.
- Threaded resolved status-line `segmentOptions` into `#buildSegmentContext` construction.
- Added regression tests for scrollback retention and append-only state transitions.
2026-06-06 22:58:33 +02:00
can1357 fde55bf927 fix(model): added bracket-affix stripping and string-keyed resolution cache
- Replaced WeakMap model cache with provider/id string keys for stable reuse.
- Returned official model ids directly when matched, before heuristics.
- Collapsed non-message token path to system prompt and tool schema totals.
2026-06-06 22:21:58 +02:00
can1357 96fd4c0f4f test(coding-agent): removed warm-cache refresh performance test case
- Removed the status-line context cache test that benchmarked warm-cache refreshes on a 200-message session.
2026-06-01 20:42:22 +02:00
can1357 d6acef8b76 perf(status-line): replaced index-based token cache with message sidecar cache
- Added Symbol-keyed sidecar on each AgentMessage to memoize estimateTokens, with a cheap content fingerprint to detect in-place mutations.
- Fixed stale cache on same-length replaceMessages, post-hoc error attachment, and branch rebuild edge cases.
- Fixed usage fetch error backoff: stamped fetchedAt on failure so the 5-min TTL also gates retries during outages.
- Extracted computeNonMessageBreakdown as shared helper to prevent drift between status-line and context panel token counts.
2026-05-25 14:06:05 +02:00
can1357 8e74996513 chore: adjust tests 2026-05-25 12:22:43 +02:00
Leo Kim 4eb2919674 fix(coding-agent): incremental per-message token cache for status-line context% (avoid 1.1s freeze on long sessions)
Root cause (verified on user's environment):

- User commit `296641213` swapped status-line's context% computation from cheap `calculatePromptTokens(lastAssistantMessage.usage)` to `computeContextBreakdown(session)`, which walks EVERY message and runs native `countTokens` (~0.5 ms per message).

- The 2-second TTL cache helps for steady-state idle but every cache MISS is a full sweep.

- `updateEditorTopBorder()` is invoked on EVERY agent event (event-controller.ts:163 — `agent_start`, `delta`, `agent_end`, `tool_*`). Each delta during streaming can trigger a cache miss.

- User session has 2,312 messages → each full sweep is ~1,120 ms blocking.

- During streaming the UI freezes for ~1.1 s every ~2 s, producing the user-visible 'jittery rendering' ("버벅거림") and 'status bar disappearing' symptoms.

Fix:

`StatusLineComponent.getCachedContextBreakdown()` (renamed from `#getCachedContextBreakdown` so unit tests can exercise it directly) now uses an incremental per-message token cache that exploits the append-only nature of `session.messages`:

  1. Message tokens (the dominant cost): cached per-index. New messages are tokenized as they arrive; previously-cached messages are reused. The LAST message is always recomputed because its content may still be growing during streaming. Compaction (messages.length shrinks) resets the cache.

  2. Non-message tokens (system prompt + tools + skills): cached separately, invalidated only when a cheap inputs-identity fingerprint changes (model swap, skill toggle, tool registration). These rarely change during a session.

Required exposing three helpers from `modes/utils/context-usage.ts` (`estimateSkillsTokens`, `estimateToolSchemaTokens`, `computeNonMessageTokens`) so the status-line cache can call them directly.

Performance (2,300-message synthetic session, measured on user's M-series Mac):

  - COLD warm-up call: ~75 ms (one-time, runs at OMP startup before any streaming)

  - WARM refresh, no new message: ~0.04 ms (20 calls = 0.7 ms total)

  - WARM refresh, 1 new message: ~0.02 ms

vs. prior implementation:

  - Per cache-miss call: ~1,120 ms blocking

  - 28,000× speedup on warm-state refresh

`computeContextBreakdown` itself is untouched — `/context` slash command continues to use it, and its output matches the status-line context% for the same session state (parity preserved).

Tests: 6 new cases in `packages/coding-agent/test/status-line-context-cache.test.ts` covering cold/warm/append/compaction/non-message-invalidation/zero-messages and a perf smoke test asserting 20 warm refreshes on a 200-message session complete in <100 ms.

Full suite: 3,199 tests, 26 pre-existing failures (status-line accent / log_experiment timing-flaky / skills / github tool / workspace-tree / tool path — all unrelated and baseline-confirmed). Lint: 1 pre-existing import-order issue in `event-controller-plan-ready.test.ts` unchanged.
2026-05-25 12:11:15 +02:00