- packages/ai/test/issue-957-repro.test.ts now tests:
- refreshKimiToken applies the 5-minute server-side skew (Kimi-specific)
- AuthStorage refreshes kimi-code credentials inside its 60s skew window
- packages/ai/test/anthropic-stream-timeout.test.ts: raise the
streamFirstEventTimeoutMs from 10ms to 5000ms so slow CI scheduling
cannot fire the first-event watchdog before the mocked events arrive.
The test still exercises the (1ms) idle path it was written for.
fix(web): allow Parallel extract via PARALLEL_API_KEY env var without storage
The fetch tool and YouTube scraper previously gated the Parallel extract
branch behind `storage && findParallelApiKey(storage)`. With no
AgentStorage the env key was never consulted, so callers that ran
without a per-session storage (e.g. ReadTool sessions in unit tests, and
in practice any caller that has only an env API key) silently fell back
to raw-html / no-ytdlp paths.
- findCredential/findParallelApiKey now accept null or undefined storage
and rely solely on the env-first path when no storage is supplied.
- searchWithParallel/extractWithParallel mirror the same nullable shape.
- Drop the redundant `storage && ` guards in fetch.ts and youtube.ts;
the inner findParallelApiKey call already returns null when no
credential is available.
- Centralized OAuth access lifecycle in `AuthStorage`, returning identity metadata and new access-result types.
- Added 60-second skew and strict expiry checks, returning undefined/throws for stale or expired OAuth credentials.
- Removed provider-local token refresh flows from Gemini, Gemini CLI, Antigravity, Kimi, and related OAuth helpers.
- Migrated web-search providers from `AgentStorage` to `AuthStorage` session-aware lookup with `authStorage`/`sessionId`/`signal` flow.
- Replaced `findAnthropicAuth`/DB auth lookup with `buildAnthropicAuthConfig` and explicit base-url override/env fallback ordering.
`ref` is a JTD-reserved keyword (RFC 8927) used by the schema-reference
form, so the JTD-to-JSON-Schema converter on releases prior to 15.3.2
silently dropped it from the generated JSON Schema and required it at
the same time. Every explore-agent invocation then failed validation
with `schema_violation: files.0.ref: must not be present`.
The converter side was hardened in #1345 (shipped in 15.3.2). This
rename is defense-in-depth at the prompt level: the explore agent's
output contract no longer relies on the converter recognising a
user-named property that collides with a JTD keyword, and the field
name now matches what it actually carries.
Fixes#1379
- Added OpenAI Codex and Gemini web search provider options with updated setup/auth descriptions.
- Updated Codex OAuth flow to refresh near-expiry tokens during web_search and persist the refreshed credentials.
- Plumbed AgentStorage through search orchestrator, scrapers, and fetch paths so providers share session credentials.
- Refactored web provider and credential helpers to accept caller-provided AgentStorage and resolve keys synchronously.
Quarantined persistent session keys only while the native cancellation promise remains unsettled, so healthy cleanup restores persistent mode and stalled cleanup cannot accumulate live shell instances.
Added coverage for both stalled and settled native cleanup paths.
Fixes#1347
Queued extension-delivered user messages when deliverAs is set and waited for session_start extension message sends before prompting subagents.
Fixes#1343
Stopped marking persistent bash sessions as permanently broken when the JavaScript abort or timeout race wins.
Stopped the Rust descendant kill-wave helper once no cancellation targets remain so later commands are not swept into old cancels.
Fixes#1347
- Updated nested task-rendering tests to use parent-qualified IDs for completed child task results.
- Updated in-flight nested snapshot expectations to verify parent-aware `Parent>Subtask` labeling.
- Documented the live nested task rendering behavior in the package changelog.
- Captured `tool_execution_update` snapshots for `task` calls into in-flight progress state for live nested rendering.
- Cleared in-flight task snapshots at task start and completion to prevent stale nested progress from persisting.
- Updated progress rendering to combine completed and in-flight task details through a dedicated nested task tree view.
Raced bash execution against the JavaScript abort signal and timeout so the tool returns even when native shell cleanup does not settle.
Added regression coverage for native cleanup stalls on ESC abort and timeout.
Fixes#1347
The report_finding tool's priority is exposed as a string enum
("P0"-"P3") for ergonomics, but the reviewer agent and every
custom review agent declare priority as `type: number` in their
JTD output schema. The cast at executor.ts:1473 lied about the
runtime shape, so the auto-injected `findings[].priority` flowed
through as strings and every yield with at least one finding was
rejected with `findings.0.priority: expected number, received string`,
forcing the run into the schema_violation exit path.
Added `toReviewFinding(details)` in tools/review.ts that maps the
priority enum to its numeric ordinal via the existing PRIORITY_INFO
table and use it at the boundary in executor.ts. Render paths still
see the original `ReportFindingDetails` shape (string priority)
through normalizeReportFindings, so display formatting is unaffected.
Fixes#1350
- Extended `invalidateCredentialMatching` to accept session-scoped options and clear cached session credentials before blocking the matched credential.
- Updated the OAuth auth-error retry flow to pass `agent.sessionId` through credential invalidation.
- Added a regression test ensuring invalidating a session-sticky OAuth key rotates to the next active credential.
- Added `checkCredentials()` with result types/options for per-credential tri-state health checks.
- Added `/v1/credentials/check` endpoint via `handleCredentialsCheck` returning `{ generatedAt, credentials }`.
- Added `omp auth-gateway check` flow with provider grouping, `--json` output, and exit status 1 on failures.
- Added command examples, changelog updates, and tests for expired OAuth refresh, null/missing config, and ordering edge cases.
Loaded marketplace lspServers metadata from Claude plugin caches and embedded it for OMP marketplace installs so config-only plugins register without package code.
Fixes#1352
The JTD-to-JSON-Schema converter post-processed convertSchema's
output with normalizeMixedSchemaNode, which walked back into the
emitted JSON Schema looking for nested JTD forms. Inside a
properties block, user-defined property names whose keys happened
to collide with JTD keywords ('ref', 'elements', 'values',
'optionalProperties', 'discriminator') were misclassified as JTD
forms and re-rewritten - corrupting properties like { ref: { type:
'string' } } into { $ref: '#/$defs/[object Object]' } and breaking
the built-in explore agent's output validator with
schema_violation: files.0.ref: must not be present.
convertSchema is already fully recursive and emits pure JSON Schema,
so the post-walk is both unnecessary and unsafe. Drop it.
Fixes#1345
- Exported `normalizeTools` so `AppendOnlyContext` uses the same tool normalization as the agent loop.
- Added `BuildOptions.intentTracing` to `build()`/`reset()`/`takeSnapshot()` so intent injection is consistent and included in the prefix fingerprint.
- Improved `#computeDigest` to cover tool_calls, tool_call_id, name, and id fields to catch in-place mutations.
- Fixed `#unsubscribeAppendOnly` leak and added no-op guard in `#syncAppendOnlyContext`.
- Extracted `mergeDiscoveredModel` so discovered baseUrl takes priority over bundled entry, fixing 401s on Xiaomi tp- token-plan streams.
- User providerOverride.baseUrl still wins over both discovered and bundled values.
- Added regression tests covering all merge priority paths.
- Added Symbol-keyed sidecar on each AgentMessage to memoize estimateTokens, with a cheap content fingerprint to detect in-place mutations.
- Fixed stale cache on same-length replaceMessages, post-hoc error attachment, and branch rebuild edge cases.
- Fixed usage fetch error backoff: stamped fetchedAt on failure so the 5-min TTL also gates retries during outages.
- Extracted computeNonMessageBreakdown as shared helper to prevent drift between status-line and context panel token counts.
- Fixed /plan and /goal history preservation by snapshotting enabled state before handlePlanModeCommand/handleGoalModeCommand executes.
- Previous check read state after the call, missing cases where the handler itself toggled the mode off (e.g., confirmed exit).
- Added tests covering confirm-exit, cancel-exit, and first-activation paths.
- Raised PowerShell timeout to 8s and swallowed reap errors to prevent unhandled throws on WSL interop.
- Fixed fallback logic so arboard is skipped when no display server is present on headless WSL.
- Added test coverage for the headless WSL short-circuit path.
- Added `recoverOrphanedBackups` to promote `.jsonl..bak` files back to their primary path when the primary is missing, preventing data loss after a mid-rename crash.
- Changed backup filename from dot-prefixed to plain `..bak` so the shared `*.bak` glob can find it on both real and in-memory storage backends.
- Surfaced the original EPERM as the error `cause` and included both original and retry messages when rollback also fails.
- Replaced baked module-load value with per-call `isWebPExcluded()` so runtime env changes take effect.
- Only `"1"` and `"true"` (case-insensitive) enable exclusion; empty string and `"0"` are treated as disabled.
- Fast path now bypassed for WebP sources when exclusion is active.
- Explicit error surfaced when decode fails and WebP exclusion cannot be honored.
Anchors are formatted by read/search as LINE+HASH|TEXT, and lines may be
prefixed with marker decoration (*, >, +, -). The parser previously required
a bare LINE+HASH and rejected verbatim copy-pasted anchors with:
line N: expected a full anchor such as "119sr", ...; got "364sp|".
Loosen LID_CAPTURE_RE to allow optional leading decoration and an optional
trailing |... body on each anchor (including each side of a range).
- Added `retry.maxDelayMs` to the settings schema and interfaces, with a default cap for provider backoff delays.
- Updated session auto-retry logic to fail fast when a requested wait exceeds the cap without fallback, emitting terminal auto-retry failure state.
- Propagated retry state and failure data into task progress and rendering so children show retry/wait details and reminder prompts stop after terminal errors.
Separated model selector provider tab labels from provider ids so human-readable labels like Ollama Cloud refresh and filter the underlying ollama-cloud models.
Fixes#1153
In a compiled binary, Bun.resolveSync(spec, import.meta.dir) throws
'Cannot find module' because import.meta.dir is inside /$bunfs/root
and the virtual FS exposes no node_modules tree at runtime.
Previously this throw propagated through rewriteLegacyPiImports ->
rewriteLegacyPiImportsForRuntime -> mirrorLegacyPiFile ->
loadLegacyPiModule -> loadExtension, which swallowed it as 'Failed to
load extension' and silently dropped any plugin whose files imported
@mariozechner/pi-ai (or any @mariozechner/pi-* whose bundled
counterpart isn't reachable via resolveSync in the binary).
Fix: wrap the resolution call in rewriteLegacyPiImports in a try/catch
and return the original match on failure. rewriteBareImportsForLegacyExtension
runs immediately afterwards in every call path and already resolves bare
specifiers against the importer's real filesystem directory, so it picks
up @mariozechner/pi-ai from the plugin's installed peer deps instead.
Apply the same fallback to resolveLegacyPiSpecifier (the Bun plugin
shim's onResolve handler) for tool/hook files loaded directly via Bun's
import system rather than through loadLegacyPiModule.
Fixes#1215
- Removed the exported formatBashFixupNotice helper from bash command fixup utilities.
- Removed BashTool's one-time bash-fixup notice tracking and stopped emitting those notices when fixups were applied.
Root cause (verified on user's environment):
- User commit `296641213` swapped status-line's context% computation from cheap `calculatePromptTokens(lastAssistantMessage.usage)` to `computeContextBreakdown(session)`, which walks EVERY message and runs native `countTokens` (~0.5 ms per message).
- The 2-second TTL cache helps for steady-state idle but every cache MISS is a full sweep.
- `updateEditorTopBorder()` is invoked on EVERY agent event (event-controller.ts:163 — `agent_start`, `delta`, `agent_end`, `tool_*`). Each delta during streaming can trigger a cache miss.
- User session has 2,312 messages → each full sweep is ~1,120 ms blocking.
- During streaming the UI freezes for ~1.1 s every ~2 s, producing the user-visible 'jittery rendering' ("버벅거림") and 'status bar disappearing' symptoms.
Fix:
`StatusLineComponent.getCachedContextBreakdown()` (renamed from `#getCachedContextBreakdown` so unit tests can exercise it directly) now uses an incremental per-message token cache that exploits the append-only nature of `session.messages`:
1. Message tokens (the dominant cost): cached per-index. New messages are tokenized as they arrive; previously-cached messages are reused. The LAST message is always recomputed because its content may still be growing during streaming. Compaction (messages.length shrinks) resets the cache.
2. Non-message tokens (system prompt + tools + skills): cached separately, invalidated only when a cheap inputs-identity fingerprint changes (model swap, skill toggle, tool registration). These rarely change during a session.
Required exposing three helpers from `modes/utils/context-usage.ts` (`estimateSkillsTokens`, `estimateToolSchemaTokens`, `computeNonMessageTokens`) so the status-line cache can call them directly.
Performance (2,300-message synthetic session, measured on user's M-series Mac):
- COLD warm-up call: ~75 ms (one-time, runs at OMP startup before any streaming)
- WARM refresh, no new message: ~0.04 ms (20 calls = 0.7 ms total)
- WARM refresh, 1 new message: ~0.02 ms
vs. prior implementation:
- Per cache-miss call: ~1,120 ms blocking
- 28,000× speedup on warm-state refresh
`computeContextBreakdown` itself is untouched — `/context` slash command continues to use it, and its output matches the status-line context% for the same session state (parity preserved).
Tests: 6 new cases in `packages/coding-agent/test/status-line-context-cache.test.ts` covering cold/warm/append/compaction/non-message-invalidation/zero-messages and a perf smoke test asserting 20 warm refreshes on a 200-message session complete in <100 ms.
Full suite: 3,199 tests, 26 pre-existing failures (status-line accent / log_experiment timing-flaky / skills / github tool / workspace-tree / tool path — all unrelated and baseline-confirmed). Lint: 1 pre-existing import-order issue in `event-controller-plan-ready.test.ts` unchanged.
Status-line's context_pct segment was computing tokens via
calculatePromptTokens(lastAssistantMessage.usage), which sums input +
cacheRead + cacheWrite from the Anthropic API usage object. The /context
slash command is computed by computeContextBreakdown, an offline estimate
over the live session state (systemPrompt + tools + skills + messages).
Both numbers are correct under their own definition, but they can
diverge by 2x+ on the same session when a turn rotates cache tiers
(e.g. 5m → 1h ephemeral re-cache) and cache_creation_input_tokens spikes.
Users read the two surfaces as one consistent dashboard and treat the
mismatch as a bug.
Repro: same session at the same moment reports 212K (21.2%) in /context
and 44.2%/1M in the status line — ~230K gap driven by per-turn
cache_creation on a system-prompt boundary.
This change makes status-line use the same computeContextBreakdown
source as /context so both surfaces stay consistent. The breakdown
result is cached with a 2s TTL inside the component so the per-frame
status-line render does not re-walk every message via
estimateMessagesTokens on long sessions. The Anthropic API per-turn
prompt size remains observable via existing token_in / cache_read /
cache_write / token_total segments.
- New 'usage' status-line segment showing Anthropic 5h/7d quota
- Background refresh (5min TTL) via fetchUsageReports
- thinking.max symbol added to UNICODE_SYMBOLS
Personal patch consolidated into branch.
Bug: Ctrl+C on the ask tool selector threw ToolAbortError, the turn
ended with stopReason === "aborted", and handleBackgroundEvent fired
sendCompletionNotification() unconditionally — producing a misleading
"Task complete" desktop toast for a turn that never actually completed.
Fix mirrors the stopReason filter already used by
#currentContextTokens, #handleMessageEnd, and the retry / TTSR /
compaction skip paths across agent-session.ts: check the most recent
assistant message via session.getLastAssistantMessage() and return
early when stopReason is "aborted" or "error".
Test coverage (event-controller-abort-guard.test.ts, 6 cases):
- aborted -> 0 sendNotification calls
- error -> 0 calls
- stop -> 1 call (normal completion)
- no last assistant message -> proceeds (defensive)
- isBackgrounded=false (foreground) -> still 0
- completion.notify=off -> still 0
Matching guard applied to the standalone desktop-notify extension
(~/.omp/agent/extensions/desktop-notify/index.ts) which is currently
the live producer of completion toasts after Phase 1 of
seed_0ca7e1143ac1.