- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
Models frequently write '+ ANCHOR| CONTENT' instead of '+ ANCHOR'
followed by '~CONTENT' on the next line. Add a compact FORMAT block
between </ops> and <rules> with a WRONG/RIGHT example that directly
addresses this failure mode.
Codex review flagged that the silent-abort sentinel
("__omp.silent_abort__") persists into AssistantMessage.errorMessage
but three downstream consumers render errorMessage verbatim:
- session-observer-overlay.ts: renders "✗ Error: __omp.silent_abort__"
when content is empty (confirmed user-visible today)
- print-mode.ts: writes marker to stderr and exits non-zero (latent;
plan-mode→compact not reachable from print mode today, but unguarded)
- acp-agent.ts: emits marker as agent_message_chunk text to ACP
clients when message has no other notifications (latent)
Add isSilentAbort() guard at each site. Extend the SILENT_ABORT_MARKER
consumer list in messages.ts doc comment to include all six consumers.
Add regression tests: overlay (2 tests), print-mode (2 tests), ACP
replay (1 test).
Op: correct
Restores: spec:silent-abort-marker-never-surfaces
When an MCP server uses OAuth Dynamic Client Registration (RFC 7591) and
no client_id is pre-configured, MCPOAuthFlow registers a fresh public
PKCE client on each authorize, captures the issued client_id into a
private field, then discards it once the flow object goes out of scope.
At refresh time, MCPManager#resolveAuthConfig calls refreshMCPOAuthToken
with auth.clientId from mcp.json — which is empty for these servers —
so providers that require client_id on the refresh grant (e.g. Linear at
mcp.linear.app/token) reject with HTTP 401 invalid_client. The user is
forced to /mcp reauth manually every time the access token expires.
This change threads the resolved/registered client credentials back out
of the OAuth flow and persists them into mcp.json so refresh has what
it needs indefinitely:
- MCPOAuthFlow exposes resolvedClientId / registeredClientSecret getters.
- MCPCommandController#handleOAuthFlow returns OAuthFlowResult with
credentialId + clientId + clientSecret, populated from the flow's
post-login state.
- The initial-connect non-wizard path and /mcp reauth path persist the
returned client credentials into both auth.{clientId,clientSecret}
(used at refresh) and oauth.{clientId,clientSecret} (used by future
/mcp reauth to skip re-registration).
- The wizard's onOAuth callback signature now returns the same shape;
#launchOAuthFlow folds the registered credentials into wizard state so
the final mcp.json entry built by #buildServerConfigWithAuth includes
them under auth.{clientId,clientSecret}.
Servers that configure a static oauth.clientId in mcp.json (Notion,
Slack, Datadog) are unaffected: #tryRegisterClient short-circuits, the
returned clientId equals the configured one, and the write-back is a
no-op.
Adds two MCPOAuthFlow unit tests covering both paths.
/simplify pass on 4882d1e38. Three small cleanups, no behavior change.
* Extracted the inline 50ms bootstrap-race guard into an exported
ACP_BOOTSTRAP_RACE_GUARD_MS constant at the top of acp-agent.ts.
Source uses it in the #scheduleBootstrapUpdates setTimeout. Tests import
it and call a new waitForBootstrapGuard() helper (constant + 30ms slack
for setTimeout drift) instead of three hardcoded Bun.sleep(80) sites —
tests now bind to the source-of-truth instead of dueling magic numbers.
* Consolidated four block comments that all narrated the same race story
into one canonical explanation at the install site (#scheduleBootstrapUpdates).
Field declaration, #registerPreparedSession, and the setSessionConfigOption
handler keep brief one/two-line pointers. Net change is roughly 30 lines
of comments removed without losing the diagnosis.
* Trimmed the handler-site thinkingHandledBySubscription comment from six
lines to three; the local-variable name carries the intent.
Verified:
* bun test test/acp-agent.test.ts: 11/11 pass
* biome check on touched files: clean
* No behavior change (no test had to be updated)
Co-Authored-By: omp <noreply@oh-my-pi.dev>
Addresses codex review on #1060: an extension session_start handler that
calls setThinkingLevel via the exposed extension action (line 1541) would
have run BEFORE #registerPreparedSession set the record into #sessions and
BEFORE the session/new response was delivered to the client, causing
config_option_update to be pushed for a session id the client did not yet
know about. This is the exact race that #scheduleBootstrapUpdates already
documents and guards for available_commands_update / session_info_update
(Zed's 'Received session notification for unknown session' drop).
Moved the session.subscribe(...) installation out of #registerPreparedSession
and into #scheduleBootstrapUpdates's 50ms timer callback so the lifetime
subscription shares the same response-delivery guard as the existing
bootstrap notifications. The pre-bootstrap thinking level is still
communicated to the client through the response payload's configOptions
(newSession / loadSession / resumeSession / unstable_forkSession all return
it), so no state is lost; it is only the notification that is deferred.
For client-driven setSessionConfigOption({thinking}) the handler now only
skips its own push when the lifetime subscription is already installed.
Pre-bootstrap the handler keeps pushing (the client knows the session id
because they passed it in), post-bootstrap the subscription pushes
exactly once. No double-push, no missing pre-bootstrap notification.
Tests:
- updated existing pushes-config-option-update test to await past the 50ms
bootstrap timer before driving the internal setThinkingLevel
- updated the single-config_option_update-per-setSessionConfigOption test
the same way
- added 'suppresses lifetime config_option_update during the bootstrap
window' regression that drives setThinkingLevel synchronously after
newSession and asserts zero notifications, then asserts notifications
resume after the bootstrap timer fires
- bun test test/acp-agent.test.ts: 11/11 pass
Co-Authored-By: omp <noreply@oh-my-pi.dev>
ACP clients (Zed, etc.) only received `config_option_update` notifications
when they themselves drove the change via `session/set_session_config_option`.
Internal thinking-level updates (slash commands, automatic model-driven
adjustments, extension UI) bypassed the notification path, so client config
panels went stale until the next user-initiated change.
AgentSession now emits a `thinking_level_changed` event from
`setThinkingLevel`, and AcpAgent installs a session-lifetime subscription on
each managed session that pushes a fresh `config_option_update` whenever the
event fires — independent of prompt-turn lifecycle. The
`session/set_session_config_option` handler no longer pushes its own
notification for the `thinking` config (lifetime subscription covers it);
the response still returns fresh `configOptions` so callers see the new
state synchronously. Subscriptions are released in `#disposeSessionRecord`.
Also consolidated four duplicate `config_option_update` send sites into a
new `#pushConfigOptionUpdate(record)` helper.
Tests: added two cases to `test/acp-agent.test.ts` — one verifying internal
`setThinkingLevel` calls produce a `config_option_update` and a no-op
re-set produces none, and one verifying client-driven
`setSessionConfigOption(thinking, …)` produces exactly one notification.
Co-Authored-By: omp <noreply@oh-my-pi.dev>
- Restored formatDimensionNote bracket form '[Image: original WxH, displayed at WxH. Multiply coordinates by S to map to original image.]' that tests assert. The Bun 1.3.14 refactor regressed it to a less informative 'Image resized from …' line.
- Broadened bash-sixel-render multi-line styling assertion to accept both truecolor (38;2;) and 256-color (38;5;) SGR runs so CI runners with TERM=dumb don't fail. The contract being tested — every line carries its own SGR — is independent of color depth.
The async job completion path copied durationMs, tokens, and extractedToolData
from SingleResult back into the AgentProgress object, but missed cost.
Add progress.cost = singleResult?.usage?.cost.total ?? 0 to the same block.
Eliminates the three-way duplication of toolCount/tokens/cost stat
appending across renderAgentProgress (running), renderAgentProgress
(completed), and renderAgentResult.
renderAgentResult (rendered after the task tool resolves) was missing the
cost display present in renderAgentProgress. Read cost directly from
result.usage?.cost.total which is already available on SingleResult.
- Updated @inquirer lockfile entries to newer releases for core, prompts, and related prompt packages, including updated dependency ranges and integrity hashes.
- Expanded minimumReleaseAgeExcludes in bunfig.toml by adding bun-types to the excluded package list.
Token counter (token_total status-line segment, subagent progress tree,
session-observer stats line) previously included cacheRead in its cumulative
sum. With Anthropic prompt caching, cacheRead per turn equals the full cached
context, so summing across N turns gives N*context_size -- a session with a 1M
context and 5 turns showed ~5M tokens despite no compaction occurring.
Fix: display shows input + output + cacheWrite per turn. cacheWrite is kept
because each byte is written once; cacheRead re-reads the same context every
turn. Dedicated cache_read/cache_write status-line segments still show cache
activity; billing cost is unaffected.
Also adds per-subagent cost display (dollar amount, statusLineCost color)
accumulated incrementally from message_end events. Hidden when cost is zero
(subscription/OAuth providers). Brings token and cost display in line with
what Claude Code shows per-agent.
- Updated BashTool's leading `cd` regex to stop matching newline characters so cwd extraction only applies to a single-line `cd ... &&` prefix.
- Added a regression test for multiline commands with a later-line `&&` to ensure each line of the script executes normally.
Keep the active page stealth setup synchronous, but make the broader CDP target UA override sweep selective and best-effort. Non-page or ephemeral Chrome targets can otherwise block worker initialization long enough for browser.open to hit the tool timeout before the tab worker sends ready.
Fixes#1053
Forward worker error and messageerror events while acquireTab waits for the initial ready/init-failed response. This prevents async worker module-load or early startup failures from being reported only as a generic tab worker init timeout.
- Added a new formatBashCommandLines helper that syntax-highlighted each command line and applied the dim prefix only to the first line.
- Updated the shell renderer to emit command output as line-based entries instead of a single dimmed string.
- Extended the bash renderer test to verify multi-line commands keep ANSI styling on every rendered line.
Op: correct
Restores: ref:44e5e0bb8 — queued /skill: chip lifecycle parity with plain-text steer
EventController.#handleMessageStart now mirrors the user-role refresh in
the custom branch, gated on readPendingDisplayTag(details). Without this,
AgentSession's tag-keyed dequeue mutated #steeringMessages /
#followUpMessages correctly but pendingMessagesContainer kept painting
the stale chip until an unrelated trigger (next user submit, dequeue key,
compaction flush) fired a refresh.
Non-queued custom variants (ttsr-injection, irc:*, async-result,
hookMessage) skip the refresh — they never registered a pending chip, so
rebuilding pendingMessagesContainer for them would be pure waste.
Pairs with the existing E4 (array splice) regression — the new E10 covers
the UI-refresh side of the same dequeue event with both positive and
negative gate assertions.
Co-Authored-By: chatgpt-codex-connector[bot] (P2 review on PR #1043)
- Removed local `abortableSleep` in favour of Node's built-in `scheduler.wait` from `node:timers/promises`.
- Consolidated per-provider retry/fetch loops into a shared `fetchWithRetry` utility in `packages/utils`.
- Moved `extractHttpStatusFromError`, `isRetryableError`, and related helpers out of `packages/ai` into `packages/utils`.
- Deleted `extractRetryDelay` in favour of `extractRetryHint` with unified header and body parsing.
When cached or freshly-discovered provider models carry UNK_CONTEXT_WINDOW
(222222) / UNK_MAX_TOKENS (8888) sentinels, #mergeResolvedModels was
replacing the bundled model wholesale — wiping out the correct values.
Switch to a field-level merge that preserves the bundled model's
contextWindow and maxTokens when the replacement only has sentinel
fallbacks. Custom models (via #mergeCustomModels) already had this
protection via ?? fallback; provider discoveries didn't.
Fixes the TUI showing 222222/8888 instead of the real context/token
limits for discovered models.
- Raised the Bun minimum version to >=1.3.14 across package metadata, install scripts, and changelog notes.
- Removed the Photon native image pipeline and added SIXEL-based `sixel` support in pi-natives.
- Migrated coding-agent image handling and resizing to `Bun.Image`, including updated tests and a JPEG quality bump to 80.
- Added HTTP/2 fetch bootstrap with HTTPS-only fallback and updated Bun build flags for autoload suppression/`--keep-names`.
The OutputSink now keeps a head budget (tools.artifactHeadBytes, default
20 KB) in addition to the tail spill window, so outputBytes can legally
reach head + tail + marker overhead. The multi-million line test still
asserted the pre-elision tail-only bound and started failing on CI.
- Changed multi-file search paging to skip whole files and page results in file windows.
- Added per-file match caps, round-robin file selection, and new file-limit truncation reporting.
- Replaced match/result limit metadata with fileLimitReached and perFileLimitReached.
- Lowered read.defaultLimit default to 300 with 1 lead and 3 trailing context lines.
- Replaced the search skip test with file-pagination coverage and added per-file cap tests.
- Added session-stats analytics tooling to classify searches, detect repeats, and render relevance plots.
- Updated read range expansion to use 1 leading and 3 trailing context lines.
- Changed read.defaultLimit from 500 to 300 in settings defaults.
- Updated read docs and tests to reflect the asymmetric context line behavior.
- Added read-selector analyzers and replay simulators to evaluate coverage and savings.
- Added plotting tools that output new session-stats PNG dashboards from local usage data.
- Added `getPriorityPremiumRequests` and used it in OpenAI and OpenAI Codex providers to increment `usage.premiumRequests` when `serviceTier: "priority"` was sent.
- Combined Copilot and priority premium counts in OpenAI completions and responses usage parsing so emitted usage uses the merged premium total.
- Updated stats parsing/backfill to re-read `service_tier_change` entries and UPSERT only `premium_requests`, so historical sessions pick up priority traffic without mutating other metrics.
- Added `tools.artifactHeadBytes` and `tools.outputMaxColumns` settings with defaults in `SETTINGS_SCHEMA`.
- Expanded `OutputSink` with `headBytes`/`maxColumns` and middle truncate logic with elision markers and tracking.
- Updated output-meta to resolve sink settings, emit truncation metrics, and use `truncateMiddle` for spills.
- Integrated head and column limits into JS/Python/Bash/SSH/read output flows, with `:raw` skipping read truncation.
- Documented new output middle-elision and column-cap behavior in `CHANGELOG.md`.
- Added truncation tests for `OutputSink`, `truncateMiddle`, and read-tool line handling.
Status-line's context_pct segment was computing tokens via
calculatePromptTokens(lastAssistantMessage.usage), which sums input +
cacheRead + cacheWrite from the Anthropic API usage object. The /context
slash command is computed by computeContextBreakdown, an offline estimate
over the live session state (systemPrompt + tools + skills + messages).
Both numbers are correct under their own definition, but they can
diverge by 2x+ on the same session when a turn rotates cache tiers
(e.g. 5m → 1h ephemeral re-cache) and cache_creation_input_tokens spikes.
Users read the two surfaces as one consistent dashboard and treat the
mismatch as a bug.
Repro: same session at the same moment reports 212K (21.2%) in /context
and 44.2%/1M in the status line — ~230K gap driven by per-turn
cache_creation on a system-prompt boundary.
This change makes status-line use the same computeContextBreakdown
source as /context so both surfaces stay consistent. The breakdown
result is cached with a 2s TTL inside the component so the per-frame
status-line render does not re-walk every message via
estimateMessagesTokens on long sessions. The Anthropic API per-turn
prompt size remains observable via existing token_in / cache_read /
cache_write / token_total segments.
- Updated the TUI shutdown slash command handler to return a `SlashCommandResult` instead of `void`.
- Returned `commandConsumed()` after clearing the editor and invoking runtime shutdown.
- Imported `SkillPromptDetails` as a type in the input controller message imports.
- Zed dispatches RPC responses and notifications on separate async tasks, so `setTimeout(0)` lost the race against the session registration handler.
- Dropped `available_commands_update` left the slash-command palette empty (#1015; zed-industries/zed#55965).
Collapses the parallel slash-commands/acp-builtins/ tree (28 files,
~1850 LOC) into one entry per command in builtin-registry.ts. Each
SlashCommandSpec has an optional 'handle' for the text-mode path used
by both TUI and ACP, and an optional 'handleTui' override for
selector/wizard UX. The two dispatchers (executeBuiltinSlashCommand
and executeAcpBuiltinSlashCommand) walk the same registry; the TUI
dispatcher synthesizes a SlashCommandRuntime from ctx when adapting
'handle', and the ACP dispatcher requires 'handle' and skips TUI-only
entries.
Helpers used by both dispatchers move under slash-commands/helpers/.
Behavior, dispatcher entry points, and test fixtures are unchanged;
all 62 ACP builtin tests and 19 TUI slash-command tests pass.
Addresses codex P2 review feedback on #1015:
- getMcpConfiguredServers (powers /mcp list, /mcp test, /mcp resources,
/mcp prompts): iterate project config first so when the same server
name is defined in both scopes the entry shown matches the one the
runtime actually loads. Matches discovery/builtin.ts where the loader
pushes project paths before user paths and capability dedupe is
first-wins.
- handleEnableDisableCommand (/mcp enable, /mcp disable): check the
project file before the user file so toggling a duplicated name flips
the effective entry. Previously '/mcp disable foo' reported success
while the active project foo stayed enabled.