- Added GoalRuntime with wall-clock and token accounting, budget steering, and lifecycle operations (create, pause, resume, drop, complete).
- Exposed goal tool as a hidden agent tool, activated only when goal mode is enabled.
- Integrated goal continuation loop in InteractiveMode with auto-submit between turns.
- Added status line segment and theme icons for goal mode state.
- Removed ExitPlanModeTool and deleted exit-plan-mode docs/tests, dropping the old approval contract outputs.
- Replaced plan-mode approval flow from exit_plan_mode to resolve across session, SDK, controllers, and discovery.
- Added standing resolve handler accessors and updated resolve routing for queued or standing approval handlers.
- Added PlanApprovalDetails and enforced normalized, validated approval titles with readable plan-file requirements.
- Extended resolve schema and invocation signatures with optional extra metadata and reason trimming behavior updates.
- Updated plan and resolve prompts and changelog guidance to require resolve action, reason, and extra.title for apply/discard.
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
ACP clients (Zed, etc.) only received `config_option_update` notifications
when they themselves drove the change via `session/set_session_config_option`.
Internal thinking-level updates (slash commands, automatic model-driven
adjustments, extension UI) bypassed the notification path, so client config
panels went stale until the next user-initiated change.
AgentSession now emits a `thinking_level_changed` event from
`setThinkingLevel`, and AcpAgent installs a session-lifetime subscription on
each managed session that pushes a fresh `config_option_update` whenever the
event fires — independent of prompt-turn lifecycle. The
`session/set_session_config_option` handler no longer pushes its own
notification for the `thinking` config (lifetime subscription covers it);
the response still returns fresh `configOptions` so callers see the new
state synchronously. Subscriptions are released in `#disposeSessionRecord`.
Also consolidated four duplicate `config_option_update` send sites into a
new `#pushConfigOptionUpdate(record)` helper.
Tests: added two cases to `test/acp-agent.test.ts` — one verifying internal
`setThinkingLevel` calls produce a `config_option_update` and a no-op
re-set produces none, and one verifying client-driven
`setSessionConfigOption(thinking, …)` produces exactly one notification.
Co-Authored-By: omp <noreply@oh-my-pi.dev>
- Removed local `abortableSleep` in favour of Node's built-in `scheduler.wait` from `node:timers/promises`.
- Consolidated per-provider retry/fetch loops into a shared `fetchWithRetry` utility in `packages/utils`.
- Moved `extractHttpStatusFromError`, `isRetryableError`, and related helpers out of `packages/ai` into `packages/utils`.
- Deleted `extractRetryDelay` in favour of `extractRetryHint` with unified header and body parsing.
Addresses the codex review comments on #1015 plus a sweep of adjacent
ACP conformance gaps surfaced while wiring them up.
Tool call + diff metadata
- acp-event-mapper: thread session cwd through and resolve every
`ToolCallLocation` (initial args, in-flight updates, result details)
to absolute paths against it; ACP requires absolute paths for
client-side file mapping.
- edit/modes/patch: emit the destination path for moves in the diff
result so post-edit "open file" actions land on the new file.
Permissions
- agent-session: pass cwd into `extractPermissionLocations` and resolve
raw `path`/`file`/etc. fields against it before sending
`session/request_permission`.
- agent-session: gate the permission wrapper on
`bridge.capabilities.requestPermission && bridge.requestPermission`,
matching the read/write/bash capability+method pattern.
acp-agent
- `authenticate`: validate `methodId` against the methods advertised by
`initialize` and reject anything else, so malformed clients fail fast.
- `setSessionConfigOption(MODE_CONFIG_ID)`: also emit
`current_mode_update` so clients tracking `modes.currentModeId` see
the same transition `session/set_mode` would produce.
- Pass `runtime.notifyConfigChanged` to builtins; emit
`available_commands_update` from a shared `reloadPlugins` helper
reused by `/reload-plugins`, `/marketplace`, and `/plugins`.
- prompt resource handling: route `resource` content with `image/*`
MIME into the `images` array instead of dropping it as an opaque
blob; non-image blobs still fall back to the URI placeholder.
- pass session cwd to the event mapper.
Builtins
- model: call `runtime.notifyConfigChanged()` after a successful
`setModel` so the ACP config selector reflects the new model
immediately.
- mcp: redact query strings and userinfo from MCP server URLs before
emitting them in `/mcp list` (prevents leaking `?exaApiKey=…` style
secrets); wire `manager.setAuthStorage(...)` before `prepareConfig`
in `/mcp test|resources|prompts` so OAuth servers can refresh tokens.
- ssh: reject non-integer `--port` values via a `^\d+$` guard instead
of silently coercing through `Number.parseInt`; list project hosts
first and dedupe user-scope duplicates to match capability-loader
precedence.
- export: reject clipboard aliases (`--copy`, `clipboard`, `copy`)
before passing them to `exportToHtml` as a filename.
- compact / force / move / browser: surface underlying failures via
`usage(errorMessage(...))` instead of letting them crash the command.
- session save|delete: route through the active SessionManager so the
persist writer is consulted and stale storage references are removed.
- marketplace / plugins / reload-plugins: call `runtime.reloadPlugins()`
on install/uninstall/upgrade and enable/disable so slash command
registries and command lists refresh consistently.
- shared.usage: make async and `await runtime.output(...)` so
`sessionUpdate` text is never dropped or reordered.
- types: document the new `reloadPlugins` and `notifyConfigChanged`
runtime hooks.
bash tool
- Use a shared `fireKill()` from the abort listener so `session/cancel`
terminates the remote command immediately instead of waiting for the
next `currentOutput()` round trip.
- Race `currentOutput()` against the abort signal so a stuck
`terminal/output` RPC cannot delay cancellation.
- Kill the terminal before reading final output on timeout so a slow
output read cannot let a timed-out command keep running past the
enforced timeout.
Tests
- acp-agent.test: extend the existing config-option assertions to
verify both `model` and `thinking_level` changes emit
`config_option_update` notifications scoped to the right session.
- acp-builtins.test: cover `/model` emitting both
`notifyTitleChanged` and `notifyConfigChanged`; lock in the parsed
`mcp add` / `ssh add` call shapes so future arg-parser regressions
fail the test instead of silently writing different configs; add a
`reloadPlugins` stub plus a typed `notifyConfigChanged` slot to the
shared test runtime factory.
- acp-stdout-hygiene.test: drain stderr in parallel and assert no
JSON-RPC frame leaks onto it; terminate the spawned process so the
stderr pump resolves deterministically.
CHANGELOG: itemize the above under `[Unreleased] > Fixed`.
CI
- bun run check: clean (TS + Rust)
- bun run test: 4128 pass / 689 skip / 0 fail (TS); 252 pass / 0 fail
(Rust nextest)
- bun run ci:test:smoke: --version / --help / `stats --help` all OK
- Introduces ClientBridge — the abstract boundary between AgentSession and external clients (ACP, TUI), defining terminal handle and permission request contracts
- Adds ACP permission gating in AgentSession for destructive tools (bash, edit, write, ast_edit): allow-once, reject-once, allow-always with caching
- Wires todo tracking, model cycling/retry-fallback chains, and auto-compaction into the session lifecycle
Introduce `CompactionCancelledError` and `CompactionOutcome` ("ok" |
"cancelled" | "failed") so callers can discriminate user-driven aborts
from generic failures via `instanceof`, instead of inspecting error
messages or `AbortError`-name strings.
`AgentSession.compact()`'s two abort-rejection sites now throw the
typed sentinel; the model-call wrapper normalizes AbortError-shaped
rejections to the sentinel only when the compaction's abort signal
is actually set, preserving every other exception unchanged so real
compaction bugs are not silently relabeled as cancellations.
`CommandController.executeCompaction` and `handleCompactCommand`
return `Promise<CompactionOutcome>`; the catch classifies via
`instanceof CompactionCancelledError`. Existing callers (`/compact`,
loop runner, auto-compact) ignore the return value — non-breaking.
Op: extend
- Replaced Python execution with a local `python -u runner.py` subprocess and NDJSON stdin/stdout framing.
- Removed shared-gateway architecture, including coordinator lifecycle APIs, `useSharedGateway` wiring, and `jupyter` CLI/actions.
- Simplified setup checks to a plain Python 3 availability probe and removed automatic dependency-install fallbacks.
- Updated kernel cancellation and display processing to use status frames, SIGINT/SIGTERM escalation, and normalized output coercion.
- Added `python-runner` integration and display tests while deleting legacy websocket and kernel lifecycle test suites.
- Added ownerId metadata to async jobs and to task/bash progress items from the session agent id.
- Extended async job registration and query methods with optional owner filters, and updated cancelAll to target matching owners.
- Updated session handoff and disposal so subagents inherit the parent manager, top-level sessions own it, and teardown cancels own jobs only.
- Added owner-aware async-job tests using hold/AbortSignal and scoped cancelAll assertions for running versus cancelled jobs.
- Added process-wide singleton instances for InternalUrlRouter, AsyncJobManager, and MCPManager.
- Changed internal URL protocols to resolve through registered sessions and scan all active roots/datasets for matches.
- Refactored agent, artifact, memory, rule, skill, jobs, and mcp handlers to use shared manager and rule/skill state.
- Removed per-session protocol/tool wiring and switched tests to initialize and reset global singleton state.
- Stored a failure counter during single-path edit execution and set isError on aggregate results when any entry edit failed.
- Set streaming-edit handling to always evaluate auto-generated-file checks, but only primed the file cache when edit.streamingAbort was enabled.
The 50ms retry loop in #scheduleBackgroundExchangeFlush had no guard
against session teardown. If dispose() was called while streaming,
the setTimeout callbacks kept firing and attempted emitExternalEvent
on a disconnected agent.
- Add #isDisposed flag; set it at the top of dispose() before any
async teardown so poll ticks that fire mid-teardown bail out cleanly
- Clear #pendingBackgroundExchanges in dispose() to drop queued
messages that will never be rendered
- Widen the attempt() bail condition to include #isDisposed so the
loop terminates even if new messages arrived between the dispose()
clear and the next tick
- Added macOS power assertion settings for idle, system, user, and display with schema defaults.
- Added idle, system, and user options to MacOSPowerAssertionOptions in TS typings and Rust, preserving display.
- Changed agent-session power flow to use begin/end/reset in-flight helpers and avoid manual counter updates.
- Updated power assertions to combine multiple kinds per native handle and support safe repeated-stop behavior.
- Updated unreleased changelog notes to document the breaking behavior and canceled-prompt unblock correction.
- Model resolver: provider-prefixed `<provider>/<id>` selectors are now
strict. If the provider is known and the exact pair does not resolve,
return undefined instead of silently crossing provider boundaries
(e.g. routing `anthropic/claude-3-7-sonnet` to amazon-bedrock when
the user only has Anthropic auth). Unqualified resolution is unchanged.
- Compaction: when the current model's provider has no credentials,
manual compaction now retries across compaction model candidates and
falls back to an authenticated role; if no usable fallback exists, it
throws a clear provider-specific pre-stream error instead of bubbling
a 503 `auth_unavailable` from the provider stream.
Fixes#986Fixes#980
Adds an opt-in onSseEvent callback across HTTP-streaming providers (Anthropic, OpenAI Responses/Completions, Azure OpenAI Responses, OpenAI Codex SSE, Google Gemini CLI, GitLab Duo, Kimi, Synthetic) so callers can inspect raw SSE frames without altering parsed output. Provider fetch wrapping only tees response bodies when an observer is wired; standalone packages/ai consumers without onSseEvent are not penalized.
Adds streamIdleTimeoutMs (env: PI_STREAM_IDLE_TIMEOUT_MS, with PI_OPENAI_STREAM_IDLE_TIMEOUT_MS as a backward-compatible alias). Anthropic now enforces a steady-state idle watchdog (default 120s) in addition to the first-event watchdog. OpenAI Responses, Azure Responses, and Codex (SSE + WebSocket) gain a semantic-progress predicate so response.in_progress-style keepalives no longer keep stalled tool calls alive forever.
Adds a coding-agent debug-panel raw SSE viewer backed by a per-session bounded buffer (1000 records / 512KB) that AgentSession populates unconditionally so users can post-hoc inspect a stuck stream from the TUI.
Real CC's getAPIMetadata includes device_id alongside session_id and
account_uuid. Rather than reading OS machine UUIDs (hardware fingerprinting)
or storing a random persistent ID, derive it as:
sha256("omp-device-id-v1:" + account_uuid).hex()
Properties:
- Indistinguishable from a randomly generated device ID on the wire
- Deterministic per account — survives reinstalls, no persistent storage
- Auditable: derived solely from the OAuth UUID already shared with Anthropic
- Zero hardware access, zero extra I/O
- Omitted for API-key callers (no account_uuid → no hash)
Anthropic counts sessions by metadata.user_id. Without this fix, OMP
generated fresh random entropy on every API request, inflating the
session count and preventing backend attribution to the authenticated
account.
Changes:
packages/ai:
- resolveAnthropicMetadataUserId() now accepts JSON-format user_id
matching real Claude Code's getAPIMetadata shape
({ session_id, account_uuid, ... }). Previously only the legacy
cloaking format was accepted on OAuth, causing stable caller-supplied
values to be silently discarded.
- AnthropicOAuthFlow.exchangeToken() and refreshAnthropicToken() now
populate OAuthCredentials.{accountId, email} from the token response
account block, removing the need for a separate /api/oauth/profile
round-trip.
- AuthStorage.getOAuthAccountId(provider, sessionId) returns the OAuth
accountId for the session-sticky credential, used to build
account_uuid in metadata.user_id. Guards against misattribution for
API-key, runtime-override, env-key, and fallback-resolver paths that
do not record a session credential.
packages/agent:
- Agent.metadataForProvider(provider) resolves request metadata for
the given provider via the installed resolver, or returns the static
metadata value. The plain metadata getter now returns only the static
value; provider-aware resolution is explicit.
- Agent.setMetadataResolver(fn) installs a (provider: string) resolver
evaluated per LLM request in agent-loop, after getApiKey records the
session-sticky credential, so account_uuid reflects the credential
actually used.
- AgentLoopConfig.metadataResolver is called with config.model.provider
after getApiKey, overriding the static metadata field.
packages/coding-agent:
- AgentSession.#syncAgentSessionId installs a metadata resolver that
builds { user_id: JSON.stringify({ session_id, account_uuid? }) },
matching the Anthropic session attribution format. account_uuid is
only included for provider="anthropic" to avoid leaking the OAuth
identity to third-party Anthropic-format-compatible providers.
- sessionId getter prefers providerSessionId when supplied via
AgentSessionConfig so all API paths (getApiKey, direct calls,
metadata resolver) share the same provider-facing session ID.
- prepareSimpleStreamOptions stamps session metadata on direct calls
(runEphemeralTurn, compaction, branch summary, title generation) so
they share the same session bucket as Agent.prompt requests.
- generateBranchSummary and generateSessionTitle accept a
(provider: string) metadata resolver evaluated after their own
getApiKey call for correct credential attribution.
- Added hideThinkingSummary options across stream, agent, and session payload paths.
- Routed Coding-Agent hideThinkingBlock toggles to agent hideThinkingSummary during session updates.
- Updated OpenAI, Azure OpenAI, and Codex requests to omit reasoning.summary when hide/ summary is null.
- Reworked system-prompt preparation with per-step timeouts, fallback defaults, and step-level warnings.
- Added optional `loadMode` and `summary` fields to `AgentTool` and related type declarations.
- Added `loadMode` and `summary` metadata to built-in tool classes for discoverable/essential behavior.
- Replaced `BUILTIN_TOOL_METADATA` with per-tool fields in discovery code paths.
- Updated `search_tool_bm25` and discovery indexing to use each tool's `summary` text.
- Updated discovery tests to validate tool `loadMode` and summary completeness.
Restore BUILTIN_TOOLS to Record<string, ToolFactory> so external SDK callers can
still invoke BUILTIN_TOOLS.read(session) directly, and move per-tool discovery
metadata (loadMode, summary) into a dedicated BUILTIN_TOOL_METADATA map. All
internal callers (computeEssentialBuiltinNames, getBuiltinDiscoverableEntries,
createTools, sdk.ts initial-tool filter, agent-session built-in collection) now
read metadata through the new map.
Restore the legacy MCP discovery API on AgentSession: getDiscoverableMCPTools()
returns DiscoverableMCPTool[] with description, and getDiscoverableMCPSearchIndex()
returns the legacy DiscoverableMCPSearchIndex whose documents expose
tool.description while remaining usable by searchDiscoverableTools (summary is
populated from description so the BM25 corpus still scores correctly). Generic
discovery via getDiscoverableTools / getDiscoverableToolSearchIndex is unchanged.
Centralize discovery cache invalidation in #invalidateDiscoveryCaches and call it
from #applyActiveToolsByName, refreshMCPTools, and refreshRpcHostTools so the
generic search index can no longer return tools that have already been activated
or registry entries that have been replaced.
Restrict #collectDiscoverableBuiltinTools to entries whose
BUILTIN_TOOL_METADATA[name].loadMode === "discoverable", which keeps hidden
tools (resolve, yield, exit_plan_mode, report_finding, report_tool_issue) and
unknown extension/custom registry entries out of the discovery corpus.
Tests: add coverage for callable BUILTIN_TOOLS factories, legacy MCP description
shape on getDiscoverableMCPTools / getDiscoverableMCPSearchIndex, stale-index
invalidation on setActiveToolsByName, and hidden-tool exclusion from
getDiscoverableTools({ source: "builtin" }).
The internal class is now SearchToolsTool but the wire name stays
"search_tool_bm25" so persisted MCP selections in user session files
continue to resolve correctly.
The BM25 corpus is now drawn from the generic DiscoverableTool list
(built-ins, MCP, extension, custom) rather than only MCP tools.
createIf fires when either tools.discoveryMode !== "off" or the
legacy mcp.discoveryMode is true.
AgentSession gains generic discovery methods (isToolDiscoveryEnabled,
getDiscoverableTools, getDiscoverableToolSearchIndex,
getSelectedDiscoveredToolNames, activateDiscoveredTools). The existing
MCP-named methods are kept as thin shims that filter by
source === "mcp" so existing callers continue to work.
Updated the tool's prompt copy to advertise discovery across all
sources rather than MCP only. Tests extended for the new shape.
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
- Added `HindsightSessionState` to `AgentSession` and bound hindsight lifecycle hooks to session state.
- Removed global hindsight state/queue handling and replaced it with per-session `HindsightRetainQueue` batching and scoped flushing.
- Reworked recall, reflect, and retain tools to use `session.getHindsightSessionState()` instead of sessionId-based lookup.
- Updated SDK/task/backend/controller flows to pass `session`/`parentHindsightSessionState`, scope `/memory` behavior, and document it in changelog.
- Added per-session retain queues with size/time auto-flush, recursive drain, and lifecycle flushes on end/clear/enqueue.
- Changed `hindsight-retain` to validate session state, enqueue writes, and return `Memory queued.` immediately.
- Added `notice` event support and handlers that route error/warning/status messages with source-aware formatting.
- Updated client and tests with shared request mapping, `RequestOptions`, `buildMemoryItem`, and expanded batch/list/doc APIs.
- Added inline hashline parse and apply support for `<` prepend and `+` append operations with prefix+suffix edits.
- Added fail-fast behavior to reject inline modify ops combined with delete or replace on same line.
- Renamed HASHLINE_* and mode symbols to HL_* in prompt tooling, read/search checks, and prompt templates.
- Standardized separators to `PI_HL_SEP`/`HL_EDIT_SEP` and fixed `HL_BODY_SEP='|'`, updating parser formatting behavior.
- Updated benchmark subtype constants and python cleanup test setup to use HL_* values and AgentRegistry mock failure injection.
- Added a new optional `beforeAgentStartPrompt` hook to `MemoryBackend` and implemented it in the Hindsight backend to recall long-term context for the first turn.
- Updated `AgentSession` startup flow to inject the recalled context into the turn-specific system prompt before the first response is generated.
- Preserved `<hindsight_memories>` tags in Hindsight developer instructions and added tests for first-turn injection and state caching.
- Added `memory.backend` and `hindsight.*` settings schema with migration from `memories.enabled` legacy mode.
- Added Hindsight memory backend runtime modules for resolved config, client creation, bank ID derivation, and state lifecycle.
- Added off/local/hindsight backends and resolver wiring across SDK, commands, and compaction context.
- Added `hindsight_recall`, `hindsight_reflect`, and `hindsight_retain` tools with schema validation and backend gating.
- Added Memory tab metadata and symbols to expose backend selection in the settings UI.
- Added package export barrels and tests for bank ID, content formatting, and hindsight config env precedence.
Manual handoff starts a fresh session and seeds it with a displayed custom handoff message, not an assistant message. Session persistence normally waits for an assistant message before creating the session file, which made the new handoff session exist only in memory until later activity.
Persist the seeded handoff session after injecting the handoff context, and record the previous session file as its parent so lineage remains discoverable.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
buildSystemPrompt injects today's date into the prompt body. The
applied-tool signature skips rebuilds when tools are byte-identical,
but did not cover the date — so a session spanning midnight with only
tool-stable MCP reconnects would keep yesterday's date indefinitely.
Append the current YYYY-MM-DD date as a suffix to the signature so any
reconnect after midnight triggers exactly one rebuild, then resumes
skipping normally for the rest of the new day.
Built-in tools whose prompt-rendered metadata depends on settings
(`TaskTool`, `SearchToolBm25Tool`, `EditTool`) expose `description`/
`label` via getters that re-evaluate on every access. The skip
optimization in `#applyActiveToolsByName` is correctness-safe for these
because `#computeAppliedToolSignature` reads `tool.description` live
each call, so a settings flip mutates the rendered string and differs
the signature on the next refresh.
This contract was implicit; a future refactor that caches per-tool
description strings would silently break it. Defending it explicitly:
- Added a regression test that wires a getter onto a CustomTool's
`description`, verifies `refreshMCPTools` skips while the underlying
state is unchanged, then mutates the state (without changing tool
object identity) and verifies the rebuild fires.
- Expanded the `#computeAppliedToolSignature` docstring to document
the getter-based coverage path and the SDK-init-time closure
constants in `sdk.ts` that genuinely cannot change at runtime
(`repeatToolDescriptions`, `eagerTasks`, `intentField`,
`mcpDiscoveryEnabled`, `secretsEnabled`).
Triggered by a review question on whether the skip breaks settings-
based prompt changes. It does not, but the property is non-obvious.