PR #3648 review (codex): the first iteration emitted every envelope
path alongside one combined digest, so a multi-file payload that added
`: any` to a README.md hunk and merely touched src/ok.ts would surface
a *.ts path in the TTSR match context and trip the bundled
tool:edit(*.ts) ts-no-any rule on text that belonged to the Markdown
hunk — aborting valid edits under interruptMode:always.
Add a per-file matcherEntries(args) hook on AgentTool / EditStreaming-
Strategy returning [{ path, digest }] entries, one per touched file
(same-path sections/hunks merged):
- replace / patch: one entry from the top-level path + matcherDigest
- hashline: regex-split by [path#TAG] section, body added-lines per
entry (tolerant of streaming partial payloads)
- apply_patch: expandApplyPatchToPreviewEntries grouped by path
AgentSession.#checkTtsrStream / #checkTtsrAstStream now prefer
matcherEntries and iterate per-file with isolated filePaths + streamKey,
so each file's buffer and repeat-tracking are independent. Tools
without matcherEntries keep the existing combined matcherDigest +
matcherPaths path.
- Standardized retry logic to ensure mocked assistant turns and generic aborts share the same reset path as live provider failures.
- Introduced a classification helper that forces retry policy alignment with the currently active session model.
- Enabled retries for generic reason-less abort sentinels and stale OpenAI response errors.
AgentSession's TTSR match context only scanned top-level path/paths
arguments, so hashline and apply_patch edit streams (whose only target
path lives inside the wire payload — section headers or envelope
markers — not as a top-level argument) arrived without any filePaths
and silently skipped path-scoped rules like the bundled ts-no-any
(scope: tool:edit(*.ts)).
Add an optional AgentTool.matcherPaths(args) hook, companion to the
existing matcherDigest(args), so tools whose wire grammar embeds paths
can surface them. Implement on each edit streaming strategy:
- replace / patch: top-level path
- hashline: parse [path#TAG] (and tag-less [path]) section headers
tolerant of streaming partial payloads
- apply_patch: parse *** Add/Update/Delete File: markers, also tolerant
of pre-End-Patch buffers
AgentSession.#getTtsrToolMatchContext consults tool.matcherPaths first,
normalising its output through the existing path-candidate helper, and
falls back to the generic top-level argument scan for tools that don't
implement it.
Fixes#3646
- Replaced functional streaming logic with a stateful `CodexStreamProcessor` class.
- Converted `CodexStreamRuntime` into a class with explicit state initializers and lifecycle methods.
- Encapsulated tool call delta tracking and interruption logic within dedicated class methods.
- Integrated `textSignature` generation into text-type output blocks using versioned encoders.
- Added textVerbosity configuration schema to allow user control over model response length.
- Integrated textVerbosity into settings-aware stream processing for "openai-responses" and "openai-codex-responses" APIs.
- Updated settings stream tests to verify that verbosity settings correctly apply while respecting caller overrides.
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
The CI gate failed in several non-overlapping ways once the full coding-agent suite ran here: tests wrote into the real $HOME (`/srv/agent-home`) which is read-only, a fixture rotated stored Anthropic API keys but Settings reloaded the user models.yml and shadowed them, an OAuth callback server bound to `hostname:"localhost"` (loopback unreachable in this runtime), a built-in tool metadata assertion only saw `github` when `gh` was installed, and the GithubTool `pr_checkout` worktree assertions assumed `~/.omp/wt` but `XDG_DATA_HOME` redirected `getWorktreesDir()` to `$XDG_DATA_HOME/omp/wt`.
Fixes:
- `packages/ai/src/registry/oauth/callback-server.ts`: drop `hostname: "localhost"` when no caller-supplied hostname overrides it. Bun on Linux refused inbound connections to the listener when bound explicitly to `localhost`; defaulting to Bun.serves default binding restores loopback connectivity.
- `packages/coding-agent/test/status-line-path.test.ts`: route the `~/Projects` fixtures through a writable temp home (spy `os.homedir()`), housed under the repo `.wt/` worktree scratch so the temp home is not classified as a status-line scratch root.
- `packages/coding-agent/test/skills.test.ts`: ditto for the `~/.pi-skills-test-*` mkdtemp in the tilde-expansion test.
- `packages/coding-agent/test/marketplace/project-scope.test.ts`: stop writing `~/.git`; build the entire home-dir guard fixture in a temp dir and spy `os.homedir()`.
- `packages/coding-agent/test/oauth-flow.test.ts`: wrap each callback fetch in a brief retry so the simulated browser redirect tolerates the few-ms gap before the Bun callback server starts accepting connections.
- `packages/coding-agent/test/tools/gh.test.ts`: extend `setupTempHome()` to clear `XDG_DATA_HOME`/`XDG_STATE_HOME`/`XDG_CACHE_HOME` for the duration of the test so the rebuilt dirs resolver routes `getWorktreesDir()` back through the spied home, then restore them on cleanup.
- `packages/coding-agent/test/tool-discovery/initial-tools.test.ts`: instantiate `GithubTool` directly in the metadata fixture so the assertion runs even when `gh` is unavailable (GithubTool.createIf returns null without `gh`).
- `packages/coding-agent/test/agent-session-retry-cap.test.ts`: pass an isolated `models.yml` path to `ModelRegistry` so the two-Anthropic-key fixture is the authoritative credential source instead of any user-level command-backed Anthropic key.
Verification:
- `bun check` → passed
- `bun run test` (full coding-agent suite, 4 buckets, all chunks) → 0 fails
Fixes#3639
Agent.#runLoop appends the user batch and a synthetic stopReason: "error" assistant turn to state.messages before resolving prompt() with state.error set. AdvisorRuntime now snapshots state.messages.length before each prompt, restores it on failure (via a new AdvisorAgent.rollbackTo hook that also resets the advisor's append-only sync cursor), and clears state.error so retries replay a clean baseline and the drop-after-3 path never leaks orphan failed turns into the next successful run's context.
Fixes#3635
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
Removed the remaining callers of StatusLineComponent.setSubagentHubHint after the status-line subagent hub hint display was simplified away. The status-line cache test already asserts the hub hint is absent, so the stale setter call only broke type checking.
Fixes#3639
AgentSession.#buildAdvisorRuntime constructed the advisor Agent without the provider-shaping options the SDK installs on the main agent: the streamFn wrapper that applies providers.openrouterVariant / providers.antigravityEndpoint / providers.maxInFlightRequests / model.loopGuard.*, the onPayload/onResponse/onSseEvent hooks, the shared providerSessionState map, transformProviderContext (snapcompact, secret obfuscation, image clamping), and a stable promptCacheKey. Advisor turns therefore dropped the OpenRouter sticky-routing variant suffix, used a different prompt_cache_key than the main turn, and skipped the per-session provider hooks — producing intermittent OpenRouter response-cache misses across consecutive advisor calls.
Extract the inline streamFn wrapper in sdk.ts into a shared createSettingsAwareStreamFn helper (packages/coding-agent/src/session/settings-stream-fn.ts) and pass it (plus transformProviderContext) through AgentSessionConfig as advisorStreamFn / transformProviderContext. #buildAdvisorRuntime now hands the advisor Agent the same streamFn, hooks, providerSessionState, promptCacheKey (= advisor session id), and transformProviderContext as the main turn. Adds getAdvisorAgent() accessor on AgentSession for diagnostics and parity tests.
Fixes#3639
- Removed the pi dialect implementation and associated source files.
- Updated dialect resolution, factory registration, and type definitions to exclude pi.
- Cleaned up settings schema and user options to remove pi-related configurations.
- Deleted corresponding test suites covering pi dialect functionality, in-band tools, and examples.
Agent.#runLoop catches provider/stream failures internally and resolves prompt() cleanly with the message recorded on state.error, so AdvisorRuntime treated the OpenRouter 404/no-endpoints turn as a success and never reached notifyFailure. Inspect state.error after each prompt and throw so the retry/notify path runs on real provider failures.
Fixes#3635
- Removed "running" status and hub hint details from the subagent badge text.
- Updated relevant status line tests to expect the simplified badge format.
Surfaced non-recovering advisor prompt failures through session notices so provider errors like OpenRouter ZDR endpoint rejection are visible in the main session.
Fixes#3635
- Updated `reconcileTitleCasing` to ignore purely uppercase words when restoring source casing.
- Restricted restoration logic to mixed-case identifiers (e.g., `TinyVMM`, `iOS`) to avoid overriding the model's clean sentence case with emphatic user shouting.
- Added regression tests to ensure all-caps input does not trigger title re-shouting.
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
On a long subagent-heavy run the TUI thought stream froze for 1.5–4 s
at a time, with the loop watchdog blaming `subagent:*` phases and the
log carrying ~1,255 `Skipping mid-run compaction because turn persistence
is out of order` debug lines per session (issue #3629).
`#persistTurnMessagesForMidRunCompaction` and the persistence-check
helpers it calls were doing two O(n²) things per `onTurnEnd`:
1. Every call rebuilt the branch path via `SessionManager.getBranch()`
(O(branchSize) `unshift` per call), once per turn message twice (once
in `#sessionMessageAlreadyPersisted`, once in
`#hasPersistedLaterTurnMessage`).
2. Every pairwise message comparison serialized full message content
through `#messageValueSignature` (=`JSON.stringify`), even though
the content was almost always orthogonal to the identity decision.
Replace the structural comparator with a stable persistence key
(timestamp + role-specific discriminators) extracted to a sibling module
`session/turn-persistence.ts` and drive the planner off one snapshot of
the branch built per `onTurnEnd`. Content equality is retained as the
slow-path tiebreaker for the rare case where two messages collide on the
cheap key (e.g. two assistant turns at the same millisecond with
`undefined` responseId — the shape the test harness emits).
The new helpers (`sessionMessagePersistenceKey`, `planTurnPersistence`,
`sameMessageContent`) are pure functions covered by direct unit tests
in `test/turn-persistence.test.ts`; the existing mid-run / eager
compaction integration tests prove the end-to-end behavior is preserved.
Fixes#3629
Coerced missing scope on installed_plugins.json entries to user before suppressing them, matching listClaudePluginRoots semantics, so users carrying registries written before the scope field still see marketplace plugins hidden from plugin list and doctor.
Fixes#3628
Derived marketplace runtime package names from plugin IDs when package.json is absent, preserving suppression for config-only marketplace installs. Added regression coverage for package-less LSP plugin installs in plugin list and doctor.
Fixes#3628
Compared marketplace runtime entries by realpath before suppressing link-only plugin-manager entries. Added a regression where a local runtime link reuses a marketplace package name but points at a different path.
Fixes#3628
Filtered marketplace-managed runtime symlinks out of the npm plugin listing and OMP extension-package status provider while keeping marketplace installs available to the runtime loader. Added regressions for both duplicate surfaces.
Fixes#3628
- Added formatAdvisorContextPrompt to render project context files into the advisor's system prompt.
- Updated AgentSession to accept and inject advisorContextPrompt into the session system prompt.
- Registered project context files for the advisor to ensure the reviewer evaluates the agent against standing project instructions like AGENTS.md.
Resolved live ACP generate_image payloads through the blob store before emitting image content while keeping rawOutput compact.
Added regression coverage for content[] image blocks and details.images entries without duplicating blob refs as fallback text.
Fixes#3623
Expanded environment variable placeholders in Claude marketplace plugin MCP url and headers before registration. Added a regression test covering context7-style HTTP server headers.\n\nFixes #3621
- advisor-toggle: the advisor role now falls back to the slow priority
chain when modelRoles.advisor is unset, so enabling the advisor with no
explicit model now resolves one and activates. Assert the inactive-but-
enabled path via an explicit unresolvable advisor model override.
- acp-builtins: dropped the /move happy-path test that ran the real
command (setProjectDir -> process.chdir into a temp dir) then removed
the dir, leaving the process cwd unlinked and poisoning every later file
in the chunk with CurrentWorkingDirectoryUnlinked when Bun.Transpiler
initialized in rewrite-imports.
- Update `acp-builtins.test.ts` to support interactive session movement and path session file testing.
- Clarify `/move` command routing in `/move` slash command tests.
- Rename internal `search` references to `grep` to align with product definitions.
- Refactor TUI render tests to use `Promise.withResolvers` for cleaner flow control.
- Added `resolveAdvisorRoleSelection` to handle the advisor role's distinct configuration logic.
- Implemented `rolePriorityDefaults` to allow the advisor role to alias the `slow` model priority chain.
- Updated `AgentSession` to utilize the new advisor-specific resolution logic during session operations.
- Added new Gemini flash-lite and 3.5-flash variants to the priority list.
- Included additional OpenAI codex model versions in the slow priority configuration.
- Update `umans-provider` test to remove references to deprecated GLM 5.1 model.
- Rename search tool reference to `grep` in `advisor` test.
- Improve test stability in TUI components by explicitly draining `setImmediate` queues before flushing terminal state.