PR #3648 review (codex): the first iteration emitted every envelope
path alongside one combined digest, so a multi-file payload that added
`: any` to a README.md hunk and merely touched src/ok.ts would surface
a *.ts path in the TTSR match context and trip the bundled
tool:edit(*.ts) ts-no-any rule on text that belonged to the Markdown
hunk — aborting valid edits under interruptMode:always.
Add a per-file matcherEntries(args) hook on AgentTool / EditStreaming-
Strategy returning [{ path, digest }] entries, one per touched file
(same-path sections/hunks merged):
- replace / patch: one entry from the top-level path + matcherDigest
- hashline: regex-split by [path#TAG] section, body added-lines per
entry (tolerant of streaming partial payloads)
- apply_patch: expandApplyPatchToPreviewEntries grouped by path
AgentSession.#checkTtsrStream / #checkTtsrAstStream now prefer
matcherEntries and iterate per-file with isolated filePaths + streamKey,
so each file's buffer and repeat-tracking are independent. Tools
without matcherEntries keep the existing combined matcherDigest +
matcherPaths path.
- Standardized retry logic to ensure mocked assistant turns and generic aborts share the same reset path as live provider failures.
- Introduced a classification helper that forces retry policy alignment with the currently active session model.
- Enabled retries for generic reason-less abort sentinels and stale OpenAI response errors.
AgentSession's TTSR match context only scanned top-level path/paths
arguments, so hashline and apply_patch edit streams (whose only target
path lives inside the wire payload — section headers or envelope
markers — not as a top-level argument) arrived without any filePaths
and silently skipped path-scoped rules like the bundled ts-no-any
(scope: tool:edit(*.ts)).
Add an optional AgentTool.matcherPaths(args) hook, companion to the
existing matcherDigest(args), so tools whose wire grammar embeds paths
can surface them. Implement on each edit streaming strategy:
- replace / patch: top-level path
- hashline: parse [path#TAG] (and tag-less [path]) section headers
tolerant of streaming partial payloads
- apply_patch: parse *** Add/Update/Delete File: markers, also tolerant
of pre-End-Patch buffers
AgentSession.#getTtsrToolMatchContext consults tool.matcherPaths first,
normalising its output through the existing path-candidate helper, and
falls back to the generic top-level argument scan for tools that don't
implement it.
Fixes#3646
- Replaced functional streaming logic with a stateful `CodexStreamProcessor` class.
- Converted `CodexStreamRuntime` into a class with explicit state initializers and lifecycle methods.
- Encapsulated tool call delta tracking and interruption logic within dedicated class methods.
- Integrated `textSignature` generation into text-type output blocks using versioned encoders.
- Added textVerbosity configuration schema to allow user control over model response length.
- Integrated textVerbosity into settings-aware stream processing for "openai-responses" and "openai-codex-responses" APIs.
- Updated settings stream tests to verify that verbosity settings correctly apply while respecting caller overrides.
- Added optional `phase` field to assistant messages to distinguish between `commentary` and `final_answer` stages.
- Updated `openai-responses-server` to parse and propagate phase metadata via text signatures during streaming and batch processing.
- Modified `encodeStream` and `buildOutputItems` to trigger new message items when an assistant message phase transition is detected.
- Extended schema definitions and test suites to validate phase-aware message handling and history replay.
- Added textVerbosity option to OpenAIResponsesOptions and SimpleStreamOptions.
- Updated stream mapping logic to propagate verbosity setting to the API request body.
- Implemented endpoint validation to ensure verbosity settings are only applied to official OpenAI endpoints.
- Added comprehensive test coverage for request payload inspection and stream event handling.
- Default the reasoning context to all_turns in the OpenAI Codex request transformer.
- Increase the default text verbosity for Codex requests to medium.
- Set a detailed reasoning summary for AI streams by default.
- Remove outdated testing logic for low verbosity defaults.
- Migrated configuration constants to use dynamic environment variable lookups.
- Removed unused helper functions and associated internal logic.
- Deleted comprehensive test suite for decommissioned WebSocket transport mechanisms.
- Update scrubPartialJson to utilize clearStreamingPartialJson for consistent tool-call cleanup.
- Adjust execution order in streamProxy to ensure partial error messages are finalized before scrubbing.
- Remove redundant test expectation comment regarding partialJson leakage.
- Resolved environment variables for Codex WebSocket settings at module load time to prevent redundant parsing during requests.
- Removed deprecated `parseCodexPositiveInteger` helper as environment variables are now resolved once globally.
- Hardened internal state chaining by implementing structural equality checks that ignore transient streaming symbol properties.
- Removed `stripStreamingBlockSymbols` utility as streaming symbol handling is now implicitly managed via key enumeration.
- Refactored `deepEqualsWithout` to rely on `for...in` string key enumeration, ensuring transient streaming symbols are ignored during structural equality checks.
- Aligned structural equality logic to treat `undefined` values and absent keys as equivalent, improving reliability when matching stream chain prefixes.
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
The CI gate failed in several non-overlapping ways once the full coding-agent suite ran here: tests wrote into the real $HOME (`/srv/agent-home`) which is read-only, a fixture rotated stored Anthropic API keys but Settings reloaded the user models.yml and shadowed them, an OAuth callback server bound to `hostname:"localhost"` (loopback unreachable in this runtime), a built-in tool metadata assertion only saw `github` when `gh` was installed, and the GithubTool `pr_checkout` worktree assertions assumed `~/.omp/wt` but `XDG_DATA_HOME` redirected `getWorktreesDir()` to `$XDG_DATA_HOME/omp/wt`.
Fixes:
- `packages/ai/src/registry/oauth/callback-server.ts`: drop `hostname: "localhost"` when no caller-supplied hostname overrides it. Bun on Linux refused inbound connections to the listener when bound explicitly to `localhost`; defaulting to Bun.serves default binding restores loopback connectivity.
- `packages/coding-agent/test/status-line-path.test.ts`: route the `~/Projects` fixtures through a writable temp home (spy `os.homedir()`), housed under the repo `.wt/` worktree scratch so the temp home is not classified as a status-line scratch root.
- `packages/coding-agent/test/skills.test.ts`: ditto for the `~/.pi-skills-test-*` mkdtemp in the tilde-expansion test.
- `packages/coding-agent/test/marketplace/project-scope.test.ts`: stop writing `~/.git`; build the entire home-dir guard fixture in a temp dir and spy `os.homedir()`.
- `packages/coding-agent/test/oauth-flow.test.ts`: wrap each callback fetch in a brief retry so the simulated browser redirect tolerates the few-ms gap before the Bun callback server starts accepting connections.
- `packages/coding-agent/test/tools/gh.test.ts`: extend `setupTempHome()` to clear `XDG_DATA_HOME`/`XDG_STATE_HOME`/`XDG_CACHE_HOME` for the duration of the test so the rebuilt dirs resolver routes `getWorktreesDir()` back through the spied home, then restore them on cleanup.
- `packages/coding-agent/test/tool-discovery/initial-tools.test.ts`: instantiate `GithubTool` directly in the metadata fixture so the assertion runs even when `gh` is unavailable (GithubTool.createIf returns null without `gh`).
- `packages/coding-agent/test/agent-session-retry-cap.test.ts`: pass an isolated `models.yml` path to `ModelRegistry` so the two-Anthropic-key fixture is the authoritative credential source instead of any user-level command-backed Anthropic key.
Verification:
- `bun check` → passed
- `bun run test` (full coding-agent suite, 4 buckets, all chunks) → 0 fails
Fixes#3639
- Centralized transient transport error patterns to enable consistent reuse across error handling modules.
- Refactored `isOpaqueStatusBody` for improved accessibility in classification logic.
- Updated retryable error detection to include the consolidated transport pattern and additional provider-specific error criteria.
Agent.#runLoop appends the user batch and a synthetic stopReason: "error" assistant turn to state.messages before resolving prompt() with state.error set. AdvisorRuntime now snapshots state.messages.length before each prompt, restores it on failure (via a new AdvisorAgent.rollbackTo hook that also resets the advisor's append-only sync cursor), and clears state.error so retries replay a clean baseline and the drop-after-3 path never leaks orphan failed turns into the next successful run's context.
Fixes#3635
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
- Simplified the crc32 function by delegating to the Bun native hash implementation.
- Removed the manual lookup table generation and bitwise iteration logic.
Removed the remaining callers of StatusLineComponent.setSubagentHubHint after the status-line subagent hub hint display was simplified away. The status-line cache test already asserts the hub hint is absent, so the stale setter call only broke type checking.
Fixes#3639
AgentSession.#buildAdvisorRuntime constructed the advisor Agent without the provider-shaping options the SDK installs on the main agent: the streamFn wrapper that applies providers.openrouterVariant / providers.antigravityEndpoint / providers.maxInFlightRequests / model.loopGuard.*, the onPayload/onResponse/onSseEvent hooks, the shared providerSessionState map, transformProviderContext (snapcompact, secret obfuscation, image clamping), and a stable promptCacheKey. Advisor turns therefore dropped the OpenRouter sticky-routing variant suffix, used a different prompt_cache_key than the main turn, and skipped the per-session provider hooks — producing intermittent OpenRouter response-cache misses across consecutive advisor calls.
Extract the inline streamFn wrapper in sdk.ts into a shared createSettingsAwareStreamFn helper (packages/coding-agent/src/session/settings-stream-fn.ts) and pass it (plus transformProviderContext) through AgentSessionConfig as advisorStreamFn / transformProviderContext. #buildAdvisorRuntime now hands the advisor Agent the same streamFn, hooks, providerSessionState, promptCacheKey (= advisor session id), and transformProviderContext as the main turn. Adds getAdvisorAgent() accessor on AgentSession for diagnostics and parity tests.
Fixes#3639
`delete` on object properties degrades V8 hidden class optimization; the new `stripVariant` util sets the property to `undefined` instead, keeping the object shape stable.
`performance.now()` is used in place of `Date.now()` for duration and TTFT measurements to get a monotonic, high-resolution clock that is unaffected by system clock adjustments.
- Added support for `<scratchpad>`, fenced markdown blocks, and specific model-specific channel tags to the thinking scan logic.
- Exposed the `ThinkingInbandScanner` via the package dialect index.
- Removed the pi dialect implementation and associated source files.
- Updated dialect resolution, factory registration, and type definitions to exclude pi.
- Cleaned up settings schema and user options to remove pi-related configurations.
- Deleted corresponding test suites covering pi dialect functionality, in-band tools, and examples.
Agent.#runLoop catches provider/stream failures internally and resolves prompt() cleanly with the message recorded on state.error, so AdvisorRuntime treated the OpenRouter 404/no-endpoints turn as a success and never reached notifyFailure. Inspect state.error after each prompt and throw so the retry/notify path runs on real provider failures.
Fixes#3635
- Centralized thinking block demotion logic using `renderDemotedThinking` across all model providers.
- Updated `transformMessages` to default to text-based demotion for foreign thinking content.
- Restricted foreign thinking preservation to explicitly supported targets with `zai` thinking formats.
- Refactored message transformations and updated test suites to validate standardized demotion outcomes.
- Removed "running" status and hub hint details from the subagent badge text.
- Updated relevant status line tests to expect the simplified badge format.
Surfaced non-recovering advisor prompt failures through session notices so provider errors like OpenRouter ZDR endpoint rejection are visible in the main session.
Fixes#3635
- Updated `reconcileTitleCasing` to ignore purely uppercase words when restoring source casing.
- Restricted restoration logic to mixed-case identifiers (e.g., `TinyVMM`, `iOS`) to avoid overriding the model's clean sentence case with emphatic user shouting.
- Added regression tests to ensure all-caps input does not trigger title re-shouting.
- Implement `renderDemotedThinking` to encapsulate reasoning from prior turns when switching models across providers.
- Render reasoning in the target model's canonical thinking dialect (e.g., Markdown fences) to preserve context while ensuring the blocks are processed as reasoning.
- Use a neutral `<think>` tag fallback for models where canonical native formatting risks leaking incompatible chat-template instrumentation tokens.
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
- Introduce `resolveWithThinkingLoopCook` to manage automated retries for detected model thinking-loop stalls.
- Update `complete` and `completeSimple` to utilize the new retry orchestrator instead of direct stream resolution.
- Configure `pi-native-server` to respect the `loopGuard` option flag.
- Enable a final unguarded pass for persistent stalls after the retry budget is exhausted to ensure generation completion.
On a long subagent-heavy run the TUI thought stream froze for 1.5–4 s
at a time, with the loop watchdog blaming `subagent:*` phases and the
log carrying ~1,255 `Skipping mid-run compaction because turn persistence
is out of order` debug lines per session (issue #3629).
`#persistTurnMessagesForMidRunCompaction` and the persistence-check
helpers it calls were doing two O(n²) things per `onTurnEnd`:
1. Every call rebuilt the branch path via `SessionManager.getBranch()`
(O(branchSize) `unshift` per call), once per turn message twice (once
in `#sessionMessageAlreadyPersisted`, once in
`#hasPersistedLaterTurnMessage`).
2. Every pairwise message comparison serialized full message content
through `#messageValueSignature` (=`JSON.stringify`), even though
the content was almost always orthogonal to the identity decision.
Replace the structural comparator with a stable persistence key
(timestamp + role-specific discriminators) extracted to a sibling module
`session/turn-persistence.ts` and drive the planner off one snapshot of
the branch built per `onTurnEnd`. Content equality is retained as the
slow-path tiebreaker for the rare case where two messages collide on the
cheap key (e.g. two assistant turns at the same millisecond with
`undefined` responseId — the shape the test harness emits).
The new helpers (`sessionMessagePersistenceKey`, `planTurnPersistence`,
`sameMessageContent`) are pure functions covered by direct unit tests
in `test/turn-persistence.test.ts`; the existing mid-run / eager
compaction integration tests prove the end-to-end behavior is preserved.
Fixes#3629
Coerced missing scope on installed_plugins.json entries to user before suppressing them, matching listClaudePluginRoots semantics, so users carrying registries written before the scope field still see marketplace plugins hidden from plugin list and doctor.
Fixes#3628
Derived marketplace runtime package names from plugin IDs when package.json is absent, preserving suppression for config-only marketplace installs. Added regression coverage for package-less LSP plugin installs in plugin list and doctor.
Fixes#3628