- Lifted the unknown-incremental-label check above the !useLastTurn guard so type: [findings], result: {} also rejects, preventing a sibling section's MAX_SCHEMA_RETRIES override from sneaking the stale section through finalization.\n- Extended the regression test with the last-turn payload case.\n\nRefs #3926
- Made unknown incremental yield labels under a closed caller schema throw on every call without consuming MAX_SCHEMA_RETRIES, so the post-mortem override no longer accepts the stale-label payload.\n- Strengthened the regression test to assert the hard failure repeats and the schema-retry budget stays intact for legitimate shape mismatches.\n\nRefs #3926
- Rejected unknown incremental yield labels when the active output schema is closed, so caller override schemas fail in-tool with retry feedback instead of post-mortem schema_violation.\n- Added regression coverage for reviewer-native labels under a caller override schema.\n\nFixes #3926
Normalized string-encoded JSON arrays in grep/search path handling so direct execute paths match validated tool-call behavior.
Added regression coverage for direct GrepTool.execute paths supplied as a JSON-array-shaped string.
Fixes#3873
The yield tool's per-call schema validator was skipped entirely for incremental
yields (`type: ["<label>"]`), so when a subagent emitted a non-conforming value
for a known section (e.g. DeepSeek-v4-pro returning "Correct"/"correct."/"approved"
for the reviewer's `overall_correctness` enum), the call succeeded locally and
the model got no retry feedback. The mismatch only surfaced post-mortem in
`finalizeSubprocessOutput` as a fatal `schema_violation` — the parent agent
lost the entire result, with no recourse for the subagent to fix it.
Build a per-label sub-validator map alongside the full-schema validator: each
entry validates one section's `data` against its top-level property's sub-schema
(items schema for array-typed labels like `findings`). The yield tool runs this
map for incremental yields and routes failures through the same MAX_SCHEMA_RETRIES
budget the terminal path uses, so the model sees up to three corrective retries
and the existing schema-override safety net accepts the value with
SUBAGENT_WARNING_SCHEMA_OVERRIDDEN after exhaustion. Unknown labels remain
unconstrained so scratchpad/streaming sections still pass.
Fixes#3870
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Always remove surfaced irc:incoming records from the pending-aside queue; the inbox tool result already injects the body, so leaving them queued would auto-inject a duplicate at the next step.
- Updated the inbox tool to drain pending asides regardless of peek.
- Added a regression test asserting a peeked pending aside does not auto-inject.
- Drained running-session IRC asides through the inbox tool before the model step consumes them.
- Added a regression test for messages delivered while the recipient is already running.
Fixes#3834
Rejoined split Windows extension module paths before launch parsing finishes and stripped extended-length Win32 prefixes before Bun import and worker spawn APIs see them.
Fixes#3804
- Adjusted hashline header formatting to preserve absolute file paths instead of truncating them to basenames.
- Ensured absolute paths are passed through shortenPath to allow resolution while keeping home directory references concise.
- Prevented edit failures when reading files outside the workspace by ensuring tags remain resolvable.
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
The SSH renderer previously declared 'provisionalPendingPreview: "collapsed"',
which only opted the COLLAPSED pending shape out of the transcript's
stable-prefix ratchet. Once the user expanded an in-flight SSH preview (ctrl+o)
and the framed block outgrew the viewport, the pending rows became
ratchet-eligible and committed to native scrollback before the result render
inserted the 'Output' section. The settled render then re-anchored the frame,
producing two distinct stranded shapes in history:
- a stale 'pending SSH: [host]' header pinned above the final '<- SSH: [host]'
frame (header variant), and
- the pending bottom border row reused in-place as the new 'Output' separator,
with a fresh '...' footer pushed below it (footer variant).
Flip 'provisionalPendingPreview' to 'true' so every pending shape — collapsed
or expanded — is treated as provisional and stays out of native scrollback
until the result render commits a settled frame. The 'collapsed'-only opt-out
remains correct for renderers (bash, eval) whose expanded pending preview is
top-anchored and survives the result render without re-anchoring.
Added two contract tests asserting expanded pending SSH is commit-unstable
and that bash/eval expanded pending preview is still commit-stable — keeping
the opt-in renderer-scoped.
Fixes#3714
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
- Replace the static "Goal/Next" status line with an ephemeral LLM-generated summary triggered after idle periods.
- Hook the recap into the agent's side-channel pipeline, using live goal and task state as context anchors for meaningful recaps.
- Implement abort logic so that active user interactions immediately cancel pending recaps and discard late-arriving responses.
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
- Extracted taskToolRenderer to a dedicated renderer file to resolve circular dependencies.
- Updated all references to the renderer to point to the new location.
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
- Added a fixed-width title slot system to serialize and persist session titles across physical files and backend storage.
- Integrated automated session title updates triggered by todo replan operations using conversation history context.
- Extended storage interfaces across memory, file-system, Redis, and SQL backends to support independent title metadata updates.
- Implemented title metadata parsing within session loaders and list utilities to ensure accurate retrieval and display.
Kept the ask tool question and options visible when switching to the Other custom input editor.
Added a regression test for the custom editor title context.
Fixes#3660
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
- Redesigned the Todo HUD as a connector tree with fixed-budget stage previews.
- Anchored status and HUD containers to prevent redundant UI elements in terminal scrollback.
- Implemented tree-based rendering for project phases and tasks while removing dynamic border rules.
- Upgraded `sherpa-onnx` and related packages to support current infrastructure.
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
- Updated `formatReadHashlineHeader` to use `path.basename` instead of the full path.
- Applied the new formatter across record and display functions in the read tool.
- Added and updated tests to verify that hashline headers now contain only the filename for nested files.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Removed the enforced minimum length for the `items` array in the tool schema to prevent unnecessary validation errors on operations that ignore the field.
- Added tests to verify schema acceptance of empty `items` arrays and confirmed that runtime errors are still correctly thrown for empty inputs during specific operations like `append`.
- Suppressed main-UI relay for sibling broadcast legs when the main agent is a direct target.
- Added `suppressRelay` option to `IrcBus.send` to allow selective disabling of relay rendering.
- Ensured broadcast fan-outs avoid rendering the same message twice in the main transcript.
The >4 MiB JS fallback re-introduced regex-dialect divergence by size. Oversized line-mode virtual resources are now searched in line-boundary chunks (each <= NATIVE_GREP_MAX_FILE_BYTES) through native grep, with matched line numbers offset by each chunk's start, so RE2 dialect parity holds for all line-mode sizes. A single line larger than the cap (un-grepable) is JS-tested individually. Multiline keeps the JS fallback, since chunk boundaries would drop cross-line matches. Test now proves (?i)NEEDLE matches a >4 MiB resource.
- search: native grep silently skips files above NATIVE_GREP_MAX_FILE_BYTES (4 MiB), so the RE2 virtual path dropped matches for large virtual resources (history://, big artifacts). searchVirtualResources now falls back to a JS RegExp matcher for oversized content (the pre-RE2 behavior) while keeping native parity for normal sizes; buildVirtualMatches still rebuilds context/ranges.
- ssh write: a trailing-slash target (ssh://h/dir/) is now rejected before staging, so the mkdir -p parent-creation no longer leaves a directory behind on a refused write.
- search: the native virtual probe capped matched-line detection at INTERNAL_TOTAL_CAP before range filtering, so a ranged virtual selector (ssh://h/log:5000-5100) over a file with >2000 earlier matches returned nothing. The probe now uses the line-count bound for ranged resources so range filtering sees every hit.
- ssh write: writeRemoteFile staged into a temp beside the destination before creating parents, so writing a new nested path failed with 'No such file or directory'. It now mkdir -p's the parent before staging, matching local write, keeping the directory/special-file refusal checks.
searchVirtualResources matched with JS RegExp, so an ssh:// (or other virtual) search diverged from local search for RE2-valid but JS-invalid patterns (e.g. (?i)x, [[:digit:]]) — throwing 'Invalid regex' or returning different matches even when local grep had already validated the pattern in a mixed scope. It now detects matched line numbers with native grep (the same RE2 matcher local search uses) and rebuilds the existing forward-only, range-trimmed context windows via buildVirtualMatches, so virtual/remote and local search share one dialect. Removed the now-dead JS-regex helpers (probeRegexDialect, compileVirtualRegex, searchVirtualResource{Lines,Multiline}, findLineIndex).
search's generic glob-char guard saw the [ ] of an IPv6 authority (ssh://[::1]/etc/hosts) and threw 'Glob patterns are not supported' before the SSH handler could strip the brackets. The check now runs only on the path portion for ssh:// URLs, so IPv6 literals resolve while a glob in the remote path (ssh://host/p*) still rejects.
- reject explicit ssh:// port 0 before connecting (Codex P2)
- keep a path-less ssh://host:port authority port out of selector peeling
- restore ResolveContext on ProtocolHandler.complete (symmetry with resolve/write)
- clarify search/read selector-parity docs + add read-side regression
- `read ssh://host/dir` lists a remote directory one level deep; `ssh://host/` lists the remote root
- add statRemotePath + listRemoteDir; resolve reads first and classifies on error (directory -> one-level listing, dirs-first, dotfiles included)
- directory resources carry isDirectory + immutable and expose no sourcePath
- search refuses a virtual (no-sourcePath) directory resource instead of grepping the listing text
- writeRemoteFile refuses a directory destination and cleans up its temp on that path
- buildSshTarget rejects destinations beginning with "-" (SSH argument-injection / local RCE guard)
- gate ssh:// read/search/write at the exec approval tier; substring scan covers search's pre-expansion delimited paths and write's hashline-wrapped paths
- validate the entire materialized buffer as UTF-8 instead of only the first 8 KiB prefix
- write peels read selectors (raw/conflicts) so it targets the same file read does, and rejects line-range/malformed selectors instead of silently stripping them
- write to a uniquely named remote temp; document symlink-replacement on write as a v1 limit
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.