- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
The SSH renderer previously declared 'provisionalPendingPreview: "collapsed"',
which only opted the COLLAPSED pending shape out of the transcript's
stable-prefix ratchet. Once the user expanded an in-flight SSH preview (ctrl+o)
and the framed block outgrew the viewport, the pending rows became
ratchet-eligible and committed to native scrollback before the result render
inserted the 'Output' section. The settled render then re-anchored the frame,
producing two distinct stranded shapes in history:
- a stale 'pending SSH: [host]' header pinned above the final '<- SSH: [host]'
frame (header variant), and
- the pending bottom border row reused in-place as the new 'Output' separator,
with a fresh '...' footer pushed below it (footer variant).
Flip 'provisionalPendingPreview' to 'true' so every pending shape — collapsed
or expanded — is treated as provisional and stays out of native scrollback
until the result render commits a settled frame. The 'collapsed'-only opt-out
remains correct for renderers (bash, eval) whose expanded pending preview is
top-anchored and survives the result render without re-anchoring.
Added two contract tests asserting expanded pending SSH is commit-unstable
and that bash/eval expanded pending preview is still commit-stable — keeping
the opt-in renderer-scoped.
Fixes#3714
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
- Replace the static "Goal/Next" status line with an ephemeral LLM-generated summary triggered after idle periods.
- Hook the recap into the agent's side-channel pipeline, using live goal and task state as context anchors for meaningful recaps.
- Implement abort logic so that active user interactions immediately cancel pending recaps and discard late-arriving responses.
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
- Extracted taskToolRenderer to a dedicated renderer file to resolve circular dependencies.
- Updated all references to the renderer to point to the new location.
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
- Added a fixed-width title slot system to serialize and persist session titles across physical files and backend storage.
- Integrated automated session title updates triggered by todo replan operations using conversation history context.
- Extended storage interfaces across memory, file-system, Redis, and SQL backends to support independent title metadata updates.
- Implemented title metadata parsing within session loaders and list utilities to ensure accurate retrieval and display.
Kept the ask tool question and options visible when switching to the Other custom input editor.
Added a regression test for the custom editor title context.
Fixes#3660
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
- Redesigned the Todo HUD as a connector tree with fixed-budget stage previews.
- Anchored status and HUD containers to prevent redundant UI elements in terminal scrollback.
- Implemented tree-based rendering for project phases and tasks while removing dynamic border rules.
- Upgraded `sherpa-onnx` and related packages to support current infrastructure.
- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
- Updated `formatReadHashlineHeader` to use `path.basename` instead of the full path.
- Applied the new formatter across record and display functions in the read tool.
- Added and updated tests to verify that hashline headers now contain only the filename for nested files.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Removed the enforced minimum length for the `items` array in the tool schema to prevent unnecessary validation errors on operations that ignore the field.
- Added tests to verify schema acceptance of empty `items` arrays and confirmed that runtime errors are still correctly thrown for empty inputs during specific operations like `append`.
- Suppressed main-UI relay for sibling broadcast legs when the main agent is a direct target.
- Added `suppressRelay` option to `IrcBus.send` to allow selective disabling of relay rendering.
- Ensured broadcast fan-outs avoid rendering the same message twice in the main transcript.
The >4 MiB JS fallback re-introduced regex-dialect divergence by size. Oversized line-mode virtual resources are now searched in line-boundary chunks (each <= NATIVE_GREP_MAX_FILE_BYTES) through native grep, with matched line numbers offset by each chunk's start, so RE2 dialect parity holds for all line-mode sizes. A single line larger than the cap (un-grepable) is JS-tested individually. Multiline keeps the JS fallback, since chunk boundaries would drop cross-line matches. Test now proves (?i)NEEDLE matches a >4 MiB resource.
- search: native grep silently skips files above NATIVE_GREP_MAX_FILE_BYTES (4 MiB), so the RE2 virtual path dropped matches for large virtual resources (history://, big artifacts). searchVirtualResources now falls back to a JS RegExp matcher for oversized content (the pre-RE2 behavior) while keeping native parity for normal sizes; buildVirtualMatches still rebuilds context/ranges.
- ssh write: a trailing-slash target (ssh://h/dir/) is now rejected before staging, so the mkdir -p parent-creation no longer leaves a directory behind on a refused write.
- search: the native virtual probe capped matched-line detection at INTERNAL_TOTAL_CAP before range filtering, so a ranged virtual selector (ssh://h/log:5000-5100) over a file with >2000 earlier matches returned nothing. The probe now uses the line-count bound for ranged resources so range filtering sees every hit.
- ssh write: writeRemoteFile staged into a temp beside the destination before creating parents, so writing a new nested path failed with 'No such file or directory'. It now mkdir -p's the parent before staging, matching local write, keeping the directory/special-file refusal checks.
searchVirtualResources matched with JS RegExp, so an ssh:// (or other virtual) search diverged from local search for RE2-valid but JS-invalid patterns (e.g. (?i)x, [[:digit:]]) — throwing 'Invalid regex' or returning different matches even when local grep had already validated the pattern in a mixed scope. It now detects matched line numbers with native grep (the same RE2 matcher local search uses) and rebuilds the existing forward-only, range-trimmed context windows via buildVirtualMatches, so virtual/remote and local search share one dialect. Removed the now-dead JS-regex helpers (probeRegexDialect, compileVirtualRegex, searchVirtualResource{Lines,Multiline}, findLineIndex).
search's generic glob-char guard saw the [ ] of an IPv6 authority (ssh://[::1]/etc/hosts) and threw 'Glob patterns are not supported' before the SSH handler could strip the brackets. The check now runs only on the path portion for ssh:// URLs, so IPv6 literals resolve while a glob in the remote path (ssh://host/p*) still rejects.
- reject explicit ssh:// port 0 before connecting (Codex P2)
- keep a path-less ssh://host:port authority port out of selector peeling
- restore ResolveContext on ProtocolHandler.complete (symmetry with resolve/write)
- clarify search/read selector-parity docs + add read-side regression
- `read ssh://host/dir` lists a remote directory one level deep; `ssh://host/` lists the remote root
- add statRemotePath + listRemoteDir; resolve reads first and classifies on error (directory -> one-level listing, dirs-first, dotfiles included)
- directory resources carry isDirectory + immutable and expose no sourcePath
- search refuses a virtual (no-sourcePath) directory resource instead of grepping the listing text
- writeRemoteFile refuses a directory destination and cleans up its temp on that path
- buildSshTarget rejects destinations beginning with "-" (SSH argument-injection / local RCE guard)
- gate ssh:// read/search/write at the exec approval tier; substring scan covers search's pre-expansion delimited paths and write's hashline-wrapped paths
- validate the entire materialized buffer as UTF-8 instead of only the first 8 KiB prefix
- write peels read selectors (raw/conflicts) so it targets the same file read does, and rejects line-range/malformed selectors instead of silently stripping them
- write to a uniquely named remote temp; document symlink-replacement on write as a v1 limit
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.
- Clarify the auto-advancement logic for the in-progress pointer in the documentation and output.
- Add an overall completion count summary to the task list view.
- Update the list format to use standard checkbox indicators and explicit tags for task statuses.
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
Accepted displayed gallery lifecycle labels as --state aliases, rejected unknown values before rendering, and updated failed fixtures to render visibly failed states.
Fixes#3473
The streaming reader's NUL check only walked completed lines collected
from streamLinesFromFile, so a binary blob whose first newline lay past
the byte budget (videos, archives, packed JSON) left collectedLines
empty and slipped through to the firstLineExceedsLimit branch — which
emitted the decoded preview as text instead of the intended refusal.
Sniff firstLinePreview alongside collectedLines so the existing refusal
fires uniformly. Also added a regression test that uses a 256 KiB blob
with no 0x0A bytes to actually exercise the firstLineExceedsLimit
path — the previous 6-byte test fit in one collected line and never
covered the bug.
Fixes#3448
Routed file-backed '/data/workspaces/can1357__oh-my-pi__3448/.omp-session/2026-06-25T07-25-13-303Z_019efdab-28d7-7000-a7a4-e20282508056/local' reads through the normal filesystem reader so binary detection, document/image handling, and streaming safeguards apply before content is materialized.
Hardened the local protocol handler to return metadata-only refusals for binary/container resources instead of decoding them with Bun.file().text().
Fixes#3448
Threaded the caller's loaded skills through internal URL resolution so skill:// handlers do not depend on process-global skill state during tool execution.
Fixes#3436
- Standardized `local://` image processing to prevent file corruption during decoding.
- Refactored local path resolution logic to enforce safety constraints and path containment.
- Implemented an image fast-path in `ReadTool` to correctly render local images before text decoding.
- Added comprehensive test coverage for image rendering, text compatibility, and path security.
- Resolved an event loop hang associated with `omp --resume` operations.
- Update the eval tool to stream stdout chunks directly into the active cell's output buffer while the process is still running.
- Prevent long-running cells from appearing empty in the UI by surfacing incremental output before the backend resolves.
- Add regression tests to ensure streamed output is captured mid-execution and reconciled with final results.
- Removed deprecated eval prelude helpers `append`, `tree`, `diff`, `sort`, `uniq`, and `counter` from all supported runtimes.
- Cleaned up runtime implementations, protocol definitions, and UI rendering logic associated with the removed helpers.
- Updated project documentation, prompts, and test suites to reflect the reduced helper API surface.
- Recorded functional changes in the package changelog.
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
- Refactored `todo` tool to accept a single operation object instead of an `ops` array.
- Implemented parameter normalization to maintain backward compatibility with legacy array-based tool calls.
- Updated tool instructions, documentation, and UI rendering components to reflect the new interface.
- Added compatibility tests to verify rendering and execution for both legacy and current operation formats.