Three-piece architecture so subagents inherit async.enabled and
bash.autoBackground.enabled instead of having both force-disabled:
- Owner-routed delivery: AsyncJobManager gains registerDeliverySink /
waitForOwnerJobs; every AgentSession registers a sink for its own agent
id, so background job results inject into the owning agent's run.
Owned deliveries with no live sink dead-letter (result retained on the
job row) instead of misrouting into the first top-level session.
- Quiescence barrier: a subagent's final yield with owner jobs still
running/undelivered is a scheduling pause, not completion. The run
driver notifies the model once (hub wait/cancel), settles owner work,
and folds results in as async-result follow-ups; teardown cancels and
awaits surviving jobs before isolation worktree capture/cleanup.
- Steering soft channel: queued steering no longer hard-aborts
non-interruptible tools; it aborts interruptible waits and raises a
cooperative ToolCallContext.steeringSignal. The mid-batch watch runs
for every batch, and auto-backgroundable bash backgrounds itself on
steer so incoming messages inject promptly with no work lost.
The exact-bounds rewrite matched reverted PR #5812; the ±context expansion
(1 leading + 3 trailing line) around explicit selectors is intended behavior
so edit anchors at range boundaries stay fresh.
- tools.test.ts still asserted the pre-#5812 ±context expansion for offset/
limit and archive-entry reads; updated to the exact-bounds contract.
- selector-controller-logout.test.ts mocked the old modelRegistry.refresh;
#5786 switched logout to a provider-scoped refreshProvider(id, 'online'),
so the mock never resolved and the test timed out.
- Treated ZIP-based .jar/.war/.ear/.apk as zip archives in archiveFormatFromPath and parseArchivePathCandidates so read/write member access works.
- Shared one archive-extension alternation between format detection and path splitting to stop them drifting.
- Derived the markit convertible-extension set from a single source of truth (utils/markit) matching the registered converters (pdf/docx/pptx/xlsx/epub), dropping legacy .doc/.ppt/.xls/.rtf that had no converter and only produced Unsupported format errors.
- Updated read/write tool prompts to document the zip-family extensions.
Fixes#5808
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
Per-line column cap trims individual lines with a `…` marker but does not
truncate the output window. OutputSink.dump() nonetheless set truncated=true
whenever a line was capped, and truncationFromSummary then reported a byte
tail-window truncation, appending a bogus "Showing lines X-Y of Z (…B limit).
Read artifact://N for full output" footer even though every line was shown.
- OutputSink no longer flips #truncated on column-cap-only drops.
- OutputSummary carries columnMax; truncationFromSummary surfaces it as the
"Some lines truncated to N chars" limit notice regardless of window state.
Fixes#4735
- Stop propagating real-time updates for backgrounded Bash jobs to avoid UI flickering once a job enters the background.
- Refine background task tracking in `EventController` to distinguish between persistent background tasks and transient backgrounded Bash commands.
- Update UI rendering to display cleaner background job metadata in the footer instead of inline text notices.
Models emit optional string args as empty strings; since ff3b0c795
(#4622) read/grep rejected a present-but-empty selector as invalid
instead of behaving like an omitted one. Normalize empty and
whitespace-only selector params to undefined before validation.
Adopted from PR #4881 minus unrelated prompt churn.
Fixes#4879
Treat timeout 0 as an explicit no-deadline contract across the bash tool, executor, async job, and PTY paths.
Signed-off-by: Christian Stewart <christian@aperture.us>
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Added `sanitizeOpenAIResponsesReasoningItemForReplay` to process reasoning-type items by stripping unique identifiers and filtering properties.
- Updated the main sanitization utility to route reasoning items through the new logic.
Batch migration of 13 fs.rmSync calls to removeSyncWithRetries across:
- core/apply-patch.test.ts (4 calls)
- bash-executor.test.ts (4 calls)
- tools.test.ts (2 calls)
- compaction-hooks.test.ts (1 call)
- compaction-thinking-model.test.ts (2 calls)
Also exports removeSyncWithRetries from @oh-my-pi/pi-utils as a
standalone function for tests that manage their own temp dirs.
All tests pass: 139 pass, 0 fail across the 5 migrated files.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.
Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init
All 522 previously-failing Windows tests now pass.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
- Queued steering now drains after session settlement, so aborted auto-continued turns no longer leave queued messages stranded.
- Resumable-state detection now treats tool-result messages as resumable so continue can process queued steering after an interrupted tool execution.
- Regression tests were added for queued steer draining after abort and after an interrupted tool result.
- Added optional `useless` flags to tool result types and payload builders.
- Added `pruneUseless` and `dropUeless` options to control uneventful result pruning.
- Changed compaction and shake passes to prune or ignore non-error useless tool results.
- Changed conversation serialization to omit useless toolCall/toolResult pairs from output.
- Added coverage for useless tagging, pruning, and serialization behavior.
Removes isBackgroundJobSupportEnabled and JobTool.createIf; the tool is now registered unconditionally via `new JobTool(s)`. `async.enabled` now gates async bash commands only — the task tool runs asynchronously regardless. Deletes the async/support module and its barrel re-export.
vault writes now rated write-tier and plan-mode enforced; .tar.gz rewrites keep gzip, are atomic, and write through symlinks; CRLF conflict detection works; conflict twins only invalidated when truly stale; ask discloses timeout auto-selection in result and transcript; todo rejects duplicate ids and stops persisting half-applied batches; auto-generated guard validates against mtime+size; ACP writes run post-write bookkeeping; irc errors set isError.
- Set `FindTool` to disable recursive glob traversal so `dir/*` stays shallow.
- Added `parseSearchPathPreferringLiteral` to prefer literal paths like `apps/[id]/page.tsx` when they exist.
- Updated `resolveToolSearchScope` to reject external URLs with a clear `read` usage error.
- Expanded plan-mode sandbox checks to allow absolute paths inside the local artifact root.
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
- Parsed zip metadata via central directory and lazy ranged reads.
- Inflated member contents only when a specific entry is read.
- Prevented large or corrupt zips from freezing directory reads.
- Updated interactive-mode plan review tests to capture shared fixtures, clear references, and run cleanup with explicit garbage collection before disposal.
- Increased the MCP HTTP transport test connection timeout from 200ms to 1,000ms.
- Adjusted the tool streaming command delay and tightened a start-pending-submission spy type in tests for better stability and type accuracy.
- Updated read tool prompt to show alternate image description when inspect_image is disabled.
- Passed INSPECT_IMAGE_ENABLED flag into prompt rendering context.
- Added test verifying description omits inspect_image references when disabled.
- Updated hashline streaming preview tests to generate snapshot-tagged section headers and use an in-memory snapshot store.
- Replaced untagged file markers in multi-section preview inputs with `formatHashlineHeader` values derived from recorded file contents.
- Changed the bash command error test to expect a returned `isError` result with exit code 1 instead of a rejected promise.
- Changed output format to group results under `# /` headers to reduce token usage for shared path prefixes.
- Clamped the `limit` parameter to 1-200 (default 200) instead of the previous 1000.
- Updated tests to assert against raw file lists instead of parsed text output.
- Measured bash wall-clock duration for direct, terminal-bridge, and interactive execution paths.
- Recorded wall time in result notices and details, then stripped the duplicated literal notice during shell rendering.
- Updated the renderer to include wall time in the status label and added tests for the new wall-time behavior.
- Dropped the `fileType: natives.FileType.File` restriction so glob searches can return directories as well as files.
- Updated the find tool prompt to document directory results and trailing-slash output.
- Added tests verifying directory matches are included and emitted with a trailing `/`.
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.
- Extended the abort scenario shell command from a 5-second sleep to 60 seconds.
- Raised the corresponding test timeout to 15,000 ms for the abort test case.
- Added `providerRetryWait` and `retryWait` hooks to stream/usage options so tests bypass real scheduler delays.
- Parameterized GitHub Copilot poll intervals and Copilot model retry base delay for fast test execution.
- Replaced `Bun.sleep`/`setTimeout` polling loops with `AbortSignal` event listeners in agent session tests.
- Consolidated auth-gateway E2E helpers into a shared `test/helpers` module, eliminating duplicated `checkGatewayAvailable` implementations.
- Migrated credential-disabled tests from SQLite-backed stores to an in-memory store, removing temp-dir lifecycle overhead.
- SearchTool now tracked the last emitted line and inserted ellipsis markers when noncontiguous match blocks were output.
- Display output gap markers were padded to align with code-frame gutters.
- Added a regression test that verified a no-context search emits an ellipsis between separated matches in the same file.
- Updated BashTool's leading `cd` regex to stop matching newline characters so cwd extraction only applies to a single-line `cd ... &&` prefix.
- Added a regression test for multiline commands with a later-line `&&` to ensure each line of the script executes normally.
- Changed multi-file search paging to skip whole files and page results in file windows.
- Added per-file match caps, round-robin file selection, and new file-limit truncation reporting.
- Replaced match/result limit metadata with fileLimitReached and perFileLimitReached.
- Lowered read.defaultLimit default to 300 with 1 lead and 3 trailing context lines.
- Replaced the search skip test with file-pagination coverage and added per-file cap tests.
- Added session-stats analytics tooling to classify searches, detect repeats, and render relevance plots.
- Updated read range expansion to use 1 leading and 3 trailing context lines.
- Changed read.defaultLimit from 500 to 300 in settings defaults.
- Updated read docs and tests to reflect the asymmetric context line behavior.
- Added read-selector analyzers and replay simulators to evaluate coverage and savings.
- Added plotting tools that output new session-stats PNG dashboards from local usage data.
- Added `tools.artifactHeadBytes` and `tools.outputMaxColumns` settings with defaults in `SETTINGS_SCHEMA`.
- Expanded `OutputSink` with `headBytes`/`maxColumns` and middle truncate logic with elision markers and tracking.
- Updated output-meta to resolve sink settings, emit truncation metrics, and use `truncateMiddle` for spills.
- Integrated head and column limits into JS/Python/Bash/SSH/read output flows, with `:raw` skipping read truncation.
- Documented new output middle-elision and column-cap behavior in `CHANGELOG.md`.
- Added truncation tests for `OutputSink`, `truncateMiddle`, and read-tool line handling.
- Added process-wide singleton instances for InternalUrlRouter, AsyncJobManager, and MCPManager.
- Changed internal URL protocols to resolve through registered sessions and scan all active roots/datasets for matches.
- Refactored agent, artifact, memory, rule, skill, jobs, and mcp handlers to use shared manager and rule/skill state.
- Removed per-session protocol/tool wiring and switched tests to initialize and reset global singleton state.
- Expanded read tool range calculations to include optional leading and trailing context lines around user-requested offsets and limits.
- Added shared context range expansion logic so line-slice and streaming reads return anchor-safe windows while preserving existing truncation behavior.
- Updated read-tool tests to assert boundary-offset, limit, and combined offset-limit reads now include the expected ±3 lines of context.
- Added `.ipynb` detection and editable-cell conversion utilities, including merge/serialize helpers in `edit/notebook`.
- Rerouted hashline, patch, and replace edit flows, plus read/write paths, through notebook-aware helpers before persistence.
- Removed the dedicated `notebook` tool, its schema flags, renderer, and built-in registration/settings checks.
- Updated notebook read behavior, docs, and tests so `.ipynb` reads return editable `# %%` cells and edits reserialize to JSON.