Commit Graph

165 Commits

Author SHA1 Message Date
can1357 29aefbca4a test: restored read-tool context-expansion contract tests
The exact-bounds rewrite matched reverted PR #5812; the ±context expansion
(1 leading + 3 trailing line) around explicit selectors is intended behavior
so edit anchors at range boundaries stay fresh.
2026-07-18 22:05:14 +02:00
can1357 32dc282517 test: aligned read-tool and logout tests with landed contracts
- tools.test.ts still asserted the pre-#5812 ±context expansion for offset/
  limit and archive-entry reads; updated to the exact-bounds contract.
- selector-controller-logout.test.ts mocked the old modelRegistry.refresh;
  #5786 switched logout to a provider-scoped refreshProvider(id, 'online'),
  so the mock never resolved and the test timed out.
2026-07-18 21:58:03 +02:00
roboomp 22a8524058 fix(tools): recognized zip-family archives and pruned unconvertible extensions
- Treated ZIP-based .jar/.war/.ear/.apk as zip archives in archiveFormatFromPath and parseArchivePathCandidates so read/write member access works.
- Shared one archive-extension alternation between format detection and path splitting to stop them drifting.
- Derived the markit convertible-extension set from a single source of truth (utils/markit) matching the registered converters (pdf/docx/pptx/xlsx/epub), dropping legacy .doc/.ppt/.xls/.rtf that had no converter and only produced Unsupported format errors.
- Updated read/write tool prompts to document the zip-family extensions.

Fixes #5808
2026-07-17 08:04:10 +00:00
can1357 206cb93fa1 test(bash): asserted timeout settles as flagged result
Local bash timeouts resolve with details.timedOut and a warning render
since #5546; ACP retains rejection semantics.
2026-07-17 05:46:57 +02:00
can1357 a049ed4ca7 merged PR #5422: fix(tools): stop column cap from faking window truncation 2026-07-16 03:32:03 +02:00
can1357 5ff277349c refactor(coding-agent): consolidated tool surface onto xd:// devices and hub
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
2026-07-15 15:16:29 +02:00
can1357 a9c038818d feat(tools): removed separate selector args from read and grep APIs
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
2026-07-15 00:04:02 +02:00
roboomp f697837659 fix(tools): stopped column cap from faking window truncation
Per-line column cap trims individual lines with a `…` marker but does not
truncate the output window. OutputSink.dump() nonetheless set truncated=true
whenever a line was capped, and truncationFromSummary then reported a byte
tail-window truncation, appending a bogus "Showing lines X-Y of Z (…B limit).
Read artifact://N for full output" footer even though every line was shown.

- OutputSink no longer flips #truncated on column-cap-only drops.
- OutputSummary carries columnMax; truncationFromSummary surfaces it as the
  "Some lines truncated to N chars" limit notice regardless of window state.

Fixes #4735
2026-07-14 16:27:06 +00:00
can1357 87a64b2f6a feat(coding-agent): improved background job lifecycle and display
- Stop propagating real-time updates for backgrounded Bash jobs to avoid UI flickering once a job enters the background.
- Refine background task tracking in `EventController` to distinguish between persistent background tasks and transient backgrounded Bash commands.
- Update UI rendering to display cleaner background job metadata in the footer instead of inline text notices.
2026-07-13 00:54:51 +02:00
can1357 2553495d00 fix(tools): treated empty read/grep selector fields as omitted
Models emit optional string args as empty strings; since ff3b0c795
(#4622) read/grep rejected a present-but-empty selector as invalid
instead of behaving like an omitted one. Normalize empty and
whitespace-only selector params to undefined before validation.
Adopted from PR #4881 minus unrelated prompt churn.

Fixes #4879
2026-07-09 18:27:20 +02:00
Christian Stewart 78d4978c51 fix(bash): support disabled command deadlines
Treat timeout 0 as an explicit no-deadline contract across the bash tool, executor, async job, and PTY paths.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-05 16:03:37 -07:00
can1357 95b91c7f73 feat(coding-agent/tools)!: replaced paths arrays with path strings
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
2026-07-02 08:30:33 +02:00
can1357 e8090bb48a feat: introduced binary file detection to prevent encoding corruption
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
2026-06-30 02:59:41 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 a6bdef85a2 feat(ai): added support for reasoning items in replay sanitization
- Added `sanitizeOpenAIResponsesReasoningItemForReplay` to process reasoning-type items by stripping unique identifiers and filtering properties.
- Updated the main sanitization utility to route reasoning items through the new logic.
2026-06-23 08:18:59 +02:00
can1357 2787b6dff7 chore: update changelogs 2026-06-23 08:18:29 +02:00
oldschoola a436dfbbeb fix(test): migrate fs.rmSync to removeSyncWithRetries in 5 more test files
Batch migration of 13 fs.rmSync calls to removeSyncWithRetries across:
- core/apply-patch.test.ts (4 calls)
- bash-executor.test.ts (4 calls)
- tools.test.ts (2 calls)
- compaction-hooks.test.ts (1 call)
- compaction-thinking-model.test.ts (2 calls)

Also exports removeSyncWithRetries from @oh-my-pi/pi-utils as a
standalone function for tests that manage their own temp dirs.

All tests pass: 139 pass, 0 fail across the 5 migrated files.
2026-06-19 17:30:58 -07:00
can1357 2de02f2119 Merge PR #3019: fix: Windows test failures — path handling, EBUSY, SQLite handle leaks (@oldschoola) 2026-06-19 17:17:04 +02:00
can1357 67f6518e42 feat: enhanced tool robustness, improve authentication flow, and update API parameters
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
2026-06-19 16:46:07 +02:00
oldschoola 14252e71cb fix: Windows test failures — path handling, EBUSY, SQLite handle leaks
Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.

Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
   with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
   and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
   that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
   from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
   now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
   added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
    cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init

All 522 previously-failing Windows tests now pass.
2026-06-18 21:32:38 -07:00
can1357 38d1d3a2f5 feat(coding-agent): vendor relevant parts of markit 2026-06-18 19:11:22 +02:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
can1357 f9a8aa1d96 fix(coding-agent): fixed queued steering drain after aborted and interrupted tool turns
- Queued steering now drains after session settlement, so aborted auto-continued turns no longer leave queued messages stranded.
- Resumable-state detection now treats tool-result messages as resumable so continue can process queued steering after an interrupted tool execution.
- Regression tests were added for queued steer draining after abort and after an interrupted tool result.
2026-06-13 21:25:48 +02:00
can1357 28df38ed37 feat(agent): added useless-result tagging and compaction dropping of tool outputs
- Added optional `useless` flags to tool result types and payload builders.
- Added `pruneUseless` and `dropUeless` options to control uneventful result pruning.
- Changed compaction and shake passes to prune or ignore non-error useless tool results.
- Changed conversation serialization to omit useless toolCall/toolResult pairs from output.
- Added coverage for useless tagging, pruning, and serialization behavior.
2026-06-12 16:26:38 +02:00
can1357 5406edeed0 refactor(coding-agent): made the job tool always available and dropped async-job gating
Removes isBackgroundJobSupportEnabled and JobTool.createIf; the tool is now registered unconditionally via `new JobTool(s)`. `async.enabled` now gates async bash commands only — the task tool runs asynchronously regardless. Deletes the async/support module and its barrel re-export.
2026-06-10 17:44:31 +02:00
can1357 82225e4444 fix(coding-agent): closed vault write approval bypass and fixed interaction tools
vault writes now rated write-tier and plan-mode enforced; .tar.gz rewrites keep gzip, are atomic, and write through symlinks; CRLF conflict detection works; conflict twins only invalidated when truly stale; ask discloses timeout auto-selection in result and transcript; todo rejects duplicate ids and stops persisting half-applied batches; auto-generated guard validates against mtime+size; ACP writes run post-write bookkeeping; irc errors set isError.
2026-06-10 01:27:17 +02:00
can1357 ee4df538f0 fix(tools): resolved search-path parsing and local sandbox checks
- Set `FindTool` to disable recursive glob traversal so `dir/*` stays shallow.
- Added `parseSearchPathPreferringLiteral` to prefer literal paths like `apps/[id]/page.tsx` when they exist.
- Updated `resolveToolSearchScope` to reject external URLs with a clear `read` usage error.
- Expanded plan-mode sandbox checks to allow absolute paths inside the local artifact root.
2026-06-08 02:04:59 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 492b454141 fix(archive): read zip central directory without inflating members
- Parsed zip metadata via central directory and lazy ranged reads.
- Inflated member contents only when a specific entry is read.
- Prevented large or corrupt zips from freezing directory reads.
2026-06-05 17:38:24 +02:00
can1357 1951081ad5 test(coding-agent): removed redundant AsyncJobManager.setInstance calls
- Relied on session-injected asyncJobManager instead of global instance.
2026-06-05 15:57:10 +02:00
can1357 07998fcc09 test(coding-agent): injected asyncJobManager via test session deps
- Passed asyncJobManager through test tool session dependencies.
- Replaced AsyncJobManager.setInstance with explicit session injection.
2026-06-05 15:55:25 +02:00
can1357 2eb39bae34 test(coding-agent): stabilized coding-agent tests with safer teardown and timing tweaks
- Updated interactive-mode plan review tests to capture shared fixtures, clear references, and run cleanup with explicit garbage collection before disposal.
- Increased the MCP HTTP transport test connection timeout from 200ms to 1,000ms.
- Adjusted the tool streaming command delay and tightened a start-pending-submission spy type in tests for better stability and type accuracy.
2026-05-31 04:25:14 +02:00
can1357 b777fe34f9 test: hardened test isolation and reliability across test suites
- Added `omfg` escape handler spies alongside existing `btw` spies in input controller tests.
- Introduced `waitForRenderedText` helper and settings lifecycle hooks to fix flaky apply-patch renderer tests.
- Raised `bash.autoBackground.thresholdMs` to avoid timing-sensitive test failures.
- Wrapped `runSearchQuery` calls with isolated `AuthStorage` instances to prevent shared state leaks.
2026-05-31 04:03:05 +02:00
can1357 725539aeeb feat(read): conditioned inspect_image docs on feature flag
- Updated read tool prompt to show alternate image description when inspect_image is disabled.
- Passed INSPECT_IMAGE_ENABLED flag into prompt rendering context.
- Added test verifying description omits inspect_image references when disabled.
2026-05-30 21:24:14 +02:00
can1357 a53acf1431 test(coding-agent): updated hashline preview tests to use snapshot-tagged headers
- Updated hashline streaming preview tests to generate snapshot-tagged section headers and use an in-memory snapshot store.
- Replaced untagged file markers in multi-section preview inputs with `formatHashlineHeader` values derived from recorded file contents.
- Changed the bash command error test to expect a returned `isError` result with exit code 1 instead of a rejected promise.
2026-05-30 17:51:11 +02:00
can1357 41d509eff0 feat(find): grouped output by directory and clamped limit to 200
- Changed output format to group results under `# /` headers to reduce token usage for shared path prefixes.
- Clamped the `limit` parameter to 1-200 (default 200) instead of the previous 1000.
- Updated tests to assert against raw file lists instead of parsed text output.
2026-05-27 13:13:11 +02:00
can1357 3e5b0b2340 feat(coding-agent/tools): added bash wall-time tracking to results and renderer
- Measured bash wall-clock duration for direct, terminal-bridge, and interactive execution paths.
- Recorded wall time in result notices and details, then stripped the duplicated literal notice during shell rendering.
- Updated the renderer to include wall time in the status label and added tests for the new wall-time behavior.
2026-05-27 01:53:34 +02:00
can1357 c31daf0805 feat(coding-agent): expanded find tool to return matching directories with slash markers
- Dropped the `fileType: natives.FileType.File` restriction so glob searches can return directories as well as files.
- Updated the find tool prompt to document directory results and trailing-slash output.
- Added tests verifying directory matches are included and emitted with a trailing `/`.
2026-05-26 20:22:25 +02:00
can1357 30793c1655 refactor: restructured hashline to use file-level hash validation with colon separators
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.
2026-05-26 13:25:22 +02:00
can1357 6ca92dd08d test(coding-agent): increased bash abort test timeout and runtime
- Extended the abort scenario shell command from a 5-second sleep to 60 seconds.
- Raised the corresponding test timeout to 15,000 ms for the abort test case.
2026-05-22 18:06:50 +09:00
can1357 90b134ca4c test: replaced real timers and sleeps with deterministic test hooks
- Added `providerRetryWait` and `retryWait` hooks to stream/usage options so tests bypass real scheduler delays.
- Parameterized GitHub Copilot poll intervals and Copilot model retry base delay for fast test execution.
- Replaced `Bun.sleep`/`setTimeout` polling loops with `AbortSignal` event listeners in agent session tests.
- Consolidated auth-gateway E2E helpers into a shared `test/helpers` module, eliminating duplicated `checkGatewayAvailable` implementations.
- Migrated credential-disabled tests from SQLite-backed stores to an in-memory store, removing temp-dir lifecycle overhead.
2026-05-17 04:02:09 +02:00
can1357 85515f35a6 fix(coding-agent): added gap separators between noncontiguous search matches
- SearchTool now tracked the last emitted line and inserted ellipsis markers when noncontiguous match blocks were output.
- Display output gap markers were padded to align with code-frame gutters.
- Added a regression test that verified a no-context search emits an ellipsis between separated matches in the same file.
2026-05-15 23:46:23 +02:00
can1357 c7d04f3a13 fix(coding-agent): constrained bash cwd auto-detect regex to single-line cd commands
- Updated BashTool's leading `cd` regex to stop matching newline characters so cwd extraction only applies to a single-line `cd ... &&` prefix.
- Added a regression test for multiline commands with a later-line `&&` to ensure each line of the script executes normally.
2026-05-13 19:19:50 +02:00
can1357 9828764cb4 feat(scripts-session-stats): added session-stats search relevance plotting
- Changed multi-file search paging to skip whole files and page results in file windows.
- Added per-file match caps, round-robin file selection, and new file-limit truncation reporting.
- Replaced match/result limit metadata with fileLimitReached and perFileLimitReached.
- Lowered read.defaultLimit default to 300 with 1 lead and 3 trailing context lines.
- Replaced the search skip test with file-pagination coverage and added per-file cap tests.
- Added session-stats analytics tooling to classify searches, detect repeats, and render relevance plots.
2026-05-13 11:38:09 +02:00
can1357 a541a63547 feat: added read-selector analyzers and coverage plotting tools
- Updated read range expansion to use 1 leading and 3 trailing context lines.
- Changed read.defaultLimit from 500 to 300 in settings defaults.
- Updated read docs and tests to reflect the asymmetric context line behavior.
- Added read-selector analyzers and replay simulators to evaluate coverage and savings.
- Added plotting tools that output new session-stats PNG dashboards from local usage data.
2026-05-13 11:33:42 +02:00
can1357 28b9ce7a0c feat(coding-agent): added middle-elision caps to OutputSink truncation
- Added `tools.artifactHeadBytes` and `tools.outputMaxColumns` settings with defaults in `SETTINGS_SCHEMA`.
- Expanded `OutputSink` with `headBytes`/`maxColumns` and middle truncate logic with elision markers and tracking.
- Updated output-meta to resolve sink settings, emit truncation metrics, and use `truncateMiddle` for spills.
- Integrated head and column limits into JS/Python/Bash/SSH/read output flows, with `:raw` skipping read truncation.
- Documented new output middle-elision and column-cap behavior in `CHANGELOG.md`.
- Added truncation tests for `OutputSink`, `truncateMiddle`, and read-tool line handling.
2026-05-13 11:19:11 +02:00
can1357 1bde755933 feat(coding-agent): added global singletons for URL protocol handlers
- Added process-wide singleton instances for InternalUrlRouter, AsyncJobManager, and MCPManager.
- Changed internal URL protocols to resolve through registered sessions and scan all active roots/datasets for matches.
- Refactored agent, artifact, memory, rule, skill, jobs, and mcp handlers to use shared manager and rule/skill state.
- Removed per-session protocol/tool wiring and switched tests to initialize and reset global singleton state.
2026-05-12 05:07:52 +02:00
can1357 7223b800b0 fix(coding-agent/tools): expanded read ranges with three-line anchor context
- Expanded read tool range calculations to include optional leading and trailing context lines around user-requested offsets and limits.
- Added shared context range expansion logic so line-slice and streaming reads return anchor-safe windows while preserving existing truncation behavior.
- Updated read-tool tests to assert boundary-offset, limit, and combined offset-limit reads now include the expected ±3 lines of context.
2026-05-09 05:22:23 +02:00
can1357 cbfdc0a212 feat(coding-agent): added notebook-aware .ipynb edit/read path helpers
- Added `.ipynb` detection and editable-cell conversion utilities, including merge/serialize helpers in `edit/notebook`.
- Rerouted hashline, patch, and replace edit flows, plus read/write paths, through notebook-aware helpers before persistence.
- Removed the dedicated `notebook` tool, its schema flags, renderer, and built-in registration/settings checks.
- Updated notebook read behavior, docs, and tests so `.ipynb` reads return editable `# %%` cells and edits reserialize to JSON.
2026-05-07 04:16:11 +02:00
can1357 112a0ad632 refactor(coding-agent): exported search default match limit for test reuse
- Exported DEFAULT_MATCH_LIMIT from the search tool module.
- Updated the search tool tests to import DEFAULT_MATCH_LIMIT from the module.
- Updated default-limit assertions and messages to use DEFAULT_MATCH_LIMIT instead of hardcoded values.
2026-05-05 16:15:21 +02:00