Commit Graph

154 Commits

Author SHA1 Message Date
can1357 95b91c7f73 feat(coding-agent/tools)!: replaced paths arrays with path strings
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
2026-07-02 08:30:33 +02:00
can1357 e8090bb48a feat: introduced binary file detection to prevent encoding corruption
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
2026-06-30 02:59:41 +02:00
can1357 ae1650d689 refactor: renamed search and find tools to grep and glob
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
2026-06-27 00:57:55 +02:00
can1357 a6bdef85a2 feat(ai): added support for reasoning items in replay sanitization
- Added `sanitizeOpenAIResponsesReasoningItemForReplay` to process reasoning-type items by stripping unique identifiers and filtering properties.
- Updated the main sanitization utility to route reasoning items through the new logic.
2026-06-23 08:18:59 +02:00
can1357 2787b6dff7 chore: update changelogs 2026-06-23 08:18:29 +02:00
oldschoola a436dfbbeb fix(test): migrate fs.rmSync to removeSyncWithRetries in 5 more test files
Batch migration of 13 fs.rmSync calls to removeSyncWithRetries across:
- core/apply-patch.test.ts (4 calls)
- bash-executor.test.ts (4 calls)
- tools.test.ts (2 calls)
- compaction-hooks.test.ts (1 call)
- compaction-thinking-model.test.ts (2 calls)

Also exports removeSyncWithRetries from @oh-my-pi/pi-utils as a
standalone function for tests that manage their own temp dirs.

All tests pass: 139 pass, 0 fail across the 5 migrated files.
2026-06-19 17:30:58 -07:00
can1357 2de02f2119 Merge PR #3019: fix: Windows test failures — path handling, EBUSY, SQLite handle leaks (@oldschoola) 2026-06-19 17:17:04 +02:00
can1357 67f6518e42 feat: enhanced tool robustness, improve authentication flow, and update API parameters
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
2026-06-19 16:46:07 +02:00
oldschoola 14252e71cb fix: Windows test failures — path handling, EBUSY, SQLite handle leaks
Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.

Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
   with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
   and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
   that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
   from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
   now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
   added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
    cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init

All 522 previously-failing Windows tests now pass.
2026-06-18 21:32:38 -07:00
can1357 38d1d3a2f5 feat(coding-agent): vendor relevant parts of markit 2026-06-18 19:11:22 +02:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
can1357 f9a8aa1d96 fix(coding-agent): fixed queued steering drain after aborted and interrupted tool turns
- Queued steering now drains after session settlement, so aborted auto-continued turns no longer leave queued messages stranded.
- Resumable-state detection now treats tool-result messages as resumable so continue can process queued steering after an interrupted tool execution.
- Regression tests were added for queued steer draining after abort and after an interrupted tool result.
2026-06-13 21:25:48 +02:00
can1357 28df38ed37 feat(agent): added useless-result tagging and compaction dropping of tool outputs
- Added optional `useless` flags to tool result types and payload builders.
- Added `pruneUseless` and `dropUeless` options to control uneventful result pruning.
- Changed compaction and shake passes to prune or ignore non-error useless tool results.
- Changed conversation serialization to omit useless toolCall/toolResult pairs from output.
- Added coverage for useless tagging, pruning, and serialization behavior.
2026-06-12 16:26:38 +02:00
can1357 5406edeed0 refactor(coding-agent): made the job tool always available and dropped async-job gating
Removes isBackgroundJobSupportEnabled and JobTool.createIf; the tool is now registered unconditionally via `new JobTool(s)`. `async.enabled` now gates async bash commands only — the task tool runs asynchronously regardless. Deletes the async/support module and its barrel re-export.
2026-06-10 17:44:31 +02:00
can1357 82225e4444 fix(coding-agent): closed vault write approval bypass and fixed interaction tools
vault writes now rated write-tier and plan-mode enforced; .tar.gz rewrites keep gzip, are atomic, and write through symlinks; CRLF conflict detection works; conflict twins only invalidated when truly stale; ask discloses timeout auto-selection in result and transcript; todo rejects duplicate ids and stops persisting half-applied batches; auto-generated guard validates against mtime+size; ACP writes run post-write bookkeeping; irc errors set isError.
2026-06-10 01:27:17 +02:00
can1357 ee4df538f0 fix(tools): resolved search-path parsing and local sandbox checks
- Set `FindTool` to disable recursive glob traversal so `dir/*` stays shallow.
- Added `parseSearchPathPreferringLiteral` to prefer literal paths like `apps/[id]/page.tsx` when they exist.
- Updated `resolveToolSearchScope` to reject external URLs with a clear `read` usage error.
- Expanded plan-mode sandbox checks to allow absolute paths inside the local artifact root.
2026-06-08 02:04:59 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 492b454141 fix(archive): read zip central directory without inflating members
- Parsed zip metadata via central directory and lazy ranged reads.
- Inflated member contents only when a specific entry is read.
- Prevented large or corrupt zips from freezing directory reads.
2026-06-05 17:38:24 +02:00
can1357 1951081ad5 test(coding-agent): removed redundant AsyncJobManager.setInstance calls
- Relied on session-injected asyncJobManager instead of global instance.
2026-06-05 15:57:10 +02:00
can1357 07998fcc09 test(coding-agent): injected asyncJobManager via test session deps
- Passed asyncJobManager through test tool session dependencies.
- Replaced AsyncJobManager.setInstance with explicit session injection.
2026-06-05 15:55:25 +02:00
can1357 2eb39bae34 test(coding-agent): stabilized coding-agent tests with safer teardown and timing tweaks
- Updated interactive-mode plan review tests to capture shared fixtures, clear references, and run cleanup with explicit garbage collection before disposal.
- Increased the MCP HTTP transport test connection timeout from 200ms to 1,000ms.
- Adjusted the tool streaming command delay and tightened a start-pending-submission spy type in tests for better stability and type accuracy.
2026-05-31 04:25:14 +02:00
can1357 b777fe34f9 test: hardened test isolation and reliability across test suites
- Added `omfg` escape handler spies alongside existing `btw` spies in input controller tests.
- Introduced `waitForRenderedText` helper and settings lifecycle hooks to fix flaky apply-patch renderer tests.
- Raised `bash.autoBackground.thresholdMs` to avoid timing-sensitive test failures.
- Wrapped `runSearchQuery` calls with isolated `AuthStorage` instances to prevent shared state leaks.
2026-05-31 04:03:05 +02:00
can1357 725539aeeb feat(read): conditioned inspect_image docs on feature flag
- Updated read tool prompt to show alternate image description when inspect_image is disabled.
- Passed INSPECT_IMAGE_ENABLED flag into prompt rendering context.
- Added test verifying description omits inspect_image references when disabled.
2026-05-30 21:24:14 +02:00
can1357 a53acf1431 test(coding-agent): updated hashline preview tests to use snapshot-tagged headers
- Updated hashline streaming preview tests to generate snapshot-tagged section headers and use an in-memory snapshot store.
- Replaced untagged file markers in multi-section preview inputs with `formatHashlineHeader` values derived from recorded file contents.
- Changed the bash command error test to expect a returned `isError` result with exit code 1 instead of a rejected promise.
2026-05-30 17:51:11 +02:00
can1357 41d509eff0 feat(find): grouped output by directory and clamped limit to 200
- Changed output format to group results under `# /` headers to reduce token usage for shared path prefixes.
- Clamped the `limit` parameter to 1-200 (default 200) instead of the previous 1000.
- Updated tests to assert against raw file lists instead of parsed text output.
2026-05-27 13:13:11 +02:00
can1357 3e5b0b2340 feat(coding-agent/tools): added bash wall-time tracking to results and renderer
- Measured bash wall-clock duration for direct, terminal-bridge, and interactive execution paths.
- Recorded wall time in result notices and details, then stripped the duplicated literal notice during shell rendering.
- Updated the renderer to include wall time in the status label and added tests for the new wall-time behavior.
2026-05-27 01:53:34 +02:00
can1357 c31daf0805 feat(coding-agent): expanded find tool to return matching directories with slash markers
- Dropped the `fileType: natives.FileType.File` restriction so glob searches can return directories as well as files.
- Updated the find tool prompt to document directory results and trailing-slash output.
- Added tests verifying directory matches are included and emitted with a trailing `/`.
2026-05-26 20:22:25 +02:00
can1357 30793c1655 refactor: restructured hashline to use file-level hash validation with colon separators
- Replaced per-line hash anchors with file-level hash validation in hashline format, changing anchor syntax from LINE+HASH to bare LINE numbers.
- Simplified hashline line separator from pipe (|) to colon (:) and replaced replace operator (->) with colon, added delete operator (!) for explicit line deletion.
- Implemented file-read snapshot caching with multi-snapshot ring buffer per path and file-hash-based recovery to detect and recover from stale edits.
- Refactored hashline grammar, parser, and execution to support file-level hash binding, anchor-scoped validation, and structural bracket warnings for delete operations.
- Updated documentation and test fixtures to reflect new hashline syntax with file hashes, colon separators, and delete operator throughout.
2026-05-26 13:25:22 +02:00
can1357 6ca92dd08d test(coding-agent): increased bash abort test timeout and runtime
- Extended the abort scenario shell command from a 5-second sleep to 60 seconds.
- Raised the corresponding test timeout to 15,000 ms for the abort test case.
2026-05-22 18:06:50 +09:00
can1357 90b134ca4c test: replaced real timers and sleeps with deterministic test hooks
- Added `providerRetryWait` and `retryWait` hooks to stream/usage options so tests bypass real scheduler delays.
- Parameterized GitHub Copilot poll intervals and Copilot model retry base delay for fast test execution.
- Replaced `Bun.sleep`/`setTimeout` polling loops with `AbortSignal` event listeners in agent session tests.
- Consolidated auth-gateway E2E helpers into a shared `test/helpers` module, eliminating duplicated `checkGatewayAvailable` implementations.
- Migrated credential-disabled tests from SQLite-backed stores to an in-memory store, removing temp-dir lifecycle overhead.
2026-05-17 04:02:09 +02:00
can1357 85515f35a6 fix(coding-agent): added gap separators between noncontiguous search matches
- SearchTool now tracked the last emitted line and inserted ellipsis markers when noncontiguous match blocks were output.
- Display output gap markers were padded to align with code-frame gutters.
- Added a regression test that verified a no-context search emits an ellipsis between separated matches in the same file.
2026-05-15 23:46:23 +02:00
can1357 c7d04f3a13 fix(coding-agent): constrained bash cwd auto-detect regex to single-line cd commands
- Updated BashTool's leading `cd` regex to stop matching newline characters so cwd extraction only applies to a single-line `cd ... &&` prefix.
- Added a regression test for multiline commands with a later-line `&&` to ensure each line of the script executes normally.
2026-05-13 19:19:50 +02:00
can1357 9828764cb4 feat(scripts-session-stats): added session-stats search relevance plotting
- Changed multi-file search paging to skip whole files and page results in file windows.
- Added per-file match caps, round-robin file selection, and new file-limit truncation reporting.
- Replaced match/result limit metadata with fileLimitReached and perFileLimitReached.
- Lowered read.defaultLimit default to 300 with 1 lead and 3 trailing context lines.
- Replaced the search skip test with file-pagination coverage and added per-file cap tests.
- Added session-stats analytics tooling to classify searches, detect repeats, and render relevance plots.
2026-05-13 11:38:09 +02:00
can1357 a541a63547 feat: added read-selector analyzers and coverage plotting tools
- Updated read range expansion to use 1 leading and 3 trailing context lines.
- Changed read.defaultLimit from 500 to 300 in settings defaults.
- Updated read docs and tests to reflect the asymmetric context line behavior.
- Added read-selector analyzers and replay simulators to evaluate coverage and savings.
- Added plotting tools that output new session-stats PNG dashboards from local usage data.
2026-05-13 11:33:42 +02:00
can1357 28b9ce7a0c feat(coding-agent): added middle-elision caps to OutputSink truncation
- Added `tools.artifactHeadBytes` and `tools.outputMaxColumns` settings with defaults in `SETTINGS_SCHEMA`.
- Expanded `OutputSink` with `headBytes`/`maxColumns` and middle truncate logic with elision markers and tracking.
- Updated output-meta to resolve sink settings, emit truncation metrics, and use `truncateMiddle` for spills.
- Integrated head and column limits into JS/Python/Bash/SSH/read output flows, with `:raw` skipping read truncation.
- Documented new output middle-elision and column-cap behavior in `CHANGELOG.md`.
- Added truncation tests for `OutputSink`, `truncateMiddle`, and read-tool line handling.
2026-05-13 11:19:11 +02:00
can1357 1bde755933 feat(coding-agent): added global singletons for URL protocol handlers
- Added process-wide singleton instances for InternalUrlRouter, AsyncJobManager, and MCPManager.
- Changed internal URL protocols to resolve through registered sessions and scan all active roots/datasets for matches.
- Refactored agent, artifact, memory, rule, skill, jobs, and mcp handlers to use shared manager and rule/skill state.
- Removed per-session protocol/tool wiring and switched tests to initialize and reset global singleton state.
2026-05-12 05:07:52 +02:00
can1357 7223b800b0 fix(coding-agent/tools): expanded read ranges with three-line anchor context
- Expanded read tool range calculations to include optional leading and trailing context lines around user-requested offsets and limits.
- Added shared context range expansion logic so line-slice and streaming reads return anchor-safe windows while preserving existing truncation behavior.
- Updated read-tool tests to assert boundary-offset, limit, and combined offset-limit reads now include the expected ±3 lines of context.
2026-05-09 05:22:23 +02:00
can1357 cbfdc0a212 feat(coding-agent): added notebook-aware .ipynb edit/read path helpers
- Added `.ipynb` detection and editable-cell conversion utilities, including merge/serialize helpers in `edit/notebook`.
- Rerouted hashline, patch, and replace edit flows, plus read/write paths, through notebook-aware helpers before persistence.
- Removed the dedicated `notebook` tool, its schema flags, renderer, and built-in registration/settings checks.
- Updated notebook read behavior, docs, and tests so `.ipynb` reads return editable `# %%` cells and edits reserialize to JSON.
2026-05-07 04:16:11 +02:00
can1357 112a0ad632 refactor(coding-agent): exported search default match limit for test reuse
- Exported DEFAULT_MATCH_LIMIT from the search tool module.
- Updated the search tool tests to import DEFAULT_MATCH_LIMIT from the module.
- Updated default-limit assertions and messages to use DEFAULT_MATCH_LIMIT instead of hardcoded values.
2026-05-05 16:15:21 +02:00
can1357 8c323666be feat: added ordered systemPrompt arrays and normalized context prompts
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
2026-05-04 15:20:26 +02:00
can1357 ebd17597f4 feat(coding-agent): added read tool summarize mode with depth limits
- Added DirectoryTree, DirectoryTreeOptions, and buildDirectoryTree exports for configurable tree rendering.
- Changed read tool directory output to use buildDirectoryTree with depth and exclusion limits.
- Added read.summarize settings and summary-mode parseable read output behavior when no selector is used.
- Added tests for truncated root/child listings and hidden or excluded entry filtering.
2026-05-04 05:58:43 +02:00
can1357 a019994091 feat: added tree-sitter summarizeCode support and N-API summary exports
- Updated read schema, path utilities, and dispatch to parse selectors from :raw/:L suffixes on path, removing standalone sel usage.
- Changed truncation/error notices to continue with :<nextOffset> and :1 guidance for read and sqlite pagination.
- Added summarizeCode support with tree-sitter summaries, read.summarize settings, and N-API Summary types/exports.
- Added tests and docs updates for path-embedded selectors, summary behavior, and explicit raw/offset SQL/read cases.
2026-05-04 04:24:33 +02:00
can1357 f0e9a830ff refactor(packages/coding-agent): migrated search args to paths arrays
- Switched search/ast_grep/ast_edit/find inputs from scalar path fields to required `paths` arrays.
- Reworked path resolution to normalize and expand each `paths` entry via explicit helpers in `src/tools/path-utils.ts`.
- Updated search and find tool-call rendering to display explicit `paths` values in mode overlays and export views.
- Updated tool prompt docs and examples to document `paths` array inputs for search, grep, find, and AST tools.
- Raised the default search match limit from 20 to 500 and updated limit-reached messaging.
2026-05-02 04:34:39 +02:00
Can Bölük 7f24cb67a3 test(coding-agent): apply biome formatting 2026-04-29 06:35:00 +02:00
Can Bölük 51047564a5 test(coding-agent): align test path expectations with formatPathRelativeToCwd output 2026-04-29 06:24:32 +02:00
Can Bölük 7f5b450511 test(coding-agent): use long lines so byte limit triggers in read tool truncation test 2026-04-29 06:01:19 +02:00
can1357 5944fe34c0 feat(coding-agent): implemented shared single-path edit request handling
- Required top-level `path` in edit requests, removed per-entry `path` fields, and updated docs/tests to match.
- Updated patch/hashline/atom streaming preview generation to return a single request-level path diff preview.
- Changed edit execution to route patch/replace/atom/hashline through single-path handlers sharing `path`.
- Added `scripts/analyze-edit-formats` Go CLI with reports to audit edit-tool usage from session JSONL logs.
2026-04-28 05:21:24 +02:00
can1357 a3f8f122cc feat(coding-agent): renamed grep to search in runtime mappings
- Renamed the built-in `grep` content-search tool to `search` across settings, schemas, and SDK exports.
- Switched execution wiring so `Task`, `Plan`, cursor, and shell mapping now invoke `search` instead of `grep`.
- Updated prompts, plan-mode docs, and example tool lists to replace `grep`/`ls` references with `search` guidance.
- Aligned `Grep*`/`grep` event, renderer, and hook types to `Search*`/`search` across runtime and tests.
- Documented and fixed `search` result rendering budget behavior and added internal-URL/path-list transcript notes.
2026-04-27 21:55:31 +02:00
can1357 c8d6efa395 style(coding-agent/tools): removed tree glyph from grouped file headers
- Simplified diagnostic grouping output in the LSP utility to use plain `## file` headers.
- Adjusted AstEdit, AstGrep, and Grep tools to emit file headings without the `└─` tree prefix.
- Removed the GrepTool context legend text and its related guard variable from output construction.
2026-04-26 19:16:36 +02:00
can1357 3b82ecf242 feat(coding-agent/tools): updated read selector syntax to support numeric and plus ranges
- Updated `read` selector parsing for file and URL reads to accept optional leading `L` and `+` count-style ranges.
- Adjusted truncation notices and schema/help text to emit and suggest `sel` offsets without the `L` prefix, including continuation and suggestion messages.
- Updated hashline/output parsing and related tests to recognize the new `sel` formatting in truncation notices.
2026-04-26 11:48:45 +02:00