Commit Graph

454 Commits

Author SHA1 Message Date
can1357 b522fde56d perf(sqlite-reader): replaced full COUNT(*) scan with bounded row probing
- Added ROW_COUNT_PROBE_CAP to limit rows scanned when counting tables, preventing JS thread freezes on large databases.
- Used sqlite_stat1 estimates for tables exceeding the cap; exact counts only for provably small tables.
- Introduced TableRowCount type with exact/estimate/atLeast variants reflected in rendered output.
2026-06-02 05:24:21 +02:00
can1357 fbdc064186 feat(lsp): added session-scoped LSP diagnostics deduplication
- Added `DiagnosticsLedger` to track diagnostics already surfaced per file, suppressing repeats within a session.
- Wired dedup into both edit and write tools via a `transformDiagnostics` hook on the writethrough pipeline.
- Added `lsp.diagnosticsDeduplicate` setting (default: true) to control the behavior.
2026-06-01 18:04:05 +02:00
can1357 dbd9489010 refactor(eval): changed timeout from inactivity to wall-clock budget
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
2026-06-01 17:17:00 +02:00
Can Bölük 890d63fdb1 Merge branch 'main' into fix-ask-option-descriptions 2026-06-01 18:14:42 +03:00
can1357 6a63b17711 Merge remote-tracking branch 'origin/farm/cdb7e4ab/fix-find-renderer-string-paths' 2026-06-01 14:36:57 +02:00
rimless-casualty 3c7d50d292 Improve ask option rendering 2026-06-01 17:37:26 +08:00
roboomp 8580ba248c fix(tool): handled string paths in find renderer
Guarded find renderer path summaries so raw pre-validation string paths render instead of throwing. Added coverage for pending, fallback, empty, and detailed result render paths.\n\nFixes #1622
2026-06-01 06:22:27 +00:00
can1357 7e11bea8f4 fix(coding-agent): repaired per-field double-encoded JSON in task tool
- Added `repairDoubleEncodedJsonString` to unescape fields double-encoded by the model (e.g. literal `\n`, `\"`, `\uXXXX` in `context`/`assignment`/`description`).
- Scoped repair to natural-language fields only, leaving code-bearing tools untouched.
- Applied repair on both render and execution paths in `TaskTool`.
2026-05-31 20:17:38 +02:00
roboomp f4ed2698dd fix(lsp): shut down servers with exit notification
Send the LSP exit notification after a successful shutdown response before falling back to process termination. Add a regression test that fails when a server receives shutdown but not exit.\n\nFixes #1593
2026-05-31 15:26:13 +00:00
can1357 a8d50879df chore: reformat 2026-05-31 13:22:50 +02:00
can1357 d016150d01 fix(ai): prevented provider retry after streaming unsafe content
- Added `!streamedReplayUnsafeContent` guard to `canRetryProviderFailure` to avoid replaying unsafe content on retry.
- Updated test fixtures to supply a required `workspaceTree` parameter via a shared `emptyWorkspaceTree` helper.
2026-05-31 13:21:15 +02:00
can1357 9fabd5e4d5 feat(coding-agent): added onStatus in eval backends for status streams
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
2026-05-31 08:45:12 +02:00
can1357 59c8749ad7 fix(tui): resolved tui status footer truncation using VT stripping
- Replaced ad-hoc ANSI/VT stripping regexes with `stripVTControlCharacters` in status text handling and related tests.
- Updated status footer rendering to truncate using `truncateToWidth` and visible width after VT stripping.
- Extended tui cursor handling and rendering to strip markers from all lines and fit repaint/append-tail lines to width.
- Expanded deterministic render tests with overlay-aware assertions and recorded the truncation/cursor-marker behavior in changelogs.
2026-05-31 07:54:23 +02:00
oldschoola cd578a86d1 fix(coding-agent,ai): file-lock race + render-utils sanitization + google named-tool routing
F1 (coding-agent): close the withFileLock mkdir-vs-writeLockInfo race that
let a losing contender wipe the winner's freshly-created lock directory.
Every lock now carries a per-process UUID token; releaseLock verifies the
token before fs.rm, and isLockStale no longer treats an info-less but
fresh dir (or a dir that vanished mid-check) as stale.

F4 (coding-agent): sanitize tabs and truncate oversized error strings in
formatErrorMessage so error renderings that embed file content
(apply_patch, hashline, etc.) cannot break terminal alignment or
overflow the line width.

F7 (ai): support named-tool routing on Google providers. Widens
GoogleSharedStreamOptions.toolChoice and GoogleGeminiCliOptions.toolChoice
to accept { mode: 'ANY'; allowedFunctionNames }. mapGoogleToolChoice
now converts ToolChoice { type: 'tool'|'function', name } to the wire
shape (mirroring mapAnthropicToolChoice). buildGoogleGenerateContentParams
and the gemini-cli request serializer honor the allow-list.
2026-05-31 04:45:57 +02:00
can1357 4ab0360f65 test(coding-agent/tools): updated internal URL resolution test fixture wording
- Updated the internal URL resolution test input pattern to use a new search phrase.
- Updated the expected assertion string to match the revised phrase in command output.
2026-05-31 04:41:36 +02:00
can1357 1dba122c53 chore: updated docs 2026-05-31 04:36:14 +02:00
can1357 756f32f687 feat(coding-agent-tools): added virtual multi-file search scope expansion
- Expanded search scope handling for virtual multi-file targets and executed grep only when searchable paths existed.
- Merged `searchVirtualResources` results into main output and rendered internal URL matches as accent lines.
- Updated grouped-file output to detect URL-like paths and keep full URL headers for root grouping.
- Documented URL path/range behavior and added tests for doc routing, missing-content errors, and `omp://` expansion.
2026-05-31 04:32:35 +02:00
can1357 f21c23277a feat(coding-agent/tools): implemented internal URL routing in SearchTool
- Added virtual internal URL path resolution in `SearchTool` via `InternalUrlRouter` for in-memory search.
- Added `omp://` root expansion so `search` resolves and scans each completion target.
- Fixed `SearchTool` handling of internal URLs without `sourcePath` by returning virtual matches instead of `Path not found`.
- Extended `edit-renderer.test.ts` coverage for normalizing raw streamed text in custom text renderers.
2026-05-31 04:28:39 +02:00
can1357 ce2e8ce7cd fix(coding-agent): added raw partialJson preview support for hashline and apply_patch tools
- Added helper logic to treat raw non-JSON `__partialJson` values as edit `input` for hashline and apply_patch args.
- Updated edit argument preparation, preview generation, and streaming fallback rendering to use that derived input.
- Added renderer tests verifying raw hashline and apply_patch partial streams render target paths and patch content.
2026-05-31 04:24:40 +02:00
can1357 b777fe34f9 test: hardened test isolation and reliability across test suites
- Added `omfg` escape handler spies alongside existing `btw` spies in input controller tests.
- Introduced `waitForRenderedText` helper and settings lifecycle hooks to fix flaky apply-patch renderer tests.
- Raised `bash.autoBackground.thresholdMs` to avoid timing-sensitive test failures.
- Wrapped `runSearchQuery` calls with isolated `AuthStorage` instances to prevent shared state leaks.
2026-05-31 04:03:05 +02:00
can1357 453e219732 fix(edit): fixed hashline preview matching for live and recovered file snapshots
- Updated hashline preview diffing to route section edits through shared apply/resolve logic and partial streaming parsing.
- Allowed previews to accept live content-hash matches immediately when snapshot records are missing.
- Enabled stale-tag recovery from the snapshot store for anchor-scoped edits and added coverage for snapshot-capture behavior and session-mismatch failures.
2026-05-31 03:48:19 +02:00
can1357 ebb9276393 feat(coding-agent): added animated pending border for bash/eval blocks
- Added clockwise sweeping dark segment animation to output block borders while bash/eval tool calls are pending/running.
- Changed bash renderCall to immediately render a full bordered block instead of a one-liner status preview, so silent commands show the framed block for their entire runtime.
- Added shimmerEnabled() helper and wired animate flag through OutputBlockOptions, CodeCellOptions, and shell/eval renderers.
2026-05-31 01:43:00 +02:00
can1357 dfa6007f36 feat(coding-agent): removed recipe tool and all runner implementations
- Deleted RecipeTool, runner logic, and all task runner backends (just, make, cargo, pkg, task).
- Removed recipe from BUILTIN_TOOLS, auto-injection in createTools, and HTML export renderer.
- Deleted recipe tool prompt template and runner module exports.
2026-05-31 01:42:47 +02:00
can1357 0eee5a4019 feat(coding-agent): added todo-write strike-through completion animation
- Added todo_write strike-frame animation timing and updated execution flow after completion finalizes.
- Added strike animation cancellation in cleanup to clear todo timer and reset frames when spinner is idle.
- Removed todo-closing state and timeout handling from interactive mode and simplified empty todo-list rendering.
- Reworked todo-write to compute completion transitions, track completedTasks, and render strike-through frames.
- Added test coverage for completedTasks, theme setup, and strike-through progression at hold-frame thresholds.
2026-05-30 18:58:48 +02:00
can1357 e831c2c758 chore: reformat 2026-05-30 18:08:51 +02:00
can1357 98de6510f0 fix(coding-agent/tools): renamed exit status label from "Status: exit N" to "Exit: N"
- Updated shell renderer footer label for non-zero exits to use the shorter `Exit: N` format.
- Updated changelog and tests to match the new label.
2026-05-30 16:35:34 +02:00
can1357 4356298182 fix(coding-agent/tools): handled non-zero bash exits as error results
- Tagged non-zero bash command completions as error results, capturing `exitCode` and keeping exit notices in returned text.
- Updated the shell renderer to hide duplicate exit notices from output while surfacing failed command status in the footer.
- Added tests for non-zero versus zero-exit bash results and footer rendering of failed commands.
2026-05-30 16:31:48 +02:00
can1357 318d553045 feat(tools): added memory inline renderers for retain/recall/reflect
- Added shared helpers and inline TUI renderers for retain, recall, and reflect tool outputs.
- Added tool registry entries for retain, recall, and reflect to use the new inline renderers.
- Updated changelog notes describing new memory inline rendering semantics and output headers.
- Added memory renderer tests for summary, truncation, streaming, and expand/collapse behavior.
2026-05-30 16:28:55 +02:00
can1357 2d49cfbfa5 fix(coding-agent): fixed hashline edit preview deduping for final stream completion
- Updated the edit preview coalescing key to include streaming state plus a hash so final non-streaming diffs are not skipped when payload bytes match.
- Used a hashed partial-stream key as fallback when tool args are not JSON-serializable, keeping deterministic dedup cache keys.
- Added a regression test that verifies single-line hashline streaming edits render the completed preview diff instead of "No changes would be made".
2026-05-30 16:11:24 +02:00
can1357 9ce250eb99 chore: remove deprecated calc tool as it is completely useless with eval 2026-05-30 04:05:26 +02:00
can1357 64daece638 fix(pi-natives): handled legacy Alt+letter pairs in mixed enhanced-keyboard mode
- Preserved Alt and Ctrl+Alt letter ESC-prefix matching when kitty_protocol_active is true to support mixed tmux and Kitty keyboard modes.
- Parsed two-byte ESC sequences before legacy sequence lookup so mixed-mode Meta pairs are treated as Alt letter keys instead of legacy aliases.
- Updated native and TUI key tests to verify Alt+letter and Alt+Shift+letter parsing and matching while enhanced mode is active.

Fixes #1511
2026-05-29 18:21:30 +02:00
can1357 01c34db450 feat(hashline): added full-file hash snapshots with 4-hex tags
- Replaced snapshot internals with full-file records and removed contiguous/sparse snapshot APIs.
- Added file-hash normalization, computed `computeFileHash`, and updated grammar/messages to 4-hex tags.
- Simplified recovery by checking whole-file hashes first, then applying merge-replay fallback after mismatches.
- Updated coding-agent tools to use `record`/`recordFileSnapshot` and skip hash headers for unsnapshotted large files.
- Expanded patcher and snapshot tests to verify 4-hex anchors, hash deduplication, and cache-capped behavior.
2026-05-29 18:12:08 +02:00
oldschoola 98fdcdcf63 feat(coding-agent): live cube, fuzzier matcher, auto-checkmark, close animation
Four refinements to the sticky Todos panel on top of the live SessionObserverRegistry linkage:

- Cube animates whenever any visible open todo is "live" (in_progress, or a still-pending todo with a matching in-flight subagent). The previous subagent-only gate left lone in_progress rows on the static '⟳' fallback; ticking on an orphan in_progress row is the correct "still open" signal.
- 'normalizeForTodoMatch' now collapses any non-alphanumeric run to one space, so subagent descriptions with '#', '.', ':' etc. match todo content that omits the punctuation. Fixes the case where 3 subagents were spawned but only 2 of 3 matching todos lit up because the matcher's normalizer collapsed whitespace but left '#' intact.
- New '#reconcileTodosWithSubagents' runs on every observer-registry change and auto-checkmarks any pending/in_progress todo whose content matches a 'status === "completed"' subagent description. Failed/aborted subagents intentionally don't auto-flip - those stay open for the user (or next agent turn) to decide.
- All-done close animation: when every visible task is closed, fold the panel away over ~1.4s. A 900ms celebratory frame holds the bright bold "Todos ✓" header so the user can read the final checkmarks, then a fade through 'muted' / 'dim' with rows progressively dropped from the bottom. '#todoClosingState' state machine plays the animation exactly once per open->all-closed transition and aborts cleanly if a new open task arrives mid-animation.

Verification:
  bun test test/tools/todo-write.test.ts → 24 pass / 0 fail (one new case for # punctuation tolerance)
  bun run check                          → biome + tsgo clean
2026-05-29 02:49:18 -07:00
oldschoola cea3ac53b0 feat(coding-agent): advance the sticky todo panel as tasks close
The always-on Todos panel above the editor pinned to the first 5 tasks of the active phase, so each todo_write flip mutated at most one row (color + strikethrough) and the +N more hint only shrank at end-of-phase. Marking task 1 done left tasks 6,7,... invisible until tasks 1-5 were all closed.

Introduce selectStickyTodoWindow(tasks, maxVisible=5) — returns up to 5 open (pending / in_progress) tasks in original phase order plus the count of remaining open tasks for +N more. When every task is closed, falls back to the trailing window (with +N more suppressed) so the panel keeps useful context until getActivePhase walks to the next phase. The collapsed branch of #renderTodoList now uses it; the expanded branch is untouched.
2026-05-29 00:16:36 -07:00
Vu Anh Nguyen 674d9b00a2 fix(codex): prefer gpt-5.5 for web search 2026-05-28 14:02:52 +07:00
can1357 7dd00c015b feat: hashline improvements for spark
- Redesigned hashline patch syntax from anchor-based (`A-B:`) to hunk-header format (`@@ A..B @@`) with unified-diff compatibility.
- Removed `autoDropPureInsertDuplicates` option and simplified apply behavior to preserve duplicated boundary and context lines.
- Changed repeat operator from `^A-B` to `&A..B` and range separator from `-` to `..` for consistency with hunk-header syntax.
- Added image resizing and dimension notes to eval tool output; improved write tool hashline header sanitation for legacy formats.
- Removed 521 lines of boundary-duplicate absorption code and simplified parser to auto-convert bare body rows and unified-diff contamination.
2026-05-28 03:06:52 +02:00
can1357 7fa55750f9 feat(hashline): introduced explicit range syntax and repeat edit kind for hashline
- Replaced anchor shorthand syntax with explicit range format (1: -> 1-1:) and removed ^/v sigils in favor of ^A-B repeat and A-B:- delete operations.
- Added repeat edit kind to support ^A-B syntax for copying lines A through B, and inline delete syntax A-B:- for range deletions.
- Removed after_anchor cursor kind and standalone delete rows; empty anchor blocks now produce blank-line replacements instead of deletions.
- Updated parser, tokenizer, and type system to discriminate literal and repeat payloads, and refactored apply/recovery logic to expand repeat edits into individual inserts.
- Updated coding-agent test fixtures and settings documentation to reflect new hashline syntax and behavior.
2026-05-27 22:51:21 +02:00
can1357 1dbd2a0659 fix(coding-agent): pin streaming diff preview to tail of the diff 2026-05-27 19:57:54 +02:00
roboomp 9362cbf13d style: bun run fix 2026-05-27 19:00:49 +02:00
roboomp efba782fa7 fix(tools): isolate read URL reader-mode fallback chain from remote stalls
A stalled Jina reader request shared the overall reader-mode AbortSignal
with the downstream trafilatura/lynx/native fallbacks. When Jina hung
until the budget timer fired, the shared signal aborted and the catch
handler's signal?.throwIfAborted() re-threw before any local fallback
ran.

- Bound Jina and Parallel extract to their own per-attempt sub-budget
  (REMOTE_READER_MAX_MS, capped at 10s) so a remote stall cannot consume
  the whole overall reader-mode budget.
- Catch handlers now rethrow only on real userSignal cancellation, not
  on remote sub-budget or overall budget expiry.
- Wrap trafilatura/lynx in their own try/catch so a subprocess failure
  or abort does not skip the in-process native renderer.
- Always attempt the native renderer last: it works on already-loaded
  HTML with no network or subprocess, so even an exhausted overall
  budget still yields a result.

Fixes #1449
2026-05-27 19:00:49 +02:00
Can Bölük f9c5484892 Merge pull request #1425 from oldschoola/fix/search-regex-error-prefix
fix(search): wrap native regex-build errors in ToolError
2026-05-27 19:53:02 +03:00
Can Bölük dc1eb8d96f Merge pull request #1446 from OutlineDriven/fix/xai-grok-oauth-stabilize
fix(ai,coding-agent): stabilize xAI Grok OAuth
2026-05-27 19:41:20 +03:00
metaphorics b76f39d3a6 fix(coding-agent): reopen approved plan on plan-mode reentry
Patch axis: extend

Displacement: net-zero; reuses existing plan reference state instead of adding persistence or overwriting approved artifacts

Rule violations averted: no approved-plan overwrite, no transcript format migration, no public CLI/API expansion

PASS/FAIL: PASS after plan-mode focused tests and package check. Note: system-prompt-templates has an unrelated HOME=/tmp path-shortening expectation failure.
2026-05-27 16:32:40 +00:00
metaphorics bb249cb911 test(coding-agent): mock authStorage + getProviderBaseUrl on image-gen xAI ctx
The Surface 1 commit added authStorage.hasNonEnvCredential as a credential
gate inside resolveXAIHttpCredentials, and the Surface 3 commit added
resolveXAIBaseURL which consults getProviderBaseUrl and getAll. The
existing image-gen xAI test mocked only getApiKeyForProvider on
modelRegistry, so the new code paths threw "undefined is not an object"
at runtime.

Add the missing mock surface:
  - authStorage.hasNonEnvCredential returns true for "xai-oauth" so the
    dedicated-credential gate routes through the xai-oauth branch (which
    the test's getApiKeyForProvider mock services).
  - getProviderBaseUrl returns undefined so resolveXAIBaseURL falls
    through to the XAI_BASE_URL / DEFAULT_BASE_URL leg, preserving the
    test's existing expectation that the request hits
    https://api.x.ai/v1/images/generations.
  - getAll returns [] so the per-model override check in
    resolveXAIBaseURL no-ops cleanly.

Op: correct
Restores: ref:feat/xai-grok-oauth@015437534 ref:feat/xai-grok-oauth@2c1abd7fa
2026-05-27 16:07:38 +00:00
can1357 5be5053696 fix(coding-agent): widen xAI test header captures via holder 2026-05-27 15:52:43 +02:00
can1357 7fd69dbb81 fix(coding-agent): format image-gen test 2026-05-27 15:50:13 +02:00
can1357 ff94f91104 feat: added extraBody support, xAI fixes, and image provider updates
- Added `extraBody` merging into OpenAI Responses request params.
- Fixed xAI OAuth redirect URI to fail fast on port conflicts.
- Exposed `antigravity` and `xai` as explicit `providers.image` options.
- Added `isImageProviderPreference` guard, replacing inline string checks.
- Fixed TTS tool to resolve output path relative to cwd and require write approval.
2026-05-27 15:20:55 +02:00
can1357 f051cc6a0e feat(tools): added multi-range line selectors and raw mode support for URLs and directories
- Added support for multi-range line selectors on URLs (e.g., `:5-10,20-30`) and combining `:raw` mode with line range selectors.
- Added support for line range selectors on directory listings with offset and limit parameters.
- Fixed `:raw` selector being ignored for JSON and feed URLs and directory listing line selectors dropping offset parameter.
- Added clear error message for line offset beyond directory listing end.
- Refactored URL parsing and directory reading to support multiple comma-separated ranges and improved line-based slicing logic.
- Added comprehensive test coverage for multi-range selectors, raw mode combinations, and directory range operations.
2026-05-27 15:03:12 +02:00
can1357 8a0329026c refactor(hashline): redesigned patch syntax from sigil-ops to anchor+payload model
- Replaced `LINE↑`/`LINE↓`/`A-B:` op sigils with unified `A-B:` anchor + `|`/`↑`/`↓` payload sigils.
- Added `mode: "replacement"` tag to insert edits so the applier distinguishes replace-bucket from insert-bucket lines.
- Removed lenient fallbacks (implicit continuation, inline payload acceptance, escaped delimiter stripping).
- Updated grammar, prompt, tokenizer, parser, applier, and messages to match the new format.
2026-05-27 14:45:01 +02:00
oldschoola db84b52c0c fix(search): wrap native regex-build errors in ToolError
The catch block at search.ts:480-485 tried to convert native regex-build
failures into clean `ToolError`s but checked for the prefix `"regex parse
error"` (lowercase). The native crate at `crates/pi-natives/src/grep.rs`
actually emits `"Regex error: "` (capital R, no "parse"). The branch was
unreachable: invalid patterns like `a[` leaked out as raw `Error` with a
stack trace instead of being wrapped in a structured `ToolError` for the
agent to feed back on.

Match against `/^regex(?: parse)? error/i` so both the actual native
prefix and any hypothetical `regex parse error: ...` variant are caught.
Rewrite the leading prefix to `Invalid regex: ` so the agent immediately
sees the failure mode.

Test: new `test/tools/search-invalid-regex.test.ts` asserts the pattern
`a[` rejects with an `instanceof ToolError` whose message matches `/regex/i`.
Fails on current main (raw `Error` escapes); passes with the fix.
2026-05-26 23:26:15 -07:00