- Improved the leaked-thinking stream projector to clone and sync native tool-call blocks directly.
- Eliminated the need for placeholder IDs and complex rekeying logic in the event controller and argument reveal module.
- Simplified native tool-call validation in owned-stream processing by requiring only a non-empty name.
- Added comprehensive unit tests to ensure tool-call IDs and partial JSON parameters remain intact during healing.
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Thin OpenAI-compatible proxies that omit context_length / max_model_len on
/v1/models made every discovered model fall back to
DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a
bundled model with a much larger intrinsic window. discoverProxyModels
and discoverLiteLLMModels already resolve ids against the bundled
reference index; discoverOpenAIModelsList (which also backs lm-studio
discovery) now does the same.
Behavior:
- Build the reference index once outside the loop and resolve each item
via resolveModelReference().
- contextWindow precedence keeps provider-reported values authoritative:
item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow
?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW.
- maxTokens uses reference?.maxTokens when available, otherwise the
api-specific discovery default, capped at contextWindow so a bundled
ref for a larger sibling can never over-request output tokens.
- name / reasoning / thinking / input inherit from the reference; native
lm-studio metadata still wins for input modality.
- Provider-specific baseUrl, headers, and local-unknown cost stay local.
- OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole
/ supportsReasoningEffort all false) to match the proxy sibling.
Also updated two pre-existing regression tests that used
deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional"
ids to exercise the default-fallback branch. Those model names have since
been added to the bundled catalog, so the tests were renamed to
vllm-lab-fork-* ids that unambiguously miss the reference index while
preserving each test's original default-fallback intent.
Fixes#3983
Added a cross-turn tool-call loop guard that hashes canonical tool names and arguments, ignores intent metadata, and injects a hidden redirect when identical calls reach the configured threshold.
Fixes#3971
- Introduced an `#editVariantCache` to memoize resolved edit modes for model variants.
- Replaced the generic `shallowStringRecord` helper with specialized, type-safe parsing methods for model variants and roles.
- Invalidated the cached edit variants during settings rebuilds and verified correct cache refreshment across project directories.
- Moved Streamable HTTP request and notify timeout cleanup after response body consumption.\n- Added regression coverage for stalled request JSON bodies and stalled notify error bodies.\n\nFixes #3974
The subagent system prompt rendered `{{jtdToTypeScript outputSchema}}` as a
bare TypeScript interface with the text "Your result MUST match this
TypeScript interface". The yield tool actually nests the user schema under
`result.data`, so the LLM pattern-matched on the visually dominant code
block and put the payload directly in `result.data`, tripping schema
validation repeatedly. In the worst reported case a subagent used all 3
retry attempts, had validation dropped, and lost its audit output entirely.
Add a `renderYieldSchema` Handlebars helper that renders the schema inside
`result: { data: … }` and swap the system-prompt block to use it, so the
model sees the exact envelope the yield tool expects. Multi-line object
schemas, scalars, unions, and array-of-object schemas all round-trip
cleanly with the new helper.
Fixes#3972
Two termination boundaries in the browser tool leaked browser-owned OS resources into the long-lived coding-agent process.
1. Aborted 'open' published an orphan. #open wrapped acquisition in untilAborted, which rejects its outer wrapper on abort but lets the inner launch resolve in the background; acquireBrowser then unconditionally stored the resolved handle in the module-global browsers map. releaseAllTabs walks tabs, not browsers, so the refCount:0 handle stayed alive to process exit.
2. Session dispose had no browser teardown. Browser/tab state lives in module-global maps, and AgentSession.dispose() had no hook to walk them, so headless/spawned Chromium the session opened survived it.
acquireBrowser now short-circuits before launch on a pre-aborted signal and disposes the handle when the launch completes after abort. TabSession records the creating session's id (opts.ownerSessionId, threaded through BrowserTool.#open), preserved across reuse so a subagent re-driving an existing tab does not yank teardown responsibility. AgentSession.dispose() invokes releaseTabsForOwner bounded by withTimeout(3s), mirroring the async-job/MCP disposal pattern.
Regression tests exercise both boundaries via spied CmuxSocketClient (no real puppeteer/socket) and cover: pre-aborted open short-circuit, aborted-mid-launch cleanup, releaseTabsForOwner reaping only owned tabs, and reuse preserving original ownership.
Fixes#3963
Two client-level paths in the LSP tool bypassed the combined tool-timeout/
caller abort signal built in `LspTool.execute`, so a wedged server hung
past the advertised tool deadline and past user cancellation:
- `getOrCreateClient` took no `AbortSignal` and its `initialize`
`sendRequest` was invoked with `signal = undefined`. With no signal
and no explicit `timeoutMs`, `sendRequest` fell back to the hard-coded
`DEFAULT_REQUEST_TIMEOUT_MS = 30000` internal timer, so a first-use
`lsp` call against a server that wedged in `initialize` ignored the
20s tool default (and any user-supplied shorter `timeout`) until the
30s internal timer fired.
- `writeMessage`/`queueWriteMessage`/`sendNotification` had no timeout
and no signal, so a `textDocument/didOpen`/`didChange`/`didSave` sent
to a server that stopped draining stdin awaited `sink.flush()`
forever. Because writes serialize through `client.writeQueue`, every
later op on the client stalled behind the stuck flush too.
Thread the caller `AbortSignal` through `getOrCreateClient` (initialize
+ initialized notification) and through `sendNotification` /
`queueWriteMessage` / `writeMessage` so the sink flush is raced against
the signal. On abort, tear the client down: kill the process and evict
it from the active-clients map so the next `getOrCreateClient` call
spawns a fresh server instead of queueing behind the wedged sink.
Update the LSP tool callsites and internal helpers
(`captureDiagnosticVersions`, `captureOpenFileVersions`,
`syncFileContent`, `notifyFileSaved`, `formatContent`,
`getDiagnosticsForFile`, `reloadServer`, and the rename didClose /
didRenameFiles path) to forward their operation signal.
Warmup keeps its short explicit `initTimeoutMs` and passes no caller
signal; `sendRequest`'s existing `timeoutMs ?? (signal ? undefined : DEFAULT)`
policy still uses that fixed timer.
Fixes#3962
- Replaced leaf-to-root unshift path assembly with push plus one reverse in buildSessionContext and SessionEntryIndex.pathTo.
- Added regression coverage that keeps deep linear context and branch paths root-to-leaf without Array.unshift work.
Fixes#3961
Emitted partial write-tool updates before filesystem, archive, SQLite, internal URL, and conflict writes so the TUI can render execution-phase progress instead of waiting for the final result.
Updated the write renderer to keep partial results pending, show the progress snapshot, and suppress diagnostics until the final result.
Fixes#3960
Stream Python eval shell helper stdout in fixed-size chunks instead of buffering through subprocess.run or newline-bound text iteration.
Added regression coverage for !cmd and newline-free %%bash streaming.
Fixes#3950
Bun.sleep(timeoutMs).then(...) leaves an uncancellable timer registered
with the event loop, so every successful handler race in the runner
leaked one — a completed tool_call/tool_result handler could delay
non-interactive CLI exit by up to the 30s default cap. Verified with a
subprocess exit-time probe: buggy pattern exits in ~5000ms for a 5s
timeout, setTimeout+clearTimeout pattern exits in ~17ms.
Extract a raceHandlerWithTimeout helper backed by setTimeout with a
finally-scoped clearTimeout, and route both #runHandlerWithTimeout
(pre-existing latent leak) and emitToolCall (introduced in the same PR)
through it. No behavior change on the timeout branch.
Addresses review on #3951 from chatgpt-codex-connector[bot].
StdioTransport.request() awaited stdin.write() and stdin.flush() before returning the internal deferred promise. When the child stopped draining stdin (wedged process, or full OS pipe buffer with no reader), Bun's FileSink returned a pending Promise that never settled — the async function got stuck above 'return promise', past the timeout timer and the abort handler. cleanup() + reject() still ran on the inner deferred, but the outer async-function promise never adopted it, so the caller's await hung forever and the deferred rejection surfaced as an unhandled promise rejection.
Send the frame without awaiting: sync EPIPE throws (Windows) still reject the request immediately; async EPIPE rejections (POSIX processTicksAndRejections) are wired to the same reject() via a guarded failFromSend handler that no-ops after cleanup(). The returned promise now settles from the response, the timer, the abort signal, or the read loop's transport-close broadcast.
Regression test spawns 'sleep 60' (POSIX only), sends a 1MB tools/call payload past the pipe buffer, and asserts the deferred rejects with the timeout error before the outer window elapses and produces no orphaned unhandled rejections.
Fixes#3945
The in-process fd builtin passed no-op heartbeats to pi_walker for
both its gitignore-respecting fallback path (`collect_with_heartbeat`
in `search`) and its fast path (`for_each_entry_with_heartbeat` in
`try_search_fast`), so cancellation of a large or slow directory walk
was deferred until traversal completed. The shell wrapper flips the
shared `AtomicBool` cancel flag when the runtime cancellation token
fires and then awaits the blocking task; with no heartbeat hookup the
walker had no way to observe the flag mid-walk and kept collecting the
whole tree before the wrapper could return exit 130.
Introduce `cancel_heartbeat(&AtomicBool)` — the walker-level heartbeat
that returns `io::ErrorKind::Interrupted` when the flag is set — and
plug it into both walker calls. Both call sites recognize the resulting
`WalkError::Interrupted` alongside `cancelled` and break silently
instead of surfacing an `fd:` diagnostic on stderr; the shell wrapper
owns the user-visible exit code.
Regression cover: a walker-level test pre-sets the cancel flag and
asserts `collect_with_heartbeat(cancel_heartbeat(&flag))` surfaces
`WalkError::Interrupted` instead of collecting the tree; two
higher-level `search` tests exercise the silent break for both the
fallback and fast paths; a fourth pins the non-cancelled contract so
the added heartbeat can't stall normal searches. Neuter the helper to
a no-op and the walker-level test fails with the pre-fix `WalkOutcome`
showing every entry scanned — the exact bug the issue reports.
Fixes#3949
emitToolCall awaited each extension handler directly (runner.ts:704-706),
bypassing the #runHandlerWithTimeout wrapper every other subscribed event
routes through. A tool_call handler that never resolves parked
ExtensionToolWrapper.execute indefinitely, freezing tool dispatch even
though the symmetric emitToolResult path has always been timeout-protected.
Race each tool_call handler against Bun.sleep(extensionHandlerTimeoutMs)
inline (the shared wrapper swallows errors, and this callsite is
fail-closed). On timeout: emit an ExtensionError with event: 'tool_call',
log a warning, and return { block: true, reason: 'Extension <path>
timed out after <ms>ms' } — symmetric with the existing per-handler error
branch. Fail-closed is the correct policy for a pre-execution gate: an
unresponsive extension MUST NOT be silent consent to run the tool.
Fixes#3948
- Added `hasUsableNativeToolCall` helper to verify that a streaming tool call has non-empty, trimmed name and id values.
- Retain projection initialization and updates on subsequent deltas if the provider emits native tool identifiers late.
- Guard tool call synchronization and late salvage logic to prevent empty or invalid placeholders from corrupting streaming state.
Kept the STT subprocess referenced while download and stream requests are pending so setup cannot exit before the worker answers.
Propagated worker download errors to setup callers and verified completed downloads leave the expected cache files.
Fixes#3939
Added a contributor-facing native crate map (docs/native-crates.md) covering pi-natives, pi-shell, pi-ast, pi-iso, pi-walker, pi_uu_grep, pi-uutils-ctx, and vendored brush crates, and linked it from natives-architecture.md and user-facing-packages.md.
Added a docs-index tool coverage test asserting every BUILTIN_TOOL_NAMES entry and injected custom tool (generate_image, tts) has a docs/tools/<name>.md page served by omp://.
Inlined tiny fail/buildPayloadText/checkDocsIndexFreshness helpers in generate-docs-index.ts per the project rule against single-expression named functions.
Fixes#3934
Retained only the requested AST search page window in native ast_grep/ast_match and the coding-agent multi-target wrapper while preserving exact totals.
Fixes#3935
Added root omp docs for memory_edit, learn, manage_skill, generate_image, and tts, plus package-level coverage for user-facing README-only CLIs.
Added a docs-index freshness check to package check and made gen:bundle generate and reset the docs embed itself.
Fixes#3934
The grep and rg in-process builtins passed no-op heartbeats to pi_walker
when recursing into a directory, so the uutils scope's cancel flag was
ignored mid-walk. The shell wrapper sets that flag on abort/timeout and
then awaits the blocking task — with no heartbeat hookup, a cancelled
recursive grep/rg waited for the whole tree to be scanned before the
shell could return exit 130.
Add pi_uutils_ctx::is_cancelled() and have grep's search_dir, rg's
search_dir, and rg's collect_filtered_files supply heartbeats that
return io::ErrorKind::Interrupted when the flag is set. The walker maps
that to WalkError::Interrupted, which the utilities now silently treat
as a harness cancellation (no spurious 'native directory scan
interrupted' on the command's stderr — the shell wrapper owns the
user-visible exit code).
Regression tests pre-set the cancel flag and assert grep -r /
rg / rg --files exit without scanning the matching file in the tree.
Fixes#3933