Excluded running async result snapshots from live partial spinner intervals so finalized background tool rows do not repaint scrollback.
Added regression coverage for async bash snapshots staying static.
Started live partial tool spinner intervals for non-static tool blocks and forwarded spinner frames through eval and shell-style renderers.
Added regression coverage for live eval and shell preview spinner frames.
Fixes#4170
Add provider-native V2 remote compaction metadata to discovered OpenAI Codex models so context-full compaction uses the streaming compaction_trigger path instead of the legacy compact endpoint.
Expose the V2 remote compaction schema fields in models.yml and cover the Codex discovery metadata contract.
Fixes#4146
AGENTS.md forbids inline `await import()` — move `fs`, `path`, and `prompt`
to top-level namespace/named imports. The `prompt-templates` module import
already registers the Handlebars helper as a side-effect, so the render
calls still resolve the new `renderYieldSchema` helper.
- Updated exact substring match assertions to use word-boundary regular expressions.
- Prevents false-positive test failures when target strings overlap with other generated text.
- Improved the leaked-thinking stream projector to clone and sync native tool-call blocks directly.
- Eliminated the need for placeholder IDs and complex rekeying logic in the event controller and argument reveal module.
- Simplified native tool-call validation in owned-stream processing by requiring only a non-empty name.
- Added comprehensive unit tests to ensure tool-call IDs and partial JSON parameters remain intact during healing.
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Preserved the referenced model's OpenAI-compatible reasoning-effort support when openai-models-list discovery enriches a thin /v1/models payload. The discovered model still keeps conservative proxy-local store and developer-role defaults, but known reasoning models like gpt-5 no longer force supportsReasoningEffort false and trigger the omitReasoningEffort request path.
Added a regression assertion that a thin proxied gpt-5 keeps supportsReasoningEffort true and omitReasoningEffort false after reference enrichment.
Fixes#3983
Thin OpenAI-compatible proxies that omit context_length / max_model_len on
/v1/models made every discovered model fall back to
DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a
bundled model with a much larger intrinsic window. discoverProxyModels
and discoverLiteLLMModels already resolve ids against the bundled
reference index; discoverOpenAIModelsList (which also backs lm-studio
discovery) now does the same.
Behavior:
- Build the reference index once outside the loop and resolve each item
via resolveModelReference().
- contextWindow precedence keeps provider-reported values authoritative:
item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow
?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW.
- maxTokens uses reference?.maxTokens when available, otherwise the
api-specific discovery default, capped at contextWindow so a bundled
ref for a larger sibling can never over-request output tokens.
- name / reasoning / thinking / input inherit from the reference; native
lm-studio metadata still wins for input modality.
- Provider-specific baseUrl, headers, and local-unknown cost stay local.
- OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole
/ supportsReasoningEffort all false) to match the proxy sibling.
Also updated two pre-existing regression tests that used
deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional"
ids to exercise the default-fallback branch. Those model names have since
been added to the bundled catalog, so the tests were renamed to
vllm-lab-fork-* ids that unambiguously miss the reference index while
preserving each test's original default-fallback intent.
Fixes#3983
Added a cross-turn tool-call loop guard that hashes canonical tool names and arguments, ignores intent metadata, and injects a hidden redirect when identical calls reach the configured threshold.
Fixes#3971
- Introduced an `#editVariantCache` to memoize resolved edit modes for model variants.
- Replaced the generic `shallowStringRecord` helper with specialized, type-safe parsing methods for model variants and roles.
- Invalidated the cached edit variants during settings rebuilds and verified correct cache refreshment across project directories.
- Moved Streamable HTTP request and notify timeout cleanup after response body consumption.\n- Added regression coverage for stalled request JSON bodies and stalled notify error bodies.\n\nFixes #3974
The subagent system prompt rendered `{{jtdToTypeScript outputSchema}}` as a
bare TypeScript interface with the text "Your result MUST match this
TypeScript interface". The yield tool actually nests the user schema under
`result.data`, so the LLM pattern-matched on the visually dominant code
block and put the payload directly in `result.data`, tripping schema
validation repeatedly. In the worst reported case a subagent used all 3
retry attempts, had validation dropped, and lost its audit output entirely.
Add a `renderYieldSchema` Handlebars helper that renders the schema inside
`result: { data: … }` and swap the system-prompt block to use it, so the
model sees the exact envelope the yield tool expects. Multi-line object
schemas, scalars, unions, and array-of-object schemas all round-trip
cleanly with the new helper.
Fixes#3972
Wrapped cmux page, browser, and tab globals with per-run abort checks so stale continuations cannot reuse the long-lived CmuxTab after timeout.
Fixes#3964
Distinguish aborts that race an active sink.flush() from aborts that happen
before a queued write starts. Only the former leaves the sink flush pending
and requires killing/evicting the LSP client; pre-write aborts should reject
that caller without disrupting unrelated in-flight operations.
Add a regression with one notification blocked in flush and a second queued
notification whose signal aborts before its write starts, asserting the shared
client is not killed and only the first message is written.
Caller cancellations and tool timeout signals are transient initialize failures.
Do not put them in the three-minute init failure backoff, so a later normal
LSP call can retry the server/cwd instead of failing fast as recently failed.
Add a regression that aborts a wedged initialize and then retries the same
server/cwd with a short explicit timeout, asserting it does not hit the
negative-cache error.
Two termination boundaries in the browser tool leaked browser-owned OS resources into the long-lived coding-agent process.
1. Aborted 'open' published an orphan. #open wrapped acquisition in untilAborted, which rejects its outer wrapper on abort but lets the inner launch resolve in the background; acquireBrowser then unconditionally stored the resolved handle in the module-global browsers map. releaseAllTabs walks tabs, not browsers, so the refCount:0 handle stayed alive to process exit.
2. Session dispose had no browser teardown. Browser/tab state lives in module-global maps, and AgentSession.dispose() had no hook to walk them, so headless/spawned Chromium the session opened survived it.
acquireBrowser now short-circuits before launch on a pre-aborted signal and disposes the handle when the launch completes after abort. TabSession records the creating session's id (opts.ownerSessionId, threaded through BrowserTool.#open), preserved across reuse so a subagent re-driving an existing tab does not yank teardown responsibility. AgentSession.dispose() invokes releaseTabsForOwner bounded by withTimeout(3s), mirroring the async-job/MCP disposal pattern.
Regression tests exercise both boundaries via spied CmuxSocketClient (no real puppeteer/socket) and cover: pre-aborted open short-circuit, aborted-mid-launch cleanup, releaseTabsForOwner reaping only owned tabs, and reuse preserving original ownership.
Fixes#3963
sink.flush()'s type is number | Promise<number>, so calling .then/.catch
directly failed under tsgo. Wrap in Promise.resolve() and annotate the
rejection handler's err parameter to satisfy strict noImplicitAny.