Add provider-native V2 remote compaction metadata to discovered OpenAI Codex models so context-full compaction uses the streaming compaction_trigger path instead of the legacy compact endpoint.
Expose the V2 remote compaction schema fields in models.yml and cover the Codex discovery metadata contract.
Fixes#4146
AGENTS.md forbids inline `await import()` — move `fs`, `path`, and `prompt`
to top-level namespace/named imports. The `prompt-templates` module import
already registers the Handlebars helper as a side-effect, so the render
calls still resolve the new `renderYieldSchema` helper.
- Updated exact substring match assertions to use word-boundary regular expressions.
- Prevents false-positive test failures when target strings overlap with other generated text.
- Improved the leaked-thinking stream projector to clone and sync native tool-call blocks directly.
- Eliminated the need for placeholder IDs and complex rekeying logic in the event controller and argument reveal module.
- Simplified native tool-call validation in owned-stream processing by requiring only a non-empty name.
- Added comprehensive unit tests to ensure tool-call IDs and partial JSON parameters remain intact during healing.
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Preserved the referenced model's OpenAI-compatible reasoning-effort support when openai-models-list discovery enriches a thin /v1/models payload. The discovered model still keeps conservative proxy-local store and developer-role defaults, but known reasoning models like gpt-5 no longer force supportsReasoningEffort false and trigger the omitReasoningEffort request path.
Added a regression assertion that a thin proxied gpt-5 keeps supportsReasoningEffort true and omitReasoningEffort false after reference enrichment.
Fixes#3983
Thin OpenAI-compatible proxies that omit context_length / max_model_len on
/v1/models made every discovered model fall back to
DISCOVERY_DEFAULT_CONTEXT_WINDOW (128K/33K), even when the id matched a
bundled model with a much larger intrinsic window. discoverProxyModels
and discoverLiteLLMModels already resolve ids against the bundled
reference index; discoverOpenAIModelsList (which also backs lm-studio
discovery) now does the same.
Behavior:
- Build the reference index once outside the loop and resolve each item
via resolveModelReference().
- contextWindow precedence keeps provider-reported values authoritative:
item.max_model_len ?? item.context_length ?? nativeMetadata?.contextWindow
?? reference?.contextWindow ?? DISCOVERY_DEFAULT_CONTEXT_WINDOW.
- maxTokens uses reference?.maxTokens when available, otherwise the
api-specific discovery default, capped at contextWindow so a bundled
ref for a larger sibling can never over-request output tokens.
- name / reasoning / thinking / input inherit from the reference; native
lm-studio metadata still wins for input modality.
- Provider-specific baseUrl, headers, and local-unknown cost stay local.
- OpenAI-compat flags stay conservative (supportsStore / supportsDeveloperRole
/ supportsReasoningEffort all false) to match the proxy sibling.
Also updated two pre-existing regression tests that used
deepseek-v4-pro / deepseek-r1 / DeepSeek-V4-Flash as stand-in "fictional"
ids to exercise the default-fallback branch. Those model names have since
been added to the bundled catalog, so the tests were renamed to
vllm-lab-fork-* ids that unambiguously miss the reference index while
preserving each test's original default-fallback intent.
Fixes#3983
Added a cross-turn tool-call loop guard that hashes canonical tool names and arguments, ignores intent metadata, and injects a hidden redirect when identical calls reach the configured threshold.
Fixes#3971
- Introduced an `#editVariantCache` to memoize resolved edit modes for model variants.
- Replaced the generic `shallowStringRecord` helper with specialized, type-safe parsing methods for model variants and roles.
- Invalidated the cached edit variants during settings rebuilds and verified correct cache refreshment across project directories.
- Moved Streamable HTTP request and notify timeout cleanup after response body consumption.\n- Added regression coverage for stalled request JSON bodies and stalled notify error bodies.\n\nFixes #3974
The subagent system prompt rendered `{{jtdToTypeScript outputSchema}}` as a
bare TypeScript interface with the text "Your result MUST match this
TypeScript interface". The yield tool actually nests the user schema under
`result.data`, so the LLM pattern-matched on the visually dominant code
block and put the payload directly in `result.data`, tripping schema
validation repeatedly. In the worst reported case a subagent used all 3
retry attempts, had validation dropped, and lost its audit output entirely.
Add a `renderYieldSchema` Handlebars helper that renders the schema inside
`result: { data: … }` and swap the system-prompt block to use it, so the
model sees the exact envelope the yield tool expects. Multi-line object
schemas, scalars, unions, and array-of-object schemas all round-trip
cleanly with the new helper.
Fixes#3972
Wrapped cmux page, browser, and tab globals with per-run abort checks so stale continuations cannot reuse the long-lived CmuxTab after timeout.
Fixes#3964
Distinguish aborts that race an active sink.flush() from aborts that happen
before a queued write starts. Only the former leaves the sink flush pending
and requires killing/evicting the LSP client; pre-write aborts should reject
that caller without disrupting unrelated in-flight operations.
Add a regression with one notification blocked in flush and a second queued
notification whose signal aborts before its write starts, asserting the shared
client is not killed and only the first message is written.