- Updated compaction, branch summarization, and session dump formatting to pass preferred model tool syntax into conversation serialization.
- Enhanced shared serializers to render assistant tool calls and tool results through grammar envelopes when syntax is available, with the prior compact format as fallback.
- Aligned prompt, preview, and test fixtures to the new transcript tags: `[Think]`, `[Tool Call]`, and `[Tool Result]`.
- Owned stream wrapping handled `toolcall_start`, `toolcall_delta`, and `toolcall_end` events from provider-native calls.
- The in-band projector tracked native calls by source index, created projected tool blocks, emitted lifecycle events, and committed final call fields so toolUse output is preserved.
- Gemini example rendering omitted the `default_api.` prefix when rendering examples, and a regression test covered native tool-call passthrough without in-band code text.
- Added `WorkerInbox` and `installWorkerInbox(port)` to queue worker messages before bind.
- Added `consumeWorkerInbox()` to replay buffered messages and clear one active inbox.
- Added buffered inbox consumption in JS and tab worker transports before direct message handlers.
- Normalized worker selector arguments to the `__omp_worker_*` naming across workers and tests.
- Added an `interruptible` field to AgentTool and documented when it is honored.
- Updated immediate-mode tool execution to poll steering during in-flight interruptible calls and abort them when steering is queued.
- Marked the coding `job` tool as interruptible and added tests covering mid-wait aborts versus boundary-only steering drain.
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
- Added a helper that renders hub output and extracts ordered agent IDs.
- Replaced positional string assertions with explicit row-order expectations.
- Awaited asynchronous theme initialization in test setup.
- Added Gemini and Gemma syntax routing by model family and owned syntax env values.
- Added Gemini and Gemma in-band parsers for tool_code and token-based tool_call streams.
- Added rendering support for Gemini fenced tool_code/tool_outputs and Gemma tool tokens.
- Fixed parsing edge cases for comments, string escapes, nested args, and truncated blocks.
- Added a streamingBehavior field to submitted user inputs and threaded it through interactive mode types and controller start-up.
- Updated interactive submission dispatch to default to followUp queueing while preserving explicit steer intent when provided.
- Changed tests to verify followUp and steer queueing behavior, preventing AgentBusyError from race-window busy sessions.
- Captured each row's initial position in AgentHubOverlayComponent on first refresh and reused it on later refreshes.
- Changed subsequent sorting to prioritize status then prior row position, so activity updates no longer reordered visible rows.
- Added a regression test that verifies row ordering stays stable when activity changes and new agents append at the end.
- #handleAgentEvent now delegates non-`agent_end` events to a new `#processAgentEvent` routine.
- For `agent_end`, it tracked a resolver promise via `#trackPostPromptTask` before awaiting event handling and resolved it in `finally`.
- This kept `#waitForPostPromptRecovery()` from returning before deferred compaction or handoff work was registered.
- Updated default model identifiers across many catalog providers to newer model versions.
- Renamed a couple OpenAI compatibility provider descriptors, including Together and Zhipu coding-plan identifiers.
- Added multiple new OpenAI-compatible specialized provider descriptors for additional model provider families.
- Added `azure` provider registration in the AI registry with API key env mapping.
- Added Azure provider descriptors with default model `gpt-4o` and catalog discovery metadata.
- Enabled Azure-specific OpenAI compatibility for developer roles and strict responses pairing.
- Added Azure models namespace using OpenAI-family IDs with `models.dev` filtering and responses transport.
Track the non-message token estimate used for provider-anchored assistant usage, then add only positive current system/tool growth during pre-prompt threshold checks. Restore released collab changelog entries to the immutable 15.13.1 section.
Fixes#2628
Include the current system prompt and active tool schemas when provider-anchored pre-prompt checks estimate threshold pressure, and cover the case with a regression test.
Fixes#2628
Use provider-anchored context usage for pre-prompt context-full threshold checks when available, then add only pending prompt tokens. This keeps OpenAI Responses encrypted reasoning payloads from overcounting local prompt pressure while the visible context usage remains below threshold.
Fixes#2628
- Added `features.unexpectedStopDetection` and `unexpectedStopModel` settings for opt-in behavior.
- Added assistant-stop handling to classify stop reasons and resume generation with retry prompts.
- Added unexpected-stop classifier logic with candidate checks, model fallback, and YES/NO parsing.
- Added retry tracking that caps auto-continues at three attempts and logs a warning when exceeded.
- Updated the bun-install action to retry `bun install --frozen-lockfile` when the first attempt fails.
- Added a fallback path that creates a temporary job-local cache directory and reruns install with `--cache-dir`.
- Emitted a warning to note the shared-store failure before the retry path is used.
- Updated worker core so it emitted `ready` only after receiving an `init` message instead of on construction.
- Changed worker startup to send `init` before waiting for readiness, enabling startup errors to fail fast and trigger fallback.
- Expanded JS worker tests with startup-error simulation and verified execution falls back to the inline worker when spawning fails.
- Initialized JS workers through a new helper that attaches message and error listeners before initialization handshake and enforces the existing ready timeout.
- Handled startup failures by rejecting on init or error events, cleaning up listeners and replacing dead workers with a retry path.
- Retried session creation with an inline worker after non-inline init failure so requests no longer stall on worker startup timeouts.
The keyed maps only see items whose `output_item.added` carries `item.id`
or `output_index`. A fully keyless add never reached either map, and the
unkeyed fallback was scanning those maps instead of the actual current
item — so `function_call_arguments.delta` / `output_item.done` for a
keyless tool call landed on null and the stored block kept `{}`. Mixed
streams also picked an older `output_index` entry before a later
id-only current item for the same reason.
`CodexStreamRuntime.currentEntry` now always points at the most recently
added `output_item.added` (whether or not it has keys). `openItemForEvent`
returns `currentEntry` when both `item_id` and `output_index` are absent,
and `closeCodexOpenItem` clears `currentEntry` (alongside the legacy
`currentItem` / `currentBlock` mirrors) when its item closes. The keyed
maps are unchanged so deliberate drop-on-mismatch for keyed events still
holds.
Regression tests cover the fully keyless function-call stream and the
mixed id-only vs `output_index`-only ordering. Existing reasoning/message
flow keeps singleton semantics through the same fallback.
Fixes#2619
- Removed the `--only-failures` flag from workspace test commands so CI runs complete test passes.
- Raised native test mode parallelism from 1 to 4 for native and integration package execution.
- Updated `init` handling to build a single-phase list from `items` when `list` is absent, using `phase` if provided or defaulting to `Tasks`.
- Kept the existing init validation behavior by emitting `Missing list for init operation` only when neither `list` nor `items` was supplied.
- Added tests for flattened init forms, including implicit phase defaulting, explicit phase use, and the missing-input error case.
- Added an `example` option to `GrammarRenderOptions` and threaded it through grammar tool-call renderers.
- Updated Harmony tool-call rendering to output raw argument JSON when example mode is enabled.
- Removed `string` attributes from generated XML `parameter` elements in tool invocation output.
Codex Responses function/custom tool call items can omit `id` while still
carrying `output_index` on `output_item.added` and `output_item.done`.
The first pass only keyed open items by `item.id`, so idless done events
could not find the stored block and the authoritative final arguments were
not copied into `output.content`.
Track open items by both `item.id` and `output_index`. Event lookup now
uses `item_id` first, then `output_index`, and only falls back to the
singleton-current path when neither key is present. Closing an item removes
both keys, so stale keyed deltas remain dropped instead of leaking into a
sibling.
Added a regression covering idless function and custom tool calls finalized
only by `output_item.done`, including out-of-order completion and per-call
`contentIndex`.
Fixes#2619
- Added `jsonSchemaToTypeScript` and `renderToolInventory` to generate tool blocks with TypeScript signatures.
- Added `examples` and `TSchema` fields to dump-tool metadata and passed them through prompt rendering.
- Changed Harmony invocation rendering to omit `<|constrain|>json` markers in tool call payloads.
- Added compact native tool list-mode inventory rendering with full `# Tool:` output elsewhere.
The Codex Responses stream runtime tracked a singleton
`currentItem`/`currentBlock`. With more than one tool call open
concurrently every `response.function_call_arguments.delta` was
appended to whichever item was added most recently, and the next
`response.output_item.done` for the earlier call overwrote the
sibling's stored arguments. On the agent loop's `task` tool this
surfaced as `tasks: Invalid input: expected array, received undefined`.
Open items are now tracked in a `Map<string, CodexOpenItem>` keyed by
`item.id`, each entry carrying its own `block` and `contentIndex`.
`response.function_call_arguments.{delta,done}`,
`response.custom_tool_call_input.{delta,done}`, and
`response.output_item.done` route through `openItemForEvent` and
operate on the matching entry's block; the legacy singleton-current
fallback only kicks in when an event omits `item_id`. A delta whose
keyed item already closed is dropped instead of leaking into a
sibling, and `toolcall_delta` / `toolcall_end` stream events emit the
right `contentIndex` for each call. Recovery sites converge on a new
`resetCodexStreamAccumulators` helper that clears the open-items map
in lockstep with `currentItem`/`currentBlock`/`nativeOutputItems`.
Regression test (`packages/ai/test/openai-codex-stream.test.ts`)
interleaves two function-call argument streams plus a stale
post-close delta and asserts per-call argument integrity and
per-call stream `contentIndex`.
Fixes#2619
- Added an abortOnFabricatedToolResult option to Agent and AgentLoopConfig to choose whether in-band fabricated tool results are aborted or drained.
- Propagated the option through agent loop wiring into wrapInbandToolStream so fabrication is aborted only when enabled.
- Exposed the setting in coding-agent as tools.abortOnFabricatedResult and wired it through session creation with a true default.
- Detected the runner environment in CI by checking SCCACHE_BUCKET and exporting an on_infra output.
- Updated the workflow to use the local ensure-sccache action on self-hosted runners and mozilla-actions/sccache-action on GitHub-hosted runners.
- Added model-to-syntax mapping in catalog with preferred tool-call syntax API.
- Added `ToolExample` typing and `ToolCallSyntax` exports across tool/grammar interfaces.
- Added syntax-aware tool example rendering through provider-specific grammar invocations.
- Added `exampleSyntax` context flow and example metadata so rendered prompts include examples.
- Added a `#stripLeadingWhitespace` state flag to track whitespace after a dropped control token.
- Trimmed leading whitespace from the scanner buffer at the start of outside-token consumption when the flag was set.
- This suppressed template-control trailing spaces from being emitted in visible text.
- Added a `truncateFrom` option to `TreeListOptions` supporting `start` and `end` modes, defaulting to `end`.
- Updated `renderTreeList` to compute candidate items and summary placement based on truncation direction.
- Set the todo list renderer to use start-side truncation so collapsed todo output shows the tail entries first.
- Replaced large-paste wrap options with a single `<attachment>` wrapper action.
- Added explicit local-file attachment and inline paste choices to the selector.
- Updated the changelog to document the new large-paste action set.
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
- Renamed line and block patch op verbs to XCHG, DEL, and INS in parsing and formatting.
- Updated grammar and tokenizer to support XCHG.BLK, DEL.BLK, and INS.PRE/POST/HEAD/TAIL forms.
- Updated diagnostics, docs, prompts, tests, and changelog to use XCHG/DEL/INS-based operators.
- Expanded session-stats parsing to normalize legacy op aliases to compact IDs.
- WelcomeComponent now lazily selected and cached a tip per instance, preserving it across re-renders.
- With the unicode preset, it showed a special nerdfont tip 10% of the time and otherwise used the regular tip rotation.
- Added tests that mocked theme preset and Math.random to verify standard and special tip selection behavior.
The pre-prompt context check ran compaction directly, so snapcompact (or
any strategy) fired before auto-promote ever got a chance — defeating
Auto-Promote Context. It now tries promotion to a larger-context model
first (mirroring the post-turn threshold path) and only compacts when no
target is available.
Auto and manual compaction now project a snapcompact result's
post-compaction size (kept history + frames at the image budget + summary
+ non-message overhead); when it still exceeds the model's usable window,
they downgrade to a context-full LLM summary instead of leaving the
session overflowing.
Native linux-x64/arm64 builds moved onto the Ubuntu 24.04 (glibc 2.39)
omp-kata runner. The x64 addon was a plain host build that linked the
runner's glibc and failed to dlopen with `version 'GLIBC_2.39' not found`
on older distros; the arm64 cross-build floated up to GLIBC_2.30. Build
the shipped linux-gnu addons through cargo-zigbuild against a pinned 2.17
floor so they load on any glibc >= 2.17.
- build-native.ts: key the tree-sitter-just `-UNDEBUG` CFLAGS off the
bare triple (cargo-zigbuild strips the `.2.17` glibc suffix before
invoking cargo) and symlink the suffixed target dir napi 3.7.0 expects
to the bare dir cargo-zigbuild writes, so postBuild copyArtifact finds
the cdylib.
- build-native action: add a `glibc` input plus a resolve step deriving
the zigbuild cross_target (suffixed) and the rustup bare_target
(stripped); gate zig/cargo-zigbuild install on cross_target so the
host-arch x64 build still runs native Rust tests.
- ci.yml: GLIBC_FLOOR=2.17 fed to the linux-x64 and linux-arm64 native
jobs.
Re-tags 15.13.1, whose release failed at the linux-x64 binary smoke
before any publish step ran.