- Implement the ai& provider registry entry with API-key authentication and login support.
- Add model descriptors, static model seeding, and openai-compatible model discovery for the ai& provider.
- Update the model catalog with ai& provider models, pricing, and updated provider model names.
- Add unit tests for the ai& provider environment resolution, metadata, and dynamic model mapping.
waitForManagedBashJob raced job completion against a bare
Bun.sleep(thresholdMs), which cannot be cancelled. When completion,
abort, or steering won the race, the losing Bun.sleep timer stayed
scheduled and ref'd, keeping Bun's event loop alive until the threshold
expired — delaying SDK/headless shutdown and accumulating timers under
fast command rates.
Replace the Bun.sleep with a Promise.withResolvers settled by a
cancellable setTimeout, and route every outcome (including the former
no-signal early return) through one try/finally that clears the timer
and removes the abort/steer listeners.
Add a child-process regression test that runs the real auto-background
path for a fast command against a 30s threshold and asserts the process
exits promptly instead of being held for the full threshold.
Fixes#7235
- Reused the agent's explicit or inherited cache identity for ephemeral side turns.
- Forwarded the same effective key through manual and automatic native compaction.
- Added regression coverage for all three secondary request paths.
Fixes#7218
The live AskDialog trusted question.question while its render helpers
(replaceTabs, renderQuestionTitle, questionTabLabel) assume a string. A
question reaching AskDialogComponent without a string question field threw
an uncaught TypeError that escaped the TUI render loop and killed the
session. The transcript renderer already normalizes the same malformed
data via normalizeRenderQuestions; the live path did not.
Normalize the questions array at dialog entry (new normalizeDialogQuestions),
coercing question/id/label to strings and options to a well-formed array,
matching the transcript path.
Fixes#7211
The Codex SSE `type:"error"` branch read only top-level `code`/`message`,
so backend rejections emitted under a nested `error` or `response.error`
object collapsed to `Codex error (): Unknown error`, hiding the cause
(e.g. a regional/model-snapshot rejection). `response.failed` similarly
dropped the error code.
Add a shared `extractCodexSseError` that reads top-level, nested `error`,
and `response.error` envelopes, and wire both error paths through it so
the backend code and message survive in `SearchProviderError`. The
existing `web_search_call` requirement is untouched.
Fixes#7200
- Extracted the chromiumCanLaunch probe from browser-tab-evaluate into a
shared test/tools/chromium-probe.ts helper.
- The two attached-navigation tests from PR #7006 launch real headless
Chrome; CI runners without Chrome system libraries (libnspr4 & co.)
cannot exec the downloaded binary, failing the native/unit bucket.
- Gave both tests 30s timeouts to survive CI cold starts.
- Update hashline block resolution formatting to correctly incorporate anchor lines within operation labels.
- Fix and update test assertions and mock contexts across coding agent tests.
- Replaced the reactive weekly-only auto-redeem predicate with a pool-wide
planner: an expiry-salvage sweep piggybacks on the 5-minute usage
heartbeat and spends any account's reset that would otherwise expire
within codexResets.salvageHorizonHours, and the blocked-turn path scans
all stored accounts with eligibility built from the exact exhausted
5h/weekly windows (openai/codex#28525), unblocking at the latest reset
among them.
- Made the live 429's parsed unblock timestamp authoritative for the
active account (pre-block snapshots survive cache invalidation via
in-flight adoption and last-good fallback), synthesizing the candidate
when no usable report exists, and overlaying live credit counts from
the dedicated credits route since a stale /wham/usage zero is never
corrected upstream.
- Treated nothing_to_reset, credit_list_failed, and thrown consumes as
non-terminal: the episode key is released and deferred 30 minutes
instead of burying a banked credit; redeemResetCredit now spends the
soonest-expiring credit.
- Added planner unit fixtures plus integration regressions driving the
real triggers end to end, with an injectable per-session coordinator
seam and a sweep settlement handle.
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
- Subagents inherit the parent's eval session id, so a child's
reset: true destroyed the co-owned kernel and every sibling's
interpreter state mid-session.
- resolveOwnerScopedSessionKey now routes a reset from a non-exclusive
owner onto a deterministic per-owner fork key: the requester gets a
fresh private kernel, co-owners keep the shared one, and the fork
stays sticky for that owner until its teardown reaps it.
- Applied across Python, JavaScript, Ruby, and Julia executors; JS
contexts gained an owner registry plus disposeVmContextsByOwner,
wired into EvalRunner.disposeKernels and SDK session teardown.
- Covered by pure key-resolution contracts and an end-to-end JS test:
co-owner reset forks, shared state survives, fork is sticky, and
per-owner dispose reaps only the fork.
- The earlier EBADF 'hardening' was chasing a Bun 1.4-canary quirk that
only triggers when bun test runs workspace-package files from the repo
root; the canonical runner (scripts/ci-test-ts.ts) always uses the
package cwd, where the original PR-authored pipe-based assertions pass.
- Restores the stronger stdout contracts from PRs #7135 and #7141.
- The IRC wake monitor from PR #7108 calls the observer on every kept-alive
subagent finalize; fakes across the executor suites predate the method
and threw during finalization (69 failures).
- Bun 1.4 canary's posix_spawn intermittently rejects pipe-backed child
stdio (EBADF) inside test workers, failing telemetry-export and
shell-snapshot probes before their assertions ran.
- Probes now spawn with ignored stdio and capture output via exit status
or temp-file redirection; all behavioral assertions retained.
- #7163's context-token anchoring re-triggered #7153's dead-end warning on
the pre-prompt pass; a one-shot marker now spans agent loops and re-arms
only when a persisted cut point appears.
- Updated scrollback tests to the Alt+L display-reset binding main adopted.