Measured each animated Loader paint and recursively scheduled the next frame with a 10% duty-cycle target, capped at 200ms to retain responsiveness.
Added a regression covering a slow synchronized-output direct-write path and documented the fix.
Fixes#7290
Routed direct custom-message conversion through the collab steering transform so side requests and compaction see the same enveloped user turn as primary requests.
Extended the regression test to exercise convertToLlm without transformContext.
Converted user-attributed collab prompt frames to prioritized user messages only on the model-facing path, preserving guest details in persisted transcript frames.
Added regression coverage for the provider role, steering envelope, and retained guest attribution.
Fixes#7288
- Stored replay-sanitized Codex response items as the append baseline.
- Disabled append state for responses without replayable output and covered oversized call IDs.
Fixes#7279
- Rewrote fetchCodexDiscoveryModels to resolve every stored openai-codex
OAuth account via getOAuthAccesses and reuse openaiCodexModelManagerOptions'
tested union/fail-closed path, so a single narrow account can no longer
authoritatively wipe sibling-account models from the bundle.
- Restored openai-codex gpt-5.4, gpt-5.6-sol, and gpt-5.3-codex-spark bundle
entries (and gpt-5.5's contextPromotionTarget) dropped by the previous
single-account regen.
- Updated the live Codex image tool-result tests off the retired
gpt-5.2-codex id to gpt-5.5.
Fixes#6265
The 17.2.1 cowork request profile appended context-1m-2025-08-07 for any
model with a 1M catalog window, but that gate only runs on the OAuth path.
Subscription (Pro/Max) credentials have no long-context credit balance, so
Anthropic hard-429s ('Usage credits are required for long context requests')
on every beta-gated 1M model regardless of prompt size, breaking all
subagents (task/smol/scout resolve to claude-sonnet-4-6).
The beta is no longer advertised on OAuth requests; subscription accounts
transparently get the standard 200k window. Natively-1M models like
claude-sonnet-5 serve their full window without the beta anyway.
Fixes#7238
- Implement the ai& provider registry entry with API-key authentication and login support.
- Add model descriptors, static model seeding, and openai-compatible model discovery for the ai& provider.
- Update the model catalog with ai& provider models, pricing, and updated provider model names.
- Add unit tests for the ai& provider environment resolution, metadata, and dynamic model mapping.
- Added `ensureSharedBrowser` and shared browser acquisition to manage project-shared broker-owned Chromium instances.
- Implemented concurrent duplicate daemon start prevention and single-flight `pendingOpens` deduplication.
- Updated browser handle disposal to disconnect from shared daemons rather than closing them.
- Updated browser documentation and launch specifications to support shared and local headless runs.
- Restricted the reasoning-wrapper removal to leading <think> blocks so a
literal think tag after real content is preserved instead of deleted.
- Added coverage asserting mid-content think tags survive cleanOutput.
Fixes#7231
waitForManagedBashJob raced job completion against a bare
Bun.sleep(thresholdMs), which cannot be cancelled. When completion,
abort, or steering won the race, the losing Bun.sleep timer stayed
scheduled and ref'd, keeping Bun's event loop alive until the threshold
expired — delaying SDK/headless shutdown and accumulating timers under
fast command rates.
Replace the Bun.sleep with a Promise.withResolvers settled by a
cancellable setTimeout, and route every outcome (including the former
no-signal early return) through one try/finally that clears the timer
and removes the abort/steer listeners.
Add a child-process regression test that runs the real auto-background
path for a fast command against a 30s threshold and asserts the process
exits promptly instead of being held for the full threshold.
Fixes#7235
- Removed MiniMax-style think wrappers in cleanOutput so reasoning-model
responses no longer leak into consolidation summaries or corrupt fact
extraction (the wrapper previously survived parsing and every stored fact
became reasoning prose).
- Added remote consolidation and remote extraction regression coverage.
Fixes#7231
Routed Codex prompt cache identity through the shared retention-aware resolver while preserving transport session identity.
Covered explicit and environment-derived opt-outs, option precedence, direct body construction, and transport headers.
Fixes#7219
- Reused the agent's explicit or inherited cache identity for ephemeral side turns.
- Forwarded the same effective key through manual and automatic native compaction.
- Added regression coverage for all three secondary request paths.
Fixes#7218
The live AskDialog trusted question.question while its render helpers
(replaceTabs, renderQuestionTitle, questionTabLabel) assume a string. A
question reaching AskDialogComponent without a string question field threw
an uncaught TypeError that escaped the TUI render loop and killed the
session. The transcript renderer already normalizes the same malformed
data via normalizeRenderQuestions; the live path did not.
Normalize the questions array at dialog entry (new normalizeDialogQuestions),
coercing question/id/label to strings and options to a well-formed array,
matching the transcript path.
Fixes#7211
The Codex SSE `type:"error"` branch read only top-level `code`/`message`,
so backend rejections emitted under a nested `error` or `response.error`
object collapsed to `Codex error (): Unknown error`, hiding the cause
(e.g. a regional/model-snapshot rejection). `response.failed` similarly
dropped the error code.
Add a shared `extractCodexSseError` that reads top-level, nested `error`,
and `response.error` envelopes, and wire both error paths through it so
the backend code and message survive in `SearchProviderError`. The
existing `web_search_call` requirement is untouched.
Fixes#7200
- Extracted the chromiumCanLaunch probe from browser-tab-evaluate into a
shared test/tools/chromium-probe.ts helper.
- The two attached-navigation tests from PR #7006 launch real headless
Chrome; CI runners without Chrome system libraries (libnspr4 & co.)
cannot exec the downloaded binary, failing the native/unit bucket.
- Gave both tests 30s timeouts to survive CI cold starts.
- Update hashline block resolution formatting to correctly incorporate anchor lines within operation labels.
- Fix and update test assertions and mock contexts across coding agent tests.
- Replaced the reactive weekly-only auto-redeem predicate with a pool-wide
planner: an expiry-salvage sweep piggybacks on the 5-minute usage
heartbeat and spends any account's reset that would otherwise expire
within codexResets.salvageHorizonHours, and the blocked-turn path scans
all stored accounts with eligibility built from the exact exhausted
5h/weekly windows (openai/codex#28525), unblocking at the latest reset
among them.
- Made the live 429's parsed unblock timestamp authoritative for the
active account (pre-block snapshots survive cache invalidation via
in-flight adoption and last-good fallback), synthesizing the candidate
when no usable report exists, and overlaying live credit counts from
the dedicated credits route since a stale /wham/usage zero is never
corrected upstream.
- Treated nothing_to_reset, credit_list_failed, and thrown consumes as
non-terminal: the episode key is released and deferred 30 minutes
instead of burying a banked credit; redeemResetCredit now spends the
soonest-expiring credit.
- Added planner unit fixtures plus integration regressions driving the
real triggers end to end, with an injectable per-session coordinator
seam and a sweep settlement handle.
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
- resolveOwnerScopedSessionKey's getOwners in the Python and JS executors
only read live sessions, so a subagent reset issued while the shared
kernel was still starting resolved to the base key, awaited the
parent's startup, and shut its brand-new kernel down.
- Python and JS starting sessions are now owner-bearing records like
Ruby/Julia's: owners attach synchronously before startup resolves,
getOwners and per-owner disposal consult them, and the final
sessions.set is identity-guarded so a disposed starting record cannot
resurrect its kernel.
- Regression: deferred PythonKernel.start proves a concurrent subagent
reset forks immediately and never reaps the parent's starting kernel
(fails with the previous getOwners).
- Subagents inherit the parent's eval session id, so a child's
reset: true destroyed the co-owned kernel and every sibling's
interpreter state mid-session.
- resolveOwnerScopedSessionKey now routes a reset from a non-exclusive
owner onto a deterministic per-owner fork key: the requester gets a
fresh private kernel, co-owners keep the shared one, and the fork
stays sticky for that owner until its teardown reaps it.
- Applied across Python, JavaScript, Ruby, and Julia executors; JS
contexts gained an owner registry plus disposeVmContextsByOwner,
wired into EvalRunner.disposeKernels and SDK session teardown.
- Covered by pure key-resolution contracts and an end-to-end JS test:
co-owner reset forks, shared state survives, fork is sticky, and
per-owner dispose reaps only the fork.
- The earlier EBADF 'hardening' was chasing a Bun 1.4-canary quirk that
only triggers when bun test runs workspace-package files from the repo
root; the canonical runner (scripts/ci-test-ts.ts) always uses the
package cwd, where the original PR-authored pipe-based assertions pass.
- Restores the stronger stdout contracts from PRs #7135 and #7141.