- Added `onTurnError` hook to `AdvisorRuntime` to handle failed turns before retries.
- Integrated credential blocking in `AgentSession` to prevent retrying usage-limited accounts when advisor turns fail.
- Included account key in `codex-auto-reset` debug logs to improve skip reason visibility.
Drained pending IRC asides before parking irc wait so replies that arrive between wait calls are returned instead of being treated only as queued interrupts.
Added regression coverage for the already-aborted queued-IRC signal path and documented the fix in the coding-agent changelog.
Fixes#4657
Reset per-turn maintenance counters before IRC wake prompts so yielded subagents do not carry stale yield termination into later wake turns.
Add regression coverage for empty-stop retry after an IRC wake following a yielded run.
Fixes#4658
Rebuilt active advisor runtimes when modelRoles.advisor changes so live sessions stop using stale advisor models.
Added regression coverage for advisor role updates reaching the live advisor without a manual /advisor restart.
Fixes#4612
`#resolveAdvisorRuntimeDescriptors` hardcoded `ThinkingLevel.Medium` when
no thinking suffix was configured. For reasoning models with no
controllable effort surface (`devin-agent`: `reasoning: true`,
`thinking: undefined` — Cascade selects effort by routing to sibling
model ids, not a wire param), that default tripped
`requireSupportedEffort` on the first advisor prompt with an empty
`Supported efforts:` list, disabling the advisor session-wide.
Route the default through `resolveThinkingLevelForModel(model, level)`
which preserves explicit `off`, clamps a concrete effort into the
model's supported range, and returns `undefined` for reasoning models
without controllable efforts — falling back to `Inherit` so no effort is
sent while reasoning stays enabled. Matches the `auto`-path fix
(`clampAutoThinkingEffort`) and the Autonomous Memory clamp
(`clampThinkingLevelForModel`).
Fixes#4579
`plan.defaultOnStartup` records a `mode_change` before the composer restores its draft; without this the draft-cleanup arm check treats the file as durable and the metadata-only JSONL leak reappears for default-plan sessions.
Added a regression case that drives a model_change + mode_change + draft-clear cycle and asserts the session file is dropped on close().
Fixes#4571
Limit empty-session close cleanup to files whose draft sidecar lifecycle
materialized an otherwise startup-metadata-only session. Direct
ensureOnDisk() callers now remain discoverable even when they have no
user/assistant messages yet, and handoff custom_message entries survive
close before the next user turn.
Added regression coverage for resumed draft cleanup, ACP-style explicit
ensureOnDisk() records, and handoff custom messages.
Fixes#4571
`SessionManager.saveDraft(text)` calls `ensureOnDisk()` so the draft
sidecar has a parent JSONL. A follow-up `saveDraft("")` only unlinks
the sidecar — the session file was left behind, and `#shouldHaveSessionFile()`
could not prune it once the load path latched `#fileIsCurrent` and
`#forceFileCreation` to true. Each draft-then-clear-then-exit cycle
leaked a ~500–750 B zombie into `~/.omp/agent/sessions/<cwd>/`
containing only the title slot, session header, and a handful of
`model_change`/`mode_change`/`thinking_level_change` entries.
`close()` now calls `#dropIfEmptyAndNoDraft()` after draining the
writer: when the file exists, holds no user/assistant messages, and
no draft sidecar is present, it removes the session file and its
artifacts directory via `deleteSessionWithArtifacts`. Real conversations,
sessions with a saved draft still on disk (needed for `--resume`), and
never-materialized sessions are untouched.
Fixes#4571
- Introduced an automated retry recovery system to track, manage, and persist recovered error states within agent sessions.
- Enabled compact transcript rendering for recovered auto-retry errors by removing heuristic commit machinery.
- Improved raw read tracking and provenance in the ReadTool to support refined file snapshot recording and hashline editing.
- Excluded recovered assistant messages from default model context and updated event controllers to handle retry recovery life cycles.
- Made resolveShapeForText choose silver16-bw for CJK-heavy auto transcripts while preserving explicit variants and unsafe glyph protection.
- Added silver16-bw to the snapcompact shape settings submenu and renamed unsupported-glyph warnings.
- Covered auto shape selection, explicit variant precedence, unsafe glyph scans, and settings option parity.
Fixes#4486
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.
Fixes#4469
A request the provider rejects (e.g. 413 oversized payload) yields a
synthesized assistant turn with empty content and stopReason 'error'.
That turn is written to session.jsonl, so on reload it replays as an
empty assistant turn and re-sends the same rejected context. Keep the
rejection UI-only (pinned error) and out of persisted history so a
reloaded session resumes from the last good turn.
- Classified OpenAI-compatible custom relays serving OpenAI model ids into the OpenAI service-tier family.
- Passed the model into OpenAI service-tier wire gating so custom relays emit service_tier when eligible.
- Reported /fast on as unavailable when the active model has no service-tier family.
Fixes#4386
Sizing `maxTokens` off the static `model.reasoning` catalog flag cannot
distinguish a thinking model catalogued `reasoning: false` (e.g. Qwen3
served locally via llama.cpp, whose bundled jinja chat template defaults
`enable_thinking: true`) from a model that never emits thinking. The
tight non-reasoning budget was consumed by the thinking preamble before
the useful output could be emitted, so every affected call silently
failed with `stopReason: "length"`.
Drop the `model.reasoning` conditional across every affected online call
site and always reserve the reasoning-safe budget. `maxTokens` is a hard
cap, not a target — non-thinking completions still return in the tiny
happy-path budget.
Sites fixed:
- utils/title-generator.ts (30 -> 1024)
- utils/commit-message-generator (60 -> 1024)
- tts/speech-enhancer (512 -> 1536)
- auto-thinking/classifier online path (8 -> 1024); classifyLocal
keeps its separate LOCAL_ANSWER_MAX_TOKENS
- session/unexpected-stop-classifier online path (16 -> 1024);
classifyLocal keeps ANSWER_MAX_TOKENS
Fixes#4355
Plan-approval's 'Approve and compact context' used to pass the rendered
plan-mode-compact-instructions prompt as the first positional argument
to handleCompactCommand -> session.compact(), which landed on the
session_before_compact extension hook as customInstructions. Extensions
treating that field as user focus (e.g. to bias a query-focused summary)
would then see plan-mode boilerplate instead of operator intent and
produce query-biased compactions.
Add CompactOptions.internalGuidance: a private summarizer-only channel.
session.compact() reads it into the fallback-model summarizer while the
session_before_compact hook payload still only carries the public
customInstructions arg (undefined for the plan-compact path). The
snapcompact-disable predicate and the /compact rejectsFocus guard cover
both fields so a directed summary is never silently downgraded.
Extend the interactive-mode handleCompactCommand facade + command
controller with a fourth internalGuidance parameter, and switch the
plan-approval callsite in interactive-mode.ts to route the plan prompt
through it.
Fixes#4359