Commit Graph
9801 Commits
Author SHA1 Message Date
can1357 ca4f0d6c56 Merge PR #6494: fix(session): cancel in-flight session_stop handlers (@roboomp) 2026-07-26 15:41:28 +02:00
can1357 e758f10a96 Merge PR #6660: fix(prewalk): apply same-model effort downgrades instead of skipping (@roboomp) 2026-07-26 15:41:28 +02:00
can1357 c5a819b162 Merge PR #6669: fix(session): retry reasonless tool-call aborts (@roboomp) 2026-07-26 15:41:28 +02:00
can1357 9b651f2869 Merge PR #6616: fix(cursor): sync native todo list from server-resolved tool calls (@quantmind-br) 2026-07-26 15:41:27 +02:00
can1357 7bbe140914 Merge PR #6626: fix(advisor): propagate advisor provider session id via metadata (@roboomp) 2026-07-26 15:38:52 +02:00
can1357 2cc6b28d98 Merge PR #6537: fix(advisor): bound Codex SSE retries (@jeffscottward) 2026-07-26 15:38:52 +02:00
roboomp 2fb9f888ad fix(ai): preserved custom anthropic web-search history
Retained complete native web-search call/result pairs through leaked-thinking projection at their original source anchors.

Kept opaque server-tool blocks atomic during persistence, stripped them on reparent, and covered custom continuations plus compaction retention.

Fixes #6703
2026-07-26 13:32:08 +00:00
Diogo Soares Rodrigues 9088fe821b fix(cursor): await error-drain transforms and sanitize mirrored todo labels
Two boundary defects on the mirrored-todo path.

1. The Agent error drain snapshotted #cursorToolResultBuffer without
   awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
   async cursorOnToolResult still running when the provider errored
   patched an entry the catch path had already detached, so the
   pre-transform payload was persisted. A provider error is exactly when
   a transform is most likely to be in flight.

2. The todo renderer interpolated mirrored provider text straight into
   terminal output. A Cursor snapshot carries model-authored task
   content, phase names and summary text verbatim, so a label holding
   ANSI/C0 sequences rewrote the terminal on every render and replay.
   sanitizeText alone is not enough - it preserves tabs, which punch
   holes in bordered output - so every display path now funnels through
   one forDisplay() helper: task labels, blocker notes, phase headers,
   the zero-task fallback, and the streaming renderCall preview. Raw
   values are untouched; content and phase name are the identity keys
   the local list is looked up by and what gets persisted.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues c7c375f5e1 fix(cursor): settle todo cards whose completion outruns the streamed block
When Cursor packs toolCallStarted and toolCallCompleted into one HTTP/2
chunk, the bridge tool_execution_end (synchronous callback, fired
mid-parse) reaches the interactive controller before the streamed
toolcall_start (queued on AssistantMessageEventStream, delivered a
microtask later). The controller found no pendingTools entry, dropped
the completion, and the card created afterwards animated forever.

Two halves, each necessary:

- Hold an early todo completion in #orphanedToolCompletions and replay
  it when the streamed block creates its component.
- Guard card creation from cumulative message_update frames with the
  turn-scoped #toolTimelineComponents map. Without this, the update
  after the replay re-lists the same toolCall block, finds pendingTools
  empty again, and spawns a second, permanently pending card. This is
  also why emitting a synthetic tool_execution_start from the bridge
  (previous attempt, reverted) could not work.

Both maps are cleared together at the existing transcript-anchor reset
sites. The normal ordering (start first) is covered by a control test.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues cc210094ef fix(cursor): refuse todo dependency graphs and sanitize failure text
Two review findings on the native todo sync.

- TodoItem.dependencies is a graph the local model cannot store: rows
  are keyed by content, carry no id, and hold no edges. An imported
  dependent row files as plain pending and nextActionableTask then
  offers work the server considers blocked. Refuse snapshots with an
  edge pointing at an unfinished row; edges whose blockers already
  finished constrain nothing and still mirror.

- The todo failure warning interpolated the provider error verbatim.
  Collapse and truncate it at the render boundary.

Also documents two known, unfixed defects: an async cursorOnToolResult
transformer resolving after the buffer drain, and the todo card
lifecycle race. Emitting a synthetic tool_execution_start for the
latter was measured and rejected -- the completion deletes the entry it
creates, so the late streamed block adds a second card.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues 0ab0aa6331 fix(cursor): refuse empty read_todos and say when a snapshot is not mirrored
An empty `read_todos` with `total_count=0` (proto3 unset or genuinely
empty) was accepted and mirrored as an authoritative wipe. Refuse empty
reads; clearing the list stays on `update_todos`.

Benign refusal text is now "Todo snapshot not mirrored" instead of
"No todo changes", which falsely described a server-accepted update that
only the local mirror declined.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues 3f2b45662e docs(cursor): list the unrepresentable snapshot among benign refusals
`CursorTodoSyncHandler`, `buildTodoToolResult`, `extractTodoError`, and
the host `todoSync` each enumerate why a snapshot may be `null`. All
four listed only filtered and truncated reads, so a duplicate-content
refusal read as undocumented -- easy to mistake for an error, or to
"fix" by mirroring it again.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues 0639246d27 fix(cursor): settle refused and failed native todo calls
Only a successful snapshot settled a native todo block. A `read_todos`
narrowed by a filter and a server `UpdateTodosError` both went
unanswered: no `tool_execution_end`, so the card animated forever, and
no `toolResult`, so `buildSessionContext` stripped the block on rebuild.

Every completed native todo call now settles. The refusal path carries
no `details.phases` -- `event-controller` feeds that straight into
`setTodos`, so echoing the current list back would let a call that
changed nothing overwrite live panel state. A server error is carried
through as a failed result instead of collapsing into the benign no-op.

Each regression is covered by a test verified to fail without its fix.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues ea023c380a fix(cursor): persist the phase-bearing todo result and close a buffer race
The previous commit paired every server-resolved todo block with a
result, but built that result in the provider from the flat snapshot.
`todoToolRenderer.renderResult` reconstructs the list exclusively from
`details.phases`, so the block survived the dangling-strip only to replay
as `Todo 0 tasks`.

Only the host computes that grouping -- the provider sees a flat list --
so `todoSync` now returns the result it already assembled and the
provider persists it verbatim. A refused snapshot never reaches the host,
so the provider's summary-only fallback still covers that path, and
exactly one result is emitted either way.

Separately, `Agent`'s Cursor buffering wrapper pushed its entry only
after awaiting the optional `cursorOnToolResult` transformer. The
provider dispatches decoded messages with `void handleServerMessage(...)`,
so a `message_end` from the same chunk could drain the buffer while a
transformer was still pending, dropping the result. The entry is now
reserved synchronously and patched in place when the transformer
resolves, keeping buffer order and still applying the customization.
Production is unaffected -- `sdk.ts` sets no transformer -- but the
option is supported and its contract returns a Promise.

Tests: a delayed-transformer case that loses the result without the
buffering change, and a replay case driving the persisted result through
`buildSessionContext` and asserting `details.phases` rebuilds a non-empty
list -- the id-pair assertion alone did not catch the empty render.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues 29e64ce608 fix(cursor): resolve and persist server-owned todo blocks
Review follow-up on two defects in the native todo bridge.

`tool_execution_end` was emitted under a freshly generated UUID, but the
interactive transcript files the visible block under the streamed
`callId` and only clears it when the ids match. The card therefore stayed
pending and animating for the rest of the session. The settled call id is
now passed to `todoSync`, making the parameter required so no caller can
silently reintroduce a mismatch.

Nothing produced a `toolResult` for these blocks either: `todoSync` only
appended a custom entry and emitted a transient event. Since
`buildSessionContext` strips any `toolCall` with no matching result, the
interaction vanished from every rebuilt transcript -- reload, branch
switch, or Ctrl+L -- leaving a "tool call elided" placeholder. A paired
result now travels the same `onToolResult` channel the other
server-resolved Cursor calls already use, including when the snapshot is
refused: the call happened, it just changed no local state.
2026-07-26 09:29:13 -03:00
Diogo Soares Rodrigues 7214951ead fix(cursor): sync native todo list from server-resolved tool calls
Cursor resolves its native `update_todos`/`read_todos` tools server-side,
so the todo list never followed the model's intent locally.

Two defects, both silent:

- `agent.v1.ToolCall` is a protobuf oneof. A decoded message exposes the
  selected variant as `tool: { case, value }` and has no flattened
  `updateTodosToolCall` property, so the bridge recognized no native todo
  call at all on the wire path.
- The synthesized `todo` block was emitted as locally runnable carrying a
  `{todos}` payload the local tool's schema rejects, turning every update
  into a validation error and driving a spurious continuation turn.

Todo calls are now read through the oneof, both native blocks are stamped
resolved, and local state is mirrored only from the server's confirmed
success snapshot. Partial `read_todos` responses -- narrowed by
`status_filter`/`id_filter`, or short of the server's own `total_count` --
are subsets, not the list, and are refused rather than deleting the tasks
they omit. `TODO_STATUS_CANCELLED` maps to `abandoned` instead of
reverting the task to `pending`.

The exec bridge mirrors each snapshot into session state, refreshes the
interactive panel via a synthetic `tool_execution_end`, and persists to
the session branch so the list survives reloads, rewinds, compaction, and
session switches. Existing phase grouping is preserved.

Regression tests drive the bridge with wire-encoded protobuf, which is
the only shape production ever sees; all six fail without this change.
2026-07-26 09:29:13 -03:00
Paolo Mazzitti 81bc0d4368 fix(coding-agent): stop the live advisor runtime when /settings disables it
Disabling the Advisor from /settings persisted the setting but left the
live Advisor runtime running until the session restarted.
SelectorController.handleSettingChange had no case for "advisor.enabled",
unlike other session-managed toggles (autoCompact, steeringMode, ...), so
the change never reached session.setAdvisorEnabled — the same call /advisor
off already uses to stop the runtime immediately.
2026-07-26 12:25:19 +00:00
roboomp e650607cad fix(coding-agent): segment bash approval commands with shell-aware tokenizer
The regex splitter only recognized `&&`, `||`, `;`, `|` and newlines, so a
single `&` (background operator) — also a command terminator — slipped a
dangerous command past a deny rule (`sleep 1 & rm -rf /tmp/x`), which under
approvalMode: yolo executed with no prompt.

Extract the shell-aware tokenizer from gh-cache-invalidation into a shared
tools/shell-tokenize.ts and reuse it for deny/prompt segmentation. It honors
every command boundary (`&&`, `||`, `;`, `|`, single `&`, subshells,
newlines) plus quoting and escapes, so both callers share one implementation.

Fixes #6695
2026-07-26 11:36:19 +00:00
roboomp 06674b2571 fix(cli): defer out-of-scope default role so extension models resolve
enabledModels is resolved at startup before extensions call
registerProvider(), so a modelRoles.default naming an extension-provided
model dropped out of the resolved scope and buildSessionOptions pinned
options.model to the first scoped model. That pre-fill marked the model
"explicit" in createAgentSession, suppressing the post-extension
default-role re-resolution and silently running the session on a
different in-scope provider's model.

A configured default that cannot be found in the startup scope is now
left unset so createAgentSession re-resolves it against the fully
registered, still enabledModels-scoped catalog once extensions load. The
first scoped model is only seeded when no default role is configured.

Fixes #6694
2026-07-26 11:33:03 +00:00
roboomp 85110c366c fix(coding-agent): match bash deny/prompt patterns per command segment
bashApprovalPatternToRegExp anchors globs with ^...$ against the whole
normalized command, so a bash.patterns deny rule only fired when the
dangerous command was first in the line. A compound command such as
`cd /tmp && rm -rf /tmp/x` bypassed the rule and, under approvalMode:
yolo, executed with no prompt -- deny is the guard that outranks yolo.

deny/prompt rules now match the whole command or any single segment
(split on &&, ||, ;, |, newlines). allow rules still require the entire
command to match and never apply to compound lines, so a narrow allow
cannot vouch for a smuggled unsafe segment.

Fixes #6695
2026-07-26 11:28:48 +00:00
roboomp 765a54305e fix(stt): preserved nearest sherpa runtime
- Loaded the nearest wrapper first and limited ancestor fallback to missing-addon failures.

- Covered nested version-conflict installs alongside the broken-hoist layout.
2026-07-26 11:18:31 +00:00
roboomp 838fad7e5c fix(stt): resolved workspace sherpa addon loading
- Selected the sherpa wrapper colocated with its platform-native package in source workspaces.

- Added a regression test for Bun workspace hoisting.

Fixes #6690
2026-07-26 11:08:06 +00:00
Wolfgang Schoenberger 3e1679f584 fix(cli): keep the Hindsight URL readable and redact only set credentials
The credential pass marked `hindsight.apiUrl` secret instead of
`hindsight.apiToken`: both share `condition: "hindsightActive"` and the flag
landed on the wrong one, so the settings panel masked an ordinary endpoint. The
token keeps its masking through the `credential` marker.

`config list` also redacted on classification alone, so a fresh configuration
reported every unset credential as though one were stored. Redaction now
depends on a value being present.

The suite asserted classification and panel metadata only, so both output
branches could be deleted while every test passed, which is how the wrong flag
shipped. It now drives `runConfigCommand` and reads the real human and JSON
output.
2026-07-26 02:58:38 -07:00
Wolfgang Schoenberger 914afc0d6c fix(cli): redact credential settings in config list
omp config list printed every configured value, including auth.broker.token,
searxng.token, searxng.basicPassword and dev.autoqaPush.token, in both the
human and --json output. Nobody asked for those specific credentials; the
command dumps everything.

Credentials are marked with a top-level credential flag rather than ui.secret,
because four of them have no settings-panel entry and so have nowhere to put a
UI-level flag. isCredential is the single accessor both the CLI and the panel
consult, so the two spellings cannot produce different behaviour on different
surfaces.

Human output shows dots. JSON omits value and marks the entry redacted instead
of substituting a placeholder, which a consumer could not distinguish from a
real value and might write back.

config get <path> is deliberately unchanged: that is an explicit request for a
single value, and masking it would break a retrieval API with no way to read
your own token back.
2026-07-26 02:58:38 -07:00
Wolfgang Schoenberger 0c5fa60c77 fix(coding-agent): cap captures for resizing transports
Use the resolved supportsImageDetailOriginal capability to constrain native computer frames whenever a Responses transport clamps screenshot detail to auto. Preserve the established Claude-family fallback for transports that do not expose this capability.

Cover Copilot GPT-5 Responses through real catalog model resolution and the controller's observable capture options.
2026-07-26 02:53:01 -07:00
Wolfgang Schoenberger fc68288477 fix(coding-agent): scope computer capture limits 2026-07-26 02:53:01 -07:00
Wolfgang Schoenberger 25dc887fb2 fix(coding-agent): preserve computer coordinates across providers 2026-07-26 02:53:00 -07:00
Éverton Toffanetto 1f4cdcbbbb fix(coding-agent): cap the local classifier and keep the Low floor
Two defects in the ceiling work, both found in review.

The local backend shared the online ceiling, so with `autoThinkingMaxEffort:
max` and a sparse ladder the clamp could snap a `hard` bucket up to `max` —
a tier the 3-bucket on-device classifier can never select. The local branch
now pins `xhigh`.

Applying the ceiling before the Low floor also broke the floor's contract:
on `["minimal", "max"]` under an `xhigh` ceiling the intersection hid `max`,
the code concluded the model "maxes out below Low", and it fell through to
`minimal`. The floor is now resolved against the model's own ladder first
and the ceiling filters that pool, so an excluded top tier yields no level
instead of a sub-Low one.

Docs and changelog now scope the guarantee to what `auto` resolves: a
`thinking.requiresEffort` model whose ladder holds nothing under the ceiling
still receives its lowest supported effort from the transport, because it
accepts nothing else. The test that claimed to prove billing is renamed to
say what it checks.

Prompt assertions now cover the `max` criteria and the tie-break exception,
not just the label, since the label alone is inert. Drops the duplicated
pool-level assertions in favour of the contract-level sparse-ladder case.
2026-07-26 05:11:53 -03:00
Larry Gordon d683cf90e7 fix(extensions): resolve tool approval against the revised tool input
Follow-up to the earlier approval re-check, which missed the
prompt-to-prompt case: if the original and revised inputs both resolve
to `prompt`, a handler could swap in different prompt-gated args that ran
under approval granted for the original.

Emit the `tool_call` event before the approval gate instead of after, so
the gate resolves policy and shows the interactive prompt against the
input that actually executes. The user always approves what runs, and
deny/allow-to-prompt/prompt-to-prompt transitions are all covered by one
gate rather than special-cased. A `deny` on the original input still
short-circuits before the runner is touched, so an already-denied tool
never emits `tool_call`.

Tests: the approval prompt reflects the revised input, `tool_call` fires
before `tool_approval_requested`, plus the existing deny/computer/
multi-handler cases.
2026-07-26 00:30:38 -07:00
Larry Gordon 8916e7de9f fix(extensions): re-check approval on revised tool input and skip computer calls in hook wrapper
Addresses PR review feedback.

- Re-gate approval (P1): the extension wrapper's approval/safety gate resolved
  against the original `params`, but execution ran with the overridden input, so
  a handler could rewrite approved args into ones a deny/critical policy would
  have blocked. After an override, re-resolve the policy on the revised input and
  block a revision that newly resolves to `deny` (or, outside yolo, newly requires
  a prompt) instead of running it unapproved.
- Computer skip in hook wrapper (P2): the hook wrapper applied the override to
  `computer` tool calls, contradicting the documented contract; it now skips them
  like the extension wrapper does.

Docs: the shared-events contract note is now accurate for both wrappers and
documents the approval re-check. Tests: added the re-gate (blocks-deny,
allows-benign), multi-handler last-wins, and hook computer-skip cases.
2026-07-26 00:09:50 -07:00
Éverton Toffanetto 3933bbd0f4 fix(coding-agent): keep the auto ceiling out of the model clamp
Capping the classifier result before `clampAutoThinkingEffort` was not
enough. The clamp seeds `chosen` with `pool[0]`, so a sparse ladder whose
tiers all sit above the request snaps upward instead of down: on
`thinking.efforts: ["max"]` an `xhigh` request returned `max`, letting the
default setting bill the top tier with no opt-in. The same upward snap made
`resolveProvisionalAutoLevel` hand back `max`, breaking the invariant its
doc comment had just claimed.

`clampAutoThinkingEffort` now takes the ceiling and intersects it with the
model's supported tiers, returning `undefined` when nothing is eligible so
auto leaves the current level alone instead of billing an excluded tier.
The classifier passes its configured ceiling and the provisional level
passes XHigh.
2026-07-26 04:04:07 -03:00
Éverton Toffanetto e3e4fd3a31 refactor(coding-agent): import THINKING_EFFORTS from the catalog module
AGENTS.md requires catalog values to come from `@oh-my-pi/pi-catalog/<module>`
rather than the pi-ai barrel, which re-exports only the types its own
signatures need.
2026-07-26 03:45:44 -03:00
Éverton Toffanetto fbd3ebffc6 feat(coding-agent): add opt-in max ceiling for auto thinking
`max` became a first-class effort tier in d435385a, but the `auto`
classifier prompt still offers only `low|medium|high|xhigh`. On a model
that exposes the tier, `auto` can therefore never reach it — only the
`ultrathink` keyword can, because it bypasses the classifier entirely.

`providers.autoThinkingMaxEffort` (`xhigh` | `max`, default `xhigh`) lifts
that ceiling. Opting in adds `max` to the classifier vocabulary, gated on
the target model actually supporting the tier, and scopes the tie-break
exception to that prompt variant so the default renders byte-for-byte as
before. A classification above the configured ceiling is clamped before
the model clamp, so a hallucinated `max` cannot cross a ceiling the user
did not opt into. The on-device 3-bucket classifier stays capped at
`xhigh`, and the provisional/fallback level still never provisions `max`.

Also corrects the two `Auto-detect per prompt (low-xhigh)` labels and the
stale `xhigh auto ceiling` comment, which the new setting makes wrong.
2026-07-26 03:16:56 -03:00
Larry Gordon ed457a9d4f feat(extensions): let tool_call handlers revise tool input
A `tool_call` handler (extension or hook) could previously only block a
tool. It can now also return `input` to replace the arguments the tool
executes with, so a handler can normalize or rewrite a built-in's input
without reimplementing the tool.

The returned object is the raw execution input passed to the tool's
`execute` (the handler owns its correctness), not the normalized
`event.input` view, which may carry derived gate-only fields (e.g.
hashline `edit` `path`/`paths`) that are not real parameters. It is
ignored when `block` is set, and not applied to `computer` tool calls
whose event input is a synthetic actions view rather than the real
params. When multiple handlers set `input`, the last one wins.

Honored in both the extension and hook tool wrappers; documented in
docs/extensions.md, docs/hooks.md, and docs/skills/authoring-hooks.md.
2026-07-25 23:16:04 -07:00
Jeff Scott Ward fcdd33d2fe fix(mcp): map mounted tools to xd routes 2026-07-26 01:02:11 -04:00
roboomp 0988efbce6 fix(session): recover cursor reasonless tool aborts
Scoped the Cursor server-execution marker gate to the stream-stall path so an unmarked client-side Cursor tool call (todo/MCP) followed by a reasonless abort recovers via its synthetic executed:false result instead of settling the turn.

Fixes #6668
2026-07-26 04:33:34 +00:00
roboomp 464eba46b3 fix(session): retried reasonless tool-call aborts
Continued from synthetic unexecuted tool results when a reasonless request abort arrives after a complete streamed tool call. Preserved deliberate user, lifecycle, and streaming-edit guard abort behavior.

Fixes #6668
2026-07-26 04:26:53 +00:00
roboomp 80df85a84b fix(coding-agent): key debug thread aggregation by owning session
Review feedback: DAP thread IDs are scoped per client session, so deduping
the tree-wide thread list by (id, name) collapsed distinct threads that
different children happen to share — e.g. multiple identical worker scripts
each exposing { id: 1, name: "worker.js" }. Key the dedupe set by the owning
session id + thread id instead, so every live thread is preserved and only an
exact repeat within a single session's response is dropped.

Fixes #6663
2026-07-26 03:50:50 +00:00
roboomp 04f61e2125 fix(coding-agent): aggregate debug threads without assuming a threadless root
Review feedback: classifying every root-with-a-live-child as a threadless
launcher dropped the root from thread aggregation, so a custom TCP adapter
whose root process owns threads and also issues startDebugging would have its
threads omitted. Drop the launcher heuristic: threads now fans out across all
live sessions in the tree and merges results. A genuinely threadless launcher
simply returns no threads (or an error we skip), so no topology guess is made.

Fixes #6663
2026-07-26 03:48:08 +00:00
roboomp 74c279c60b fix(coding-agent): route js-debug commands to the stopped script child
js-debug launches are session trees: a threadless root launcher spawns
child sessions (main script, [worker N]) via reverse startDebugging. Every
stateful command routed through a single #activeSessionId set on the last
registration or stop, so a worker attaching after the script stopped stole
focus. threads listed only the worker, post-launch breakpoints read back as
unbound, and step/continue/evaluate could not target the script thread.

Focus now follows stops, not registrations: a new session claims the active
pointer only when no live, stopped session already holds it. threads
aggregates every live thread across the tree, skipping the threadless
launcher while real children are alive.

Fixes #6663
2026-07-26 03:40:51 +00:00
roboomp 7ac1e69f9f fix(advisor): kept quarantine failure latched until recovery
Internal advisor context resets cleared failureNotified immediately after the quarantine warning, causing the status to return to running and repeated warnings every two quarantines.

Keep the notification latch through internal context re-primes. Clear it only on a successful advisor turn, explicit reset, or seed, and cover both deduplication and recovery in the quarantine regression test.
2026-07-26 03:17:33 +00:00
roboomp 377e9ff5d9 fix(advisor): notified the user when a quarantined turn drops advice
The advisor drain loop's quarantine branch reset context and re-primed silently with no bound, so an advisor that called an ungranted tool (e.g. bash) had its whole turn discarded before dispatch and its advice never reached the primary. Every other non-recovering failure branch calls notifyFailure -> emitNotice; quarantine was the one path with no main-UI signal, leaving supervision failures visible only in advisor diagnostics.

Count consecutive quarantines and, past MAX_QUARANTINE_RETRIES, surface the failure via notifyFailureOnce and drop the batch instead of looping silently. Reset the counter on any successful turn and on reset().

Fixes #6661
2026-07-26 03:12:59 +00:00
roboomp 7a6265f801 fix(prewalk): distinguish auto from inherit before declaring a no-op
prewalkWouldBeNoop collapsed both auto and a fixed :inherit selector to an
undefined clamped effort, so a same-model prewalk targeting :inherit while the
session ran auto was dropped as a no-op. Applying :inherit clears per-turn
classification, so that hand-off is a real change. Compare auto/fixed mode before
comparing clamped efforts so an auto<->fixed transition always switches.

Fixes #6659
2026-07-26 02:34:17 +00:00
roboomp 846a71b7a2 fix(prewalk): compare model-clamped efforts in no-op check
prewalkWouldBeNoop compared raw selectors, so a target the model cannot honor
(e.g. :xhigh on a model capped at high while running high) read as a change and
triggered an ephemeral model reset plus the plan/checklist nudges even though
setThinkingLevel clamps it straight back to the active effort. Compare the
target- and current-level efforts AFTER model clamping via
resolveThinkingLevelForModel so a clamp-equal target is recognized as a no-op.

Fixes #6659
2026-07-26 02:27:44 +00:00
roboomp 583ff590f2 fix(prewalk): apply same-model effort downgrades instead of skipping
The prewalk arm/switch guard compared model identity only (modelsAreEqual /
provider+id), discarding the resolved thinkingLevel. A legal same-model target
at a cheaper effort (e.g. prewalk: "@task" resolving to the active model at a
lower level) was dropped as a no-op, so the session ran the expensive effort for
the whole run while still paying the plan/continue nudges — silently on the
session path, logger.debug only on the subagent path.

Compare (provider, id, effective thinking level) via a shared prewalkWouldBeNoop
helper. Effort-only deltas on the same model now switch; a genuine no-op emits a
user-visible notice on the session path and never arms on the subagent path.

Fixes #6659
2026-07-26 02:20:03 +00:00
Will 7e00fa5a76 docs: clarify Exa authentication options 2026-07-25 19:14:48 -04:00
roboomp a676e3f29c fix(secrets): stopped buffering single-dollar streamed text
- Buffered only a lone trailing dollar or a $$-introduced body during streaming.
- Let $HOME/$100-style single-dollar text stream through immediately.
- Added regression coverage for the passthrough.
2026-07-25 23:08:54 +00:00
Will 19d7d14a94 feat(ai): add Exa API key login 2026-07-25 19:08:02 -04:00
roboomp 431a5509cc fix(secrets): removed hash placeholder fallback
- Removed hash-delimited placeholder parsing and stored-session aliases.
- Simplified replay and display restoration to the double-dollar format.
- Replaced legacy compatibility tests with an inert-token regression.
2026-07-25 23:03:19 +00:00
roboomp 7d8ae857c9 fix(coding-agent): rejected missing legacy auth keys
Returned the historical missing-key error when request authentication resolves without a credential.

Covered the non-throwing empty-auth result through the real registry path.
2026-07-25 22:23:16 +00:00