Commit Graph

2945 Commits

Author SHA1 Message Date
can1357 a4d8860a6c feat: added google reasoning controls mcp stream resumption and tar support
- Added Google provider thinking configuration parameters and force-reasoning-off controls.
- Implemented MCP SSE stream resumption using Last-Event-ID and `SSEResumeError`.
- Added support for TAR old-GNU sparse extension blocks, path length checks, and archive entry overrides.
- Restricted external thinking support to specific models and added semver fallback parsing.
2026-08-12 02:32:45 +02:00
can1357 20732b5355 Merge PR #8273: fix(tui): hide remaining activity in hidden-tool mode (@dannyboy-ai) 2026-08-12 01:53:53 +02:00
can1357 ce13985746 Merge PR #8276: fix(coding-agent): stop print mode hanging on plan.defaultOnStartup (@roboomp) 2026-08-12 01:53:44 +02:00
can1357 fe837038e9 Merge PR #8262: fix(coding-agent): copy local artifacts across handoff session boundary (@roboomp) 2026-08-12 01:53:44 +02:00
Daniel Anderson-Little e4dc5c9637 refactor(coding-agent): clarify warning activity policy 2026-08-11 17:52:05 -04:00
Daniel Anderson-Little d50e47d2af fix(coding-agent): align activity visibility contract 2026-08-11 17:26:33 -04:00
Daniel Anderson-Little cbd7dc2e70 fix(coding-agent): centralize hidden activity state 2026-08-11 17:07:20 -04:00
Daniel Anderson-Little 3a4ea6a376 fix(coding-agent): hide internal tool activity blocks 2026-08-11 16:59:53 -04:00
roboomp 3ebe60d8b3 fix(coding-agent): stopped print mode hanging on plan.defaultOnStartup
Headless `omp -p` armed an interactive plan-review flow whenever
plan.defaultOnStartup was set. Its only headless exit was a watcher that
fired on a successful `xd://propose` execute-dispatch, so any turn where
the model did not emit exactly that dispatch stranded until --max-time,
printing nothing. --plan-yolo could not rescue it because print mode's
arming clashed with the prewalk coordinator's plan-yolo handoff.

Print mode no longer honors the startup default: headless has no surface
to review, approve, or exit a plan. It writes a one-line stderr note and
runs the prompt normally. --plan-yolo remains the deterministic headless
plan flow and no longer clashes.

Fixes #8272
2026-08-11 19:39:40 +00:00
can1357 10fd42289c feat: introduced external thinking support and private scratchpad think tool
- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
2026-08-11 20:39:57 +02:00
roboomp dbc199d2db fix(coding-agent): copy local artifacts across handoff session boundary
/handoff mints a fresh session via newSession(), producing a new
artifactsDir and an empty local/ root. The handoff document routinely
references plans and scratch files under '/data/workspaces/can1357__oh-my-pi__8261/.omp-session/2026-08-11T16-39-09-489Z_019ff1b1-31b1-7000-81f5-c540f4ebf43d/local/,' so every reference
became a dangling pointer in the new session. The plan approve-and-execute
path already copies artifacts across the boundary; handoff did not.

Extracted the plan-approve copy helper into a shared copyLocalArtifacts()
in local-protocol.ts and invoke it across the handoff session switch
(best-effort, since the switch is already committed).

Fixes #8261
2026-08-11 16:45:15 +00:00
can1357 f3a3073a3d Merge PR #7959: fix(tui): show edit previews before approval (@roboomp) 2026-08-11 15:18:16 +02:00
can1357 b807d67e1d Merge PR #8238: fix(tui): render platform-aware modifier labels on macOS (@roboomp) 2026-08-11 15:15:45 +02:00
can1357 64baa7c1bd chore(format): applied biome formatting and removed dead code from merged prs 2026-08-11 15:14:15 +02:00
can1357 706371bb91 fix(extension-api): preserved overlay option factories 2026-08-11 15:08:39 +02:00
Vanko ed820703a7 fix(extension-api): restore overlayOptions/onHandle passthrough in ui.custom
showHookCustom hardcoded the overlay geometry and never read
overlayOptions/onHandle, which regressed the v0.45.6 API (PR
badlogic/pi-mono#667). Forward overlayOptions to showOverlay (keeping the
full-cover defaults as fallback), invoke onHandle with the returned
OverlayHandle, widen the options type via a shared ExtensionCustomOptions,
and re-export OverlayHandle/OverlayOptions from the extension API.
2026-08-11 15:08:39 +02:00
can1357 600bac9615 Merge PR #8213: Fix/8030 osc133 command start (@Mustaqeem66) 2026-08-11 15:07:43 +02:00
can1357 24c348862f Merge PR #8197: fix(status-line): recognized kimi-code subscription windows in the usage segment (@tehfiend) 2026-08-11 15:06:16 +02:00
can1357 4b8dbcee07 Merge PR #8172: fix: make terminal Mermaid state diagrams readable (@shawnkoh) 2026-08-11 15:06:16 +02:00
can1357 da680b6e67 Merge PR #8150: perf(rpc): reuse serialized output frames (@MikeeI) 2026-08-11 15:06:15 +02:00
can1357 d57656f283 Merge PR #8074: fix(coding-agent): display subagent registration time locally (@anatoli-tsinovoy) 2026-08-11 15:06:13 +02:00
can1357 7896a58406 Merge PR #8050: fix(tui): keep pasted images when a mode command submits the draft (@fatihaziz) 2026-08-11 15:06:13 +02:00
can1357 dedddb86f5 Merge PR #7998: fix(ai): map Cursor plan rails to dashboard percents and show mo N% (@dnth) 2026-08-11 15:06:12 +02:00
can1357 66a55fb916 Merge PR #7983: fix(vibe): restore the pre-vibe toolset when switching between vibe sessions (@iskWang) 2026-08-11 15:06:12 +02:00
can1357 6592b799b3 Merge PR #7980: fix(session): attribute a run to the model that produced its output (@enieuwy) 2026-08-11 15:06:12 +02:00
can1357 f6d4f1e38b Merge PR #7965: fix(coding-agent): surface todo progress in the collapsed panel (@z80dev) 2026-08-11 15:06:12 +02:00
can1357 d13121fbf8 Merge PR #7856: fix(coding-agent): answer panel commands immediately during a turn (@mvid) 2026-08-11 15:06:11 +02:00
Fatih Al-Aziz 8a22872ece fix(tui): preserve mode submissions across async input 2026-08-11 16:46:45 +07:00
roboomp 097845bff0 fix(tui): render platform-aware modifier labels on macOS
Key hints resolved modifier tokens through a static, platform-agnostic label map with no `super` entry, so on macOS the shipped `super+v` paste default rendered 'Super+V' (no such key on a Mac) and `alt` always rendered 'Alt' instead of 'Option'. The static /hotkeys navigation rows were hardcoded with macOS 'Option'/'Cmd' names on every platform, so Linux/Windows users saw 'Cmd+Left'.

Modifier labels are now platform-aware: on darwin `alt` renders 'Option' and `super` renders 'Cmd'; every other platform keeps 'Alt'/'Super'. The platform is resolved through a single seam (setKeyHintPlatform/keyHintPlatform) mirroring the TUI's setKittyProtocolActive, keeping hint output deterministic in tests without mutating process.platform. The static /hotkeys rows use the same convention and drop the macOS-only Cmd line-start/end fragments (which map to no binding) off darwin.

Fixes #8235
2026-08-11 09:28:52 +00:00
Fatih Al-Aziz de9ee6d412 fix(tui): preserve mode attachments across async paths 2026-08-11 16:25:20 +07:00
Fatih Al-Aziz d194f2d76c fix(tui): honor transformed mode attachments 2026-08-11 15:52:58 +07:00
Muhammad Mustaqeem d5e2bb9299 fix(tui): close the OSC 133 prompt zone so terminals clear input state
OSC 133;B latches a sticky `.input` cursor semantic in Ghostty (and
Ghostty-derived terminals such as cmux) that only a command-start marker
clears. UserMessageComponent emitted A/B and never C, so for the rest of
the session `cursorIsAtPrompt()` stayed true and every painted cell was
tagged `.input`. With `cursor-click-to-move = true` (Ghostty's default)
every left-click released inside the pane injects a burst of arrow keys
into omp's pty and the editor caret jumps to column 0.

Emit C immediately followed by D;0 at the end of the bubble. The zone is
opened and closed within the same render, so the sticky input state is
cleared while the grouping behaviour the original comment protected is
preserved: no later assistant/tool output can be grouped under this
prompt because the command zone never stays open.

Refs #8030, #6115
2026-08-11 10:29:38 +05:00
Bejay Cole e7f487381b fix(status-line): recognized kimi-code subscription windows in the usage segment
Kimi's usage provider built window ids from the raw duration/timeUnit ("300time_unit_minute") and left the aggregate quota on "default", so the status-line usage segment, which matches canonical "5h"/"7d" windows, never rendered for kimi-code sessions.

Derive window ids from the reported span instead (300 minutes -> "5h", 7 days -> "7d"), mirroring the minimax-code convention, and tag the aggregate quota as the weekly window. The payload carries only resetTime for the aggregate, but the observed reset horizon matches the plan's weekly cycle. The segment aggregator also falls back to the reported span (duration within a minute of 5h or 7d) when the id isn't canonical, so cache rows written by older builds still light up the 5h meter in mixed-version fleets.
2026-08-10 17:01:41 -06:00
Shawn Koh a42868f550 fix: make terminal state diagrams readable 2026-08-11 01:00:35 +08:00
Bonobo 9de7862a29 perf(rpc): reuse serialized output frames
Why:
RpcFrameEncoder serializes normal frames once for protocol routing and again
for output, while v2 snapshots serialize message payloads separately.

Changes:
- Feed the existing JSON into frame-size enforcement.
- Build v2 message snapshots from the serialized frame.

Evidence:
- End-to-end paired medians improved 10.72% for v1 and 15.72% for v2
  across 25 pairs, with identical output hashes.

Refs #8118
2026-08-10 04:08:41 +02:00
Anatoli Tsinovoy 5d8214ea77 fix(coding-agent): keep lineage timezone visible 2026-08-09 16:49:13 +03:00
Anatoli Tsinovoy d725ab10c1 fix(coding-agent): display subagent registration time locally 2026-08-09 16:35:07 +03:00
Josh Mini 66648fb518 Merge remote-tracking branch 'upstream/main' into fix/vibe-restore-toolset-on-resume 2026-08-09 12:51:42 +08:00
Josh Mini bb5bfb2c04 fix(vibe): only override the vibe snapshot when teardown lost the live toolset
The persisted pre-vibe snapshot was applied on every reconciliation that
re-entered vibe mode, including cold resumes and switches in from a non-vibe
session. Those paths build their toolset from the current CLI flags and
settings, so replacing it with a historical snapshot silently drops tools the
session was started with: resuming a session that entered vibe under
--tools read with --tools read,bash restored only read on exit.

Gate the override on the one case the snapshot exists for: the vibe -> vibe
switch, where #clearTransientModeState kept the already-reduced live set.
2026-08-09 11:23:07 +08:00
Fatih Al-Aziz 276f1dd3b2 fix(tui): keep pasted images when a mode command submits the draft
`/goal <objective>`, `/plan <prompt>` and `/vibe <prompt>` promote the
composer draft into the first turn, but built their submission from the
draft *text* only:

    this.onInputCallback(this.startPendingSubmission({ text: objective }));

The editor-submit path in `InputController` passes
`editor.pendingImages`/`pendingImageLinks` alongside the text; these four
call sites did not. A draft holding pasted screenshots therefore reached
the model with its positional `[Image #N, WxH]` markers intact and every
image payload missing, so the agent saw markers pointing at nothing and
`read "Image #1"` resolved against an empty list.

The payload was not only dropped, it also outlived the draft: the mode
commands cleared the composer with `editor.setText("")`, which leaves
`pendingImages` attached. The orphans then rode along with whatever the
user typed next, one index off, which is how a later message can attach a
screenshot the user never re-pasted.

Measured on 259 image-bearing user messages across 10 local session logs
(v17.2.x): 36 of 37 messages submitted as a goal objective lost every
image, against 190 of 198 preserved on the ordinary submit path.

Fix:
- `#takeDraftImages()` detaches the composer's pending images and links,
  and all four mode-command submissions spread it into
  `startPendingSubmission` (`cancelPendingSubmission` already restores
  them when a submission is cancelled).
- `/goal`, `/guided-goal`, `/plan` and `/vibe` clear the draft with
  `editor.clearDraft()` instead of `editor.setText("")`, so images can
  never outlive the text they were pasted into (the streaming branch of
  `/goal` never submits, so its draft must die whole).

Tests: two regression cases in `goal-mode-integration.test.ts` assert the
objective submission carries the image and empties the composer; both
fail on the previous behaviour with `images: undefined`. The `/plan`,
`/goal` and `/guided-goal` slash stubs now model `clearDraft`.
2026-08-09 08:38:09 +07:00
dnth aa92c7e97c fix(ai): address Cursor usage review feedback
Gate status-line monthly rendering to Cursor, prefer personal dashboard
rails over legacy /auth/usage request fractions, fall back from unusable
overall buckets to plan, ignore disabled plan percent fields, and keep
on-demand when the included bucket is empty. Also add changelog PR/author
attribution for the external contribution.
2026-08-08 17:36:19 +08:00
dnth 69f790fa6e fix(ai): map Cursor plan rails to dashboard percents
Cursor Pro/Pro+ usage-summary still uses individualUsage.plan, but the
web dashboard percent is autoPercentUsed / apiPercentUsed — not
plan.used/limit cents. Prefer those rails for Cursor Models / Other
Models, keep overall+cents fallbacks, floor monthly status-line %, and
repaint after async usage fetch so mo N% does not stay blank.
2026-08-08 17:26:40 +08:00
dnth cc3835342e fix(ai): parse Cursor plan/onDemand usage and show monthly status-line
Cursor's /api/usage-summary now returns individualUsage.plan (and optional
onDemand) for Pro/Pro+/Ultra instead of the overall bucket that #7613
parsed. Fall back to plan when overall is absent, keep overall preferred
when both exist, and render monthly Cursor quotas in the status-line usage
segment.

Verified locally against a live Pro+ account and with focused unit tests.
2026-08-08 16:46:01 +08:00
roboomp 92e574cb02 fix(session): surface empty handoff generation as failure not cancel
The #7904 fix stopped masking provider errors as "Handoff cancelled", but
an empty or whitespace-only generation still fell through: whitespace-only
text passed the `!handoffText` guard and produced a bogus handoff, while
empty text returned undefined which the interactive /handoff caller mapped
to "Handoff cancelled" with no detail and no log entry.

Treat empty/whitespace-only output as a real failure: a user-initiated
handoff throws "Handoff generation produced no content" (surfaced as
"Handoff failed: ...") and logs it; auto-handoff keeps returning undefined
so maintenance falls back to context-full compaction. Also log genuine
handoff failures in the command controller so they persist for debugging.

Fixes #7993
2026-08-08 08:21:16 +00:00
Josh Mini 1e142aace9 fix(vibe): restore the pre-vibe toolset when switching between vibe sessions
Entering /vibe snapshotted the live toolset into an in-memory field only.
When a session already in vibe mode switches into another session that is also
in vibe mode, #clearTransientModeState takes the removeVibeToolsPreservingActive
path, which deliberately keeps the live active set instead of applying the
source snapshot. #reconcileModeFromSession then re-enters vibe mode, and the
live toolset is by then the reduced vibe set, so the new snapshot was that
reduced set and exiting restored it instead of the target's real pre-vibe
toolset. bash, edit, write, grep, glob, task, and hub were silently gone for
the rest of the session.

Neither a cold start nor switching in from a non-vibe session is affected: the
teardown path does not run, so the live toolset is still the full one when the
snapshot is taken.

Record the snapshot on the vibe mode_change entry and read it back from
sessionContext.modeData on the re-entry path, mirroring how plan mode persists
planFilePath. Sessions written by older versions carry no snapshot and fall
back to the previous behaviour.
2026-08-08 13:39:01 +08:00
enieuwy a1e60c3450 refactor(session): fold the pending-fallback surface into servingModel
Final review found `pendingRetryFallbackModel` unreachable. `servingModel`
returns `undefined` only when the session has no model at all, and the
pending getter required one, so the badge term guarding on it could never
fire. Its case — a fallback armed before anything has served — is already
answered by `servingModel`'s bootstrap, which names the current model and
flags it as fallback-routed. Removed, the same duplicate-surface cleanup
that removed `retryFallbackModel`.

Attribution now anchors on the session id rather than the session file. An
unpersisted session has no file, so two `undefined`s compared equal and
stale attribution survived `/new` and branch switches there; every real
switch mints a new id, persisted or not.

The cooldown-expiry restore keeps `#fallbackRouted` when the stored primary
selector cannot be parsed. Nothing is restored on that path, so the session
is still running on the fallback and its remaining turns are still fallback
work; clearing the flag reported them as the configured primary.

`executor-prewalk`'s fake session predates this work and never set
`servingModel`, so the prewalk hand-off stopped advancing the reported
model once the executor began reading attribution from the session. It now
mirrors the hand-off the way the other executor fixtures do.
2026-08-08 13:16:41 +08:00
enieuwy 0a075662ac fix(session): attribute a run to the model that produced its output
An Agent Hub row reported a subagent as having run on a model that never
spoke. All 97 of its requests, 421K tokens and $6.19 of cost were served
by the primary; a transient stall then armed a fallback, that fallback
errored on its first request with an exhausted quota, and the run died.
Attribution followed the routing switch rather than the output.

Three surfaces lied independently, each re-deriving "the current model"
and calling it the run's model: the executor's progress snapshot, the
session's fallback selector that the hub row reads first, and the
transcript walk behind a settled row.

Sessions now own attribution. `AgentSession.servingModel` names the model
that produced this session's output, holding the last model that actually
served while a candidate is armed but unproven. A switch is a routing
decision, not evidence the target can produce anything, so the answer only
moves once a turn on the target settles.

Consumers read it instead of reconstructing it. The executor's observer
dropped its own event bookkeeping: that bus also carries advisor turns
running on a different model, and it was reading `retry_fallback_applied`
as proof of service. The hub row reads the same getter, so the main
session — which has no executor progress and no persisted history — stops
rendering an unproven candidate as its plain configured model. A fallback
armed before anything has served is still shown, marked as a fallback,
because there is no earlier work to miscredit there.

One predicate decides "this turn produced output", shared by the live
session and the offline replay so they cannot disagree. `error` and
`aborted` are both failures — a stalled stream is finalized as `aborted`
with its partial block still attached, so a stop reason alone proves
nothing — and a turn needs actionable content, which a `length` stop
burning its budget on unsigned thinking does not have. It tolerates
malformed content blocks: transcripts outlive the shapes that wrote them,
and one bad line previously blanked a whole row's history.

Ordering matters at two swap sites. Both the chain advance and the
cooldown-expiry restore move the model and fan `model_changed` out to
subscribers synchronously, so each now updates fallback state before the
swap rather than after; otherwise an observer reading attribution inside
that window sees the incoming candidate carrying the outgoing one's proof.

A startup-selected fallback owns the run from its first request only on a
fresh session. A resumed transcript already holds turns another model
produced, so there the candidate stays unproven until it answers.

`retryFallbackModel` is removed: every consumer reads `servingModel`, and
keeping a parallel derived getter alive for tests is the duplicate surface
this change set exists to remove.
2026-08-08 13:16:11 +08:00
can1357 7cebe901b7 refactor(coding-agent): narrowed over-exported internal symbols
- 28 symbols across discovery, mcp header policy, agent-hub projection and
  rendering, the agent registry, shell tokenizing and changelog comparison
  were exported but referenced only inside their own module; they are now
  module-private, shrinking the deep-import surface.
- Kept AGENT_PLUGIN_MANIFEST_SCHEMA, AGENT_PLUGIN_MCP_SCHEMA,
  parseAgentPluginManifest, clearAgentPluginRootCache and mergeMCPHeaders
  exported: each is a seam for tests that defend real parsing or header
  precedence behavior.
- Nothing reachable from an explicit exports entry or public barrel changed.
2026-08-08 06:32:01 +02:00
can1357 e0ec404de6 refactor(coding-agent): split theme module by responsibility
- theme.ts mixed symbol presets, JSON schema, color math, the Theme class,
  loading, global state, appearance handling and TUI adapters in 3171 lines.
- Symbols, schema, color, theme-class, loader and tui-adapters are now
  siblings; theme.ts keeps global state, the watcher, appearance handling and
  HTML export at 745 lines, with all 44 exports intact.
- Left appearance and export-colors in place: both read private mutable
  auto-theme state, so extracting them would have required new exported
  internals or DI rather than a straight move.
2026-08-08 06:32:01 +02:00
z80 40b2d534a0 fix(coding-agent): surface todo progress in the collapsed panel
While the agent worked through a plan, every sub-todo rendered unchecked
no matter how far along the run was: the phase header highlighted, the
task rows below it looked untouched. Three separate causes, all on the
collapsed path that is the default view.

`selectCollapsedTodos` dropped every closed row while a phase held open
work, so finishing a task only ever *removed* a line — the panel never
rendered a checked box until the whole phase settled. That also made the
card's completion animation dead code: `details.completedTasks` drives a
14-frame strike reveal at 65ms with a component render per tick, against
a row the viewport had already discarded. The existing animation test
missed it by asserting on `expanded: true`.

The viewport now keeps the newest closed task as a checked lead row,
additive to the open-task cap so it never evicts open work, and the
strike sweep lands where users actually see it.

Second, the card gave a `done/total` count to every collapsed untouched
phase but not to the active one, so the phase being worked in was the
single phase reporting no progress. Extracted `formatPhaseProgress` and
put it on every phase header.

Third, the todo auto-clear (`tasks.todoClearDelay`, default 60s) armed
on any list holding a closed task and physically deleted those tasks
from the HUD's copy. An in-flight phase at `3/4` silently became `0/1`
sixty seconds later, fully-closed phases vanished, and stage roman
numerals renumbered off the filtered index — until the next `todo` call
restored the real snapshot. It now fires only once the whole list is
settled, which is the case the setting exists for; the walking viewport
already hides closed rows while work remains.

Progress counters also count closed tasks rather than only completed
ones. The viewport hides abandoned tasks too, so counting only
completions left a phase reading permanently stuck.
2026-08-07 19:49:31 -04:00