Commit Graph

112 Commits

Author SHA1 Message Date
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
can1357 6a44844c57 Merge PR #7377: fix(hub): wake owners when supervised processes exit (@paralin) 2026-08-03 15:12:03 +02:00
roboomp 4997423f10 fix(agent): keep queued non-interruptible tools alive on peer-IRC interrupt
A peer-IRC interrupt (e.g. a subagent message) aborts only interruptible
waits and leaves already-running non-interruptible foreground work alone.
But runTool's `interruptState.triggered` early-return skipped every
not-yet-started tool regardless of source, so a non-interruptible tool
queued behind an interruptible wait in the same batch (a batched todo/write
after `hub wait`) was dropped with "Skipped due to pending peer interrupt".

Exclude non-interruptible tools from the early skip on the IRC path; user
and system steering still preempt all queued work.

Fixes #7493
2026-08-03 12:24:10 +00:00
Christian Stewart 302148523f fix(hub): wake owners when supervised processes exit
Publish terminal daemon completions to the session that started the
process so idle agents can resume without polling hub status.

Persist every unacknowledged generation with a stable completion ID and
immutable snapshot. Replay the collection after reconnect or broker
recovery, and clear each event only after the owning client acknowledges
it.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-08-03 01:31:39 -07:00
can1357 4e65a685ae fix(ai): addressed anthropic stream truncation errors for tool call recovery
- Added `isStreamEnvelopeErrorText` to packages/ai/src/error/flags.ts to recognize stream envelope truncation errors.
- Updated `streamAnthropicOnce` in packages/ai/src/providers/anthropic.ts to throw an envelope error when streams die mid-generation without a terminal frame.
- Updated `recoverTransientErrorToolTurn` in packages/agent/src/agent-loop.ts to recognize Anthropic stream envelope truncation errors for tool call salvage.
2026-08-02 06:14:01 +02:00
roboomp af343f70d0 fix(agent): distinguished started steering aborts
Track entry into tool.execute separately from tool event emission. Never-started skips retain SyntheticToolResultDetails with executed:false; in-flight aborts now use distinct interrupted metadata with execution:started so consumers do not assume no partial work occurred.

Keep both interrupt states neutral in the TUI and cover the agent metadata boundary plus rendering behavior.

Fixes #7199
2026-07-31 21:16:28 +00:00
can1357 daeb683528 feat(agent): restructured tool call dispatch to validate arguments earlier
- Added `prepareToolCallDispatch` and `PreparedToolCall` to handle argument validation and `beforeToolCall` before message snapshotting.
- Implemented `preparedDispatchByMessage` WeakMap to store pre-dispatch results for streamed messages.
- Updated `executeToolCalls` to consume pre-computed dispatch preparation results.
- Updated documentation and changelog to specify `beforeToolCall` timing on the streamed path.
2026-07-27 23:01:50 +02:00
can1357 da6d11de0e feat(agent): introduced prepareToolCall phase supporting argument replacement
- Added a prepareToolCall phase to the agent loop running before tool scheduling for validation and hooks.
- Updated BeforeToolCallContext and result types to support argument replacement instead of in-place mutation.
- Updated coding-agent extension handling and runner to track emitted tool calls and re-evaluate approvals on input revisions.
- Added comprehensive test coverage for argument replacement, concurrency resolution, and schema validation.
2026-07-27 22:55:20 +02:00
can1357 ae01a76136 fix(agent): hardened pre-model-call gate state cleanup and API surface
- Cleared the retained soft-requirement lifecycle alongside the deferred
  hard choice: clearDeferredToolDirectives() owns both, is called from
  clearAllQueues/reset and session-scoped tool-state cleanup, with a
  regression covering reminder re-injection after a queue clear.
- Allowed void-returning pre-model gates via the named AgentBeforeModelCall
  type and normalized gate results in the loop and Agent dispatcher.
- Documented that the first gate installed mid-run applies from the next
  run; corrected the onToolChoiceRejected contract docs; documented the
  cross-run lifetime of ToolChoiceQueue's in-flight claim.
- Removed the unused addBeforeModelContextBuild hook.
- Relocated both packages' changelog entries out of the released 17.1.4
  sections into Unreleased with PR attribution, folding the never-shipped
  Fixed bullet into Added and noting the input-event timing change.
2026-07-27 14:08:53 +02:00
Christian Stewart 01ecab6df7 fix(agent): clear deferred choices on branches
A pre-model gate can defer a claimed hard tool choice for the next call. Branch transitions cleared the coding-agent queue but left that agent-owned value alive, allowing an obsolete forced tool to cross into the replacement transcript.

Expose the narrow deferred-choice reset at the Agent owner and invoke it from the shared session-scoped tool-state cleanup used by both branch paths. Failed session switches retain their existing rollback behavior.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:05:30 -07:00
Christian Stewart c5e5f6dbc3 fix(agent): close aborted harmony retry turns
A Harmony retry keeps its logical turn open while the next provider call is prepared and gated. An abort during that gate previously ended the agent stream directly, leaving observers with an unmatched turn_start event.

Route the aborted gate through the existing pre-model stop owner so it emits the synthetic aborted message and closes the open turn before ending the stream. Fresh turns retain their existing no-provider-call cancellation behavior.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:02:18 -07:00
Christian Stewart e79eabc3b6 feat(agent): add a pre-model-call gate that can stop the turn
The agent loop had no place to refuse a provider request. A host that needs to
act on the assembled context before it is billed, checking that the prompt still
fits the window, that a budget boundary has not been crossed, or that the
session should hand off instead of spending, could only observe the request
after the fact, when the tokens were already committed.

Add `AgentLoopConfig.beforeModelCall`, asked once per turn beside the deadline
check and before `turn_start` is emitted. A `stop` result ends the stream with
no turn open, so nothing has to synthesise a cancellation event and no consumer
is left holding a half-open turn. Placing it there also keeps `turn_end`'s
contract intact: that event carries the assistant message for a completed turn,
and a gated stop has no assistant message to report.

`syncContextBeforeModelCall` keeps its existing void contract and its job of
refreshing prompt and tool state, so implementations typed as returning void are
unaffected.

`Agent.setBeforeModelCall` installs the host's callback, and `addBeforeModelCall`
registers an additional callback without displacing the host's, returning a
disposer so an extension can attach and detach independently. A supplied
`reason` is logged where the loop stops.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-26 12:02:18 -07:00
Christian Stewart d363ca0968 feat(agent): wake steering on an event instead of polling for it
While a turn is running, the loop watched for out-of-band steering by
re-checking the queue on a fixed interval. The interval sets a floor on how long
a message waits and burns a wakeup on every tick that finds nothing, and the
delay is worst exactly when it matters: a cancellation or a correction issued
during a long tool loop sits until the next tick.

Wait on the queue instead. `Agent` keeps a set of steering waiters and notifies
them whenever a message is enqueued, and the loop awaits
`waitForSteeringMessages`, re-checking only when woken. The timer path remains
for the IRC interrupt queue, which is session-owned and has no wake callback, so
nothing regresses where no event source exists.

Every wait races against local abort. The callback contract does not require an
implementation to observe the signal, and one that resolves only on the next
queue event would otherwise never settle once a batch finished, so awaiting it
during teardown would hang a batch that simply had no steer.

Which controllers a steer raises is unchanged: this is a timing change, not a
change to what an interrupt does.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-25 03:54:46 -07:00
can1357 9f8aa87dbf feat(task): removed per-call model override from task tool
- Removes `model` field from task item/schema, TaskParams, and TaskItem types.
- Removes model selector validation, formatting, and approval display logic.
- Updates task tool priority docs to reflect that model is no longer per-call overridable.
- Updates eval agent() helper docs and prompt templates to remove model parameter.
- Updates tests to reflect removal of model override capability.
2026-07-24 14:51:48 +02:00
can1357 c780662881 feat(agent): added resolveFallbackTool option for routing unadvertised tool calls
- Add `resolveFallbackTool` callback to `AgentOptions` and `AgentLoopConfig` that resolves tool calls not found in the advertised set.
- Use the callback as a third lookup step after `name` and `customWireName` match, enabling side transports like `xd://` device mounts.
- Add test coverage verifying the fallback resolves known devices and preserves "not found" errors for unknown names.
- Wire the coding agent's device registry as `resolveFallbackTool` in both `createAgentSession` and `streamAgentSession` paths.
2026-07-24 12:18:17 +02:00
usr-bin-roygbiv b9504f65e7 feat: add native Codex computer use 2026-07-24 01:40:04 +00:00
usr_bin_roygbiv 7c6d691c10 fix(agent): recover tools after stream parse errors 2026-07-20 17:22:25 -05:00
roboomp a713b941dc fix(agent): preserved side-effecting hub outcomes
Resolved tool interruptibility from each call's raw arguments so mixed-operation tools can keep side-effecting calls non-interruptible.

Restricted the unified hub to interrupt passive waits and followed logs while preserving start, send, and lifecycle operation results.

Fixes #5995
2026-07-18 14:41:19 +00:00
roboomp ed4ddcba6c fix(session): await session.dispose() on non-interactive exit paths
Print-mode assistant-error/aborted exit, RPC pi.shutdown() and stdin-EOF
shutdowns, and the extension command-context shutdown() called
process.exit() before (or racing) session.dispose(), skipping the bounded
browser reaper (releaseTabsForOwner) installed in dispose(). An OMP-owned
Chromium could survive the parent and reparent to PID 1.

Route all four graceful paths through the idempotent, promise-memoized
session.dispose() and await it before the final exit. The RPC
performShutdown no longer emits session_shutdown directly (dispose() emits
it), avoiding a double emit.

Fixes #5643
2026-07-16 04:29:55 +02:00
can1357 e28197c694 Revert "merged PR #5602: fix(agent-loop): reclassify empty toolUse stop as retryable error"
This reverts commit 2baaea8782, reversing
changes made to 6783c3a473.
2026-07-16 04:09:44 +02:00
can1357 1df790a5d2 Revert "fix(agent-loop): discard incomplete sibling tool calls"
This reverts commit 7a2e34988b.
2026-07-16 04:08:59 +02:00
can1357 2c355102ce chore: applied biome formatting to merged sources 2026-07-16 03:52:48 +02:00
can1357 7a2e34988b fix(agent-loop): discard incomplete sibling tool calls 2026-07-16 03:31:57 +02:00
roboomp 4200dec047 fix(agent-loop): strip incomplete tool calls on empty toolUse stop
A dropped stream that emitted toolcall_start/delta but never toolcall_end
leaves an incomplete toolCall block in content. The prior guard bailed on
any toolCall block, so the outer loop dispatched empty or partially
parsed arguments instead of retrying the transport failure.

Track streamed tool-call ids (toolcall_start/delta) alongside completed
ones (toolcall_end): a call streamed but never completed is incomplete.
reclassifyEmptyToolUseStop now reclassifies unless a usable (atomic or
completed) tool call remains, stripping incomplete blocks first. Atomic
deliveries (single done/end(result) message, e.g. Cursor) emit no
granular events and stay usable.

Fixes #5600
2026-07-15 18:06:13 +00:00
roboomp 94fc54859f fix(agent-loop): reclassify empty toolUse in result-only completion
Streams finalized via end(result) with no terminal done/error event fall
through to the trailing-result branch, which returned response.result()
unchanged. An empty toolUse turn completing that way stayed a silent
success and never retried. Apply the same
retainCompletedToolCalls/recoverTransientErrorToolTurn/
reclassifyEmptyToolUseStop chain to the trailing branch.

Fixes #5600
2026-07-15 17:52:25 +00:00
roboomp f64a17c529 fix(agent-loop): reclassified empty toolUse stop as retryable error
A provider stream that closes after the thinking block but before the
tool call JSON is emitted finalizes with stopReason=toolUse and zero
toolCall content blocks. The loop treated this as a successful turn:
dispatched no tools, rendered an empty tool widget, and never retried.

reclassifyEmptyToolUseStop stamps such a turn as stopReason=error with
the Transient classifier bit set explicitly, so AgentSession's standard
retry-with-backoff path fires regardless of message-text matching.

Fixes #5600
2026-07-15 17:44:53 +00:00
can1357 b58fa6e1d3 Merge PR #5513: fix(agent): keep completed tool results from false "skipped" placeholder (@roboomp) 2026-07-14 23:11:08 +02:00
roboomp b591f3d054 fix(agent): kept completed tool results from false skipped placeholder
The interrupt-clobber branch in executeToolCalls replaced a tool's real
result with the "Skipped due to queued user message" placeholder whenever
an interrupt fired, the tool's own signal aborted, and the result was an
error. It ignored whether tool.execute() actually completed, so steering
while a tool was in flight could discard a genuine error result (e.g. a
command exiting non-zero) that the tool had already produced.

Gate the clobber on !completedToolExecution so a tool that ran to
completion keeps its real result; only tools cut off before returning are
reported as skipped. Align the aborted telemetry status the same way.

Fixes #4752
2026-07-14 19:57:38 +00:00
roboomp 8e6d26b1e8 fix(eval): honored unlimited cell timeouts
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.

Fixes #5250
2026-07-14 17:33:56 +00:00
roboomp 74c63fa6c5 fix(agent): labeled system steering skips accurately
- Carried steering queue origin through mid-batch interrupt polling.
- Preserved queued-user skip wording while labeling advisor/system steering as system advisory skips.
- Added regression coverage for advisor steering skip wording.

Fixes #5074
2026-07-10 14:00:59 +00:00
can1357 17c7c6d0f8 fix(agent): preserve external abort boundaries 2026-07-10 12:37:46 +02:00
can1357 82f3706410 Merge remote-tracking branch 'origin/farm/6f26b11c/distinguish-synthetic-tool-failure' 2026-07-02 23:42:38 +02:00
roboomp 107503d201 test(agent-loop): align tool result details type with parameter shape 2026-07-02 21:27:32 +00:00
roboomp eb64e7e5d1 fix(agent-loop): skip cursor exec-resolved toolCall blocks to avoid double-execution
Codex review on PR #4351: synthesizing toolCall content blocks for
Cursor's exec-channel native tools made the shared agent loop treat the
finalized assistant message as a fresh runnable tool turn. Because
executeToolCalls filters message.content for any toolCall block on
stop/toolUse, bash/write/delete/etc. ran a second time after Cursor
already executed them server-side via the bridge, duplicating side
effects and appending conflicting toolResults.

- packages/ai/src/utils/block-symbols.ts: add `kCursorExecResolved`
  symbol and `CursorExecResolvedCarrier` carrier type. Symbol-keyed so
  the marker never leaks into JSONL; rebuild pairs blocks with toolResult
  messages by id.
- packages/ai/src/providers/cursor.ts: stamp the marker onto every
  block `synthesizeCursorExecToolCall` emits and extend `ToolCallState`.
- packages/agent/src/agent-loop.ts: filter marked blocks out of the
  runnable-toolCall extraction in both the main runnable path and the
  error/aborted placeholder path, plus defense-in-depth inside
  `executeToolCalls`. Marked blocks stay in `assistantMessage.content`
  for persistence + rebuild rendering; they just never re-execute.
- packages/agent/test/agent-loop.test.ts: two regression tests — one
  proves a marked block yields zero `tool.execute` calls and no
  `tool_execution_*` events from the loop, the other verifies mixed
  batches still run the unmarked blocks unchanged.

Fixes #4348
2026-07-02 21:26:37 +00:00
roboomp 3d86935a5e fix(agent): distinguish synthetic placeholder tool results from real tool failures
When an assistant turn ends with stopReason="error" after a tool call
was already streamed, agent-loop synthesizes a placeholder tool result
via createAbortedToolResult() to preserve the tool_use / tool_result
pairing the provider API requires. The previous wording ("Tool
execution failed due to an error: <upstream>") and event shape
(normal tool_execution_start / tool_execution_end with empty details)
were indistinguishable from a real local tool failure — a Codex
websocket close mid-turn showed up in the CLI as a broken Edit panel,
misattributing provider-transport faults to the local tool.

Reword the "error" placeholder to state explicitly that the tool
never ran ("Tool call was not executed because the provider stream
ended with an error before the tool could run: <upstream>") and thread
a SyntheticToolResultDetails discriminator ({ __synthetic: true,
source: "assistant_stop_error" | "assistant_stop_aborted" |
"assistant_stop_skipped" | "assistant_stop_length", executed: false,
upstreamError }) through both the ToolResultMessage.details and the
tool_execution_end event's result.details, so downstream UI/telemetry/
ACP consumers can render "provider transport failed, tool not
executed" without string-matching content.

Fixes #4321
2026-07-02 15:19:30 +00:00
can1357 21f5728ef8 fix(agent): prevented consuming legacy steering queue during mid-batch interrupts
- Stopped calling the consuming `getSteeringMessages` getter during mid-batch interrupt polls to prevent stranding or dropping messages before they reach the injection boundary.
- Skip subsequent steering checks in the poll loop once an interrupt has already triggered.
- Added a regression test to ensure legacy steering remains queued until the injection boundary when no non-consuming peek exists.
2026-07-02 03:34:14 +02:00
can1357 021d4fc1e3 Merge PR #4161: fix(agent): interrupt waits for IRC delivery (@roboomp) 2026-07-01 21:53:19 +02:00
can1357 23ea5e0808 test(agent): require queued skip text content 2026-07-01 21:47:55 +02:00
roboomp 1754c108df fix(agent): scoped irc-only aborts to interruptible tools
Previously an IRC-only interrupt shared the batch-wide abort controller with
user steering, so a peer message that landed while an interruptible wait ran
alongside a foreground non-interruptible tool (e.g. bash) killed the foreground
tool too. Split the batch signal into a shared steering/external channel and an
interruptible-only IRC channel; each record picks its per-tool signal based on
the tool's interruptible flag, and only that signal is used for validation,
before/after hooks, and execute. User steering still upgrades an in-flight IRC
interrupt to a full batch abort.
2026-07-01 17:14:04 +00:00
roboomp 619bfda3eb fix(agent): interrupted irc waits
Fixes #4160
2026-07-01 16:56:08 +00:00
Wolfgang Schoenberger 0d9ef89549 fix(agent): clarify queued skipped tool results 2026-06-29 19:46:13 -07:00
ben dc17aead5f fix(agent): narrow stream recovery review fixes 2026-06-28 16:18:46 +08:00
ben cb4e68888f fix(agent): recover completed tools after stream read errors 2026-06-28 16:18:46 +08:00
can1357 0b5e2d8276 fix(ai-providers): normalized Anthropic tool call IDs
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
2026-06-21 06:04:49 +02:00
roboomp 19d103b708 style: bun run fix 2026-06-19 15:14:50 +00:00
can1357 67f6518e42 feat: enhanced tool robustness, improve authentication flow, and update API parameters
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
2026-06-19 16:46:07 +02:00
can1357 d46bd0139b refactor(coding-agent): fixed system prompt customization path
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
2026-06-19 16:16:37 +02:00
can1357 865694df7e fix: resolved streaming stability and reference isolation issues
- Prevent object reference sharing between agent snapshots and stream events by deep-cloning tool-call arguments.
- Stabilize GFM tables and Mermaid diagrams during streaming by delaying transcript block commits until content finalization.
- Implement session resume safety to prevent crashes when working directories are missing.
- Add comprehensive test suites to verify streaming commit stability and immutable snapshot isolation.
2026-06-19 16:09:00 +02:00
can1357 4f03180ae6 refactor(deps): moved intent field constant to pi-wire
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
2026-06-19 04:58:29 +02:00
can1357 291b3c74c2 feat: enhanced model reasoning, schema normalization, and loop guarding
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
2026-06-18 04:51:43 +02:00