Commit Graph

606 Commits

Author SHA1 Message Date
can1357 f8f8136021 refactor(coding-agent): privatized the legacy nextToolChoice method to
- Privatized the legacy `nextToolChoice` method to `#nextHardToolChoice` to ensure all tool-choice directives flow through the unified `nextToolChoiceDirective` entry point.
- Eliminated redundant dual entry points for fetching tool choices, which previously bypassed the soft pending-preview lifecycle.
- Updated test suites to consume `nextToolChoiceDirective` where appropriate to maintain consistency with internal agent-loop logic.
2026-06-19 22:24:12 +02:00
roboomp 49317a0f1c fix(agent): cleared stale pending preview gates
Cleared pending-preview markers when resolve has no runnable handler so a stale gate cannot keep forcing resolve after the invoker is gone.

Added regression coverage for apply and discard draining stale pending markers.

Fixes #3061
2026-06-19 17:54:37 +00:00
can1357 b2dc706ee7 feat: enabled auto-retry for AI thinking loops
- Improved thinking loop detection logic by refining text normalization and treating stalls as retryable errors.
- Instrumented the agent session to recognize thinking loop markers within retryable error conditions.
- Automated the clearing of stale error banners upon successful auto-retry execution.
- Added comprehensive test coverage for chunked thinking loop errors and banner management.
2026-06-19 17:51:53 +02:00
can1357 480aa1763c Merge PR #3034: fix(mnemopi): isolate local embeddings worker in subprocess (@roboomp) 2026-06-19 17:16:36 +02:00
can1357 a1dae58c34 fix(coding-agent/session): prevented emission of unhandled agent events
- Added an early return to stop event propagation if no handlers are registered for the event type.
- Updated the `agent_start` logic to return early, ensuring consistent flow control.
2026-06-19 16:06:17 +02:00
cagedbird043 5ff5f110af fix(agent): dynamically fallback from snapcompact to text summary on high CJK/non-ASCII rates 2026-06-19 21:50:16 +08:00
roboomp bf56a76e59 fix(mnemopi): isolated local embeddings worker in subprocess to skip onnxruntime napi crash
Moved mnemopi's local embedding provider out of the main agent process by
spawning a dedicated `Bun.spawn` child for fastembed + onnxruntime-node. The
agent CLI gains a hidden `__omp_worker_mnemopi_embed` dispatch; `loadMnemopi`
/ `loadMnemopiCore` install the subprocess-backed initializer through the
newly-exposed `setLocalModelInitializer` seam so every `embed()` call
round-trips through IPC instead of loading the NAPI module. The parent
SIGKILLs the child on dispose so the destructor that segfaults Bun on
Windows shutdown (NAPI finalizer at exit on npm installs, `process.dlopen`
constructor at session start on standalone binaries) never runs in any
address space the agent owns. Mirrors the tiny-model fix from #1607.

Adds `smokeTestMnemopiEmbedWorker` to `omp --smoke-test`, a new
`test/issue-3031-repro.test.ts` that pins the spawn/dispatch/signal-exit
contract and forbids re-importing `fastembed-runtime` from the agent surface,
and changelog entries.

Fixes #3031
2026-06-19 08:09:50 +00:00
can1357 7864a1363c feat(coding-agent): consolidated macOS power settings into enum
- Replaced four individual power boolean settings with a single `power.sleepPrevention` enum for improved configuration management.
- Implemented automatic migration logic in `Settings.init` to map legacy macOS power booleans to the new enum levels.
- Updated `AgentSession` to utilize the new `sleepPrevention` enum for macOS power assertion lifecycle management.
- Changed default value of `display.cacheMissMarker` to `false` to suppress cache-miss markers.
- Added comprehensive unit tests in `settings-manager.test.ts` to verify backward-compatible migration behavior.
2026-06-19 08:54:20 +02:00
can1357 65f0d0a245 refactor(coding-agent): renamed repeatToolDescriptions to inlineToolDescriptors
- Renamed `repeatToolDescriptions` to `inlineToolDescriptors` throughout configuration, SDK, and internal session management.
- Set the default value to `true` and updated descriptions to clarify the descriptor inlining behavior.
- Fixed the `/dump` command to prevent duplicate tool inventory output when inlining is enabled.
2026-06-19 08:05:49 +02:00
can1357 ed73429faa feat: implemented soft tool requirement and escalation system
- Introduced `SoftToolRequirement` to support non-invasive tool enforcement with lifecycle management and escalation.
- Added `ToolChoiceDirective` to coordinate hard and soft tool requirements within the agent loop.
- Optimized preview workflows in `coding-agent` by replacing forced tool choices with non-forcing pending invokers.
- Enhanced `CompactionSummaryMessage` to prioritize structured rendering for tool requirement reminders.
2026-06-19 07:51:27 +02:00
can1357 b137764f0c feat: implemented text-first foveated snapcompact archive layout
- Implemented text-first archive structure that incorporates bounded source text with foveated image frames (HQ edges, LQ middle) to improve context quality.
- Updated snapcompact core logic to use `historyBlocks` for reconstruction, transitioning away from reliance on previous PNG frame inheritance.
- Increased default frame limits to 80 and adjusted token estimation constants (5024) to enhance context budget accuracy.
- Added `pruneToolDescriptions` and improved multi-block summary token estimation to optimize tool spec usage.
2026-06-19 07:38:54 +02:00
can1357 dcb16d86f6 perf(compaction): kept per-turn tool-result pruning inside the warm prompt-cache prefix
- Added `keepBoundaryId` and `cacheWarmSuffixTokens` guards to `pruneToolOutputs` and `pruneSupersededToolResults` so superseded/useless results sitting in the already-sent cached prefix are no longer rewritten mid-session, with new `computeMessageSuffixTokens`/`resolveBoundaryIndex` helpers.
- Added a `keepBoundaryId` floor to `collectShakeRegions` so shake skips entries summarized away by the latest compaction.
- Threaded `firstKeptEntryId`, `PRUNE_CACHE_WARM_SUFFIX_TOKENS` (8k) and `PRUNE_IDLE_FLUSH_MS` (90m, above the 1h cache TTL) from `#pruneToolOutputs`, `#pruneStaleToolResults` and `shake` in `agent-session.ts`.
- Added six boundary tests in `supersede-prune.test.ts` covering warm-prefix protection, tail-case pruning, and the pre-boundary floor.
2026-06-19 06:57:53 +02:00
can1357 96151e6c30 feat: implemented tool description pruning to reduce token usage
- Added utility functions to strip descriptions from JSON schemas and tool definitions for optimized token output.
- Integrated `pruneToolDescriptions` configuration across agent loops and sessions to enable optional schema pruning.
- Updated agent context and snapshot logic to propagate pruning settings and maintain fingerprinting integrity.
- Verified schema structural integrity and removal of annotation descriptions through new unit tests.
2026-06-19 06:55:19 +02:00
can1357 f01b5e4c27 refactor(snapcompact): updated MAX_FRAMES_DEFAULT to 80 to better
- Update `MAX_FRAMES_DEFAULT` to 80 to better utilize high-capacity model context windows.
- Remove `providerFrameBudget` from `snapcompact` to decouple archival limits from provider-specific image caps.
- Update `compact` logic to treat `maxFrames` as a hard upper bound rather than a provider-clamped limit.
2026-06-19 06:33:19 +02:00
can1357 29d250fae2 feat(coding-agent): supported advisor transcript persistence
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
2026-06-19 03:50:00 +02:00
can1357 71144825ec feat(coding-agent): extended compact command with submode support
- Added a `mode` property to `CompactOptions` to allow fine-grained control over compaction strategies.
- Implemented `soft`, `remote`, and `snapcompact` submode overrides for the `/compact` command.
- Integrated `parseCompactArgs` to enable robust subcommand routing and validation, including focus instruction rejection for specific modes.
- Established a `CompactMode` registry to manage compaction strategies and verify remote availability.
2026-06-19 02:55:21 +02:00
can1357 5168e6d6af Merge PR #2791: fix(tool): resolve pasted image attachments in inspect_image (@roboomp)
# Conflicts:
#	packages/coding-agent/src/tools/inspect-image.ts
2026-06-19 00:59:03 +02:00
can1357 5d72ce1237 Merge PR #2729: fix(coding-agent): accept max thinking alias (@roboomp)
# Conflicts:
#	packages/coding-agent/test/model-resolver.test.ts
#	packages/coding-agent/test/sdk-model-selection.test.ts
2026-06-19 00:58:52 +02:00
can1357 6d5f548147 fix(mcp): bound owned-manager disconnect on dispose
The dispose() disconnect added by this PR awaited mcpManager.disconnectAll()
unbounded. An owned manager holding an HTTP/SSE server whose session-
termination DELETE hangs would block dispose for the full MCP request timeout
(30s default, unbounded when OMP_MCP_TIMEOUT_MS=0), stalling /exit and
print-mode shutdown on a broken remote endpoint.

Wrap the disconnect in withTimeout(..., 3_000) — mirroring the bounded
async-job teardown two lines above and the startup bound from issue #2100.
stdio close (the subprocess reap this PR targets) completes well within the
bound; a slow transport close is left to finish detached.

Adds a regression test driving the real MCPManager.disconnectAll() with a
stalled transport close.
2026-06-18 23:32:55 +02:00
can1357 469e4626d0 Merge PR #2839: fix(mcp): disconnect owned MCP manager on AgentSession.dispose() (@jms830) 2026-06-18 23:32:54 +02:00
can1357 7c635de10a Merge PR #2914: fix(coding-agent): preserve queued steers/follow-ups during auto-compaction (@metaphorics) 2026-06-18 23:32:54 +02:00
can1357 a1dbbb50c9 Merge PR #2986: fix(session): prevent goal mode pause during compaction/switch (@roboomp) 2026-06-18 22:53:39 +02:00
can1357 dbf4b53163 fix(coding-agent): re-inject approved plan reference after compaction
After plan approval the executor delivers the plan-mode-reference exactly
once and sets `#planReferenceSent = true`. Both compaction paths — `compact()`
and `#runAutoCompaction()` — replace the conversation history that carried
that reference but never cleared the flag, so `#buildPlanReferenceMessage()`
short-circuited to null on every subsequent turn and the executor permanently
lost the plan it was working on (exactly the long-session failure reported).

Clear `#planReferenceSent` right after `replaceMessages()` in both paths so the
next turn re-reads the plan from disk and re-injects it. The reset is a no-op
for ordinary sessions: the default plan path (PLAN.md in session-local scratch)
has no file on disk, so `#buildPlanReferenceMessage()` still returns null there.

Adds a deterministic regression test (short-circuited compaction, mock stream)
that fails before this change and passes after, plus a guard proving normal
sessions get no spurious plan injection.

Fixes #1246
2026-06-18 22:52:29 +02:00
usr_bin_roygbiv c16a41fd95 fix(session): prevent goal mode pause during compaction/switch 2026-06-18 14:44:39 -05:00
can1357 81563e0155 feat(coding-agent): improved retry safety for partial message streams
- Refactor replay safety logic to specifically prevent retries when a tool call is present in the assistant message.
- Enable retries for transient stream errors occurring during partial text or thinking sequences, ensuring consistency when no tool call has been completed.
- Add regression tests to verify that transient socket closures are successfully recovered while completed tool calls remain protected from redundant retries.
2026-06-18 18:58:15 +02:00
can1357 30a133cef7 feat(coding-agent): prevented stale advisor data leakage across conversations
- Add an epoch counter to `AdvisorRuntime` to discard in-flight advisor batches when a reset or disposal occurs.
- Introduce `resetAdvisorSessionState` to clear advisor-specific queues, latches, and pending cards, ensuring pre-reset state does not interfere with new conversations.
- Extend `YieldQueue.clear` to support conditional clearing by entry kind.
2026-06-18 18:58:01 +02:00
can1357 8fcb45b9fe Merge remote-tracking branch 'origin/farm/07fc7017/fix-plan-refine-abort' 2026-06-18 16:58:30 +02:00
Can Bölük bdb9bd03e0 Merge pull request #2947 from can1357/farm/130a952c/opencode-go-usage
fix(providers): add OpenCode Go usage reporting
2026-06-18 16:17:45 +02:00
roboomp 8c0dce29c3 fix(coding-agent): hid plan refine approval abort
Mark the plan approval abort as an internal transition so choosing Refine plan returns to the editor without rendering Operation aborted.

Fixes #2971
2026-06-18 13:32:25 +00:00
roboomp 9e7bb85411 fix(usage): expired opencode cache after cost writes
- Expired the per-credential usage report cache after recording observed OpenCode Go spend.\n- Threaded provider base URL into OpenCode Go cost recording so the invalidation targets the same cache key /usage uses.\n- Added regression coverage for refreshing cached OpenCode Go limits immediately after a completed turn.\n\nFixes #2942
2026-06-18 07:52:45 +00:00
roboomp 2b9047b79c fix(providers): added opencode go usage tracking
- Added an OpenCode Go usage provider that synthesizes 5h, weekly, and monthly cap windows from OMP-observed request costs.\n- Recorded OpenCode Go assistant request costs against the active credential so /usage can report local cap utilization.\n- Added regression coverage for fresh keys and observed spend aggregation.\n\nFixes #2942
2026-06-18 07:22:56 +00:00
usr_bin_roygbiv 2634bbf1d2 fix(coding-agent): bypass stream-interrupted guard for Gemini malformed function calls 2026-06-18 00:42:22 -05:00
can1357 291b3c74c2 feat: enhanced model reasoning, schema normalization, and loop guarding
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
2026-06-18 04:51:43 +02:00
can1357 4a1f3483c7 feat(coding-agent): promote completed /btw answers into branches (#2894) 2026-06-18 02:50:23 +02:00
can1357 39cd8d19c0 Merge remote-tracking branch 'origin/farm/9eb49b02/fast-fallback-compaction-timeouts' 2026-06-18 02:31:04 +02:00
roboomp 492c528aa3 fix(agent): skipped compaction timeout retries
Stopped auto context-full maintenance from retrying repeated summarization timeouts on the same model before fallback. Added a regression test for timeout fast-fallback behavior.\n\nFixes #2913
2026-06-18 00:25:17 +00:00
metaphorics 46f4993f65 fix(coding-agent): preserved queued steers/follow-ups during auto-compaction
Two defects dropped the first steering/follow-up message typed as
auto-compaction began:

- The compaction AbortController (which backs isCompacting) was installed
  AFTER auto_compaction_start was emitted. The emit awaits extension
  delivery and yields to the event loop, so a message typed as the loader
  appeared was read while isCompacting was false and mis-routed into the
  core agent queue (which the handoff reset then wiped). Install the
  controller before the emit, and move the emit to the first statement
  inside the existing try so the catch/finally cleanup still runs, in both
  #runAutoCompaction and #runAutoShake. The handoff branch now passes the
  run's local abort signal instead of the mutable controller field, so a
  superseded run bails at the handoff entry check rather than resetting the
  session.

- handoff() calls agent.reset(), which clears the core steering/follow-up
  queues. Capture both queues immediately before reset and restore them
  immediately after (synchronous, no await gap), so queued steers and
  follow-ups -- including in-flight RPC/SDK steer()/followUp() and a hidden
  user companion such as an ultrathink notice -- survive the new-session
  reset instead of being silently dropped.

Adds regression tests for both defects (controller-before-emit for the
context-full and shake paths, queue preservation across the reset for both
pre-enqueued and in-flight messages).
2026-06-18 09:15:01 +09:00
can1357 3f304ee5a8 refactor(agent-core): isolated token counting into a new local tokenizer
- Extracted native token counting into a new localized `tokenizer.ts` wrapping `@oh-my-pi/pi-natives`.
- Introduced a lightning-fast byte-length estimation logic for token counting when accurate counting is disabled.
- Diverted token calculations to the faster estimator during test environments and when `PI_TOKENIZER_ACCURATE` is falsy.
- Updated agent base and coding-agent sessions to consume the new localized `countTokens` utility.
2026-06-18 01:57:23 +02:00
can1357 17ce678468 fix(coding-agent): handled provider error finish reasons and preserve subprocess failure codes
- Added detection for provider error finish reasons occurring before tool calls to identify fatal messages.
- Prevented subprocess tool execution finalization from resetting a non-zero exit code when yield items exist.
- Ensured a default error message is set in stderr when a subprocess fails after yielding a result.
2026-06-18 01:11:10 +02:00
can1357 a050474af7 feat: migrated validation schemas and tool definitions from Zod to ArkType
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
2026-06-18 00:59:53 +02:00
can1357 d4317d3d20 feat(ai): consolidated OpenAI-family streaming and add OpenRouter API support
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.

Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
    - Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
    - Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
    - Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
2026-06-17 21:36:48 +02:00
KamijoToma da29a25e54 fix(coding-agent): cancel btw post-prompt work 2026-06-18 03:10:09 +08:00
can1357 eb67863e75 feat(coding-agent/tools): secured browser stealth scripts against detection
- Bound native Reflect methods to local variables in the browser launch script to prevent detection.
- Updated stealth injection scripts to use the bound Reflect methods instead of global Reflect calls.
- Fixed a type assertion issue in the AgentSession tool proxy.
2026-06-17 20:59:20 +02:00
KamijoToma b051bcf627 fix(coding-agent): block btw branch during maintenance 2026-06-18 02:59:16 +08:00
KamijoToma 2bbb16c49b fix(coding-agent): harden btw branch state 2026-06-18 02:35:33 +08:00
KamijoToma 9e67a507b9 fix(coding-agent): stabilize btw branch context 2026-06-18 01:37:50 +08:00
KamijoToma d4df729396 fix(coding-agent): harden btw branch promotion 2026-06-18 01:15:07 +08:00
KamijoToma 52b3155cbd feat(coding-agent): branch completed btw answers 2026-06-18 00:58:23 +08:00
can1357 af6e83a651 fix(coding-agent): replaced JSON string equality checks with deep equality
- Changed provider delta input comparisons in `buildResponsesDeltaInput` to use `Bun.deepEquals` instead of stringified JSON checks.
- Updated session message diffing to return early on length changes and compare normalized entries with deep equality.
- Replaced JSON-string assertions in the issue-966 repro test with `Bun.deepEquals` for stable equality checks.
2026-06-17 13:45:11 +02:00
can1357 54d212ec33 feat(coding-agent): added advisor immune-turn window for concern/blocker interruptions
- Tracked completed primary turns in `AgentSession` and started an immune-turn window after each interrupting advisor steer.
- Routed follow-on `concern`/`blocker` notes to the aside channel while the immune window is active, while preserving prior auto-resume-suppressed handling.
- Added the `advisor.immuneTurns` setting and tests for immune-turn detection and delivery-channel decisions.
2026-06-17 13:21:47 +02:00