Commit Graph

67 Commits

Author SHA1 Message Date
can1357 e429166673 fix(coding-agent): queued extension sendUserMessage as steer while streaming
Extension sendUserMessage() without deliverAs fell through to prompt(),
which throws AgentBusyError during an active stream; the message was
dropped and surfaced as 'Extension sendUserMessage failed'. Route the
omitted-deliverAs path through prompt() with streamingBehavior 'steer'
so streaming queues a steer with normal prompt-flow side effects
(keyword notices, advisor auto-resume reset) and idle still starts a
turn.

ACP skill-command prompts now pass streamingBehavior 'steer'; the RPC
skill fast-path honors the prompt command's streamingBehavior field
(default steer) like the plain-prompt path already did. Documented the
extension-facing delivery semantics.

Synthesized from PR #4942 (prompt-flow steer routing, docs, tests) and
PR #4922 (RPC streamingBehavior threading, steer regression test);
dropped PR #4942's unrelated workflow-notice.md ellipsis churn.

Fixes #4923

Co-authored-by: roboomp <omp@can.ac>
Co-authored-by: metaphorics <metaphorics@users.noreply.github.com>
2026-07-09 18:27:22 +02:00
can1357 53df3c82b7 style: applied biome formatting to merged sources 2026-07-08 15:37:43 +02:00
can1357 1bb29873ea fix(agent): adapted scoped TTSR abort labels to completed-call retention
- derive per-tool abort labels from a tool-scoped abort signal for provider-built aborted messages
- restore main's single-call TTSR label test dropped by the merge
- complete the innocent read in the sibling-label test; incomplete matched calls mint no placeholder under the retention policy
2026-07-08 15:34:37 +02:00
can1357 32158b74ea merge PR #4542: fix(coding-agent): scoped TTSR abort reason to matching tool call 2026-07-08 15:27:25 +02:00
roboomp 8b17f0d3a1 fix(coding-agent): scoped TTSR abort reason to matching tool call
TTSR stream-interrupt aborts now carry a per-tool reason so the placeholder
loop labels only the tool call whose stream matched the rule with the rule
name and gives sibling committed tool calls a neutral "TTSR interrupt on
another tool call" reason. Previously the single `message.errorMessage`
was stamped onto every retained tool-call block, so unrelated read/edit
calls read as violating a rule they never matched and misled the model's
own reasoning about which call fired.

Threads the matched `toolcall:<id>` extracted from the TTSR match context
through `agent.abort(...)` as a `ToolScopedAbortReason` object; the agent
loop unwraps it in `emitAbortedAssistantMessage` into a
`toolCallAbortMessages` map on the aborted `AssistantMessage`, and the
`stopReason === "aborted"` fanout in `runAgentLoop` prefers the per-tool
message when one exists.

Fixes #2783
2026-07-04 17:38:51 +00:00
oldschoola a2854ba768 fix: migrate coding-agent tests from fs.rm to removeWithRetries
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.

The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
2026-06-23 15:28:05 -07:00
can1357 076fd0de81 test(coding-agent): realign discovery and TTSR fixtures with current tool invariants
- sdk-mcp-discovery: `find` became an essential tool (2eef88978), so it can no longer be hidden/rediscovered under `tools.discoveryMode: all`. Switch the discoverable-tool assertions to `search` (still `loadMode: discoverable`).
- agent-session-concurrent: the agent loop now drops tool calls that never reached `toolcall_end` from an aborted turn (0890b2be6, partial args are unsafe to replay). Emit `toolcall_end` before the TTSR rule-driven abort so the labeled placeholder result is minted.
2026-06-21 08:37:37 +02:00
can1357 a050474af7 feat: migrated validation schemas and tool definitions from Zod to ArkType
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
2026-06-18 00:59:53 +02:00
can1357 6e918712fd fix(coding-agent): guard session stop continuations
- Guarded `session_stop` continuations against aborted or superseded hook results before queueing hidden follow-up turns.
- Reported compaction recovery continuations from `#checkCompaction` and skipped stop hooks while internal recovery owns the next turn.
- Let terminal empty-stop retry caps fall through to `session_stop` and added regression coverage for aborts, empty-stop caps, and promotion recovery.
2026-06-17 12:26:52 +02:00
ben c40bc8365e fix(coding-agent): honor session stop reason fallback 2026-06-17 12:26:52 +02:00
ben bd14fed678 fix(coding-agent): finalize session stop lifecycle 2026-06-17 12:26:52 +02:00
ben c93774f892 fix(coding-agent): implement session stop hook semantics 2026-06-17 12:26:52 +02:00
Ogrodev 0123a46f83 Merge remote-tracking branch 'upstream/main' into feat/profiles-and-alias
# Conflicts:
#	packages/coding-agent/src/cli/args.ts
2026-06-14 19:10:31 -03:00
can1357 24c8bb24c6 feat(session): added modular session APIs and rebuilt listing/persistence behavior
- Added session-domain modules and exports for session-entries, context, listing, loader, and migrations.
- Changed persistence to async append writes plus writeTextAtomic, removing sync line APIs.
- Added compaction-aware session context rebuild with dangling tool-call cleanup.
- Added resumable session resolution with status inference, id/stem/suffix matching, and backup recovery.
2026-06-14 02:02:53 +02:00
Ogrodev 6450d3b46a Merge remote-tracking branch 'upstream/main' into feat/profiles-and-alias
# Conflicts:
#	packages/coding-agent/src/cli/args.ts
2026-06-13 18:05:16 -03:00
can1357 705750453d fix(coding-agent-turn-interrupt/queue-ux): resolved steering abort state
- Replaced queued-message interrupt flow with session abort calls on empty submit and escape.
- Removed interrupting state and notifyInterrupting teardown paths from abort handling.
- Updated AgentSession queue operations to use shared steering and follow-up queue views.
- Propagated isAborting through session state and collab payloads to suppress late updates.
2026-06-13 17:31:25 +02:00
can1357 39efdfe91b fix(coding-agent): coalesced queued steer flushes to avoid AgentBusyError interrupt races
- Coalesced repeated interrupt-and-flush calls into one queued-steer resume flow.
- Retried queued continues on AgentBusyError after waitForIdle up to a 30s timeout.
- Handled queue-flush failures in input controllers with warning logs and TUI error display.
2026-06-13 15:52:25 +02:00
can1357 64aa558e62 chore: consistency 2026-06-13 00:03:27 +02:00
can1357 84175ce4b2 fix(agent): preserved queued steering across externally aborted runs
Interrupting mid-tool execution (e.g. Enter with a pending steer) drained the steering queue into the dying run — it landed in history without a response and the post-abort resume saw an empty queue, so the agent stopped instead of continuing. Steering/follow-up/aside queue polls in runLoopBody and the post-tool-call check in executeToolCalls are now skipped once the run's abort signal fires, leaving the queue intact for Agent.continue().
2026-06-10 17:43:51 +02:00
handlecusion 9820a5233f test(coding-agent): stabilize CI expectations 2026-06-10 16:21:43 +09:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 1a9d898b8a feat(coding-agent): added immediate steering flush on empty streaming submit
- Updated InputController streaming submit handling to interrupt and resume when queued steering exists, refreshing the pending-message UI and render.
- Added AgentSession.interruptAndFlushQueuedMessages to abort active work and continue processing queued messages immediately.
- Added tests for queued steering interruption in both agent-session concurrency and input-controller keybinding flows.
2026-06-07 07:18:02 +02:00
can1357 793251fa2e fix(coding-agent): labeled TTSR-aborted tool placeholders with matched rule names
- Passed a formatted TTSR match message into the agent abort call when streaming is interrupted by TTSR rules.
- Added a `#formatTtsrAbortReason` helper to include matched rule names in the abort reason.
- Added a concurrent-session test confirming aborted tool results mention the matched TTSR rule instead of `Request was aborted`.
2026-06-07 06:56:16 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 001a6ad564 tests: remove useless assertations 2026-06-06 16:00:31 +02:00
can1357 0340604f64 Merge remote-tracking branch 'origin/farm/9c8b047b/fix-async-job-manager-singleton-overwrite' 2026-06-05 11:41:15 +02:00
roboomp ac304eacdc fix(sdk): scoped async job snapshots to sessions
AgentSession now stores the same scoped AsyncJobManager reference that tools receive: owning top-level sessions use their constructed manager, subagents inherit the parent's manager, and secondary in-process top-level sessions get no manager when a singleton is already live.

getAsyncJobSnapshot and ACP delivery drains now use that scoped manager instead of AsyncJobManager.instance(), so secondary sessions cannot report or drain the primary session's background jobs. The regression test covers a secondary session created while the primary has a Main-owned running job.
2026-06-05 09:29:30 +00:00
can1357 0f83efdf78 fix(coding-agent): relativized rule path in TTSR injections
- Stopped leaking absolute home directory to the model in ttsr-interrupt and ttsr-tool-reminder blocks.
- Rendered rule paths as cwd-relative in-project, `~`-relative under home, else raw.
2026-06-05 10:53:01 +02:00
roboomp 1217091557 fix(agent): prevented subagent session_start busy race
Queued extension-delivered user messages when deliverAs is set and waited for session_start extension message sends before prompting subagents.

Fixes #1343
2026-05-25 18:48:06 +00:00
can1357 bfae4d46c3 fix(coding-agent): tracked acp tool args by session for replay
- Tracked ACP tool-call inputs per session and replayed them via `toolArgsById`/`getToolArgs` plumbing.
- Merged ACP tool execution end content from start and result events so command output replay preserves original args.
- Scoped ACP async-job draining by session `ownerId` and `agentId` with in-flight tracking and permission-gated deferred turns.
- Refactored compaction telemetry and async tests with per-test telemetry setup and asynchronous teardown resets.
2026-05-17 13:15:07 +02:00
jiwangyihao 6e53286ce7 fix(coding-agent): scope async delivery drains 2026-05-17 15:40:04 +08:00
jiwangyihao 05e1e702dd fix(coding-agent): keep ACP async continuations owned 2026-05-17 15:30:09 +08:00
can1357 8f204539b8 fix(coding-agent): deferred agent_end emission until prompt unwinds
- Held wire-level agent_end until #promptInFlightCount drops to 0, preventing AgentBusyError when subscribers fire the next prompt synchronously from agent_end.
- Added #pendingAgentEndEmit field and #flushPendingAgentEnd(), called from #endInFlight and #resetInFlight.
- Added regression test covering re-entrant prompt() from agent_end listener.
2026-05-17 08:53:41 +02:00
can1357 90b134ca4c test: replaced real timers and sleeps with deterministic test hooks
- Added `providerRetryWait` and `retryWait` hooks to stream/usage options so tests bypass real scheduler delays.
- Parameterized GitHub Copilot poll intervals and Copilot model retry base delay for fast test execution.
- Replaced `Bun.sleep`/`setTimeout` polling loops with `AbortSignal` event listeners in agent session tests.
- Consolidated auth-gateway E2E helpers into a shared `test/helpers` module, eliminating duplicated `checkGatewayAvailable` implementations.
- Migrated credential-disabled tests from SQLite-backed stores to an in-memory store, removing temp-dir lifecycle overhead.
2026-05-17 04:02:09 +02:00
can1357 7ea9e16408 feat(coding-agent): changed TTSR non-interrupt tool matches to fold into toolResult
- Non-interrupting tool-source TTSR matches now prepend a system-reminder to the matched tool's `toolResult` content instead of queuing a loop-wide deferred follow-up turn.
- Text/thinking source matches retain the previous deferred-injection behavior.
- Added deduplication so one rule attaches to exactly one sibling tool call per batch.
- Stale per-tool injections are cleared on abort/error before tools produce results.
2026-05-16 20:15:53 +02:00
can1357 2867e1f4e3 feat(deps): added pi.zod exports and removed TypeBox package exports
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
2026-05-15 14:46:54 +02:00
can1357 1e601b9094 test(agent): replaced agent stream mocks with createMockModel responses
- Replaced custom MockAssistantStream helpers with createMockModel streams across agent tests.
- Removed manual queueMicrotask stream-event scripting in favor of scripted mock responses.
- Consolidated helper fixtures by deleting local aliases and reusing shared user-message/model helpers.
- Updated test assertions to use mock.calls and mock.model metadata for call and context validation.
2026-05-15 14:46:54 +02:00
can1357 8c323666be feat: added ordered systemPrompt arrays and normalized context prompts
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
2026-05-04 15:20:26 +02:00
can1357 a815a25afe fix(coding-agent): reject empty write & align promotion tests to gpt-5.5
- Add explicit validation rejecting write:"" with the expected error message in normalizeChunkEditOperations.
- Update spark context-promotion tests (it_1, it_2, concurrent it_5) to expect promotion to gpt-5.5 (the new chain target on openai-codex), since spark variants now promote to gpt-5.5 rather than the base codex model.
2026-04-25 17:52:33 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 2c93655796 feat(autoresearch): added auto-resume, path validation, and security guards
- Added auto-resume mechanism with state tracking to automatically resume pending experiment runs and prevent duplicate resumptions.
- Added contract path validation to reject unsafe path specifications with absolute paths and parent directory traversal attempts.
- Added secondary metrics input to autoresearch setup flow for specifying tradeoff metrics alongside primary objectives.
- Enhanced command parsing with shell operator detection to reject piped, redirected, or chained autoresearch.sh commands.
- Added prototype pollution guards in object cloning functions to prevent injection via __proto__, constructor, and prototype keys.
- Fixed boundary duplication warnings in hashline detection to properly report multiple overlapping hashline references.
2026-03-23 02:16:39 +01:00
can1357 003f46f42c feat(coding-agent/autoresearch): added contract validation and run tracking system
- Added contract system for validating benchmark commands, metrics, scope paths, constraints, and off-limits paths.
 - Contract validation enforces matching initialization parameters against autoresearch.md before init_experiment.
 - Segment fingerprinting detects configuration drift and warns when metrics are not directly comparable.
 - Added pending run detection and recovery to resume incomplete experiments from .autoresearch/runs/.
 - Run directories organize artifacts with benchmark logs and optional checks logs for traceability.
 - Extended experiment state to track run number, command, scope, off-limits, constraints, and fingerprint.
2026-03-23 01:49:03 +01:00
can1357 2f151fea9a fix(tests): added resource cleanup methods and initiatorOverride support
- Added `close()` method to SessionManager and AuthStorage for proper resource cleanup and finalization of prepared statements.
- Added `initiatorOverride` option support in OpenAI and Anthropic providers for message attribution control.
- Fixed resource leaks in RpcClient timeout handling by centralizing timeout creation with unref() and adding explicit clearTimeout() calls.
- Fixed AgentSession disposal to call SessionManager's `close()` method for guaranteed resource cleanup instead of fallback flush.
- Updated all test suites to properly dispose AuthStorage instances in cleanup hooks to prevent resource leaks between tests.
2026-03-14 11:25:40 +01:00
can1357 967d8cc05b fix(coding-agent): changed session.dispose() to await 2026-03-14 11:03:55 +01:00
can1357 17181b2497 feat(coding-agent): added deterministic session APIs and unified recovery orchestration
- Added public `waitForIdle()` and `getLastAssistantMessage()` APIs to AgentSession for deterministic session state access.
- Refactored deferred continuation scheduling from raw `setTimeout()` to centralized post-prompt task tracking system for concurrent recovery operations.
- Fixed race conditions between deferred TTSR/context-promotion continuations and `prompt()` completion via shared recovery orchestrator.
- Replaced `#waitForRetry()` with `#waitForPostPromptRecovery()` to unify retry and TTSR resume gate handling.
2026-02-28 05:27:45 +01:00
can1357 b0dc5f6869 fix(coding-agent): corrected TTSR violations aborting subagent runs
- Fixed TTSR violations during subagent execution aborting the entire subagent run; `#waitForPostPromptRecovery()` now awaits agent idle after TTSR/retry gates resolve, preventing `prompt()` from returning while fire-and-forget `agent.continue()` is still streaming.
- Added comprehensive test case verifying `prompt()` blocks until TTSR continuation with tool calls completes, preventing premature session disposal.
2026-02-28 05:27:45 +01:00
can1357 c0af47d841 fix(coding-agent): corrected TTSR resume gate to prevent prompt race conditions
- Implemented TTSR resume gate to ensure `prompt()` blocks until TTSR interrupt continuations complete, preventing race conditions between TTSR injections and subsequent prompts.
- Replaced `#waitForRetry()` with `#waitForPostPromptRecovery()` to handle both retry and TTSR resume gates, ensuring prompt completion waits for all post-prompt recovery operations.
- Added comprehensive test coverage for TTSR resume gate behavior under interrupt and deferred injection modes.
2026-02-28 05:27:44 +01:00
can1357 a1efc5f4e1 refactor(coding-agent): simplified error handling and import organization
- Reorganized import statements in openai-compat.ts for consistency.
- Consolidated multi-line ternary expression into single line in openai-compat.ts.
- Simplified AgentBusyError instantiation to use default message.
- Updated test assertion to check error type instead of message content.
2026-02-18 17:47:59 +01:00
can1357 bc69fd207d feat: implemented dynamic model resolution across all providers with ModelManager API
- Added ModelManager API with createModelManager() factory for managing bundled and dynamically discovered models with configurable refresh strategies.
- Exported discovery utilities for fetching models from Antigravity, Codex, Cursor, Gemini, and OpenAI-compatible endpoints with provider-specific model manager configuration helpers.
- Renamed public API functions for clarity: getModel() -> getBundledModel(), getModels() -> getBundledModels(), getProviders() -> getBundledProviders().
- Added on-disk model caching with TTL-based invalidation and resolveProviderModels() function for runtime model resolution with source precedence.
- Refactored model discovery script to dynamically fetch models from Codex, Cursor, and Antigravity using OAuth credentials instead of hardcoded lists.
2026-02-18 01:38:05 +01:00
can1357 f8950c22b4 refactor(auth-storage): cleanup auth.json mentions
- Migrated authentication storage from JSON-based (auth.json) to database-based (agent.db) format across test files and configuration.
- Simplified auth storage discovery in sdk.ts by removing manual path construction and fallback logic in favor of centralized getAgentDbPath() function.
- Removed dbPath instance property from AuthStorage class as database path is now managed centrally.
- Updated environment variable precedence documentation to reflect agent.db instead of auth.json as the lowest priority source.
- Removed OAuth provider section header comments from multiple test files for cleaner test organization.
- Added support for JSON and JSONC configuration file formats without requiring migration to YAML.
2026-02-05 14:30:22 +01:00