Commit Graph

443 Commits

Author SHA1 Message Date
can1357 316722397a Merge pull request #1896: fix(coding-agent): respect user shell for ! shortcuts 2026-06-10 09:52:28 +02:00
can1357 1821e167f4 fix(coding-agent): deobfuscate secrets in ephemeral turns and obfuscate via provider-context hook 2026-06-10 09:51:47 +02:00
can1357 730e9c8e0c fix(acp): honor the explicit autoApprove session flag when skipping the permission gate 2026-06-10 09:51:47 +02:00
roboomp cafd957f5d fix(providers): disabled ollama thinking for off turns
Propagated explicit thinking-off state through the agent loop so provider requests receive disableReasoning instead of an undefined effort. Added Ollama and agent-session regressions for the :off path.\n\nFixes #2239
2026-06-10 07:41:58 +00:00
handlecusion 2d7d717029 fix(coding-agent): tighten user shell routing 2026-06-10 16:07:59 +09:00
handlecusion 8b9c4fa1b9 fix(coding-agent): respect user shell for shortcuts 2026-06-10 16:07:59 +09:00
can1357 11c53051bf Merge pull request #2097: fix(acp): skip permission gate when yolo mode is explicitly enabled 2026-06-10 08:32:09 +02:00
can1357 d30dca302f Merge pull request #2044: fix(coding-agent): apply enabledModels filter to ACP model list 2026-06-10 08:32:09 +02:00
can1357 3fd09d5770 fix(acp): require explicit yolo opt-in before skipping the client permission gate
The schema default for tools.approvalMode is already "yolo", so checking the
resolved setting alone disabled the ACP permission gate for every
default-config session (17 existing tests in
agent-session-acp-permission.test.ts fail). The skip now requires an
explicitly configured approval mode — the --yolo/--auto-approve runtime
override or a user-set tools.approvalMode — via the new
Settings.isConfigured(), keeping default ACP sessions gated.

Addresses review feedback on #2097.
2026-06-10 08:31:31 +02:00
Theo Mathieu fbb48faf81 fix: reflect --auto-approve flag in settings override so #wrapToolForAcpPermission sees yolo mode 2026-06-10 08:31:30 +02:00
Theo Mathieu c110a624c3 fix(acp): skip permission gate in yolo mode when effective policy is allow 2026-06-10 08:31:30 +02:00
Theo 9e28ec4b7e fix: apply enabledModels filter to ACP model list
getAvailableModels() was calling modelRegistry.getAvailable() directly,
which skips the enabledModels setting. The setting was only applied
during session init to pick the starting model, not to the list
advertised to ACP clients (Zed, etc.).

Add filterAvailableModelsByEnabledPatterns() to model-resolver.ts - a
synchronous subset of resolveAllowedModels() that handles the patterns
used in real configs (exact provider/modelId, canonical ids, bare model
ids, thinking-level suffixes). Glob patterns fall back to showing all
models rather than accidentally emptying the picker.

Update getAvailableModels() to call it, so the ACP model dropdown in
Zed (and any other ACP client) respects the users enabledModels config.
2026-06-10 08:31:30 +02:00
can1357 7ce58fe95c Merge pull request #2204: feat(eval): add python interpreter setting 2026-06-10 08:26:59 +02:00
can1357 62ad6e73ae Merge pull request #2200: fix(ai): rotate antigravity credentials on Individual quota reached 429s 2026-06-10 08:26:59 +02:00
can1357 e19f2f689c Merge pull request #2147: fix(coding-agent): hide secrets in provider requests 2026-06-10 08:26:58 +02:00
Theo Mathieu 416c9947bd fix(acp): finish prompt turn when extension/custom command is handled locally
Extension commands (e.g. /sonnet) and TypeScript custom commands that
consume the input without calling the LLM return early from
session.prompt() with no agent turn. In ACP mode this left the pending
prompt promise unresolved, hanging the client forever.

Change session.prompt() to return Promise<boolean>: true when the LLM
was invoked, false when the command was fully handled locally.
#runPromptOrCommand calls #finishPrompt immediately on a false return so
the ACP turn completes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 08:26:07 +02:00
can1357 ec81b927fa fix(coding-agent): redact ephemeral side-channel context and preserved remote compaction history
runEphemeralTurn (IRC//btw) called streamSimple directly with the raw
system prompt, bypassing the SDK-level obfuscateProviderContext wrapper;
and #obfuscatePreparationForProvider skipped previousPreserveData, so a
pre-fix openaiRemoteCompaction.replacementHistory could resend raw
secrets on the next remote compaction.

Addresses review feedback on #2147.
2026-06-10 08:26:02 +02:00
roboomp d39e4b0d38 fix(coding-agent): hid previous compaction summary
Obfuscated preparation.previousSummary, hook prompt/context, before forwarding to compact() so prior pi- or extension-supplied summaries do not leak verbatim secrets on subsequent compactions.

Fixes #2146
2026-06-10 08:26:02 +02:00
roboomp 5517115f1a fix(coding-agent): hid handoff instructions
Obfuscated custom instructions before handoff and related side-request provider calls, then deobfuscated generated handoff output before persistence.

Fixes #2146
2026-06-10 08:26:02 +02:00
roboomp aa4cd0ab2d fix(coding-agent): preserved tool schemas while redacting
Converted provider-facing tool parameters to wire JSON Schema before redaction so live Zod instances are not deep-cloned into plain objects.

Fixes #2146
2026-06-10 08:26:02 +02:00
roboomp 2a90108893 fix(coding-agent): hid secrets in provider requests
Redacted configured secrets across provider-facing system prompts, tool definitions, developer reminders, and assistant tool-call payloads before LLM requests.

Fixes #2146
2026-06-10 08:26:01 +02:00
can1357 24cbc93913 fix(eval): thread python.interpreter from session settings and expand ~
Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.

Addresses review feedback on #2204.
2026-06-10 08:26:00 +02:00
roboomp 8c3149e5a9 fix(ai): scoped antigravity quota blocks by model family
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.

Fixes #2198
2026-06-10 08:26:00 +02:00
can1357 661587e110 feat(coding-agent): raised retry limits to ten with capped exponential backoff
- Raised Anthropic provider retries to 10 attempts and used a shared jittered exponential backoff for each retry.
- Updated coding-agent retry defaults and session delay calculation to a 500ms base with an 8,000ms jittered cap.
- Added tests that verify capped ten-step backoff sequences and recovery after repeated 502 errors.
2026-06-10 07:49:12 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 7124c74067 fix(ai): waited for sibling unblock over provider retry window
- Returned earliest sibling block expiry from markUsageLimitReached.
- Capped usage-limit retry to whichever frees up first, avoiding multi-hour waits.
- Added 1s buffer so retry lands after the block actually lapses.
2026-06-10 02:44:40 +02:00
Can Bölük 15f597b71d Merge branch 'main' into farm/44203ced/shake-fallback-when-threshold-stays-above 2026-06-08 22:36:14 +02:00
roboomp a190298414 fix(coding-agent): fell back to context-full when shake cannot drop below threshold
Threshold-driven auto-compaction with strategy=shake auto-continued even when
the shake reclaimed nothing material, and the next agent turn re-triggered the
same shake (which had nothing new to drop on a second pass), spinning forever.
After shake completes, recompute the post-shake context estimate; when it
still exceeds the auto-compact threshold (or shake reclaimed nothing on overflow
recovery), emit a one-shot fallback warning and hand off to the summarization
driven context-full path so progress actually resumes. Idle is exempt — its
60s+ timer self-throttles and cannot dead-loop on its own.

Fixes #2119
2026-06-08 20:34:16 +00:00
roboomp ef49122afe fix(coding-agent): honored skill prompt magic keywords
Skill prompt custom messages now scan user-authored skill args for magic keywords and turn budgets, matching normal prompt steering behavior without scanning skill body text.

Added a regression test for workflowz and hard turn-budget args on skill prompts.

Fixes #2128
2026-06-08 20:32:22 +00:00
can1357 9edf773b81 feat(coding-agent): centralized tool filtering for discovery mode with forced tool activation
- Introduced filterInitialToolsForDiscoveryAll() to centralize tool filtering when discovery mode is "all", replacing inline logic in createAgentSession.
- Added forceActive parameter to ensure tools required by forced tool_choice features (e.g., eager todo) remain active in the request, preventing provider 400 errors.
- Updated eager todo enforcement to check active tool set instead of registry, ensuring it respects tool discovery hiding.
2026-06-08 19:08:31 +02:00
can1357 522a7d0f5d fix(ai): scoped Claude device_id to install id and account
- Derived device_id from both install id and account UUID via v2 hash domain.
- Fell back to v1 install-only hash when no account UUID is present.
- Threaded account UUID through Anthropic metadata user_id resolution.
2026-06-08 18:28:07 +02:00
can1357 eae8b932ef fix: fixed incomplete tool-call abort handling and image content guards
- Normalized image-content preprocessing in agent sessions now returns early when `content` is missing or not a string/array, avoiding invalid normalization paths.
- Updated agent-loop tests to verify partial tool calls are dropped when an assistant aborts before `toolcall_end`, with run completion reported via an assistant `stopReason` of `aborted`.
- Adjusted Anthropic alignment expectations to reflect a single trailing cache-control breakpoint on the final system block.
2026-06-08 15:07:26 +02:00
can1357 a84da2a61e feat(coding-agent): added AST-condition matching for interrupt flow rule handling
- Added astCondition to rule frontmatter parsing and rule metadata, with AST-grep normalization.
- Updated TTSR bucketing so astCondition-only rules are treated as interruptible matches.
- Added ts-redundant-clear-guard as a built-in JS/TS tool rule for guarded clear* calls.
- Added AST snapshot matching in agent sessions with per-stream cache throttling and cleanup.
2026-06-08 14:57:17 +02:00
can1357 0890b2be61 fix: fixed tool-call recovery, stream parsing, and image normalization flows
- Fixed tool result validation to reject invalid blocks and append explicit error diagnostics.
- Handled aborted and leaked tool calls by returning abort results and dropping partial leaked output.
- Fixed Anthropic streaming by enforcing strict SSE parsing and content-block lifecycle checks.
- Fixed image handling by normalizing model-context inputs and preserving images on resize failure.
2026-06-08 14:40:18 +02:00
can1357 462c2b749c fix(coding-agent/task): show agent type in task result header
Append the dispatched agent type to the task result frame header so it reads `Task 16 agents: Reviewer` instead of just `Task 16 agents`.
2026-06-08 14:22:49 +02:00
can1357 0753a7a433 feat(agent): removed max tool-call caps from agent loop and session flow
- Removed `maxToolCallsPerTurn` from `AgentOptions`, `AgentLoopConfig`, and config serialization.
- Removed stream-loop cap enforcement, including the `toolcall_end` abort path and capped assistant messages.
- Removed Anthropic Opus 4.8 batch-cap resolver and agent-session sync logic from coding-agent.
- Updated tests and changelogs to align with uncapped tool-call behavior and dropped cap-specific cases.
2026-06-08 14:15:46 +02:00
can1357 28dade85c3 fix(agent): resolved deferred asides at injection to drop stale messages
- Introduced `AsideMessage` as a message-or-thunk union so aside providers can defer injection decisions.
- Updated agent loop handling to resolve aside thunks at injection time and skip entries that returned `null`, then switched the session yield queue to `drainLazy` for deferred message building.
- Added tests validating lazy aside evaluation and staleness-aware dropping when everything becomes stale after dequeueing.
2026-06-08 06:46:41 +02:00
can1357 de03d7f3d1 feat(agent): added non-interrupting aside message support
- Added a new aside-message source on `Agent` and exposed it in `AgentLoopConfig` as `getAsideMessages`.
- Updated the agent loop to poll aside messages after tool batches and before yielding, merging them with follow-up messages before continuing.
- Changed coding-agent session and yield queue handling to pull queued background messages via `drainMessages` at step boundaries instead of streaming-only injection.
2026-06-08 06:08:43 +02:00
can1357 caa9c69309 feat(coding-agent): added provider-priority model selection and refined fallback ordering
- Added first-party-first provider priority defaults for model ranking.
- Consolidated model resolution to use getModelMatchPreferences from session settings.
- Prioritized providerPriorityRank ahead of usage rank when picking preferred models.
- Added second-pass fallback to default-model or API-key-valid matching order.
2026-06-08 01:32:39 +02:00
can1357 248f14cc3b Merge remote-tracking branch 'origin/farm/e318626d/retry-generic-upstream-error' 2026-06-08 00:29:12 +02:00
roboomp 56240e527d fix(agent): retried generic upstream gateway failures
Classified generic upstream gateway failures as transient so configured session auto-retry starts when auth-gateway surfaces upstream_error: Upstream request failed. Added a focused AgentSession regression covering the retry event and recovery path.\n\nFixes #2056
2026-06-07 12:36:58 +00:00
DarkPhilosophy be71d53842 fix(coding-agent): optimize orphaned tool-use stop guard 2026-06-07 15:08:47 +03:00
DarkPhilosophy 4e675f4c5f Merge remote-tracking branch 'can1357/main' into fix/empty-stop-guard-tooluse 2026-06-07 15:04:44 +03:00
can1357 20487d8cb6 chore: fix stale tests 2026-06-07 08:11:38 +02:00
can1357 0914379f49 fix(coding-agent): rendered shared task context as markdown and froze async block borders
- Updated task call and result rendering to process shared context with the Markdown renderer, so context sections are now displayed with proper Markdown formatting.
- Stopped shimmer animation on pending bash/eval/task blocks once async state is `running`, preventing the committed frame from freezing a transient dark border segment.
- Adjusted rule path display to fall back to a root-relative path when cwd-relative resolution is unavailable.
2026-06-07 08:01:52 +02:00
can1357 1a9d898b8a feat(coding-agent): added immediate steering flush on empty streaming submit
- Updated InputController streaming submit handling to interrupt and resume when queued steering exists, refreshing the pending-message UI and render.
- Added AgentSession.interruptAndFlushQueuedMessages to abort active work and continue processing queued messages immediately.
- Added tests for queued steering interruption in both agent-session concurrency and input-controller keybinding flows.
2026-06-07 07:18:02 +02:00
can1357 793251fa2e fix(coding-agent): labeled TTSR-aborted tool placeholders with matched rule names
- Passed a formatted TTSR match message into the agent abort call when streaming is interrupted by TTSR rules.
- Added a `#formatTtsrAbortReason` helper to include matched rule names in the abort reason.
- Added a concurrent-session test confirming aborted tool results mention the matched TTSR rule instead of `Request was aborted`.
2026-06-07 06:56:16 +02:00
can1357 9bd9e3127e feat(coding-agent): added /tan background forking with prompt cache inheritance
- Added `/tan` slash command registration and interactive handling.
- Added TanCommandController validation and async task scheduling for `/tan` dispatch.
- Added session cloning that suppresses breadcrumbs, copies artifacts, and handles abort cleanup.
- Added `promptCacheKey` support in Agent and inherited `providerPromptCacheKey` in session creation.
2026-06-07 06:52:15 +02:00
can1357 a78f4c6766 feat(coding-agent): enabled abort reasons to propagate and surface through streaming messages
- Added optional `reason` parameters to `Agent.abort` and `AgentSession.abort` APIs.
- Passed abort reasons through interrupt flows into underlying agent cancellation.
- Replaced hard-coded abort text with `resolveAbortLabel` for streaming and replayed messages.
- Fell back to generic `Request was aborted` text when no abort reason was provided.
2026-06-07 06:19:55 +02:00
can1357 f1e2e51a4b feat(coding-agent): added /fresh to reset provider state while keeping session files
- Added `AgentSession.freshSession()` to rotate provider-facing IDs and prune provider stream state.
- Added `/fresh` command handling in the builtin registry and mode command flow.
- Kept persisted session metadata intact during `/fresh` and cleared transient IDs on session switches.
- Invalidated `appendOnlyContext` and provider caches when refreshing provider state.
2026-06-07 03:37:48 +02:00