Commit Graph

191 Commits

Author SHA1 Message Date
roboomp 8b22d6de65 style: bun run fix 2026-05-15 01:07:19 +00:00
roboomp 1b40197b8c fix(discovery): respect disabledProviders in discoverAgents for claude-plugins
discoverAgents() called listClaudePluginRoots() unconditionally, so agents
from Claude Code marketplace plugins appeared in /agents and the Agent
Control Center even when claude-plugins was listed in disabledProviders.

Guard the listClaudePluginRoots() call with isProviderEnabled("claude-plugins"),
returning an empty roots array when the provider is disabled — matching how
filterProviders() handles every other capability's provider set.

Added regression test that verifies both the enabled path (agents visible)
and the disabled path (agents absent) using a real temp-directory plugin
registry fixture.

Fixes #1075
2026-05-15 01:06:59 +00:00
Miroslav Drbal 8c77d2b15a fix(coding-agent): copy cost into async task progress on completion
The async job completion path copied durationMs, tokens, and extractedToolData
from SingleResult back into the AgentProgress object, but missed cost.
Add progress.cost = singleResult?.usage?.cost.total ?? 0 to the same block.
2026-05-13 20:19:53 +02:00
Miroslav Drbal 5cc8095cd1 refactor(coding-agent): extract appendAgentStats helper in render.ts
Eliminates the three-way duplication of toolCount/tokens/cost stat
appending across renderAgentProgress (running), renderAgentProgress
(completed), and renderAgentResult.
2026-05-13 20:07:09 +02:00
Miroslav Drbal 08551b83a5 fix(coding-agent): show cost in final subagent result line
renderAgentResult (rendered after the task tool resolves) was missing the
cost display present in renderAgentProgress. Read cost directly from
result.usage?.cost.total which is already available on SingleResult.
2026-05-13 20:04:55 +02:00
Miroslav Drbal 084488b680 fix(coding-agent): exclude cacheRead from token display, add per-subagent cost
Token counter (token_total status-line segment, subagent progress tree,
session-observer stats line) previously included cacheRead in its cumulative
sum. With Anthropic prompt caching, cacheRead per turn equals the full cached
context, so summing across N turns gives N*context_size -- a session with a 1M
context and 5 turns showed ~5M tokens despite no compaction occurring.

Fix: display shows input + output + cacheWrite per turn. cacheWrite is kept
because each byte is written once; cacheRead re-reads the same context every
turn. Dedicated cache_read/cache_write status-line segments still show cache
activity; billing cost is unaffected.

Also adds per-subagent cost display (dollar amount, statusLineCost color)
accumulated incrementally from message_end events. Hidden when cost is zero
(subscription/OAuth providers). Brings token and cost display in line with
what Claude Code shows per-agent.
2026-05-13 19:26:38 +02:00
can1357 205bb60ab4 fix(coding-agent/task): sanitized review preview and finding titles before rendering
- Flattened newlines and tabs in summary explanations before generating the review preview text, and trimmed the extracted sentence before truncating.
- Sanitized finding titles during render by normalizing tabs/newlines to spaces after stripping priority prefixes.
2026-05-13 05:59:35 +02:00
can1357 a498d45900 feat(task): updated task launch responses with live ids and coordination guidance
- Updated the task tool output to list newly started background jobs by live task id with optional descriptions.
- Extended the async task prompt guidance to distinguish IRC-enabled versus standard coordination and cancellation behavior.
2026-05-13 05:32:31 +02:00
David Marshall 6872a73977 feat(ai,coding-agent): credential_disabled extension event via multi-subscriber AuthStorage
Adds `pi.on("credential_disabled", handler)` so extensions can react to
soft-disabled credentials (e.g. OAuth invalid_grant) without regex-matching
`agent_end` errorMessages.

`AuthStorage.onCredentialDisabled(listener)` returns an unsubscribe function;
multiple listeners fire for every event with per-listener exception isolation
and FIFO buffer-and-replay (cap 32) when none are attached. The constructor
option from #991 stays as sugar for an immediate permanent subscription.

`createAgentSession()` subscribes the per-session extension runner to
`modelRegistry.authStorage` immediately after resolution and unsubscribes on
dispose / startup failure. Events are forwarded via
`ExtensionRunner.emitCredentialDisabled(event)`, which buffers (cap 32,
drop-oldest) until `runner.initialize(...)` runs in the mode controller so
extension handlers see real UI/runtime context, not the constructor no-op
defaults.

Supersedes #997. Builds on #991.

Co-Authored-By: omp <noreply@oh-my-pi.dev>
2026-05-13 02:18:57 +02:00
can1357 1cffb3473a fix: corrected worktree baseline capture with synthetic tree diff
- Added untracked worktree baseline capture via `untrackedPatch` and synthetic tree diffing.
- Added fallback-aware backend ordering by collecting host candidates and retrying alternates on unavailable PAL.
- Hardened overlay mount lifecycle by removing stale overlays and deleting upper/work dirs after unmount.
- Refined ZFS clone deletion to validate ownership and remove dataset+origin only when checks pass.
- Updated ProjFS integration to use extended-info callbacks and symlink metadata reads.
- Updated rcopy path handling to absolutize paths, and added `writeTree` plus combined shell timeout checks.
2026-05-12 14:36:31 +02:00
can1357 d37d48d11d feat(pi-iso): added unified pi-iso backend resolver with auto probe
- Added unified isolation primitives (BackendKind, ProbeResult, IsoError) and resolve fallback selection logic.
- Added Diff, FileChange, and ChangeKind with default_diff choosing git mode or filesystem walk based on repository state.
- Added APFS/btrfs/zfs/reflink/overlayfs/projfs/rcopy/block-clone backends with platform-aware start, stop, and probe.
- Mapped backend operations to canonicalized paths, recursive clone helpers, rollback cleanup, and unavailable error mapping.
- Added unified native exports isoBackend/isoProbe/isoResolve/isoStart/isoStop/isoDiff and removed projfsOverlay APIs.
- Replaced task isolation resolver flow with ensureIsolation/cleanupIsolation and migrated mode handling to auto plus legacy-mode aliases.
2026-05-12 13:39:45 +02:00
Can Bölük 976280a177 Merge branch 'main' into fix/esm-circular-import-tdz 2026-05-12 06:20:20 +02:00
can1357 ec649278bd fix(coding-agent): resolved async job owner filters for scoped cancelAll
- Added ownerId metadata to async jobs and to task/bash progress items from the session agent id.
- Extended async job registration and query methods with optional owner filters, and updated cancelAll to target matching owners.
- Updated session handoff and disposal so subagents inherit the parent manager, top-level sessions own it, and teardown cancels own jobs only.
- Added owner-aware async-job tests using hold/AbortSignal and scoped cancelAll assertions for running versus cancelled jobs.
2026-05-12 05:53:18 +02:00
can1357 4135897e19 feat(coding-agent/prompts): updated task prompts to branch on IRC mode
- Passed the `irc.enabled` session setting into task prompt rendering so templates can branch for IRC mode.
- Updated orchestrator and subagent-task prompts with IRC-specific instruction paths for skip-all checks, coordination, and file-scoped assignments.
2026-05-12 05:51:52 +02:00
can1357 9ed81977f7 feat(coding-agent): shared artifact manager and flat output directory across subagent sessions
- Added parent-to-subagent artifact manager adoption so subagents reuse the parent `ArtifactManager` and write artifacts into a shared directory with shared IDs.
- Passed the shared artifact manager through tool/session context into subagent executor startup and exposed it via `SessionManager` and `ToolSession` for lookup.
- Updated kernel environment and artifact-resolution paths to prefer `PI_ARTIFACTS_DIR`, falling back to existing session-file-based behavior when absent.
2026-05-12 05:17:53 +02:00
can1357 1bde755933 feat(coding-agent): added global singletons for URL protocol handlers
- Added process-wide singleton instances for InternalUrlRouter, AsyncJobManager, and MCPManager.
- Changed internal URL protocols to resolve through registered sessions and scan all active roots/datasets for matches.
- Refactored agent, artifact, memory, rule, skill, jobs, and mcp handlers to use shared manager and rule/skill state.
- Removed per-session protocol/tool wiring and switched tests to initialize and reset global singleton state.
2026-05-12 05:07:52 +02:00
can1357 032e1aa040 fix(coding-agent/task): skipped context file for IRC-enabled subagents
- Skipped writing compact conversation context files when IRC is enabled so subagents avoid using stale markdown snapshots.
- In runSubprocess, passed an undefined context file for IRC-enabled paths and kept context file prompting for non-IRC executions.
2026-05-12 01:22:27 +02:00
David Marshall 1f18e0bf40 fix(coding-agent): break task↔tools ESM cycle causing TDZ at module load
Two-edge fix for a runtime circular-import TDZ that manifested as
`ReferenceError: Cannot access 'TaskTool' / 'SUBAGENT_WARNING_*' /
'MAX_OUTPUT_BYTES' before initialization.` whenever the executor module
graph and the task tool module graph were evaluated together (e.g.
`bun test executor-warnings.test.ts task-simple-mode.test.ts` in either
order, or the wider `task/ discovery/ task-simple-mode` combination).

The runtime cycle is
  task/index.ts → ./executor → ../sdk → ./tools → ../task
closed by `tools/index.ts:285` eagerly dereferencing `TaskTool.create`
while the `task` module's body had not yet reached its `export class
TaskTool` declaration. The throw aborted `tools/index.ts`, which
propagated back through `sdk → executor`, leaving executor's body
suspended before its post-import `const`s (`SUBAGENT_WARNING_*`,
`MAX_OUTPUT_BYTES`) were initialized.

Two structural changes:

1. `tools/index.ts:285`: replace `task: TaskTool.create` with
   `task: s => TaskTool.create(s)`. Defers the `TaskTool` binding
   dereference to factory-call time, by which point the cycle has
   fully unwound. Matches the lazy-factory shape every other
   `BUILTIN_TOOLS` entry already uses.

2. `task/executor.ts:31`: split the `"../tools"` import. Source
   `truncateTail` directly from its leaf module
   `../session/streaming-output` (the barrel was just re-exporting
   it). Keep `ContextFileEntry` as a type-only import — erased at
   runtime, so no participation in the cycle.

Neither change touches tests, test runners, or removes the cycle in
source. They eliminate the eager dereferences that turned a benign
linker-level cycle into a TDZ at evaluation time.

Verified:
- `bun test executor-warnings.test.ts task-simple-mode.test.ts`
  passes in both orderings (11/11).
- Wider `bun test packages/coding-agent/test/task/
  packages/coding-agent/test/discovery/
  packages/coding-agent/test/tools/task-simple-mode.test.ts` now
  100/100 (was 14 fail / 1 error).
- `bun --cwd=packages/coding-agent run check` clean apart from the
  pre-existing unrelated `anthropic.ts:1206 'stop_details'` error
  from 4e0ca3c0e.

Co-Authored-By: omp <noreply@oh-my-pi.dev>
2026-05-11 12:47:59 -05:00
can1357 32e1b41889 feat(coding-agent): added prompt markers to system prompt assembly
- Added explicit prompt markers and wrappers across system templates, including `[env]`, `[role]`, `[coop]`, `[closure]`, and `[now]`.
- Removed `renderTemplate` and `sectionSeparator` flows, deleted `task/template.ts`, and switched to per-task `renderSubagentUserPrompt` rendering.
- Updated system prompt assembly to `shortenPath`-normalize `cwd`, append rendered now metadata, and preserve trailing `[now]` blocks.
- Removed legacy template tests and added prompt-composition tests for ordered `[contract]`->`[project]`->`[now]` blocks and context-only system placement.
- Reworked shared prompt utilities by collapsing consecutive blank lines and removing obsolete `OPENING_HBS`/`LIST_ITEM` helper behavior.
2026-05-10 12:27:45 +02:00
can1357 2e46257e9e feat: added listWorkspace binding and moved AGENTS.md lookup into tree
- Added `listWorkspace` native binding and API types, exporting bounded workspace trees with AGENTS.md candidates.
- Reworked `buildWorkspaceTree` and `buildDirectoryTree` to call `listWorkspace` with 5s timeout defaults.
- Replaced startup AGENTS.md discovery with workspace-tree-only scanning and removed legacy AgentsMdSearch session plumbing.
- Updated `WorkspaceTree` and system prompt context to expose `agentsMdFiles` and aligned tests/changelog expectations.
2026-05-10 07:59:53 +02:00
can1357 7dec6953ac fix(coding-agent/task): skipped refreshing modelRegistry when inherited from parent
- Introduced a flag to detect when an explicit modelRegistry was supplied to runSubprocess.
- Skipped modelRegistry.refresh() when reusing the parent registry and retained refresh when creating a new one.
- Added a debug log to indicate when a parent modelRegistry is reused and refresh is bypassed.
2026-05-10 07:19:26 +02:00
can1357 5d830a5453 fix(task): fixed task renderer handling of missing tasks arrays during streaming
- Updated task call rendering to use `args.tasks?.length ?? 0` when displaying agent counts.
- Prevented runtime warnings in streaming task calls when the `tasks` array was still undefined.
2026-05-10 06:45:02 +02:00
can1357 e17e64004d fix(coding-agent/task): fall back to parent active model when subagent model has no auth
Reporter screenshot showed a parent session on DeepSeek V4 Pro dispatching
a task subagent that resolved to `qwen3.6-plus-free` — an opencode-zen
model the user had no working credentials for. The dispatch hit a
provider that could not serve the model and surfaced a confusing API
rejection instead of using the parent's already-authenticated model.

Adds `resolveModelOverrideWithAuthFallback`, an auth-aware wrapper
around `resolveModelOverride` that checks the resolved subagent model's
credentials via `modelRegistry.getApiKey` + `isAuthenticated` and
falls back to the parent session's active model pattern when the
primary has no working auth. The parent's active model is plumbed
through `ExecutorOptions.parentActiveModelPattern` from `TaskTool`
into `runSubprocess`. If neither has working auth (or they resolve to
the same model), the primary resolution is preserved so the existing
error path still surfaces a meaningful failure downstream.

Fixes #985
2026-05-10 05:17:39 +02:00
can1357 8266ab2697 fix(coding-agent): enforced turn-level tool-call enforcement and final-retry reminder handling
- Updated system prompts to require every active turn to end with a tool call and clarified that reminders should default to resuming work unless the task was complete or genuinely blocked.
- Reworked the yield-reminder text to disallow fake blocker reasons and to permit continued tool-calling instead of forced immediate yields.
- In `runSubprocess`, applied `toolChoice` only on the final yield retry by adding an `isFinalRetry` check before sending the reminder.
2026-05-08 20:40:26 +02:00
can1357 5c35759526 fix(coding-agent): inherit AGENTS.md search and workspace tree from parent in subagents
Subagents previously re-ran buildAgentsMdSearch and buildWorkspaceTree on
every spawn, repeating the slowest part of system-prompt construction for
each task tool invocation. On large/pathological repos those scans
exceeded the 5s preparation deadline and tripped the per-subagent
'system prompt preparation timed out' warning.

Forward the parent's already-resolved AgentsMdSearch and WorkspaceTree
through createAgentSession (alongside the existing contextFiles, skills,
and promptTemplates inheritance):

- Add agentsMdSearch and workspaceTree to CreateAgentSessionOptions;
  createAgentSession short-circuits the parallel scan promises when
  these are provided.
- Resolve them with contextFiles before constructing ToolSession; expose
  on ToolSession so the task tool can read the parent's values.
- Thread them through ExecutorOptions (task/executor.ts) into the
  subagent's createAgentSession call, and pass them from the task tool
  (task/index.ts) on both the worktree-isolated and non-isolated paths.
2026-05-07 23:33:07 +02:00
can1357 77682eac93 feat(coding-agent): added orchestrate command prompt and registered in embedded commands
- Added a new `orchestrate.md` prompt under `src/prompts/commands` defining an orchestrator workflow, verification gates, and anti-patterns.
- Imported and rendered the prompt in `src/task/commands.ts` as `orchestrateMd`.
- Appended the new prompt to `EMBEDDED_COMMANDS` so it is registered with existing embedded command templates.
2026-05-07 04:55:57 +02:00
can1357 2f871b6d23 feat(coding-agent): added loadMode and summary to AgentTool discovery
- Added optional `loadMode` and `summary` fields to `AgentTool` and related type declarations.
- Added `loadMode` and `summary` metadata to built-in tool classes for discoverable/essential behavior.
- Replaced `BUILTIN_TOOL_METADATA` with per-tool fields in discovery code paths.
- Updated `search_tool_bm25` and discovery indexing to use each tool's `summary` text.
- Updated discovery tests to validate tool `loadMode` and summary completeness.
2026-05-06 19:14:39 +02:00
can1357 ba2affae7f fix(coding-agent): restore path alias equivalence
Fixes #935
2026-05-06 17:23:33 +02:00
can1357 d892a16020 fix(coding-agent): avoid ProjFS on Windows ARM emulation
Fixes #949
2026-05-06 17:13:11 +02:00
can1357 8c323666be feat: added ordered systemPrompt arrays and normalized context prompts
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
2026-05-04 15:20:26 +02:00
can1357 0e761ef9e0 feat(coding-agent/hindsight): added per-session HindsightSessionState
- Added `HindsightSessionState` to `AgentSession` and bound hindsight lifecycle hooks to session state.
- Removed global hindsight state/queue handling and replaced it with per-session `HindsightRetainQueue` batching and scoped flushing.
- Reworked recall, reflect, and retain tools to use `session.getHindsightSessionState()` instead of sessionId-based lookup.
- Updated SDK/task/backend/controller flows to pass `session`/`parentHindsightSessionState`, scope `/memory` behavior, and document it in changelog.
2026-05-03 10:27:06 +02:00
can1357 ca84ee760d perf: optimized ai providers with lazy-loading imports and cached init
- Consolidated AI provider imports through register-builtins and moved Gemini/Antigravity header helpers to a shared module.
- Added lazy loading for heavy providers and SDK-backed modules with cached initialization to trim startup cost.
- Converted markdown conversion helpers to async and awaited htmlToBasicMarkdown in affected scraper and kernel output paths.
- Parsed bundled agent definitions on-demand and moved BrowserTool prompt rendering behind a memoized getter.
- Added cached validation/error handling paths by replacing AJV runtime checks with Value.Check and trimming validation error output.
2026-05-01 18:54:21 +02:00
can1357 cf60e6df51 feat(coding-agent): implemented eval framework and replaced python tool
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
2026-04-30 18:08:37 +02:00
can1357 712f69b823 feat(coding-agent): enabled subagents to share parent local:// protocol options
- Added a localProtocolOptions field to session and executor option types for configurable local:// behavior.
- Passed localProtocolOptions through to LocalProtocolHandler creation and subagent session bootstrap.
- Propagated parent session local protocol settings from TaskTool so subprocess subagents share the same local:// artifacts and session context.
2026-04-29 04:36:11 +02:00
can1357 a3f8f122cc feat(coding-agent): renamed grep to search in runtime mappings
- Renamed the built-in `grep` content-search tool to `search` across settings, schemas, and SDK exports.
- Switched execution wiring so `Task`, `Plan`, cursor, and shell mapping now invoke `search` instead of `grep`.
- Updated prompts, plan-mode docs, and example tool lists to replace `grep`/`ls` references with `search` guidance.
- Aligned `Grep*`/`grep` event, renderer, and hook types to `Search*`/`search` across runtime and tests.
- Documented and fixed `search` result rendering budget behavior and added internal-URL/path-list transcript notes.
2026-04-27 21:55:31 +02:00
can1357 fe25549d1f refactor(coding-agent): updated hashline/grep anchors to * and |-separated output
- Updated `formatMatchLine` to emit `*` for matched lines, a leading space for context, and a `|` anchor/content separator.
- Revised grep/hashline mismatch messages and prompts to describe the new marker and separator format.
- Aligned affected atom and hashline tests with the updated match-line prefixes and separators.
2026-04-26 10:41:07 +02:00
can1357 413e517c5e feat(coding-agent): added AgentRegistry for IRC session peer lookup
- Added `AgentRegistry` singleton with session registration/unregistration and IRC routing metadata for peer lookups.
- Added IRC messaging prompts and tooling with `irc.enabled` setting, `list/send` tool paths, and peer roster rendering.
- Changed `/btw` to session-side `runEphemeralTurn`, added background IRC exchange flushing, and fixed empty-input checks.
- Added unit tests for IRC tool and BtwController ephemeral behavior, including disabled, busy, not-found, and abort cases.
2026-04-26 10:28:22 +02:00
can1357 ecd1554eba feat: renamed subagent handoff flow to use yield instead of submit_result
- Renamed subagent completion flow from `submit_result` to `yield` across SDK tools, prompts, and docs.
- Updated executor/task handling to require and parse `yield` calls, replacing legacy submit-result extraction and state flags.
- Added `subagent-yield-reminder` and updated system prompts to require `yield` with `result.data` or `result.error`.
- Renamed hidden-tool and registration plumbing to `yield`, including discovery helpers and renderer/test surface.
2026-04-26 00:29:52 +02:00
can1357 d82377cda8 revert: "read-to-open"
This reverts commit c48d2e6080.
2026-04-24 22:36:44 +02:00
can1357 c48d2e6080 feat(coding-agent): implemented read-to-open tool aliasing in runtime
- Canonicalized file and CLI defaults from `read` to `open` across tool registration and prompts.
- Added `resolveToolAlias()` and applied alias-normalized tool selection so legacy `read` maps to `open`.
- Updated runtime, UI, and export layers to treat `open` as first-class while preserving `read` compatibility.
- Renamed read prompt docs to `open.md`/`open-chunk.md` and refreshed system guidance to recommend `open`.
- Updated tool-related tests and expectations from `read` to `open` (including test fixtures and aliases).
2026-04-24 19:13:11 +02:00
Can Bölük 77bf79e79a Merge pull request #509 from apoc/fix/mcp-manager-propagation
fix(mcp): propagate mcpManager to nested subagent sessions
2026-04-24 07:37:44 +02:00
can1357 a9ca2daaed fix(coding-agent): fail structured subagents without submit_result
Fixes #729
2026-04-24 01:02:18 +02:00
Miroslav Drbal bf313a2acf fix(mcp): propagate mcpManager to nested subagent sessions
Subagents created with enableMCP=false had toolSession.mcpManager
unset, causing depth-2+ sub-subagents to re-discover and spawn
duplicate MCP server processes.

- Add mcpManager option to CreateAgentSessionOptions
- Set toolSession.mcpManager unconditionally after MCP block
- Pass options.mcpManager from executor to createAgentSession
- Guard callback registration: only wire onToolsChanged/onPromptsChanged/
  onResourcesChanged when the session owns the manager (created via
  discovery), not when reusing a parent's — prevents child sessions
  from clobbering the parent's live MCP refresh handlers

Latent since 91da560cc (in-process subagent migration), observable
since f82a5d121 added task.maxRecursionDepth allowing depth-2 agents.
2026-04-24 00:46:56 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 c4ea6c920f fix: restore Copilot prompt budgets and align task/model-registry expectations with opus 4.7
Three fixes to make CI green after the opus 4.7 and auto-bump landed:

1. github-copilot model mapper: prefer capabilities.limits.max_prompt_tokens
   over the root-level context_length field (which mirrors max_context_window_tokens, i.e.
   total window). Copilot's real /models response returns both for the gpt-5.x family, and
   context_length inflates contextWindow with the output budget. Also restore the bundled
   Copilot limits (claude-opus-4.6, gpt-5.2, gpt-5.4, gpt-5.4-mini, grok-code-fast-1) to
   the values the fixed mapper produces so tests that depend on truthful offline fallbacks
   pass. Update the two Copilot discovery tests whose payloads conflated context_length
   with prompt capacity.

2. coding-agent task schema: make the per-task assignment description context-mode-aware.
   The previous description unconditionally told agents that 'shared background belongs
   in context', which is wrong for independent mode where shared context is disabled.

3. coding-agent model-registry test: update the anthropic-latest canonical collapse case
   to claude-opus-4-7 since opus 4.7 is now the newest official opus in models.json.
2026-04-17 19:24:03 +02:00
can1357 8f1da70d79 feat(coding-agent): enabled simple-mode routing in task tool execution
- Added task.simple to settings and schema with default, schema-free, and independent modes.
- Added mode-aware task schema and validation to enforce context/schema rules per simple mode.
- Updated prompts and template rendering to tailor headers and guidance for each simple mode.
- Added simple-mode capabilities and updated TaskTool execution for mode-aware context behavior.
- Added tests for independent rendering and mode-specific rejection of invalid context or schema inputs.
2026-04-15 18:13:47 +02:00
can1357 5caddcebd6 feat(cross-cutting): added fd crate export and moved fuzzy-find bindings
- Removed `SearchDb` APIs and `searchDb` fields, dropping db-backed state from native and agent sessions.
- Replaced crate export `fff` with `fd`, moving fuzzy-find bindings into `fd.rs`.
- Removed `SearchDb`/picker fast-path logic from `glob` and `grep`, simplifying scan flow and dropping db args.
- Removed `SearchDb`/`getSearchDb` wiring from extension, tool, and task context constructors across coding-agent.
- Added over-indentation validation warnings in chunk-edit normalization for suspicious `~` body line formatting.
- Removed `bytes`, `fff-grep`, `fff-search`, and `blake3` deps, adding `grep-searcher = "0.1"`.
2026-04-13 21:25:50 +02:00
can1357 212d56bc11 feat: added strict-mode fallback for OpenAI tool calls with all_strict
- Added `toolStrictMode` support with `all_strict`/`none`/`mixed` options to OpenAI compatibility.
- Fixed OpenAI-completion strict-mode flows by capturing failed HTTP responses and retrying once as non-strict.
- Fixed completion error reporting by surfacing captured status, headers, and JSON `type`/`param`/`code` details.
- Improved strict-schema enforcement with WeakMap memoization and circular-schema detection in sanitization.
- Fixed OpenRouter provider lookup by resolving fallback model IDs for suffix and date variants in registry resolution.
- Refactored benchmark tooling and added async RPC error-window tracking for scheduled run execution.
2026-04-13 15:46:06 +02:00
can1357 c62ab2d53e chore(benchmarks-misc-fixes): cleaned benchmark task/pid validation
- Added `vim` as an edit variant in benchmark CLI/config and rating script coverage.
- Expanded benchmark execution so `vim` is treated as a mutation tool for retries, stats, and edit intent checks.
- Adjusted `TaskTool` output schema precedence so explicit params override agent frontmatter.
- Fixed `TaskTool` success counting by excluding aborted tasks from success totals.
- Improved validation guidance in `SubmitResultTool`/`TodoWriteTool` for clearer recovery when payloads are missing or invalid.
- Added background command PID regression coverage in `executeBash` to confirm a real, terminateable PID is returned.
2026-04-13 12:27:10 +02:00
can1357 4bc79b9a67 fix: fixed session naming context and native import paths
- Passed the user context when setting the session name during subprocess execution.
- Updated native build scripts to import detectHostAvx2Support from the correct shared module paths.
2026-04-13 01:12:15 +02:00