Commit Graph

80 Commits

Author SHA1 Message Date
can1357 cf60e6df51 feat(coding-agent): implemented eval framework and replaced python tool
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
2026-04-30 18:08:37 +02:00
can1357 712f69b823 feat(coding-agent): enabled subagents to share parent local:// protocol options
- Added a localProtocolOptions field to session and executor option types for configurable local:// behavior.
- Passed localProtocolOptions through to LocalProtocolHandler creation and subagent session bootstrap.
- Propagated parent session local protocol settings from TaskTool so subprocess subagents share the same local:// artifacts and session context.
2026-04-29 04:36:11 +02:00
can1357 fe25549d1f refactor(coding-agent): updated hashline/grep anchors to * and |-separated output
- Updated `formatMatchLine` to emit `*` for matched lines, a leading space for context, and a `|` anchor/content separator.
- Revised grep/hashline mismatch messages and prompts to describe the new marker and separator format.
- Aligned affected atom and hashline tests with the updated match-line prefixes and separators.
2026-04-26 10:41:07 +02:00
can1357 413e517c5e feat(coding-agent): added AgentRegistry for IRC session peer lookup
- Added `AgentRegistry` singleton with session registration/unregistration and IRC routing metadata for peer lookups.
- Added IRC messaging prompts and tooling with `irc.enabled` setting, `list/send` tool paths, and peer roster rendering.
- Changed `/btw` to session-side `runEphemeralTurn`, added background IRC exchange flushing, and fixed empty-input checks.
- Added unit tests for IRC tool and BtwController ephemeral behavior, including disabled, busy, not-found, and abort cases.
2026-04-26 10:28:22 +02:00
can1357 ecd1554eba feat: renamed subagent handoff flow to use yield instead of submit_result
- Renamed subagent completion flow from `submit_result` to `yield` across SDK tools, prompts, and docs.
- Updated executor/task handling to require and parse `yield` calls, replacing legacy submit-result extraction and state flags.
- Added `subagent-yield-reminder` and updated system prompts to require `yield` with `result.data` or `result.error`.
- Renamed hidden-tool and registration plumbing to `yield`, including discovery helpers and renderer/test surface.
2026-04-26 00:29:52 +02:00
Can Bölük 77bf79e79a Merge pull request #509 from apoc/fix/mcp-manager-propagation
fix(mcp): propagate mcpManager to nested subagent sessions
2026-04-24 07:37:44 +02:00
can1357 a9ca2daaed fix(coding-agent): fail structured subagents without submit_result
Fixes #729
2026-04-24 01:02:18 +02:00
Miroslav Drbal bf313a2acf fix(mcp): propagate mcpManager to nested subagent sessions
Subagents created with enableMCP=false had toolSession.mcpManager
unset, causing depth-2+ sub-subagents to re-discover and spawn
duplicate MCP server processes.

- Add mcpManager option to CreateAgentSessionOptions
- Set toolSession.mcpManager unconditionally after MCP block
- Pass options.mcpManager from executor to createAgentSession
- Guard callback registration: only wire onToolsChanged/onPromptsChanged/
  onResourcesChanged when the session owns the manager (created via
  discovery), not when reusing a parent's — prevents child sessions
  from clobbering the parent's live MCP refresh handlers

Latent since 91da560cc (in-process subagent migration), observable
since f82a5d121 added task.maxRecursionDepth allowing depth-2 agents.
2026-04-24 00:46:56 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 5caddcebd6 feat(cross-cutting): added fd crate export and moved fuzzy-find bindings
- Removed `SearchDb` APIs and `searchDb` fields, dropping db-backed state from native and agent sessions.
- Replaced crate export `fff` with `fd`, moving fuzzy-find bindings into `fd.rs`.
- Removed `SearchDb`/picker fast-path logic from `glob` and `grep`, simplifying scan flow and dropping db args.
- Removed `SearchDb`/`getSearchDb` wiring from extension, tool, and task context constructors across coding-agent.
- Added over-indentation validation warnings in chunk-edit normalization for suspicious `~` body line formatting.
- Removed `bytes`, `fff-grep`, `fff-search`, and `blake3` deps, adding `grep-searcher = "0.1"`.
2026-04-13 21:25:50 +02:00
can1357 4bc79b9a67 fix: fixed session naming context and native import paths
- Passed the user context when setting the session name during subprocess execution.
- Updated native build scripts to import detectHostAvx2Support from the correct shared module paths.
2026-04-13 01:12:15 +02:00
can1357 78c888ec16 merge: PR #687 2026-04-13 00:31:39 +02:00
can1357 23a5afe8d9 feat(coding-agent): added auto-backgrounding for long bash jobs to avoid blocking
- Added bash.autoBackground.enabled and bash.autoBackground.thresholdMs settings with defaults for background job behavior.
- Added auto-backgrounding support for long foreground bash commands with managed job state and timeout threshold.
- Updated bash prompts, job-protocol messages, and tool activation checks to use async and auto-background support.
- Added background bash completion integration with AsyncJobManager and end-to-end tests for short/long auto-background scenarios.
2026-04-11 11:05:16 +02:00
djdembeck 72e9e20441 feat: add session name getter/setter to extension API
- Document session name getter/setter methods in extensions.md
- Add stub methods to ExtensionProxy that throw if called before init
- Add delegating implementations to ExtensionProxyInit
- Wire up getSessionName and setSessionName in extension runtime
- Add method signatures to ExtensionContext interface and type
- Add getSessionName to session manager API
- Add getSessionName and setSessionName to all mode contexts:
  - ACP agent
  - Extension UI controller (also updates terminal title)
  - Print mode
  - RPC mode
2026-04-11 01:09:08 -05:00
can1357 a21a542afd refactor(prompt-templates): migrated prompt utilities to pi-utils package
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).
2026-04-08 05:47:35 +02:00
Patrik Sundberg 94d3636bc5 fix(coding-agent): address session observer PR review feedback
- Move lifecycle end event after finalizeSubprocessOutput() so
  submit_result({status:'aborted'}) is correctly reflected in the
  observer instead of showing completed/failed (P1)
- Reset observer registry on session resume and /new command so
  stale subagent sessions from prior conversations are cleared (P2)
- Cache parsed JSONL transcript incrementally: track byte offset,
  only read and parse new bytes on each refresh instead of reparsing
  the entire file, avoiding render loop blocking on large sessions (P2)
2026-04-04 15:53:31 +02:00
Patrik Sundberg 4d641a52b4 feat(coding-agent): add session observer (Ctrl+S) to view subagent sessions
Wire EventBus through task runtime so subagent lifecycle and progress
events propagate to the TUI. Add SessionObserverRegistry to track
active sessions, and a session observer overlay accessible via Ctrl+S
that shows a picker of running subagents and a read-only transcript
viewer that reads the subagent's session JSONL file to display
thinking, text, tool calls, and results.

- Add TASK_SUBAGENT_LIFECYCLE_CHANNEL for start/end events
- Add SubagentProgressPayload.sessionFile for session file tracking
- Pass eventBus through sdk.ts -> main.ts -> InteractiveMode
- SessionObserverRegistry with multi-listener onChange pattern
- Overlay picker preserves selection position across live refreshes
- Viewer renders full transcript from session JSONL (thinking, text,
  tool calls with smart arg summaries, inline tool results)
2026-04-04 15:53:31 +02:00
can1357 49cb500175 feat(pi-natives): added unified picker coordination for file search operations
- Added `wait_for_picker_scan()` function to search_db module with cancellation token support for polling picker scan completion.
- Integrated picker-based file search into glob matching logic with `collect_files_from_picker()` helper to reuse shared SearchDb picker results.
- Refactored fff and grep modules to use centralized `wait_for_picker_scan()` wrapper instead of direct FilePicker calls, improving cancellation handling.
- Propagated SearchDb instance through agent session initialization and input controller to enable unified file picker coordination across search operations.
2026-03-27 11:40:50 +01:00
can1357 dbc1c4af1d feat(coding-agent): added attribution option to control billing and initiator tracking
- Added `attribution` option to `PromptOptions` for controlling billing and initiator attribution.
- Updated subagent prompts to explicitly set `attribution: "agent"` for accurate billing attribution.
- Refactored message attribution logic to use configurable `promptAttribution` with fallback defaults.
- Added test coverage for `attribution` option in subagent reminder prompts.
Fixes #439.

feat(coding-agent): added attribution option and explicit session directory control

- Added `attribution` option to `PromptOptions` for explicit billing and initiator attribution control.
- Updated `SessionManager.create()` to require both `cwd` and `sessionDir` parameters for explicit session directory control.
- Changed session directory naming for temporary working directories from `--` format to `-tmp-` prefix.
- Made `cwd` and `sessionDir` fields mutable in SessionManager to support session relocation.
- Fixed automatic migration of legacy session directories to new `-tmp-` prefixed naming scheme.
- Updated all test fixtures to pass both `cwd` and `sessionDir` parameters to `SessionManager.create()`.
2026-03-15 22:56:18 +01:00
can1357 09c0d28193 fix(coding-agent): corrected boolean coercion in fetch and executor modules
- Fixed boolean type coercion in fetch and executor modules by wrapping truncation flags with Boolean() cast.
- Removed maxBytes property from truncation metadata to simplify output metadata structure.
- Normalized optional result properties with explicit fallbacks in output-meta module.
- Updated test expectations to reflect undefined truncation properties instead of false/null values.
2026-03-13 15:05:11 +01:00
can1357 a1be87ad9b feat(coding-agent/task): added optional assignment field to track raw task text separately from templates
- Added optional `assignment` field to task result and progress interfaces to track raw per-task assignment text separately from full templated task.
- Updated task rendering to display assignment text instead of full task template when available, improving clarity of task display.
- Modified task section rendering to show trimmed assignment text with fallback to task field if assignment is not available.
- Propagated assignment field through task executor, template renderer, and result objects to maintain consistency across task processing pipeline.
2026-03-11 01:11:20 +01:00
can1357 bb0026cb32 feat(coding-agent): added eager todo configuration and per-turn tool choice overrides
- Added 'todo.eager' configuration setting to automatically create a comprehensive todo list after the first user message.
- Added 'buildNamedToolChoice' utility function to build provider-aware tool choice constraints for named tools.
- Modified tool choice resolution to support per-turn tool choice overrides via consumeNextToolChoiceOverride() method.
- Implemented eager todo enforcement mechanism that injects a synthetic prompt to encourage todo creation when conditions are met.
- Extracted tool choice building logic into reusable utility module for better code organization.
- Added comprehensive test coverage for eager todo enforcement functionality in AgentSession.
2026-03-11 00:54:55 +01:00
can1357 8e3e0ebf9e feat: introduced Effort enum and ThinkingConfig for model-aware reasoning
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
2026-03-06 12:35:01 +01:00
can1357 10242a445a refactor(ai): renamed reasoningEffort to reasoning across providers
- Updated option interfaces, param builders, and stream functions to use unified `reasoning` field.
- Added `resolveOpenAiReasoningEffort()` to centralize xhigh clamping logic.
- Replaced type casts with `castApi()` helper and fixed test option names.
- Updated coding-agent and benchmark references for consistency.
2026-03-05 00:25:12 +01:00
can1357 e1897ce013 refactor: migrated thinking configuration to centralized pi-ai module
- Extracted thinking module with ThinkingEffort, ThinkingLevel, and ThinkingMode types to centralize reasoning configuration across packages.
- Migrated ThinkingLevel type from pi-agent-core to pi-ai package with new validation functions parseThinkingLevel() and getAvailableThinkingLevel().
- Consolidated thinking level constants and descriptions into reusable exports (ALL_THINKING_LEVELS, THINKING_MODE_DESCRIPTIONS) for consistent UI display.
- Removed local thinking-effort-label utility and replaced formatThinkingEffortLabel() with centralized formatThinking() function from pi-ai.
- Refactored thinking mode handling to distinguish ThinkingSelector (user-facing with 'off' option) from ThinkingEffort (provider-level).
2026-03-05 00:03:40 +01:00
can1357 1427e93183 Merge pr-223: thinking suffix per-role overrides 2026-03-04 23:12:49 +01:00
Miroslav Drbal [ApoC] a0fe672ca2 fix: dereference MCP tool schema $ref/$defs and suppress Ajv format warnings (#276)
- Add dereferenceJsonSchema() that inlines local $ref pointers and strips
  $defs/definitions from MCP tool schemas before they reach LLM providers.
  Previously, Anthropic's convertTools() extracted only properties/required,
  dropping $defs and leaving dangling $ref — the LLM never saw the actual
  type definitions (e.g. SourceAnchorInput enum values from nucleus).

- Silence Ajv logger (logger: false) on all three instances that use
  strict: false. MCP servers may declare non-standard format keywords
  (e.g. "uint") that caused console.warn() to corrupt TUI output.

- Cache compiled Ajv validators per schema object identity in validation.ts,
  eliminating redundant recompilation on every tool call.

Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
2026-03-03 22:52:41 +01:00
maximhar 8b7893d042 feat(coding-agent): add per-role thinking specs and inline badge effort display 2026-03-03 06:12:51 +01:00
can1357 a88dd79f9c feat(task): surface explicit abort reasons for subagent results
- Add a dedicated abortReason field to SingleResult and thread it through task execution paths so aborted subagents show actionable context instead of a generic badge.
- Populate abortReason for signal cancellation, pre-start cancellation, submit_result aborted status, and missing submit_result after reminders

Fixes #248
2026-03-03 04:14:04 +01:00
can1357 17181b2497 feat(coding-agent): added deterministic session APIs and unified recovery orchestration
- Added public `waitForIdle()` and `getLastAssistantMessage()` APIs to AgentSession for deterministic session state access.
- Refactored deferred continuation scheduling from raw `setTimeout()` to centralized post-prompt task tracking system for concurrent recovery operations.
- Fixed race conditions between deferred TTSR/context-promotion continuations and `prompt()` completion via shared recovery orchestrator.
- Replaced `#waitForRetry()` with `#waitForPostPromptRecovery()` to unify retry and TTSR resume gate handling.
2026-02-28 05:27:45 +01:00
can1357 2e20d87151 feat(coding-agent): simplified skill management by removing per-task pinning
- Removed per-task skill pinning feature; subagents now inherit session skill set instead of per-task selection.
- Removed `preloadedSkills` option from `CreateAgentSessionOptions` interface and related system prompt plumbing.
- Removed `skills` field from Task schema; task execution now passes available session skills directly to subagents.
- Simplified task execution pipeline by removing skill resolution logic and preloaded skills handling from system prompt templates.
2026-02-26 20:47:54 +01:00
can1357 476b858b3a feat(coding-agent): added async background job execution with configurable concurrency limits
- Added async background job execution for bash and task tools with configurable concurrency limits and automatic result delivery.
- Added cancel_job tool and /jobs slash command to manage and inspect running background jobs with status display.
- Added jobs:// internal protocol handler for querying job status and retrieving job execution details.
- Added async.enabled and async.maxJobs settings to control background job execution behavior.
- Enhanced status line to display count of running background jobs with visual indicator.
- Implemented AsyncJobManager with exponential backoff retry delivery, job lifecycle tracking, and automatic eviction.

Fixes #56.
2026-02-22 12:44:57 +01:00
can1357 c199779a05 fix(task): corrected submit_result to terminate only on success
- Fixed submit_result tool to only terminate on successful execution instead of always terminating.
- Removed deferred termination logic and simplified abort behavior to call requestAbort immediately.
- Added submitResultCalled flag tracking to properly manage tool execution state.
- Added test coverage for submit_result tool retry behavior after execution errors.
2026-02-20 20:03:37 +01:00
can1357 a15c39ae7d feat(coding-agent/task): added subprocess output finalization with submit_result validation
- Exported `finalizeSubprocessOutput()` function and `SubmitResultItem` interface for subprocess output finalization with submit_result validation.
- Added automatic reminders (up to 3) when subagent stops without calling submit_result tool, aborting with exit code 1 after final reminder.
- Extracted subprocess output finalization logic into dedicated `finalizeSubprocessOutput()` function for improved testability and reusability.
- Added comprehensive test coverage for subagent warning injection, reminder behavior, and abort handling in executor module.
2026-02-20 17:15:02 +01:00
can1357 83e914d075 fix(submit-result): corrected submit_result validation to prevent false completion flags
- Added validation to submit_result tool to ensure status field is present and correctly typed as 'success' or 'aborted'.
- Fixed executor to only mark submitResultCalled when submit_result tool succeeds or aborts, preventing false positives on validation failures.
- Added type guards and error state checks to prevent setting completion flags on malformed tool execution results.
- Added comprehensive test coverage for submit_result extraction with valid and malformed payload validation.
2026-02-20 15:23:33 +01:00
can1357 b09ded42a0 feat(coding-agent): render traced intent in task progress/output 2026-02-20 11:05:19 +01:00
can1357 918bf9b043 style: formatting 2026-02-15 09:44:56 +01:00
can1357 8526491221 fix(task): filtered todo_write from active tools instead of full registry at subagent startup 2026-02-15 09:23:04 +01:00
can1357 3b890dbc9a fix(task): filtered parent-owned tools from subagent setActiveTools handler 2026-02-15 09:18:07 +01:00
Chris Watson 315aa32115 fix(coding-agent): improve todo progress tracking reliability 2026-02-15 09:13:56 +01:00
can1357 733a46c226 fix(coding-agent): discovered ollama models at runtime 2026-02-13 12:56:22 +01:00
can1357 9f94d39789 feat(coding-agent/mcp): added abort signal support to MCP request handling for cancellation
- Added MCPRequestOptions interface with signal property for request cancellation via AbortSignal.
- Added abort signal support to MCP tool execution enabling request cancellation via Escape-to-interrupt or other abort mechanisms.
- Enhanced MCP request handling with abort signal propagation through HTTP, SSE, and stdio transports with proper cleanup.
- Improved stdio transport request handling to use Promise.withResolvers for cleaner async flow and better abort signal integration.
- Updated HTTP transport to combine operation abort signals with timeout signals using AbortSignal.any() for unified cancellation.
- Modified SSE response parsing to support abort signals and distinguish between timeout and user-initiated cancellation.
2026-02-10 17:19:50 +01:00
can1357 1eaf40b784 feat(coding-agent): added temperature configuration setting for LLM sampling control
- Added `temperature` configuration setting to control LLM sampling temperature with values from 0 (deterministic) to 1 (creative) to -1 (provider default).
- Added temperature option selector in settings UI with preset values: Default, 0, 0.2, 0.5, 0.7, 1.
- Renamed `settingsInstance` parameter to `settings` in `CreateAgentSessionOptions` for consistency.
- Updated all internal references from `settingsInstance` to `settings` throughout SDK and components.
- Integrated temperature setting into agent configuration and selector controller.
2026-02-10 15:25:10 +01:00
can1357 a5d9bbb4ca chore(coding-agent): migrated console logging to structured logger and updated import styles
- Migrated console.error() calls to structured logger.warn() and logger.error() throughout codebase.
- Updated os.tmpdir() import style in browser tool from destructured to namespace import.
2026-02-10 11:26:42 +01:00
can1357 acf8ab5225 style: stylistic changes 2026-02-10 07:39:39 +01:00
can1357 5339369abd style: replaced ASCII ellipsis with Unicode character for improved typography
- Replaced ASCII ellipsis characters (three dots '...') with Unicode ellipsis character ('...') throughout the codebase for improved typography.
- Adjusted string truncation logic to account for single-character Unicode ellipsis instead of three-character ASCII ellipsis, reducing reserved space from 3 to 1 character in truncation calculations.
- Updated truncation offsets in multiple files (session-manager, agent, executor, footer) to preserve 2 additional characters before ellipsis due to more compact Unicode representation.
2026-02-06 02:51:07 +01:00
can1357 a5db81b90b fix(coding-agent/task): fixed task executor to handle null data in submit_result and added schema mismatch warnings
- Fixed task executor to properly handle agents calling `submit_result` with null data by treating it as missing and attempting to extract output from conversation text rather than silently failing.
- Added documentation caution in task.md about schema vs agent mismatch causing null output, with guidance on using `schema` parameter to override built-in schemas.
- Added warning message when subagent calls submit_result with null data to help diagnose schema mismatch issues.
2026-02-06 02:14:53 +01:00
can1357 bfe175fc01 feat(natives): backported fixes from pi-mono (82d7da878..9ce00079) 2026-02-06 01:55:42 +01:00
can1357 881737fec4 refactor(coding-agent): restructured web search and config systems with class-based architecture and type generalization
- Renamed web search types and functions to remove 'Web' prefix for broader applicability (WebSearchProvider -> SearchProviderId, WebSearchResponse -> SearchResponse, WebSearchTool -> SearchTool, etc.).
- Refactored web search provider system from object-based configuration to class-based architecture with abstract SearchProvider base class and concrete provider implementations.
- Refactored ModelRegistry to use direct constructor instantiation instead of discoverModels() helper function, simplifying model discovery pattern.
- Refactored config system to use new ConfigFile class with schema validation, caching, and multi-format support (JSON, JSONC, YAML).
- Simplified Exa API key discovery to check environment variables only, removing .env file reading logic.
2026-02-05 13:56:38 +01:00
can1357 39634776a7 feat(coding-agent): refactored task API to use 'assignment' field and structured context separation
- Refactored task API to use 'assignment' field instead of 'args' for per-task instructions, enabling clearer separation between shared context and task-specific work.
- Introduced structured context/assignment separation pattern with '<swarm_context>' wrapper for template rendering, replacing placeholder-based substitution.
- Changed agent frontmatter field from 'thinkingLevel' to 'thinking-level' (kebab-case) for consistency with YAML conventions.
- Removed 'context' parameter from ExecutorOptions as context is now prepended at template level rather than executor level.
- Removed 'args' field from AgentProgress and SingleResult interfaces, simplifying task result tracking.
- Updated task rendering to display full task text instead of formatted args, improving clarity in progress output.
2026-02-05 07:43:53 +01:00