Commit Graph

39 Commits

Author SHA1 Message Date
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
can1357 a872d77068 chore: cleanup dumb tests 2026-08-02 20:39:23 +02:00
can1357 58a2acd330 test(coding-agent): align ctx fixtures with settled-component cache and todo commit-on-execute
- Added transcriptMessageComponents to every InteractiveMode ctx test literal; #6033's reuse cache made the field required and addMessageToChat populates it unconditionally.
- Wired setTodoPhases into the eager-todo ToolSession fixture to mirror sdk.ts; the test relied on the stale message_end todo replay #6148 removed.
2026-07-22 21:32:12 +02:00
roboomp 53a3aef39a fix(subagent): keep title refresh for focusable subagents
Review follow-up: a live subagent focused from the Agent Hub renders its
session name in the status line (session_name segment reads
sessionManager.getSessionName()), so the blanket agentKind === "sub" skip
made the user-enabled title.refreshOnReplan silently ineffective and left
focused subagents untitled after their first todo replan.

Focus only exists in an interactive host, and subagents run in-process, so
gate the skip on a process-global interactive-host flag: subagents skip the
replan title refresh only in non-interactive hosts (print/RPC/ACP/eval/SDK/
CI) where no session tree is focusable. The interactive entrypoint declares
the host via setInteractiveHost(isInteractive); the flag defaults false, so
bun test and headless embedders keep the optimization without leaking state.

Fixes #5910
2026-07-17 21:10:58 +00:00
roboomp b00197b568 fix(subagent): skip session title generation for headless subagents
Subagent sessions run `todo init` per the eager-todo prelude, which
triggered `#scheduleReplanTitleRefresh()` and a tiny-model title
generation call. The result is written to JSONL but never displayed —
subagents surface their registry id and generated task label, not a
session title.

Short-circuit `#scheduleReplanTitleRefresh()` when `#agentKind === "sub"`.
Uses the session-level subagent marker rather than `hasUI` so print/RPC
top-level sessions keep persisting their auto title for `--resume`.

Fixes #5910
2026-07-17 20:42:08 +00:00
roboomp b6b947bdbe fix(openai): rendered native response images
- Normalized completed image_generation_call results into assistant image blocks.
- Persisted image bytes through the session blob store and rendered them in live, replay, ACP, proxy, telemetry, and HTML paths.
- Added response normalization, persistence, and TUI rendering regressions.

Fixes #4768
2026-07-14 16:15:53 +00:00
can1357 04783381b4 feat(coding-agent): transitioned session title generation to xml markers
- Replaced tool-based `set_title` invocation with XML-style `<title>` marker tags for session title discovery.
- Implemented robust JSON-unwrapping logic to handle and sanitize title generation outputs.
- Updated model registry in catalog with new model support, provider prefixes, and metadata adjustments.
- Synchronized system prompt documentation and test suites to reflect the new marker-based generation flow.
2026-07-06 17:38:00 +02:00
roboomp 482091dc73 fix(coding-agent): replan title refresh honors TITLE_SYSTEM.md override
The replan-driven title refresh (title.refreshOnReplan, fired after a
`todo init`) called `generateSessionTitle()` without the user's
`TITLE_SYSTEM.md` override, silently falling back to the bundled
`prompts/system/title-system.md` and overwriting auto titles with the
default policy. The override was only ever discovered by main.ts and
passed into the first-input title path on InteractiveMode, never into
`AgentSession.#refreshTitleAfterReplan`. Most visible in Plan Mode,
which initializes todos early.

`AgentSession` now owns the resolved title prompt:
- New `CreateAgentSessionOptions.titleSystemPrompt` threaded by
  `createAgentSession()` into the constructor.
- New `AgentSessionConfig.titleSystemPrompt` stored on
  `#titleSystemPrompt` with a `get titleSystemPrompt` /
  `setTitleSystemPrompt(...)` pair.
- `#refreshTitleAfterReplan` passes `#titleSystemPrompt` as
  `customSystemPrompt` to `generateSessionTitle()`.
- `input-controller.ts` reads from `session.titleSystemPrompt`, and
  the duplicate `InteractiveMode.titleSystemPrompt` field /
  constructor arg / `InteractiveModeContext` field / `runInteractiveMode`
  parameter are removed. `InteractiveMode.refreshTitleSystemPrompt`
  now calls `session.setTitleSystemPrompt(...)` so a `/move`-style cwd
  change keeps the override in sync.

Regression test asserts the prompt handed to `completeSimple()` from
`#refreshTitleAfterReplan` is the configured override, not the bundled
`title-system.md`.

Fixes #3734
2026-06-28 16:32:15 +00:00
can1357 46278588ea test(session): implemented persistent session management and testing
- Implemented mutable session titles with audit tracking, including storage persistence for SQL and Redis backends.
- Added comprehensive session management features such as idle recap triggers, incremental subagent yield submissions, and automated title refreshing.
- Enhanced task tracking in the TodoTool with progress prioritization and improved session cleanup logic.
- Introduced citation tag handling for OpenAI-compatible source markers and improved edit parsing.
2026-06-28 07:27:01 +02:00
can1357 899c0ef08b feat: simplified todo tool to single operation interface
- Refactored `todo` tool to accept a single operation object instead of an `ops` array.
- Implemented parameter normalization to maintain backward compatibility with legacy array-based tool calls.
- Updated tool instructions, documentation, and UI rendering components to reflect the new interface.
- Added compatibility tests to verify rendering and execution for both legacy and current operation formats.
2026-06-23 00:54:54 +02:00
can1357 f8f8136021 refactor(coding-agent): privatized the legacy nextToolChoice method to
- Privatized the legacy `nextToolChoice` method to `#nextHardToolChoice` to ensure all tool-choice directives flow through the unified `nextToolChoiceDirective` entry point.
- Eliminated redundant dual entry points for fetching tool choices, which previously bypassed the soft pending-preview lifecycle.
- Updated test suites to consume `nextToolChoiceDirective` where appropriate to maintain consistency with internal agent-loop logic.
2026-06-19 22:24:12 +02:00
can1357 a050474af7 feat: migrated validation schemas and tool definitions from Zod to ArkType
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
2026-06-18 00:59:53 +02:00
can1357 3efebf8805 fix: harden merged provider, agent-loop, eager, and autolearn paths
- agent-loop: raise repetition-detection floor to 180 chars and clear thinking
  replay anchors when collapsing a detected loop.
- providers/google: ignore empty text parts, retain terminal thoughtSignatures,
  and stop function-call signatures clobbering the prior block.
- autolearn: capture goal-mode at the turn boundary; harden managed-skill writes
  against hard-links/symlinks (O_NOFOLLOW + nlink); refuse minting managed skills
  whose name an authored skill already claims.
- eager tasks: thread agentKind through the session so a custom top-level agentId
  still gets always-mode delegation; split Eager Tasks prompt into hard vs soft.
- title-generator: race the online title model against a local tiny-model fallback.
- eager-todo: keep the soft reminder aligned with the todo init schema.
- mcp/stdio: keep close() detaching the read loop instead of awaiting it.
- stream loop: fix collapsing and tool-call thought-signature handling.
2026-06-14 17:09:59 +02:00
can1357 c03566cd01 Merge PR #2540: feat(coding-agent): three-level eagerness enum for task.eager and todo.eager 2026-06-14 17:09:41 +02:00
roboomp a72fc212ca style: bun run fix 2026-06-14 12:05:38 +00:00
roboomp f8a8be8664 fix(agent): corrected eager todo init prompt
Removed schema-incompatible task metadata instructions from the eager todo prelude so GPT-5.5 can satisfy the forced initial todo call.

Added regression coverage that keeps the eager init prompt aligned with the todo init schema.

Fixes #2561
2026-06-14 12:05:28 +00:00
metaphorics 42626b8c44 feat(coding-agent): make task.eager and todo.eager three-level eagerness enums 2026-06-14 11:09:29 +09:00
can1357 64aa558e62 chore: consistency 2026-06-13 00:03:27 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 3b7a44cb49 test(coding-agent): realigned stale tests with intentional behavior changes
These three suites encoded pre-refactor behavior and broke the 15.10.3 release CI.

- event-controller read grouping: the group header count now reflects aggregated
  display rows (distinct files), and a mid-turn visible-reasoning break finalizes
  the prior group but keeps it live until pending reads settle (97ae0fd3, b5eff5b3).
- gallery harness: lsp no longer attaches an instance renderer (custom rendering
  removed in d9d06134), so the custom-branch regression guard is retargeted to
  task, which still attaches its renderer and merges call+result.
- eager-todo enforcement: the prelude reminder converts to a developer-role
  message now that auxiliary messages map to developer for compaction (e13f2de5).
2026-06-08 07:17:08 +02:00
can1357 dc4aeb7b88 refactor(coding-agent): renamed todo_write tool to todo
- Renamed `TodoWriteTool` to `TodoTool` and its source/prompt files.
- Updated tool registration, schema, renderers, and gating to `todo`.
- Adjusted cursor provider native tool names and tests to match.
- Renamed strike-animation constants and `todo-error-reminder` type.
2026-06-04 02:45:30 +02:00
can1357 2867e1f4e3 feat(deps): added pi.zod exports and removed TypeBox package exports
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
2026-05-15 14:46:54 +02:00
can1357 1e601b9094 test(agent): replaced agent stream mocks with createMockModel responses
- Replaced custom MockAssistantStream helpers with createMockModel streams across agent tests.
- Removed manual queueMicrotask stream-event scripting in favor of scripted mock responses.
- Consolidated helper fixtures by deleting local aliases and reusing shared user-message/model helpers.
- Updated test assertions to use mock.calls and mock.model metadata for call and context validation.
2026-05-15 14:46:54 +02:00
can1357 2740ef65a0 fix(tests): removing useless tests 2026-05-07 09:33:50 +02:00
can1357 8c323666be feat: added ordered systemPrompt arrays and normalized context prompts
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
2026-05-04 15:20:26 +02:00
can1357 5003996a01 feat(coding-agent): added content-based todo matching for write/commands
- Removed `id` fields from todo models/fixtures and switched session clones to content-based task identity.
- Replaced `/todo_write` `replace` with `init`, updated setup schemas to `list`/`phase`, and append content-only items.
- Updated `/todo` command flows to match phases and tasks by names/content (exact/prefix/substr, case-insensitive), with no ID targeting.
- Updated rendering/output labels to `# Todos`, `formatPhaseDisplayName`, and Roman-numeral phase headings across todo views.
- Aligned prompts, changelog, and todo tests/fixtures with the new init and content-based todo-write contract.
2026-04-30 06:40:53 +02:00
can1357 52da3674a4 feat(coding-agent): added auto-rebase for stale atom/hashline anchors
- Added auto-rebasing for stale atom and hashline anchors within ±2 lines, with warning diagnostics.
- Removed atom range-locator support and dropped `between` ops, updating docs/tests for `sub` over `set` nudges.
- Changed no-op handling to track unchanged hashline edits and emit contextual hints for unchanged ranges.
- Expanded Anthropic strict error handling to retry on schema-too-complex and compiled-grammar-too-large errors.
- Added anchor-retargeting warning checks and updated edit tests to expect locator and range-locator rejections.
- Updated python tool-call fixtures by adding `title` to executed-cell payloads and adjusting related test expectations.
2026-04-26 05:22:15 +02:00
can1357 652f49442f refactor(coding-agent/tools): wrapped todo_write parameters in an ops object
- Changed the todo-write parameter schema from a bare operations array to an object containing an `ops` array.
- Updated request parsing and renderer logic to consume `params.ops` and `args.ops` when processing todo operations.
- Updated session and unit tests to invoke todo_write with wrapped `{ ops: [...] }` payloads.
2026-04-26 05:11:59 +02:00
can1357 e3f7496deb feat(coding-agent): added ordered todo_write ops and sequential execution
- Changed `todo_write` to an ordered `op`-array model with `replace`, `start`, `done`, `rm`, `drop`, `append`.
- Removed legacy multi-field todo payloads and updated tests/fixtures to use ordered `{op, task?, phase?, items?}[]` args.
- Reworked todo operation execution to apply entries sequentially and validate missing or unknown task/phase IDs.
- Updated todo rendering to use `todo.content` only and changed bash artifact labels from `full result` to `raw output`.
2026-04-26 05:00:54 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
Can Bölük 4a731f1ecc feat(coding-agent): added todo-write mutation fields in place of ops
- Replaced todo_write's `ops` payload with top-level mutation fields (`phases`, `complete`, `start`, `add_notes`, etc.).
- Removed in-place task content and note updates; moved note writes to append-only `add_notes` calls.
- Matched `add_tasks.phase` by phase ID or name and appended new note text to existing notes.
- Changed execution to drop strict in-progress update sequencing, support `start`, and auto-promote next pending task.
- Updated tests and prompt docs to use direct payloads (`phases`, `complete`, `add_tasks`) instead of `ops`.
2026-04-15 18:53:48 +02:00
can1357 5f304cccb5 fix(coding-agent): limited eager todo enforcement to first user message
- Added a guard in AgentSession to detect prior user messages and skip eager todo injection after the initial user turn.
- Updated session tests to assert that only the first user message triggers eager todo enforcement while follow-up prompts do not.
- Documented the behavior change in the package changelog under Unreleased.
2026-04-11 17:16:59 +02:00
can1357 ea9c46262f feat(coding-agent): introduced /force command and ToolChoiceQueue for directive management
- Added `/force` slash command and `ToolChoiceQueue` system for managing tool-choice directives with lifecycle callbacks and requeue semantics.
- Added `setForcedToolChoice()`, `peekQueueInvoker()`, `buildToolChoice()`, `steer()`, and `getToolChoiceQueue()` methods to AgentSession and ToolSession APIs.
- Refactored tool-choice override mechanism from simple override to queue-based system with generator directives, lifecycle callbacks, and requeue preservation.
- Removed `PendingActionStore` class and replaced with `ToolChoiceQueue`; updated `ResolveTool` and custom tool loader to use queue invokers.
- Fixed tool-choice queue cleanup on agent loop abort and requeue semantics to preserve callbacks across abort cycles.
2026-04-08 22:10:37 +02:00
can1357 d19d51770f feat(coding-agent): enabled eager todo skip for queries and commands
- Eager todo enforcement now skips prompts ending with question marks or exclamation marks, treating them as queries or commands rather than statements requiring task planning.
- Added 2 test cases validating that eager todo enforcement is skipped for prompts ending with question and exclamation marks.
2026-04-08 03:48:46 +02:00
can1357 2f151fea9a fix(tests): added resource cleanup methods and initiatorOverride support
- Added `close()` method to SessionManager and AuthStorage for proper resource cleanup and finalization of prepared statements.
- Added `initiatorOverride` option support in OpenAI and Anthropic providers for message attribution control.
- Fixed resource leaks in RpcClient timeout handling by centralizing timeout creation with unref() and adding explicit clearTimeout() calls.
- Fixed AgentSession disposal to call SessionManager's `close()` method for guaranteed resource cleanup instead of fallback flush.
- Updated all test suites to properly dispose AuthStorage instances in cleanup hooks to prevent resource leaks between tests.
2026-03-14 11:25:40 +01:00
can1357 c8d4946d70 feat(coding-agent): refactored eager todo messaging for cleaner categorization
- Changed eager todo reminder message role from 'developer' to 'custom' with customType field for better message categorization.
- Removed userRequest parameter from eager todo prelude generation to simplify prompt template rendering.
- Updated eager todo prompt to avoid redundant todo_write calls unless task state materially changed.
- Modified eager todo reminder message to use string content with display: false property instead of array format.
2026-03-11 03:08:32 +01:00
can1357 1932b35062 refactor(session): restructured eager todo injection to prepended message pattern
- Refactored eager todo injection from recursive prompt call to prepended message pattern.
- Removed state fields for todo injection tracking and consolidated logic into message composition.
- Added prependMessages option to promptWithMessage for composing messages before main prompt.
- Updated system prompt to require todo creation before substantive work on user requests.
2026-03-11 02:52:39 +01:00
can1357 1a5bbc3e51 feat(coding-agent): added details field to TodoItem for storing implementation specifics
- Added optional `details` field to TodoItem type for storing implementation specifics, file paths, and edge cases.
- Enhanced todo item display to show multi-line details with automatic indentation in interactive and reminder modes.
- Updated eager-todo system prompt to enforce separation of short task content (5-10 words) from detailed implementation information.
- Extended TodoWriteTool to support creating and updating tasks with details field via add_task and update operations.
- Added comprehensive test coverage for details field handling across todo operations (replace, add_task, update).
2026-03-11 00:56:54 +01:00
can1357 bb0026cb32 feat(coding-agent): added eager todo configuration and per-turn tool choice overrides
- Added 'todo.eager' configuration setting to automatically create a comprehensive todo list after the first user message.
- Added 'buildNamedToolChoice' utility function to build provider-aware tool choice constraints for named tools.
- Modified tool choice resolution to support per-turn tool choice overrides via consumeNextToolChoiceOverride() method.
- Implemented eager todo enforcement mechanism that injects a synthetic prompt to encourage todo creation when conditions are met.
- Extracted tool choice building logic into reusable utility module for better code organization.
- Added comprehensive test coverage for eager todo enforcement functionality in AgentSession.
2026-03-11 00:54:55 +01:00