Commit Graph

52 Commits

Author SHA1 Message Date
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
roboomp 1b6588e520 fix(tools): cap per-tool default timeout with tools.maxTimeout
clampTimeout resolved the per-tool default (bash 300s) whenever the agent
omitted `timeout` and only enforced the tool's own min/max, so the
tools.maxTimeout global ceiling — applied solely in sdk.ts on explicitly
numeric args — was bypassed on the common default-fallback path.

Thread maxTimeout into clampTimeout so the resolved effective timeout,
including the default path, is capped before the per-tool floor/ceiling
apply. Explicit values below the cap still win; maxTimeout <= 0 stays
no-cap. Applied at every call site (bash, eval, browser, debug, lsp,
fetch, and the session-level bash executor), and the bash clamp notice
now names the global ceiling when it is the binding limit.

Fixes #6294
2026-07-22 15:18:27 +00:00
can1357 1678607617 fix(eval): preserve column truncation metadata 2026-07-16 03:32:03 +02:00
roboomp 8e6d26b1e8 fix(eval): honored unlimited cell timeouts
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.

Fixes #5250
2026-07-14 17:33:56 +00:00
roboomp 8614b4c086 fix(task): respected restricted spawn defaults
Resolved eval agent() and task tool defaults from the active spawn policy so restricted agents advertise and execute an allowed default.

Fixes #3973
2026-07-01 02:44:08 +00:00
can1357 c4e23fed15 fix(coding-agent/tools): enabled live stdout streaming for running cells
- Update the eval tool to stream stdout chunks directly into the active cell's output buffer while the process is still running.
- Prevent long-running cells from appearing empty in the UI by surfacing incremental output before the backend resolves.
- Add regression tests to ensure streamed output is captured mid-execution and reconciled with final results.
2026-06-23 01:46:45 +02:00
can1357 060f4004e7 feat(coding-agent): refactored eval tool to single-step execution
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
2026-06-23 00:59:58 +02:00
can1357 bda98c63ef feat(coding-agent): elevated eval to an essential tool to ensure
- Elevated `eval` to an essential tool to ensure availability across all discovery modes.
- Updated system and tool prompts to mandate the use of `eval` for non-trivial shell operations like conditionals, loops, heredocs, and complex pipelines.
- Restricted `bash` usage to simple binary invocations and single-fact computation to reduce shell-escaping and execution errors.
2026-06-23 00:01:05 +02:00
can1357 c926eb381d feat(coding-agent/tools): made ruby and julia backends opt-in
- Changed default evaluation backend configuration to only enable Python and JavaScript by default.
- Implemented dynamic tool parameter generation to hide Ruby and Julia from the model's schema when they are disabled in settings.
- Updated tool summary and field descriptions to reflect the currently enabled runtime backends.
2026-06-22 06:13:01 +02:00
can1357 33e2594f03 feat(coding-agent): added support for Julia and display language icons in code cells
- Added Julia language support to the theme symbol maps.
- Enabled language icons in code cell headers for the eval tool renderer.
2026-06-22 06:13:01 +02:00
can1357 1f3f3cf5d1 feat: added ruby and julia language support to coding-agent
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
2026-06-22 06:13:01 +02:00
can1357 0abcd76101 docs(coding-agent-prompts): refined agent instructions and tool documentation
- Simplified system and personality prompts for improved conciseness and clarity.
- Streamlined tool instruction sets and parameter descriptions across all agent modules.
- Refactored prompt documentation in `hashline` to clarify terminology and task-specific constraints.
- Updated tool metadata in TypeScript service definitions to align with reduced documentation verbosity.
2026-06-19 16:06:17 +02:00
can1357 9d2728a455 ux(coding-agent): standardized truncation formatting for elided content
- Updated log and error truncation messages to use a consistent `[...Nch elided...]` format.
- Standardized diagnostics and prompt text output to improve consistency in reporting truncated information.
2026-06-19 04:45:08 +02:00
can1357 e92db73ec6 feat(coding-agent/tools): added explicit ArkType schema descriptions to agent tools
- Added explicit ArkType schema descriptions across all coding agent tool definitions.
- Updated schema definitions in autoresearch and commit tools with descriptive wrappers.
- Documented tool schema enhancements in the packages/coding-agent CHANGELOG.
2026-06-18 00:59:55 +02:00
can1357 a050474af7 feat: migrated validation schemas and tool definitions from Zod to ArkType
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
2026-06-18 00:59:53 +02:00
can1357 7687810c5d feat: added syntax-aware tool example rendering across catalog, AI, and agent modules
- Added model-to-syntax mapping in catalog with preferred tool-call syntax API.
- Added `ToolExample` typing and `ToolCallSyntax` exports across tool/grammar interfaces.
- Added syntax-aware tool example rendering through provider-specific grammar invocations.
- Added `exampleSyntax` context flow and example metadata so rendered prompts include examples.
2026-06-15 07:33:25 +02:00
can1357 92fba8d486 fix: addressed model-aware image handling for tool image WebP filtering
- Added model-aware session image normalization by passing active model.
- Propagated model-derived WebP exclusion into browser snapshot capture and resizing.
- Applied model-derived WebP exclusion to eval, fetch, read, and inspect image paths.
2026-06-13 16:17:56 +02:00
can1357 64aa558e62 chore: consistency 2026-06-13 00:03:27 +02:00
can1357 d7eae06830 fix(coding-agent): fixed eval artifact double-writes and kernel I/O capture
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
2026-06-10 01:27:41 +02:00
can1357 347f857f03 feat(eval): raised per-cell timeout ceiling from 600s to 3600s
- Matched the `bash` tool's max in the Zod schema and runtime clamp.
- Allowed heavy local-compute cells to request budgets above 10 minutes.
- Updated docs and prompt to reflect the new 1-3600s range.
2026-06-08 18:28:08 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
can1357 49aa6e5839 fix(eval): corrected eval LLM calls and spawn-aware tool descriptions
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.
2026-06-08 02:04:59 +02:00
can1357 6efe86c07b fix(coding-agent): honored boolean env flag overrides over settings
- Made PI_INTENT_TRACING, PI_AUTO_QA, PI_PY, and PI_JS take precedence when set.
- Fell back to config when the env flag is unset instead of ORing.
- Surfaced PI_PY=0/PI_JS=0 in the disabled-backend error messages.
2026-06-06 16:30:09 +02:00
can1357 76ca917bfd fix(coding-agent-eval): suspended idle timeout during delegated bridge calls
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
2026-06-06 14:02:04 +02:00
can1357 503d29f408 fix(eval): resolved TDZ crash by splitting eval renderer into eval-render.ts
- Extracted TUI rendering from `eval.ts` into a dependency-light `eval-render.ts` to break the circular initialization chain.
- `renderers.ts` now imports `evalToolRenderer` from `eval-render` directly, avoiding re-entry into the root barrel while `eval.ts` is still initializing.
- `eval.ts` re-exports `evalToolRenderer` and `EVAL_DEFAULT_PREVIEW_LINES` for backward compatibility.
- Increased first-event timeout test budget from 50ms to 5000ms to prevent CI scheduler jitter from tripping the watchdog on success cases.
2026-06-01 20:31:46 +02:00
can1357 c069136eca refactor(exec): extracted eval-backends module and improved shell quarantine
- Moved EvalBackendsAllowance and related functions to a dedicated eval-backends.ts module.
- Replaced ad-hoc brokenShellSessions tracking with a quarantineShellSession helper that also awaits the abort cleanup promise.
- Applied quarantine on timeout and cancellation paths, not just errors.
2026-06-01 20:15:21 +02:00
can1357 dbd9489010 refactor(eval): changed timeout from inactivity to wall-clock budget
- Only bridge heartbeats (`agent()`/`llm()`) now re-arm the watchdog; compute, stdout, `log()`/`phase()`, and ordinary tool calls count against the budget.
- Emitted an immediate heartbeat at bridge call start to avoid early abort near budget edge.
- Removed `idle` flag and "of inactivity" suffix from timeout annotation strings.
- Updated docs, prompts, and comments to reflect the new wall-clock semantics.
2026-06-01 17:17:00 +02:00
can1357 e18e4ada71 fix(eval): keep idle watchdog armed during in-flight agent()/llm() calls
The per-cell `timeout` is an inactivity budget that only re-arms on status
events, but host-side bridge calls can run long stretches with no
intermediate status (a subagent's time-to-first-token on a reasoning
model, a long quiet nested tool, or an entire oneshot llm() request).
The watchdog mistook that for a stall and aborted working subagents
mid-flight.

Pump a lightweight heartbeat while a bridge call awaits, re-arming the
watchdog through the existing emitStatus -> onStatus channel. The
heartbeat is a pure keepalive: forwarded to bump the timer but never
stored or rendered, so a genuinely stalled cell is still interrupted
once the call settles.

- eval/heartbeat.ts: withBridgeHeartbeat() + EVAL_HEARTBEAT_OP
- agent-bridge/llm-bridge: wrap runSubprocess / completeSimple
- js+py executors: forward heartbeat to onStatus, drop from displayOutputs
- tools/eval.ts: bump on heartbeat, skip persist/render
2026-06-01 16:38:32 +02:00
can1357 2003d7382e feat(eval): added per-cell inactivity timeout budgets in eval executors
- Changed eval timeout behavior from hard wall-clock deadlines to per-cell inactivity budgets in all executors.
- Added IdleTimeout watchdog support, including bumps on status/tool activity and timer cleanup after execution.
- Updated executor option plumbing to replace deadlineMs with idleTimeoutMs and emit inactivity timeout annotations.
- Added IdleTimeout and shared-executor tests and updated prompt/repl docs for the new timeout contract.
2026-05-31 10:11:59 +02:00
can1357 9fabd5e4d5 feat(coding-agent): added onStatus in eval backends for status streams
- Added optional `onStatus` callback wiring across eval backends and JS/Python executors for live status streams.
- Added collectDisplay-based forwarding so `emitStatus` and `onDisplay` route status outputs consistently.
- Expanded agent status payloads with preview/model/token-cost context and kept completion updates single-pass.
- Added status upsert and render adjustments in `tools/eval.ts` to coalesce agent events with progress stats.
- Added status/progress test coverage for running/completed agent events, final metric retention, and parallel placement.
- Updated CHANGELOG Unreleased notes to record live progress updates and completion-status metric fixes.
2026-05-31 08:45:12 +02:00
can1357 4ab40764d1 refactor(coding-agent): removed args global from eval runtimes
- Dropped `args` input from eval tool schema, JS/Python executors, and worker protocol.
- Removed per-call `args` injection from JS runtime and Python kernel/runner.
- Deleted related tests and updated docs to reflect removal.
2026-05-31 07:54:52 +02:00
can1357 7613c2a913 feat(coding-agent): expanded eval execution and bridge paths with workflow helpers
Expand eval execution/bridge paths with args/log/phase/budget support, add workflow helpers (parallel/pipeline), expose usage statistics, and add eval integration tests and docs.
2026-05-31 07:40:18 +02:00
can1357 ebb9276393 feat(coding-agent): added animated pending border for bash/eval blocks
- Added clockwise sweeping dark segment animation to output block borders while bash/eval tool calls are pending/running.
- Changed bash renderCall to immediately render a full bordered block instead of a one-liner status preview, so silent commands show the framed block for their entire runtime.
- Added shimmerEnabled() helper and wired animate flag through OutputBlockOptions, CodeCellOptions, and shell/eval renderers.
2026-05-31 01:43:00 +02:00
can1357 8715ed207c feat(coding-agent-eval): added oneshot llm helper and __llm__ bridge
- Added one-shot `llm(prompt, opts)` helpers in JS and Python eval runtimes.
- Added `__llm__` eval bridge wiring for synthetic LLM tool dispatch and status/event output.
- Added `runEvalLlm` with tier-to-model resolution, effort handling, and oneshot completion execution.
- Added structured schema output handling via `respond` tool and JSON fallback parsing.
- Documented new llm behavior in eval docs/changelog and added tests for tier mapping and error cases.
2026-05-30 00:27:04 +02:00
can1357 04d6e831dc refactor(eval): replaced inline JSON tree renderer with shared renderJsonTreeLines
- Dropped local `renderJsonTree` and `formatJsonScalar` in favor of the shared `renderJsonTreeLines` used by tool args, MCP results, and subagent output.
- Removed `Object(N)`/`Array(N)` type labels and per-output `JSON output N` headers; type icons and bare keys are used instead.
- The `display[N]` header is now shown only when a cell emits more than one `display()` value.
2026-05-30 00:27:03 +02:00
can1357 7dd00c015b feat: hashline improvements for spark
- Redesigned hashline patch syntax from anchor-based (`A-B:`) to hunk-header format (`@@ A..B @@`) with unified-diff compatibility.
- Removed `autoDropPureInsertDuplicates` option and simplified apply behavior to preserve duplicated boundary and context lines.
- Changed repeat operator from `^A-B` to `&A..B` and range separator from `-` to `..` for consistency with hunk-header syntax.
- Added image resizing and dimension notes to eval tool output; improved write tool hashline header sanitation for legacy formats.
- Removed 521 lines of boundary-duplicate absorption code and simplified parser to auto-convert bare body rows and unified-diff contamination.
2026-05-28 03:06:52 +02:00
can1357 e4a16451ec feat(coding-agent): added coding-agent approval types and mode options
- Added `ToolTier`, `ToolApproval`, and `ToolApprovalDecision` types and exported approval APIs.
- Updated approval-mode options from `auto|prompt|custom` to `always-ask|write|yolo` and defaulted mode to `yolo`.
- Changed approval resolution to apply per-tool decisions first, then mode-tier limits, with legacy-mode migration.
- Assigned read/write/exec `approval` and approval-detail prompts across built-in, custom, extension, and MCP tools.
2026-05-26 21:52:16 +02:00
can1357 8a5b3e9552 feat(eval): added shared executor inheritance for subagents with concurrent async cells
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
2026-05-26 14:37:56 +02:00
can1357 84ec8fba49 feat(coding-agent/eval): implemented JSON cell-based eval tool inputs
- Removed the legacy `parseEvalInput` parser module and `eval.lark`, eliminating `*** Cell` stream parsing.
- Replaced eval tool arguments from single `input` strings to ordered `cells` arrays in tool calls and schema.
- Updated execution to resolve language explicitly, map `py` to `python`, and apply timeout/reset defaults.
- Removed backend sniffing and `ABORT_WARNING` suffix handling, then updated docs and tests to the new JSON cells format.
2026-05-16 19:33:44 +02:00
can1357 2867e1f4e3 feat(deps): added pi.zod exports and removed TypeBox package exports
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
2026-05-15 14:46:54 +02:00
can1357 f0ff398607 fix(coding-agent/tools): stripped duplicate output notices from TUI tool renderers
- Added stripOutputNotice to output-meta to remove appended truncation notices when output metadata is available.
- Updated bash, eval, browser, read, and ssh renderers to strip the notice before display so the styled warning line is not duplicated.
- Left fallback behavior unchanged so outputs without a notice continue through unchanged.
2026-05-15 03:16:16 +02:00
can1357 28b9ce7a0c feat(coding-agent): added middle-elision caps to OutputSink truncation
- Added `tools.artifactHeadBytes` and `tools.outputMaxColumns` settings with defaults in `SETTINGS_SCHEMA`.
- Expanded `OutputSink` with `headBytes`/`maxColumns` and middle truncate logic with elision markers and tracking.
- Updated output-meta to resolve sink settings, emit truncation metrics, and use `truncateMiddle` for spills.
- Integrated head and column limits into JS/Python/Bash/SSH/read output flows, with `:raw` skipping read truncation.
- Documented new output middle-elision and column-cap behavior in `CHANGELOG.md`.
- Added truncation tests for `OutputSink`, `truncateMiddle`, and read-tool line handling.
2026-05-13 11:19:11 +02:00
can1357 028a442f18 feat(coding-agent): added parser support for t/rst in Cell eval
- Introduced canonical `*** Cell` headers with `t:` and `rst` attributes in eval prompts, schema, and docs.
- Updated parser and grammar to parse `*** Cell` blocks, stop on `*** End`/next header/EOF, and handle invalid `rst` with errors.
- Added quote-aware attribute tokenizers and split HTML eval parsing into `Cell` and legacy `Begin` handlers with `py` defaults.
- Expanded parsing behavior and tests for `rst` booleans, title aliases, abort boundaries, and stray-line skips between cells.
2026-05-12 10:01:56 +02:00
can1357 a4d86a075a feat(coding-agent): added verbatim unicode rule to hashline prompt 2026-05-12 05:25:55 +02:00
can1357 2268fcc5bd feat(coding-agent/tools): surfaced display() outputs in eval tool responses
- Added helpers to stringify and truncate display() JSON values before including them in text responses.
- Updated eval execution to append formatted display output text alongside stdout and emit images as dedicated content blocks.
- Adjusted no-output messaging for image-only runs and removed detail-level image payload duplication when content already includes image blocks.
2026-05-12 03:23:47 +02:00
can1357 8b92ec937e feat: gpt-5 harmony errata fixes
- Replaced `===== ... =====` eval cell headers with `*** Begin ` / `*** End ` markers; legacy format remains renderable in HTML exports.
- Replaced hashline patch grammar with `*** Begin Patch` / `*** End Patch` envelope; old inputs without the envelope are still accepted.
- Extracted `sniffEvalLanguage` into a shared `sniff.ts` module reused by the parser and tool.
- Added `docs/ERRATA-GPT5-HARMONY.md` and `scripts/session-stats/harmony_backtest.py` documenting and backtesting the GPT-5 Harmony-header leak defect.
2026-05-10 19:52:44 +02:00
can1357 0e8b797abd feat(coding-agent/tools): added intent synthesis for eval tool inputs
- Added an intent function to EvalTool that generated a label from provided input.
- Parsed input with parseEvalInput and joined each cell title or language fallback into a newline string.
- Returned "evaluating" when input was missing or could not be parsed.
2026-05-09 02:34:13 +02:00
can1357 2f871b6d23 feat(coding-agent): added loadMode and summary to AgentTool discovery
- Added optional `loadMode` and `summary` fields to `AgentTool` and related type declarations.
- Added `loadMode` and `summary` metadata to built-in tool classes for discoverable/essential behavior.
- Replaced `BUILTIN_TOOL_METADATA` with per-tool fields in discovery code paths.
- Updated `search_tool_bm25` and discovery indexing to use each tool's `summary` text.
- Updated discovery tests to validate tool `loadMode` and summary completeness.
2026-05-06 19:14:39 +02:00
can1357 d064170c56 fix: chatgpt really doesnt like sandwitching lark 2026-05-02 04:54:56 +02:00
can1357 9fc762cbad refactor(packages/coding-agent): reorganized eval helper APIs in prelude
- Removed deprecated file helper APIs (find, glob, grep, rgrep, sed, and stat) from JS and Python eval preludes.
- Removed status icons and formatting branches for find/grep/rgrep/glob/stat/sed from tools/eval.ts.
- Updated eval helper docs and Python prelude tests to match the reduced exposed helper surface.
2026-05-02 04:34:39 +02:00