- Replaced Python execution with a local `python -u runner.py` subprocess and NDJSON stdin/stdout framing.
- Removed shared-gateway architecture, including coordinator lifecycle APIs, `useSharedGateway` wiring, and `jupyter` CLI/actions.
- Simplified setup checks to a plain Python 3 availability probe and removed automatic dependency-install fallbacks.
- Updated kernel cancellation and display processing to use status frames, SIGINT/SIGTERM escalation, and normalized output coercion.
- Added `python-runner` integration and display tests while deleting legacy websocket and kernel lifecycle test suites.
- Removed the JS and Python eval prelude `run` helpers, including their shell execution and timeout/cwd option handling.
- Updated the JS VM helper set to expose `Bun` and removed the deleted `run` entry from the prelude.
- Revised eval docs to drop `run` from the helper surface and note the new JS `Bun` global.
- Replaced `===== ... =====` eval cell headers with `*** Begin ` / `*** End ` markers; legacy format remains renderable in HTML exports.
- Replaced hashline patch grammar with `*** Begin Patch` / `*** End Patch` envelope; old inputs without the envelope are still accepted.
- Extracted `sniffEvalLanguage` into a shared `sniff.ts` module reused by the parser and tool.
- Added `docs/ERRATA-GPT5-HARMONY.md` and `scripts/session-stats/harmony_backtest.py` documenting and backtesting the GPT-5 Harmony-header leak defect.
RenderMermaid is disabled by default and emits terminal text rather
than SVG/PNG; complex sequence diagrams with alt/else blocks become
hard to read at default widths. Add docs/render-mermaid.md covering
enablement (renderMermaid.enabled), config knobs (useAscii, paddingX,
paddingY, boxBorderPadding), output expectations, and limitations.
Link the new guide from the coding-agent README.
Fixes#961
- Removed PI_STRICT_EDIT_MODE gating from edit-mode resolution so model fallbacks now always apply.
- Stopped injecting PI_STRICT_EDIT_MODE in edit-benchmark.py and rate-edit-tool.py execution environments.
- Removed PI_STRICT_EDIT_MODE from environment-variable documentation and strict-mode test coverage.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
Exposes model.compat.disableStrictTools (already supported by the anthropic
transport since #826) via models.yml so users can configure it without code
changes.
Set disableStrictTools: true at the provider level to disable strict tool
schemas for third-party Anthropic-compatible endpoints (AWS Bedrock, Vertex
AI proxies, custom gateways) that reject the strict field.
- Add disableStrictTools to ProviderConfigSchema
- Merge { disableStrictTools: true } into provider compat override when set,
flowing through the existing compat pipeline to model.compat.disableStrictTools
- disableStrictTools alone is sufficient for an override-only provider entry
- Update docs/models.md with field reference, Bedrock example, and proxy note
- Add tests covering provider-level propagation, built-in override, and
overlay merge
- Renamed subagent completion flow from `submit_result` to `yield` across SDK tools, prompts, and docs.
- Updated executor/task handling to require and parse `yield` calls, replacing legacy submit-result extraction and state flags.
- Added `subagent-yield-reminder` and updated system prompts to require `yield` with `result.data` or `result.error`.
- Renamed hidden-tool and registration plumbing to `yield`, including discovery helpers and renderer/test surface.
- Removed `process.env.PI_DEV` checks from `packages/natives/native/index.js`, dropping per-candidate debug load logging and load-error reporting from the native addon loader.
- Pruned `PI_DEV` and `--dev` usage from native build/run and runtime docs, including setup, variant, and troubleshooting guidance.
- Updated `packages/natives/CHANGELOG.md` to note removal of the `PI_DEV` loader diagnostic environment variable and associated console logging.
- Enabled Bedrock runtime and AWS credential calls to use a proxy-aware HTTP/1 handler when HTTPS_PROXY, HTTP_PROXY, or ALL_PROXY (including lowercase variants) is configured.
- Added a retry path that recreates the Bedrock client with HTTP/1 transport when an initial HTTP/2-related error occurs before streaming starts.
- Updated package and lock dependencies to include Bedrock credential-provider and proxy-agent support, and documented the new Bedrock proxy environment variables.
Slots a new "apply_patch" variant alongside the existing edit modes
(replace, patch, hashline, chunk, vim). The mode accepts a single input
string containing a Codex *** Begin Patch / *** End Patch envelope,
parses it with a new lenient parser (heredoc-tolerant), and fans each
file-op out to the existing executePatchSingle so LSP writethrough,
plan-mode guards, fs-cache invalidation and diagnostics are shared
with the patch mode.
Exposes both tool shapes from the spec: the JSON function-tool variant
(§1.2, {input: string}) and the OpenAI custom-tool / Lark-grammar
"freeform" variant (§1.1, raw patch string). The edit tool advertises
a Lark grammar via customFormat and a wire name via customWireName;
openai-responses emits it as a grammar-constrained custom tool when a
model opts in with applyPatchToolType: "freeform" in models.json.
custom_tool_call / custom_tool_call_output are plumbed end-to-end
through the shared responses code (emission, streaming, history
replay), and the agent-loop dispatcher matches tool calls by either
name or customWireName so returned calls route correctly.
Also threads preview/diff rendering for apply_patch through the TUI
(tool-execution + edit renderer) so streaming patches show per-file
diffs like the other edit modes.
Default edit mode is unchanged (hashline); opt in via edit.mode or
PI_EDIT_VARIANT=apply_patch.
- Added canonical model equivalence types, cache helpers, and registry APIs for provider variant lookup.
- Changed model resolution to apply canonical ID overrides/excludes with provider order before fallback matching.
- Added canonical and provider model views in list-models and selector UI with canonical sorting/persistence.
- Updated role/model persistence to store selectors while runtime now resolves concrete canonical-backed provider models.
- Document session name getter/setter methods in extensions.md
- Add stub methods to ExtensionProxy that throw if called before init
- Add delegating implementations to ExtensionProxyInit
- Wire up getSessionName and setSessionName in extension runtime
- Add method signatures to ExtensionContext interface and type
- Add getSessionName to session manager API
- Add getSessionName and setSessionName to all mode contexts:
- ACP agent
- Extension UI controller (also updates terminal title)
- Print mode
- RPC mode
- Migrated all package tsconfig files to extend tsconfig.workspace.json for unified TypeScript configuration across monorepo.
- Consolidated build and check scripts across 10+ packages to use biome for linting/formatting with separate type checking via tsgo.
- Renamed build scripts from build:native and build:binary to build for simplified command naming across packages/natives and packages/coding-agent.
- Refactored CI workflow to invoke bun tasks instead of inline shell scripts, reducing workflow complexity by 40+ lines.
- Removed sync-exports.ts and repro-stuck.ts scripts; deleted path aliases from tsconfig.base.json in favor of workspace-based configuration.
- Updated turbo.json with new task definitions (check:types, lint, fmt, fix) and removed build:native/embed:native tasks.
- Refactored chunk anchor format from bracket notation to marker-based syntax with +++ and --- delimiters.
- Renamed chunk edit regions from @container/@prologue/@body/@epilogue to @head/@inner/@tail for clearer semantics.
- Simplified line-number selector resolution to auto-convert line targets to chunk paths with checksum format.
- Removed PI_DEV debug addon loading logic; PI_DEV now only enables loader diagnostics without changing candidate filenames.
- Added full-name aliases for chunk kind patterns (decls, expr, fn, meth, mod, params, ret, stmts, var) alongside abbreviated forms.
- Refactored sanitize_node_kind() to return &str instead of String; adjusted identifier passing logic across AST modules.
- Removed explicit js_name attributes from NAPI macros across all modules, relying on automatic snake_case to camelCase conversion.
- Renamed internal functions in keys.rs for clarity: parse_kitty_sequence_napi() -> parse_kitty_sequence() and parse_kitty_sequence() -> parse_kitty_sequence_bytes().
- Renamed text.rs functions for consistency: set_env_tab_width() -> set_default_tab_width() and get_env_tab_width() -> get_default_tab_width().
- Updated documentation to reflect automatic NAPI naming convention and simplified export mapping tables.
- Added host tool execution framework with HostTool, HostToolContext, and host_tool() factory for custom tool integration.
- Added RpcConcurrencyError exception and _PromptLifecycleCoordinator to enforce single-flight constraint on prompt lifecycle methods.
- Enhanced JSON parsing with 10 validation helpers and enum frozensets for safe field extraction with detailed error messages.
- Replaced manual event/error list management with _BoundedHistory for bounded-size history with offset tracking.
- Added deep JSON cloning to prevent external mutations of stored payloads and improved UTF-8 error handling in subprocess stderr.
- Added custom_tools parameter to RpcClient and set_custom_tools() method for runtime tool registration.
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).
- Added typed event listeners and granular event handling for all RPC notification types.
- Added set_todos RPC command and todoPhases session state field for todo phase management.
- Added RpcClient initialization parameters (thinking, tools, no_session, rpc_defaults) for startup configuration.
- Added install_headless_ui() method and todo management methods (get_todos, set_todos, clear_todos).
- Added TodoItem and TodoPhase dataclasses with parser functions for structured todo representation.
- Added RPC mode behavior: disables session title generation by default and resets workflow settings to built-in defaults.
rules with alwaysApply: true were parsed by all providers and used to
exclude the rule from rulebookRules, but the inclusion half was never
built — rule content was silently dropped. now:
- full content is injected directly into the system prompt (before the
rulebook rules section) in both default and custom prompt templates
- rules remain addressable via rule:// for re-reading
- ttsr rules still take priority (condition + alwaysApply goes to ttsr only)
updated rulebook-matching-pipeline.md to reflect the three-bucket split
(ttsr > always-apply > rulebook) and corrected the rule:// resolution
docs.
- Added support for provider-level OpenAI compatibility configuration enabling reasoning effort mapping and streaming usage fallback across models.
- Added 10 new AI models (DeepSeek V3.2, Llama 3.1 405B, Mistral Large 3, Pixtral Large, and others) with updated pricing and context windows.
- Fixed autocomplete to preserve ./ prefix in relative file/directory path completions and paste marker expansion to handle regex tokens literally.
- Changed system prompt date format to ISO 8601 and tool download timeout from 15s to 120s for improved cross-platform compatibility.
- Refactored OpenAI completions provider to extract token parsing logic and support choice-level usage fallback with improved message serialization.
* add llama.cpp as local provider
* use responses api instead of messages
* use api-keys correctly for llama.cpp provider
---------
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
- Added Tavily web search provider with API key authentication and credential discovery from environment or database.
- Integrated Tavily as highest-priority search provider in fallback chain with structured response mapping and error handling.
- Added Tavily OAuth login flow in CLI and auth-storage with manual API key input and validation.
- Added comprehensive test suite for Tavily provider covering registration, response mapping, error handling, and credential validation.
Fixes#313
* feat: Add LM Studio as a supported model provider with OpenAI-compatible API fetching and discovery.
Add LM Studio as a supported AI provider with optional API key, environment variables, and discovery.
* feat: Add rustup as a dev dependency.
* fix: Refine LM Studio API key handling to conditionally send authorization headers during model discovery based on whether the key is a default local token or a custom key, and add new tests.
* feat: Enhance OAuth token and account ID resolution for model providers in the model registry.
* rebase for packages/ai/CHANGELOG.md
* feat: improve implicit model discovery to independently auto-detect Ollama and LM Studio, and refine LM Studio base URL handling.
---------
Co-authored-by: Can Bölük <can1357@users.noreply.github.com>
PendingAction.apply() and reject() now receive the reason string that was
passed to resolve(). This lets custom tools surface the agent's rationale
in their apply/discard output or use it for logging.
- PendingAction interface: apply(reason) and reject?(reason)
- CustomToolPendingAction: same signatures, reject is optional
- CustomToolLoader: threads reject through when building PendingAction
- AstEditTool: accepts _reason (unused, reserved for future tracing)
- resolve.test: covers reason forwarding on apply and reject paths,
and verifies reject return value replaces the default discard message
- docs/resolve-tool-runtime.md: updated interface table, built-in
producer description, usage example, and developer guidance