- Centralized `line-hash.ts` hash regex sources and resolved `atom.lark` via `resolveLarkLidPlaceholders`.
- Expanded `computeLineHash` to emit `>[a-z]` and `[a-z]<` hashes for brace-context anchors.
- Replaced atom/hashline parsers' hard-coded lid regex with shared `HASHLINE_HASH_RE_SRC` and lax counterparts.
- Removed `\\TEXT` continuation handling in atom rewrites and switched multi-line replacements to `+TEXT`.
- Added brace-body insertion warning when `@Lid` on `{`-ending lines inserts at non-body-safe indent.
- Suppressed duplicate auto-rebase warnings and kept unmatched `-`/`+` ranges separate in compact previews.
- Removed deprecated file helper APIs (find, glob, grep, rgrep, sed, and stat) from JS and Python eval preludes.
- Removed status icons and formatting branches for find/grep/rgrep/glob/stat/sed from tools/eval.ts.
- Updated eval helper docs and Python prelude tests to match the reduced exposed helper surface.
- Removed PI_STRICT_EDIT_MODE gating from edit-mode resolution so model fallbacks now always apply.
- Stopped injecting PI_STRICT_EDIT_MODE in edit-benchmark.py and rate-edit-tool.py execution environments.
- Removed PI_STRICT_EDIT_MODE from environment-variable documentation and strict-mode test coverage.
- Added branch-aware session loading so autoresearch state only rehydrates for current branch.
- Replaced user-specified experiment commands with fixed `bash autoresearch.sh` execution flow.
- Enforced safer setup checks, including missing `autoresearch.sh` and uncommitted-worktree errors.
- Added branch-specific storage helpers, baseline-commit persistence, and expanded tests for dirty-path cases.
- Adjusted atom mutation conflict validation to skip throwing on repeated delete edits for the same anchor line.
- Added tests confirming duplicate delete edits on one anchor are idempotent and do not trigger conflicts.
- Added coverage ensuring explicit deletes within replace ranges are treated as redundant and ignored.
- Optimized `repairJson` to scan JSON input incrementally and only patch invalid escape sequences or raw control characters.
- Reworked escape validation with byte-level constants and stricter `\uXXXX` handling for truncated or non-hex sequences.
- Added coverage for valid escapes, control-character repair, literal invalid-escape preservation, and whitespace-only streaming JSON inputs.
- Added `disableReasoning` to `SimpleStreamOptions` and OpenAI completions, sending `reasoning: { enabled: false }` for OpenRouter requests to prevent reasoning models from consuming the full output budget on small calls like title generation.
- Fixed `canAppend` to accept `response.completed` as a terminal event, restoring websocket append reuse after codex sessions end.
- Replaced async blob-decoding and `addEventListener` with synchronous `onmessage`/`onopen`/`onerror`/`onclose` handlers and `binaryType = "nodebuffer"` for simpler, reliable message decoding.
- Simplified title generator to discard per-role thinking level and always pass `disableReasoning: true`.
- Removed title-source aware branching from session terminal-title and accent helpers, and updated callers to use session name plus cwd only.
- Dropped UUID-based recent-session naming by preferring explicit header titles or first user prompts and generating an "Untitled · <time>" fallback.
- Adjusted welcome session-row rendering for width-aware name truncation and disabled reasoning in title generation requests to keep terminal titles concise.
- Replaced file-backed autoresearch contracts with sqlite-backed session/run storage in `~/.omp/autoresearch`.
- Added `AutoresearchStorage` and rewired `init_experiment`, `run_experiment`, and `log_experiment` to persist sessions and runs.
- Added `update_notes` tool with `body`/`append_idea` inputs and updated prompts to use active-session context.
- Removed `autoresearch.md` contract parsing and checks flow, including `runChecks`, `force`, and timeout schema options.
- Updated autoresearch state/types to persist `goal`, `notes`, `branch`, and `baselineCommit` plus run justification/flag metadata.
- Fixed abort-source handling so caller aborts always win and local reasons only attach to matching request signals.
- Fixed agent-loop streaming to race event reads against abort signals and emit an aborted assistant message.
- Fixed Anthropic request construction to honor thinkingEnabled=false and omit temperature/top_p/top_k for Opus non-thinking.
- Fixed OpenAI Codex request handling by normalizing response URLs, decoding non-string websocket frames, and cleaning handshake headers.
- Added regression tests for abort precedence, Anthropic alignment cases, and Codex stream/header normalization.
- Documented the cancellation and provider behavior fixes in package changelogs.
- Added a `reset` helper in git utilities to execute `git reset` with optional flags.
- Added support for `hard`, `mixed`, and `soft` mode options, an optional target revision, and signal forwarding to `runEffect`.
- Fixed path resolution to emit per-target fanout when common base collapses to filesystem root, preventing full-filesystem scans.
- Fixed match/file counts and pagination to aggregate correctly across all targets.
- Fixed returned paths to be normalized relative to the original search scope.
- Added a runtime export for `OpenAICodexResponsesOptions` and `AnthropicCompat` capability flags.
- Added Anthropic raw SSE decoding, envelope filtering, repaired JSON parsing, and thinking/eager-input support.
- Added `repairJson`/`parseJsonWithRepair` and OAuth `postJson()` with `formatErrorDetails()` for richer stream/auth failures.
- Changed OpenAI Codex streaming to default verbosity to `low`, send `previous_response_id` on continuations, and record websocket debug stats.
- Added raw SSE and websocket tests for unknown trace events, malformed payload repair, continuation IDs, and debug counters.
- Defined ASI as an object with explicit `hypothesis`, `rollback_reason`, and `next_action_hint` fields while allowing additional keys.
- Updated `validateAsiRequirements` to clarify guidance when ASI data is missing or missing a valid hypothesis.
- Adjusted autoresearch state tests to assert the revised ASI validation error messages.
- Adjusted deepseek reasoning tests to use the AssistantMessage content type.
- Enabled synthetic reasoning-content support in openai-completions compatibility test fixtures for tool-call handling.
- Added the synthetic-reasoning flag to both openai completion compatibility and tool-result image test configs.
- Consolidated AI provider imports through register-builtins and moved Gemini/Antigravity header helpers to a shared module.
- Added lazy loading for heavy providers and SDK-backed modules with cached initialization to trim startup cost.
- Converted markdown conversion helpers to async and awaited htmlToBasicMarkdown in affected scraper and kernel output paths.
- Parsed bundled agent definitions on-demand and moved BrowserTool prompt rendering behind a memoized getter.
- Added cached validation/error handling paths by replacing AJV runtime checks with Value.Check and trimming validation error output.
DeepSeek V4 requires reasoning_content on EVERY assistant turn once any
prior turn included it, not just tool-call turns. Previously the replay
logic only triggered for tool-call turns, which caused 400 errors on
plain-text assistant follow-ups in multi-turn conversations.
Key change: the replay conditions now use needsReasoningField (all turns
for strict providers) instead of toolCalls.length > 0.
Codex review feedback: thinkingSignature can be an opaque value (Anthropic
encrypted signature, OpenAI Responses JSON item ID, etc.) not just a field
name. Using it as a property name on the wire message writes to an arbitrary
key and marks hasReasoningField=true, skipping both the empty-string fallback
and the synthetic placeholder — the outgoing message still misses
reasoning_content and triggers the same 400 on DeepSeek follow-ups.
Both the pre-existing nonEmptyThinkingBlocks path and the new Tier 1 path
now validate against recognized keys: reasoning_content, reasoning,
reasoning_text.
16 tests covering all three failure modes:
- reasoningEffortMap xhigh→max for DeepSeek-family on any provider
- allowsSyntheticReasoningContentForToolCalls flag detection
- Tier 1: signature recovery from empty thinking blocks
- Tier 2: empty-string fallback when no thinking blocks exist
- Tier 3: synthetic placeholder for non-DeepSeek providers (Kimi)
Updates existing issue-883 tests to match new behavior (empty string
instead of synthetic "." for DeepSeek).
Three independent failure modes observed in HTTP 400 logs:
1. reasoningEffortMap only mapped xhigh→max for the deepseek provider,
not for DeepSeek-family models on NVIDIA/OpenCode-Go/etc. Sending
reasoning_effort: "xhigh" to these endpoints caused 400 errors.
2. convertMessages filtered thinking blocks by nonEmptyThinkingBlocks,
excluding blocks with valid thinkingSignature but empty text. These
signatures identify the correct field name for reasoning_content
replay but were lost in the filter.
3. When a proxy (OpenCode-Go, NVIDIA) returns a tool-call response
without any reasoning_content at all, no thinking blocks exist to
recover from. Added empty-string fallback so the required field is
present even when no reasoning was captured.
Adds allowsSyntheticReasoningContentForToolCalls compat flag to
distinguish DeepSeek (rejects synthetic "." placeholder) from Kimi
and OpenRouter (accept it).
Credit: builds on the approach from PR #902 by @edmand46.
- Added a cached provider/model index with a WeakMap for resolver calls.
- Indexed each provider and model key to either a model entry or an ambiguity sentinel.
- Updated exact and fallback OpenRouter lookups to use map-based resolution and return undefined on ambiguous matches.
- Tracked `models.json` modification time in the registry to skip redundant static model reloads when unchanged.
- Reworked model overlay merges and package-runner detection to use indexed lookups plus parallel file/JSON scans instead of sequential searches.
- Cached compiled prompt templates and reduced startup work by bypassing up-to-date changelog parsing and deferring background model refresh.
- Replaced the eager `markit-ai` singleton import with a lazy async factory in `markit.ts`.
- Initialized and memoized a `Markit` instance on first use, with later calls reusing it.
- Updated conversion helpers to receive the initialized `Markit` instance and execute conversions through a shared runner.
- Updated AGENTS.md discovery to use glob search honoring .gitignore, depth limits, and deduped results.
- Updated eval tool flow so Python preflight runs only when needed and exec now maps to eval when available.
- Deferred canonical model-index rebuilds during refresh/rebuildProvider and replayed pending rebuilds after resume.
- Added memoized model-equivalence resolution with trailing-marker and canonical reference caches.
- Optimized frontmatter key normalization to keep unchanged keys/arrays/objects without extra cloning.
- Updated JS executor tests to use base-path concatenation for nested fixture filesystem calls.
The warmup path no longer produces prelude docs, so the cached-session
warmup it implemented added no value over the create-on-first-execute
path that withKernelSession already covers. Remove warmPythonEnvironment,
the backend warm() hook, the eval-tool warmup loop, the createTools
warmup preflight, and the forcePythonWarmup option. Simplify
ExecutorBackendCallOptions into ExecutorBackendExecOptions since execute
is now the only consumer.
Include the current ContextUsage payload in get_state responses so RPC clients can render context-window status without in-process AgentSession access.
Manual handoff starts a fresh session and seeds it with a displayed custom handoff message, not an assistant message. Session persistence normally waits for an assistant message before creating the session file, which made the new handoff session exist only in memory until later activity.
Persist the seeded handoff session after injecting the handoff context, and record the previous session file as its parent so lineage remains discoverable.
The Vercel AI Gateway discovery path was hitting
https://ai-gateway.vercel.sh/models, which currently returns 404. Vercel
only serves the catalog on /v1/models, so discovery silently failed and
generate-models fell back to the stale bundled snapshot. Recently
published models such as deepseek/deepseek-v4-pro and
deepseek/deepseek-v4-flash never made it into models.json under
vercel-ai-gateway.
Split the configured base URL into two roles:
- catalog URL is <base>/v1 so discovery requests <base>/v1/models
- runtime URL stays at <base> for the gateway's Anthropic-compatible
request path
A custom baseUrl ending in /v1 is normalized so discovery never produces
/v1/v1/models.
Regenerated packages/ai/src/models.json per AGENTS.md; the bundle diff
also picks up incidental drift from other catalog providers refreshed by
the same generator run.
Drives AgentSession + Agent + BashTool end-to-end through the patched
pi-natives binding. Asserts a spawned child reports its own session id
(setsid was called) and that pipelines still produce both stages'
output. Validated by reverting commands.rs and confirming the test
fails with a named diagnostic before restoring.
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
- Added `rename_file` LSP action with path checks, capped pair enumeration, and apply preview flow.
- Added `request` LSP action with auto-built params, optional JSON `payload`, and method-not-found handling.
- Added `capabilities` LSP action to inspect server capabilities for a file or all servers.
- Added schema updates and client file-operation capabilities to advertise rename support.
- Added regression tests for rename behavior, preview mode, and request/capabilities flows.
The signature in #computeAppliedToolSignature hashed the full raw
instructions string, but rebuildSystemPrompt truncates each server
instruction at 4000 chars before embedding it. A change past character
4000 produced an identical prompt but a different signature, causing a
spurious rebuild and a cache miss on every such reconnect.
Fix: hoist MAX_MCP_INSTRUCTIONS_LENGTH to module scope in sdk.ts and
apply the same truncation in the getMcpServerInstructions callback
before returning. The session now hashes exactly the strings that end
up in the prompt.
Regression test: changes only past char 4000 do not trigger a rebuild;
changes within the first 4000 chars do.
- Set a shared 60s Puppeteer protocol timeout and applied it to browser launch/connect.
- Added a combined cancel-or-timeout signal in `runInTab` and raced user code against it for prompt cancellation behavior.
- Reduced the browser tool max timeout to 30s to keep execution limits aligned with tool-level deadlines.
buildSystemPrompt injects today's date into the prompt body. The
applied-tool signature skips rebuilds when tools are byte-identical,
but did not cover the date — so a session spanning midnight with only
tool-stable MCP reconnects would keep yesterday's date indefinitely.
Append the current YYYY-MM-DD date as a suffix to the signature so any
reconnect after midnight triggers exactly one rebuild, then resumes
skipping normally for the rest of the new day.
- Documented and removed `utils/oauth` from the `ai` package entrypoint, noting it as a breaking change.
- Refactored `cli`, `auth-storage`, and `utils/oauth` to load provider modules via scoped dynamic `import()` calls.
- Removed top-level provider imports and barrel exports from `utils/oauth/index.ts`, streamlining oauth module loading.
- Consolidated OAuth symbol, type, and provider imports in coding-agent and tests to `@oh-my-pi/pi-ai/utils/oauth` modules.
- Defined `DEFAULT_LOCAL_TOKEN` locally in model-registry and removed its cross-package OAuth import usage.
Built-in tools whose prompt-rendered metadata depends on settings
(`TaskTool`, `SearchToolBm25Tool`, `EditTool`) expose `description`/
`label` via getters that re-evaluate on every access. The skip
optimization in `#applyActiveToolsByName` is correctness-safe for these
because `#computeAppliedToolSignature` reads `tool.description` live
each call, so a settings flip mutates the rendered string and differs
the signature on the next refresh.
This contract was implicit; a future refactor that caches per-tool
description strings would silently break it. Defending it explicitly:
- Added a regression test that wires a getter onto a CustomTool's
`description`, verifies `refreshMCPTools` skips while the underlying
state is unchanged, then mutates the state (without changing tool
object identity) and verifies the rebuild fires.
- Expanded the `#computeAppliedToolSignature` docstring to document
the getter-based coverage path and the SDK-init-time closure
constants in `sdk.ts` that genuinely cannot change at runtime
(`repeatToolDescriptions`, `eagerTasks`, `intentField`,
`mcpDiscoveryEnabled`, `secretsEnabled`).
Triggered by a review question on whether the skip breaks settings-
based prompt changes. It does not, but the property is non-obvious.
A tool's wire-visible name (`customWireName`) is rendered into the
system prompt body via `toolPromptNames`, but the applied-tool signature
only hashed name+label+description. A future tool whose wire name varied
without touching the other fields would silently produce a stale system
prompt that advertises the wrong callable name to the model — desyncing
prompt guidance from actual tool routing.
Today the only mutation path (edit-mode toggle) is also covered by a
description change and an explicit `refreshBaseSystemPrompt` from
`#syncEditToolModeAfterModelChange`, so this is a defensive fix rather
than a live bug. Including `customWireName` makes the signature a
self-consistent model of the prompt inputs.
- Extended `describeTool` in `#computeAppliedToolSignature` to include
`tool.customWireName ?? ""`. Applies to both the active-tool segment
and the (mcpDiscoveryEnabled) registry segment via the shared helper.
- Updated the docstring to call out wire-name coverage.
- Added a regression test that mutates `customWireName` between
identical-metadata refreshes and asserts the rebuild fires.
Per Codex review on #890.
Two cache-stability fixes for Anthropic prompt caching during MCP server
reconnects, which happen routinely (~5 min per server) in long sessions
due to SSE transport keepalive timeouts.
1) MCPManager: deterministic tool ordering
`#tools` is now sorted by name after every mutation. The previous
filter-out + push-to-end pattern in `#replaceServerTools` moved the
reconnecting server's tools to the end of the array, producing a new
byte order whenever the reconnect sequence differed from the initial
discovery sequence. With multiple healthy servers, each reconnect of
the non-last server flipped the order and invalidated the tools
cache breakpoint sent to Anthropic.
Sort applies in `discoverAndConnect` (initial population) and
`#replaceServerTools` (used by `reconnectServer` and
`refreshServerTools`). The comparator is character-code based,
locale-independent and deterministic. `sortMCPToolsByName` is
exported as a small generic helper and unit-tested.
2) AgentSession: skip system-prompt rebuild when inputs are unchanged
`#applyActiveToolsByName` (called from `refreshMCPTools` after every
reconnect) used to unconditionally call `rebuildSystemPrompt` and
`setSystemPrompt` even when the resulting prompt was byte-identical.
This wasted CPU on every flap and risked silent cache invalidation
if the rebuild path ever became non-deterministic.
Now `#applyActiveToolsByName` computes a stable signature of the
inputs `rebuildSystemPrompt` reads and skips the rebuild when the
signature matches the last successful one. The signature covers:
- active tool names in render order
- active tool labels and descriptions (rendered as `{{label}}:
\`{{name}}\`` in the prompt body)
- when MCP discovery is on, every registry tool's name + label +
description (the prompt summarizes discoverable-but-inactive
MCP tools)
- per-server MCP `instructions` text (embedded under "## MCP
Server Instructions" in the appended prompt; can change on
server upgrade while tool list stays identical)
Server instructions are read via a new optional
`getMcpServerInstructions` callback on `AgentSessionConfig`, wired
from the SDK as `() => mcpManager.getServerInstructions()`.
`refreshBaseSystemPrompt()` continues to rebuild unconditionally and
refreshes the cached signature, so explicit refreshes still pick up
ambient changes (edit-mode toggles, memory writes, etc.) that the
signature does not cover.
Signature inputs deliberately NOT covered: tool input schemas, memory
instructions read from disk, and other ambient state. Callers that
mutate those must call `refreshBaseSystemPrompt()` explicitly; existing
hooks (`#syncEditToolModeAfterModelChange`, memory hooks, `/clear`)
already do.