Commit Graph
991 Commits
Author SHA1 Message Date
can1357 6b4de15353 fix(coding-agent): retry sync git metadata reads on EINTR
POSIX permits short syscalls (incl. open/read/stat) to be interrupted
by a signal and surface as EINTR. Optional sync git metadata helpers
(used by the status line) propagated the raw error and crashed the
agent. Add a small bounded retry around the sync stat/read calls and
classify persistent EINTR as 'metadata unavailable' so optional reads
fall back to undefined instead of throwing.

Fixes #899
2026-05-02 07:55:30 +02:00
can1357 5d1ad6e80b fix(coding-agent): load extensions before --list-models
The --list-models handler in runRootCommand short-circuited to
listModels() right after Settings.init and modelRegistry.refresh,
exiting before extension loading ran in createAgentSession. As a
result, providers contributed via pi.registerProvider() (from -e
paths or settings.extensions) never appeared in the listing.

Extract a runListModelsCommand entry point in cli/list-models.ts
that loads extensions (CLI -e paths and settings.extensions) into
the supplied ModelRegistry, mirroring sdk.ts's handoff of pending
provider registrations, and then delegates to listModels. The load
is intentionally narrow: no agent loop, no MCP servers, no custom
tools.

Fixes #905
2026-05-02 07:48:30 +02:00
can1357 1c1d68f526 feat(coding-agent/tools): added contextual github render status messages
- Updated `githubToolRenderer.renderCall` to use operation-specific titles and contextual metadata.
- Fixed `githubToolRenderer.renderResult` fallback path to show the real operation and clearer status messages.
- Changed non-`run_watch` fallback output to truncate lines by terminal width and show `+N more lines` hints.
2026-05-02 07:44:58 +02:00
can1357 2264165567 feat: hashline v3 2026-05-02 07:01:57 +02:00
can1357 d064170c56 fix: chatgpt really doesnt like sandwitching lark 2026-05-02 04:54:56 +02:00
can1357 f0e9a830ff refactor(packages/coding-agent): migrated search args to paths arrays
- Switched search/ast_grep/ast_edit/find inputs from scalar path fields to required `paths` arrays.
- Reworked path resolution to normalize and expand each `paths` entry via explicit helpers in `src/tools/path-utils.ts`.
- Updated search and find tool-call rendering to display explicit `paths` values in mode overlays and export views.
- Updated tool prompt docs and examples to document `paths` array inputs for search, grep, find, and AST tools.
- Raised the default search match limit from 20 to 500 and updated limit-reached messaging.
2026-05-02 04:34:39 +02:00
can1357 ffdec04af2 refactor(packages/coding-agent): reorganized atom LID parsing rules
- Centralized `line-hash.ts` hash regex sources and resolved `atom.lark` via `resolveLarkLidPlaceholders`.
- Expanded `computeLineHash` to emit `>[a-z]` and `[a-z]<` hashes for brace-context anchors.
- Replaced atom/hashline parsers' hard-coded lid regex with shared `HASHLINE_HASH_RE_SRC` and lax counterparts.
- Removed `\\TEXT` continuation handling in atom rewrites and switched multi-line replacements to `+TEXT`.
- Added brace-body insertion warning when `@Lid` on `{`-ending lines inserts at non-body-safe indent.
- Suppressed duplicate auto-rebase warnings and kept unmatched `-`/`+` ranges separate in compact previews.
2026-05-02 04:34:39 +02:00
can1357 9fc762cbad refactor(packages/coding-agent): reorganized eval helper APIs in prelude
- Removed deprecated file helper APIs (find, glob, grep, rgrep, sed, and stat) from JS and Python eval preludes.
- Removed status icons and formatting branches for find/grep/rgrep/glob/stat/sed from tools/eval.ts.
- Updated eval helper docs and Python prelude tests to match the reduced exposed helper surface.
2026-05-02 04:34:39 +02:00
can1357 2826ec2ea0 refactor(scripts): restructured edit-mode fallbacks to ignore strict mode
- Removed PI_STRICT_EDIT_MODE gating from edit-mode resolution so model fallbacks now always apply.
- Stopped injecting PI_STRICT_EDIT_MODE in edit-benchmark.py and rate-edit-tool.py execution environments.
- Removed PI_STRICT_EDIT_MODE from environment-variable documentation and strict-mode test coverage.
2026-05-02 04:34:39 +02:00
can1357 31ae1b96fd feat(coding-agent/autoresearch): added branch-specific state restore
- Added branch-aware session loading so autoresearch state only rehydrates for current branch.
- Replaced user-specified experiment commands with fixed `bash autoresearch.sh` execution flow.
- Enforced safer setup checks, including missing `autoresearch.sh` and uncommitted-worktree errors.
- Added branch-specific storage helpers, baseline-commit persistence, and expanded tests for dirty-path cases.
2026-05-02 03:31:27 +02:00
can1357 0698ad3ba3 fix(coding-agent/core): allowed duplicate delete operations on atom anchor lines
- Adjusted atom mutation conflict validation to skip throwing on repeated delete edits for the same anchor line.
- Added tests confirming duplicate delete edits on one anchor are idempotent and do not trigger conflicts.
- Added coverage ensuring explicit deletes within replace ranges are treated as redundant and ignored.
2026-05-02 02:42:00 +02:00
can1357 d0db126a48 feat(ai): added disableReasoning option and fixed codex websocket reuse
- Added `disableReasoning` to `SimpleStreamOptions` and OpenAI completions, sending `reasoning: { enabled: false }` for OpenRouter requests to prevent reasoning models from consuming the full output budget on small calls like title generation.
- Fixed `canAppend` to accept `response.completed` as a terminal event, restoring websocket append reuse after codex sessions end.
- Replaced async blob-decoding and `addEventListener` with synchronous `onmessage`/`onopen`/`onerror`/`onclose` handlers and `binaryType = "nodebuffer"` for simpler, reliable message decoding.
- Simplified title generator to discard per-role thinking level and always pass `disableReasoning: true`.
2026-05-02 01:57:19 +02:00
can1357 7467add484 refactor(coding-agent): unified session title and accent handling around session names
- Removed title-source aware branching from session terminal-title and accent helpers, and updated callers to use session name plus cwd only.
- Dropped UUID-based recent-session naming by preferring explicit header titles or first user prompts and generating an "Untitled · <time>" fallback.
- Adjusted welcome session-row rendering for width-aware name truncation and disabled reasoning in title generation requests to keep terminal titles concise.
2026-05-02 01:42:06 +02:00
can1357 3293231a9f feat(coding-agent): implemented sqlite-backed storage for autoresearch
- Replaced file-backed autoresearch contracts with sqlite-backed session/run storage in `~/.omp/autoresearch`.
- Added `AutoresearchStorage` and rewired `init_experiment`, `run_experiment`, and `log_experiment` to persist sessions and runs.
- Added `update_notes` tool with `body`/`append_idea` inputs and updated prompts to use active-session context.
- Removed `autoresearch.md` contract parsing and checks flow, including `runChecks`, `force`, and timeout schema options.
- Updated autoresearch state/types to persist `goal`, `notes`, `branch`, and `baselineCommit` plus run justification/flag metadata.
2026-05-02 01:10:00 +02:00
can1357 2173619e99 fix(coding-agent): fixed multi-target search fanout for cross-tree paths
- Fixed path resolution to emit per-target fanout when common base collapses to filesystem root, preventing full-filesystem scans.
- Fixed match/file counts and pagination to aggregate correctly across all targets.
- Fixed returned paths to be normalized relative to the original search scope.
2026-05-02 00:54:26 +02:00
can1357 c6cb17e8d8 refactor(coding-agent/autoresearch): simplified ASI requirements validation
- Defined ASI as an object with explicit `hypothesis`, `rollback_reason`, and `next_action_hint` fields while allowing additional keys.
- Updated `validateAsiRequirements` to clarify guidance when ASI data is missing or missing a valid hypothesis.
- Adjusted autoresearch state tests to assert the revised ASI validation error messages.
2026-05-02 00:46:03 +02:00
Can BölükandGitHub 4a7b227b2a Merge pull request #895 from tvrmsmith/fix/brush-detach-when-embedded
fix(brush-core-vendored): detach child session whenever stdin isn't a terminal
2026-05-01 19:10:43 +02:00
Can BölükandGitHub 333510aced Merge pull request #890 from apoc/fix/mcp-tool-cache-stability
fix(coding-agent/mcp): stabilize tool ordering and skip redundant prompt rebuilds
2026-05-01 19:09:23 +02:00
Can BölükandGitHub 049bbd9483 Merge pull request #901 from fulara/pr/rpc-session-state
feat(coding-agent/rpc): expose session continuation state
2026-05-01 19:05:49 +02:00
can1357 ca84ee760d perf: optimized ai providers with lazy-loading imports and cached init
- Consolidated AI provider imports through register-builtins and moved Gemini/Antigravity header helpers to a shared module.
- Added lazy loading for heavy providers and SDK-backed modules with cached initialization to trim startup cost.
- Converted markdown conversion helpers to async and awaited htmlToBasicMarkdown in affected scraper and kernel output paths.
- Parsed bundled agent definitions on-demand and moved BrowserTool prompt rendering behind a memoized getter.
- Added cached validation/error handling paths by replacing AJV runtime checks with Value.Check and trimming validation error output.
2026-05-01 18:54:21 +02:00
can1357 83b4c0cdb5 perf(coding-agent): implemented canonical model-index rebuild during resume
- Updated AGENTS.md discovery to use glob search honoring .gitignore, depth limits, and deduped results.
- Updated eval tool flow so Python preflight runs only when needed and exec now maps to eval when available.
- Deferred canonical model-index rebuilds during refresh/rebuildProvider and replayed pending rebuilds after resume.
- Added memoized model-equivalence resolution with trailing-marker and canonical reference caches.
- Optimized frontmatter key normalization to keep unchanged keys/arrays/objects without extra cloning.
- Updated JS executor tests to use base-path concatenation for nested fixture filesystem calls.
2026-05-01 15:41:16 +02:00
can1357 6dccdf9193 refactor(coding-agent/eval): drop python warmup path
The warmup path no longer produces prelude docs, so the cached-session
warmup it implemented added no value over the create-on-first-execute
path that withKernelSession already covers. Remove warmPythonEnvironment,
the backend warm() hook, the eval-tool warmup loop, the createTools
warmup preflight, and the forcePythonWarmup option. Simplify
ExecutorBackendCallOptions into ExecutorBackendExecOptions since execute
is now the only consumer.
2026-05-01 15:27:40 +02:00
aleksander 86e83dc73c fix(coding-agent/session): persist handoff sessions immediately
Manual handoff starts a fresh session and seeds it with a displayed custom handoff message, not an assistant message. Session persistence normally waits for an assistant message before creating the session file, which made the new handoff session exist only in memory until later activity.

Persist the seeded handoff session after injecting the handoff context, and record the previous session file as its parent so lineage remains discoverable.
2026-05-01 10:05:41 +02:00
Trevor Smith e96719bc88 test(coding-agent): e2e cover brush embedded-host session detach
Drives AgentSession + Agent + BashTool end-to-end through the patched
pi-natives binding. Asserts a spawned child reports its own session id
(setsid was called) and that pipelines still produce both stages'
output. Validated by reverting commands.rs and confirming the test
fails with a named diagnostic before restoring.
2026-04-30 17:56:59 -05:00
can1357 76fe4b4995 test(coding-agent/eval): drop tests for removed helpers and stale warmup count 2026-04-30 18:44:07 +02:00
can1357 b0a31a5956 fix: outdated tests 2026-04-30 18:40:26 +02:00
can1357 cf60e6df51 feat(coding-agent): implemented eval framework and replaced python tool
- Added a unified eval framework with parser grammar, backend interfaces, and JS/Python execution result types.
- Added eval tool docs and updated prompts for fenced cells, `eval.py`/`eval.js`, and fallback behavior.
- Replaced the built-in `python` tool with `eval` across registry, rendering, interactive modes, and tool settings.
- Migrated Python execution runtime from `src/ipy` to `src/eval/py`, renamed state fields, and removed legacy introspection.
- Refactored browser tooling from in-process VM helpers to worker-managed tab supervisors and protocol transport.
- Added eval parser fallback and JS tool-bridge tests, updated imports, and removed obsolete python-mode suites.
2026-04-30 18:08:37 +02:00
can1357 d2c4958296 feat(coding-agent): added coding-agent lsp rename/request/capabilities
- Added `rename_file` LSP action with path checks, capped pair enumeration, and apply preview flow.
- Added `request` LSP action with auto-built params, optional JSON `payload`, and method-not-found handling.
- Added `capabilities` LSP action to inspect server capabilities for a file or all servers.
- Added schema updates and client file-operation capabilities to advertise rename support.
- Added regression tests for rename behavior, preview mode, and request/capabilities flows.
2026-04-30 16:42:52 +02:00
Miroslav Drbal 484e0e815f fix(coding-agent/mcp): truncate server instructions in getMcpServerInstructions callback
The signature in #computeAppliedToolSignature hashed the full raw
instructions string, but rebuildSystemPrompt truncates each server
instruction at 4000 chars before embedding it. A change past character
4000 produced an identical prompt but a different signature, causing a
spurious rebuild and a cache miss on every such reconnect.

Fix: hoist MAX_MCP_INSTRUCTIONS_LENGTH to module scope in sdk.ts and
apply the same truncation in the getMcpServerInstructions callback
before returning. The session now hashes exactly the strings that end
up in the prompt.

Regression test: changes only past char 4000 do not trigger a rebuild;
changes within the first 4000 chars do.
2026-04-30 16:28:07 +02:00
Miroslav Drbal b8e43e7603 fix(coding-agent/mcp): include calendar date in applied-tool signature
buildSystemPrompt injects today's date into the prompt body. The
applied-tool signature skips rebuilds when tools are byte-identical,
but did not cover the date — so a session spanning midnight with only
tool-stable MCP reconnects would keep yesterday's date indefinitely.

Append the current YYYY-MM-DD date as a suffix to the signature so any
reconnect after midnight triggers exactly one rebuild, then resumes
skipping normally for the rest of the new day.
2026-04-30 15:30:10 +02:00
can1357 3a46b748a3 perf(ai): dynamic oauth imports per provider
- Documented and removed `utils/oauth` from the `ai` package entrypoint, noting it as a breaking change.
- Refactored `cli`, `auth-storage`, and `utils/oauth` to load provider modules via scoped dynamic `import()` calls.
- Removed top-level provider imports and barrel exports from `utils/oauth/index.ts`, streamlining oauth module loading.
- Consolidated OAuth symbol, type, and provider imports in coding-agent and tests to `@oh-my-pi/pi-ai/utils/oauth` modules.
- Defined `DEFAULT_LOCAL_TOKEN` locally in model-registry and removed its cross-package OAuth import usage.
2026-04-30 15:14:07 +02:00
Miroslav Drbal 792b799e16 test(coding-agent/mcp): defend getter-based tool descriptions against signature regression
Built-in tools whose prompt-rendered metadata depends on settings
(`TaskTool`, `SearchToolBm25Tool`, `EditTool`) expose `description`/
`label` via getters that re-evaluate on every access. The skip
optimization in `#applyActiveToolsByName` is correctness-safe for these
because `#computeAppliedToolSignature` reads `tool.description` live
each call, so a settings flip mutates the rendered string and differs
the signature on the next refresh.

This contract was implicit; a future refactor that caches per-tool
description strings would silently break it. Defending it explicitly:

- Added a regression test that wires a getter onto a CustomTool's
  `description`, verifies `refreshMCPTools` skips while the underlying
  state is unchanged, then mutates the state (without changing tool
  object identity) and verifies the rebuild fires.
- Expanded the `#computeAppliedToolSignature` docstring to document
  the getter-based coverage path and the SDK-init-time closure
  constants in `sdk.ts` that genuinely cannot change at runtime
  (`repeatToolDescriptions`, `eagerTasks`, `intentField`,
  `mcpDiscoveryEnabled`, `secretsEnabled`).

Triggered by a review question on whether the skip breaks settings-
based prompt changes. It does not, but the property is non-obvious.
2026-04-30 15:05:01 +02:00
Miroslav Drbal 08cb3f8555 fix(coding-agent/mcp): include customWireName in applied-tool signature
A tool's wire-visible name (`customWireName`) is rendered into the
system prompt body via `toolPromptNames`, but the applied-tool signature
only hashed name+label+description. A future tool whose wire name varied
without touching the other fields would silently produce a stale system
prompt that advertises the wrong callable name to the model — desyncing
prompt guidance from actual tool routing.

Today the only mutation path (edit-mode toggle) is also covered by a
description change and an explicit `refreshBaseSystemPrompt` from
`#syncEditToolModeAfterModelChange`, so this is a defensive fix rather
than a live bug. Including `customWireName` makes the signature a
self-consistent model of the prompt inputs.

- Extended `describeTool` in `#computeAppliedToolSignature` to include
  `tool.customWireName ?? ""`. Applies to both the active-tool segment
  and the (mcpDiscoveryEnabled) registry segment via the shared helper.
- Updated the docstring to call out wire-name coverage.
- Added a regression test that mutates `customWireName` between
  identical-metadata refreshes and asserts the rebuild fires.

Per Codex review on #890.
2026-04-30 15:05:01 +02:00
Miroslav Drbal b5ca55e79f fix(coding-agent/mcp): stabilize tool ordering and skip redundant prompt rebuilds
Two cache-stability fixes for Anthropic prompt caching during MCP server
reconnects, which happen routinely (~5 min per server) in long sessions
due to SSE transport keepalive timeouts.

1) MCPManager: deterministic tool ordering

   `#tools` is now sorted by name after every mutation. The previous
   filter-out + push-to-end pattern in `#replaceServerTools` moved the
   reconnecting server's tools to the end of the array, producing a new
   byte order whenever the reconnect sequence differed from the initial
   discovery sequence. With multiple healthy servers, each reconnect of
   the non-last server flipped the order and invalidated the tools
   cache breakpoint sent to Anthropic.

   Sort applies in `discoverAndConnect` (initial population) and
   `#replaceServerTools` (used by `reconnectServer` and
   `refreshServerTools`). The comparator is character-code based,
   locale-independent and deterministic. `sortMCPToolsByName` is
   exported as a small generic helper and unit-tested.

2) AgentSession: skip system-prompt rebuild when inputs are unchanged

   `#applyActiveToolsByName` (called from `refreshMCPTools` after every
   reconnect) used to unconditionally call `rebuildSystemPrompt` and
   `setSystemPrompt` even when the resulting prompt was byte-identical.
   This wasted CPU on every flap and risked silent cache invalidation
   if the rebuild path ever became non-deterministic.

   Now `#applyActiveToolsByName` computes a stable signature of the
   inputs `rebuildSystemPrompt` reads and skips the rebuild when the
   signature matches the last successful one. The signature covers:
     - active tool names in render order
     - active tool labels and descriptions (rendered as `{{label}}:
       \`{{name}}\`` in the prompt body)
     - when MCP discovery is on, every registry tool's name + label +
       description (the prompt summarizes discoverable-but-inactive
       MCP tools)
     - per-server MCP `instructions` text (embedded under "## MCP
       Server Instructions" in the appended prompt; can change on
       server upgrade while tool list stays identical)

   Server instructions are read via a new optional
   `getMcpServerInstructions` callback on `AgentSessionConfig`, wired
   from the SDK as `() => mcpManager.getServerInstructions()`.

   `refreshBaseSystemPrompt()` continues to rebuild unconditionally and
   refreshes the cached signature, so explicit refreshes still pick up
   ambient changes (edit-mode toggles, memory writes, etc.) that the
   signature does not cover.

Signature inputs deliberately NOT covered: tool input schemas, memory
instructions read from disk, and other ambient state. Callers that
mutate those must call `refreshBaseSystemPrompt()` explicitly; existing
hooks (`#syncEditToolModeAfterModelChange`, memory hooks, `/clear`)
already do.
2026-04-30 15:04:33 +02:00
Can BölükandGitHub bbff8d5aff Merge pull request #850 from cgrossde/fix/strict-tools-non-native-endpoints
feat(coding-agent): add disableStrictTools provider option for anthropic-messages endpoints
2026-04-30 14:02:13 +02:00
Christoph Gross f908ee9496 feat(coding-agent): add disableStrictTools provider option for anthropic-messages endpoints
Exposes model.compat.disableStrictTools (already supported by the anthropic
transport since #826) via models.yml so users can configure it without code
changes.

Set disableStrictTools: true at the provider level to disable strict tool
schemas for third-party Anthropic-compatible endpoints (AWS Bedrock, Vertex
AI proxies, custom gateways) that reject the strict field.

- Add disableStrictTools to ProviderConfigSchema
- Merge { disableStrictTools: true } into provider compat override when set,
  flowing through the existing compat pipeline to model.compat.disableStrictTools
- disableStrictTools alone is sufficient for an override-only provider entry
- Update docs/models.md with field reference, Bedrock example, and proxy note
- Add tests covering provider-level propagation, built-in override, and
  overlay merge
2026-04-30 08:28:03 +02:00
can1357 729db0310f fix(coding-agent/tools): fixed plan mode to redirect PLAN.md writes to local plan artifact
- Added plan-mode path resolution that routed bare and cwd-relative PLAN.md targets to the canonical local plan artifact when plan mode was enabled.
- Preserved existing resolution for non-matching basenames and disabled-plan-mode sessions by returning the raw path unchanged.
- Expanded local plan-mode guard tests to cover redirects for PLAN.md and enforcement of non-plan and delete restrictions.
- Updated atom edit prompt guidance to remove obsolete small-drift hash-auto-rebase instructions.
2026-04-30 07:23:27 +02:00
can1357 5003996a01 feat(coding-agent): added content-based todo matching for write/commands
- Removed `id` fields from todo models/fixtures and switched session clones to content-based task identity.
- Replaced `/todo_write` `replace` with `init`, updated setup schemas to `list`/`phase`, and append content-only items.
- Updated `/todo` command flows to match phases and tasks by names/content (exact/prefix/substr, case-insensitive), with no ID targeting.
- Updated rendering/output labels to `# Todos`, `formatPhaseDisplayName`, and Roman-numeral phase headings across todo views.
- Aligned prompts, changelog, and todo tests/fixtures with the new init and content-based todo-write contract.
2026-04-30 06:40:53 +02:00
can1357 bbd789fe40 fix(coding-agent): added blank-line forgiveness for atom range replacements
- Added a lookahead helper to detect `\` continuation lines after blank lines during range replacements.
- Converted interior blank lines to explicit continuation-sentinel inserts when followed by a continuation, preserving open replacement state.
- Extended atom parser tests to cover implicit blank-`\` handling for range and single-anchor replacements, including trailing-blank termination.
2026-04-30 06:06:07 +02:00
can1357 4562b3f365 feat: compact diffs, remove garbage hint 2026-04-30 06:03:29 +02:00
can1357 55bdb4119f fix(coding-agent): added auto-retry handling for unexpected socket close errors
- Integrated `isUnexpectedSocketCloseMessage` into transient error detection so Bun socket-closure failures are treated as retryable.
- Added a retry fallback test that simulates a Bun socket close error and verifies the request is retried successfully with matching retry start/end events and recovered output.
2026-04-30 05:58:38 +02:00
can1357 961a7f053b fix(coding-agent/edit): extended duplicate auto-fix to remove multiple adjacent duplicate lines
- Updated adjacent-duplicate auto-fix logic to remove a line from each detected pair when bracket balance changed.
- Adjusted the validation to accept the corrected file only after all removals restore original bracket balance.
- Added a regression test for one edit creating duplicate block closers in two unrelated segments and asserting the auto-fix warning.
2026-04-30 05:58:08 +02:00
can1357 5ecd041bfe fix: patched bash interceptor, LSP shutdown, and concurrent command tracking
- Fixed bash interceptor to check both raw and cwd-normalized commands, catching commands hidden behind leading `cd ... &&` wrappers.
- Fixed LSP client shutdown to await graceful shutdown with a 5s timeout before killing the process, and parallelized `shutdownAll` via `Promise.allSettled`.
- Fixed concurrent bash command tracking by replacing a single abort controller with a Set, preventing premature cancellation of parallel commands.
- Removed `./hooks` and `./hooks/*` export entries from the coding-agent package exports map.
- Updated pinned Rust nightly toolchain from `nightly-2026-03-27` to `nightly-2026-04-29` in `rust-toolchain.toml` and CI workflow.
- Replaced custom already-published detection in `ci-release-publish.ts` with `bun publish --tolerate-republish` flag.
2026-04-30 05:51:01 +02:00
can1357 bf1faf8842 test(coding-agent): drop api filter from getOpenAICompat fixture helper
The helper guarded `model.api === "openai-completions"` and returned undefined
for openai-responses models. The discoverable-custom-compat test sets
`api: "openai-responses"` on a custom model with `compat.extraBody`, so the
post-refresh assertion saw `undefined` instead of the configured proxy hint.

The OpenAICompatSchema gates user-facing custom-model compat regardless of the
underlying api wire format, so reading the field as OpenAICompat for any api
matches what the registry actually stores.
2026-04-30 05:35:00 +02:00
can1357 fed95ce524 fix(ai,coding-agent): narrow Model.compat consumers after AnthropicCompat split
Commit a190397d8 made `Model.compat` resolve to `OpenAICompat | AnthropicCompat`
under the default `TApi = any`. The widened union broke every site that treated
`compat` as openai-shaped: model-registry deep-merge, openai-completions resolved
compat, and ~20 test fixtures. This restores the assumption locally instead of
papering over it with casts.

- getBundledModel is now generic on TApi so test fixtures that spread it into
  `Model<"openai-completions">` get the narrow compat back.
- mergeCompat is generic over TBase/TOverride; the schema-driven model-registry
  override path keeps its OpenAICompat-shaped merge fields, anthropic overrides
  pass through untouched.
- OpenAICompatSchema gains the openai-only fields it was missing
  (requiresMistralToolIds, reasoningContentField, requiresReasoningContent*,
  thinkingFormat, requiresThinkingAsText, disableReasoningOnForcedToolChoice).
- resolveOpenAICompat fills in disableReasoningOnForcedToolChoice so the
  Required<OpenAICompat> shape stays satisfied.
- Anthropic tool-result block id assignment uses the proper unknown double-cast.
- isForcedToolChoice accepts unknown so it can read `params.tool_choice` whose
  type comes from the OpenAI SDK ChatCompletionToolChoiceOption (now wider than
  our local OpenAICompletionsToolChoice).
- Test fixtures and Required<OpenAICompat> literals updated for the field set.

Fixes CI red on main.
2026-04-30 05:23:20 +02:00
can1357 d9376c5866 fix(coding-agent): keep steer preview honest when post-compaction flush hits AgentBusyError
flushCompactionQueue() fires session.prompt(text) on the first non-slash
queued message. If the session is still streaming when compaction-end
lands (race between isStreaming flipping false and the event arriving),
prompt() throws AgentBusyError, restoreQueue() dumps the message back
into compactionQueuedMessages, and it stalls there: nothing drains that
array except the next compaction-end. The user sees the steer preview
but cannot deliver the message (Alt+Up consults session.clearQueue, not
compactionQueuedMessages).

Pass streamingBehavior derived from the queued message's mode so
prompt() routes into the steer/follow-up queue when busy and runs as a
fresh prompt when idle.

Fixes #825
2026-04-30 04:51:23 +02:00
can1357 08403be71f fix(coding-agent): preserve explicit default model on session resume
buildSessionContext walked the entry path and unconditionally overwrote
models.default from every assistant message's reported model. Temporary
fallbacks (retry fallback, context promotion) and codex-side model
downgrades both produce assistant messages tagged with a different model
id, which clobbered the user's explicit /model pick on resume and made
the session silently revert to the older model.

Treat assistant-message inference as a legacy fallback that only fills
in models.default when no explicit `model_change` with role="default"
has been seen on the path.

Fixes #849
2026-04-30 04:51:23 +02:00
can1357 381f30233e fix(coding-agent): log stage1 memory job failures
Phase1 caught every per-claim failure, recorded the reason in
jobs.last_error, and surfaced only an aggregate failed count via the
phase1 completion debug line. Users hitting setup-time failures (e.g.
WSL2 stale rollout paths producing ENOENT before any LLM call) had no
diagnostic. Emit logger.error per failed claim with threadId,
rolloutPath, and reason so the actual error is visible in omp.log.

Fixes #846
2026-04-30 04:51:23 +02:00
can1357 0306f937ea feat(ai): support per-model thinking defaultLevel
Add optional defaultLevel to ThinkingConfig schema/type so models.yml can
declare a preferred starting thinking level per model. On model switch
the agent session adopts model.thinking.defaultLevel when present (with
explicit caller-supplied level still winning); otherwise current behavior
is preserved. SDK initial selection prefers the model's defaultLevel
before falling back to the global defaultThinkingLevel setting.

Fixes #775
2026-04-30 04:51:23 +02:00
can1357 895ff6f3a7 fix(coding-agent): clear pending plan-role model switch on plan-mode exit
When entering plan mode while the session is streaming, #applyPlanModeModel
defers the switch into #pendingModelSwitch and snapshots the previous model.
On exit, the snapshot was restored but the deferred switch was left queued,
so the next agent_end flush landed the session on the plan-role model after
the user had already left plan mode.

Drop the pending switch in #exitPlanMode when its target matches the
plan-role resolution; leave any other queued switch alone.

Fixes #816
2026-04-30 04:51:23 +02:00