When the Codex Responses API synthesizes an answer without emitting
url_citation annotations, previously-empty sources made cited results
look ungrounded. Add:
- tool_choice: { type: "web_search" } so Codex must call the tool
- markdown-link + bare-URL extraction from the answer as a fallback
that only runs when no structured citations were returned
The model-resolution path is unchanged; getBundledModels already
returns a usable Codex catalog on main.
Closes#724
- Register claude-opus-4-7 model entry in models.json
- Cover adaptive thinking/sampling payload shape in anthropic-alignment test
- Assert Opus 4.7 surfaces in ModelRegistry available models
The SQLite read helper is documented as a structured selector path
with explicit pagination, while raw SQL already has a separate
q=SELECT escape hatch. This change rejects SQL control syntax outside
quoted strings so helper filters cannot override LIMIT/OFFSET, while
still allowing semicolons inside quoted literals such as LIKE '%;%'.
Constraint: Preserve the documented q=SELECT raw SQL path unchanged
Rejected: Replace where= with a new filter DSL | too broad for a regression fix
Confidence: high
Scope-risk: narrow
Directive: Keep table?where=... as a structured helper; if broader SQL is needed, route it through q=SELECT instead
Tested: bun --cwd=packages/natives run build; bun --cwd=packages/coding-agent run check; bun --cwd=packages/coding-agent test test/tools/sqlite.test.ts
Not-tested: Manual interactive omp read invocation against a live SQLite file
The previous mismatch text ('1 line has changed since last read. Use
the updated LINE#ID references shown below.') reads like a successful
edit followed by an informational note. Providers whose tool-result
plumbing does not surface `isError: true` to the model (e.g. Qwen via
the OpenAI-completions shim, where mitmproxy traces show the model only
receives the content text) treat the response as a success and proceed
with stale state.
Reword the message to start with 'Edit rejected:' and explicitly state
'The edit was NOT applied' so the failure is unambiguous in plaintext.
The structural payload (>>> markers, updated LINE#IDs) is unchanged.
Fixes#742
Mirror the existing `commands.enableClaudeUser`/`commands.enableClaudeProject`
schema entries so the OpenCode discovery provider exposes the same
user/project toggle surface as Claude. Default remains true to preserve
current behavior.
Fixes#661
The structured SQLite helper interpolates `where=` directly into SQL.
A crafted clause like `where=1=1 LIMIT 1000000 --` could comment out
the helper's bound `LIMIT ? OFFSET ?`, returning the full table in
violation of the documented pagination contract.
Validate where= at the selector boundary and reject SQL comments,
statement terminators, and pagination/attach/pragma keywords. Raw SQL
remains available via ?q=SELECT... for callers that need it.
Fixes#735
Extracts a single `formatMatchPath` helper used by both the
fast-glob/native code paths and the streaming onMatch callback so
relative paths, trailing-slash handling, and directory markers are
produced consistently. Also drops the retry-without-gitignore fallback
when the gitignored pass returns zero matches, so a broad hidden-file
pattern that is fully ignored stays fully ignored instead of silently
flipping gitignore off on the second attempt.
resolveMultiSearchPath now reports `exactFilePaths` when every token
resolves to a plain file (no globs, no suffix glob) and accepts a
single resolvable token so partially-missing lists still search the
resolvable subset. grep iterates those exact files individually
instead of collapsing them into a brace-union glob, which preserves
the user's explicit file set even when siblings share a basename.
Also adds a small `[grep] match lines use ':'; context lines use '-'`
banner when context lines are rendered, and splits the per-file
rendering helpers so files with no remaining matches no longer emit
empty headers.
Apply now recomputes per-file replacement counts from the actual apply
pass and compares them against the preview. If totals or per-file
counts drift (file changed between preview and apply, or apply matched
nothing), the tool returns an isError result explaining that the
preview is stale instead of silently claiming success with mismatched
numbers.
When a bash tool call requests a timeout outside the allowed 1-3600s
range, the effective clamped value and the originally requested value
are now emitted as a notice appended to the tool output and exposed on
BashToolDetails via requestedTimeoutSeconds. The renderer shows the
clamped+requested pair inline in the timeout badge.
- Queued outbound JSON-RPC messages behind a per-client promise queue to serialize writes.
- Added a project-load gate for project-aware LSP operations before diagnostics and reference lookups.
- Retried declaration-only references with a short delay until project metadata is available, then proceeded with normal results.
- Relaxed chunk-mode parameter validation to accept `{path}`-only edits as valid delete operations.
- Updated the invalid-parameters help text to document accepted chunk delete payloads.
- Added a test that verifies a bare `{path}` edit removes the targeted chunk when null values are stripped.
- Added a streaming parser path for apply_patch envelopes that tolerates incomplete patch bodies.
- Updated apply patch preview expansion to return best-effort hunks when the renderer is in partial mode.
- Added a renderer test confirming streaming apply_patch input shows file paths without end-marker parse errors.
- Added built-in model entries for gpt-5.5 and gpt-image-2, including updated context windows, token limits, and pricing.
- Updated generated-model policy application to set or clear applyPatchToolType based on inferred GPT-5 freeform rules.
- Changed Spark edit-mode resolution to return apply_patch by default, honoring explicit replace and strict-mode overrides.
- Added tests for GPT-5 freeform policy inference and Spark edit-mode default, variant, and strict-mode behavior.
Slots a new "apply_patch" variant alongside the existing edit modes
(replace, patch, hashline, chunk, vim). The mode accepts a single input
string containing a Codex *** Begin Patch / *** End Patch envelope,
parses it with a new lenient parser (heredoc-tolerant), and fans each
file-op out to the existing executePatchSingle so LSP writethrough,
plan-mode guards, fs-cache invalidation and diagnostics are shared
with the patch mode.
Exposes both tool shapes from the spec: the JSON function-tool variant
(§1.2, {input: string}) and the OpenAI custom-tool / Lark-grammar
"freeform" variant (§1.1, raw patch string). The edit tool advertises
a Lark grammar via customFormat and a wire name via customWireName;
openai-responses emits it as a grammar-constrained custom tool when a
model opts in with applyPatchToolType: "freeform" in models.json.
custom_tool_call / custom_tool_call_output are plumbed end-to-end
through the shared responses code (emission, streaming, history
replay), and the agent-loop dispatcher matches tool calls by either
name or customWireName so returned calls route correctly.
Also threads preview/diff rendering for apply_patch through the TUI
(tool-execution + edit renderer) so streaming patches show per-file
diffs like the other edit modes.
Default edit mode is unchanged (hashline); opt in via edit.mode or
PI_EDIT_VARIANT=apply_patch.
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
- Added a new read.toolResultPreview boolean setting defaulting to false to control inline read result rendering.
- Updated read tool group rendering to show inline previews only when enabled and to hide duplicate summary rows for previewed entries.
- Passed the setting through event and UI helper constructors and added tests for default-off behavior and the duplicate-preview summary case.
- Refactor path normalization to combine expandPath and normalizeLocalScheme
- Add validation in utils.ts to reject local:// paths as filesystem paths
- Fix bash-skill-urls regex to handle hyphen-prefixed local:/ patterns
- Add tests for hyphen-prefixed and @local: patterns
- Add negative lookbehind to regex in bash-skill-urls to prevent matching local:/
inside paths like /repo/local:/PLAN.md
- Normalize local scheme before expanding paths in path-utils
- Add test cases for both changes
Expands the regex pattern to match local:/ (single-slash) URLs in addition to local:// (triple-slash), preventing potential Linux path leaks.
- Add regex patterns for single-quoted, double-quoted, and unquoted local:/ URLs
- Add test coverage for all three quote styles
Removed unconditional topic:news coupling in buildRequestBody that scoped
Tavily index to news publications whenever recency was set. Technical queries
with --recency now search the general index filtered by time only.
Tightened SearchParams.recency contract in base.ts: providers MUST interpret
recency as a pure time filter and MUST NOT change topic scope as a side effect.
Three fixes to make CI green after the opus 4.7 and auto-bump landed:
1. github-copilot model mapper: prefer capabilities.limits.max_prompt_tokens
over the root-level context_length field (which mirrors max_context_window_tokens, i.e.
total window). Copilot's real /models response returns both for the gpt-5.x family, and
context_length inflates contextWindow with the output budget. Also restore the bundled
Copilot limits (claude-opus-4.6, gpt-5.2, gpt-5.4, gpt-5.4-mini, grok-code-fast-1) to
the values the fixed mapper produces so tests that depend on truthful offline fallbacks
pass. Update the two Copilot discovery tests whose payloads conflated context_length
with prompt capacity.
2. coding-agent task schema: make the per-task assignment description context-mode-aware.
The previous description unconditionally told agents that 'shared background belongs
in context', which is wrong for independent mode where shared context is disabled.
3. coding-agent model-registry test: update the anthropic-latest canonical collapse case
to claude-opus-4-7 since opus 4.7 is now the newest official opus in models.json.
- Updated todo start handling to set the requested task to in_progress while demoting all other in_progress tasks to pending.
- Added task note rendering in summary output by prefixing each note line with "Note:".
- Set GIT_OPTIONAL_LOCKS to 0 in git execution options and added tests for out-of-order start jumps and note summaries.
The hashline prefix stripping logic uses an "all-or-nothing" heuristic:
it only strips N#XX: anchors when every non-empty line carries one.
When the Read tool truncation notice (e.g. [Showing lines 1-300 of
332. Use sel=L301 to continue]) is present in content passed to the
Edit or Write tool, that marker line breaks the heuristic, causing ALL
anchors to be written to disk verbatim. On subsequent reads, those
persisted anchors get double-prefixed (1#NX:1#BQ:---), compounding
the corruption.
This can happen when the AI model pastes Read tool output (including
the truncation marker) back into an Edit/Write call. While this is
arguably a model-level issue, the toolchain should be resilient against
it rather than silently corrupting files.
Changes:
- Add READ_TRUNCATION_NOTICE_RE to detect Read tool truncation markers
- Update collectLinePrefixStats to exclude truncation marker lines from
the non-empty line count so they no longer break the stripping heuristic
- Add stripLeadingHashlinePrefixes helper for recursive stripping of
nested/double-prefixed anchors (handles already-corrupted content)
- Add filterTruncationNotices helper to remove truncation marker lines
- Update both stripNewLinePrefixes and stripHashlinePrefixes to filter
truncation markers and use recursive anchor stripping
- Add 5 new test cases covering truncation markers and nested anchors
- Replaced todo_write's `ops` payload with top-level mutation fields (`phases`, `complete`, `start`, `add_notes`, etc.).
- Removed in-place task content and note updates; moved note writes to append-only `add_notes` calls.
- Matched `add_tasks.phase` by phase ID or name and appended new note text to existing notes.
- Changed execution to drop strict in-progress update sequencing, support `start`, and auto-promote next pending task.
- Updated tests and prompt docs to use direct payloads (`phases`, `complete`, `add_tasks`) instead of `ops`.
- Reworked `idle-iterator` to replace the first-event watchdog wrapper with a shared timeout handle passed through stream iteration.
- Updated Anthropic and OpenAI/Azure response-completion streams to pass the watchdog into `iterateWithIdleTimeout` and rely on unified first-event timeout logic.
- Raised the default first-stream-event timeout from 60 seconds to 100 seconds.