- Updated the todo-write prompt to require phase names to use roman-numeral ordinals and reject nonconforming identifiers.
- Refreshed the sample todo JSON invocations to use prefixed phase names such as `I. Foundation`, `II. Auth`, and `III. Verification`.
- Aligned the todo-write phase schema examples with the new roman-numeral phase naming convention.
- Added a new `async.pollWaitDuration` setting with enum values for 5s to 5m and defaulted it to 30s.
- Updated PollTool to parse the configured wait duration and include a timeout in the job-promise race before returning.
- Revised the poll tool prompt to describe timeout return behavior and discourage indefinite unproductive polling.
- Removed the default interleaved-thinking beta and only added it when interleaved thinking is requested on models without adaptive thinking display support.
- Stopped setting request temperature in Anthropic params and simplified Opus 4.7 API cleanup by dropping its removal branch.
- Passed `disableStrictTools || model.provider === "github-copilot"` when converting tools to disable strict mode for Copilot models.
- Replaced service-tier gating in OpenAI provider request builders with provider-aware checks before writing params.service_tier.
- Updated the shared service-tier helper to accept provider context and only return true for flex, scale, or priority on openai/openai-codex providers.
- Removed built-in `taplo` from `src/lsp/defaults.json`, so TOML files no longer start a default LSP server.
- Condensed many `defaults.json` arrays and option blocks to a compact single-line style.
- Session metadata now stores on-disk size bytes, populated from file stats when collecting sessions.
- The session selector metadata line now renders file size using formatBytes instead of message counts.
- Session test fixtures were updated with the new size property so they match the revised SessionInfo shape.
- Added `sed` atom editing with `g`, `i`, and `F` flags, path/anchor parsing, and conflict handling updates.
- Changed hashline and grep/read output to `LINE+ID|content` with `>` match prefixes and `:` context prefixes.
- Fixed atom anchor parsing for path-qualified locs, hyphenated single anchors, and content hints after `|` or `:`.
- Updated prompts, changelog, and tests to document and validate the new hashline and `sed` formats.
- In `normalizeOptionalNullsForSchema`, unknown object keys were removed when their value was `null` or `"null"` and `additionalProperties` was false.
- Non-null unknown fields were left intact so malformed extra properties still triggered validation errors.
- Consolidated all former gh_* tools into a single GithubTool that routes execution by a required op field.
- Replaced gh_* tool registrations and render dispatch with `github`, including renderer key and header updates.
- Removed deprecated gh-* tool prompt files and added a unified github.md prompt covering per-op inputs and outputs.
- Updated settings-schema and ci-green prompt guidance to reference the unified github tool and `run_watch` operation.
- Updated tests and tool imports to use `GithubTool`/`githubToolRenderer` with op-based payloads and assertions.
- Computed an enforceProRequirement flag per session so Pro filtering is skipped when no Pro candidate exists.
- Updated credential selection and usage validation to apply the Pro plan gate only when that flag is enabled.
- Adjusted Codex Spark OAuth tests to verify Plus accounts are selected when no Pro account is connected.
- Removed the atom tool's `sub` verb from its schema, parser, conflict checks, and edit application.
- Updated atom documentation, examples, and tests to reject `sub`-based replacements in favor of `set` replacements.
- Adjusted Anthropic strict-tool handling by dropping `write` from the allowlist and updating alignment expectations.
- Added auto-rebasing for stale atom and hashline anchors within ±2 lines, with warning diagnostics.
- Removed atom range-locator support and dropped `between` ops, updating docs/tests for `sub` over `set` nudges.
- Changed no-op handling to track unchanged hashline edits and emit contextual hints for unchanged ranges.
- Expanded Anthropic strict error handling to retry on schema-too-complex and compiled-grammar-too-large errors.
- Added anchor-retargeting warning checks and updated edit tests to expect locator and range-locator rejections.
- Updated python tool-call fixtures by adding `title` to executed-cell payloads and adjusting related test expectations.
- Detected transport failures by checking failed run errors for "Timeout exhausted".
- Counted those failures in benchmark summaries and included them in ghost-like run classification, reducing effective run totals accordingly.
- Updated benchmark reporting to display excluded transport-failure runs when present.
- Changed the todo-write parameter schema from a bare operations array to an object containing an `ops` array.
- Updated request parsing and renderer logic to consume `params.ops` and `args.ops` when processing todo operations.
- Updated session and unit tests to invoke todo_write with wrapped `{ ops: [...] }` payloads.
- Limited strict tool candidates to a named allowlist instead of all opt-in tools.
- Fixed `minItems` stripping to target object-typed schema nodes, not just arrays.
- Enabled strict mode on all remaining coding-agent tools now that the allowlist guards eligibility.
- Adjusted hashline tests to expect colon-prefixed anchors in format, diff preview, and parse assertions.
- Updated ast-edit test regexes and split logic to parse colon-delimited hashline prefixes.
- Applied HASHLINE_CONTENT_SEPARATOR in AstEditTool output so hashline entries share a shared delimiter.
- Set AstEditTool and AstGrepTool strict defaults to false for more permissive tool call input handling.
- Added `HASHLINE_CONTENT_SEPARATOR` and switched hashline formatting to colon separators across hashline helpers.
- Updated hashline prefix regexes in `edit/modes/hashline.ts` to require `:` and stop accepting tab separators.
- Updated `read.ts` and `write.ts` hashline prepending/stripping to match `LINE+ID:` content prefixes.
- Updated `match-line-format.ts` and hashline read-mode docs to document and emit `LINE+ID:content` lines.
- Updated GrepTool result-limit messaging to include a concrete next skip hint.
- Updated grep tool test assertion to verify the new skip-hint limit message.
- Changed `todo_write` to an ordered `op`-array model with `replace`, `start`, `done`, `rm`, `drop`, `append`.
- Removed legacy multi-field todo payloads and updated tests/fixtures to use ordered `{op, task?, phase?, items?}[]` args.
- Reworked todo operation execution to apply entries sequentially and validate missing or unknown task/phase IDs.
- Updated todo rendering to use `todo.content` only and changed bash artifact labels from `full result` to `raw output`.
- Consolidated grep, ast-grep, and ast-edit on required `path`, replacing `glob`/`lang`/`sel` with inline file, dir, glob, list, and URL targets.
- Changed ast-grep and grep schemas to require a single `pat` string and use `skip` pagination instead of array patterns or offsets.
- Updated argument validation to reject empty `path` and invalid `skip`, and routed grep context to session settings only.
- Updated tool prompts and tests to reflect new path globbing semantics and `first N` truncation output text.
- Updated inline argument formatting in formatArgsInline to truncate key/value pairs by width and show ellipsis when values overflow.
- Fixed JSON-tree output to hide harness-internal INTENT_FIELD and __partialJson keys from top-level tool argument rendering.
- Fixed multiline scalar rendering and empty-structure display paths in renderJsonTreeLines after tree-prefix traversal changes.
- Changed the intent marker from "i" to "_i" in the agent loop path.
- Updated intent extraction to always destructure out the intent key and return stripped arguments even when intent is not a string.
- Disabled strict schema enforcement on AskTool to relax request validation.
- Removed the optional `cwd` field from Python tool parameters and used `session.cwd` for warmup and execution context.
- Updated atom edit test fixtures to pass `set` as arrays instead of scalar strings.
- Updated the agent loop intent marker constant to use `i` instead of `_i` for tracing fields.
- Updated intent tracing documentation to describe a generic string marker field and cleanup behavior.
- Adjusted related Anthropic and coding-agent tests to match the renamed intent field via shared `INTENT_FIELD` usage.
- Updated `CustomToolAdapter` to read `strict` from the wrapped tool instead of hard-coding it.
- Disabled strict validation for `LspTool` and `ResolveTool` by setting their `strict` flags to false.
- Updated `YieldTool` to set `strict` false in its fallback path and remove the local strict accumulator.
- Added checks to skip writing error messages to stderr when stopReason is 'error' or 'aborted', as these cases already handle error output separately.
- This prevents duplicate error reporting in print mode when the assistant encounters terminal error conditions.
- Added maxItems to the set of unsupported tool schema fields.
- Removed minItems constraints that are not 0 or 1, as Anthropic only supports these specific values.
- Added test to verify maxItems and non-standard minItems are stripped while minItems: 1 is preserved.
- Added locator-based atom edits with required `loc`, including inline `file:line`, range, and `^`/`$` selectors.
- Added path resolution via top-level path, entry path, or `path:loc` selectors and flattened multi-target atom execution.
- Removed `del`; required `set/pre/post` to use line arrays and mapped `set: []` to deletion semantics.
- Changed `sub` payloads to required `[find, replace]` tuples and tightened parsing to accept null optional verbs and array-or-null lines.
- Added a regression test for tool argument coercion that preserves quoted edit arrays while stripping optional null fields.
- Updated footer branch-watcher initialization to clear cached branch state when resolving git head fails.
- Set the report_tool_issue tool definition to non-strict mode for looser argument handling.
- Updated generateDiffString to preserve bounded context at both edges of unchanged regions between adjacent changes by adding a middle skip and ellipsis.
- Added a test covering distant edit hunks to verify unchanged lines in the gap are collapsed while change lines remain visible.
- Added an optional strict field to CustomTool definitions to support non-strict execution mode.
- Set strict to false across built-in AgentTool and custom tool registrations, including browser, calculator, GitHub, image generation, and related utilities.
- Updated inspect-image tests to reflect the relaxed strict setting on the tool.
- Extended optional git metadata checks to treat ENFILE and EMFILE like missing paths.
- Updated sync and async helper functions to return null for those filesystem errors instead of throwing.
- Documented the status-line branch-rendering fallback for ENFILE/EMFILE failures in the changelog.
- Added `CodeFrameMarker` and `formatCodeFrameLine()` to centralize code-frame gutter formatting.
- Extended diff rendering to preserve `|` and `│` separators, aligning gutter markers and line numbers.
- Reworked AST, grep, hashline, Vim, and diff renderers to use shared line formatting with computed `lineNumberWidth`.
- Updated atom editing flow and tests, including `resolveAtomEntryPaths` migration and new loc-based/edge-case coverage.
- Adjusted benchmark runner early-stop configuration by passing `buildEarlyStop` through prompt collection.
- Added `openai` and `openai-codex` as image providers and let `providers.image=auto` prefer GPT images.
- Updated settings, selector, and SDK wiring so OpenAI image providers pass through `setPreferredImageProvider`.
- Replaced Gemini-only image tooling with `image-gen` and added OpenAI/Codex hosted-image execution with SSE parsing.
- Added image-gen and handoff tests, including final-yield no-compaction regression and OpenAI payload/header assertions.
- Added tracking of the last successful non-error yield tool call when a yield execution ends.
- Cleared the tracked yield ID and skipped post-turn maintenance when appropriate, including when the last assistant message was a successful yield.
- Standardized unchanged-result error messages in atom and replace edit modes.
- Changed hashline/atom parsing to require full `line+suffix` anchors and emit full-anchor guidance on failures.
- Updated atom/hashline prompt docs, `atom.test.ts`, and changelog entries to require exact full anchors like `160sr`.
- Added preformatted `displayContent` for read/grep/ast tools and switched renderers to prefer it in TUI output.
- Updated mismatch and grep/ast renderers to show context with `*` markers and gutter lines, removing legacy helpers.
- Converted `tool.parameters` handling in `convertTools` to validate schema fields with `isRecord` before use.
- Filtered tool `required` entries to strings and defaulted missing `properties` to an empty object when building the schema.
- Added conditional strict-mode output so `strict: true` is included only when not disabled by `NO_STRICT` and the tool does not set `strict: false`.
- Expanded hashline and chunk bigram tables to 647 entries and moved chunk checksums to a 40-item namespace.
- Changed hashline anchors from `LINE#ID`/`:` to concatenated `LINEID`\t forms across parsing and tool outputs.
- Removed line-number padding and routed diff/read/grep/renderer output through raw numbers, tabs, and `toDisplayLine` formatting.
- Renamed atom ops to `pre`/`post`, removed `ins`, and updated schemas, prompts, and tests for new insertion behavior.
DeepSeek rejects assistant messages with content: null when reasoning_content
is present. Set content to "" instead of null when hasReasoningField is true.
- Guarded minimized output handling so the minimized text was only applied when it changed from the original output.
- Replaced the artifact footer text with a shorter `[full result: artifact://...]` marker when a minimized save artifact was created.
- Added a --no-early-stop-on-match CLI option, passed into benchmark configuration as earlyStopOnMatch.
- Added early-stop support that verified expected files after mutation-tool completion and aborted the prompt loop on a match.
- Captured an earlyStopped flag in task results and emitted an early_stop event when match-based termination occurred.