Added a cross-turn tool-call loop guard that hashes canonical tool names and arguments, ignores intent metadata, and injects a hidden redirect when identical calls reach the configured threshold.
Fixes#3971
The subagent system prompt rendered `{{jtdToTypeScript outputSchema}}` as a
bare TypeScript interface with the text "Your result MUST match this
TypeScript interface". The yield tool actually nests the user schema under
`result.data`, so the LLM pattern-matched on the visually dominant code
block and put the payload directly in `result.data`, tripping schema
validation repeatedly. In the worst reported case a subagent used all 3
retry attempts, had validation dropped, and lost its audit output entirely.
Add a `renderYieldSchema` Handlebars helper that renders the schema inside
`result: { data: … }` and swap the system-prompt block to use it, so the
model sees the exact envelope the yield tool expects. Multi-line object
schemas, scalars, unions, and array-of-object schemas all round-trip
cleanly with the new helper.
Fixes#3972
- Removed instructions prohibiting mocks and testing defaults in favor of using a tester agent when available.
- Streamlined delivery contract rules by removing the explicit instruction against suppressing tests.
- Updated verification claim guidelines to emphasize smoke testing.
- Added a hidden system notice prompt to instruct the model to break repetitive behaviors when a thinking or response loop is detected.
- Injected the redirect notice into the retried turn's context when resetting the active context after a loop retry.
- Added unit tests to verify the custom redirect message is appended, configured as non-displaying, and visible in subsequent LLM contexts.
- Removed the oracle agent prompt configuration and its references across the codebase.
- Added a new tester agent prompt markdown file with directives for test authoring and techniques.
- Registered the new tester agent in the task runner while removing the oracle agent.
- Renamed the quick-task agent configuration and prompt references to sonic.
- Renamed references to the `quick_task` subagent to `sonic` across docs, agent definitions, prompts, and test files.
- Updated the parallel file analysis tool to spawn `sonic` subagents instead of `quick_task`.
- Documented the breaking change in the changelog along with additions and removals of other built-in subagents.
Stopped passive autolearn from adding hidden conversation messages and froze local memory developer instructions per session so learn writes land in future sessions instead of mutating the active Anthropic prompt prefix.
Fixes#3743
- Removed support for `history://` URI schemes used to read agent transcripts from system and tool prompts.
- Updated IRC tool instructions to remove references to reading agent history for peer information.
- Replace the static "Goal/Next" status line with an ephemeral LLM-generated summary triggered after idle periods.
- Hook the recap into the agent's side-channel pipeline, using live goal and task state as context anchors for meaningful recaps.
- Implement abort logic so that active user interactions immediately cancel pending recaps and discard late-arriving responses.
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
- Added `INTERRUPTED_THINKING_MESSAGE_TYPE` and `demoteInterruptedThinking` in `session/messages.ts` for stripping unfinished thinking into durable context.
- Persisted hidden interrupted-thinking context from `session/agent-session.ts` immediately after user-interrupted assistant turns.
- Added the `prompts/system/interrupted-thinking.md` envelope used for replaying interrupted reasoning.
- Added focused interrupted-thinking coverage in `test/agent-session-interrupted-thinking.test.ts` and `test/session/interrupted-thinking-demote.test.ts`.
- Implemented a monitoring system to identify Gemini model runaway behavior during thinking steps using consecutive header detection.
- Introduced a configurable tool-call reminder prompt to inject corrective context when reasoning stalls.
- Added session-level logic to automatically interrupt and prune stalled assistant turns from the conversation history.
- Provided comprehensive test coverage for the detection logic and stream interruption scenarios.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Added instructions for writing sections as cohesive multi-line blocks when performing edit operations.
- Clarified that block operations require multi-line sections to avoid falling back to standard editing behavior.
- Added tree-sitter markdown support to resolve headings into full sections in `pi-ast`.
- Enabled block operations (`SWAP.BLK`, `DEL.BLK`, `INS.BLK.POST`) on markdown headings so they encompass the entire section, including nested deeper headings.
- Updated system prompt to guide agents in using structured markdown heading edits for plans.
- Fixed `plan-mode-guard` to correctly resolve local protocol options for subagents.
With `autolearn.autoContinue` on, the controller fires a synthetic
turn whose only user-role payload is `autolearn-nudge.md`. The old
prompt opened with "Before you finish:" and gave no terminal
contract, so after the `learn`/`manage_skill` call the agent read
its own unanswered prior question (e.g. "Want me to commit and
push?") as accepted and continued — pushing commits, running tools,
etc. — without the user ever answering.
Split the nudge into two prompts and pick at fire time:
- passive (rides the user's real next message) keeps additive
framing — "answer the user normally; the capture is in addition
to, not a replacement for, the work the user just asked for".
- auto-continue (`autolearn-nudge-autocontinue.md`) is explicitly
terminal — "not a user reply; do not treat this as approval or
acceptance of any pending action; capture, then stop; do not run
other tools, resume prior work, or answer your own pending
questions; wait for the user's next prompt".
Attribution stays `user` so llama.cpp keeps reusing the warm
prefix (#3456). Regression tests pin the load-bearing terminal
language in both branches.
Fixes#3504
- Transitioned the eval tool from batch multi-cell execution to a single-step input structure with flat parameters.
- Updated core agent logic, UI components, and documentation to support state persistence across incremental eval calls.
- Restricted bash tool capabilities by requiring explicit use of `read` or `find` instead of `ls` or `find`.
- Added support for Ruby and Julia language runtimes to the eval tool and associated web renderers.
- Elevated `eval` to an essential tool to ensure availability across all discovery modes.
- Updated system and tool prompts to mandate the use of `eval` for non-trivial shell operations like conditionals, loops, heredocs, and complex pipelines.
- Restricted `bash` usage to simple binary invocations and single-fact computation to reduce shell-escaping and execution errors.
The eval agent() helper used `agent_type`/`return_handle` (snake_case) in
Python/Ruby/Julia and `agentType`/`returnHandle` (camelCase) in JS, forcing
the prelude docs to repeat every option twice ("JS same but camelcased").
Both are now single lowercase words identical across all four runtimes, and
`agent` matches the `task` tool's existing agent-selection parameter.
- Renamed across py/js/rb/jl preludes (signatures, forwarding, docstrings).
- Renamed the `__agent__` bridge wire protocol + `EvalAgentArgs` (`agentType`
→ `agent`, `returnHandle` → `handle`) so no prelude-side remap is needed.
- Updated prompt docs (workflow-notice.md, tools/eval.md), repo docs
(docs/tools/eval.md, docs/python-repl.md), and all bridge/prelude tests.
- CHANGELOG: Breaking Changes entry under [Unreleased].
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.
Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
forwards isolated/apply/merge (as booleans) plus returnHandle, and the
return_handle node carries isolated/patch_path/branch_name/
nested_patches/changes_applied/isolation_summary.
Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
the stale "defaults track task.isolation.mode" wording to the final
strict opt-in behavior, and note all four runtimes.
Fixes#3196
Per maintainer ruling on #3196, eval agent() now defaults to non-isolated regardless of task.isolation.mode, mirroring the task tool. isolated=true is the only way to turn it on; isolated=true while task.isolation.mode === "none" still throws the same clear error.
Updated tests, workflow-notice.md, and Python agent() docstring to reflect the strict opt-in contract. Existing isolation tests now pass isolated:true explicitly; the inherit-from-settings assertion is replaced with a default-off + isolated=true opt-in regression.
Fixes#3196
- Updated system and tool prompts to explicitly forbid using shell utilities like grep, rg, awk, and find for tasks better suited to specialized tools.
- Clarified that bash should be reserved for terminal operations and computational pipelines that produce facts not available through existing specialized tool outputs.
Branch-mode isolation can capture nested repository changes without creating a root branch. Eval agent() with apply=false previously treated that shape as no captured changes and returned no recoverable nested patch payload after the isolation worktree was removed.
Expose captured nested patches in EvalAgentResult details and copy them onto JS/Python returnHandle nodes (nestedPatches / nested_patches). Document the return_handle escape hatch and add regression coverage for branch-mode nested-only apply=false runs.
Fixes#3196
- Updated system and tool prompts to present dedicated tools as preferred defaults rather than absolute prohibitions.
- Relaxed the hard-forbidding of shell equivalents for file operations, searching, and editing.
- Retained guidance on prioritizing tools for their gitignore semantics, structure, and line-anchoring capabilities.
When agent() ran with schema and apply=false, the bridge correctly returned the captured patch/branch in details, but the preludes only forwarded id/agent/handle/data on the returnHandle node. Structured workflows had no way to recover the artifact for a manual apply.
Both runtimes now copy isolated, patchPath/branchName, changesApplied, and isolationSummary onto the returnHandle node (snake_case in Python, camelCase in JS), keeping null changesApplied so apply=false stays distinguishable from a successful apply. Updated the workflow notice and the Python agent() docstring to point callers at return_handle as the artifact escape hatch for isolated+apply=false runs. Added prelude tests locking the new node shape in both runtimes.
Fixes#3196
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.
Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.
Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.
Fixes#3196
- Inlined the temporary model status formatting logic directly into the controller.
- Removed the unused `formatTemporaryModelStatus` utility function and its associated test.
- Added `includeWorkspaceTree` configuration setting to optionally render the workspace directory tree.
- Configured the system prompt template and SDK to respect this toggle, allowing users to disable the tree to prevent prompt cache invalidation.
- Added `buildSideRequestContext` to the `Agent` class to generate prompt-cache-friendly provider contexts.
- Updated ephemeral side-channel turns to forward the full tool catalog to maintain prompt cache hit rates.
- Injected a `developer` role reminder into ephemeral turns to instruct the model to suppress tool calls.
- Implemented automatic post-processing to strip any tool calls from ephemeral turn responses.
- Exported message and dialect helper functions in `agent-loop.ts` to support context construction.
- Add a dedicated `<parallel-reflex>` section to the system prompt to discourage serial work habits and enforce parallelization by default.
- Refine task-spawning guidance to emphasize intentional delegation, agent specialization, and clear assignment criteria.
- Update `task.md` parallelization heuristics and rule definitions to clarify when subagents should be deployed concurrently versus sequentially.
- Simplified system and personality prompts for improved conciseness and clarity.
- Streamlined tool instruction sets and parameter descriptions across all agent modules.
- Refactored prompt documentation in `hashline` to clarify terminology and task-specific constraints.
- Updated tool metadata in TypeScript service definitions to align with reduced documentation verbosity.
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
- Added default system-prompt guidance to scan skill descriptions and read applicable skill:// content before work.
- Covered the rendered prompt contract with a frontend-design skill regression test.
Fixes#2829
- Added own-line display-math tokenizers for `$$...$$` and `\[...]` markdown blocks.
- Added delimiter-free `\begin...\end` math parsing in text conversion with multiline row preservation.
- Added ANSI rendering for `\textcolor`, `\color`, `\colorbox`, and `\fcolorbox` with scoped color resets.
- Added tests for LaTeX color and math parsing edge cases, including inline code and nested scopes.
- Added LaTeX math and Mermaid allowances in terminal and final-chat prompts.
- Added inline math tokenization in TUI for $, $$, \(\), and \[\].
- Added LaTeX-to-Unicode conversion helpers and exports for math rendering.
- Fixed inline math detection to skip escaped dollars and currency-like spans.
- Added `features.unexpectedStopDetection` and `unexpectedStopModel` settings for opt-in behavior.
- Added assistant-stop handling to classify stop reasons and resume generation with retry prompts.
- Added unexpected-stop classifier logic with candidate checks, model fallback, and YES/NO parsing.
- Added retry tracking that caps auto-continues at three attempts and logs a warning when exceeded.
- Added `jsonSchemaToTypeScript` and `renderToolInventory` to generate tool blocks with TypeScript signatures.
- Added `examples` and `TSchema` fields to dump-tool metadata and passed them through prompt rendering.
- Changed Harmony invocation rendering to omit `<|constrain|>json` markers in tool call payloads.
- Added compact native tool list-mode inventory rendering with full `# Tool:` output elsewhere.
- agent-loop: raise repetition-detection floor to 180 chars and clear thinking
replay anchors when collapsing a detected loop.
- providers/google: ignore empty text parts, retain terminal thoughtSignatures,
and stop function-call signatures clobbering the prior block.
- autolearn: capture goal-mode at the turn boundary; harden managed-skill writes
against hard-links/symlinks (O_NOFOLLOW + nlink); refuse minting managed skills
whose name an authored skill already claims.
- eager tasks: thread agentKind through the session so a custom top-level agentId
still gets always-mode delegation; split Eager Tasks prompt into hard vs soft.
- title-generator: race the online title model against a local tiny-model fallback.
- eager-todo: keep the soft reminder aligned with the todo init schema.
- mcp/stdio: keep close() detaching the read loop instead of awaiting it.
- stream loop: fix collapsing and tool-call thought-signature handling.