Caught recall and retention failures at their fire-and-forget event boundaries, with bank and operation diagnostics.
Added agent-start and agent-end regression coverage for unavailable session data and bank access.
Fixes#8351
- Added shared auth storage and model registry for concurrent and eager compaction tests.
- Replaced per-test auth storage creation and cleanup routines with shared lifecycle management.
- Removed local auth storage variables and individual cleanup hooks from test harness instances.
- Added `splitAddressableFileLines` to strip terminal newlines from line addressability without removing genuine blank lines.
- Updated coding-agent read tool context parsing to use addressable file lines.
- Updated eager compaction and plan reference tests to track call indices and task delegation markers instead of text strings.
- Removed obsolete context message marker checks, vibe mode assertions, and prompt gating test cases.
- Simplified prewalk, workflow, and Gemini instruction test expectations across agent modules.
- Removed the system prompt personality test suite entirely.
- Added archive and member size assertion limits along with path byte-length checks for PAX and GNU metadata targets.
- Added support for global PAX attributes, old-GNU name records, and signed GNU base-256 numeric header fields.
- Updated archive reading in WriteTool to accept a filesystem path instead of buffered bytes.
- Added test coverage for signed GNU base-256 values, PAX extensions, overlong path rejections, and oversized archives.
- Bound PAX sparse record memory overhead by caching sparse markers and specific keys.
- Update system prompt phrasing and tests for tool inventory and date displays.
- Added Google provider thinking configuration parameters and force-reasoning-off controls.
- Implemented MCP SSE stream resumption using Last-Event-ID and `SSEResumeError`.
- Added support for TAR old-GNU sparse extension blocks, path length checks, and archive entry overrides.
- Restricted external thinking support to specific models and added semver fallback parsing.
- Added the `--external-thinking` CLI flag alongside model capability checks to gate external thinking tool availability.
- Updated Anthropic and Google transports to honor `forceReasoningOff` for native thinking-off controls.
- Renamed the `thoughts` property and parameter to `notes` across think fixtures, tools, and tests.
- Updated system prompt instructions and test suites to verify transport-specific thinking and tool activation.
- Capped directory-alias rewrites per lookup (ELOOP-style, 40) so a directory
symlink targeting its own subtree (a -> a/b) throws a catchable ToolError
instead of looping forever growing the path.
- Deferred pending tar link resolution while any directory on the target path
is itself an unresolved link, and rewrote targets through established
directory aliases before the exact-path lookup, so file symlinks routed
through directory aliases materialize instead of dangling.
- Regression tests reproduce both shapes: pre-fix the aliased symlink read
failed with 'cannot be materialized' and the self-cycle read hung.