Retry: widened agent dequeue-hook deadline budgets from 25ms to 1s — the run loop checks the deadline before invoking dequeue hooks, so a cold or CPU-starved mock roundtrip expired the deadline first and the hooks never ran (deterministic failure in isolation, flaky under CI parallel load).
prepareCompaction walked the branch from the last compaction and ignored reset_boundary markers, so /compact (and auto-compaction) resurrected pre-/clear turns into the summary even though buildSessionContext already starts the model context after the boundary.
Model reset_boundary as a first-class agent-core session entry and start the summarization window after the latest boundary, dropping the superseded pre-reset compaction summary. A boundary before the last compaction stays superseded by it.
Fixes#8718
- Retained tool-search server calls and opaque results in signed assistant history across direct streams, gateways, and custom-endpoint projection.
- Added replay regressions for interleaved thinking and client tool continuations.
Fixes#8559
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
Each is a single, side-effect-free completion whose result is parsed after
it resolves, so one transient provider failure previously aborted the
whole operation - for /compact that left the user's context full.
Adds SummaryOptions.oneshotRetry, because both compaction paths call the
same generateSummary and the policy therefore cannot be a constant inside
it. Manual /compact has no outer loop and gets retry by default;
auto-compaction passes false because session-maintenance already retries
the whole attempt, and nesting would multiply the budget (10 outer x 3
inner) while stacking each outer wait on an inner backoff.
instrumentedCompleteSimple is the single funnel for every oneshot LLM
call in the agent, so retry belongs here rather than in a try/catch above
each caller - the failure arrives as a resolved AssistantMessage.
Opt-in rather than default-on: oneshotKind is free-form and callers may
pass arbitrary ctx.tools, so the funnel cannot itself prove a request is
replay-safe.
Response headers are captured per attempt and cleared between attempts,
so a stale retry-after can never be reused for a later failure.
- Refactored and condensed numerous system prompts, agent instructions, and tool documentation files across packages.
- Streamlined workflow rules, formatting constraints, and execution guidelines for improved clarity and brevity.
- Updated discovery rules, recommendation criteria, and syntax standards in prompt templates.
RESCUE_SHAKE_CONFIG spreads AGGRESSIVE_SHAKE_CONFIG, so the new 4k manual
tail leaked into dead-end recovery and could block eliding the very
result that caused the dead end. Override protectTokens back to 0 in the
rescue preset and add the coding-agent changelog entry for the manual
/shake behavior change.
Addresses review on #8067.
Cursor bash/grep frames wrote optional kwargs as present-undefined, and
tools.format gemini projectors dropped kCursorExecResolved so settled
calls ran twice.
Co-authored-by: Cursor <cursoragent@cursor.com>
Manual /shake used protectTokens: 0, stripping every eligible tool result
including the ones the agent is still working from. Keep a small 4k-token
recent window (matching the automatic shake mechanism, at a quarter of its
16k budget) so the full escape hatch stays aggressive without destroying
the live tail. Two matcher-focused tests that implicitly relied on the
zero window now pin protectTokens: 0 explicitly.
Fixes#7776
When an xd:// device is dispatched through the write tool, the outer
approval gate now consults tools.approval.<deviceName> before falling
back to tools.approval.write. This lets users scope allow/deny/prompt
to a single device mount without changing the blanket write tool policy.
The write tool's approval function returns { tier, policyKey: deviceName }
for xd:// device dispatches. resolveApproval uses the policyKey to look
up the user override on the device name, falling back to the invoking
tool's own policy when the device has none configured.
Adds:
- ToolApprovalDecision.policyKey field (optional, additive)
- policyKey-aware lookup in resolveApproval and requiresApproval
- Updated error messages naming the correct config key
- Unit tests for policyKey resolution and WriteTool integration
Fixescan1357/oh-my-pi#7923
- Reused the Responses replay lifecycle policy for V1 and V2 compaction input.
- Covered persisted native history and converted assistant replay items.
Fixes#7742
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.