The post-interrupt immune-window downgraded end-of-turn blockers to non-interrupting asides, so a blocker that means the agent handed off broken work never woke a new turn. Concerns keep the cooldown; blockers now bypass it and steer a triggered turn, consistent with the #5628 blocker-after-terminal-answer exception.
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
After a prewalk hand-off plus mid-run compaction, the openai-responses input
builder re-encoded replayed assistant turns via convertResponsesAssistantMessage,
demoting their reasoning to <think> output_text and emitting no reasoning item.
DeepSeek (opencode-go) then rejected the thinking-mode continuation with
400 "The reasoning_text in the thinking mode must be passed back to the API"
and OMP looped on the unchanged request.
The encoder now synthesizes a reasoning_text reasoning item for every replayed
assistant turn when the target requires reasoning replay in thinking mode
(requiresReasoningContentForAllAssistantTurns / requiresReasoningContentForToolCalls),
carrying surviving thinking text when present, mirroring the chat-completions
reasoning_content safety net. Gated on reasoning being active for the request,
so non-DeepSeek Responses targets and reasoning-disabled turns are unaffected.
Fixes#8248
A classifier refusal returned before the `onTurnError` hook that owns
model fallback, so `AdvisorRuntime` treated one provider's policy verdict
as terminal: `Refusal (cyber)` on the advisor model disabled the advisor
outright even with a fallback chain configured. Its only recovery was
stripping echoed primary reasoning and resending once, which does nothing
for a refusal about the content itself.
Route a refusal that outlives the strip through the same hook the generic
failure path uses, mirroring its epoch guard, session-transition requeue,
and requeue-on-recovery. The cascade walks the chain to exhaustion and
only reports the advisor unavailable once the host runs out of candidates,
matching what turn-recovery already allows for the primary.
Each cascade visits a model at most once. A switch re-arms
`#includeThinking` through `#syncModelIdentity`, so chain keys that point
back at each other (A to B, B to A) would otherwise strip-and-resend
against the same pair forever. A successful turn or a reset starts a
fresh walk.
Also stop `/advisor status` throwing when a live advisor has no roster
entry: `formatAdvisorStatus` guarded only the inactive case before
dereferencing `stats.advisors[0]`, and `#ensureAdvisors` clears
`#advisorStatuses` before repopulating it, so a status call landing in
that window hit `undefined.contextWindow`.
Addresses Codex P2: the obfuscation fake in delta-split-obfuscation.test.ts
used an `as any` escape. RenderAdvisorDeltaChunksOptions.obfuscator is now
AdvisorObfuscator — a structural interface exposing only obfuscate() — so the
fake is typed without any, while a real SecretObfuscator stays assignable and
the production call site is unchanged. Also drops two leftover console.logs
from the test (matching the cleanup roboomp requested elsewhere).
The advisor sends its whole Session update as a single ever-growing user
message. Provider prompt caches are prefix-based: a single user message whose
text keeps growing invalidates the entire message on every turn, so cache_read
stays pinned at the instructions/tools boundary (observed 14491 tokens in
production, 11066 in tests) instead of growing with the session.
Split the update into multiple user messages — one per source message —
delivered via a single Agent.prompt(AgentMessage[]) call, so the provider
caches each appended message incrementally. Verified end-to-end: cache_read
grows 0 -> 11126 -> 11457 -> 11583 across turns with the split, versus pinned
11066 on the old single-message behavior.
- delta-split.ts: pure renderAdvisorDeltaChunks using chunked
formatSessionHistoryMarkdown (shared toolResultIndex/consumedToolCallIds/
watchedRoleState) so toolCall/result pairing and role collapsing stay
byte-identical to the old single-block render (equivalence-tested).
- session-history-format.ts: add HistoryFormatOptions.watchedRoleState so
chunked renders collapse consecutive same-role messages exactly like the
single-block render.
- runtime.ts: #prepareBatch does a single dedup+render pass; #drain delivers
agent.prompt(preparedMessages) (array), falling back to the string.
- Keep field-selective fingerprint (candidate 1) + wip-marker-at-tail
(candidate 3) as complementary wins.
Tests: advisor suite 209 pass / 0 fail; type check clean; lint clean.
Affected subsets (342 tests) green; full suite hits WSL EMFILE fd limit.
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
- Remove `annotateForStaleness` and `hasFreshBacklog` from the advisor runtime.
- Stop appending staleness warnings to delivered advisor notes when newer primary turns queue.
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.
The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.