- Removed the streaming guard that previously rejected /tan while the parent response was still generating.
- Passed "deliverAs: \"nextTurn\"" when sending the background dispatch breadcrumb and kept "triggerTurn: false" so an in-flight turn is not steered.
- Skipped rebuilding chat messages during streaming sessions and updated tests to cover the non-blocking dispatch path.
- The AgentSession retry fallback test now validates the assistant message before reading it.
- It now verifies the first content block is text before checking the recovered message text.
- Added --recover CLI option to rebuild changelog fixes from tagged history.
- Pruned Unreleased bullets that match historical released items across all tags.
- Normalized output by sorting release sections by version and compacting bullet spacing.
- Fixed OAuth credentials to keep unknown fields in schema while preserving existing shape checks.
- Fixed MCP OAuth IDs to be profile-scoped and avoid deleting credentials from non-active profiles.
- Fixed string-flag parsing so PROFILE_BOOTSTRAP_BOUNDARY tokens are not consumed as values.
- Fixed active-profile directory resolution to refresh after env updates so profile .env overrides apply.
ExtensionRunner.emit shared the generic 30s EXTENSION_HANDLER_TIMEOUT_MS budget with every event, including the fire-and-forget session_shutdown teardown event extensions cannot observe. A hung third-party handler — observed on Windows with omp-discord-presence 0.1.2 waiting on a stuck Discord IPC pipe — held AgentSession.dispose() for the full window, making Ctrl+C look ignored for 30s.
session_shutdown now uses a dedicated 2s SESSION_SHUTDOWN_HANDLER_TIMEOUT_MS cap routed through a per-event handlerTimeoutForEvent() lookup so generic and shutdown budgets are independently configurable. The interactive-mode Ctrl+C path adds a defence-in-depth hard-exit: when isShuttingDown is true a fresh Ctrl+C exits with code 130 (the session JSONL has already been sync-flushed by the first press) instead of stacking another no-op shutdown() call.
Fixes#2600
- Added runtime hook resolvers so each `JsRuntime` instance can expose hooks for its active run.
- Patched `process.stdout` and `process.stderr` writes once per stream to route output chunks through active run text hooks and preserve existing worker logging when no run is active.
- Added chunk-to-string conversion for write payloads and encoding-aware forwarding while keeping callback semantics intact.
scripts/ci-release-notes.ts previously extracted only the single
target-version section, so changelog entries finalized under tags pushed
without a GitHub Release (e.g. v15.12.5 and v15.12.6 — collateral from
the pre-#2564 release-cancellation bug) were stranded out of the next
published release body.
The generator now walks the range (latest-published-release, target],
resolved via 'gh release list', and merges every in-range '## [X.Y.Z]'
section per package — grouped by '### <category>' with bullet-level
dedup so post-release changelog flattening cannot surface the same
entry twice. Versions iterate newest-first so newer phrasing wins on
dup resolution, and categories are sorted into the canonical
Breaking/Added/Changed/Fixed/Removed order regardless of source order.
Falls back to legacy single-version extraction when 'gh' is unavailable
or no prior published release resolves (safe no-op);
'OMP_RELEASE_NOTES_FLOOR=v15.12.4' overrides the lookup for manual
re-runs (empty string forces legacy mode).
Adds scripts/ci-release-notes.test.ts covering: range inclusion above
floor, target-inclusive boundary, dedup of bullets flattened forward
into multiple versions, canonical category ordering, the null-floor
legacy fallback, empty version sections skipped, and no empty-category
emission when dedup drains a bucket. Wired into 'bun run test:scripts'.
Fixes#2596
The workspace tree shown in the system prompt renders per-entry modification
times as render-time relative ages ("9m ago") computed from Date.now() on
every build. Those strings drift between sessions ("9m ago" -> "10m ago",
"59m ago" -> "1h ago") while the files themselves are unchanged. Because the
tree sits ahead of the (multi-thousand-token) tool block and KV cache is
contextual, that one early change invalidates the cached prefix for everything
after it, forcing a full prompt re-prefill on the first request of every new
session — even when nothing in the workspace actually changed.
Fix: render a deterministic absolute UTC timestamp (YYYY-MM-DD HH:MM) derived
purely from the file's mtime for the cached system-prompt tree, so the rendered
block is byte-identical across sessions and only changes when a file actually
changes. Scoped via a new internal AssembleOptions.ageMode:
- buildWorkspaceTree (cached system prompt) -> "absolute"
- buildDirectoryTree (read-tool output, not cached) -> "relative" (unchanged)
renderNode now takes a per-pass age formatter instead of reading Date.now()
directly.
Measured on a local llama.cpp server (single user, prompt cache on): with a
file whose age ticks between two back-to-back sessions, the unpatched build
re-prefills the full prefix on session 2 (27,124 prompt tokens, 46s); with this
change session 2 is a cache hit (13 tokens, 2s). Existing tests are unaffected
(they assert on filenames/order/elision, not on age strings); two regression
tests added.
- Collapsed the model list while the role/action menu is open and restored it when the menu closes.
- Limited rendered menu options to a terminal-derived visible window centered around the selected item.
- Added overflow handling via ScrollView with dynamic width and scrollbar when the option list exceeds available rows.
- Added beforeEach and afterEach hooks in mnemopi tests to set and clear MNEMOPI_NO_EMBEDDINGS so embeddings are skipped during those runs.
- Updated the bun-install cache script to archive only node_modules paths that exist as directories.
- Applied title-casing to `normalizeGeneratedTitle` outputs using a new internal helper.
- Adjusted tiny text and title generator tests to assert the new title-cased results.
- Added a new built-in `title` model role with `hidden` metadata and updated role definitions and schema.
- Updated title generation to resolve models in `title`, `commit`, then `smol` order and added test coverage for that precedence.
- Filtered hidden roles from selector badges and documented the new built-in role in model/settings docs.
- Updated the default `omp bench` prompt text to emphasize full-spectrum reasoning before answering.
- Required explicit enumeration of all four-table join orders with cost comparisons across nested-loop and hash joins plus index-scan versus full-scan tradeoffs.
- Expanded the changelog rationale to document sustained deliberation and exhaustive costing to avoid short-circuit benchmark responses.
- Updated CI dependency install flow to share bun cache orchestration across jobs.
- Added RustFS-backed bun cache restore/save script keyed by bun.lock hash.
- Updated the benchmark prompt to request a concrete, schema-driven query-optimization walkthrough with explicit selectivity, cardinality, join-order, and operator-cost calculations.
- Adjusted the output constraints to require plain-paragraph analysis output with no headings, lists, code fences, or tables.
- Documented the default benchmark prompt replacement in the package changelog under the Changed section.
The previous `style: bun run fix` auto-commit (a53ca08370) re-hit the known fix-changelogs failure mode: the local checkout's tags lag v15.13.0, so the script keeps comparing against an older tag and promoting the entire 15.13.0 release block into Unreleased every time it runs (same root cause as 1d931506d6).
Restored CHANGELOG.md to the pre-fix state and re-applied only the #2590 entry. Pushing with skip_checks=true to bypass the deterministic auto-fix loop; diff vs origin/main is exactly one inserted line.
The previous `style: bun run fix` run (commit 63657246dc) hit the known fix-changelogs failure mode (see commit 1d931506d6): the local checkout's tags lag the v15.13.0 release tag, so the script compared against an older tag and promoted the entire 15.13.0 release block into Unreleased.
Restored CHANGELOG.md to its pre-fix state and re-applied only the #2590 Unreleased entry. Diff vs origin/main is now exactly one inserted line.
`ToolResultContent` is not exported from `@oh-my-pi/pi-ai`; `ToolResultMessage.content` is `(TextContent | ImageContent)[]`. Test only needs text, so type the content array as `TextContent[]`.
`#checkTodoCompletion` used to append a `<system-reminder>` and then call
`#scheduleAgentContinue`, so a text-only acknowledgement ("paused at your
instruction") triggered another `agent_end` that re-ran the same check and
fired the next reminder — counter ticked 1/3 → 2/3 → 3/3 inside a single user
pause without any user input. The user perceived three back-to-back reminders
appear from nowhere; the agent felt implicit pressure to invent busy-work or
take destructive ops to silence the loop.
Added `#todoReminderAwaitingProgress`: a reminder sets it, any `toolResult`
(real tool-level progress) or a new user prompt clears it, and
`#checkTodoCompletion` stays silent while it is set. Reset alongside
`#todoReminderCount` on user prompts, session reset, handoff, and the no-op
short-circuits in `#checkTodoCompletion` so the field never gets stuck.
Escalation through `todo.reminders.max` still works when the agent makes
tool-level progress between stops — that is the case the cap was designed for.
Regression coverage in `agent-session-todo-reminder-loop.test.ts` drives a
mocked `agent.continue` to mirror the bug-reported model behaviour and pins
the contract: exactly one reminder per user pause when the agent only
acknowledges; re-escalation when the agent actually calls a tool between stops.
Fixes#2590
Loaded Kokoro's side-installed transformers runtime by absolute path before requiring kokoro-js, avoiding host/workspace onnxruntime libraries in the worker process.
Kept runtime-cache bare module requests inside the registered runtime cache when the parent module is already inside that cache, and covered the resolver boundary with a regression test.
Fixes#2591
- Added a setup-system-deps action with preloaded-runner guards and apt fallbacks.
- Updated CI workflows to download Linux x64 native artifacts and gate on native job success.
- Renamed coding-agent fast mode to singleton in scripts and test partitioning logic.
- Added settings test-state begin/restore helpers with recursive cleanup in affected tests.
- Normalized agent `setSystemPrompt` to wrap string inputs into one-item arrays.
- Updated session creation to accept string `systemPrompt` values and normalize callback or direct results to string arrays.
- Adjusted extension result handling and test fixtures to accept string `systemPrompt` and missing `assistant_message` fields without crashing.
- Updated ToolExecutionComponent.isTranscriptBlockCommitStable to return true when a tool result exists, so streaming results are treated as commit-stable.
- Limited provisionalPendingPreview handling to the pending call phase so only pre-result previews remain non-committed, preventing collapsed streams from dropping their top rows.
Confirmed root cause of the CI test hang: bun --parallel spawns one isolated
worker per core, and each worker loads the 116MB pi-natives addon plus a large
JS heap. On the 16GB hosted runner that exceeds memory, triggering swap thrash
(100s+ event-loop stalls, transient file-read failures) that looks like a hang.
Capping to 2 workers keeps per-file memory recycling (isolation) while halving
peak memory so it fits the runner.
Restored the released coding-agent changelog sections that the previous pre-publish fix run flattened into Unreleased. The local checkout only had tags through v15.5.15, so fix-changelogs compared against an old tag and promoted entire newly-added release sections.
- Kept newly-added release-section hunks out of collectPromotableAddedItemLines so the fixer still promotes individual new items added under an existing released section, but does not rewrite whole release blocks.
- Added regression coverage for the release-section case.
- Restored the coding-agent changelog history and kept only the #2582 entry under Unreleased.
Fixes#2582
fix(coding-agent): avoid placeholder crash before theme init
Resolved CHANGELOG conflict: 15.13.0 was re-opened into [Unreleased] on
main, so the fix entry goes under the existing [Unreleased] Fixed section
rather than resurrecting the released 15.13.0 heading.
Two CustomEditor/streaming-preview tests timed out under bun test --parallel
because they reset+reinitialised the process-global Settings singleton in
beforeEach. 50+ other test files do the same dance, so under parallelism the
proxy points at whichever instance won the latest race rather than the one
under test (issue #2582).
- Added CustomEditor#magicKeywordsEnabledOverride: an instance-level test/host
injection that short-circuits the global lookup at the read site.
Production wiring still reads from Settings.
- Switched the magic-keyword shimmer-disabled test to set
magicKeywordsEnabledOverride directly instead of mutating the singleton.
- Removed the gratuitous reset+init from the 'streaming tool call preview
height (bounded across renderers)' describe -- bash/ssh/eval pending
previews never read settings.*; only initTheme() (idempotent) is kept.
Fixes#2582