- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
- Renamed line and block patch op verbs to XCHG, DEL, and INS in parsing and formatting.
- Updated grammar and tokenizer to support XCHG.BLK, DEL.BLK, and INS.PRE/POST/HEAD/TAIL forms.
- Updated diagnostics, docs, prompts, tests, and changelog to use XCHG/DEL/INS-based operators.
- Expanded session-stats parsing to normalize legacy op aliases to compact IDs.
- WelcomeComponent now lazily selected and cached a tip per instance, preserving it across re-renders.
- With the unicode preset, it showed a special nerdfont tip 10% of the time and otherwise used the regular tip rotation.
- Added tests that mocked theme preset and Math.random to verify standard and special tip selection behavior.
The external editor flow (Ctrl+G, plan editor, /todo edit) warned 'No
editor configured' on Windows because getEditorCommand() returned
undefined whenever neither $VISUAL nor $EDITOR was set — the default
state for most Windows shells.
Fall back to 'notepad' on win32 after consulting $VISUAL/$EDITOR
(always present in %SystemRoot%\\System32) and trim env values so
accidentally padded strings still resolve. POSIX still returns
undefined so the warning continues to nudge users to configure an
editor.
Fixes#2604
- Tracked seen-line provenance in snapshots and propagated it from read/search/ast-grep rows.
- Rejected hashline edits on unseen lines before patching, throwing unseen-line errors.
- Rejected single-line block anchors in strict mode and dropped them in unresolved lenient mode.
- Trimmed one-sided keeper-echo duplicates during multi-line replacements with warning output.
- Removed the streaming guard that previously rejected /tan while the parent response was still generating.
- Passed "deliverAs: \"nextTurn\"" when sending the background dispatch breadcrumb and kept "triggerTurn: false" so an in-flight turn is not steered.
- Skipped rebuilding chat messages during streaming sessions and updated tests to cover the non-blocking dispatch path.
- The AgentSession retry fallback test now validates the assistant message before reading it.
- It now verifies the first content block is text before checking the recovered message text.
- Fixed OAuth credentials to keep unknown fields in schema while preserving existing shape checks.
- Fixed MCP OAuth IDs to be profile-scoped and avoid deleting credentials from non-active profiles.
- Fixed string-flag parsing so PROFILE_BOOTSTRAP_BOUNDARY tokens are not consumed as values.
- Fixed active-profile directory resolution to refresh after env updates so profile .env overrides apply.
ExtensionRunner.emit shared the generic 30s EXTENSION_HANDLER_TIMEOUT_MS budget with every event, including the fire-and-forget session_shutdown teardown event extensions cannot observe. A hung third-party handler — observed on Windows with omp-discord-presence 0.1.2 waiting on a stuck Discord IPC pipe — held AgentSession.dispose() for the full window, making Ctrl+C look ignored for 30s.
session_shutdown now uses a dedicated 2s SESSION_SHUTDOWN_HANDLER_TIMEOUT_MS cap routed through a per-event handlerTimeoutForEvent() lookup so generic and shutdown budgets are independently configurable. The interactive-mode Ctrl+C path adds a defence-in-depth hard-exit: when isShuttingDown is true a fresh Ctrl+C exits with code 130 (the session JSONL has already been sync-flushed by the first press) instead of stacking another no-op shutdown() call.
Fixes#2600
The workspace tree shown in the system prompt renders per-entry modification
times as render-time relative ages ("9m ago") computed from Date.now() on
every build. Those strings drift between sessions ("9m ago" -> "10m ago",
"59m ago" -> "1h ago") while the files themselves are unchanged. Because the
tree sits ahead of the (multi-thousand-token) tool block and KV cache is
contextual, that one early change invalidates the cached prefix for everything
after it, forcing a full prompt re-prefill on the first request of every new
session — even when nothing in the workspace actually changed.
Fix: render a deterministic absolute UTC timestamp (YYYY-MM-DD HH:MM) derived
purely from the file's mtime for the cached system-prompt tree, so the rendered
block is byte-identical across sessions and only changes when a file actually
changes. Scoped via a new internal AssembleOptions.ageMode:
- buildWorkspaceTree (cached system prompt) -> "absolute"
- buildDirectoryTree (read-tool output, not cached) -> "relative" (unchanged)
renderNode now takes a per-pass age formatter instead of reading Date.now()
directly.
Measured on a local llama.cpp server (single user, prompt cache on): with a
file whose age ticks between two back-to-back sessions, the unpatched build
re-prefills the full prefix on session 2 (27,124 prompt tokens, 46s); with this
change session 2 is a cache hit (13 tokens, 2s). Existing tests are unaffected
(they assert on filenames/order/elision, not on age strings); two regression
tests added.
- Added beforeEach and afterEach hooks in mnemopi tests to set and clear MNEMOPI_NO_EMBEDDINGS so embeddings are skipped during those runs.
- Updated the bun-install cache script to archive only node_modules paths that exist as directories.
- Applied title-casing to `normalizeGeneratedTitle` outputs using a new internal helper.
- Adjusted tiny text and title generator tests to assert the new title-cased results.
- Added a new built-in `title` model role with `hidden` metadata and updated role definitions and schema.
- Updated title generation to resolve models in `title`, `commit`, then `smol` order and added test coverage for that precedence.
- Filtered hidden roles from selector badges and documented the new built-in role in model/settings docs.
- Updated CI dependency install flow to share bun cache orchestration across jobs.
- Added RustFS-backed bun cache restore/save script keyed by bun.lock hash.
`ToolResultContent` is not exported from `@oh-my-pi/pi-ai`; `ToolResultMessage.content` is `(TextContent | ImageContent)[]`. Test only needs text, so type the content array as `TextContent[]`.
`#checkTodoCompletion` used to append a `<system-reminder>` and then call
`#scheduleAgentContinue`, so a text-only acknowledgement ("paused at your
instruction") triggered another `agent_end` that re-ran the same check and
fired the next reminder — counter ticked 1/3 → 2/3 → 3/3 inside a single user
pause without any user input. The user perceived three back-to-back reminders
appear from nowhere; the agent felt implicit pressure to invent busy-work or
take destructive ops to silence the loop.
Added `#todoReminderAwaitingProgress`: a reminder sets it, any `toolResult`
(real tool-level progress) or a new user prompt clears it, and
`#checkTodoCompletion` stays silent while it is set. Reset alongside
`#todoReminderCount` on user prompts, session reset, handoff, and the no-op
short-circuits in `#checkTodoCompletion` so the field never gets stuck.
Escalation through `todo.reminders.max` still works when the agent makes
tool-level progress between stops — that is the case the cap was designed for.
Regression coverage in `agent-session-todo-reminder-loop.test.ts` drives a
mocked `agent.continue` to mirror the bug-reported model behaviour and pins
the contract: exactly one reminder per user pause when the agent only
acknowledges; re-escalation when the agent actually calls a tool between stops.
Fixes#2590
- Added a setup-system-deps action with preloaded-runner guards and apt fallbacks.
- Updated CI workflows to download Linux x64 native artifacts and gate on native job success.
- Renamed coding-agent fast mode to singleton in scripts and test partitioning logic.
- Added settings test-state begin/restore helpers with recursive cleanup in affected tests.
Two CustomEditor/streaming-preview tests timed out under bun test --parallel
because they reset+reinitialised the process-global Settings singleton in
beforeEach. 50+ other test files do the same dance, so under parallelism the
proxy points at whichever instance won the latest race rather than the one
under test (issue #2582).
- Added CustomEditor#magicKeywordsEnabledOverride: an instance-level test/host
injection that short-circuits the global lookup at the read site.
Production wiring still reads from Settings.
- Switched the magic-keyword shimmer-disabled test to set
magicKeywordsEnabledOverride directly instead of mutating the singleton.
- Removed the gratuitous reset+init from the 'streaming tool call preview
height (bounded across renderers)' describe -- bash/ssh/eval pending
previews never read settings.*; only initTheme() (idempotent) is kept.
Fixes#2582
The tts/stt suite spawns sherpa-onnx worker subprocesses that hang on the
headless CI runner (zero-output stall → SIGTERM), and the resulting event-loop
starvation tipped real-time TUI tests (streaming-preview, custom-editor shimmer,
ask timeouts) past bun's 5s default. Remove the tts/stt tests and give the
timing-sensitive tests an explicit 30s timeout.
- Added a lifecycle test case verifying executePython retries execution with a fresh kernel when the prior session kernel dies during run.
- Removed the now-redundant python-executor-session.test.ts file that defined session lifecycle tests.
Read vLLM max_model_len and OpenAI-compatible context_length metadata during model discovery, route providers.vllm.baseUrl into built-in discovery before cached models exist, and avoid sending local placeholder bearer tokens.
Scope the vLLM model cache to the discovery base URL so endpoint changes refetch immediately, and add focused regression coverage for configured and built-in vLLM discovery.