- Added the `--external-thinking` CLI flag alongside model capability checks to gate external thinking tool availability.
- Updated Anthropic and Google transports to honor `forceReasoningOff` for native thinking-off controls.
- Renamed the `thoughts` property and parameter to `notes` across think fixtures, tools, and tests.
- Updated system prompt instructions and test suites to verify transport-specific thinking and tool activation.
- Added support for external thinking and forced reasoning disablement across AI provider options and request transformers.
- Implemented the private scratchpad think tool along with its renderer, system prompt rules, and schema configuration.
- Updated agent session management and SDK tools to support dynamic runtime activation of the think tool via the externalThinking setting.
- Added comprehensive unit tests covering reasoning fallbacks, tool activation, and rendering behavior.
- Define a centralized `USER_AGENT` constant in `@oh-my-pi/pi-utils` formatted as `omp/<version>`.
- Replace hardcoded and platform-specific user agent strings across AI providers, catalog scrapers, tools, and search providers with the unified `USER_AGENT`.
- Add unit tests for update-cli binary release distribution gating.
Key hints resolved modifier tokens through a static, platform-agnostic label map with no `super` entry, so on macOS the shipped `super+v` paste default rendered 'Super+V' (no such key on a Mac) and `alt` always rendered 'Alt' instead of 'Option'. The static /hotkeys navigation rows were hardcoded with macOS 'Option'/'Cmd' names on every platform, so Linux/Windows users saw 'Cmd+Left'.
Modifier labels are now platform-aware: on darwin `alt` renders 'Option' and `super` renders 'Cmd'; every other platform keeps 'Alt'/'Super'. The platform is resolved through a single seam (setKeyHintPlatform/keyHintPlatform) mirroring the TUI's setKittyProtocolActive, keeping hint output deterministic in tests without mutating process.platform. The static /hotkeys rows use the same convention and drop the macOS-only Cmd line-start/end fragments (which map to no binding) off darwin.
Fixes#8235
Address review feedback on #8186:
- add ty.toml to ty rootMarkers so ty.toml-only projects pass the
loadConfig marker gate (Codex P2, roboomp should-fix)
- add ty.toml to PYTHON_ROOT_MARKERS so project-local .venv/bin/ty
resolution runs for ty.toml-only projects (roboomp should-fix)
- add a ty.toml-only regression test covering both the detection
gate and the local venv bin resolution path
- move the changelog entry from the released [16.3.11] section into
[Unreleased] (roboomp should-fix)
Add ty (ty server) to defaults.json behind pyright/basedpyright/pylsp
in primary selection order and ahead of ruff (linter). ty uses the
generic LSP client path with no adapter code; settings.ty pass-through
works via the existing settings field.
- defaults.json: insert ty entry between pylsp and ruff
- lsp-regressions.test.ts: 4 tests (selection order for .py/.pyi,
auto-detect via $which + pyproject.toml, coexistence with ruff)
- CHANGELOG.md: Added entry under [Unreleased] referencing #4617Closes#4617
- Expanded web search and fetching providers with robust parsing, authentication storage integration, and response validation.
- Added support for new configurations including SearXNG safesearch, Cloudflare AI Gateway endpoints, and dynamic Firecrawl base URLs.
- Implemented comprehensive test suites covering error handling, content filtering, and provider-specific response behaviors.
Why:
Local timeout cleanup removes pending state but leaves the language server
working on an abandoned request.
Changes:
- Send $/cancelRequest after an issued request reaches its client timeout.
- Preserve the existing timeout error and pending-request cleanup.
Evidence:
- The workspace-readiness regression observes cancellation for the dropped
status request before polling succeeds.
Refs #8116
Why:
The module-level idle interval remains referenced after every client shuts
down and keeps short-lived hosts alive.
Changes:
- Stop the idle checker before draining active and pending clients.
- Preserve explicit reconfiguration through setIdleTimeout().
Evidence:
- A child-process regression probe exits after shutdown with the checker
configured.
Refs #8115
While the agent worked through a plan, every sub-todo rendered unchecked
no matter how far along the run was: the phase header highlighted, the
task rows below it looked untouched. Three separate causes, all on the
collapsed path that is the default view.
`selectCollapsedTodos` dropped every closed row while a phase held open
work, so finishing a task only ever *removed* a line — the panel never
rendered a checked box until the whole phase settled. That also made the
card's completion animation dead code: `details.completedTasks` drives a
14-frame strike reveal at 65ms with a component render per tick, against
a row the viewport had already discarded. The existing animation test
missed it by asserting on `expanded: true`.
The viewport now keeps the newest closed task as a checked lead row,
additive to the open-task cap so it never evicts open work, and the
strike sweep lands where users actually see it.
Second, the card gave a `done/total` count to every collapsed untouched
phase but not to the active one, so the phase being worked in was the
single phase reporting no progress. Extracted `formatPhaseProgress` and
put it on every phase header.
Third, the todo auto-clear (`tasks.todoClearDelay`, default 60s) armed
on any list holding a closed task and physically deleted those tasks
from the HUD's copy. An in-flight phase at `3/4` silently became `0/1`
sixty seconds later, fully-closed phases vanished, and stage roman
numerals renumbered off the filtered index — until the next `todo` call
restored the real snapshot. It now fires only once the whole list is
settled, which is the case the setting exists for; the walking viewport
already hides closed rows while work remains.
Progress counters also count closed tasks rather than only completed
ones. The viewport hides abandoned tasks too, so counting only
completions left a phase reading permanently stuck.
Preserved zero-width readiness and wait matches across the daemon wire protocol, and isolated malformed completion events from unrelated pending RPCs.
Fixes#7908
When an xd:// device is dispatched through the write tool, the outer
approval gate now consults tools.approval.<deviceName> before falling
back to tools.approval.write. This lets users scope allow/deny/prompt
to a single device mount without changing the blanket write tool policy.
The write tool's approval function returns { tier, policyKey: deviceName }
for xd:// device dispatches. resolveApproval uses the policyKey to look
up the user override on the device name, falling back to the invoking
tool's own policy when the device has none configured.
Adds:
- ToolApprovalDecision.policyKey field (optional, additive)
- policyKey-aware lookup in resolveApproval and requiresApproval
- Updated error messages naming the correct config key
- Unit tests for policyKey resolution and WriteTool integration
Fixescan1357/oh-my-pi#7923
The `cd <path> && ...` extractor matched everything up to the first `&&`
with a greedy regex, so a redirect or extra argument before the `&&` was
swallowed into the structured cwd. `cd /tmp 2>/dev/null && echo ok` became
cwd `/tmp 2>/dev/null`, which failed fs.stat and killed the command before
the shell ran.
Replace the regex with `extractLeadingCdTarget`, a quote/escape-aware
scanner in shell-tokenize.ts that captures exactly one path token and
bails (leaving the command for the shell) when anything else — a redirect,
extra argument, shell expansion, or a non-`&&` separator — precedes the
top-level `&&`.
Fixes#7883
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
Closed worker and cmux run signals before yielding for floating-rejection drainage. Stale promise continuations can no longer begin page navigation after evaluated code returns.
Classified only marked browser failures and evaluated-run stack frames as run-owned rejections. Unrelated tab-worker failures now remain on the worker guard's fatal path.
Used a run-scoped Promise subclass instead of mutating native combinator methods. Evaluated code can now freeze its Promise constructor without breaking cleanup or later browser runs.
Observed Promise.all and Promise.race results derived from browser calls during each evaluated run. User catch continuations that rethrow browser failures now fail the owning run without changing native await behavior.
Logged late user continuation failures in cmux runs and delayed worker rejection folding until request-interception cleanup completed. This closes both windows where missing awaits could be silently dropped.
Logged user continuation rejections that settle after their browser run has ended. This preserves the completed result while making missing awaits visible instead of silently dropping them.
Tracked whether user continuation callbacks create each descendant rejection. Browser errors that user code rethrows now fail the owning run instead of being contained as propagated helper failures.
Scoped browser-error markers to each run and contained only propagated browser failures. Routed floated user continuations into failed runs and added worker coverage for native await plus every continuation method.
Observed every browser facade continuation so fire-and-forget helper
timeouts cannot wedge or kill a tab worker. Preserved native Promise
identity for callers and test matchers.