Reviewer flagged that ToolExecutionComponent.isTranscriptBlockCommitStable\ntreated any block with a defined #result as commit-stable, so a long\nstreaming SSH partial render could sit on its stable pending header until\nderiveLiveCommitState's stable-prefix ratchet promoted it to native\nscrollback, then the final SSH glyph render would land below and strand a\nduplicate pending header above the settled frame.\n\n- Added ToolRenderer.provisionalPartialResult opt-in for renderers whose\n partial-result chrome differs from the final render.\n- Honored it in ToolExecutionComponent.isTranscriptBlockCommitStable so the\n block stays commit-unstable while isPartial holds and flips stable as soon\n as the result settles.\n- Set provisionalPartialResult: true on the SSH renderer.\n- Added test/tools/ssh-commit-stability.test.ts covering the partial/final\n transition and the non-regression for bash.\n\nFixes #3177
- Kept SSH result frames in the pending state while partial output is streaming.\n- Added renderer coverage for the partial-to-final SSH header transition.\n\nFixes #3177
- sdk-mcp-discovery: `find` became an essential tool (2eef88978), so it can no longer be hidden/rediscovered under `tools.discoveryMode: all`. Switch the discoverable-tool assertions to `search` (still `loadMode: discoverable`).
- agent-session-concurrent: the agent loop now drops tool calls that never reached `toolcall_end` from an aborted turn (0890b2be6, partial args are unsafe to replay). Emit `toolcall_end` before the TTSR rule-driven abort so the labeled placeholder result is minted.
- Implement ID-based guards for Codex WebSocket frames to reject stale or unauthorized interleaved frames from previous turns.
- Check sequence numbers within Codex responses to detect and handle out-of-order frame delivery.
- Restrict tool result consumption to occurrences appearing after the associated tool call, preventing the reuse of orphaned results from earlier conversation turns.
- Fixed streaming output being lost from native scrollback when an unstable "barrier" block preceded it.
- Adjusted the engine commit logic to ensure all scrolled-off content reaches history, treating the `windowTop` as the persistent commit floor.
- Updated the committed-prefix logic to perform range-aware auditing, ensuring forced-overflow rows are accurately re-anchored rather than dropped upon barrier finalization.
- Inlined the temporary model status formatting logic directly into the controller.
- Removed the unused `formatTemporaryModelStatus` utility function and its associated test.
- Refactored committed prefix auditing into distinct byte-stable, durable, and forced-overflow zones.
- Integrated a hard scan detector and refined boundary management to prevent data loss during commit-unstable barrier shifts.
- Implemented robust regression testing via the streaming-scrollback-defer harness to verify frame continuity.
- Updated audit logic to independently manage boundaries, ensuring forced-overflow rows remain correctly finalized.
- Update `shouldRetry` to treat `EISDIR` and `ENOTDIR` as terminal errors, preventing unnecessary retries when encountering Git reference directory conflicts.
- Add a test suite to verify graceful resolution of branches in scenarios where a packed ref conflicts with a directory path in the filesystem.
Prioritized rate-limit wording before auth wording in auth-gateway error classification so DashScope throttling messages that contain unauthorized do not invalidate credentials.
Fixes#3172
Classified OpenCode Go 401 Insufficient balance responses as quota exhaustion so usage-limit retry, credential rotation, and fallback chains engage. Added regression coverage for both parsing and usage-limit detection.\n\nFixes #3169
- Added `includeWorkspaceTree` configuration setting to optionally render the workspace directory tree.
- Configured the system prompt template and SDK to respect this toggle, allowing users to disable the tree to prevent prompt cache invalidation.
- Promoted `write` and `find` tools to `essential` status to ensure they are always available regardless of discovery mode.
- Updated `DEFAULT_ESSENTIAL_TOOL_NAMES` to include these tools by default.
- Updated documentation and tests to reflect the change in default essential tool availability.
Fixes#3165
- Added error handling to docker container and network removal commands.
- Print yellow warning messages to stdout when cleanup commands fail to exit cleanly.
- Synchronize tool arguments with the component state upon receipt of `tool_execution_start` to ensure visual consistency when final update events are missed.
- Terminate active argument reveal streams to prevent late ticks from overwriting valid, fully-materialized tool arguments with stale partial data.
- Add test coverage to verify that tool UI components render finalized arguments even in the absence of intermediate streaming updates.
- Implemented `tab.ariaSnapshot()` to capture and represent page structures as ARIA-tree YAML.
- Introduced `tab.ref()` and ref-based selector parsing to enable precise element interaction via unique ARIA identifiers.
- Integrated automated script bundling for cross-environment evaluation of ARIA snapshot logic.
- Updated browser action methods to resolve and target elements using ARIA-ref handles.
- Replaced `partial-json` with a high-performance `RelaxedJson` parser capable of handling LLM-specific malformations.
- Enabled native support for unquoted keys, single quotes, comments, and Python literals (`True`/`False`/`None`).
- Improved streaming robustness by rejecting non-finite numbers (NaN/Infinity) and rolling back incomplete trailing tokens.
- Optimized tool-call execution by ensuring malformed or truncated arguments are skipped instead of surfaced as undefined/NaN.
- Added comprehensive test suite for both strict repair parsing and partial streaming scenarios.
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
- Added --env flag to forward host environment variables into Harbor containers.
- Implemented collectForwardEnv to sanitize and inject DI-specific variables via OMP_TB_FORWARD_ENV.
- Introduced standalone terminal-bench cleanup command to prune stale Docker containers and networks.
- Refactored Docker cleanup logic into a reusable utility function for consistent lifecycle management.
- Replace loose type check on the operation schema with an explicit equality check.
- Add comments clarifying the wire optimization behavior for const-union to enum collapse.
- Kept provider discovery defaults authoritative for api and provider when layering bundled or models.dev metadata.
- Added a LiteLLM regression covering a deepseek-v4-flash collision with the ollama-cloud catalog.
Fixes#3162
- Mocked stdout rows for the AgentTranscriptViewer tests to ensure consistent terminal height.
- Added an afterEach cleanup to restore the original property descriptor for process.stdout.rows.
- Overhauled stealth spoofing mechanisms for WebGL, screen dimensions, Web workers, and iframe contexts using prototype-aware injection.
- Centralized function string representation patching to improve mimicry of native browser behavior across global objects.
- Patched puppeteer-core to remove detectable evaluation markers and implement lazy, pull-style execution context management.
- Enabled support for capturing LLM request JSON dumps and adjusted launcher flags to improve organic request patterns.
- Implemented `collapseConstUnionAnyOf` to simplify `anyOf` branches into compact `enum` structures when values share a scalar type.
- Preserved existing schema annotations by skipping collapse when per-variant descriptions differ or conflict with root descriptions.
- Added comprehensive unit tests in `schema-wire.test.ts` to verify lossless transformation of described literal unions.
- Included validation for user-message history payload integrity within `openai-responses-history-payload.test.ts` to ensure message order remains consistent during processing.
- Replaced the `/debug dump-next-request` command with an updated `/dump` command that exports LLM request context to JSON sidecar files.
- Removed persistent debug path state and manual path configuration in favor of automated generation.
- Updated session logic to handle serializing LLM request context to temporary directories.
- Refactored testing suites to remove path-based debug tests and verify dynamic request file generation.
- Updated array normalization to unwrap double-encoded object keys within JSON-stringified arrays.
- Added a regression test to verify successful coercion of double-encoded objects inside string-encoded union types.
- Implemented recursive key normalization to unwrap object keys incorrectly serialized by LLMs (e.g., `{"\"op\"": "done"}`).
- Integrated normalization pass into `validateToolArguments` execution flow to prevent keys from being dropped by schema repairs.
- Added support for multi-layered encoding peels and protection against sibling key collisions during renaming.
- Handled edge cases where double-encoded keys are exposed only after internal JSON-string payload parsing.
- Added comprehensive testing for various scenarios, including nested objects and string-wrapped JSON payloads.
- Update `llama.cpp` base URL to remove the `/v1` suffix.
- Consolidate local provider token validation in `coding-agent` using a `Set`.
- Add test coverage to ensure catalog model IDs are preserved verbatim on the wire.
`omp --approval-mode=yolo acp` was rewritten to `launch --approval-mode=yolo
acp`, swallowing `acp` as a launch prompt so the yolo override never reached the
ACP command path (the ACP permission gate from #2097 stayed in always-ask).
`resolveCliArgv` only inspected `argv[0]`, so any leading global option flag hid
the real subcommand. It now scans past leading flags using the launch parser's
value-consumption contract (a flag's value is never mistaken for the subcommand,
e.g. `--model acp`) and hoists the recognized subcommand to the front with the
flags preserved as its own argv. Genuine launch prompts are untouched.
The flag value-consumption rule is factored into `cli/flag-tables.ts`
(`flagConsumesValue` + the shared `isUnknownLongValueCandidate`) so the resolver
and the profile bootstrap share one source of truth.
Fixed permission mode not respected in ACP mode ([#2970](https://github.com/can1357/oh-my-pi/issues/2970))
Fixes#2970
The background message reader matched every incoming message against the
pending client-request map by id before checking for a `method`. Server
request ids live in the server's own id space and routinely collide with
the client's in-flight request ids, so a server-originated
`workspace/configuration` pull whose id matched a pending request (e.g. a
basedpyright pull landing while a `documentSymbol` request with the same
id was open) was swallowed as a bogus response: the client request
resolved with `undefined` and the pull was never answered, wedging
servers that gate analysis on configuration.
Route any message carrying a `method` as a server request (or
notification) before id-matching responses, so every config pull is
answered under lazy init -- parity with the warmup/reload path, which
escaped the bug only because it issues no concurrent semantic request
while the cold-start pulls drain. lsp.lazy default is unchanged.
Fixes#3001
The plugins docs advertise `omp list`/`omp remove` as top-level commands,
but only `omp install` is registered. `resolveCliArgv` rewrote any
unregistered first-arg to `launch`, so `omp list` silently started an
interactive agent session with "list" as the LLM prompt instead of
managing plugins.
Reserve the bare, documented-but-unregistered plugin verbs `list` and
`remove` with a helpful hint pointing at the real `omp plugin list` /
`omp plugin uninstall <name>` commands (same `extensions` mechanism /
same class as the `install` leak fixed in #1496/#1498). Multi-word
prompts that merely begin with those words still route to `launch`, so
genuine prompts are unaffected.
Fixes#2935
- Added `MOONSHOT_BASE_URL` environment variable support to redirect requests to the Kimi China platform.
- Enabled `KIMI_API_KEY` as an alias for the moonshot provider to improve user experience.
Fixes#2883
Updated /mcp enable and /mcp disable so they connect or disconnect only the named server instead of reloading every MCP server in the session. Added regression coverage for both toggle directions and updated the coding-agent changelog.
Fixes#3157
- Updated the test configuration to specify an explicit compaction strategy.
- Enabled context-full compaction to match expected behavior during tests.
- Updated `render`, `renderMany`, and native snapcompact methods to return promises, ensuring scalable async execution.
- Refactored `transformProviderContext` and `buildSideRequestContext` to support asynchronous operations in agent loops.
- Integrated `Promise.all` for improved concurrency when processing frame rendering and rendering batch operations.
- Updated all internal call sites, SDK hooks, and test suites to accommodate the asynchronous API signatures.
- Removed the logic that forced sequential tool calling when custom tools were present in both Codex and OpenAI providers.
- Exported `buildTransformedCodexRequestBody` for testing purposes.