- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
- Assert the eval tool hides disabled backends from the model-facing wire
schema (language enum + field descriptions), summary, and description by
default (rb/jl off), and advertises them once enabled — including the
enabled-subset case.
- Update the env-flag fallback test for the new rb/jl opt-in defaults and
extend the env guard to PI_RB/PI_JL so the suite is shell-independent.
- Note the opt-in default and dynamic advertising in the changelog.
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
- Centralized draft state and image management by migrating fields from context to the CustomEditor component.
- Standardized transcript row construction by introducing shared helpers for background jobs, IRC traffic, and file mentions.
- Refactored redundant UI logic and helper functions into reusable utility modules to streamline message submission and component rendering.
- Standardized event handler types by consolidating lifecycle definitions into a shared module while maintaining public API stability.
- Removed the `readHashLines` setting to consolidate hashline display logic.
- Simplified `resolveFileDisplayMode` to derive hashline visibility solely from the active edit mode.
- Added automatic cleanup of the `readHashLines` key from existing configuration files.
- Implemented the Devin inference provider, including OAuth flow with PKCE, Connect protocol integration, and streaming support for chat requests.
- Integrated comprehensive Protobuf-based service definitions and generated TypeScript clients for Devin's API infrastructure, including model management and workspace operations.
- Updated the AI and Catalog modules to support dynamic model discovery, provider-specific configuration, and authentication.
- Standardized tool call arguments as `Record<string, unknown>` across provider implementations to ensure type safety.
- Refined obfuscation logic to use granular, typed transformations instead of generic object traversal.
- Enforced an 8-character minimum for secret patterns and restricted redaction to user-authored content to prevent false positives.
- Preserved system prompts, tool schemas, and opaque remote replay data to maintain provider context and data integrity.
- Integrated protected snapshot exports with targeted redaction to safeguard sensitive information in shared sessions.
Re-added the isConfigured guard inside applyDefaultSettingOverrides so the
runtime override applied at RPC/ACP startup only fills holes — explicit
embedder/project/--config/global values now survive for task.isolation.{mode,
merge,commits}, task.{eager,batch,maxConcurrency,maxRecursionDepth,disabledAgents,
agentModelOverrides}, memory.backend, memories.enabled, advisor.{enabled,
subagents,syncBacklog,immuneTurns}, plus the RPC-only async.{enabled,maxJobs}
and bash.autoBackground.{enabled,thresholdMs}. The guard was originally added
for #2598 and had quietly regressed, so the override re-asserted the schema
default and clobbered every explicit value — matching the contract problem
#2828 already fixed for the todo paths.
Replaced the prior 'default-disables advisor' test (which codified the
clobbering as the contract) with a comprehensive 'honors explicit
host-defaulted settings' regression covering every contested path across
rpc/rpc-ui/acp.
Fixes#3207
- Implemented `waitForSelector` and `waitForNavigation` methods in the browser tool API.
- Introduced per-operation fail-fast budget management with dynamic timeout clamping.
- Added validation for selector engines to reject unsupported Playwright-only platform features.
- Standardized error handling to provide descriptive, named timeouts for stalled browser operations.
- Updated model detection to exclude WebP format for Codex-based providers, which do not support it.
- Forced the image resize pipeline to encode to PNG or JPEG for incompatible models to resolve transmission errors.
- Added test coverage to verify that WebP images are re-encoded when using the Codex Responses backend.
Added clearOptimisticUserMessage and replaceOptimisticUserMessage stubs to the EventController message_start test double so the optimistic-submission case no longer throws when the controller reconciles the signature.
Fixes#3199
Skipped optimistic replacement when a user message_start matches another recorded local submission, preserving the pending prompt bubble until its own expanded event arrives.
Added coverage for the queued-message drain race between startPendingSubmission and prompt dispatch.
Fixes#3199
Updated optimistic replay to track the replacement component handles created during transcript rebuilds, so expanded slash prompts still replace the raw replayed message.
Extended the regression test to cover the rebuild window called out in review.
Fixes#3199
Replaced raw optimistic slash-command transcript entries with the canonical user message emitted by AgentSession when prompt expansion changes the text.
Added coverage for prompt-template expansion reconciliation so the transcript keeps one expanded user message.
Fixes#3199
When the user set `providers.tinyModel` to a local key, `generateSessionTitle`
still raced local against the online `smol` path with a 10 s timeout and
silently fired the online request whenever the local worker returned `null`
(unknown key, model not downloaded, transformers.js failure). The online path
resolves the `smol` role through `priority.json` (haiku → flash → mini → …);
with an `OPENROUTER_API_KEY` picked up from env, that silently billed
OpenRouter without consent.
Drop the race entirely for local choices: honor the user's setting, log a
warning on local failure, leave the session untitled. The `raceFirstNonNull`
helper and `TITLE_LOCAL_FALLBACK_DELAY_MS` had no other consumer and are
removed; the obsolete \"silently bills online when local fails\" tests are
flipped into regressions that lock the no-fallback contract, including the
unknown-key path (e.g. \"ollama:gpt-oss\") which previously also leaked
straight through to the online billing path.
Fixes#3187
Mapped OpenCode MCP array commands to stdio command plus args and accepted environment as the provider-native env key.\n\nAdded regression coverage for array command normalization, environment mapping, env fallback, and empty args omission.\n\nFixes #3180
- sdk-mcp-discovery: `find` became an essential tool (2eef88978), so it can no longer be hidden/rediscovered under `tools.discoveryMode: all`. Switch the discoverable-tool assertions to `search` (still `loadMode: discoverable`).
- agent-session-concurrent: the agent loop now drops tool calls that never reached `toolcall_end` from an aborted turn (0890b2be6, partial args are unsafe to replay). Emit `toolcall_end` before the TTSR rule-driven abort so the labeled placeholder result is minted.
- Inlined the temporary model status formatting logic directly into the controller.
- Removed the unused `formatTemporaryModelStatus` utility function and its associated test.
- Update `shouldRetry` to treat `EISDIR` and `ENOTDIR` as terminal errors, preventing unnecessary retries when encountering Git reference directory conflicts.
- Add a test suite to verify graceful resolution of branches in scenarios where a packed ref conflicts with a directory path in the filesystem.
- Added `includeWorkspaceTree` configuration setting to optionally render the workspace directory tree.
- Configured the system prompt template and SDK to respect this toggle, allowing users to disable the tree to prevent prompt cache invalidation.
- Promoted `write` and `find` tools to `essential` status to ensure they are always available regardless of discovery mode.
- Updated `DEFAULT_ESSENTIAL_TOOL_NAMES` to include these tools by default.
- Updated documentation and tests to reflect the change in default essential tool availability.
Fixes#3165
- Synchronize tool arguments with the component state upon receipt of `tool_execution_start` to ensure visual consistency when final update events are missed.
- Terminate active argument reveal streams to prevent late ticks from overwriting valid, fully-materialized tool arguments with stale partial data.
- Add test coverage to verify that tool UI components render finalized arguments even in the absence of intermediate streaming updates.
- Implemented `tab.ariaSnapshot()` to capture and represent page structures as ARIA-tree YAML.
- Introduced `tab.ref()` and ref-based selector parsing to enable precise element interaction via unique ARIA identifiers.
- Integrated automated script bundling for cross-environment evaluation of ARIA snapshot logic.
- Updated browser action methods to resolve and target elements using ARIA-ref handles.
Tracked current-registry built-in provenance through AgentSession so plan mode
only force-activates the built-in write implementation. Extension or SDK tools
that shadow the name `write` stay inactive, preserving plan mode's read-only
contract through the built-in write/edit guard.
Added a regression that registers a shadowing write tool without built-in
provenance and verifies plan mode does not activate it.
Plan-mode entry only added `resolve` to the active toolset, so when
`tools.discoveryMode: "all"` left `write` hidden behind
`search_tool_bm25` the agent was stuck with `edit` to create the plan
file — which fails on a non-existent path and stalls the planning
turn. `#enterPlanMode` now augments the active set with both `resolve`
and `write` whenever the registry built them, matching what
`plan-mode-active.md` instructs the model to use, and `#exitPlanMode`
still restores the pre-plan toolset verbatim.
Fixes#3165
- Mocked stdout rows for the AgentTranscriptViewer tests to ensure consistent terminal height.
- Added an afterEach cleanup to restore the original property descriptor for process.stdout.rows.
- Replaced the `/debug dump-next-request` command with an updated `/dump` command that exports LLM request context to JSON sidecar files.
- Removed persistent debug path state and manual path configuration in favor of automated generation.
- Updated session logic to handle serializing LLM request context to temporary directories.
- Refactored testing suites to remove path-based debug tests and verify dynamic request file generation.
`omp --approval-mode=yolo acp` was rewritten to `launch --approval-mode=yolo
acp`, swallowing `acp` as a launch prompt so the yolo override never reached the
ACP command path (the ACP permission gate from #2097 stayed in always-ask).
`resolveCliArgv` only inspected `argv[0]`, so any leading global option flag hid
the real subcommand. It now scans past leading flags using the launch parser's
value-consumption contract (a flag's value is never mistaken for the subcommand,
e.g. `--model acp`) and hoists the recognized subcommand to the front with the
flags preserved as its own argv. Genuine launch prompts are untouched.
The flag value-consumption rule is factored into `cli/flag-tables.ts`
(`flagConsumesValue` + the shared `isUnknownLongValueCandidate`) so the resolver
and the profile bootstrap share one source of truth.
Fixed permission mode not respected in ACP mode ([#2970](https://github.com/can1357/oh-my-pi/issues/2970))
Fixes#2970
The background message reader matched every incoming message against the
pending client-request map by id before checking for a `method`. Server
request ids live in the server's own id space and routinely collide with
the client's in-flight request ids, so a server-originated
`workspace/configuration` pull whose id matched a pending request (e.g. a
basedpyright pull landing while a `documentSymbol` request with the same
id was open) was swallowed as a bogus response: the client request
resolved with `undefined` and the pull was never answered, wedging
servers that gate analysis on configuration.
Route any message carrying a `method` as a server request (or
notification) before id-matching responses, so every config pull is
answered under lazy init -- parity with the warmup/reload path, which
escaped the bug only because it issues no concurrent semantic request
while the cold-start pulls drain. lsp.lazy default is unchanged.
Fixes#3001
The plugins docs advertise `omp list`/`omp remove` as top-level commands,
but only `omp install` is registered. `resolveCliArgv` rewrote any
unregistered first-arg to `launch`, so `omp list` silently started an
interactive agent session with "list" as the LLM prompt instead of
managing plugins.
Reserve the bare, documented-but-unregistered plugin verbs `list` and
`remove` with a helpful hint pointing at the real `omp plugin list` /
`omp plugin uninstall <name>` commands (same `extensions` mechanism /
same class as the `install` leak fixed in #1496/#1498). Multi-word
prompts that merely begin with those words still route to `launch`, so
genuine prompts are unaffected.
Fixes#2935