- Added `fetchUsageData` and fleet token summation to aggregate token burn across clients.
- Changed window grouping logic from window labels to limit IDs for distinct separation.
- Updated `ProviderWindowInsight` properties and added tests for label suffixing and token summation.
The ask-card note renderer (#8786) grew tool-views.generated.js by 126 bytes; CI regenerates the asset so the byte-pinned template contract (bytes/chars/sha256) moved. Values verified identical between CI and a fresh local gen:tool-views.
bazel test sandboxes stage only the crate's declared inputs, so the repository sweep found zero sources and tripped its own evidence guard ('corpus sample too small to be evidence: 0'). Skip on a missing or empty corpus; any non-empty scan still enforces the >40-file evidence floor.
`resolveCliArgv` hoisted a subcommand hidden behind leading global launch
flags to the front and forwarded those flags to the subcommand's own
parser (#2970). Launch-shaped commands (`launch`/`acp`) share the launch
flag surface, but strict-parsing subcommands like `update` declare only
their own flags, so a leading `--cwd` reached `node:util.parseArgs` and
threw `Unknown option '--cwd'`. This bit users whose shell alias/wrapper
runs `omp --cwd <dir> update`.
Leading launch-global flags are now stripped when the hoisted subcommand
is not launch-shaped, and still forwarded for `launch`/`acp`. Shared the
launch-command set with the profile bootstrap to keep one source of truth.
Fixes#8891
parseFileDiffs split the captured `git diff --cached --binary` on
"
diff --git ", consuming the newline that terminates each file block, and
patch.join ended with `.replace(/\n+$/, "")`. Both dropped the blank line
that terminates a `GIT binary patch` block, so rebuilding a split-commit
patch produced a corrupt binary patch rejected by `git apply --binary`.
Split on a line-start lookahead so blocks keep their terminators verbatim,
and concatenate join parts without stripping trailing newlines. Both
trailing and mid-diff binary blocks now round-trip byte-exact.
Fixes#8899
GMI Cloud's /v1/models returns only bare {id} rows, so dynamic discovery
resolved every model except the single bundled DeepSeek-V4-Flash seed with
a null context window, zero pricing, and no reasoning/thinking config.
Give the gmi-cloud mapper the same cross-provider canonical fallback that
SiliconFlow uses for the identical open-weight models: recover context
window, output limit, reasoning, and thinking ladder from the bundled
reference index while never borrowing another provider's pricing.
Fixes#8890
`/mcp reauth <name>` probed the server with `{ oauth: false }` and treated a
successful unauthenticated `initialize` as proof that OAuth was unnecessary,
hard-erroring with "Server connection succeeded without OAuth; reauthorization
is not required." Per the MCP spec a server may allow unauthenticated
`initialize` while requiring a bearer token for `tools/call`, so this left no
way to acquire a credential for such servers.
When the handshake succeeds without an in-band tool challenge, fall back to
`discoverOAuthEndpoints(config.url)` and proceed with the flow if the server
advertises OAuth metadata; only refuse when no OAuth endpoint is discoverable.
Fixes#8922
Unanchored regex mapped BCP-47 script subtags to bogus regions
(lang:zh-hans -> location=HA) and prefix-matched 3+ letter codes
(lang:eng -> language=en). Require a subtag boundary, matching the
Perplexity provider's parsing.
Cursor advertises Grok 4.5/4.6 as per-effort sibling ids
(cursor-grok-4.6-low|-medium|-high|-xhigh plus -fast variants), but
VARIANT_COLLAPSE_TABLES had no cursor entry, so the model hub showed 14
unrouted siblings instead of one logical model with effort routing.
GetUsableModels ships no thinkingDetails and the bundled references read
reasoning:false, so the picker also treated them as non-reasoning.
- Add CURSOR_VARIANT_COLLAPSE_TABLE folding each service-tier lane
(standard + -fast) into one logical model with effort routing onto the
live wire ids, mirroring Devin's grok-4-5 collapse.
- Rename the generic devinTierFamily/DevinTierRoutes helper to
tierFamily/TierRoutes now that both Devin and Cursor tables use it.
- Mark versioned cursor-grok-<version> ids as reasoning during discovery
(grok-code-* coding models stay non-reasoning).
Fixes#8803
The Roles panel renderer looped from the first row with a hard height cap and no scroll offset, so roles and model-keyed fallback chains past the visible height were never drawn and unreachable by keyboard, with no truncation indicator. Unlike the sibling provider list (model-browser windows via #windowStart/#ensureSelectedVisible), the Roles panel had no equivalent.
Window #renderRolesView around #roleIndex via #ensureRoleVisible, offset mouse hit-testing by the scroll start bounded to the visible count, and draw an up/down '+N more' hint when the list is clipped.
Fixes#8817
ExtensionRunner.emitAfterProviderResponse accepted the response model but
discarded it, calling createContext() with no model. Response-scoped hooks
therefore saw the primary session model in ctx.model and ctx.models.current()
even when the response came from a cross-provider side request, so an extension
that revokes a credential on an HTTP 402 could target the wrong provider.
Call createContext(model) to match emitBeforeProviderRequest, plus a regression
test asserting both fields expose the response model.
Fixes#8955
The shared query pipeline parses a lang:/language: directive into StructuredQuery.lang, which sibling providers (DuckDuckGo, Perplexity, SearXNG) map onto their native locale params. The TinyFish provider dropped parsed.lang entirely, so every request fell back to the API's US/English default and non-US locales were silently lost.
Map parsed.lang onto TinyFish location (ISO 3166-1 alpha-2) and language (ISO 639-1): lang:it-it yields location=IT&language=it, lang:it yields language=it only. Behaviour is unchanged when no locale directive is present.
Fixes#8913
Remote OAuth MCP servers dropped out of /mcp under `omp auth-broker
serve` once their access token expired: neither the client nor the
broker could complete the refresh.
- Client: the MCP manager threw on the broker-redacted refresh sentinel
(REMOTE_REFRESH_SENTINEL) instead of asking the broker to refresh. It
now routes redacted MCP refreshes through
AuthStorage.forceRefreshCredentialById, which calls back to the broker
(the real refresh token never leaves the broker host).
- Broker: the serve process had no mcp_oauth:* refresh path, so
POST /v1/credential/:id/refresh answered "Unknown OAuth provider". Its
AuthStorage is now built with a refreshOAuthCredential override that
refreshes MCP credentials with a generic refresh_token grant from the
credential's embedded token endpoint and client id. The background
refresher keeps MCP tokens live through the same path.
Extract shared refreshManagedMcpOAuthCredential and
mcpOAuthServerUrlFromCredentialId helpers so both paths use identical
refresh material selection and RFC 8707 fallback-resource logic.
Fixes#8933
The `::1` companion listener added in #8081 cannot bind on hosts with
IPv6 disabled at the kernel (ipv6.disable=1). Bun reports that failure
with its generic "Is port X in use?" message (oven-sh/bun#7187), which
isAddressInUse misread as a real collision, tearing down the healthy
IPv4 listener and throwing a bogus "port 1455 is in use"
ConfigurationError that blocked Codex login.
#createServer now probes os.networkInterfaces() for an internal IPv6
loopback up front and serves IPv4 alone when none exists, instead of
relying on Bun error classification the message ambiguity defeats.
Fixes#8814
GitHub Copilot serves grok-4.6 / grok-4.6-1m only via /responses, but
isCopilotResponsesModelId matched grok-4.5 exactly, so both the static
generator and dynamic discovery classified grok-4.6 as openai-completions
and requests 400d with unsupported_api_for_model.
- match grok-4.6 in isCopilotResponsesModelId
- add grok-4.6 / grok-4.6-1m to COPILOT_CACHE_INVALIDATED_MODEL_IDS so
stale cached completion routes drop on refresh
- regenerate github-copilot/grok-4.6 to api openai-responses (compat
block dropped, xhigh effort added by the responses policy)
- cover discovery routing and cache migration in tests
Fixes#8807
The OpenCode Go gateway serves muse-spark-1.2 and
muse-spark-1.2-contributor only at /zen/go/v1/responses, but the
/zen/go/v1/models discovery omits the provider.npm hint, so the
resolver fell through to openai-completions. The completions parser
then closed the stream without a finish_reason on every tool-call turn.
Pin both ids to openai-responses in OPENCODE_GO_API_RESOLUTION, mirroring
the existing deepseek-v4-flash override, and add a resolver regression.
Fixes#8957
InputController.#invokeSkillCommand cleared the draft and awaited the full
promptCustomMessage dispatch with no transcript render. AgentSession.#promptWithMessage
runs awaited preflight (memory recall, before_agent_start hooks, auto-thinking
classification, pre-prompt compaction) before the message reaches the agent, so a slow
step such as a Hindsight auto-recall timeout left the composer cleared with no pending
row, unlike a normal prompt's optimistic row from startPendingSubmission.
Idle skill submissions now paint an optimistic skill row before the awaited dispatch;
the canonical message_start reconciles it in place via EventController instead of
appending a duplicate. Streaming submissions still queue and show their chip.
Fixes#8895
getContextBreakdown used message position (anchorIndex >= pending.cutoffCount) as a proxy for usage freshness. After a mid-run compaction rebased the in-flight snapshot, an in-flight provider response whose request predated the compaction landed past the rebase cutoff carrying pre-compaction usage, so it out-ranked the rebased estimate and reported the pre-compaction token count (~2.6x the real one). That phantom overflow tripped the "freed too little context to make progress" guard and drove the frame-rescue path on a byte-identical tokensBefore.
Assistant context snapshots now carry a monotonic compaction epoch, bumped in rebaseAfterCompaction and stamped at message-record time. A post-cutoff anchor whose epoch predates the pending snapshot's epoch is no longer trusted over the rebased estimate.
Fixes#8887
The multiline editor's key dispatch checked a hardcoded Ctrl/Shift+Enter
-> newline branch before the config-driven tui.input.submit branch, so a
user remap of submit onto Ctrl+Enter was swallowed as a newline and never
submitted. Gate the hardcoded newline fallbacks behind an explicit submit
binding; the bare-LF (iTerm2 Shift+Enter) case stays exempt because its
canonical form is indistinguishable from plain Enter.
Fixes#8906
Split-commit captured the staged diff with `git diff --cached --binary`,
whose stdout is hard-capped at GIT_COMMAND_OUTPUT_LIMIT_BYTES (8 MiB) by
readCappedText. A single large binary (base85-encoded inline) crossed the
cap; the capture was truncated silently, so files sorting after the binary
were absent from the parsed diff and stage.hunks threw a misleading
`No diff found for <path>` naming an innocent file.
Surface truncation as GitCommandResult.truncated, add a requireComplete
diff option that throws the new GitOutputTruncatedError instead of
returning a silently truncated diff, and have runSplitCommit request a
complete diff and abort with a clear message pointing at the real cause.
Fixes#8897
The shared SQLite model cache wrapped every read/write in a blanket catch that swallowed unrecoverable SQLITE_CORRUPT*/SQLITE_NOTADB failures as best-effort misses, and getSharedDb cached the broken handle. A physically corrupt models.db therefore permanently disabled cached catalogs across processes: a successful live discovery could never overwrite the corrupt cache, so a runtime extension with no bundled catalog was stuck with only its bootstrap model.
On those unrecoverable codes the cache now self-heals: close the handle, quarantine models.db(+-wal/-shm) to models.db.corrupt-<ts>, recreate a fresh database, and retry the operation once. SQLITE_BUSY, permission, and unrelated errors keep their existing best-effort paths. The SQLITE_BUSY/corruption classifiers moved to @oh-my-pi/pi-utils so the credential store and model cache share one implementation.
Fixes#8867
The OpenAI-wire transport called fetchWithRetry with maxAttempts: 6 and
the default 60s maxDelayMs cap, so a LiteLLM max_parallel_requests
rejection (HTTP 429, Retry-After: 60) was slept-and-retried up to six
times before TurnRecovery ever saw it. A 60s hint equals the cap, so
fetchWithRetry never bailed early and one turn could stall ~300s,
bypassing the user's retry.maxDelayMs/maxRetries and the session-level
CONCURRENT_LIMIT backoff + model fallback.
postOpenAIStream now opts out of transport-level retry for this
concurrency-admission response class via fetchWithRetry's shouldRetryResponse
gate, detecting the rate_limit_type=max_parallel_requests marker in the
response header or structured body. The 429 surfaces on the first attempt
so session recovery owns retry/fallback. Genuine RPM/quota 429s carry no
such marker and keep honoring Retry-After.
Fixes#8854
TUI.render merged live-region seams with a topmost-seam-wins rule that
adopted only that seam's pin policy for the whole frame. While the primary
turn streams, the transcript reports the topmost, unpinned seam, so the
frame-wide pin flag becomes false and a pinned AnchoredLiveContainer below
it (the /btw panel, the working HUD) loses its pinning: once its content
grows past the viewport, the emit path commits the scrolled-off rows as
frozen snapshots on every growth frame, piling duplicates into native
scrollback.
Track a pinnedBoundary (start row of the topmost pinned region)
independently of which seam wins the topmost merge, and cap every commit
ceiling at it. Equivalent to the prior behavior for a fully-pinned frame
and for a frame with no pinned region; only the mixed case changes.
Fixes#8793
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.
buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.
Fixes#8789
bash.patterns only feeds the bash tool's approval decision. The eval
tool declares the exec tier and can spawn a shell via subprocess, so a
deny rule there does nothing for the same command run through eval;
under yolo the exec call resolves to allow. Note the scope and point at
tools.approval.eval as the lever that closes the path in
bash-tool-runtime.md, approval-mode.md, and settings.md.
Fixes#8838