- The mux pane-growth oracle treated every physical scroll as a logical append, but immutable-history recovery can recommit a corrected suffix after an off-screen mutation without advancing the shadow tape; exempt only changed shared history prefixes.
- The frame-neutral oracle compared prepared rows only; at narrow widths distinct raw rows prepare identically, so the renderer's raw-prefix divergence recovery recommits legitimately. Snapshot raw frames and allow declared transient growth.
- OSC66 spacer preservation intentionally composes six bounded context rows above the resize viewport; assert that exact bound instead of zero above-fold rendering.
Both oracle false positives reproduce identically at the PR head that introduced the harness (4cc9725037); three full randomized stress passes green after the fix.
- Set generation temperature to zero for online title generation to prevent garbled names.
- Update system prompt to instruct exact copying of technical terms and names.
- Reject generated titles containing no word characters to prevent punctuation-only sessions.
Ollama Cloud serves the DeepSeek V4 family over the ollama-chat API,
which bypassed the deepseek effort-ladder branch in model-thinking.ts.
deepseek-v4-flash exposed the generic minimal/low/medium/high/xhigh
ladder with no max tier instead of the low/high/max contract the
catalog encodes on every other host.
Broadened the branch to also cover the ollama-cloud ollama-chat surface:
Flash keeps low/high/max, V4 Pro and older reasoners top out at
high/max. Regenerated models.json accordingly.
Fixes#8334
Built-in OpenAI-compatible provider managers call fetchOpenAICompatibleModels with neither a signal nor a timeoutMs, and the no-timeout branch issued the request with signal: undefined. A stalled /models endpoint left the fetch pending forever, so createAgentSession's awaited resolveModelDiscoveryFallback discovery pass never returned and startup hung.
Apply a default 10s deadline (DEFAULT_OPENAI_COMPATIBLE_DISCOVERY_TIMEOUT_MS) when the caller supplies neither signal nor timeoutMs, matching the coding-agent remote-discovery budget. Callers passing their own signal keep owning its lifecycle.
Fixes#8315
- Bound PAX sparse record memory overhead by caching sparse markers and specific keys.
- Update system prompt phrasing and tests for tool inventory and date displays.
- Added Google provider thinking configuration parameters and force-reasoning-off controls.
- Implemented MCP SSE stream resumption using Last-Event-ID and `SSEResumeError`.
- Added support for TAR old-GNU sparse extension blocks, path length checks, and archive entry overrides.
- Restricted external thinking support to specific models and added semver fallback parsing.
- Define a centralized `USER_AGENT` constant in `@oh-my-pi/pi-utils` formatted as `omp/<version>`.
- Replace hardcoded and platform-specific user agent strings across AI providers, catalog scrapers, tools, and search providers with the unified `USER_AGENT`.
- Add unit tests for update-cli binary release distribution gating.
Adds zai-org/GLM-5.2-Fast to the Baseten reasoning allowlists, matching
the sibling zai-org/GLM-5.2. Also fixes parseGlmModel to handle uppercase
GLM model IDs (used by Baseten, CoreWeave, HuggingFace, etc.), so the
identity deriver correctly classifies them as GLM-5.2 reasoning models
during catalog generation — previously the case-sensitive regex caused
rebakeModelThinking to fall back to the generic effort ladder.
Regenerated models.json with a live Baseten API key: GLM-5.2 and
GLM-5.2-Fast now bundle reasoning:true with the correct high/max effort
ladder, and other uppercase GLM-5.2 resellers (CoreWeave, HuggingFace,
Synthetic, Together, Wafer) get the corrected minimal..max ladder.
- Update packages/catalog/test/meta-provider.test.ts to expect
three META_MUSE_STATIC_MODELS entries (1.1, 1.2, contributor)
with input ["text","image"] — fixes blocking failure.
- Add ## [Unreleased] entry in packages/catalog/CHANGELOG.md per
AGENTS.md.
- Remove trailing newline from models.json to match
generate-models.ts (Bun.write without \n).
Co-authored-by: roboomp <roboomp@users.noreply.github.com>
Meta Muse Spark 1.2 and its contributor variant support vision
inputs (text, image) like 1.1, but were missing from
META_MUSE_STATIC_MODELS and bundled models.json. Discovery
without a bundled reference falls back to text-only with zero
cost, so omp models listed them as not image enabled.
Add both models to META_MUSE_STATIC_MODELS with the same
cost/thinking/vision metadata as 1.1 (contributor uses its
discounted pricing) and regenerate the bundled catalog.
- The regenerated bundle (merged with PR #8021) now ships a
responses-route github-copilot/grok-4.5, so the id legitimately
resurfaces from the bundle when the migration refresh fails; the
contract worth defending is that the stale cached completions route
never returns and the unbundled long-context variant stays dropped.
- The OpenCode Go gateway does not serve DSV4-Flash at
/zen/go/v1/chat/completions; /zen/go/v1/responses works (user-verified
against the live gateway). Added a per-id override in
OPENCODE_GO_API_RESOLUTION so both bundled generation and the runtime
/v1/models refresh route it to openai-responses; deepseek-v4-pro keeps
chat completions.
- Regenerated models.json from the resolver source.
Applied curated Alibaba Token Plan seeds after generic models.dev fallback so bundled capabilities cannot be overwritten by incomplete upstream metadata.
- Added one internal generic builder covering the repeated apiKey/baseUrl
resolution, bundled reference map, providerId and fetchDynamicModels closure.
- Migrated openai, cerebras, novita, aimlApi, alibabaCodingPlan, venice,
baseten and moonshot, and routed createSimpleOpenAICompletionsOptions
through it; deleted the duplicate responses-side helper.
- Preserved fetchDynamicModels key PRESENCE per site, since the apiKey spread
guard omits the key entirely rather than setting it undefined.
- Replace the compat cast with an annotated CompatOf<Api> local narrowed
via the in operator: assignment up-cast, compiler-checked field type,
no as assertion.
- Cover the 0 sentinel end-to-end through the lazy wrapper with fake
timers (mirrors the direct-Anthropic 0-disable test): advance 400s
past the generic budget, assert no watchdog abort, then cancel
cleanly.
- Add Unreleased changelog entries for pi-ai and pi-catalog with
external attribution for #7892.
The lazy provider wrapper ignored model.compat.streamIdleTimeoutMs, so
Bedrock reasoning models sat on the generic 300s idle watchdog despite
ConverseStream sending no ping keepalives; long quiet thinking runs died
with "Provider stream stalled while waiting for the next event" during
plan writing and todo execution (issue #4758's Bedrock variant, worst on
Fable 5 where the display default flipped to omitted).
- catalog: BedrockCompat gains streamIdleTimeoutMs; reasoning models get
a 600s floor, adaptive-thinking Claude (Opus 4.7+, Sonnet/Opus 5,
Fable/Mythos 5) 900s to match direct Anthropic's ping-extended
tolerance; explicit compat overrides still win (0 disables).
- ai: forwardStream resolves options -> env -> model.compat -> default,
and lazy terminal errors carry the structural errorId classification
so session auto-retry classifies stalls without text matching.
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
DeepSeek's API accepts reasoning_effort low/high/max and only
deepseek-v4-flash supports all three tiers (V4 Pro is high/max). The
identity-derived effort fallback blanket-applied high/max to every
direct-DeepSeek reasoning model, hiding the low tier flash accepts.
Added isDeepseekV4FlashModelId and route flash to the low/high/max ladder
on every host; non-flash DeepSeek keeps high/max (high-only on OpenRouter).
Fixes#7668
Bare anthropic.claude-* Bedrock rows already derived eu.* selectors; also
emit us-gov.* so GovCloud accounts can resolve system inference profiles
without requiring a full partition ARN.