- Added `rewrite-changelog.ts` and `fix-changelogs.ts` utilities to automate the consolidation of release notes using LLM-assisted processing.
- Updated multiple internal changelog files by consolidating redundant entries and improving phrasing for readability.
- Implemented `previewLine` utility in `coding-agent` to prevent visual spillover in status rows by managing text truncation and whitespace.
- Updated `package.json` with new workflow scripts for managing package-level change histories and documentation indexes.
Direct Anthropic Claude Sonnet 4.5 and Haiku 4.5 (plus their Cloudflare,
Vertex, GitLab-Duo, Copilot, OpenCode-Zen, and Bedrock cross-region passes)
were classified as anthropic-budget-effort, which made the Anthropic provider
serialize output_config.effort alongside the thinking.budget_tokens block.
Anthropic only honors output_config.effort on Opus 4.5 and adaptive (4.6+)
Messages-API models — Sonnet 4.5 and Haiku 4.5 reject every request with
HTTP 400 'This model does not support the effort parameter.', so the advisor
(and any agent on a Sonnet/Haiku 4.5 SKU) failed every turn.
inferThinkingControlMode now gates anthropic-budget-effort to
parsedModel.kind === 'opus' && semverGte(version, '4.5') on both
anthropic-messages and bedrock-converse-stream. Sonnet/Haiku 4.5 fall through
to mode: 'budget' (effort still scales the per-tier thinking budget via
ANTHROPIC_THINKING[reasoning]); Opus 4.5 keeps anthropic-budget-effort and
continues to emit output_config.effort. anthropic-budget-effort also remains
in use for Anthropic-compatible third-party backends that natively support
the field (Umans GLM 5.2).
Regenerated models.json so the 19 Sonnet/Haiku 4.5 first-party entries flip
to mode: 'budget' and the 12 Opus 4.5 entries stay on anthropic-budget-effort.
Regression tests cover both sides: anthropic-alignment.test.ts asserts the
Sonnet 4.5 wire body omits output_config and the Opus 4.5 wire body emits
output_config.effort: 'medium'.
Fixes#3497
Ollama Cloud rejects any chat request whose options.num_predict exceeds
65536 with HTTP 400, but the existing safety relied on the load-time
omitMaxOutputTokens policy in model-registry.ts. Stale models.db rows
predating that policy (and custom modelOverrides re-enabling output
caps) carried maxTokens: 1048576 forward to the wire layer untouched,
so every request to deepseek-v4-pro / deepseek-v4-flash 400'd with:
max_tokens (1048576) exceeds model's maximum output tokens (65536)
for model deepseek-v4-pro
createChatBody now resolves num_predict through resolveNumPredict,
which clamps every ollama-cloud request at the documented cap before
serialization — independent of the cached spec, on top of the existing
omitMaxOutputTokens path. Self-hosted ollama traffic is unaffected.
Fixes#3392
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.
- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
forces input=['text'] whenever the resolved baseUrl is not the
canonical personal-Copilot host, so the upstream's vision flag is
honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
OR-upgrading with the bundled reference) when the merged baseUrl
differs from the bundled one, so a bundled spec pinned to the
personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
rasterizer for any github-copilot model whose baseUrl is non-personal,
catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
pi-catalog/wire/github-copilot so catalog and coding-agent share one
canonical check.
Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).
Fixes#3387
- Update `sakanaModelManagerOptions` to provide `dropCachedModelIdsOnStaticMismatch` configuration using identified static model IDs.
- Add a test case to verify that stale cached model metadata is correctly cleared when bundled model specs are updated.
- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
- Implemented Sakana AI and Fugu provider integration including authentication, API base URL resolution, and dynamic model discovery.
- Configured static model definitions and reasoning metadata for the Fugu model series within the catalog.
- Added environment variable support for API configuration and base URL overrides via `SAKANA_*` and `FUGU_*` variables.
- Verified service integration and provider registry registration through comprehensive test suites in both AI and catalog packages.
Updated the Umans provider regression test to import Effort from the catalog effort submodule instead of the root catalog barrel, keeping the package-local test on the narrow dependency path.
Refs #3192
Forward anthropic-budget-effort selections through mapOptionsForApi so buildParams serializes output_config.effort with budget-token thinking. Mark Umans GLM-5.2 as budget-effort and map the UI xhigh tier back to Umans's max wire value.
Refs #3192
Removed the xhigh -> max effortMap baked onto the Umans GLM-5.2 spec; mode="budget" routes thinking depth via thinking.budget_tokens and never consults effortMap, so the wire shape is unchanged. The picker still surfaces high and xhigh via getModelDefinedEfforts, and the dynamic discovery still maps the upstream max level to Effort.XHigh.
Refs #3192
Mapped Umans GLM-5.2's upstream max reasoning level to the internal xhigh effort and preserved the max wire value in dynamic discovery and bundled catalog metadata.
Added resolver coverage for the high/max ladder and verified the xhigh request maps back to max.
Fixes#3192
Upgraded installs can carry an Umans model cache written before the
GLM via-handoff catalog correction. If dynamic discovery is skipped or
fails, resolveProviderModels merges cache rows over the corrected static
catalog; mergeDynamicModel preserves image support when either side has
it, so a stale cached ["text", "image"] GLM row can re-add native image
support until the next successful refresh.
Add a provider-scoped cache drop hook for model ids whose cached rows
are unsafe across static fingerprint changes, opt Umans into it for
`umans-glm-5.1` and `umans-glm-5.2`, and cover the offline stale-cache
upgrade path with a regression test.
Fixes#3184
`umans-glm-5.1` / `umans-glm-5.2` advertise themselves on the Umans
`models/info` endpoint with `supports_vision: "via-handoff"`. That
sentinel means image inputs are routed through a separate vision
handoff pre-analysis step; the GLM endpoint itself rejects raw image
blocks with `400 This model does not support image inputs`.
`umansSupportsVision` was returning `true` for any non-empty string,
so dynamic discovery mapped the GLM models to `input: ["text",
"image"]` and the agent sent images straight to GLM. The bundled
`umans-glm-5.1` / `umans-glm-5.2` rows in `models.json` carried the
same stale `["text","image"]` from a previous regen.
- Tighten `umansSupportsVision` to `value === true`; document the
sentinel contract.
- Correct the two bundled rows to `input: ["text"]` so the vision
handoff path runs.
- Add resolver- and bundle-level regression tests that cover
`supports_vision: "via-handoff"` alongside the native-vision
`umans-coder` case.
Fixes#3184
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
- Kept provider discovery defaults authoritative for api and provider when layering bundled or models.dev metadata.
- Added a LiteLLM regression covering a deepseek-v4-flash collision with the ollama-cloud catalog.
Fixes#3162
- Update `llama.cpp` base URL to remove the `/v1` suffix.
- Consolidate local provider token validation in `coding-agent` using a `Set`.
- Add test coverage to ensure catalog model IDs are preserved verbatim on the wire.
- Updated `buildOpenAICompat` to override the `qwen` thinking format for Fireworks-hosted models, ensuring they use `openai` thinking parameters instead.
- Prevented invalid `enable_thinking` payload errors by ensuring Fireworks-hosted Qwen requests conform to their strict schema.
- Updated `AgentSession` to allow Fireworks fast-fallback logic to execute even when standard retries are disabled.
The MiniMax-M3 long-context policy in generated-policies.ts only
covered the anthropic-messages providers `minimax` and `minimax-cn`.
The MiniMax Coding/Token Plan (international and China) endpoints
serve the same model through `minimax-code` and `minimax-code-cn`
on openai-completions, and shipped with the upstream 512K pricing
boundary baked into models.json. Switching to MiniMax-M3 under the
Coding Plan therefore still showed a 512K context window in the
status bar.
Broadens the policy carve-out to all four providers, re-bakes both
affected entries in the bundled models.json, and extends the
generated-policies / bundled-catalog tests to assert 1M for the two
newly covered providers.
Fixes#3097
Follow-up to #3071 review (codex + @PGupta-Git): the family rewrite alone
did not heal users running off the bundled catalog or stale SQLite cache
rows, which still routed Sonnet 4.6 thinking efforts to the 404
`claude-sonnet-4-6-thinking` wire id. `collapseEffortVariants` treats
those collapsed snapshots as authoritative and `refreshCollapsedThinking`
exits early for families without `effortBudgets` (Claude pairs), so the
new empty routing in the hand table never reached the snapshot.
- Declared the dead wire ids as `retiredMembers` on the Claude 4.6
families (`claude-sonnet-4-6-thinking` on the Sonnet family,
`claude-opus-4-6` on the Opus family). This triggers
`reconcileRetiredRouting` to rewrite every `effortRouting` entry that
targets a retired id to a live wire id (Sonnet falls back to the bare
member; Opus falls back to `-thinking`).
- Refreshed the bundled `packages/catalog/src/models.json` Sonnet 4.6
entry so fresh installs do not boot with the dangling routing — the
surgical diff matches what the generator would emit; the rest of the
catalog is left untouched to keep the bug-fix PR scoped.
- Added regression tests in `variant-collapse.test.ts` for both
reconciliation paths and a bundled-catalog test
(`issue-3067-repro.test.ts`) that pins the live-wire-id resolution end
to end through `buildModel` for every effort tier.
Fixes#3067
The daily Cloud Code Assist backend (`daily-cloudcode-pa`) exposes Claude 4.6
asymmetrically: `claude-sonnet-4-6` has no `-thinking` twin and
`claude-opus-4-6` has only the `-thinking` twin. The shared
`thinkingPair("claude-sonnet-4-6", …)` family (with `preserveAbsentEffortRoutes`)
kept every effort routed to `claude-sonnet-4-6-thinking` even when discovery
only returned the bare id, so any reasoning-on request 404'd with
`Requested entity was not found`. Claude on Antigravity also caps
`maxOutputTokens` at 64000, while `ANTIGRAVITY_MODEL_WIRE_PROFILES` had no
Claude entries — discovery's 65536 propagated to the wire and 400'd with
`Request contains an invalid argument`.
- Replaced the two `thinkingPair` calls for Claude 4.6 in `SHARED_CCA_FAMILIES`
with bespoke single-wire families. Sonnet collapses to the bare wire id,
Opus collapses to the `-thinking` wire id, and per-effort thinking is
carried by the request body's `thinkingBudget` on the single shared wire id.
Listing both candidate ids in `members` (priority order) keeps the collapse
correct if the backend mix ever rebalances.
- Added `claude-sonnet-4-6` and `claude-opus-4-6-thinking` entries to
`ANTIGRAVITY_MODEL_WIRE_PROFILES` capping `maxOutputTokens` at 64000.
- Made `AntigravityModelWireProfile.modelEnum` optional — Anthropic-backed
wire ids are accepted without a captured `labels.model_enum` token. The
request builder now emits the label only when the profile defines one.
- Regression tests in `variant-collapse.test.ts` (routing/wire-id resolution
for both 4.6 families across all three discovery permutations) and
`google-gemini-cli-alignment.test.ts` (request builder caps Claude
`maxOutputTokens` at 64000 and omits the unset `model_enum` label).
Fixes#3067
- Refactor GLM-5.2 effort mapping to accommodate specific requirements for Z.ai, OpenRouter, and general OpenAI-compatible hosts.
- Apply host-specific logic to ensure `xhigh` UI tiers are correctly resolved to the required `max` budget for supported providers.
- Add test coverage verifying expected effort mappings across different model hosts.
- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.
GitHub Copilot's `anthropic-messages` proxy (api.githubcopilot.com) forwards
to signature-enforcing Anthropic and returns full thinking signatures, but the
compat builder classified it as a non-signing reasoning endpoint via the
generic `reasoning && !official` default (`replayUnsignedThinking: true`).
When a checkpoint/branch-return turn is an abandoned tool-use turn (adaptive
Opus emits a tool call then ends on `stop`/`end_turn`), `transformMessages`
correctly strips its end_turn-bound, unreplayable signature. On a
`replayUnsignedThinking` endpoint the encoder then re-emitted that block as
`{ type: "thinking", signature: "" }`. An empty signature is rejected by the
signature-enforcing backend with `400 Invalid signature`, which corrupts the
session and re-trips on every full history re-send (e.g. after toggling MCP
servers).
Exclude github-copilot from `replayUnsignedThinking` so unsigned/stripped
thinking degrades to text exactly like the official Anthropic API — wire-valid
and lossless of the tool_use pairing. Z.AI / DeepSeek / other 3p reasoning
endpoints (#2005) and cross-model preservation (#2257/#2265) are unaffected.
Tests:
- packages/catalog/test/anthropic-copilot-signing-compat.test.ts: copilot
(incl. enterprise copilot-api.* hosts) -> replayUnsignedThinking false;
generic 3p reasoning -> true; official -> false. Fails before / passes after.
- packages/ai/test/anthropic-copilot-checkpoint-thinking-signature.test.ts:
a signing copilot model never emits an empty-signature thinking block for a
historical checkpoint turn (demotes to text, keeps tool_use), and still
replays a clean signed historical thinking block natively.