Commit Graph

25 Commits

Author SHA1 Message Date
can1357 e0509d5d62 fix(coding-agent): preserve bench catalog-first resolution 2026-07-24 16:26:59 +02:00
roboomp 8191cf41d3 fix(coding-agent): used authenticated registry models by default
Required CLI model registries to expose getAvailable() and used that authenticated
set whenever callers omit availableModels. Deferred SDK and bench/dry-balance
resolution now lets configured roles beat unauthenticated catalog id collisions.

Updated resolver test registries and made the #6508 regression omit the explicit
availableModels option, covering the deferred-caller path from the review.

Fixes #6508
2026-07-24 11:41:58 +00:00
can1357 5ded27b0b1 fix(cli): stopped $-pattern expansion when injecting cache prefix 2026-07-24 02:25:25 +02:00
Alexander Kirilin 4274100bbe fix(cli): preserve cache prefix whitespace 2026-07-23 16:52:20 -04:00
Alexander Kirilin f25cc6cf09 fix(cli): harden prompt cache benchmark diagnostics 2026-07-23 16:37:39 -04:00
Alexander Kirilin 88296f9c55 fix(cli): preserve cache benchmark affinity diagnostics 2026-07-23 15:40:07 -04:00
Alexander Kirilin 9c471a5088 feat(cli): benchmark prompt cache reuse 2026-07-23 15:18:37 -04:00
Christian Stewart 670304eafa fix(coding-agent): resolve bare model role aliases
Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-18 03:51:09 -07:00
can1357 93cc1fed1c merged PR #5417: fix(openai): render native response images
# Conflicts:
#	packages/coding-agent/src/modes/components/chat-transcript-builder.ts
#	packages/coding-agent/test/agent-session-skill-keywords.test.ts
2026-07-16 03:48:21 +02:00
roboomp b6b947bdbe fix(openai): rendered native response images
- Normalized completed image_generation_call results into assistant image blocks.
- Persisted image bytes through the session blob store and rendered them in live, replay, ACP, proxy, telemetry, and HTML paths.
- Added response normalization, persistence, and TUI rendering regressions.

Fixes #4768
2026-07-14 16:15:53 +00:00
can1357 b6559861d0 feat(cli): standardized throughput calculation to total duration
- Removed the `computeTokensPerSecond` helper function in favor of a direct calculation.
- Standardized throughput to use total request duration instead of post-TTFT decode time.
- Updated UI usage display to calculate tokens per second against total duration.
- Prevented over-inflation of throughput metrics caused by hidden reasoning tokens.
2026-07-12 01:06:24 +02:00
can1357 51684b4b1d refactor(coding-agent): streamlined codebase by deduplicating helper logic and shims
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
2026-07-02 02:40:08 +02:00
roboomp a4ae4c130c fix(coding-agent): preserved explicit :auto suffix in modelRoles
The model selector's persistence path dropped the `:auto` selector when parsing role values, producing a warning ('Invalid thinking level "auto"') and rendering the badge as `inherit` instead of `auto`. Reload of the default role also lost the auto state whenever the role value carried an explicit `:auto` suffix instead of relying on `defaultThinkingLevel`.

Widen the resolver chain (`parseThinkingSuffix`, `splitThinkingSuffix`, `parseModelString`, `parseModelPattern*`, `ResolvedModelRoleValue`, `ResolvedRoleModel`, `ResolveCliModelResult`) to carry the `AUTO_THINKING` sentinel end to end, and coerce it back to `undefined` at concrete-only boundaries (glob scope patterns, retry fallback, advisor, commit pipeline, guided-goal, bench).

Regression tests cover:

- `resolveModelRoleValue("provider/model:auto")` returns explicit auto without a warning.

- `ModelSelector` renders `DEFAULT (auto)` and `SMOL (auto)` when the role value has `:auto`.

- `cycleRoleModels` activates auto thinking on entering a `:auto` role.

- Startup resume activates auto thinking when `modelRoles.default` carries `:auto`.

Fixes #4128
2026-07-01 08:18:42 +00:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 d20e6c0829 feat: migrated service tier settings to a per-model-family architecture
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
2026-06-30 04:14:48 +02:00
can1357 1eb32a5df1 feat(coding-agent/bench-cli): tracked and close provider session states
- Added management of provider session states during benchmark execution.
- Implemented a teardown process to close and clear session states after request completion.
2026-06-29 16:31:21 +02:00
can1357 312b71b0a5 feat(sdk): enabled websocket transport configuration for sdk and cli
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
2026-06-29 08:05:02 +02:00
can1357 c8c24b0888 feat(bench): enabled concurrent execution and service tier selection
- Added --par flag to execute benchmark runs concurrently with a default degree of 4.
- Added --service-tier flag to allow overriding the provider service tier per benchmark.
- Increased default benchmark run count from 1 to 10 to provide more robust averaging.
- Updated benchmarking logic to process requests in a concurrency-limited pool while preserving output order.
- Implemented pre-flight credential checks to prevent unnecessary worker spawning when authentication is missing.
2026-06-29 06:51:13 +02:00
can1357 d10a5356ca feat(coding-agent): added provider extension support to CLI tools
- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
2026-06-19 19:23:46 +02:00
can1357 67f6518e42 feat: enhanced tool robustness, improve authentication flow, and update API parameters
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
2026-06-19 16:46:07 +02:00
can1357 48decd15d7 fix(coding-agent): fixed context usage tracking to keep status and selector totals in sync
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
2026-06-17 12:24:20 +02:00
can1357 3f32971944 feat(ai): align OpenRouter max_tokens with caller intent
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
2026-06-17 12:24:19 +02:00
can1357 f0c6a54f51 fix: handled unknown model limits as null to avoid artificial token caps
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
2026-06-13 15:35:40 +02:00
can1357 ce112bf9ff fix(coding-agent): fixed omp bench cached OpenRouter replays and Codex instruction 400s
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
2026-06-12 08:21:25 +02:00
can1357 300c1ada30 feat(coding-agent): added bench command and updated default compaction shapes
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.
2026-06-12 08:00:03 +02:00