Commit Graph
10 Commits
Author SHA1 Message Date
can1357 1eb32a5df1 feat(coding-agent/bench-cli): tracked and close provider session states
- Added management of provider session states during benchmark execution.
- Implemented a teardown process to close and clear session states after request completion.
2026-06-29 16:31:21 +02:00
can1357 312b71b0a5 feat(sdk): enabled websocket transport configuration for sdk and cli
- Added `preferWebsockets` option to `AgentSessionConfig` to expose transport preferences.
- Updated `AgentSession` to manage and forward websocket preferences to sub-sessions.
- Enabled websocket transport by default for benchmark CLI requests.
2026-06-29 08:05:02 +02:00
can1357 c8c24b0888 feat(bench): enabled concurrent execution and service tier selection
- Added --par flag to execute benchmark runs concurrently with a default degree of 4.
- Added --service-tier flag to allow overriding the provider service tier per benchmark.
- Increased default benchmark run count from 1 to 10 to provide more robust averaging.
- Updated benchmarking logic to process requests in a concurrency-limited pool while preserving output order.
- Implemented pre-flight credential checks to prevent unnecessary worker spawning when authentication is missing.
2026-06-29 06:51:13 +02:00
can1357 d10a5356ca feat(coding-agent): added provider extension support to CLI tools
- Enable shared extension provider loading in bench and dry-balance CLI commands to ensure custom providers are registered.
- Surface benchmark failures for empty streams that return no content and zero usage tokens instead of treating them as successful.
2026-06-19 19:23:46 +02:00
can1357 67f6518e42 feat: enhanced tool robustness, improve authentication flow, and update API parameters
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
2026-06-19 16:46:07 +02:00
can1357 48decd15d7 fix(coding-agent): fixed context usage tracking to keep status and selector totals in sync
- Added context snapshot metadata to AssistantMessage for prompt and non-message token history.
- Anchored context usage calculations on assistant snapshots and computed percent numerically.
- Updated status-line, /context, selector, and interactive mode flows to share session usage totals.
- Extended status-line cache fingerprinting and invalidation for assistant usage and prompt/tool/skill changes.
2026-06-17 12:24:20 +02:00
can1357 3f32971944 feat(ai): align OpenRouter max_tokens with caller intent
OpenRouter previously omitted `max_tokens` entirely (except for specific models) to prevent unintended provider filtering when a model's catalog default `maxTokens` exceeded an upstream's actual capacity. This could lead to incorrect routing if a model had a high catalog cap but individual providers under OpenRouter did not.
2026-06-17 12:24:19 +02:00
can1357 f0c6a54f51 fix: handled unknown model limits as null to avoid artificial token caps
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
2026-06-13 15:35:40 +02:00
can1357 ce112bf9ff fix(coding-agent): fixed omp bench cached OpenRouter replays and Codex instruction 400s
- Sent `X-OpenRouter-Cache: false` on bench requests in `bench-cli.ts`: pi-ai opts every OpenRouter request into 1h response caching, so repeated byte-identical runs replayed a cached generation with zeroed usage as "tokens 0, TPS 0.0" successes.
- Added a minimal default `systemPrompt` to bench's request context, matching eval's completion-bridge guard against Codex's HTTP 400 `{"detail":"Instructions are required"}`.
- Checked in `src/prompts/bench.md`, the default bench prompt `bench-cli.ts` already imports (left untracked by 300c1ada30).
- Added both bench changelog entries.
2026-06-12 08:21:25 +02:00
can1357 300c1ada30 feat(coding-agent): added bench command and updated default compaction shapes
- Added a new `bench` CLI command with multi-model selectors and new options.
- Implemented `runBenchCommand` validation, per-run session handling, and failure exit reporting.
- Updated default compaction shapes to `8x8r-bw` and `doc-8on16-sent-dim` in code and schema.
- Documented `bench` flags, per-run errors, failure counts, and exit behavior.
2026-06-12 08:00:03 +02:00