Commit Graph
84 Commits
Author SHA1 Message Date
can1357 4e6f04220f chore: reformat 2026-05-31 13:04:25 +02:00
can1357 6cee7a666d test(tests): added deterministic yield tests and viewport mutation regression coverage
- Replaced real-time yield assertions with a mocked clock and scheduler wait spy to validate gate timing deterministically.
- Updated fact consolidation conflict tests to capture console warnings and verify duplicate-resolution messaging.
- Added TUI offscreen-expansion regressions for unknown-viewports, covering deferred rebuild and user-driven rebuild behavior.
2026-05-31 13:02:32 +02:00
can1357 417a1a1d32 feat(agent): added shake compaction strategy primitives
Introduce the shake compaction strategy in the core agent: exports, summary prompt, implementation, and unit coverage.
2026-05-31 07:39:38 +02:00
can1357 ff9a6826dd fix(compaction): made read tool-results prunable except for skill:// paths
- Replaced flat `protectedTools: string[]` with `ProtectedToolMatcher[]` supporting predicate functions.
- Regular file/URL `read` calls are now eligible for pruning and shake compaction.
- `read` calls whose `path` starts with `skill://` remain protected like native `skill` results.
- Added `collectToolCallsById` to correlate tool results with their originating call arguments.
2026-05-31 07:12:27 +02:00
can1357 5db3bcabad fix(agent): snapshot initial mutable state
Addresses review feedback on #1507.
2026-05-31 04:49:17 +02:00
oldschoolaandcan1357 3a733c480b perf(agent,ai,coding-agent): in-place state mutation + per-delta json-parse throttle + drop hot-path structuredClone
F3 (agent): mutate state.messages and state.pendingToolCalls in place on
appendMessage/popMessage/clearMessages/reset/tool_execution_start/_end
instead of allocating a fresh array/Set on every transition. Subscribers
that capture state.messages by reference now observe updates directly.
Public type signature unchanged.

F5 (ai): add parseStreamingJsonThrottled to utils/json-parse — a per-delta
wrapper around parseStreamingJson that skips the re-parse until the
buffer has grown by minGrowthBytes (default 256). Wired into every
provider's tool-call argument accumulator (anthropic, amazon-bedrock,
openai-completions, openai-codex-responses, openai-responses-shared) so
per-delta cost becomes O(N) in total buffer length instead of O(N²).
Every provider's toolcall_end still runs a final unthrottled parse, so
the published block.arguments is unchanged.

F8 (coding-agent): drop the per-delta structuredClone of streaming tool
arguments in ToolExecutionComponent.updateArgs. event-controller.ts and
ui-helpers.ts already spread their input into a fresh object on each
delta, so cloning here was dead work on the rendering hot path. Added a
reference-equality short-circuit so repeat calls with the same args
object skip the preview-diff and display refresh.
2026-05-31 04:46:06 +02:00
can1357 e831c2c758 chore: reformat 2026-05-30 18:08:51 +02:00
can1357 ae905fb3cf feat(agent): added agent tool-call cap enforcement to stream loop
- Added `maxToolCallsPerTurn` support to `AgentOptions` and `AgentLoopConfig`, with Agent getter/setter and serialized state wiring.
- Implemented stream-loop cap handling by normalizing bad values and halting after `toolcall_end` reaches the limit.
- Added `ANTHROPIC_TOOL_CALL_BATCH_CAP`=8 and wired session cap sync on init, model changes, and restore.
- Added tests that truncated a 10-call stream to 8 tool calls, and verified non-Claude models resolve no cap.
2026-05-30 04:47:58 +02:00
can1357 c4f93eca20 fix(agent): patched compaction 401/403 fallback to copy error status
- Centralized compaction stop-reason error throws via createSummarizationError().
- Set compaction thrown errors to copy response.errorStatus into Error.status.
- Expanded compaction auth detection to treat HTTP 401/403 as auth failures with regex fallback preserved.
- Added regression tests for 401/403 status propagation and compaction fallback auth behavior.
- Documented both package fixes in Unreleased Fixed changelog entries.
2026-05-28 12:50:52 +02:00
cognitiveandmetaphorics 0455c164f9 fix(agent): compact() propagates thinkingLevel to all three summarizers
The thinking-level fix (e07b47ee4) added SummaryOptions.thinkingLevel
and threaded it from agent-session.ts into compact(), but the
field-by-field rebuild of summaryOptions inside compact() (and a
second inline rebuild for generateShortSummary) silently dropped it.

Effect on every call site that fans through compact():
  generateSummary, generateTurnPrefixSummary, generateShortSummary all
  see options?.thinkingLevel === undefined => resolveCompactionEffort
  falls back to Effort.High => user's /model :off selection is
  silently overridden, and xai-oauth/grok-build still trips on the
  unsupported-effort path even though fix #2 strips it at the wire
  layer of the openai-responses mapper.

Add `thinkingLevel: options?.thinkingLevel` at both rebuild sites in
compact(): the summaryOptions literal feeding generateSummary and
generateTurnPrefixSummary, and the inline options literal feeding
generateShortSummary.

Extend compaction-thinking-level.test.ts with four compact()-level
cases driving isSplitTurn:true so all three summarizers fire:
  - Off                  -> every fan-out call gets reasoning=undefined
  - Low                  -> every fan-out call gets reasoning="low"
  - <unset>              -> every fan-out call gets reasoning="high" (default)
  - grok-build + High    -> every fan-out call gets reasoning=undefined (clamp)

TDD red-green verified: stashing the source fix flips the Off and Low
cases to fail with received="high" (exactly the reviewer's prediction);
restoring the fix returns all four to green. Suite: 131 pass / 0 fail
(baseline 127 + 4 new). biome + tsgo --noEmit clean.

Op: correct
Restores: spec:compaction-honors-session-thinking-level
(cherry picked from commit 9b501e369b820cb992345ceb3f2a207ccc9ee233)
2026-05-27 15:01:24 +00:00
cognitiveandmetaphorics 25c6794cd5 fix(agent): compaction honors session thinking level and silent-clamps unsupported-effort models
Triple-stacked failure on the same axis (thinking effort) produced the
user-visible

    Error: Compaction failed: Thinking effort high is not supported by
           xai-oauth/grok-build.
    Supported efforts:

(empty list after the colon) whenever the active model was a curated
xAI catalog entry with compat.supportsReasoningEffort: false.

Three defects lined up. (1) Behavior: compaction at four call sites
in packages/agent/src/compaction/compaction.ts hardcoded
reasoning: Effort.High and never threaded session.thinkingLevel —
the user's /model :off selection (and any explicit low/medium) was
silently overridden. On every other model this was invisible.
(2) Validation: requireSupportedEffort threw at the openai-flavored
mapper layer before the wire-side omitReasoningEffort gate in
providers/xai-responses.ts ever ran; two contradictory guards on the
same wire param. (3) Message: when getSupportedEfforts returned [],
the rendered error tail was 'Supported efforts: ' with nothing after
the colon — disappears as a side-effect of fix #2.

Fix #1 — thread ThinkingLevel | undefined end-to-end. Add
SummaryOptions.thinkingLevel and HandoffOptions.thinkingLevel.
Convert via a single exhaustive switch (effortFromThinkingLevel) in
the new resolveCompactionEffort helper:
  - Off            → undefined  (omit reasoning entirely)
  - undefined/Inherit → Effort.High → clamp per model (preserves the
                                       historical default for users
                                       who never touched the dial)
  - explicit Effort → respect user → clamp per model

resolveCompactionEffort lives in compaction.ts; all four call sites
(generateSummary, generateHandoff, generateShortSummary,
generateTurnPrefixSummary) route through it. agent-session.ts threads
this.thinkingLevel into all three production compaction entry points
(manual /compact at L6201, auto-compaction at L6458 — the most-fired
path, originally missed in plan review — and direct generateHandoff
at L5465). The audit-gate test
(test/agent-session-compaction-thinking-threading.test.ts) scans the
file with a brace-balanced extractor and refuses any unthreaded site.

Fix #2 — silent-clamp at the openai-flavored mapper layer. Extract
exported modelOmitsReasoningEffort(model) in model-thinking.ts as the
single source of truth for compat.supportsReasoningEffort: false on
openai-responses* APIs. getSupportedEfforts now calls it instead of
inlining the check (pure refactor — observable behavior preserved).
resolveOpenAiReasoningEffort in stream.ts early-returns undefined
when the predicate is true, so the wire-side omitReasoningEffort
gate (providers/xai-responses.ts:78) becomes the single source of
truth for the actual strip — no redundant throw.

Three regression tests pin the contract:
  - packages/ai/test/xai-oauth-effort-strip.test.ts (5 tests):
    modelOmitsReasoningEffort returns true for grok-build and
    grok-4.20-0309-reasoning, false for grok-4.3 / Anthropic /
    openai-completions.
  - packages/agent/test/compaction-thinking-level.test.ts (5 tests):
    every ThinkingLevel outcome through generateHandoff — Off stays
    undefined (not coerced to High), Low stays Low, Inherit / undefined
    default to High, grok-build clamps to undefined regardless of
    requested level. Covers the Codex-caught Off-vs-not-provided
    distinction.
  - packages/coding-agent/test/agent-session-compaction-thinking-threading.test.ts
    (2 tests): brace-balanced source scan asserts every direct
    compact() / generateHandoff() in agent-session.ts threads
    'thinkingLevel: this.thinkingLevel'; floor of 3 threaded sites.

TDD red-green verified for fix #1: temporarily reverted the handoff
call-site back to hardcoded Effort.High → compaction-thinking-level
went 2 pass / 3 fail (Off coerced, Low overridden, grok-build throws);
restored → 5 pass / 0 fail.

Verified:
  - packages/agent:  127 pass / 0 fail
  - packages/ai:     1061 pass / 337 skip / 0 fail
  - packages/coding-agent (focused): 179 pass / 5 skip / 0 fail
  - biome + tsgo --noEmit clean across all three packages

Out of scope (follow-ups):
  - branch-summarization.ts:307 already passes no reasoning — no edit.
  - The empty-list error message at model-thinking.ts:296 is now
    structurally unreachable from the openai-responses path.
  - modelOmitsReasoningEffort and grokSupportsReasoningEffort
    (xai-responses.ts:22) overlap; collapse into a single predicate
    in a future commit.

Op: correct
Restores: spec:compaction-honors-session-thinking-level
Restores: spec:xai-oauth-grok-build-compaction-no-throw
(cherry picked from commit e07b47ee46769053c658819437e2478389a4cee0)
2026-05-27 15:01:20 +00:00
can1357 1d8ee5a891 refactor(yield): migrated to scheduler.wait with abort support and gate
- Replaced Bun.sleep with scheduler.wait for Node-compatible cancellable sleeps.
- Added module-level timestamp gate to skip yields within 50ms of the last one.
- Threaded AbortSignal through ExponentialYield.sleep to cancel losing timers in race.
- Added tests covering gate behaviour and stray-timer cancellation.
2026-05-26 20:37:39 +02:00
can1357 80186e341c feat(agent): threaded intentTracing option through append-only context
- Exported `normalizeTools` so `AppendOnlyContext` uses the same tool normalization as the agent loop.
- Added `BuildOptions.intentTracing` to `build()`/`reset()`/`takeSnapshot()` so intent injection is consistent and included in the prefix fingerprint.
- Improved `#computeDigest` to cover tool_calls, tool_call_id, name, and id fields to catch in-place mutations.
- Fixed `#unsubscribeAppendOnly` leak and added no-op guard in `#syncAppendOnlyContext`.
2026-05-25 14:06:44 +02:00
Brit 79973d69de fix: preserve system prompt array structure in fingerprint 2026-05-24 23:01:41 +02:00
Brit e71d2ff49e fix: detect content rewrites in syncMessages, reset append-only cache on model switch 2026-05-24 22:39:20 +02:00
Brit ee04dd8e21 chore: bun check 2026-05-24 22:17:15 +02:00
Brit 2a86049f0f feat(agent): add append-only context mode for DeepSeek prefix-cache stability
ImmutablePrefix caches system prompt + tool specs after first build()
so subsequent turns reuse identical byte sequences. AppendOnlyLog
converts messages once via syncMessages() and only appends deltas
on further turns — prior-turn bytes stay stable.

- New module: packages/agent/src/append-only-context.ts
  StablePrefix, AppendOnlyLog, AppendOnlyContextManager
- AppendOnlyContextManager added to AgentLoopConfig
- Wired into streamAssistantResponse in agent-loop.ts
- Toggleable via provider.appendOnlyContext setting (auto/on/off)
- Default auto enables for deepseek provider
- 38 tests covering prefix, log, sync, compaction handling
- /session info surfaces current active state
2026-05-24 22:08:53 +02:00
can1357andCan Bölük 796f963da9 feat(coding-agent): added coding-agent follow-up queue with onBeforeYield
- Added optional `onBeforeYield` configuration and `setOnBeforeYield` in Agent, executed before follow-up checks.
- Added `YieldQueue` to `AgentSession`, with setup/teardown and streaming/idle flush via `setOnBeforeYield`.
- Replaced immediate async-result follow-up dispatch with queued batch entries, including stale-state suppression.
- Added MCP follow-up queueing in SDK, deduplicating updates by `serverName` and `uri`.
- Added changelog entries for `onBeforeYield`, async-result batching, MCP dedupe, and `display.shimmer` modes.
- Added yield queue unit tests for streaming emission, debounced idle batches, stale filtering, and error isolation.
2026-05-22 13:08:51 +09:00
can1357 bfae4d46c3 fix(coding-agent): tracked acp tool args by session for replay
- Tracked ACP tool-call inputs per session and replayed them via `toolArgsById`/`getToolArgs` plumbing.
- Merged ACP tool execution end content from start and result events so command output replay preserves original args.
- Scoped ACP async-job draining by session `ownerId` and `agentId` with in-flight tracking and permission-gated deferred turns.
- Refactored compaction telemetry and async tests with per-test telemetry setup and asynchronous teardown resets.
2026-05-17 13:15:07 +02:00
jiwangyihao 0aca488267 Merge remote-tracking branch 'origin/main' into acp-todo-plan-sync 2026-05-17 11:47:14 +08:00
can1357 1de85ee708 test(agent): added trace/context reset before otel test setup
- Reset global trace and context state before each test suite to prevent stale providers from leaking across test runs.
2026-05-17 05:10:56 +02:00
jiwangyihao ff6e92e8e4 test(agent): isolate compaction telemetry tracer 2026-05-17 03:05:45 +08:00
can1357 32453aaff0 feat(agent): added AgentTelemetry across compaction and branch-summary
- Added optional AgentTelemetry to summary, handoff, branch-summary, and compact option types.
- Replaced one-shot `completeSimple` usage with `instrumentedCompleteSimple` across compaction, summary, and branch-summary calls and passed `oneshotKind`.
- Added `PiGenAIAttr.OneshotKind`, `InstrumentedChatSpanOptions`, and response-header forwarding in telemetry span lifecycle.
- Added `resolveTelemetry` propagation in coding-agent session and inspect-image paths to pass request-scoped telemetry.
- Added compaction telemetry test harness and span assertions for success, no-telemetry, and error cases.
2026-05-16 18:03:07 +02:00
can1357 feb6c74c84 feat(agent): implemented header-based gateway detection in agent telemetry
- Added MockResponse metadata fields and invoked onResponse pre-stream with lowercased headers, status fallback, and requestId.
- Wrapped request onResponse in agent-loop, captured response headers, and forwarded them with baseUrl to finish/fail span handling.
- Added detectGatewayFromHeaders export and pi.gen_ai.gateway.* span attributes via header-based gateway detection.
- Extended telemetry event/span payloads with responseHeaders and validated detection-priority and onResponse forwarding in tests.
2026-05-15 23:46:23 +02:00
can1357 8b1364d4c2 feat(coding-agent): implemented one-shot generateHandoff in coding-agent
- Replaced event-driven handoff with one-shot `generateHandoff(...)` via `completeSimple`.
- Added cancellable `/handoff` command handling with a loader and Escape-to-abort flow.
- Removed legacy `compaction/handoff.ts` exports and added `generateHandoff(messages, model, apiKey, options)`.
- Fixed pre-cancelled handoff behavior to return `Handoff cancelled` and propagate abort signals.
- Updated handoff tests/mocks to assert `generateHandoff` invocation details and `AgentSession.handoff()` options.
2026-05-15 18:49:53 +02:00
can1357 e1aaf78874 refactor(compaction): moved compaction APIs to @oh-my-pi/pi-agent-core
- Relocated compaction, branch-summarization, pruning, and utils from coding-agent to packages/agent/src/compaction.
- Moved OpenAI remote compaction helpers from packages/ai to the new compaction module.
- Added handoff.ts with extractHandoffDocument, createHandoffContext, and renderHandoffPrompt helpers.
- Exposed new entries.ts with standalone SessionEntry types so coding-agent no longer owns them.
2026-05-15 18:31:12 +02:00
can1357 3c6d2346ad fix(agent): corrected run telemetry counters, cached input tokens, and span hook safety
- Skipped tool double-count: runTool's interrupt early-return no longer records the skipped tool inline; the post-batch tail sweep now handles accounting once per record so a single queued steering cancellation no longer logs N+1 skips.
- Aborted/errored assistant messages with embedded tool calls now record a collector orphan with status 'aborted' or 'error', so coverage.toolsInvoked and tools counters reflect them.
- run-collector chat record now stores inputTokens = input + cacheRead + cacheWrite, matching ChatUsageEvent and the public AgentRunSummary contract.
- onSpanStart and onSpanEnd hook invocations are wrapped in safeOnSpanStart/safeOnSpanEnd; thrown user callbacks surface via onTelemetryWarning (on_span_start_failed / on_span_end_failed) instead of leaking through finishChatSpan/finishExecuteToolSpan/finishInvokeAgentSpan/recordHandoff.
- summarizeTelemetryValue gained a depth+ancestor guard for arrays (matching the existing object recursion guard); cycles return '[Circular]' and over-depth returns the bounded {kind:'array',length} sentinel.
2026-05-15 17:48:02 +02:00
can1357 2867e1f4e3 feat(deps): added pi.zod exports and removed TypeBox package exports
- Added canonical `pi.zod` schema API exports and removed TypeBox package exports/imports.
- Migrated Tool schema typing from TypeBox to shared `TSchema`/Zod flow with legacy TypeBox compatibility.
- Updated AI provider adapters and MCP/agent builders to convert tool params through `toolWireSchema()`.
- Reworked schema validation from AJV to Zod-safe parsing with `fromTypeBox`, `toolWireSchema`, and meta schema checks.
2026-05-15 14:46:54 +02:00
can1357 80103f383d feat(agent): added OpenAI/PiGenAI OTEL constants and run summaries
- Swapped deprecated telemetry keys for OpenAIAttr/PiGenAIAttr constants, including tool intent keys.
- Normalized provider names via mapProviderNameToOtel and emitted OTEL pi.gen_ai fields on chat/request/response spans.
- Revised usage and aggregate telemetry to include cache-token totals plus failed and skipped step counts in run summaries.
- Updated OTEL and run-summary tests to use z.object schemas and renamed GenAI/PiGenAI attribute assertions.
2026-05-15 14:46:54 +02:00
can1357 1e601b9094 test(agent): replaced agent stream mocks with createMockModel responses
- Replaced custom MockAssistantStream helpers with createMockModel streams across agent tests.
- Removed manual queueMicrotask stream-event scripting in favor of scripted mock responses.
- Consolidated helper fixtures by deleting local aliases and reusing shared user-message/model helpers.
- Updated test assertions to use mock.calls and mock.model metadata for call and context validation.
2026-05-15 14:46:54 +02:00
can1357 d1412eac98 feat(agent): added generic telemetry extension hooks
- Added resolveAttributes, normalizeProvider, normalizeAgentName, onCostDelta, onTelemetryWarning, and contentSerializer hooks to AgentTelemetryConfig.
- Added "summary" capture mode emitting bounded dashboard-friendly span attributes without full payloads.
- Added recordManualChatTelemetry for instrumenting non-loop model calls.
- Introduced TelemetryAttributeContext as a base for TelemetryHookContext.
2026-05-15 14:46:54 +02:00
can1357 716b6c8234 feat(agent): added run-end tracking to emit telemetry.onRunEnd only once
- Added `runEnded`/`markRunEnded()` tracking and only fired `telemetry.onRunEnd` once per run.
- Added `agent_end` telemetry support for per-run `telemetry`/`coverage` and `agentLoopDetailed()` with `detailed()`.
- Added `aggregateAgentRunSummaries`/`aggregateAgentRunCoverage` and mapped `execute_tool` outcomes to `blocked` and `skipped`.
- Updated `finishInvokeAgentSpan` to derive failure `error.type`/status text from run status and exception state.
- Added run-summary test helpers covering `agent_end`, aggregation, and `onRunEnd` warning/compatibility scenarios.
2026-05-15 14:46:54 +02:00
can1357 e4e2389f83 refactor(agent): converted GenAI telemetry attributes to a const enum
- Replaced `GenAIAttr` with an `export const enum` in telemetry while preserving all GenAI attribute constants.
- Updated the OTEL stream test fixture to emit an `error` event with an `error` payload instead of a `done` event.
- Added runSubprocess telemetry propagation tests for inheriting parent telemetry and handling missing parent telemetry.
2026-05-15 14:46:54 +02:00
can1357 4679789cf1 feat(agent): added OTEL spans for agent invoke/chat/tool flows
- Added opt-in telemetry configuration to Agent and session APIs, including Agent#setTelemetry mutator.
- Implemented OpenTelemetry spans for invoke_agent, chat, execute_tool, and handoff paths with metadata and step tracking.
- Added a telemetry helper module, OpenTelemetry request/usage types, and dependency wiring with no-op behavior when tracer SDK is absent.
- Added OTEL end-to-end tests and fixed coding-agent OutputSink realignment and artifact-link newline output issues.
2026-05-15 14:46:53 +02:00
can1357 c29fe28dbb feat(agent): add beforeToolCall and afterToolCall hooks
Mirrors the pi-mono API surface so apps can preflight tool execution
(block or mutate validated args) and post-process tool results
(override content/details/isError) without wrapping tools.

- `AgentLoopConfig.beforeToolCall` runs after argument validation. Return
  `{ block: true, reason }` to short-circuit with a tool-error result.
  Mutations to `context.args` are forwarded to `tool.execute` without
  revalidation, matching pi-mono semantics.
- `AgentLoopConfig.afterToolCall` runs after execution and before
  `tool_execution_end` / tool-result message emission. Returned fields
  override the executed result; omitted fields fall through. Hook
  exceptions surface as tool errors and do not abort the batch.
- `Agent` exposes both hooks as public, reassignable fields so extension
  reloads can swap implementations mid-session.

Compatibility: fully additive. Both hooks default to undefined and the
loop behaves identically when neither is set. The internal
`executeToolCalls` signature was collapsed to `(context, message,
signal, stream, config)` -- it is not exported, so this is not a public
API change. Pi-mono's `terminate` field on `AfterToolCallResult` is
omitted because our `AgentToolResult` has no batch-level early-stop
contract.
2026-05-15 09:16:44 +02:00
can1357 f1f6516056 refactor: reorganized exports and removed obsolete helper branches
- Removed export leakage by demoting many helper and const symbols to module-local scope.
- Renamed underscore-prefixed internals and cache fields, then updated related references and `satisfies never` checks.
- Deleted obsolete logic branches and helpers, including harmony-stream interruption flow and unused benchmark runtime helpers.
- Updated Biome config and manifests by broadening lint coverage and removing an unused `@napi-rs/cli` dev dependency.
- Adjusted tests and utilities to use renamed test helpers and remove redundant private test-only helpers/locals.
2026-05-14 04:36:19 +02:00
can1357 975941aba4 chore: remove garbage tests 2026-05-12 04:09:33 +02:00
can1357 8b92ec937e feat: gpt-5 harmony errata fixes
- Replaced `===== ... =====` eval cell headers with `*** Begin ` / `*** End ` markers; legacy format remains renderable in HTML exports.
- Replaced hashline patch grammar with `*** Begin Patch` / `*** End Patch` envelope; old inputs without the envelope are still accepted.
- Extracted `sniffEvalLanguage` into a shared `sniff.ts` module reused by the parser and tool.
- Added `docs/ERRATA-GPT5-HARMONY.md` and `scripts/session-stats/harmony_backtest.py` documenting and backtesting the GPT-5 Harmony-header leak defect.
2026-05-10 19:52:44 +02:00
Miroslav Drbal fc70a45c46 fix(ai): stable metadata.user_id per session for Anthropic OAuth
Anthropic counts sessions by metadata.user_id. Without this fix, OMP
generated fresh random entropy on every API request, inflating the
session count and preventing backend attribution to the authenticated
account.

Changes:

packages/ai:
- resolveAnthropicMetadataUserId() now accepts JSON-format user_id
  matching real Claude Code's getAPIMetadata shape
  ({ session_id, account_uuid, ... }). Previously only the legacy
  cloaking format was accepted on OAuth, causing stable caller-supplied
  values to be silently discarded.
- AnthropicOAuthFlow.exchangeToken() and refreshAnthropicToken() now
  populate OAuthCredentials.{accountId, email} from the token response
  account block, removing the need for a separate /api/oauth/profile
  round-trip.
- AuthStorage.getOAuthAccountId(provider, sessionId) returns the OAuth
  accountId for the session-sticky credential, used to build
  account_uuid in metadata.user_id. Guards against misattribution for
  API-key, runtime-override, env-key, and fallback-resolver paths that
  do not record a session credential.

packages/agent:
- Agent.metadataForProvider(provider) resolves request metadata for
  the given provider via the installed resolver, or returns the static
  metadata value. The plain metadata getter now returns only the static
  value; provider-aware resolution is explicit.
- Agent.setMetadataResolver(fn) installs a (provider: string) resolver
  evaluated per LLM request in agent-loop, after getApiKey records the
  session-sticky credential, so account_uuid reflects the credential
  actually used.
- AgentLoopConfig.metadataResolver is called with config.model.provider
  after getApiKey, overriding the static metadata field.

packages/coding-agent:
- AgentSession.#syncAgentSessionId installs a metadata resolver that
  builds { user_id: JSON.stringify({ session_id, account_uuid? }) },
  matching the Anthropic session attribution format. account_uuid is
  only included for provider="anthropic" to avoid leaking the OAuth
  identity to third-party Anthropic-format-compatible providers.
- sessionId getter prefers providerSessionId when supplied via
  AgentSessionConfig so all API paths (getApiKey, direct calls,
  metadata resolver) share the same provider-facing session ID.
- prepareSimpleStreamOptions stamps session metadata on direct calls
  (runEphemeralTurn, compaction, branch summary, title generation) so
  they share the same session bucket as Agent.prompt requests.
- generateBranchSummary and generateSessionTitle accept a
  (provider: string) metadata resolver evaluated after their own
  getApiKey call for correct credential attribution.
2026-05-09 09:48:10 +02:00
can1357 8c323666be feat: added ordered systemPrompt arrays and normalized context prompts
- Converted systemPrompt APIs and state types to ordered `string[]` across agent, AI, and coding-agent surfaces.
- Added `normalizeSystemPrompts` and applied it to context normalization before building provider request payloads.
- Updated AI providers to emit separate normalized prompt blocks/messages instead of a single merged system prompt.
- Removed dedicated `projectPrompt` state and remapped that context into system-context buckets in session, dump, and token accounting.
- Aligned tests and changelogs to pass and assert `systemPrompt` as arrays with ordered prompt semantics.
2026-05-04 15:20:26 +02:00
can1357 a60e492fba feat(agent): enabled dynamic reasoning override per model call
- Added an optional `getReasoning` callback to `AgentLoopConfig` to resolve reasoning effort dynamically for each LLM call.
- Updated the agent loop to resolve reasoning via `getReasoning` and use it in place of static `reasoning` when provided.
- Added a test confirming a run re-reads the thinking level between consecutive model calls when it changes mid-run.
2026-05-03 07:39:52 +02:00
can1357 b51ca3f7ee fix(ai): corrected AI abort precedence and provider stream normalization
- Fixed abort-source handling so caller aborts always win and local reasons only attach to matching request signals.
- Fixed agent-loop streaming to race event reads against abort signals and emit an aborted assistant message.
- Fixed Anthropic request construction to honor thinkingEnabled=false and omit temperature/top_p/top_k for Opus non-thinking.
- Fixed OpenAI Codex request handling by normalizing response URLs, decoding non-string websocket frames, and cleaning handshake headers.
- Added regression tests for abort precedence, Anthropic alignment cases, and Codex stream/header normalization.
- Documented the cancellation and provider behavior fixes in package changelogs.
2026-05-02 01:06:53 +02:00
can1357 ecd1554eba feat: renamed subagent handoff flow to use yield instead of submit_result
- Renamed subagent completion flow from `submit_result` to `yield` across SDK tools, prompts, and docs.
- Updated executor/task handling to require and parse `yield` calls, replacing legacy submit-result extraction and state flags.
- Added `subagent-yield-reminder` and updated system prompts to require `yield` with `result.data` or `result.error`.
- Renamed hidden-tool and registration plumbing to `yield`, including discovery helpers and renderer/test surface.
2026-04-26 00:29:52 +02:00
can1357 d24d11a274 fix: resolved AI/OAuth helper duplication via shared modules
- Standardized missing-file read errors and now return `File not found: <path>` for absent edit targets.
- Centralized AI provider, usage, and OAuth helpers into shared modules to remove duplicated logic.
- Migrated OAuth/API-key login flows to shared factory helpers and removed inline prompt/token-exchange code.
- Reused shared tools and formatter utilities for discovery, stream tails, LSP batching, and source formatting.
- Consolidated repeated test helpers and fixtures into shared modules, replacing inline helper duplicates.
2026-04-23 21:02:14 +02:00
can1357 7fb18faf4c fix(coding-agent): backported pi-mono changes (1feccfed..b21b42d0)
packages/ai:
- feat: expose provider responseId on AssistantMessage
- feat: lazy-load provider modules for faster startup
- fix: hash foreign Responses API tool call IDs exceeding 64-char limit
- fix: ignore null chunks in openai-completions streams
- fix: keep image tool results inline for Gemini 3+ and Antigravity
- fix: correct Bedrock Claude 4.6 context window to 200k
- fix: support prompt caching for Bedrock application inference profiles
- fix: add OpenRouter reasoning payload format
- fix: ignore placeholder Vertex API keys
- fix: skip AJV validation in restricted runtimes
- fix: Anthropic OAuth client injection and responseId extraction
- fix: Codex incomplete/failed response status handling

packages/agent:
- fix: defer steering until after tool execution completes

packages/tui:
- feat: namespaced keybinding IDs with KeybindingsManager conflict detection
- feat: configurable select list column sizing (#2154 by @markusylisiurunen)
- fix: stream truncateToWidth for large strings
- fix: skip Termux height redraws
- fix: stop evicting unrelated default keybindings
- fix: resolve raw backspace ambiguity on Windows Terminal
- fix: clear stale scrollback on session switch (#2155 by @Perlence)
- fix: remove trailing markdown block spacing (#2152 by @markusylisiurunen)

packages/coding-agent:
- feat: add resizable share sidebar (#2435 by @dmmulroy)
- feat: emit OSC 133 command-executed marker
- feat: reload custom themes from disk watcher
- feat: add --fork session flag
- feat: file mutation queue for serialized writes
- feat: initial message consolidation utility
- fix: keybindings migrated to namespaced IDs
- fix: resolve waitForRetry() race when auto-retry produces tool calls
- fix: handle slash-delimited /model refs
- fix: refresh active model after provider updates
- fix: extended transient error patterns for retry
2026-03-22 18:28:40 +01:00
94ad651d17 feat: add MCP tool discovery search (#352)
* Add MCP tool discovery search and live refresh

* Fix MCP discovery review feedback

* Address remaining MCP discovery review comments

* feat: compact MCP discovery search results

* fix: align MCP discovery search contract

* feat: add MCP server tool counts to discovery hints

* fix(agent): corrected stale toolChoice validation against active tools

- Fixed stale forced toolChoice passed to provider after mid-turn tool refresh by validating against active tools.
- Added refreshToolChoiceForActiveTools() to filter invalid tool choices when available tools change.
- Changed getToolChoice config to use computed function instead of static property for dynamic validation.
- Fixed MCP tool selection tracking in coding-agent to distinguish between discovery-enabled and non-discovery sessions.
- Updated search_tool_bm25 to filter already-selected tools before applying limit parameter.

---------

Co-authored-by: can1357 <me@can.ac>
2026-03-16 13:43:43 +01:00
can1357 02a32d5b9a refactor: migrated OpenAI stream processing to shared module
- Extracted OpenAI Responses API stream processing logic into shared module with 424 lines of reusable utilities.
- Consolidated text signature encoding, tool call normalization, and message conversion functions into openai-responses-shared module.
- Refactored openai-responses.ts and azure-openai-responses.ts to delegate stream processing to shared processResponsesStream() helper.
- Removed 473 lines of duplicated stream event handling and utility functions across OpenAI provider implementations.
2026-03-10 07:43:36 +01:00
can1357 1962938fa5 feat: added payload interception and signature metadata for provider requests
- Added `onPayload` callback option to intercept and transform provider request payloads before transmission across agent and AI packages.
- Added structured text signature metadata with phase information to OpenAI and Azure OpenAI providers for enhanced response tracking.
- Added `before_provider_request` extension event to coding-agent for chaining payload transformations across multiple handlers.
- Improved error messages in `response.failed` events with detailed error codes, messages, and incomplete reasons from provider responses.
2026-03-10 07:38:13 +01:00
can1357 8e3e0ebf9e feat: introduced Effort enum and ThinkingConfig for model-aware reasoning
- Introduced Effort enum and ThinkingConfig metadata for per-model reasoning capabilities with min/max effort levels.
- Migrated thinking level API from string-based ThinkingLevel to structured Effort enum across agent and AI packages.
- Added model-thinking module with effort mapping, policy application, and semantic versioning utilities for provider-specific thinking modes.
- Removed supportsXhigh() function and replaced effort clamping with model-aware validation using ThinkingConfig metadata.
- Expanded models.json with thinking configuration objects for 50+ models including Claude, Gemini, and OpenAI variants.
- Added Python analysis scripts for edit tool usage patterns and tool invocation stream processing.
2026-03-06 12:35:01 +01:00
can1357 8c89b5c38d refactor(agent): extracted tool result emission into dedicated function
- Extracted tool result emission logic into dedicated emitToolResult function to eliminate duplication.
- Consolidated tool execution result handling to emit results immediately after execution rather than deferring to post-processing loop.
- Simplified post-execution loop by delegating result emission to emitToolResult and removing redundant message construction.
2026-03-01 06:04:29 +01:00