Commit Graph
881 Commits
Author SHA1 Message Date
can1357 0eb21efa1a Merge remote-tracking branch 'origin/farm/7ef98714/snapcompact-copilot-vision-gate' 2026-06-24 21:00:37 +02:00
can1357 0f2737f16e Merge remote-tracking branch 'origin/farm/e322f828/eval-agent-yield-terminal' 2026-06-24 20:59:54 +02:00
roboomp 997b2b24ff fix(session): suppressed empty-stop retry after successful yield
Trailing empty assistant 'stop' arriving after a successful 'yield'
revived the already-yielded subagent. AgentSession.agent_end maintenance
compared #assistantEndedWithSuccessfulYield(msg) against the trailing
empty-stop message — not the yield-bearing one — so the empty-stop
recovery path appended a retry reminder and scheduled agent.continue().

Track a sticky #yieldTerminationPending flag set when the yield tool
finishes without error and cleared on the next #promptWithMessage. The
agent_end routing extends the existing successful-yield branch: when the
flag is set, or the current message ended with yield, short-circuit
empty-stop / unexpected-stop / compaction continuations for the rest of
the run, so a successful yield is terminal regardless of trailing stops.

Fixes #3389
2026-06-24 17:08:49 +00:00
roboomp 714051d795 fix(catalog,coding-agent): disable vision on non-personal copilot endpoints
GitHub Copilot's /models response advertises supports.vision = true for
Claude/GPT chat models on every host, but only the canonical personal
endpoint (https://api.githubcopilot.com) actually accepts image inputs;
the business (api.business.githubcopilot.com) and enterprise
(copilot-api.{domain}) hosts respond '400 vision is not supported'.
snapcompact then injected rasterized transcript frames after compaction
and permanently broke every business-Copilot session.

- Catalog discovery (githubCopilotModelManagerOptions.mapModel) now
  forces input=['text'] whenever the resolved baseUrl is not the
  canonical personal-Copilot host, so the upstream's vision flag is
  honoured only where it actually works.
- mergeDynamicModel honours the dynamic input value (instead of
  OR-upgrading with the bundled reference) when the merged baseUrl
  differs from the bundled one, so a bundled spec pinned to the
  personal host can no longer taint a business-resolved merge.
- snapcompact-inline's canSendImages helper short-circuits the
  rasterizer for any github-copilot model whose baseUrl is non-personal,
  catching stale cached specs that still advertise vision.
- Helper isPersonalGitHubCopilotBaseUrl exported from
  pi-catalog/wire/github-copilot so catalog and coding-agent share one
  canonical check.

Regression coverage in github-copilot-model-limits.test.ts (vision
endpoint policy + full merge) and snapcompact-inline.test.ts (#3387
business/enterprise case).

Fixes #3387
2026-06-24 16:27:14 +00:00
can1357 9c56956631 fix(tui): kept recent turns visible when collapsing remote compactions
With collapseCompactedHistory the live display fell into the LLM compaction
branch, which skips the firstKeptEntryId..compaction turns whenever an OpenAI
remote-compaction replacementHistory payload is present. That payload feeds the
provider only and is not rendered, so a remotely-compacted session showed just
the summary plus post-compaction rows, hiding recent turns that were visible
before. Emit the kept SessionEntry rows in transcript mode regardless. Adds a
regression.
2026-06-24 18:26:20 +02:00
can1357 c04747c5ac Merge PR #3259: fix(tui): reduce large transcript stalls (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 3117a20ab9 Merge PR #3249: fix(agent): size snapcompact maxFrames by the live model window (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 ef67f685ac Merge PR #3380: fix(coding-agent): clamp auto thinking to undefined for models without controllable effort (@roboomp) 2026-06-24 18:26:19 +02:00
can1357 2b42f19456 Merge PR #3232: fix(agent): clamp provider context images (@roboomp) 2026-06-24 18:23:48 +02:00
roboomp b0e07f52d2 fix(coding-agent): clamped auto thinking to undefined for models without controllable effort
Devin provider models (devin-agent) advertise reasoning: true but no
thinking.efforts metadata — Cascade selects effort by routing to sibling
model ids, not a wire param. getSupportedEfforts(model) therefore returns
[]. clampAutoThinkingEffort previously short-circuited that empty supported
list by returning the requested effort as-is, so the auto-thinking
classifier-resolved level (e.g. low) reached stream.ts:1163 where
requireSupportedEffort threw 'Thinking effort low is not supported by
devin/<id>. Supported efforts: '. In --print mode the user saw the error
text; in the TUI it was silently swallowed, producing the reported
'working then empty response' symptom.

Returns undefined when supported is empty so the result mirrors
clampThinkingLevelForModel's behavior on the same shape (the explicit
--thinking low / high paths already worked because of this). Updates
classifyDifficulty's return type to Effort | undefined and threads through
to the existing #applyAutoThinkingLevel undefined-effort early-return.
#applyAutoThinkingLevel also short-circuits the classifier call up front
for these models — there is no effort to pick.

Fixes #3356
2026-06-24 14:12:18 +00:00
roboomp 24adad2890 fix(providers): honored llama cpp model context
Read per-model llama.cpp meta.n_ctx values during discovery, refresh selected models after lazy load, and bypass fresh cache reuse for llama.cpp refreshes so server restarts update context windows.\n\nFixes #3310
2026-06-23 12:23:59 +00:00
can1357 44b62419c5 Merge remote-tracking branch 'origin/farm/fc89a4e3/goal-auto-compaction-still-not-triggering' 2026-06-23 00:01:47 +02:00
roboomp 29e69f5364 style: bun run fix 2026-06-22 20:02:46 +00:00
roboomp a305e68a53 fix(compaction): compact goal runs between tool turns
Active goal loops can stay inside one agent run while the model keeps
emitting tool calls, so the normal agent_end threshold maintenance never
runs. That lets context grow past the soft threshold until provider
overflow or user abort.

Run threshold maintenance from the per-turn onTurnEnd hook for active
goals, splice the compacted agent state back into the live loop message
array, and suppress queued continuations because the current run is
already continuing. Cover the mid-run tool-call path and the non-goal
control case.

Refs #3174
2026-06-22 20:02:33 +00:00
can1357 5c21b28786 feat: optimized handoff generation and harden request safety
- Introduced `generateHandoffFromContext` to enable provider-aware oneshot generation and improved cache hit rates via the live-turn pipeline.
- Updated `buildSideRequestContext` to support pinning custom system prompts, preventing per-turn hook leakage during handoff.
- Added concurrency guards across CLI and RPC modes to block manual `/handoff` requests while a session is actively streaming.
- Standardized handoff execution to force `toolChoice: "none"` and enforce consistent cache-routing behavior.
2026-06-22 20:05:40 +02:00
can1357 1082b6316a Merge remote-tracking branch 'origin/farm/5ce2a200/hide-secrets-blocks-advisor' 2026-06-22 17:44:48 +02:00
can1357 beb61b1831 Merge remote-tracking branch 'origin/farm/f3af36ee/clear-openai-completions-on-switch' 2026-06-22 17:43:34 +02:00
roboomp 45b200cd09 fix(coding-agent): evicted resolved completions session URLs
The completions provider stores session state under the request-time resolved base URL, which can differ from the catalog baseUrl for Moonshot, Alibaba Coding Plan, Azure deployments, and similar provider overrides. The model-switch cleanup now evicts the previous provider prefix whenever the switch leaves that completions backend, so those resolved-url keys cannot survive the switch.
2026-06-22 13:39:48 +00:00
roboomp 596a2d1317 style: bun run fix 2026-06-22 13:30:10 +00:00
roboomp 8cebe98d2f fix(coding-agent): evicted openai-completions provider session state on backend switch
`AgentSession.#closeProviderSessionsForModelSwitch` only handled
`openai-codex-responses` and `openai-responses:<provider>` keys. The
`openai-completions:<provider>:<baseUrl>:<modelId>` entries — which cache
strict-tools disable scopes and reasoning-effort fallbacks tied to the
upstream backend — survived /model switches between different providers or
base URLs, so the next request to that backend (e.g. on /model toggle
back) replayed stale decisions made against an entirely different
transport.

Switching to a model whose `(provider, baseUrl)` differs from the current
openai-completions model now evicts every cached entry sharing the old
prefix. Same-backend model toggles keep their cached state, matching the
existing codex/responses semantics.

Fixes #3260
2026-06-22 13:30:01 +00:00
roboomp 3bcbf1515d fix(tui): reduced large transcript stalls
Tail appended transcript JSONL instead of rebuilding rendered history on every poll, collapse compacted history for live chat rendering, and replace synchronous session rewrites so tailers detect historical changes.

Fixes #3258
2026-06-22 12:01:34 +00:00
roboomp 232994496d fix(agent): size snapcompact cap reserve from live shape's text-edge cost
chatgpt-codex third-pass review on #3249: the 4k SUMMARY_TEXT_RESERVE
in the cap math undersized the actual textHead+textTail cost a frame-
bearing archive carries (the projection separately bills
'countTokens(summary + textHead + textTail)'). At ~120k headroom on
Anthropic 11on16-bw, the cap picked maxFrames=23, but
'23 * 5024 + 2 * 13916 chars (≈7k tokens) + 2k summary template ≈ 124.5k'
still exceeded the same 120k headroom — the cap chose a value the
projection then immediately rejected, re-opening the warning loop.

#computeSnapcompactMaxFrames now resolves the live snapcompact shape
(same call the auto/manual paths pass to snapcompact.compact) and sizes
the cap reserve from 'geometry(shape).capacity':

  textEdgeTokens = ceil(2 * capacity * 1.15 / 4)   // 1.15 absorbs
                                                   // tokenizer drift
  capReserve     = textEdgeTokens + 2000           // + summary template

For the default per-provider winners that resolves to ~10k (Anthropic
Sonnet), ~14k (Opus 4.7), ~16k (Gemini 2.x), and ~10k (OpenAI) — all
larger than the prior fixed 4k. Skip decision stays separate
(baseTokens >= totalBudget), so positive sub-reserve headroom still
runs snapcompact's text-only path.

Test 1 retuned to baseline kept-recent ≈ 100k tokens with a strengthened
assertion verifying the FULL projection invariant (frames + worst-case
text edges + summary template + base ≤ budget). Confirmed test fails
against the previous 4k-reserve helper by exactly the reviewer's
predicted margin (174,271 vs 170,000 budget = 4,271 token overshoot).
2026-06-22 10:38:24 +00:00
roboomp db57efc3d3 fix(agent): split snapcompact skip reserve from frame-cap reserve
chatgpt-codex second-pass review on #3249: the previous helper folded
the 4k SUMMARY_TEXT_RESERVE into both the maxFrames cap math AND the
skip decision (return 0 when frameBudget < 0). That made any residual
headroom below 4k fall negative and force the LLM-summarizer fallback,
even though a text-only snapcompact archive (the 'text.length <= 2 *
edgeCap' short-circuit in planArchive) typically costs only a few
hundred tokens of summary lead-in and would have fit cleanly.

The two reserves now serve their own jobs:

- Skip iff 'baseTokens >= totalBudget' (kept-recent + non-message
  already eats the entire window − reserve envelope). No reserve
  fudge here; positive residual is always worth attempting.
- Cap reserve (4k) is applied ONLY to the maxFrames calculation so
  the projection still passes once frames land. When the frame budget
  goes negative under that reserve but residual headroom is positive,
  the helper now returns maxFrames=1 instead of 0 so snapcompact's
  frame-less planArchive branch can still produce a valid archive.

Updated regression test to pin the new contract directly: kept-recent
tuned for 1500 tokens of headroom (well below the 4k cap reserve), the
old helper returned 0 and skipped to the LLM summarizer, the new helper
invokes snapcompact with maxFrames=1.
2026-06-22 10:24:36 +00:00
roboomp 65f945f1b7 fix(agent): preserve snapcompact text-only path when budget is near full
chatgpt-codex review on #3249: the helper returned 0 when frameBudget
< FRAME_TOKEN_ESTIMATE, causing the caller to skip snapcompact entirely.
But snapcompact.planArchive has a 'text.length <= 2 * edgeCap' short-
circuit that produces a valid frames:[] archive when the discarded
history is small enough — and the projection charges 0 for that. Hard
return-0 blocked that opportunity, forcing the LLM summarizer fallback
in offline/no-credential sessions where the text-only path would have
landed cleanly.

#computeSnapcompactMaxFrames now distinguishes two near-full cases:
  - frameBudget < 0 → return 0 (kept-recent already exhausted budget;
    no text-only summary can fit either) → caller still skips outright.
  - 0 ≤ frameBudget < FRAME_TOKEN_ESTIMATE → return 1 → snapcompact runs
    and picks the frame-less planArchive branch automatically for small
    discarded histories; the projection guard rejects any actual
    frame-bearing archive that overflows.

Added regression test pinning maxFrames=1 (not 0) in the near-full
window case.
2026-06-22 10:16:25 +00:00
roboomp 5cce507582 fix(agent): size snapcompact maxFrames by the live model window
Snapcompact's bundled MAX_FRAMES_DEFAULT (80) × FRAME_TOKEN_ESTIMATE (5024)
≈ 402k tokens worth of frames. AgentSession was calling snapcompact.compact()
with no maxFrames override, so the post-render projection inside #runAuto
Compaction / compact() always overflowed the budget on any sub-1M-token
window (Claude Sonnet 4.5's 200k = 170k usable, the 80-frame projection
alone clears that 2.4×), looping the 'snapcompact could not bring the
context under the limit — using an LLM summary instead' warning on every
threshold tick.

AgentSession.#computeSnapcompactMaxFrames now sizes the frame cap from
the resolved budget — (window − reserve − non-message − kept-recent −
summary-text reserve) / FRAME_TOKEN_ESTIMATE, clamped to MAX_FRAMES_DEFAULT
— and threads it into snapcompact.compact() in both the auto-compaction
and manual /compact paths. When the kept-recent slice already exceeds the
budget, snapcompact is skipped outright instead of running just to be
rejected: the projection guard remains as a defensive check.

Fixes #3247
2026-06-22 10:03:44 +00:00
roboomp 7dadeab331 fix(coding-agent): hid secrets from advisor prompts
Stopped restoring placeholders inside opaque assistant thinking blocks and threaded the configured secret obfuscator into advisor session-update prompts.

Added regression coverage for thinking preservation and advisor prompt redaction.

Fixes #3237
2026-06-22 07:05:23 +00:00
roboomp ce7a4e4562 fix(agent): preserved clamped tool image results
Added a textual omission marker when provider image clamping removes every block from a successful tool result, keeping the serialized tool_result meaningful and protocol-safe.\n\nFixes #3230
2026-06-22 06:05:38 +00:00
can1357 4e54e557bc feat: improved context usage display and update model configurations
- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
2026-06-22 08:04:42 +02:00
roboomp eb88ecd7bb fix(agent): clamped provider context images
Dropped oldest outgoing image blocks above the active provider budget so umans requests honor the shipped 10-image cap even when snapcompact is disabled. Added regression coverage for preserving text and newest images.\n\nFixes #3230
2026-06-22 05:57:30 +00:00
can1357 33e2594f03 feat(coding-agent): added support for Julia and display language icons in code cells
- Added Julia language support to the theme symbol maps.
- Enabled language icons in code cell headers for the eval tool renderer.
2026-06-22 06:13:01 +02:00
can1357 1f3f3cf5d1 feat: added ruby and julia language support to coding-agent
- Implemented persistent execution backends for Ruby and Julia using dedicated kernel processes and NDJSON-based IPC.
- Integrated language-specific prelude environments, runtime path resolution, and security-focused environment variable filtering.
- Exposed configuration options, tool schema updates, and lifecycle management for seamless agent interaction with both languages.
- Added comprehensive integration tests and updated prompt documentation to support the new evaluation capabilities.
2026-06-22 06:13:01 +02:00
can1357 7302a7ae96 feat(coding-agent): improved secret obfuscation and data protection
- Refined obfuscation logic to use granular, typed transformations instead of generic object traversal.
- Enforced an 8-character minimum for secret patterns and restricted redaction to user-authored content to prevent false positives.
- Preserved system prompts, tool schemas, and opaque remote replay data to maintain provider context and data integrity.
- Integrated protected snapshot exports with targeted redaction to safeguard sensitive information in shared sessions.
2026-06-22 01:13:52 +02:00
can1357 838b5162bc feat(coding-agent): standardized html export styling with brand palette
- Introduced `web-palette.ts` to implement the collab-web pink/purple brand identity for HTML exports.
- Updated `generateThemeVars` to support a `palette` option, allowing users to choose between the brand-web aesthetic and a specific TUI theme.
- Configured public exports and the share-viewer script to default to the brand-web palette rather than inheriting the user's terminal theme.
- Refactored `AgentSession.exportToHtml` to align with the new branding defaults while allowing for per-export theme overrides.
- Adjusted `dark.json` background values to ensure better visual consistency across internal surfaces.
2026-06-21 18:31:03 +02:00
can1357 f353ae0067 chore: biome format after PR integration 2026-06-21 17:20:22 +02:00
can1357 0ca62f1b8b Merge PR #1865: fix(session): keep auto thinking mode active across session resume (@msimon) 2026-06-21 17:09:50 +02:00
roboomp eef2f7b7e9 fix(compaction): resolve retry before goal continuation return
Active-goal threshold compaction can pre-empt the normal post-turn tail
and return once it schedules a deferred handoff or auto-continue. When
that turn is the successful response from an auto-retry, returning there
skips the later retry-gate cleanup and leaves isRetrying stuck.

Resolve the completed retry gate before the compaction-continuation
return, and cover the retry-success-over-threshold path so future
changes cannot strand prompt()/waitForIdle() behind a stale retry state.

Refs #3174
2026-06-21 09:08:59 +00:00
roboomp e70c71077e fix(compaction): log agent_end maintenance routing for goal turns
Reporter on #3174 still sees no auto-compaction with thresholdTokens
lowered to 32768 against a 70k+ visible context, and reports the
existing `Auto-compaction threshold decision` log never appears.
Either the goal turn never reaches `#checkCompaction`, or the log
fires but is filtered out of their view (winston is at debug level so
it writes to ~/.omp/logs/omp.<DATE>.log, not the TUI).

Add an `agent_end maintenance routing` debug log at every branch of
the `agent_end` handler — entered/no-message,
skip-post-turn-maintenance, successful-yield (active goal vs not),
empty-stop-handled, active-goal pre-empt (and whether it scheduled a
continuation), unexpected-stop-handled, and bottom checkCompaction —
together with stopReason, provider/model, content shape, goal
state, and `successfulYield`. Combined with the existing
`Auto-compaction threshold decision` log, the next no-start report
identifies the exact early-return branch and the inputs that fed
`shouldCompact`.

Refs #3174
2026-06-21 08:53:54 +00:00
roboomp 00d14accfb fix(compaction): keep empty-stop cleanup before active-goal compaction continuation
Codex review on #3175: the active-goal compaction pre-empt I added in
8ab754f636 short-circuited #handleEmptyAssistantStop. That handler is
the only path that strips an orphan toolUse assistant (stopReason
"toolUse" with no toolCall block) from both active context and the
session branch via #removeEmptyStopFromActiveContext. With the pre-empt
ordering, an over-threshold goal turn that returned an empty toolUse
left the orphan as the session leaf, and the compaction auto-continue
prompt fed it back into the next Anthropic turn as a tool_use with no
matching tool_result — the exact history-corruption pattern the
existing cleanup comment defends against.

Move #handleEmptyAssistantStop back ahead of the active-goal
compaction probe. Empty stops still self-retry and never reach the
threshold pre-empt; non-empty stops (the reporter's failure in #3174)
still hit threshold maintenance before the unexpected-stop classifier.

Regression test seeds a goal-mode empty toolUse stop billed at 91k
against thresholdTokens 76384 and asserts the threshold compaction
never starts and the orphan is no longer in the session branch.

Fixes #3174
2026-06-21 08:49:28 +00:00
roboomp 8ab754f636 fix(compaction): run goal threshold maintenance before retry continuations
Active goal turns that stopped with text could hit the empty/unexpected-stop
continuation guards before threshold maintenance. When those guards scheduled
another goal turn, #checkCompaction never ran, so no auto_compaction_start was
emitted even while visible context stayed above thresholdTokens.

Run threshold maintenance once before those active-goal self-continuations and
log the threshold decision inputs: billed context, stored estimate, resolved
trigger tokens, post-maintenance tokens, strategy, threshold, promotion state,
and shouldCompact.

Also pass post-prune maintenance tokens into the shake recovery-band check so
supersede/drop-useless savings are preserved when deciding whether shake still
needs to fall back to context-full compaction.

Fixes #3174
2026-06-21 08:35:58 +00:00
roboomp b74f7bb720 fix(compaction): trigger goal-mode threshold on billed context, not post-prune estimate
Pruning frees bytes for the NEXT prompt — it does not change the size of
the prompt the LLM just billed for. Subtracting the per-turn
`#pruneStaleToolResults` / `#pruneToolOutputs` savings from the
threshold input let a long-running `/goal` session sit above
`compaction.thresholdTokens` indefinitely: the visible context
(anchored to the same provider billing) showed >threshold, but
`shouldCompact` no-op'd because the subtraction dropped the input below
the trigger. The `compactionContextTokens` floor against the post-prune
local estimate is still applied, so a payload-compression hook still
can't deflate the trigger.

Regression test seeds one large `useless` tool result whose suffix sits
inside the 8k cache-warm window so `#pruneStaleToolResults` actually
returns ≥20k savings, then asserts compaction fires when the final turn
bills 91k tokens against the reporter's `thresholdTokens: 76384`.

Fixes #3174
2026-06-21 07:21:22 +00:00
roboomp 6f3d6ba2e4 fix(coding-agent): restricted plan-mode write activation to built-ins
Tracked current-registry built-in provenance through AgentSession so plan mode
only force-activates the built-in write implementation. Extension or SDK tools
that shadow the name `write` stay inactive, preserving plan mode's read-only
contract through the built-in write/edit guard.

Added a regression that registers a shadowing write tool without built-in
provenance and verifies plan mode does not activate it.
2026-06-21 04:04:14 +00:00
can1357 4d96bcf6b6 feat: replaced manual debug request path with automated llm dump
- Replaced the `/debug dump-next-request` command with an updated `/dump` command that exports LLM request context to JSON sidecar files.
- Removed persistent debug path state and manual path configuration in favor of automated generation.
- Updated session logic to handle serializing LLM request context to temporary directories.
- Refactored testing suites to remove path-based debug tests and verify dynamic request file generation.
2026-06-21 03:18:20 +02:00
can1357 29f7b57fde feat: migrated core rendering and context transformation to async
- Updated `render`, `renderMany`, and native snapcompact methods to return promises, ensuring scalable async execution.
- Refactored `transformProviderContext` and `buildSideRequestContext` to support asynchronous operations in agent loops.
- Integrated `Promise.all` for improved concurrency when processing frame rendering and rendering batch operations.
- Updated all internal call sites, SDK hooks, and test suites to accommodate the asynchronous API signatures.
2026-06-21 00:32:13 +02:00
can1357 aa9709c47f feat(coding-agent): added snapcompact safety checks and UI integration
- Added validation to scan for non-ASCII characters before performing snap-compaction, falling back to LLM-based summarization if the unrenderable ratio is too high.
- Updated event handling and status reporting to explicitly support snapcompact actions, including specific error warnings and cancellation states in the UI.
- Updated session logic to default to snapcompact strategy when auto-compaction is enabled.
2026-06-21 00:08:12 +02:00
can1357 3c356fc0e0 fix(coding-agent/session): improved auto-compaction strategy selection
- Adjusted the fallback strategy when enabling auto-compaction to respect the system default instead of hardcoding a specific value.
2026-06-21 00:06:36 +02:00
can1357 64e44fd30f fix(coding-agent): preprompt / triggerContextTokens 2026-06-20 23:21:57 +02:00
DarkPhilosophyandcan1357 8c9ae9ef67 fix(compaction): exclude encrypted reasoning from the compaction floor
CI caught that flooring by the raw local estimate falsely triggers compaction on
thinking-heavy turns: estimateTokens counts the opaque thinkingSignature /
redactedThinking payloads (providers bill them on replay, #2275), but their local
byte size diverges wildly from what the provider actually charges — so a turn
with a large encrypted-reasoning blob but small provider usage would trip the
floor (broke agent-session-handoff 'provider-anchored usage' test).

estimateTokens now takes { excludeEncryptedReasoning } and the compaction floor
(#estimateStoredContextTokens) uses it: the floor counts only reliably-countable,
on-wire-compressible content (text, tool results, tool calls), while the provider
usage arm of compactionContextTokens still accounts for encrypted reasoning. This
keeps the encrypted-reasoning case provider-anchored while still flooring upward
when a before_provider_request hook compresses tool results.
2026-06-20 23:21:36 +02:00
DarkPhilosophyandcan1357 7e8450192c fix(compaction): floor context tokens by local estimate so payload compression can't suppress auto-compaction
A before_provider_request extension (a context-compression proxy like Headroom,
an obfuscator, or inline snapcompact) can shrink the outgoing request below the
real stored conversation. The provider then reports deflated prompt tokens, so
the auto-compaction threshold never fires and the stored history grows unbounded
until it overflows the context window and can no longer be compacted at all.

Add compactionContextTokens(provider, storedEstimate) = max(provider, estimate)
and apply it to both the pre-prompt and post-response compaction decisions,
flooring the provider-reported tokens by the agent's own estimate of the stored
conversation. Display and cost accounting still use exact provider usage; only
the compaction trigger takes the floor.
2026-06-20 23:21:36 +02:00
can1357 b5437c2cad feat(coding-agent): improved prompt caching for ephemeral side-channel requests
- Added `buildSideRequestContext` to the `Agent` class to generate prompt-cache-friendly provider contexts.
- Updated ephemeral side-channel turns to forward the full tool catalog to maintain prompt cache hit rates.
- Injected a `developer` role reminder into ephemeral turns to instruct the model to suppress tool calls.
- Implemented automatic post-processing to strip any tool calls from ephemeral turn responses.
- Exported message and dialect helper functions in `agent-loop.ts` to support context construction.
2026-06-20 23:19:10 +02:00
can1357 b88d97cd25 fix(coding-agent/session): prevented duplicate persistence of rewound tool results
- Added a set to track tool result IDs that have been rewound to ensure they are not re-appended to the session history.
- Updated message handling logic to conditionally skip persistence for tool results associated with active rewind operations.
2026-06-20 23:01:01 +02:00