Commit Graph

5588 Commits

Author SHA1 Message Date
roboomp eef2f7b7e9 fix(compaction): resolve retry before goal continuation return
Active-goal threshold compaction can pre-empt the normal post-turn tail
and return once it schedules a deferred handoff or auto-continue. When
that turn is the successful response from an auto-retry, returning there
skips the later retry-gate cleanup and leaves isRetrying stuck.

Resolve the completed retry gate before the compaction-continuation
return, and cover the retry-success-over-threshold path so future
changes cannot strand prompt()/waitForIdle() behind a stale retry state.

Refs #3174
2026-06-21 09:08:59 +00:00
roboomp e70c71077e fix(compaction): log agent_end maintenance routing for goal turns
Reporter on #3174 still sees no auto-compaction with thresholdTokens
lowered to 32768 against a 70k+ visible context, and reports the
existing `Auto-compaction threshold decision` log never appears.
Either the goal turn never reaches `#checkCompaction`, or the log
fires but is filtered out of their view (winston is at debug level so
it writes to ~/.omp/logs/omp.<DATE>.log, not the TUI).

Add an `agent_end maintenance routing` debug log at every branch of
the `agent_end` handler — entered/no-message,
skip-post-turn-maintenance, successful-yield (active goal vs not),
empty-stop-handled, active-goal pre-empt (and whether it scheduled a
continuation), unexpected-stop-handled, and bottom checkCompaction —
together with stopReason, provider/model, content shape, goal
state, and `successfulYield`. Combined with the existing
`Auto-compaction threshold decision` log, the next no-start report
identifies the exact early-return branch and the inputs that fed
`shouldCompact`.

Refs #3174
2026-06-21 08:53:54 +00:00
roboomp 00d14accfb fix(compaction): keep empty-stop cleanup before active-goal compaction continuation
Codex review on #3175: the active-goal compaction pre-empt I added in
8ab754f636 short-circuited #handleEmptyAssistantStop. That handler is
the only path that strips an orphan toolUse assistant (stopReason
"toolUse" with no toolCall block) from both active context and the
session branch via #removeEmptyStopFromActiveContext. With the pre-empt
ordering, an over-threshold goal turn that returned an empty toolUse
left the orphan as the session leaf, and the compaction auto-continue
prompt fed it back into the next Anthropic turn as a tool_use with no
matching tool_result — the exact history-corruption pattern the
existing cleanup comment defends against.

Move #handleEmptyAssistantStop back ahead of the active-goal
compaction probe. Empty stops still self-retry and never reach the
threshold pre-empt; non-empty stops (the reporter's failure in #3174)
still hit threshold maintenance before the unexpected-stop classifier.

Regression test seeds a goal-mode empty toolUse stop billed at 91k
against thresholdTokens 76384 and asserts the threshold compaction
never starts and the orphan is no longer in the session branch.

Fixes #3174
2026-06-21 08:49:28 +00:00
roboomp 8ab754f636 fix(compaction): run goal threshold maintenance before retry continuations
Active goal turns that stopped with text could hit the empty/unexpected-stop
continuation guards before threshold maintenance. When those guards scheduled
another goal turn, #checkCompaction never ran, so no auto_compaction_start was
emitted even while visible context stayed above thresholdTokens.

Run threshold maintenance once before those active-goal self-continuations and
log the threshold decision inputs: billed context, stored estimate, resolved
trigger tokens, post-maintenance tokens, strategy, threshold, promotion state,
and shouldCompact.

Also pass post-prune maintenance tokens into the shake recovery-band check so
supersede/drop-useless savings are preserved when deciding whether shake still
needs to fall back to context-full compaction.

Fixes #3174
2026-06-21 08:35:58 +00:00
roboomp b74f7bb720 fix(compaction): trigger goal-mode threshold on billed context, not post-prune estimate
Pruning frees bytes for the NEXT prompt — it does not change the size of
the prompt the LLM just billed for. Subtracting the per-turn
`#pruneStaleToolResults` / `#pruneToolOutputs` savings from the
threshold input let a long-running `/goal` session sit above
`compaction.thresholdTokens` indefinitely: the visible context
(anchored to the same provider billing) showed >threshold, but
`shouldCompact` no-op'd because the subtraction dropped the input below
the trigger. The `compactionContextTokens` floor against the post-prune
local estimate is still applied, so a payload-compression hook still
can't deflate the trigger.

Regression test seeds one large `useless` tool result whose suffix sits
inside the 8k cache-warm window so `#pruneStaleToolResults` actually
returns ≥20k savings, then asserts compaction fires when the final turn
bills 91k tokens against the reporter's `thresholdTokens: 76384`.

Fixes #3174
2026-06-21 07:21:22 +00:00
can1357 5207a98fd8 refactor(coding-agent): simplified temporary model status message
- Inlined the temporary model status formatting logic directly into the controller.
- Removed the unused `formatTemporaryModelStatus` utility function and its associated test.
2026-06-21 07:42:46 +02:00
can1357 c5cf47aa7f fix(coding-agent/utils): handled EISDIR and ENOTDIR errors during ref resolution
- Update `shouldRetry` to treat `EISDIR` and `ENOTDIR` as terminal errors, preventing unnecessary retries when encountering Git reference directory conflicts.
- Add a test suite to verify graceful resolution of branches in scenarios where a packed ref conflicts with a directory path in the filesystem.
2026-06-21 07:36:25 +02:00
can1357 07c910210b feat(coding-agent): allowed toggling workspace tree in system prompt
- Added `includeWorkspaceTree` configuration setting to optionally render the workspace directory tree.
- Configured the system prompt template and SDK to respect this toggle, allowing users to disable the tree to prevent prompt cache invalidation.
2026-06-21 07:03:45 +02:00
can1357 2eef88978b feat(coding-agent): made write and find tools essential
- Promoted `write` and `find` tools to `essential` status to ensure they are always available regardless of discovery mode.
- Updated `DEFAULT_ESSENTIAL_TOOL_NAMES` to include these tools by default.
- Updated documentation and tests to reflect the change in default essential tool availability.

Fixes #3165
2026-06-21 07:02:44 +02:00
can1357 984c8dd2f6 fix(coding-agent): reconciled tool arguments on execution start
- Synchronize tool arguments with the component state upon receipt of `tool_execution_start` to ensure visual consistency when final update events are missed.
- Terminate active argument reveal streams to prevent late ticks from overwriting valid, fully-materialized tool arguments with stale partial data.
- Add test coverage to verify that tool UI components render finalized arguments even in the absence of intermediate streaming updates.
2026-06-21 06:51:37 +02:00
can1357 e649017322 feat(coding-agent): added ARIA snapshot support for browser tools
- Implemented `tab.ariaSnapshot()` to capture and represent page structures as ARIA-tree YAML.
- Introduced `tab.ref()` and ref-based selector parsing to enable precise element interaction via unique ARIA identifiers.
- Integrated automated script bundling for cross-environment evaluation of ARIA snapshot logic.
- Updated browser action methods to resolve and target elements using ARIA-ref handles.
2026-06-21 06:39:14 +02:00
can1357 57cb7447e0 feat(coding-agent): enhanced browser stealth and automation capabilities
- Overhauled stealth spoofing mechanisms for WebGL, screen dimensions, Web workers, and iframe contexts using prototype-aware injection.
- Centralized function string representation patching to improve mimicry of native browser behavior across global objects.
- Patched puppeteer-core to remove detectable evaluation markers and implement lazy, pull-style execution context management.
- Enabled support for capturing LLM request JSON dumps and adjusted launcher flags to improve organic request patterns.
2026-06-21 04:14:26 +02:00
can1357 4d96bcf6b6 feat: replaced manual debug request path with automated llm dump
- Replaced the `/debug dump-next-request` command with an updated `/dump` command that exports LLM request context to JSON sidecar files.
- Removed persistent debug path state and manual path configuration in favor of automated generation.
- Updated session logic to handle serializing LLM request context to temporary directories.
- Refactored testing suites to remove path-based debug tests and verify dynamic request file generation.
2026-06-21 03:18:20 +02:00
can1357 320f514f77 fix(ai): adjusted llama-cpp base url and refactor token validation
- Update `llama.cpp` base URL to remove the `/v1` suffix.
- Consolidate local provider token validation in `coding-agent` using a `Set`.
- Add test coverage to ensure catalog model IDs are preserved verbatim on the wire.
2026-06-21 02:45:48 +02:00
oldschoola 13bfc7b9c0 Remove Wafer Pass provider 2026-06-21 02:45:48 +02:00
can1357 e31bfa4b96 Merge remote-tracking branch 'origin/farm/7a22e7a4/mcp-toggle-reconnects-others' 2026-06-21 02:18:29 +02:00
can1357 0f976084b8 fix(coding-agent/cli): route leading global flags before a subcommand to that subcommand
`omp --approval-mode=yolo acp` was rewritten to `launch --approval-mode=yolo
acp`, swallowing `acp` as a launch prompt so the yolo override never reached the
ACP command path (the ACP permission gate from #2097 stayed in always-ask).

`resolveCliArgv` only inspected `argv[0]`, so any leading global option flag hid
the real subcommand. It now scans past leading flags using the launch parser's
value-consumption contract (a flag's value is never mistaken for the subcommand,
e.g. `--model acp`) and hoists the recognized subcommand to the front with the
flags preserved as its own argv. Genuine launch prompts are untouched.

The flag value-consumption rule is factored into `cli/flag-tables.ts`
(`flagConsumesValue` + the shared `isUnknownLongValueCandidate`) so the resolver
and the profile bootstrap share one source of truth.

Fixed permission mode not respected in ACP mode ([#2970](https://github.com/can1357/oh-my-pi/issues/2970))

Fixes #2970
2026-06-21 02:17:54 +02:00
can1357 725aa2222b fix(coding-agent/lsp): drain config pulls on lazy cold start
The background message reader matched every incoming message against the
pending client-request map by id before checking for a `method`. Server
request ids live in the server's own id space and routinely collide with
the client's in-flight request ids, so a server-originated
`workspace/configuration` pull whose id matched a pending request (e.g. a
basedpyright pull landing while a `documentSymbol` request with the same
id was open) was swallowed as a bogus response: the client request
resolved with `undefined` and the pull was never answered, wedging
servers that gate analysis on configuration.

Route any message carrying a `method` as a server request (or
notification) before id-matching responses, so every config pull is
answered under lazy init -- parity with the warmup/reload path, which
escaped the bug only because it issues no concurrent semantic request
while the cold-start pulls drain. lsp.lazy default is unchanged.

Fixes #3001
2026-06-21 02:16:26 +02:00
can1357 b0c6ad825c fix(coding-agent): hint at omp plugin list/uninstall instead of leaking bare plugin verbs to launch
The plugins docs advertise `omp list`/`omp remove` as top-level commands,
but only `omp install` is registered. `resolveCliArgv` rewrote any
unregistered first-arg to `launch`, so `omp list` silently started an
interactive agent session with "list" as the LLM prompt instead of
managing plugins.

Reserve the bare, documented-but-unregistered plugin verbs `list` and
`remove` with a helpful hint pointing at the real `omp plugin list` /
`omp plugin uninstall <name>` commands (same `extensions` mechanism /
same class as the `install` leak fixed in #1496/#1498). Multi-word
prompts that merely begin with those words still route to `launch`, so
genuine prompts are unaffected.

Fixes #2935
2026-06-21 02:16:17 +02:00
roboomp c18d400bdf fix(cli): scoped mcp toggles to one server
Updated /mcp enable and /mcp disable so they connect or disconnect only the named server instead of reloading every MCP server in the session. Added regression coverage for both toggle directions and updated the coding-agent changelog.

Fixes #3157
2026-06-20 23:58:03 +00:00
can1357 29f7b57fde feat: migrated core rendering and context transformation to async
- Updated `render`, `renderMany`, and native snapcompact methods to return promises, ensuring scalable async execution.
- Refactored `transformProviderContext` and `buildSideRequestContext` to support asynchronous operations in agent loops.
- Integrated `Promise.all` for improved concurrency when processing frame rendering and rendering batch operations.
- Updated all internal call sites, SDK hooks, and test suites to accommodate the asynchronous API signatures.
2026-06-21 00:32:13 +02:00
can1357 51f5919486 feat(coding-agent): improved advisor communication constraints
- Updated the system prompt to explicitly instruct the advisor never to send the same advice twice.
2026-06-21 00:14:06 +02:00
can1357 829d6c66dc Merge PR #2903: fix(mnemopi): make proactive linking configurable from host settings (@wolfiesch) 2026-06-21 00:09:20 +02:00
can1357 aa9709c47f feat(coding-agent): added snapcompact safety checks and UI integration
- Added validation to scan for non-ASCII characters before performing snap-compaction, falling back to LLM-based summarization if the unrenderable ratio is too high.
- Updated event handling and status reporting to explicitly support snapcompact actions, including specific error warnings and cancellation states in the UI.
- Updated session logic to default to snapcompact strategy when auto-compaction is enabled.
2026-06-21 00:08:12 +02:00
can1357 3c356fc0e0 fix(coding-agent/session): improved auto-compaction strategy selection
- Adjusted the fallback strategy when enabling auto-compaction to respect the system default instead of hardcoding a specific value.
2026-06-21 00:06:36 +02:00
can1357 850d55e2ea config(coding-agent/config): updated default agent settings
- Increase default immune turns for advisor concerns from 1 to 3.
- Update default compaction strategy from context-full to snapcompact.
2026-06-21 00:05:43 +02:00
can1357 aa87848f1a fix(coding-agent): prevented accidental cancellation of background maintenance
- Stop the Esc key from aborting active background maintenance (compaction, handoff, or retry) while a subagent is focused.
- Remove "(esc to cancel)" hints from maintenance loaders when a subagent is active to avoid false affordance.
- Ensure that main-session maintenance remains cancellable via Esc when no subagent is focused.

Fixes #2819
2026-06-21 00:00:35 +02:00
can1357 9f34c45d94 fix(coding-agent/modes): restricted branch key input to focused editor
- Added a check to ensure the editor is focused before allowing the branch keyboard shortcut.
2026-06-20 23:57:29 +02:00
can1357 aaac4e414f Merge PR #2950: feat(coding-agent): add copy affordance to completed /btw answers (@wolfiesch)
Adds 'c copy' to the completed /btw panel footer (alongside b branch / Esc
dismiss), copying the sanitized visible answer to the clipboard. The copy
shortcut is guarded by canCopyBtw + main-editor focus + empty editor.
2026-06-20 23:56:24 +02:00
roboomp 75b21611f7 fix(mcp): sanitized startup server names
Sanitized MCP server names before status formatting so configured keys cannot leak home paths, tabs, newlines, or oversized text into the TUI.

Fixes #3150
2026-06-20 23:43:16 +02:00
can1357 63655a5b71 Merge PR #3050: feat: show live slash command status in autocomplete (@wolfiesch)
Shows live per-command status (plan/goal/loop/model/advisor/collab/jobs/
context/...) in the slash-command autocomplete.

Adjustments on merge:
- Exclude the two .github/pr-assets/*.png screenshots.
- Resolve the autocomplete.ts conflict keeping main's skill-command empty-prefix
  boost alongside the new static-vs-display description split.
- Compute the live display description lazily — only once a command actually
  matches (name or alias) — instead of for every command on each keystroke, and
  guard the getter with a typeof check. getAutocompleteDescription reads live
  session state, so the eager call was O(commands) work per refresh.
2026-06-20 23:36:17 +02:00
can1357 a2f8c97e10 Merge remote-tracking branch 'origin/farm/96f37e17/mcp-connection-status' 2026-06-20 23:29:19 +02:00
roboomp 8b2012420f fix(mcp): sanitized startup failure status
Sanitized MCP failure text before it reaches the startup status formatter, including tabs, newlines, home paths, and long server output.

Fixes #3150
2026-06-20 21:23:27 +00:00
can1357 64e44fd30f fix(coding-agent): preprompt / triggerContextTokens 2026-06-20 23:21:57 +02:00
DarkPhilosophy 8c9ae9ef67 fix(compaction): exclude encrypted reasoning from the compaction floor
CI caught that flooring by the raw local estimate falsely triggers compaction on
thinking-heavy turns: estimateTokens counts the opaque thinkingSignature /
redactedThinking payloads (providers bill them on replay, #2275), but their local
byte size diverges wildly from what the provider actually charges — so a turn
with a large encrypted-reasoning blob but small provider usage would trip the
floor (broke agent-session-handoff 'provider-anchored usage' test).

estimateTokens now takes { excludeEncryptedReasoning } and the compaction floor
(#estimateStoredContextTokens) uses it: the floor counts only reliably-countable,
on-wire-compressible content (text, tool results, tool calls), while the provider
usage arm of compactionContextTokens still accounts for encrypted reasoning. This
keeps the encrypted-reasoning case provider-anchored while still flooring upward
when a before_provider_request hook compresses tool results.
2026-06-20 23:21:36 +02:00
DarkPhilosophy 7e8450192c fix(compaction): floor context tokens by local estimate so payload compression can't suppress auto-compaction
A before_provider_request extension (a context-compression proxy like Headroom,
an obfuscator, or inline snapcompact) can shrink the outgoing request below the
real stored conversation. The provider then reports deflated prompt tokens, so
the auto-compaction threshold never fires and the stored history grows unbounded
until it overflows the context window and can no longer be compacted at all.

Add compactionContextTokens(provider, storedEstimate) = max(provider, estimate)
and apply it to both the pre-prompt and post-response compaction decisions,
flooring the provider-reported tokens by the agent's own estimate of the stored
conversation. Display and cost accounting still use exact provider usage; only
the compaction trigger takes the floor.
2026-06-20 23:21:36 +02:00
can1357 b5437c2cad feat(coding-agent): improved prompt caching for ephemeral side-channel requests
- Added `buildSideRequestContext` to the `Agent` class to generate prompt-cache-friendly provider contexts.
- Updated ephemeral side-channel turns to forward the full tool catalog to maintain prompt cache hit rates.
- Injected a `developer` role reminder into ephemeral turns to instruct the model to suppress tool calls.
- Implemented automatic post-processing to strip any tool calls from ephemeral turn responses.
- Exported message and dialect helper functions in `agent-loop.ts` to support context construction.
2026-06-20 23:19:10 +02:00
roboomp 655fed1e48 fix(mcp): updated startup connection status
Emitted MCP connection lifecycle events through the startup event bus so the TUI can replace the initial connecting banner with connected, pending, or failed server state.

Added manager and interactive-mode coverage for mixed success/failure MCP startup updates.

Fixes #3150
2026-06-20 21:14:41 +00:00
can1357 b88d97cd25 fix(coding-agent/session): prevented duplicate persistence of rewound tool results
- Added a set to track tool result IDs that have been rewound to ensure they are not re-appended to the session history.
- Updated message handling logic to conditionally skip persistence for tool results associated with active rewind operations.
2026-06-20 23:01:01 +02:00
can1357 81a4af6684 perf(coding-agent): deduplicated primary context messages in advisor history
- Implemented message deduplication in `AdvisorRuntime` to collapse verbatim re-injected primary context (plan rules/approved plans) between turns.
- Reduced token consumption by replacing identical primary context segments with a status marker.
- Updated history formatter to allow expansion of load-bearing context types specifically, while retaining one-line summaries for other custom messages.
2026-06-20 22:59:30 +02:00
can1357 60b043adb7 fix(coding-agent): fixed session history desynchronization during
- Fixed session history desynchronization during `rewind` by performing a full context rebuild after applying branch changes.
- Prevented `rewind` tool output and assistant side-channel data from polluting the prompt cache by flushing and sanitizing session state.
- Added explicit test coverage for context reconstruction and assistant message sanitization after rewind events.
2026-06-20 22:58:23 +02:00
can1357 66d9df6e1f ux(coding-agent): removed static pending icons from edit headers
- Remove the static "pending" hourglass icon from edit and write tool headers to reduce visual noise.
- Update multi-file status lines to use the active spinner icon directly instead of replacing a static icon, ensuring consistent liveness cues.
2026-06-20 22:46:12 +02:00
Can Bölük 47bb5172fe Merge pull request #3138 from KamijoToma/save-command-history
feat(coding-agent): save slash command text for /btw /tan /omfg /memory /rename /move to TUI history
2026-06-20 22:39:47 +02:00
can1357 321b51517a Merge PR #3022: refactor(coding-agent): use Promise.withResolvers in AsyncDrain (@oldschoola) 2026-06-20 22:36:13 +02:00
can1357 d8ec46eee5 fix(tiny): guard unsupportedReason access on const model-spec union (#3133 integration typecheck) 2026-06-20 22:25:29 +02:00
can1357 9722505a2b Merge PR #3017: fix(coding-agent): handle todo paths and prompt inventory (@oldschoola)
# Conflicts:
#	packages/coding-agent/test/system-prompt-inventory.test.ts
2026-06-20 22:19:28 +02:00
can1357 ffc842ee7c fix(agent): reset promoted memory prompt on session fork
fork() reset mnemopi conversation tracking directly but skipped the shared new-transcript reset, so the folded/promoted first-turn memory stayed in #baseSystemPrompt. The next turn re-recalled and the change-detection saw no diff, taking the fallback promotion path and injecting the <memories> block twice into the forked session prompt. Route fork() through #resetMemoryContextForNewTranscript() like the other reset paths and add a regression test asserting the forked prompt contains recalled memory exactly once.
2026-06-20 22:16:58 +02:00
can1357 10ed2d4efa Merge PR #3113: fix(agent): preserve memory recall in append-only prompt cache (@roboomp) 2026-06-20 22:16:58 +02:00
can1357 1fadcdb8e3 Merge PR #3147: fix(agent): compact active goal yields (@roboomp) 2026-06-20 22:16:57 +02:00
can1357 7bf8db9215 Merge PR #2946: fix(coding-agent): show omfg dismiss hint (@wolfiesch) 2026-06-20 22:16:57 +02:00