Commit Graph
622 Commits
Author SHA1 Message Date
roboomp aba01289d0 fix(mnemopi): consolidate working memory on session shutdown
MnemopiSessionState.dispose() now drains pending fact extractions and
runs sleepAllSessions on every owned bank before closing handles, and
AgentSession.dispose() awaits the result. This matches the manual
`/memory enqueue` slash command which was the only caller of the
consolidation pipeline.

Without the shutdown hook, episodic_memory, gists, consolidation_log,
graph_edges, and triples stayed empty for every deployment that never
typed `/memory enqueue|rebuild` — working memory accumulated forever
and long-term recall never formed.

Factored the consolidate step out of `mnemopiBackend.enqueue` into a
shared MnemopiSessionState#consolidate so the slash command and the
shutdown path share one implementation. Added two regression tests
verifying owned-bank consolidation order (flush -> sleep -> close, per
bank) and that aliased subagent dispose stays a no-op against the
parent.

Fixes #2320
2026-06-11 18:24:11 +00:00
can1357 926dfdea39 fix(coding-agent): restored side-channel IRC auto-reply for busy recipients when async execution is disabled
A subagent's send await:true to Main during a blocking task spawn was a structural deadlock: deliverIrcMessage queues mid-turn messages as step-boundary asides, but Main's next boundary requires the sender's own batch to finish, so the sender always burned the full irc.timeoutMs. Awaited sends now pass expectsReply through IrcBus.send; a mid-turn recipient with async.enabled off generates an ephemeral no-tools reply via runEphemeralTurn, records an irc:autoreply aside in its own history, and delivers it back over the bus with replyTo threading so the sender's waiter resolves.
2026-06-11 18:12:22 +02:00
can1357 965c5d4000 Merge PR #2262: feat(rpc): expose slash command metadata 2026-06-11 17:59:40 +02:00
can1357 9dbc55bcb9 Merge PR #2283: feat(prompting): add magic keyword toggles 2026-06-11 17:58:45 +02:00
roboomp dba543f226 fix(retry): capped classifier fallback attempts
Kept classifier refusals eligible for model fallback, but restored the retry.maxRetries guard so fallback chains cannot consume extra provider calls after the turn budget is exhausted.

Fixes #2290
2026-06-11 05:44:20 +00:00
roboomp 37d9eaabe5 fix(retry): fell back on anthropic refusals
Preserved Anthropic stop_details on assistant messages so the agent can distinguish classifier refusals from transport failures.

Taught AgentSession to use configured retry fallback chains for refusal and sensitive stops without same-model retries, then pin the fallback for the conversation.

Fixes #2290
2026-06-11 05:37:53 +00:00
danzaio 7450acd4f0 feat(prompting): add magic keyword toggles 2026-06-10 19:43:18 -03:00
roboomp dfc3ef52d5 style: bun run fix 2026-06-10 21:19:08 +00:00
roboomp 7fa0eeb43d fix(agent): break shake auto-continue loop on token-metric divergence
The shake-strategy post-shake threshold check was reading
#estimatePendingPromptTokens([]) while #checkCompaction triggered on
calculateContextTokens(assistantMessage.usage). The local estimator
ignored block.thinkingSignature payloads (OpenAI Responses encrypted
reasoning items, Anthropic signed thinking blocks, etc.), so on a
thinking-heavy session the estimate sat ~0.9–2× below provider-reported
usage. Once the two straddled the threshold, the #2119 dead-loop guard
never fired, shake reported 'handled', and #scheduleAutoContinuePrompt
re-injected the auto-continue developer prompt every turn — 53 injections
in a real 25-minute repro session before an external timeout.

Thread the trigger's provider-anchored contextTokens through
#runAutoCompaction → #runAutoShake for the threshold and incomplete
paths, then evaluate residual pressure as triggerContextTokens −
result.tokensFreed with an 80% recovery-band hysteresis. Re-checking
against the raw threshold (even on the corrected metric) would still let
shake reclaim a trickle of the previous turn's elidable blocks and land
just under the line every turn; the band closes that oscillation.

As defense in depth, estimateTokens() now charges thinkingSignature and
redactedThinking.data alongside the visible thinking text so every
other site that uses the estimator (idle compaction, pre-prompt check,
status line) tracks provider usage on replay.

New regression test pins the contract; existing dispatch test bumped
its mocked tokensFreed so its happy-path scenario lands inside the new
recovery band.

Fixes #2275
2026-06-10 21:18:46 +00:00
can1357 08a941a14e feat: added standalone snapcompact package and model-specific frame shaping
- Added a new @oh-my-pi/snapcompact package and redirected compaction call sites to it.
- Added provider-aware snapcompact shape resolution for model-specific mixed-frame behavior.
- Added optional image detail support by extending ImageContent and passing hints through OpenAI providers.
- Added native snapcompact render options, including 5x8/8x8 font loading and palette/geometry controls.
2026-06-10 21:50:03 +02:00
danzaio 32d1142c47 fix(rpc): complete available command lifecycle 2026-06-10 15:33:46 -03:00
danzaio 1f2fcff15f feat(rpc): expose slash command metadata 2026-06-10 14:10:02 -03:00
can1357 3aa1cedd73 feat(coding-agent): enforced an inline byte cap at the bash and browser tool boundary
Adds enforceInlineByteCap() in streaming-output and applies it to bash and browser tool results: oversized outputs are elided head/tail with an artifact:// footer pointing at the full capture, closing paths that previously let 100KB+ inline results past the minimizer. Defense at the tool-result boundary (no-op for already-bounded output).
2026-06-10 17:53:26 +02:00
can1357 9f62c7904a feat(coding-agent): integrated snapcompact strategy and per-turn supersede pruning
Adds compaction.strategy: "snapcompact" to the schema and the AgentSession routing: when chosen, both manual /compact (without custom instructions) and auto compaction call snapcompactCompact() to archive history as PNG frames instead of an LLM summary. Falls back to context-full with a visible warning notice when the current model is text-only or when /compact gets custom instructions. CustomTool and shared-event payloads carry the new action through. \n\nAlso wires the per-turn supersede pass: #pruneSupersededReads() runs every turn before threshold gating (cache-aware: only fires when the post-candidate suffix is small or the prompt cache is cold), prunes older read results superseded by a newer read of the same file, rewrites the session, and accounts the saved tokens in the next compaction decision. Gated by compaction.supersedeReads (default on).\n\nsession/messages.ts now delegates the core role conversion to agent-core's convertMessageToLlm so snapcompact image blocks flow through the LLM-context conversion path.
2026-06-10 17:52:49 +02:00
can1357 a64ff00cb8 feat(coding-agent): added history:// internal URL for agent transcripts
Registers a HistoryProtocolHandler with the internal URL router: history:// lists every registered agent (id, status, kind, last activity) and history://<agentId> renders a concise markdown transcript (tool calls collapsed to one line each, thinking elided). Live refs render from the in-memory message array; parked refs load read-only from the JSONL session file via loadSessionMessagesReadOnly (no writer, no lock). System prompt + read tool prompt + docs/tools/read.md learn the new scheme.
2026-06-10 17:52:13 +02:00
can1357 2b28816010 feat(coding-agent): replaced session observer with Agent Hub and inline compaction divider
The session-observer overlay is gone. The Agent Hub (ctrl+s, alt+a, or double-tap left arrow on an empty editor) presents one overlay with two views: a live registry table (status, unread irc count, current task, last activity; j/k to navigate, r to revive, x to abort/release) and per-agent chat (transcript + input line) — submitting revives a parked agent and steers it via the normal prompt path. Main is the ambient chat and stays out of the table.\n\nrenderInitialMessages no longer takes a prebuilt context: every redraw now reaches for AgentSession.buildTranscriptSessionContext() (full-history transcript with each compaction emitted inline at the point it fired, snapcompact frames re-attached on rebuild). UiHelpers drops the deferred-compaction render and the IRC autoreply branch that the new mailbox bus deprecated. CompactionSummaryMessage renders as a slim divider (── 📷 compacted · ctrl+o ──), expanding to the summary + snapcompact frame count. session-manager.buildSessionContext gains a { transcript: true } mode; an exported buildSessionContextFromFile() reads a session file without taking the writer lock so the hub chat view can tail any agent (parked or live). Theme picks up icon.camera + tool.irc symbol entries.
2026-06-10 17:51:42 +02:00
can1357 a92d2ce989 feat(coding-agent): removed context argument from eval agent() spawn
Shared background now flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each prompt instead of a context string forwarded into the subagent's system prompt. The JS and Python preludes drop the context kwarg from agent(), the subagent system prompt drops the {{#if context}} block and the conversation-context file pointer, and runEvalAgent no longer writes a per-call conversation context file. AgentSession sheds the now-unused formatCompactContext() helper that supplied the file's body, and ToolSession.getCompactContext is removed alongside it.
2026-06-10 17:49:01 +02:00
can1357 fcb8663de8 feat(coding-agent): reworked irc to a send/wait/inbox/list mailbox bus
Replaces the blocking auto-reply IRC turn with a process-global IrcBus and a four-op tool (send/wait/inbox/list). send is fire-and-forget with per-recipient delivery receipts (injected/woken/revived/failed); replies become real turns by the recipient, observed via wait (or the send await:true sugar). Bounded per-agent mailboxes (cap 100) drop oldest on overflow; AgentSession.deliverIrcMessage folds an in-flight delivery in as a non-interrupting aside at the next step boundary, or starts a real wake turn when idle. AgentRegistry/lifecycle handle the idle→woken / parked→revived transitions, so messaging a non-running peer brings it back. Adds a dedicated TUI renderer (directional headers, delivery-outcome coloring, quoted bodies, per-recipient receipt trees, status-badged peer lists with unread counts). irc.timeoutMs is now the default timeout for wait / send await:true. AgentSession sheds the agentRegistry config field, the dedupeIrcReply → dedupeEphemeralReply rename (now used by /btw and /omfg), and the background-channel exchange queue / forwardIrcRelayToMain plumbing the auto-reply model needed.
2026-06-10 17:47:47 +02:00
can1357 316722397a Merge pull request #1896: fix(coding-agent): respect user shell for ! shortcuts 2026-06-10 09:52:28 +02:00
can1357 c08e6f8a19 ux(coding-agent): default artifact captures back to unbounded with opt-in capping 2026-06-10 09:51:47 +02:00
can1357 1821e167f4 fix(coding-agent): deobfuscate secrets in ephemeral turns and obfuscate via provider-context hook 2026-06-10 09:51:47 +02:00
can1357 730e9c8e0c fix(acp): honor the explicit autoApprove session flag when skipping the permission gate 2026-06-10 09:51:47 +02:00
roboomp cafd957f5d fix(providers): disabled ollama thinking for off turns
Propagated explicit thinking-off state through the agent loop so provider requests receive disableReasoning instead of an undefined effort. Added Ollama and agent-session regressions for the :off path.\n\nFixes #2239
2026-06-10 07:41:58 +00:00
handlecusion 2d7d717029 fix(coding-agent): tighten user shell routing 2026-06-10 16:07:59 +09:00
handlecusion 8b9c4fa1b9 fix(coding-agent): respect user shell for shortcuts 2026-06-10 16:07:59 +09:00
can1357 11c53051bf Merge pull request #2097: fix(acp): skip permission gate when yolo mode is explicitly enabled 2026-06-10 08:32:09 +02:00
can1357 d30dca302f Merge pull request #2044: fix(coding-agent): apply enabledModels filter to ACP model list 2026-06-10 08:32:09 +02:00
can1357 3fd09d5770 fix(acp): require explicit yolo opt-in before skipping the client permission gate
The schema default for tools.approvalMode is already "yolo", so checking the
resolved setting alone disabled the ACP permission gate for every
default-config session (17 existing tests in
agent-session-acp-permission.test.ts fail). The skip now requires an
explicitly configured approval mode — the --yolo/--auto-approve runtime
override or a user-set tools.approvalMode — via the new
Settings.isConfigured(), keeping default ACP sessions gated.

Addresses review feedback on #2097.
2026-06-10 08:31:31 +02:00
Theo Mathieuandcan1357 fbb48faf81 fix: reflect --auto-approve flag in settings override so #wrapToolForAcpPermission sees yolo mode 2026-06-10 08:31:30 +02:00
Theo Mathieuandcan1357 c110a624c3 fix(acp): skip permission gate in yolo mode when effective policy is allow 2026-06-10 08:31:30 +02:00
Theoandcan1357 9e28ec4b7e fix: apply enabledModels filter to ACP model list
getAvailableModels() was calling modelRegistry.getAvailable() directly,
which skips the enabledModels setting. The setting was only applied
during session init to pick the starting model, not to the list
advertised to ACP clients (Zed, etc.).

Add filterAvailableModelsByEnabledPatterns() to model-resolver.ts - a
synchronous subset of resolveAllowedModels() that handles the patterns
used in real configs (exact provider/modelId, canonical ids, bare model
ids, thinking-level suffixes). Glob patterns fall back to showing all
models rather than accidentally emptying the picker.

Update getAvailableModels() to call it, so the ACP model dropdown in
Zed (and any other ACP client) respects the users enabledModels config.
2026-06-10 08:31:30 +02:00
can1357 7ce58fe95c Merge pull request #2204: feat(eval): add python interpreter setting 2026-06-10 08:26:59 +02:00
can1357 62ad6e73ae Merge pull request #2200: fix(ai): rotate antigravity credentials on Individual quota reached 429s 2026-06-10 08:26:59 +02:00
can1357 e19f2f689c Merge pull request #2147: fix(coding-agent): hide secrets in provider requests 2026-06-10 08:26:58 +02:00
can1357 c03075c9ae Merge pull request #2083: fix(coding-agent): broke runaway edit loops and capped bash artifact spew 2026-06-10 08:26:57 +02:00
416c9947bd fix(acp): finish prompt turn when extension/custom command is handled locally
Extension commands (e.g. /sonnet) and TypeScript custom commands that
consume the input without calling the LLM return early from
session.prompt() with no agent turn. In ACP mode this left the pending
prompt promise unresolved, hanging the client forever.

Change session.prompt() to return Promise<boolean>: true when the LLM
was invoked, false when the command was fully handled locally.
#runPromptOrCommand calls #finishPrompt immediately on a false return so
the ACP turn completes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 08:26:07 +02:00
can1357 ec81b927fa fix(coding-agent): redact ephemeral side-channel context and preserved remote compaction history
runEphemeralTurn (IRC//btw) called streamSimple directly with the raw
system prompt, bypassing the SDK-level obfuscateProviderContext wrapper;
and #obfuscatePreparationForProvider skipped previousPreserveData, so a
pre-fix openaiRemoteCompaction.replacementHistory could resend raw
secrets on the next remote compaction.

Addresses review feedback on #2147.
2026-06-10 08:26:02 +02:00
roboompandcan1357 d39e4b0d38 fix(coding-agent): hid previous compaction summary
Obfuscated preparation.previousSummary, hook prompt/context, before forwarding to compact() so prior pi- or extension-supplied summaries do not leak verbatim secrets on subsequent compactions.

Fixes #2146
2026-06-10 08:26:02 +02:00
roboompandcan1357 5517115f1a fix(coding-agent): hid handoff instructions
Obfuscated custom instructions before handoff and related side-request provider calls, then deobfuscated generated handoff output before persistence.

Fixes #2146
2026-06-10 08:26:02 +02:00
roboompandcan1357 aa4cd0ab2d fix(coding-agent): preserved tool schemas while redacting
Converted provider-facing tool parameters to wire JSON Schema before redaction so live Zod instances are not deep-cloned into plain objects.

Fixes #2146
2026-06-10 08:26:02 +02:00
roboompandcan1357 2a90108893 fix(coding-agent): hid secrets in provider requests
Redacted configured secrets across provider-facing system prompts, tool definitions, developer reminders, and assistant tool-call payloads before LLM requests.

Fixes #2146
2026-06-10 08:26:01 +02:00
can1357 24cbc93913 fix(eval): thread python.interpreter from session settings and expand ~
Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.

Addresses review feedback on #2204.
2026-06-10 08:26:00 +02:00
roboompandcan1357 8c3149e5a9 fix(ai): scoped antigravity quota blocks by model family
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.

Fixes #2198
2026-06-10 08:26:00 +02:00
can1357 dbf1be2096 fix(coding-agent): count head-retained bytes against the artifact cap
After rebasing onto a13e9827f, #createFileSink's head-retention flush
wrote directly to the sink, bypassing #emitToSink's budget accounting —
the on-disk artifact could grow past artifactMaxBytes by up to the head
window. Route the flush through #emitToSink and pin it with a
regression test (fails 24B vs 16B cap without the fix).

Addresses review feedback on #2083.
2026-06-10 08:26:00 +02:00
roboompandcan1357 e18a8649f0 fix(coding-agent): suppressed artifact truncation notice when nothing was elided
Codex review on PR #2083 caught a corruption case in `OutputSink`:
when a stream exceeds `artifactHeadBytes` but still fits below
`artifactMaxBytes`, post-head bytes flow into the tail ring without any
eviction (`droppedBytes === 0`) — yet `#flushArtifactTailIfCapped`
unconditionally injected `[ARTIFACT TRUNCATED: kept first … + last … of
…; 0 B elided from the middle]` between head and tail. The resulting
artifact-on-disk is then no longer verbatim and falsely advertises
truncation for outputs that actually fit. With the default 4 MiB / 3 MiB
split, every ~3–4 MiB bash capture in this band tripped the bug.

The notice is now gated on `droppedBytes > 0`. The tail ring is still
flushed unconditionally so head + tail still equal the verbatim stream
in this band. Regression pinned by a new test in
`test/streaming-output.test.ts` that pushes 24 bytes into a 16-head /
16-tail cap and asserts the file equals the payload byte-for-byte with
no `[ARTIFACT TRUNCATED:` marker.

Refs #2081
2026-06-10 08:26:00 +02:00
roboompandcan1357 77a68b1070 fix(coding-agent): broke runaway edit loops and capped bash artifact spew
Two pathologies surfaced in the same captured failure (#2081): a subagent
spent 16 minutes hammering 205 `edit` calls (182 byte-identical no-ops)
against a file that already matched its payload, while a sibling bash
invocation persisted 7.6MB of PowerShell rich-object metadata to
`~/.omp/agent/artifacts/<id>.bash.log` from what was intended as a small
tail. Both are addressed independently here:

- Hashline executor now consults a per-ToolSession `noopLoopGuard` that
  hashes the raw patch input and tracks consecutive no-ops per canonical
  path. After NOOP_HARD_LIMIT (3) repeats of the same payload the soft
  "byte-identical" hint escalates to a thrown ToolError, which the agent
  loop surfaces as a tool failure rather than success-with-text — far
  more effective at breaking the loop than the soft hint alone. A
  non-noop commit (or any variant payload) resets the counter; state is
  isolated per ToolSession so subagents cannot inherit each other's
  history.
- OutputSink artifact-on-disk writes are now bounded by
  `artifactMaxBytes` (default 4 MiB = 3 MiB head + 1 MiB rolling tail).
  Once the head budget is exhausted, subsequent chunks divert into a
  fixed-size tail ring; `dump()` replays the ring behind a single
  `[ARTIFACT TRUNCATED: kept first … + last … of …; … elided from the
  middle]` notice before closing the sink. Setting `artifactMaxBytes: 0`
  restores the historical unbounded behavior. Sized comfortably above
  anything a model would reasonably scroll through via the artifact URL
  scheme while preventing the captured 7.6MB spray from sitting on disk.

The terminal-typing lag the reporter observed has multiple compounding
causes (transcript-render freezing is disabled on win32; the bash result
renderer lacks the per-render cache that the eval renderer already has).
Those land in a follow-up — the loop guard + artifact cap address the
root pathologies that turned the session into a multi-MB transcript in
the first place.

Fixes #2081
2026-06-10 08:26:00 +02:00
can1357 661587e110 feat(coding-agent): raised retry limits to ten with capped exponential backoff
- Raised Anthropic provider retries to 10 attempts and used a shared jittered exponential backoff for each retry.
- Updated coding-agent retry defaults and session delay calculation to a 500ms base with an 8,000ms jittered cap.
- Added tests that verify capped ten-step backoff sequences and recovery after repeated 502 errors.
2026-06-10 07:49:12 +02:00
can1357 b0fec42218 feat(coding-agent): surfaced lazy LSP servers as available in welcome screen
- Added "available" status so recognized servers show under lazy mode without warmup.
- Rendered full welcome box as pre-TUI splash with fixed slot heights to avoid layout shift.
- Reported lazy servers as available in /status instead of omitting the section.
2026-06-10 07:22:22 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 7124c74067 fix(ai): waited for sibling unblock over provider retry window
- Returned earliest sibling block expiry from markUsageLimitReached.
- Capped usage-limit retry to whichever frees up first, avoiding multi-hour waits.
- Added 1s buffer so retry lands after the block actually lapses.
2026-06-10 02:44:40 +02:00