Commit Graph
839 Commits
Author SHA1 Message Date
can1357 4d96bcf6b6 feat: replaced manual debug request path with automated llm dump
- Replaced the `/debug dump-next-request` command with an updated `/dump` command that exports LLM request context to JSON sidecar files.
- Removed persistent debug path state and manual path configuration in favor of automated generation.
- Updated session logic to handle serializing LLM request context to temporary directories.
- Refactored testing suites to remove path-based debug tests and verify dynamic request file generation.
2026-06-21 03:18:20 +02:00
can1357 29f7b57fde feat: migrated core rendering and context transformation to async
- Updated `render`, `renderMany`, and native snapcompact methods to return promises, ensuring scalable async execution.
- Refactored `transformProviderContext` and `buildSideRequestContext` to support asynchronous operations in agent loops.
- Integrated `Promise.all` for improved concurrency when processing frame rendering and rendering batch operations.
- Updated all internal call sites, SDK hooks, and test suites to accommodate the asynchronous API signatures.
2026-06-21 00:32:13 +02:00
can1357 aa9709c47f feat(coding-agent): added snapcompact safety checks and UI integration
- Added validation to scan for non-ASCII characters before performing snap-compaction, falling back to LLM-based summarization if the unrenderable ratio is too high.
- Updated event handling and status reporting to explicitly support snapcompact actions, including specific error warnings and cancellation states in the UI.
- Updated session logic to default to snapcompact strategy when auto-compaction is enabled.
2026-06-21 00:08:12 +02:00
can1357 3c356fc0e0 fix(coding-agent/session): improved auto-compaction strategy selection
- Adjusted the fallback strategy when enabling auto-compaction to respect the system default instead of hardcoding a specific value.
2026-06-21 00:06:36 +02:00
can1357 64e44fd30f fix(coding-agent): preprompt / triggerContextTokens 2026-06-20 23:21:57 +02:00
DarkPhilosophyandcan1357 8c9ae9ef67 fix(compaction): exclude encrypted reasoning from the compaction floor
CI caught that flooring by the raw local estimate falsely triggers compaction on
thinking-heavy turns: estimateTokens counts the opaque thinkingSignature /
redactedThinking payloads (providers bill them on replay, #2275), but their local
byte size diverges wildly from what the provider actually charges — so a turn
with a large encrypted-reasoning blob but small provider usage would trip the
floor (broke agent-session-handoff 'provider-anchored usage' test).

estimateTokens now takes { excludeEncryptedReasoning } and the compaction floor
(#estimateStoredContextTokens) uses it: the floor counts only reliably-countable,
on-wire-compressible content (text, tool results, tool calls), while the provider
usage arm of compactionContextTokens still accounts for encrypted reasoning. This
keeps the encrypted-reasoning case provider-anchored while still flooring upward
when a before_provider_request hook compresses tool results.
2026-06-20 23:21:36 +02:00
DarkPhilosophyandcan1357 7e8450192c fix(compaction): floor context tokens by local estimate so payload compression can't suppress auto-compaction
A before_provider_request extension (a context-compression proxy like Headroom,
an obfuscator, or inline snapcompact) can shrink the outgoing request below the
real stored conversation. The provider then reports deflated prompt tokens, so
the auto-compaction threshold never fires and the stored history grows unbounded
until it overflows the context window and can no longer be compacted at all.

Add compactionContextTokens(provider, storedEstimate) = max(provider, estimate)
and apply it to both the pre-prompt and post-response compaction decisions,
flooring the provider-reported tokens by the agent's own estimate of the stored
conversation. Display and cost accounting still use exact provider usage; only
the compaction trigger takes the floor.
2026-06-20 23:21:36 +02:00
can1357 b5437c2cad feat(coding-agent): improved prompt caching for ephemeral side-channel requests
- Added `buildSideRequestContext` to the `Agent` class to generate prompt-cache-friendly provider contexts.
- Updated ephemeral side-channel turns to forward the full tool catalog to maintain prompt cache hit rates.
- Injected a `developer` role reminder into ephemeral turns to instruct the model to suppress tool calls.
- Implemented automatic post-processing to strip any tool calls from ephemeral turn responses.
- Exported message and dialect helper functions in `agent-loop.ts` to support context construction.
2026-06-20 23:19:10 +02:00
can1357 b88d97cd25 fix(coding-agent/session): prevented duplicate persistence of rewound tool results
- Added a set to track tool result IDs that have been rewound to ensure they are not re-appended to the session history.
- Updated message handling logic to conditionally skip persistence for tool results associated with active rewind operations.
2026-06-20 23:01:01 +02:00
can1357 81a4af6684 perf(coding-agent): deduplicated primary context messages in advisor history
- Implemented message deduplication in `AdvisorRuntime` to collapse verbatim re-injected primary context (plan rules/approved plans) between turns.
- Reduced token consumption by replacing identical primary context segments with a status marker.
- Updated history formatter to allow expansion of load-bearing context types specifically, while retaining one-line summaries for other custom messages.
2026-06-20 22:59:30 +02:00
can1357 60b043adb7 fix(coding-agent): fixed session history desynchronization during
- Fixed session history desynchronization during `rewind` by performing a full context rebuild after applying branch changes.
- Prevented `rewind` tool output and assistant side-channel data from polluting the prompt cache by flushing and sanitizing session state.
- Added explicit test coverage for context reconstruction and assistant message sanitization after rewind events.
2026-06-20 22:58:23 +02:00
can1357 321b51517a Merge PR #3022: refactor(coding-agent): use Promise.withResolvers in AsyncDrain (@oldschoola) 2026-06-20 22:36:13 +02:00
can1357 ffc842ee7c fix(agent): reset promoted memory prompt on session fork
fork() reset mnemopi conversation tracking directly but skipped the shared new-transcript reset, so the folded/promoted first-turn memory stayed in #baseSystemPrompt. The next turn re-recalled and the change-detection saw no diff, taking the fallback promotion path and injecting the <memories> block twice into the forked session prompt. Route fork() through #resetMemoryContextForNewTranscript() like the other reset paths and add a regression test asserting the forked prompt contains recalled memory exactly once.
2026-06-20 22:16:58 +02:00
can1357 10ed2d4efa Merge PR #3113: fix(agent): preserve memory recall in append-only prompt cache (@roboomp) 2026-06-20 22:16:58 +02:00
can1357 1fadcdb8e3 Merge PR #3147: fix(agent): compact active goal yields (@roboomp) 2026-06-20 22:16:57 +02:00
roboomp e00483b2c1 fix(agent): compact active goal yields
Run threshold compaction maintenance when an active goal turn ends through a successful yield, while preserving the final-yield skip for non-goal completions.

Fixes #3146
2026-06-20 19:01:26 +00:00
roboomp 1b197956a4 fix(agent): reset memory prompts on new transcripts
Restored and refreshed promoted memory prompts through the shared new-transcript reset path used by new sessions, handoff, branch, /btw, and cross-session switches.\n\nAdded coverage for newSession after first-turn memory recall so stale recalled memories cannot leak into the next transcript when recall returns no context.\n\nFixes #3111
2026-06-20 08:50:30 +00:00
roboomp 7d7cc1d1ba fix(agent): cleared promoted memory on session switch
Reset memory recall state before rebuilding the prompt during session switches, and clear fallback-promoted memory prompts when moving to another session.\n\nAdded a regression test that switches sessions after first-turn memory recall and verifies the next session does not receive stale memories.\n\nFixes #3111
2026-06-20 08:45:25 +00:00
roboomp 3d864fa763 fix(agent): preserved memory prompt cache
Promoted first-turn memory recall into the stable base prompt so append-only sessions do not drop the memory block on the next turn and rebuild the provider prefix.\n\nAdded a regression test covering a memory backend that recalls once before the first model request.\n\nFixes #3111
2026-06-20 08:32:48 +00:00
can1357 221f4102fb fix: resolved Fireworks Qwen models to openai thinking format
- Updated `buildOpenAICompat` to override the `qwen` thinking format for Fireworks-hosted models, ensuring they use `openai` thinking parameters instead.
- Prevented invalid `enable_thinking` payload errors by ensuring Fireworks-hosted Qwen requests conform to their strict schema.
- Updated `AgentSession` to allow Fireworks fast-fallback logic to execute even when standard retries are disabled.
2026-06-20 09:25:36 +02:00
can1357 1afa6ba68a feat(catalog): supported fireworks fast serving path
- Added support for "Fast" serving-path variants for select Fireworks models.
- Updated compatibility logic to route `-fast` suffixes to the appropriate router wire format.
- Extended the model generation catalog to include these Fast variants with their respective pricing.
- Updated AI types to allow the `priority` service tier for Fireworks providers.
2026-06-20 09:21:01 +02:00
oldschoola ac04400462 fix(coding-agent): Fixed AsyncDrain losing buffered entries by clearing queue early
- AsyncDrain reset #queue before invoking the handler, risking lost pushes.
- Moved queue reset to after the handler runs and wrapped exec in a promise.
- Added history-storage drain tests covering batch flush and fresh-batch behavior.
2026-06-19 17:35:27 -07:00
can1357 f8f8136021 refactor(coding-agent): privatized the legacy nextToolChoice method to
- Privatized the legacy `nextToolChoice` method to `#nextHardToolChoice` to ensure all tool-choice directives flow through the unified `nextToolChoiceDirective` entry point.
- Eliminated redundant dual entry points for fetching tool choices, which previously bypassed the soft pending-preview lifecycle.
- Updated test suites to consume `nextToolChoiceDirective` where appropriate to maintain consistency with internal agent-loop logic.
2026-06-19 22:24:12 +02:00
roboomp 49317a0f1c fix(agent): cleared stale pending preview gates
Cleared pending-preview markers when resolve has no runnable handler so a stale gate cannot keep forcing resolve after the invoker is gone.

Added regression coverage for apply and discard draining stale pending markers.

Fixes #3061
2026-06-19 17:54:37 +00:00
can1357 b2dc706ee7 feat: enabled auto-retry for AI thinking loops
- Improved thinking loop detection logic by refining text normalization and treating stalls as retryable errors.
- Instrumented the agent session to recognize thinking loop markers within retryable error conditions.
- Automated the clearing of stale error banners upon successful auto-retry execution.
- Added comprehensive test coverage for chunked thinking loop errors and banner management.
2026-06-19 17:51:53 +02:00
can1357 9478e3cc5c refactor: replaced ReturnType<typeof setTimeout> with Timer type
- Replaced usage of `ReturnType<typeof setTimeout>` and `ReturnType<typeof setInterval>` with the explicit `Timer` type across the codebase.
- Updated several type definitions and function signatures to use concrete types instead of inferred return types for improved clarity and maintainability.
2026-06-19 17:38:07 +02:00
can1357 2de02f2119 Merge PR #3019: fix: Windows test failures — path handling, EBUSY, SQLite handle leaks (@oldschoola) 2026-06-19 17:17:04 +02:00
can1357 480aa1763c Merge PR #3034: fix(mnemopi): isolate local embeddings worker in subprocess (@roboomp) 2026-06-19 17:16:36 +02:00
can1357 67f6518e42 feat: enhanced tool robustness, improve authentication flow, and update API parameters
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
2026-06-19 16:46:07 +02:00
can1357 a1dae58c34 fix(coding-agent/session): prevented emission of unhandled agent events
- Added an early return to stop event propagation if no handlers are registered for the event type.
- Updated the `agent_start` logic to return early, ensuring consistent flow control.
2026-06-19 16:06:17 +02:00
can1357 0dfeac8a75 feat: added auth discovery broker and expand model support
- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.
2026-06-19 16:06:16 +02:00
can1357 6d7dab0c68 fix(coding-agent): prevented crash when resuming session in deleted directory
- Export `directoryExists` in `utils` to safely validate working directories before traversal.
- Update `SessionManager` and startup logic to fallback to the launch directory if a session's recorded working directory no longer exists.
- Add regression tests to ensure sessions now correctly adopt the launch directory instead of crashing on missing paths.
2026-06-19 16:06:01 +02:00
can1357 81d0993915 feat(coding-agent/session): prevented adoption of missing session directories
- Add `directoryExists` check to validate session working directories.
- Skip updating the session cwd if the recorded directory no longer exists on disk to prevent runtime errors.
2026-06-19 16:06:00 +02:00
cagedbird043 5ff5f110af fix(agent): dynamically fallback from snapcompact to text summary on high CJK/non-ASCII rates 2026-06-19 21:50:16 +08:00
roboomp bf56a76e59 fix(mnemopi): isolated local embeddings worker in subprocess to skip onnxruntime napi crash
Moved mnemopi's local embedding provider out of the main agent process by
spawning a dedicated `Bun.spawn` child for fastembed + onnxruntime-node. The
agent CLI gains a hidden `__omp_worker_mnemopi_embed` dispatch; `loadMnemopi`
/ `loadMnemopiCore` install the subprocess-backed initializer through the
newly-exposed `setLocalModelInitializer` seam so every `embed()` call
round-trips through IPC instead of loading the NAPI module. The parent
SIGKILLs the child on dispose so the destructor that segfaults Bun on
Windows shutdown (NAPI finalizer at exit on npm installs, `process.dlopen`
constructor at session start on standalone binaries) never runs in any
address space the agent owns. Mirrors the tiny-model fix from #1607.

Adds `smokeTestMnemopiEmbedWorker` to `omp --smoke-test`, a new
`test/issue-3031-repro.test.ts` that pins the spawn/dispatch/signal-exit
contract and forbids re-importing `fastembed-runtime` from the agent surface,
and changelog entries.

Fixes #3031
2026-06-19 08:09:50 +00:00
can1357 7864a1363c feat(coding-agent): consolidated macOS power settings into enum
- Replaced four individual power boolean settings with a single `power.sleepPrevention` enum for improved configuration management.
- Implemented automatic migration logic in `Settings.init` to map legacy macOS power booleans to the new enum levels.
- Updated `AgentSession` to utilize the new `sleepPrevention` enum for macOS power assertion lifecycle management.
- Changed default value of `display.cacheMissMarker` to `false` to suppress cache-miss markers.
- Added comprehensive unit tests in `settings-manager.test.ts` to verify backward-compatible migration behavior.
2026-06-19 08:54:20 +02:00
can1357 82c4c14fd7 feat(coding-agent): added tracking for expected cache invalidations
- Added tracking for expected cache invalidations during model changes, compactions, and plan-mode transitions.
- Included `cacheMissExplainedAt` metadata in session context to prevent displaying misleading cache miss warnings in the transcript.
- Updated controller logic to reset assistant usage markers when mode-switching or performing actions that invalidate the prompt cache.
2026-06-19 08:08:14 +02:00
can1357 65f0d0a245 refactor(coding-agent): renamed repeatToolDescriptions to inlineToolDescriptors
- Renamed `repeatToolDescriptions` to `inlineToolDescriptors` throughout configuration, SDK, and internal session management.
- Set the default value to `true` and updated descriptions to clarify the descriptor inlining behavior.
- Fixed the `/dump` command to prevent duplicate tool inventory output when inlining is enabled.
2026-06-19 08:05:49 +02:00
can1357 ed73429faa feat: implemented soft tool requirement and escalation system
- Introduced `SoftToolRequirement` to support non-invasive tool enforcement with lifecycle management and escalation.
- Added `ToolChoiceDirective` to coordinate hard and soft tool requirements within the agent loop.
- Optimized preview workflows in `coding-agent` by replacing forced tool choices with non-forcing pending invokers.
- Enhanced `CompactionSummaryMessage` to prioritize structured rendering for tool requirement reminders.
2026-06-19 07:51:27 +02:00
can1357 b137764f0c feat: implemented text-first foveated snapcompact archive layout
- Implemented text-first archive structure that incorporates bounded source text with foveated image frames (HQ edges, LQ middle) to improve context quality.
- Updated snapcompact core logic to use `historyBlocks` for reconstruction, transitioning away from reliance on previous PNG frame inheritance.
- Increased default frame limits to 80 and adjusted token estimation constants (5024) to enhance context budget accuracy.
- Added `pruneToolDescriptions` and improved multi-block summary token estimation to optimize tool spec usage.
2026-06-19 07:38:54 +02:00
can1357 80af40eb55 perf(coding-agent): stabilized steering message wrapping across turns
- Changed `wrapSteeringForModel` to wrap every `steering:true` user message regardless of position, so a steer's wire bytes stay identical once buried instead of reverting to raw and busting the prompt cache.
- Neutralized the present-tense framing in `user-interjection.md` so always-wrapping no longer leaves stale "current task" wording on buried interjections.
- Updated `session-messages.test.ts` to assert buried steering messages are wrapped too.
2026-06-19 06:58:06 +02:00
can1357 dcb16d86f6 perf(compaction): kept per-turn tool-result pruning inside the warm prompt-cache prefix
- Added `keepBoundaryId` and `cacheWarmSuffixTokens` guards to `pruneToolOutputs` and `pruneSupersededToolResults` so superseded/useless results sitting in the already-sent cached prefix are no longer rewritten mid-session, with new `computeMessageSuffixTokens`/`resolveBoundaryIndex` helpers.
- Added a `keepBoundaryId` floor to `collectShakeRegions` so shake skips entries summarized away by the latest compaction.
- Threaded `firstKeptEntryId`, `PRUNE_CACHE_WARM_SUFFIX_TOKENS` (8k) and `PRUNE_IDLE_FLUSH_MS` (90m, above the 1h cache TTL) from `#pruneToolOutputs`, `#pruneStaleToolResults` and `shake` in `agent-session.ts`.
- Added six boundary tests in `supersede-prune.test.ts` covering warm-prefix protection, tail-case pruning, and the pre-boundary floor.
2026-06-19 06:57:53 +02:00
can1357 96151e6c30 feat: implemented tool description pruning to reduce token usage
- Added utility functions to strip descriptions from JSON schemas and tool definitions for optimized token output.
- Integrated `pruneToolDescriptions` configuration across agent loops and sessions to enable optional schema pruning.
- Updated agent context and snapshot logic to propagate pruning settings and maintain fingerprinting integrity.
- Verified schema structural integrity and removal of annotation descriptions through new unit tests.
2026-06-19 06:55:19 +02:00
can1357 f01b5e4c27 refactor(snapcompact): updated MAX_FRAMES_DEFAULT to 80 to better
- Update `MAX_FRAMES_DEFAULT` to 80 to better utilize high-capacity model context windows.
- Remove `providerFrameBudget` from `snapcompact` to decouple archival limits from provider-specific image caps.
- Update `compact` logic to treat `maxFrames` as a hard upper bound rather than a provider-clamped limit.
2026-06-19 06:33:19 +02:00
oldschoola 14252e71cb fix: Windows test failures — path handling, EBUSY, SQLite handle leaks
Fix all Windows-specific test failures caused by path handling problems
and EBUSY errors from unclosed SQLite database handles.

Root causes fixed:
1. POSIX path assumptions: replaced hard-coded file:///tmp, /repo, etc.
   with pathToFileURL/path.resolve/path.join computed expectations
2. shortenPath() now normalizes backslashes to forward slashes after ~
   and respects home directory boundaries
3. HistoryStorage.resetInstance() leaked its Database — added #close()
   that finalizes all prepared statements and closes the DB
4. AgentStorage gained the same resetInstance()/#close() pattern
5. SqliteAuthCredentialStore.close() leaked one-off prepared statements
   from inline this.#db.prepare() calls — wrapped each in try/finally
6. model-cache.ts used a process-global DB even for custom dbPath —
   now opens/closes per-call via withModelCacheDb
7. createAgentSession leaked AuthStorage on construction failure —
   added ownsAuthStorage cleanup in catch block
8. MnemopiBackend.removeDbFiles() now truly best-effort (catches errors)
9. TempDir retry window expanded from 4x10ms to 40x25ms
10. TempDir prefix convention: non-@ prefixes created dirs relative to
    cwd instead of os.tmpdir() — all test temp dirs now use @ prefix
11. Shell-escaped interpolated paths in bash tool tests
12. git core.autocrlf false in autoresearch test repo init

All 522 previously-failing Windows tests now pass.
2026-06-18 21:32:38 -07:00
can1357 4f03180ae6 refactor(deps): moved intent field constant to pi-wire
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
2026-06-19 04:58:29 +02:00
can1357 c224ef4df5 feat: standardized elision markers and improve transcript viewer robustness
- Standardized elision markers across all tool outputs and filters to use cohesive `[...N [type] elided...]`, `[...Nln elided...]`, and `[...xB elided...]` syntax.
- Updated documentation, prompts, and test expectations to reflect the unified elision format.
- Improved transcript viewer robustness by preventing content aliasing through path-inclusive signature hashing.
- Added logic to clear stale transcript content when associated session files are deleted, accompanied by verifying test cases.
2026-06-19 04:35:18 +02:00
can1357 29d250fae2 feat(coding-agent): supported advisor transcript persistence
- Implemented `AdvisorTranscriptRecorder` to persist advisor sessions to append-only `__advisor.jsonl` files.
- Integrated transcript recording into agent sessions with managed flushing, atomic file switching, and synthetic turn attribution.
- Restricted advisor-kind agents by excluding them from rosters, history protocols, messaging, and interactive agent commands.
- Reserved the `__advisor` filename stem across the output manager and task registry to prevent task ID collisions.
2026-06-19 03:50:00 +02:00
can1357 71144825ec feat(coding-agent): extended compact command with submode support
- Added a `mode` property to `CompactOptions` to allow fine-grained control over compaction strategies.
- Implemented `soft`, `remote`, and `snapcompact` submode overrides for the `/compact` command.
- Integrated `parseCompactArgs` to enable robust subcommand routing and validation, including focus instruction rejection for specific modes.
- Established a `CompactMode` registry to manage compaction strategies and verify remote availability.
2026-06-19 02:55:21 +02:00
can1357 5168e6d6af Merge PR #2791: fix(tool): resolve pasted image attachments in inspect_image (@roboomp)
# Conflicts:
#	packages/coding-agent/src/tools/inspect-image.ts
2026-06-19 00:59:03 +02:00