Commit Graph
124 Commits
Author SHA1 Message Date
can1357 24c8bb24c6 feat(session): added modular session APIs and rebuilt listing/persistence behavior
- Added session-domain modules and exports for session-entries, context, listing, loader, and migrations.
- Changed persistence to async append writes plus writeTextAtomic, removing sync line APIs.
- Added compaction-aware session context rebuild with dangling tool-call cleanup.
- Added resumable session resolution with status inference, id/stem/suffix matching, and backup recovery.
2026-06-14 02:02:53 +02:00
can1357 64aa558e62 chore: consistency 2026-06-13 00:03:27 +02:00
can1357 9629842f33 feat(coding-agent): accepted a model directly in ModelRegistry.resolver 2026-06-12 02:33:46 +02:00
can1357 72c12ff9c2 feat(task): migrated task tool to batch-first mode with shared context
- Replaced task-simple-mode with a `task.batch` setting enabled by default.
- Updated task schema to use batch `{agent, context, tasks[]}` payloads.
- Migrated task execution to spawn one async job per task and merge outputs.
- Removed per-call schema passing while preserving legacy flat task calls.
2026-06-10 23:47:28 +02:00
can1357 9d99ae1af0 feat(coding-agent): rewrote the task tool to spawn one persistent subagent per call
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
2026-06-10 17:54:47 +02:00
can1357 a92d2ce989 feat(coding-agent): removed context argument from eval agent() spawn
Shared background now flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each prompt instead of a context string forwarded into the subagent's system prompt. The JS and Python preludes drop the context kwarg from agent(), the subagent system prompt drops the {{#if context}} block and the conversation-context file pointer, and runEvalAgent no longer writes a per-call conversation context file. AgentSession sheds the now-unused formatCompactContext() helper that supplied the file's body, and ToolSession.getCompactContext is removed alongside it.
2026-06-10 17:49:01 +02:00
can1357 4068bfc304 fix(coding-agent): key python kernel sessions by resolved interpreter 2026-06-10 09:51:45 +02:00
can1357 7ce58fe95c Merge pull request #2204: feat(eval): add python interpreter setting 2026-06-10 08:26:59 +02:00
can1357 24cbc93913 fix(eval): thread python.interpreter from session settings and expand ~
Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.

Addresses review feedback on #2204.
2026-06-10 08:26:00 +02:00
roboompandcan1357 8c3149e5a9 fix(ai): scoped antigravity quota blocks by model family
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.

Fixes #2198
2026-06-10 08:26:00 +02:00
danzaioandcan1357 caff395d07 feat(eval): add python interpreter setting 2026-06-10 08:26:00 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 8583421362 perf(coding-agent): deferred heavy imports until first feature use
- Added lazy async module loaders for @babel/parser, linkedom, puppeteer/browsers, @mozilla/readability, @xterm/headless, and mnemopi to avoid loading them during cold startup.
- Added an interactive startup splash before session construction and skipped it for resume/fork/continue, quiet mode, timing mode, or non-TTY runs.
- Updated JS import-rewrite and memory tests to match async parser loading and preloaded mnemopi modules for sync state helpers.
2026-06-10 03:57:32 +02:00
can1357 dc5c93462f feat: rerouted worker subprocesses through the bundled CLI host entrypoint
- Rerouted sync, tab, js-eval, and tiny workers to re-enter CLI modes via `__omp_*` selectors.
- Adjusted `cli.ts` startup to dispatch worker entrypoints before parsing and exit 1 on uncaught errors.
- Bundled CLI as `dist/cli.js` in prepack, switching `omp` binary and published files.
- Removed explicit Bun `--compile` worker entrypoints from build/release scripts in favor of host-entry dispatch.
- Added `declareWorkerHostEntry()` and `workerHostEntry()` environment helpers and `PI_COMPILED` binary detection.
2026-06-10 03:57:31 +02:00
can1357 f638a5b3e0 refactor(coding-agent): added lazy loading for heavy dependencies
- Introduced memoized dynamic import loaders for Babel parser, mnemopi modules, puppeteer, and HTML-related packages.
- Refactored eval import-rewrite helpers and runtime call sites to use asynchronous wrapping and parsing flows.
- Shifted fetch and web-scraper linkedom usage to on-demand imports so heavy modules load only when needed.
2026-06-10 03:09:09 +02:00
can1357 d7eae06830 fix(coding-agent): fixed eval artifact double-writes and kernel I/O capture
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
2026-06-10 01:27:41 +02:00
can1357 3e990c5bea fix(coding-agent): added guarded dynamic import shim for worker fallback
- Rewrote dynamic `import(...)` rewriting to emit a guarded callee that prefers `__omp_import__` and falls back to native `import` when the helper is unavailable.
- Added a shared shim constant and updated import-rewrite tests to verify routed dynamic imports work both with the injected helper and after serializing into a realm without it.
2026-06-09 22:57:19 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
can1357 cfd72ffb40 fix(coding-agent): fixed eval helpers treating local:// as plain paths
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
2026-06-08 18:28:06 +02:00
can1357 59cb767374 perf(coding-agent/eval): forced eval agent subagents to skip LSP startup
- Changed runEvalAgent to always pass enableLsp: false when launching bridge subagents.
- Added a regression test asserting runSubprocess received enableLsp as false even when LSP is enabled by default.
2026-06-08 12:25:53 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
basedcorp99 1e2bbe71aa fix(eval): use withResolvers for worker close 2026-06-08 11:49:23 +02:00
basedcorp99 1136ba81a1 fix(eval): exit graceful-close JS worker 2026-06-08 11:16:49 +02:00
basedcorp99 a6c8014697 fix(coding-agent): close JS eval workers gracefully 2026-06-08 11:16:49 +02:00
can1357 65bddab719 feat(coding-agent/eval): enforced JS eval helper options as trailing object literals
- Added strict option parsing for JavaScript helpers so read/sort/uniq/counter/tree/llm/agent now reject positional options and non-object option values.
- Introduced `optionsArg` and `isPlainObject` to enforce a single trailing options object and throw clear TypeError messages on invalid calls.
- Updated the eval tool prompt to document that JavaScript helpers require one trailing options object and no extra positional arguments.
2026-06-08 03:33:45 +02:00
can1357 088fb7fb75 fix(eval): resolved JS/Python resets by awaiting in-flight operations
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
2026-06-08 02:04:59 +02:00
can1357 49aa6e5839 fix(eval): corrected eval LLM calls and spawn-aware tool descriptions
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.
2026-06-08 02:04:59 +02:00
can1357 f73942eb28 fix(eval/py): corrected read() positional offset/limit support and tests
- Updated `packages/coding-agent/src/eval/py/prelude.py` to accept positional `offset` and `limit` in `read()`.
- Added regression coverage in `packages/coding-agent/src/eval/py/__tests__/prelude.test.ts` for positional read signatures.
2026-06-08 02:04:59 +02:00
can1357 caa9c69309 feat(coding-agent): added provider-priority model selection and refined fallback ordering
- Added first-party-first provider priority defaults for model ranking.
- Consolidated model resolution to use getModelMatchPreferences from session settings.
- Prioritized providerPriorityRank ahead of usage rank when picking preferred models.
- Added second-pass fallback to default-model or API-key-valid matching order.
2026-06-08 01:32:39 +02:00
Can BölükandGitHub e9e2d02223 Merge pull request #1982 from AsafMah/fix/js-eval-worker-init-flake
fix(eval): floor JS worker init timeout to stop terminate-mid-init CI flake
2026-06-08 00:19:33 +02:00
can1357 20487d8cb6 chore: fix stale tests 2026-06-07 08:11:38 +02:00
can1357 0914379f49 fix(coding-agent): rendered shared task context as markdown and froze async block borders
- Updated task call and result rendering to process shared context with the Markdown renderer, so context sections are now displayed with proper Markdown formatting.
- Stopped shimmer animation on pending bash/eval/task blocks once async state is `running`, preventing the committed frame from freezing a transient dark border segment.
- Adjusted rule path display to fall back to a root-relative path when cwd-relative resolution is unavailable.
2026-06-07 08:01:52 +02:00
can1357 76af35dcd6 fix(tui): preserved TUI initial scrollback unless clear was requested
- Updated the initial render intent to carry a `clearScrollback` flag.
- Adjusted first-frame rendering logic to preserve existing scrollback by default and clear only when requested.
- Added `resolver` stubs to coding-agent and web-search test registries for interface compatibility.
2026-06-07 07:54:50 +02:00
can1357 c10eb5e50e feat(coding-agent): added resolver-based auth retries to image and search tools
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
2026-06-07 04:49:19 +02:00
can1357 cba6641299 feat: enabled resolver-based API key retries with refresh and rotation
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
2026-06-07 04:20:48 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 5721034739 fix(ui): forced full replay on tool output expand toggle
- Replaced viewport-only repaint with resetDisplay so committed scrollback reflects new heights.
- Added per-server rust-analyzer workspace-ready timing overrides as a test seam.
- Added clearSuppressedSelectors to reset retry-fallback cooldown state.
- Removed obsolete shared eval executors test.
2026-06-06 21:56:00 +02:00
can1357 3b5b182553 Merge remote-tracking branch 'origin/farm/10b5f6a2/surface-subagent-abort-reason' 2026-06-06 21:34:33 +02:00
can1357 133137c9a6 fix(eval): surfaced subagent abort reasons and disabled runtime cap
- Used `||` so empty stderr falls through to abortReason in agent bridge.
- Preferred assistant errorMessage over "Cancelled by caller" on internal aborts.
- Forced `maxRuntimeMs: 0` for eval subagents via ExecutorOptions override.
2026-06-06 21:33:07 +02:00
can1357 bdbbfa9778 fix(eval): surfaced subagent abort reason
- Used `||` so empty stderr no longer masks the real abort reason.
2026-06-06 20:54:00 +02:00
roboomp cab465cba4 fix(eval): surfaced subagent abort reason through python agent() bridge
Python eval agent() collapsed every subagent runtime-limit abort into a
generic 'RuntimeError: bridge call __agent__ failed' instead of the real
reason. runEvalAgent built its failure message with:

  result.error ?? result.stderr ?? result.abortReason ?? <default>

? is nullish-coalescing, so result.stderr = "" (the executor's value for
a runtime-limit abort) short-circuited the chain and never reached
abortReason. The host bridge then shipped {ok: false, error: ""}, and
prelude.py's '<msg> or <fallback>' picked the named-bridge fallback.

Extracted buildSubagentFailureMessage(): aborted subagents prefer the
trimmed abortReason; otherwise fall through error, stderr (trimmed),
abortReason, and the named-bridge default. Empty/whitespace strings no
longer mask anything. The failure-detection condition also accepts
result.aborted so an abort with exitCode 0 (theoretically) still flows
the abort reason out.

Added a regression test asserting that runtime-limit aborts, whitespace
stderr/error, and totally blank aborts all produce non-empty messages
matching the executor's abortReason text.

Fixes #2006
2026-06-06 18:15:04 +00:00
can1357 c057b0ef33 fix(coding-agent/eval): prevented subagent eval session deadlock
- Stopped sharing parentEvalSessionId with bridge-spawned subagents.
- Sharing it deadlocked since the parent's kernel blocks on the bridge call.
- Each subagent now gets its own eval session with an independent kernel.
2026-06-06 16:55:05 +02:00
can1357 001a6ad564 tests: remove useless assertations 2026-06-06 16:00:31 +02:00
can1357 ecd80120a3 feat(coding-agent): added framed tool rendering with capped streaming previews
- Wrapped tool call/result renderers with framed block markers for full-width display.
- Added preview-line caps via capPreviewLines to bound multiline outputs with truncation hints.
- Added code-cell tail rendering so streaming output shows capped tail slices with markers.
- Added setPaddingX in Box and applied framed-inline padding adjustments for tight layouts.
2026-06-06 15:19:06 +02:00
can1357 76ca917bfd fix(coding-agent-eval): suspended idle timeout during delegated bridge calls
- Added timeout pause/resume control ops and helper to suspend idle timers during bridge work.
- Fixed TimeoutError on long delegated agent()/llm() calls by pausing timeout while silent.
- Updated IdleTimeout with reference-counted pauses, ignored checks while paused, and resumed fresh.
- Replaced heartbeat keepalives with timeout-control events in bridge paths and status routing.
2026-06-06 14:02:04 +02:00
Asaf Mahlev ea8d58d3c6 docs(eval): clarify indirect-eval cross-reference in worker-init comment
Address review: shared/indirect-eval.ts documents the vm.runInContext mid-execution terminate-race, not a mid-init one. Reword so a reader doesn't grep that file for an init-specific note.
2026-06-06 12:43:53 +03:00
Asaf Mahlev 236bba96b9 fix(eval): floor JS worker init timeout to stop terminate-mid-init flake
Worker-ready wait reused Bun's 5s default per-test timeout as its floor, so a slow cold-start under --isolate + high CI concurrency was aborted at 5s. The catch then terminate()s a still-initializing Bun worker -- the documented SIGILL/SIGTRAP crash trigger -- crashing the whole test file and intermittently failing unrelated PRs.

Introduce WORKER_INIT_TIMEOUT_MS=15s as a fixed infrastructure floor (independent of, still dominated by, a larger per-cell timeout) and set a 20s file-local setDefaultTimeout in js-executor/js-workflow-helpers tests so cold starts complete instead of being torn down.
2026-06-06 11:56:23 +03:00
roboomp 867ee74a3b fix(eval): probed Win32 console directly instead of inferring from stdio
The TTY-OR heuristic still mis-classifies the all-stdio-redirected case:
`omp -p "..." < in.txt > out.txt 2> err.log` from a real Windows Terminal
session has `stdin.isTTY === stdout.isTTY === stderr.isTTY === false`,
but the host process can still own a console that the kernel could
inherit. The TTY signals reflect handle redirection, not console
attachment — `GetConsoleWindow()` is the authoritative Win32 signal.

`spawn-options.ts` now:

- Calls `kernel32!GetConsoleWindow()` via `bun:ffi` on Windows. A non-NULL
  HWND means the host has a console regardless of how the standard
  streams are wired, so `windowsHide` stays `false` and the kernel
  inherits the console — which is what fixes the #1960
  numpy/pandas `LoadLibraryExW` hang and lets SIGINT recover via
  `GenerateConsoleCtrlEvent`.
- Falls back to the TTY-OR heuristic when the FFI probe is unavailable
  or off-Windows. That keeps the predicate working on non-Bun-FFI
  runtimes and on POSIX, where `windowsHide` is a no-op anyway.
- Caches the probe result; console attachment is stable for the host's
  lifetime in practice, and dlopening kernel32 on every kernel spawn
  would be wasteful.

The pure helpers (`shouldHideKernelWindow`, `consoleAttachedViaTTY`)
stay separately exported so they remain unit-testable. The integration
boundary `hostHasInheritableConsole()` is what the kernel spawn site
calls, and its return is asserted to be a concrete boolean (kernel spawn
must commit to a `windowsHide` value).
2026-06-06 01:04:13 +00:00
roboomp 69ffc80368 style: bun run fix 2026-06-06 00:59:23 +00:00