Commit Graph
132 Commits
Author SHA1 Message Date
can1357 3ab675a83b fix: added buffered worker inboxing and standardized worker selectors
- Added `WorkerInbox` and `installWorkerInbox(port)` to queue worker messages before bind.
- Added `consumeWorkerInbox()` to replay buffered messages and clear one active inbox.
- Added buffered inbox consumption in JS and tab worker transports before direct message handlers.
- Normalized worker selector arguments to the `__omp_worker_*` naming across workers and tests.
2026-06-15 11:59:44 +02:00
can1357 6385afdfb7 test(coding-agent): replaced Bun.sleep and wall-clock timing
- Replaced Bun.sleep and wall-clock timing with fake timers (vi.useFakeTimers), release gates, and deterministic polling across 15+ test files to eliminate flakiness and improve speed.
- Consolidated per-test fixture setup into beforeAll/afterAll lifecycle hooks across 20+ test files, reducing redundant initialization and improving test performance by reusing shared immutable fixtures.
- Stubbed network calls in ModelRegistry and test discovery to prevent unintended outbound requests during test execution.
- Replaced subprocess-based test coordination (file markers, Bun.sleep polling) with in-memory fakes (FakeWebSocket, FakeLspServer, VirtualClock) for deterministic, fast test execution.
2026-06-15 11:48:55 +02:00
can1357 eaffa00177 fix(coding-agent/eval): routed JS worker ready signaling through init response
- Updated worker core so it emitted `ready` only after receiving an `init` message instead of on construction.
- Changed worker startup to send `init` before waiting for readiness, enabling startup errors to fail fast and trigger fallback.
- Expanded JS worker tests with startup-error simulation and verified execution falls back to the inline worker when spawning fails.
2026-06-15 08:49:00 +02:00
can1357 ebe05c3fc2 fix(coding-agent/eval): fixed JS eval worker initialization with inline fallback on failure
- Initialized JS workers through a new helper that attaches message and error listeners before initialization handshake and enforces the existing ready timeout.
- Handled startup failures by rejecting on init or error events, cleaning up listeners and replacing dead workers with a retry path.
- Retried session creation with an inline worker after non-inline init failure so requests no longer stall on worker startup timeouts.
2026-06-15 08:45:15 +02:00
can1357 7b319f0f02 fix(coding-agent/eval): routed process stdio writes to active runtime text sink
- Added runtime hook resolvers so each `JsRuntime` instance can expose hooks for its active run.
- Patched `process.stdout` and `process.stderr` writes once per stream to route output chunks through active run text hooks and preserve existing worker logging when no run is active.
- Added chunk-to-string conversion for write payloads and encoding-aware forwarding while keeping callback semantics intact.
2026-06-15 03:04:10 +02:00
can1357 2be8256af0 ci(workflows): shared bun caching in CI via backend-aware install action
- Updated CI dependency install flow to share bun cache orchestration across jobs.
- Added RustFS-backed bun cache restore/save script keyed by bun.lock hash.
2026-06-15 00:18:43 +02:00
can1357 932de0a1dd test(coding-agent): updated test import path and removed stale ffmpegAssetName tests
- Changed helper tests to import TempDir from @oh-my-pi/pi-utils/temp.
- Removed the obsolete ffmpegAssetName test file under tools-manager.
2026-06-14 22:02:03 +02:00
can1357 74bfcabc35 fix(coding-agent): fixed JS eval helper positional args and URI read slicing
- Updated JS eval helper option parsing to accept positional optional args or a trailing plain-object options argument, while rejecting mixed or invalid forms.
- Updated `read` to route non-`local://` URI paths through the `read` tool with line selectors so offset/limit slicing works for artifact-style resources.
- Expanded JS executor tests for positional reads, nullable slot skipping, and delegated URI read slicing.
2026-06-14 07:27:16 +02:00
can1357 24c8bb24c6 feat(session): added modular session APIs and rebuilt listing/persistence behavior
- Added session-domain modules and exports for session-entries, context, listing, loader, and migrations.
- Changed persistence to async append writes plus writeTextAtomic, removing sync line APIs.
- Added compaction-aware session context rebuild with dangling tool-call cleanup.
- Added resumable session resolution with status inference, id/stem/suffix matching, and backup recovery.
2026-06-14 02:02:53 +02:00
can1357 64aa558e62 chore: consistency 2026-06-13 00:03:27 +02:00
can1357 9629842f33 feat(coding-agent): accepted a model directly in ModelRegistry.resolver 2026-06-12 02:33:46 +02:00
can1357 72c12ff9c2 feat(task): migrated task tool to batch-first mode with shared context
- Replaced task-simple-mode with a `task.batch` setting enabled by default.
- Updated task schema to use batch `{agent, context, tasks[]}` payloads.
- Migrated task execution to spawn one async job per task and merge outputs.
- Removed per-call schema passing while preserving legacy flat task calls.
2026-06-10 23:47:28 +02:00
can1357 9d99ae1af0 feat(coding-agent): rewrote the task tool to spawn one persistent subagent per call
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
2026-06-10 17:54:47 +02:00
can1357 a92d2ce989 feat(coding-agent): removed context argument from eval agent() spawn
Shared background now flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each prompt instead of a context string forwarded into the subagent's system prompt. The JS and Python preludes drop the context kwarg from agent(), the subagent system prompt drops the {{#if context}} block and the conversation-context file pointer, and runEvalAgent no longer writes a per-call conversation context file. AgentSession sheds the now-unused formatCompactContext() helper that supplied the file's body, and ToolSession.getCompactContext is removed alongside it.
2026-06-10 17:49:01 +02:00
can1357 4068bfc304 fix(coding-agent): key python kernel sessions by resolved interpreter 2026-06-10 09:51:45 +02:00
can1357 7ce58fe95c Merge pull request #2204: feat(eval): add python interpreter setting 2026-06-10 08:26:59 +02:00
can1357 24cbc93913 fix(eval): thread python.interpreter from session settings and expand ~
Resolve the explicit interpreter from the session's Settings instance
(ToolSession.settings / AgentSession.settings) instead of re-reading the
process-global Settings.init() singleton, so project-scoped and cloned
session settings take effect. The availability cache is now keyed by
cwd + interpreter, and PythonKernel.start/executePython accept the
resolved interpreter as an option. Also expand home-relative paths
(~/...) before resolving against cwd, and document the contract of
resolveExplicitPythonRuntime.

Addresses review feedback on #2204.
2026-06-10 08:26:00 +02:00
roboompandcan1357 8c3149e5a9 fix(ai): scoped antigravity quota blocks by model family
- Added CredentialRankingStrategy scope hooks so providers can rank and block only the limits relevant to the requested model.
- Scoped Antigravity usage reports by model family: Gemini/Gemma use Google counters, Claude uses Anthropic counters, and GPT/OpenAI models use OpenAI counters.
- Added scoped backoff keys so a Gemini quota block no longer suppresses healthy Claude/OpenAI Antigravity sessions on the same OAuth credential.
- Threaded modelId through coding-agent API-key resolvers and usage-limit rotation paths.
- Added regression coverage proving a Google/Gemini exhaustion block still allows Claude selection on the same credential.

Fixes #2198
2026-06-10 08:26:00 +02:00
danzaioandcan1357 caff395d07 feat(eval): add python interpreter setting 2026-06-10 08:26:00 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00
can1357 8583421362 perf(coding-agent): deferred heavy imports until first feature use
- Added lazy async module loaders for @babel/parser, linkedom, puppeteer/browsers, @mozilla/readability, @xterm/headless, and mnemopi to avoid loading them during cold startup.
- Added an interactive startup splash before session construction and skipped it for resume/fork/continue, quiet mode, timing mode, or non-TTY runs.
- Updated JS import-rewrite and memory tests to match async parser loading and preloaded mnemopi modules for sync state helpers.
2026-06-10 03:57:32 +02:00
can1357 dc5c93462f feat: rerouted worker subprocesses through the bundled CLI host entrypoint
- Rerouted sync, tab, js-eval, and tiny workers to re-enter CLI modes via `__omp_*` selectors.
- Adjusted `cli.ts` startup to dispatch worker entrypoints before parsing and exit 1 on uncaught errors.
- Bundled CLI as `dist/cli.js` in prepack, switching `omp` binary and published files.
- Removed explicit Bun `--compile` worker entrypoints from build/release scripts in favor of host-entry dispatch.
- Added `declareWorkerHostEntry()` and `workerHostEntry()` environment helpers and `PI_COMPILED` binary detection.
2026-06-10 03:57:31 +02:00
can1357 f638a5b3e0 refactor(coding-agent): added lazy loading for heavy dependencies
- Introduced memoized dynamic import loaders for Babel parser, mnemopi modules, puppeteer, and HTML-related packages.
- Refactored eval import-rewrite helpers and runtime call sites to use asynchronous wrapping and parsing flows.
- Shifted fetch and web-scraper linkedom usage to on-demand imports so heavy modules load only when needed.
2026-06-10 03:09:09 +02:00
can1357 d7eae06830 fix(coding-agent): fixed eval artifact double-writes and kernel I/O capture
single OutputSink owner per cell artifact; JS parallel() honors its documented barrier (allSettled) instead of orphaning in-flight thunks; Python subprocesses no longer inherit the NDJSON frame pipe (stdout captured and forwarded); JS timeouts annotate the VM reset; console bridge implements dir/time/group/assert/trace; python availability probe cached; runner frames coalesce per write.
2026-06-10 01:27:41 +02:00
can1357 3e990c5bea fix(coding-agent): added guarded dynamic import shim for worker fallback
- Rewrote dynamic `import(...)` rewriting to emit a guarded callee that prefers `__omp_import__` and falls back to native `import` when the helper is unavailable.
- Added a shared shim constant and updated import-rewrite tests to verify routed dynamic imports work both with the injected helper and after serializing into a realm without it.
2026-06-09 22:57:19 +02:00
can1357 bda3102451 ux(coding-agent): updated status glyphs and fixed extension model discovery refresh
- Added status.done and tool.* symbols to theme mappings and presets.
- Replaced generic success glyphs with contextual +/-, tool icons, and warnings.
- Mapped tool/task/job completions to status.done or status.enabled with icon overrides.
- Triggered runtime provider refresh after extension registration and warned on failure.
2026-06-08 18:28:15 +02:00
can1357 cfd72ffb40 fix(coding-agent): fixed eval helpers treating local:// as plain paths
- Substituted injected on-disk roots for `local://` in read/write/append (Python and JS).
- Pinned `local://` to the session's own root so eval writes land where reads resolve.
- Rejected path traversal and unknown `scheme://` paths instead of creating junk `local:` dirs.
- Added unit and integration tests covering resolution, guards, and plain-path passthrough.
2026-06-08 18:28:06 +02:00
can1357 59cb767374 perf(coding-agent/eval): forced eval agent subagents to skip LSP startup
- Changed runEvalAgent to always pass enableLsp: false when launching bridge subagents.
- Added a regression test asserting runSubprocess received enableLsp as false even when LSP is enabled by default.
2026-06-08 12:25:53 +02:00
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
basedcorp99 1e2bbe71aa fix(eval): use withResolvers for worker close 2026-06-08 11:49:23 +02:00
basedcorp99 1136ba81a1 fix(eval): exit graceful-close JS worker 2026-06-08 11:16:49 +02:00
basedcorp99 a6c8014697 fix(coding-agent): close JS eval workers gracefully 2026-06-08 11:16:49 +02:00
can1357 65bddab719 feat(coding-agent/eval): enforced JS eval helper options as trailing object literals
- Added strict option parsing for JavaScript helpers so read/sort/uniq/counter/tree/llm/agent now reject positional options and non-object option values.
- Introduced `optionsArg` and `isPlainObject` to enforce a single trailing options object and throw clear TypeError messages on invalid calls.
- Updated the eval tool prompt to document that JavaScript helpers require one trailing options object and no extra positional arguments.
2026-06-08 03:33:45 +02:00
can1357 088fb7fb75 fix(eval): resolved JS/Python resets by awaiting in-flight operations
- Coalesced concurrent JS and Python reset requests by awaiting in-flight promises instead of throwing.
- Aligned non-reset execution calls to wait for in-progress resets before running on a recreated session.
2026-06-08 02:04:59 +02:00
can1357 49aa6e5839 fix(eval): corrected eval LLM calls and spawn-aware tool descriptions
- Ensured eval LLM calls always include a non-empty system prompt to avoid 400s.
- Aligned eval tool docs with session spawn policy by omitting agent() when spawns are disallowed.
- Added regression coverage for default system prompts and spawn-aware agent() description behavior.
2026-06-08 02:04:59 +02:00
can1357 f73942eb28 fix(eval/py): corrected read() positional offset/limit support and tests
- Updated `packages/coding-agent/src/eval/py/prelude.py` to accept positional `offset` and `limit` in `read()`.
- Added regression coverage in `packages/coding-agent/src/eval/py/__tests__/prelude.test.ts` for positional read signatures.
2026-06-08 02:04:59 +02:00
can1357 caa9c69309 feat(coding-agent): added provider-priority model selection and refined fallback ordering
- Added first-party-first provider priority defaults for model ranking.
- Consolidated model resolution to use getModelMatchPreferences from session settings.
- Prioritized providerPriorityRank ahead of usage rank when picking preferred models.
- Added second-pass fallback to default-model or API-key-valid matching order.
2026-06-08 01:32:39 +02:00
Can BölükandGitHub e9e2d02223 Merge pull request #1982 from AsafMah/fix/js-eval-worker-init-flake
fix(eval): floor JS worker init timeout to stop terminate-mid-init CI flake
2026-06-08 00:19:33 +02:00
can1357 20487d8cb6 chore: fix stale tests 2026-06-07 08:11:38 +02:00
can1357 0914379f49 fix(coding-agent): rendered shared task context as markdown and froze async block borders
- Updated task call and result rendering to process shared context with the Markdown renderer, so context sections are now displayed with proper Markdown formatting.
- Stopped shimmer animation on pending bash/eval/task blocks once async state is `running`, preventing the committed frame from freezing a transient dark border segment.
- Adjusted rule path display to fall back to a root-relative path when cwd-relative resolution is unavailable.
2026-06-07 08:01:52 +02:00
can1357 76af35dcd6 fix(tui): preserved TUI initial scrollback unless clear was requested
- Updated the initial render intent to carry a `clearScrollback` flag.
- Adjusted first-frame rendering logic to preserve existing scrollback by default and clear only when requested.
- Added `resolver` stubs to coding-agent and web-search test registries for interface compatibility.
2026-06-07 07:54:50 +02:00
can1357 c10eb5e50e feat(coding-agent): added resolver-based auth retries to image and search tools
- Routed image-gen, inspect-image, and web search providers through `withAuth`.
- Used `reuseInitialApiKey`/`createAuthStorageResolver` for force-refresh and rotate retries.
- Attached HTTP status to thrown errors so the retry classifier detects retryable failures.
2026-06-07 04:49:19 +02:00
can1357 cba6641299 feat: enabled resolver-based API key retries with refresh and rotation
- Added `ApiKeyResolver`/`ApiKey` types and exported auth-retry helpers.
- Changed stream and gateway auth retry handling to use resolver steps.
- Added initial-key, force-refresh, and rotate credential retries for auth failures.
- Updated agent and coding-agent integrations to use context-aware API-key resolvers.
2026-06-07 04:20:48 +02:00
can1357 20d19e8002 test: replaced blind sleeps with shared fixtures and condition polling
- Shared immutable model registries and auth storage via beforeAll/afterAll.
- Swapped fixed-delay settle sleeps for predicate polling and signals.
- Stubbed network/timers to drop wall-clock waits in registry and history tests.
- Added resetDisplay invalidation tests and startup-timing breakdown lines.
2026-06-06 22:09:04 +02:00
can1357 5721034739 fix(ui): forced full replay on tool output expand toggle
- Replaced viewport-only repaint with resetDisplay so committed scrollback reflects new heights.
- Added per-server rust-analyzer workspace-ready timing overrides as a test seam.
- Added clearSuppressedSelectors to reset retry-fallback cooldown state.
- Removed obsolete shared eval executors test.
2026-06-06 21:56:00 +02:00
can1357 3b5b182553 Merge remote-tracking branch 'origin/farm/10b5f6a2/surface-subagent-abort-reason' 2026-06-06 21:34:33 +02:00
can1357 133137c9a6 fix(eval): surfaced subagent abort reasons and disabled runtime cap
- Used `||` so empty stderr falls through to abortReason in agent bridge.
- Preferred assistant errorMessage over "Cancelled by caller" on internal aborts.
- Forced `maxRuntimeMs: 0` for eval subagents via ExecutorOptions override.
2026-06-06 21:33:07 +02:00
can1357 bdbbfa9778 fix(eval): surfaced subagent abort reason
- Used `||` so empty stderr no longer masks the real abort reason.
2026-06-06 20:54:00 +02:00
roboomp cab465cba4 fix(eval): surfaced subagent abort reason through python agent() bridge
Python eval agent() collapsed every subagent runtime-limit abort into a
generic 'RuntimeError: bridge call __agent__ failed' instead of the real
reason. runEvalAgent built its failure message with:

  result.error ?? result.stderr ?? result.abortReason ?? <default>

? is nullish-coalescing, so result.stderr = "" (the executor's value for
a runtime-limit abort) short-circuited the chain and never reached
abortReason. The host bridge then shipped {ok: false, error: ""}, and
prelude.py's '<msg> or <fallback>' picked the named-bridge fallback.

Extracted buildSubagentFailureMessage(): aborted subagents prefer the
trimmed abortReason; otherwise fall through error, stderr (trimmed),
abortReason, and the named-bridge default. Empty/whitespace strings no
longer mask anything. The failure-detection condition also accepts
result.aborted so an abort with exitCode 0 (theoretically) still flows
the abort reason out.

Added a regression test asserting that runtime-limit aborts, whitespace
stderr/error, and totally blank aborts all produce non-empty messages
matching the executor's abortReason text.

Fixes #2006
2026-06-06 18:15:04 +00:00