Commit Graph

13 Commits

Author SHA1 Message Date
can1357 f74c6d892a feat(coding-agent-eval): renamed the eval helper API from llm() to completion()
- Renamed eval oneshot helper from llm() to completion() across JS/Python APIs.
- Remapped eval bridge internals to completion semantics (__completion__, runEvalCompletion, completion status op).
- Updated docs, prompts, and timeout guidance to describe completion() usage and behavior.
- Adjusted completion defaults for active-session model preference, fallback parsing, and slow-tier effort handling.
2026-06-08 12:22:47 +02:00
can1357 384a206737 refactor(task): replaced numeric-prefix ids with name-first agent output ids
- Changed `AgentOutputManager` to use requested names verbatim, adding `-2`/`-3` suffixes only on repeats (e.g. `Anna`, `Anna-2`).
- Renamed main agent id from `0-Main` to `Main`; nested ids now use dot notation without numeric prefix (e.g. `Parent.Child`).
- Updated task widget to render dotted hierarchy as `Parent>Child` breadcrumb without leading index.
- Resume scan now tracks seen names instead of a counter to avoid clobbering prior outputs.
2026-06-02 06:50:03 +02:00
can1357 cf621d0abf feat(coding-agent-eval): added runEvalAgent bridge for agent plan checks
- Added `agent()` in JS/Python preludes to call host bridge and parse returned text when schema is set.
- Added JS `parallel()` and `pipeline()` with bounded `__pool()` pools and concurrency normalization.
- Added `runEvalAgent` bridge logic with argument parsing plus plan-mode, allowlist, depth, and artifacts checks.
- Added tool routing and tests documenting new `agent/parallel/pipeline` behavior, defaults, and validation failures.
2026-05-31 06:57:02 +02:00
can1357 1dba122c53 chore: updated docs 2026-05-31 04:36:14 +02:00
can1357 8715ed207c feat(coding-agent-eval): added oneshot llm helper and __llm__ bridge
- Added one-shot `llm(prompt, opts)` helpers in JS and Python eval runtimes.
- Added `__llm__` eval bridge wiring for synthetic LLM tool dispatch and status/event output.
- Added `runEvalLlm` with tier-to-model resolution, effort handling, and oneshot completion execution.
- Added structured schema output handling via `respond` tool and JSON fallback parsing.
- Documented new llm behavior in eval docs/changelog and added tests for tier mapping and error cases.
2026-05-30 00:27:04 +02:00
can1357 8a5b3e9552 feat(eval): added shared executor inheritance for subagents with concurrent async cells
- Removed per-session run queues from JS and Python backends, allowing async cells on the same session id to interleave.
- Introduced `getEvalSessionId` on ToolSession so subagents spawned via `task` inherit the parent's executor id and share JS VM and Python kernel state.
- Switched JS runtime state from module-level fields to AsyncLocalStorage so concurrent runs route output and tool calls to their own context.
- Changed Python runner to an asyncio event loop with per-request tasks and ContextVar-based run id tracking for concurrent execution.
- Added mtime-based module cache eviction to preserve singleton state across re-imports of unchanged local files.
2026-05-26 14:37:56 +02:00
can1357 c049613cab docs: added auth-broker, schema-normalize, install-id, and eval docs
- Added auth-broker-gateway.md covering remote OAuth vault, gateway forward-proxy, usage cache layering, and env surface.
- Added ai-schema-normalize.md documenting the unified tool-schema normalization pipeline and strict-mode edge cases.
- Added install-id.md describing the per-install UUID persistence and consumer contract.
- Rewrote eval.md to reflect structured JSON cells schema, removing the legacy `*** Cell` parser and Lark grammar.
- Updated environment-variables.md, models.md, sdk.md, secrets.md, lsp.md, session-tree-plan.md, ttsr-injection-lifecycle.md, and natives docs to match code changes.
2026-05-17 02:15:55 +02:00
can1357 16542063c4 docs(tools): added guidance note for eval tool usage in docs
- Added an explicit notice in the eval tool docs cautioning against using one-off `-c`/`-e` shell executions.
- Documented that the eval tool should be used instead for persistent runtimes, structured outputs, and cancellation or timeout support.
2026-05-15 04:20:21 +02:00
can1357 1ccc56aca6 fix(coding-agent/eval): routed JS import calls through session-aware import helper
- Updated JS import rewriting to route top-level `import` declarations through `__omp_import__` with support for import attributes.
- Added AST traversal to replace `import(...)` call callee nodes with `__omp_import__` so dynamic imports resolve via session-aware helper.
- Updated runtime `__omp_import__` to accept an optional options object and pass it through to `import(target, options)`.
2026-05-13 04:47:49 +02:00
can1357 028a442f18 feat(coding-agent): added parser support for t/rst in Cell eval
- Introduced canonical `*** Cell` headers with `t:` and `rst` attributes in eval prompts, schema, and docs.
- Updated parser and grammar to parse `*** Cell` blocks, stop on `*** End`/next header/EOF, and handle invalid `rst` with errors.
- Added quote-aware attribute tokenizers and split HTML eval parsing into `Cell` and legacy `Begin` handlers with `py` defaults.
- Expanded parsing behavior and tests for `rst` booleans, title aliases, abort boundaries, and stray-line skips between cells.
2026-05-12 10:01:56 +02:00
can1357 8d144e17ec feat(coding-agent/eval): added local python-runner subprocess execution
- Replaced Python execution with a local `python -u runner.py` subprocess and NDJSON stdin/stdout framing.
- Removed shared-gateway architecture, including coordinator lifecycle APIs, `useSharedGateway` wiring, and `jupyter` CLI/actions.
- Simplified setup checks to a plain Python 3 availability probe and removed automatic dependency-install fallbacks.
- Updated kernel cancellation and display processing to use status frames, SIGINT/SIGTERM escalation, and normalized output coercion.
- Added `python-runner` integration and display tests while deleting legacy websocket and kernel lifecycle test suites.
2026-05-12 09:09:24 +02:00
can1357 7d233725b6 refactor(coding-agent/eval): removed eval shell run helpers from JS and Python preludes
- Removed the JS and Python eval prelude `run` helpers, including their shell execution and timeout/cwd option handling.
- Updated the JS VM helper set to expose `Bun` and removed the deleted `run` entry from the prelude.
- Revised eval docs to drop `run` from the helper surface and note the new JS `Bun` global.
2026-05-12 05:43:37 +02:00
can1357 aa0d0ad4ed docs: tool behaviour 2026-05-11 00:38:35 +02:00