chore: update stale docs

This commit is contained in:
can1357
2026-08-03 16:37:05 +02:00
parent fc04aa6fa7
commit ebd5e3f86f
120 changed files with 5246 additions and 4691 deletions
+134 -233
View File
@@ -1,286 +1,187 @@
# eval
> Execute Python or JavaScript code in persistent cell-based runtimes.
> Execute one Python, JavaScript, Ruby, or Julia cell in a persistent language runtime. One tool call is one cell; state survives later calls.
> **Notice:** Do not shell out to `python -c`/`python -e`, `bun -e`, or `node -e` via the `bash` tool for ad-hoc code execution. Use this tool instead — it gives you persistent state across cells, structured `display()` output, image/JSON capture, and proper cancellation/timeout handling that one-shot `-e`/`-c` invocations cannot provide.
> **Notice:** Do not shell out to `python -c`, `ruby -e`, `julia -e`, `bun -e`, or `node -e` through `bash` for ad-hoc code. `eval` provides retained state, structured `display()` capture, tool/subagent bridges, streaming, cancellation, and artifact-backed truncation.
## Source
- Entry: `packages/coding-agent/src/tools/eval.ts`
- Entry and dynamic schema: `packages/coding-agent/src/tools/eval.ts`
- Backend enablement: `packages/coding-agent/src/tools/eval-backends.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/eval.md`
- Key collaborators:
- `packages/coding-agent/src/eval/backend.ts` — backend execution contract
- `packages/coding-agent/src/eval/agent-bridge.ts` — host-side `agent()` bridge into the subagent executor
- `packages/coding-agent/src/eval/js/executor.ts` — JS backend adapter
- `packages/coding-agent/src/eval/js/worker-core.ts` — JS execution, VM context, display/log capture
- `packages/coding-agent/src/eval/js/shared/prelude.txt` — JS global helper installer
- `packages/coding-agent/src/eval/js/shared/helpers.ts` — JS filesystem/text/env helper implementations
- `packages/coding-agent/src/eval/py/index.ts` — Python backend adapter
- `packages/coding-agent/src/eval/py/executor.ts` — kernel session retention, reset, cleanup
- `packages/coding-agent/src/eval/py/kernel.ts` — subprocess NDJSON runner protocol, display capture
- `packages/coding-agent/src/eval/py/prelude.py` — Python helper functions and status events
- `packages/coding-agent/src/session/streaming-output.ts` — truncation, artifacts, streamed chunks
- `docs/python-repl.md` — Python kernel/runner internals
- Shared contracts: `packages/coding-agent/src/eval/backend.ts`, `types.ts`, `executor-base.ts`, `kernel-base.ts`
- Host bridges: `packages/coding-agent/src/eval/agent-bridge.ts`, `completion-bridge.ts`, `concurrency-bridge.ts`, `budget-bridge.ts`
- JavaScript: `packages/coding-agent/src/eval/js/`
- Python: `packages/coding-agent/src/eval/py/`
- Ruby: `packages/coding-agent/src/eval/rb/`
- Julia: `packages/coding-agent/src/eval/jl/`
- Output/truncation: `packages/coding-agent/src/session/streaming-output.ts`
- Python internals: `docs/python-repl.md`
## Inputs
Tool parameters are a JSON object with a single `cells` field — an ordered array of cell objects. Each cell is a structured record; there is no `*** Cell` header parsing, no language sniffing, and no implicit single-cell fallback. Cells run in array order; state persists within each language across cells and across tool calls.
The params object is one cell. There is no `cells` array, header parser, language sniffing, or implicit fallback. Run incremental steps as separate tool calls; each language keeps its own state.
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `cells` | `EvalCellInput[]` | Yes | Cells executed in order. At least one cell is required (`.min(1)`). |
| `language` | `"py" \| "js" \| "rb" \| "jl"` | Yes | Explicit backend token. Normally the live schema includes only enabled runtimes; see the all-disabled edge case below. |
| `code` | `string` | Yes | Cell body, verbatim. |
| `title` | `string` | No | Short transcript label. |
| `timeout` | `number` | No | Runtime-work timeout in seconds. Default 30; `0` disables the cell timeout. Nonzero values are clamped by the tool timeout policy and `tools.maxTimeout`. |
| `reset` | `boolean` | No | Recreate this language's retained runtime before execution. Other language runtimes are untouched. Default `false`. |
Each `EvalCellInput` (from `evalCellSchema` in `packages/coding-agent/src/tools/eval.ts`):
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `language` | `"py" \| "js"` | Yes | Backend selector. `"py"` maps to the IPython-style subprocess kernel (`python` backend); `"js"` maps to the persistent JavaScript VM. |
| `code` | `string` | Yes | Cell body, verbatim. JSON-encoded — embed newlines, quotes, and indentation directly; no fences, no headers. |
| `title` | `string` | No | Short label rendered in the transcript (e.g. `"imports"`, `"load config"`). |
| `timeout` | `integer` | No | Per-cell timeout in seconds, clamped to `1..3600`. Defaults to 30 when omitted. |
| `reset` | `boolean` | No | Wipe this cell's language kernel before running. Reset is per-language: a `py` cell's reset does not touch the JS VM and vice versa. Defaults to `false`. |
Minimal example matching the live schema:
Example across three calls:
```json
{
"cells": [
{ "language": "py", "title": "imports", "timeout": 10, "code": "import json\nfrom pathlib import Path" },
{ "language": "py", "title": "load config", "code": "data = json.loads(read('package.json'))\ndisplay(data)" },
{ "language": "js", "title": "summary", "reset": true, "code": "const data = JSON.parse(await read('package.json'));\ndisplay(data);\nreturn data.name;" }
]
}
{"language":"py","title":"imports","code":"import json\nfrom pathlib import Path"}
```
```json
{"language":"py","title":"load config","code":"data = json.loads(read('package.json'))\ndisplay(data)"}
```
```json
{"language":"py","title":"reuse state","code":"display(sorted(data['dependencies']))"}
```
## Backend availability
`resolveEvalBackends(...)` combines settings with environment overrides:
| Token | Runtime | Setting/default | Environment override | Additional prerequisite |
| --- | --- | --- | --- | --- |
| `py` | retained IPython-style Python kernel | `eval.py=true` | `PI_PY` | usable configured Python interpreter/kernel |
| `js` | retained Bun worker VM | `eval.js=true` | `PI_JS` | bundled JS runtime |
| `rb` | retained Ruby kernel | `eval.rb=false` | `PI_RB` | usable `ruby.interpreter` or discovered Ruby |
| `jl` | retained Julia kernel | `eval.jl=false` | `PI_JL` | usable `julia.interpreter` or discovered Julia |
Ruby and Julia are opt-in. When at least one runtime is enabled, disabled runtimes are removed from the session-scoped wire schema and model prompt. If **all four** are disabled, the current `parameters` fallback returns the full static union even though every execution is rejected by `resolveBackend(...)`; this contradicts the nearby source comment that disabled backends never reach the model. A requested unavailable runtime raises `ToolError`; the tool never substitutes another language.
## Outputs
Final result from `EvalTool.execute()` is single-shot, but `onUpdate` streams partial text and `details` while cells run.
`execute()` returns one text content block plus any image blocks. `onUpdate` streams the active cell's output and details while it runs.
Returned shape:
- Text is stdout/stderr plus model-visible JSON `display()` values and image dimension notes.
- Image-only success reports `(displayed N image(s); no text output)`; a cell with no visible output reports `(no output)`.
- A nonzero backend exit appends `Command exited with code N`, marks the cell `error`, and sets `details.isError`.
- Cancellation returns the captured output or `Command aborted`, with `details.isError=true`.
- `content`: one text block containing combined cell output, `(displayed N image(s); no text output)` when only images exist, or `(no output)` when nothing visible was produced; image outputs are appended as additional image content blocks.
- `details` (`EvalToolDetails` from `packages/coding-agent/src/eval/types.ts`):
- `cells`: per-cell code, status (`pending`/`running`/`complete`/`error`), output, duration, exit code, status events, markdown flag
- `language`: first backend used
- `languages`: distinct backends used, in first-use order
- `jsonOutputs`: structured values emitted via `display(...)`
- `statusEvents`: aggregated helper/tool status events
- `notice`: backend fallback notice (currently unused; reserved for future per-cell notices)
- `meta`: truncation metadata
- `isError`: set on cell failure or cancellation
`EvalToolDetails`:
Renderer behavior in `packages/coding-agent/src/tools/eval.ts`:
- `cells`: a one-element `EvalCellResult[]` with `index`, `title?`, `code`, backend `language`, `output`, `status`, `durationMs?`, `exitCode?`, `statusEvents?`, and `hasMarkdown?`.
- `language`: the backend used; `languages`: the distinct backend list. These retain the historical multi-cell-compatible shape, but a current call has one backend.
- `jsonOutputs`: values captured through structured display.
- `images`: present on live updates when images have arrived; final images are content blocks.
- `statusEvents`: deduplicated helper/tool status events.
- `notice`: optional backend notice.
- `meta`: output truncation/artifact metadata supplied by `toolResult(...)`.
- `isError`: set for backend failure or cancellation.
- call preview renders each cell's `code` with syntax highlighting based on its declared `language`
- result view renders each cell separately, including status, duration, and output
- markdown outputs are rendered with the Markdown component instead of plain text
- `jsonOutputs` render as a tree, collapsed or expanded depending on UI state
- timeout / truncation notices render as dim metadata lines
- images are returned as content image blocks; live updates may also carry `details.images` while execution is in progress
The renderer merges call and result inline, syntax-highlights from the declared language, renders markdown and JSON trees specially, and shows timeout/truncation metadata. `session.allocateOutputArtifact?.("eval")` backs spilled output; `artifact://...` in `meta` reaches the full capture.
Side-channel artifacts:
## Execution flow
- `session.allocateOutputArtifact?.("eval")` may allocate an `artifact://...` backing store for spilled output.
- Truncated output metadata points at that artifact when available.
1. `EvalTool` builds a session-specific schema from enabled languages. It is essential, strict, `approval="exec"`, and `concurrency="exclusive"` within one agent session.
2. `execute()` maps `py/js/rb/jl` to `python/js/ruby/julia`, resolves availability, and wraps the single input in the renderer-compatible internal cell list.
3. It obtains the retained executor id from `session.getEvalSessionId?.()` or `defaultEvalSessionId(session)`, allocates the output sink/artifact, and registers the run through `trackEvalExecution?.(...)`.
4. The timeout defaults to 30 seconds. `0` creates no watchdog. Otherwise `IdleTimeout` is combined with tool and session abort signals.
5. `agent()`, `parallel()`, and `completion()` emit pause/resume status operations: time spent in those host bridges does not consume the cell's runtime-work budget. Compute, output, status helpers, and ordinary `tool.*` calls do consume it.
6. The selected backend receives cwd, retained session id, session file, kernel owner, reset flag, callbacks, and cancellation signal.
7. Output chunks stream into an artifact-aware `OutputSink` and live tail. Rich displays are separated into JSON, image, markdown, and status channels.
8. Success, nonzero exit, and cancellation are assembled into the result shapes above. The output sink is finalized even when execution fails.
## Flow
## Runtime behavior
1. `EvalTool.execute()` in `packages/coding-agent/src/tools/eval.ts` receives `params.cells` already validated by the Zod schema — no string parsing step.
2. For each cell, `execute()` maps `cell.language` to an `EvalLanguage` (`"py"` → `"python"`, `"js"` → `"js"`) and calls `resolveBackend(session, language)`:
- `python` is gated on `resolveEvalBackends(session).python` (the `eval.py` setting, overridden by the `PI_PY` env flag) and `pythonBackend.isAvailable(session)`.
- `js` is gated on `resolveEvalBackends(session).js` (the `eval.js` setting, overridden by the `PI_JS` env flag).
- A disabled or unavailable requested backend throws `ToolError`; there is no auto-fallback or sniffing.
3. The tool allocates an `OutputSink`, a `TailBuffer`, per-cell result objects, and a `sessionAbortController`. `session.trackEvalExecution?.(...)` can wrap the whole run for external cancellation tracking.
4. It resolves the executor session id from `session.getEvalSessionId?.()`, falling back to `defaultEvalSessionId(session)`. Subagents inherit the parent's id so both sides share the same JS VM and Python kernel for each backend.
5. Cells execute sequentially within one eval tool call. For each cell, `execute()`:
- clamps `cell.timeout ?? 30` seconds through `clampTimeout("eval", ...)`
- wraps the clamped budget in an `IdleTimeout` and combines its signal with the tool signal and the session abort controller (`AbortSignal.any`). The per-cell `timeout` is a runtime-work budget, not a wall clock: `EVAL_TIMEOUT_PAUSE_OP`/`EVAL_TIMEOUT_RESUME_OP` status events pause and resume the idle timer so host-side `agent()`/`parallel()`/`completion()` calls do not spend it
- marks the cell `running` and emits an update
- calls the backend's `execute()` with `cwd`, `sessionId`, `sessionFile`, `kernelOwnerId`, `session`, `idleTimeoutMs`, `reset` (defaults to `false`), the combined signal, and chunk/status callbacks
6. JS cells dispatch through `packages/coding-agent/src/eval/js/index.ts` into `executeJs()`; Python cells dispatch through `packages/coding-agent/src/eval/py/index.ts` into `executePython()`.
7. Backend text chunks stream into the shared `OutputSink`; rich outputs are accumulated separately as JSON, images, markdown markers, and status events.
8. After each cell:
- text output is trimmed and stored on that cell result
- multi-cell runs prefix text with `[i/n]` and the optional title
- cancellations return early with `isError: true` and a cell-specific abort message
- non-zero exit codes return early with `isError: true` and a message naming the failed cell
- later cells are skipped after the first error, but earlier cell state persists in the underlying runtime
9. On success, the tool joins all cell outputs, synthesizes `(no text output)` or `(no output)` when needed, and attaches truncation metadata from `summarizeFinal()`.
10. The renderer uses `details.cells`, `details.jsonOutputs`, and `details.statusEvents` to build notebook-style output. `mergeCallAndResult = true` and `inline = true`, so call and result render together in the transcript.
### JavaScript (`js`)
## Modes / Variants
- Persistent worker VM keyed by `js:${sessionId}`; `reset` recreates the VM and is destructive to concurrent users of that session id.
- Runs under Bun and exposes host globals including `Bun`, `Buffer`, `fetch`, `process`, `require`, `createRequire`, `fs`, and Web Crypto.
- Top-level `await` and bare `return` work through async wrapping.
- Static top-level imports and dynamic imports are rewritten through the local module loader. Local filesystem imports are cache-busted between cells; bare package and scheme/URL imports retain normal cache identity.
- Awaited regions can interleave with another session sharing the executor; synchronous code still blocks the worker event loop.
### Backend selection
### Python (`py`)
Backend choice is **explicit per cell** — there is no auto-detection.
- Retained kernels are keyed by `python:${sessionId}`, normalized cwd, and interpreter. `python.kernelMode="per-call"` instead creates and shuts down a fresh kernel for each invocation.
- The runner uses one persistent asyncio event loop, so top-level `await` works; `asyncio.run(...)` is invalid there.
- MIME frames support status, PNG, JSON, markdown, plain text, and HTML-to-markdown conversion.
- Interactive stdin is rejected with `Kernel requested stdin; interactive input is not supported.`
- Synchronous blocks use the default executor with copied ContextVars; Python bytecode still contends on the GIL.
- `language: "py"` → Python (IPython-style subprocess kernel) backend
- `language: "js"` → JavaScript VM backend
### Ruby (`rb`)
If the requested backend is disabled or unavailable, the tool throws `ToolError` for that cell. The caller chooses; the tool does not silently substitute.
- Retained kernels are keyed by `ruby:${sessionId}`, normalized cwd, and interpreter.
- Cells evaluate in persistent `TOPLEVEL_BINDING`; locals, methods, and constants survive. A trailing value is displayed like IRB unless it is nil, an assignment, or a definition.
- Rich display supports the OMP MIME convention and IRuby-compatible MIME hooks, using the shared kernel display pipeline.
- `reset` replaces the retained Ruby kernel.
### JavaScript runtime
### Julia (`jl`)
Implemented in `packages/coding-agent/src/eval/js/worker-core.ts`, `packages/coding-agent/src/eval/js/shared/prelude.txt`, and `packages/coding-agent/src/eval/js/shared/helpers.ts`.
- Retained kernels are keyed by `julia:${sessionId}`, normalized cwd, and interpreter.
- Cells evaluate in persistent `Main`; a value-bearing trailing expression is displayed unless suppressed by statement form.
- Julia's display stack is bridged into the same MIME/status pipeline.
- `reset` replaces the retained Julia kernel.
- Persistent worker-backed VM sessions keyed by `js:${sessionId}`
- `reset: true` calls `resetVmContext(sessionKey)` before the cell executes; reset is destructive for all live runs on that JS session
- Top-level `await` and bare `return` are supported by wrapping code in an async IIFE when `wrapCode()` sees `await` or `return`
- Top-level static `import ... from ...` and dynamic `import(...)` calls are routed through `rewriteImports()`, which sends them via `__omp_import__` so the specifier resolves against the session cwd. Dynamic-import call sites are swapped for a guarded shim (`typeof __omp_import__ === "function" ? __omp_import__ : (s, o) => import(s, o)`) rather than the bare helper identifier: functions handed to puppeteer (`tab.evaluate`, `page.evaluate`, ...) are serialized with `Function.prototype.toString()` and re-evaluated inside the browser page, where the worker-injected helper does not exist, so the shim falls back to native dynamic import there
- Module cache is busted for **local** imports between cells so edits to source files are picked up without restarting the runtime. `__omp_import__` deletes `require.cache[absPath]` before re-importing whenever the original specifier is a filesystem path: relative (`./x`, `../x`, `.`, `..`), POSIX-absolute (`/...`), home-prefixed (`~/...`), or Windows drive-letter (`C:\...` / `C:/...`). Bare specifiers (`react`, `lodash/x`) and URL/scheme specifiers (`node:fs`, `file://...`, `https://...`) are left in cache so package identity stays stable across cells. The cache-bust only fires when the resolved target is an absolute path — unresolved bare-package fallbacks (`resolveImportSpecifier()` returning the original specifier) skip it.
- The prelude installs globals:
- `display`, `print`, and a `console` bridge
- `read`, `write`, `env`, `output`
- `tool.<name>(args)` proxy for arbitrary session tool calls
- `completion(prompt, opts?)` for oneshot, stateless model calls (see _Oneshot completion helper_ below)
- `agent(prompt, opts?)` for a single subagent call, plus `parallel()` / `pipeline()` bounded-pool helpers (see _Subagent helper_ below)
- `log(message)`, `phase(title)`, and `budget` (live token-budget view via async `budget.total()` / `budget.spent()` / `budget.remaining()` / `budget.hard()`)
- JS host/runtime helpers (`read`, `write`, `output`) are async and `await`able; `env` returns synchronously.
- JS helper options may be passed either positionally in the Python order or as a trailing options object. `null` and `undefined` skip positional slots:
- `await read(path, offset?, limit?)` or `await read(path, { offset?, limit? })`
- `await agent(prompt, agent?, model?, label?, schema?)` or `await agent(prompt, { agent?, model?, label?, schema?, handle? })`
- `await parallel([() => agent("a"), () => agent("b")])`
- `await pipeline(items, stage1, stage2)`
- `display(value)` behavior:
- plain objects/arrays become JSON outputs
- `{ type: "image", data, mimeType }` becomes an image output
- scalars become text
- The VM runs in the host worker's global scope: user code gets the worker's real `process` (intentionally not subsetted — subsetting it segfaulted alongside puppeteer/worker_threads), the injected `fs`, `require`, `createRequire`, and `webcrypto`, plus host globals like `Buffer`, `fetch`, `Blob`, `File`, `Headers`, `Request`, and `Response`
- Concurrent runs on the same VM are not queued end-to-end. Synchronous JS still runs on the single event loop; awaited regions can interleave with sibling runs.
## Prelude helpers
### Python runtime
All enabled runtimes expose equivalent helpers where the language permits:
Implemented in `packages/coding-agent/src/eval/py/executor.ts`, `packages/coding-agent/src/eval/py/kernel.ts`, and `packages/coding-agent/src/eval/py/prelude.py`. See `docs/python-repl.md` for kernel and runner details.
- `display(value)`, `print(...)`
- `read(path, offset?, limit?)`, `write(path, content)`, `env(...)`, `output(...)`
- `tool.<name>(args)` for a normal session tool call
- `completion(...)`, `agent(...)`, `parallel(...)`, `pipeline(...)`
- `log(message)`, `phase(title)`, `budget`
- Default mode is retained `session` kernels keyed by `python:${sessionId}` plus normalized cwd and interpreter
- Optional `python.kernelMode = "per-call"` creates a fresh kernel for each cell and shuts it down afterward
- `reset: true` disposes the retained kernel for that session before the cell runs; later Python cells in the same tool call reuse the fresh kernel
- Startup path:
- availability check
- create/connect kernel
- initialize cwd / env / `sys.path`
- execute `PYTHON_PRELUDE`
- Python cells run in the runner's persistent asyncio event loop, so top-level `await` works; the prompt warns not to use `asyncio.run(...)`
- The Python prelude defines helpers with the same surface as JS where practical, including `tool.<name>(args)`, `completion(...)`, and `agent(...)` through a per-run loopback bridge
- Synchronous statement blocks run in the default executor with ContextVar state copied in; the GIL still serializes bytecode execution, but awaited regions can interleave with sibling cells
- Kernel `display` / `result` frames map to:
- `application/x-omp-status` → status event
- `image/png` → image output
- `application/json` → JSON output
- `text/markdown` → markdown output
- `text/plain` → text output
- `text/html` → HTML converted to markdown with `htmlToBasicMarkdown()`
- Interactive stdin is rejected: a stdin-flagged result returns exit code `1` with `Kernel requested stdin; interactive input is not supported.`
JS filesystem/bridge helpers are asynchronous; Python, Ruby, and Julia helpers are synchronous. `read()` delegates non-`local://` schemes to the registered read tool, resolves `local://` through injected roots, and reads regular paths relative to cwd. `write()` accepts regular and `local://` paths but rejects other protocol URLs.
### Oneshot completion helper (`completion`)
`display()` captures JSON-compatible structures, images, markdown, or text according to the backend. Ruby and Julia additionally auto-display eligible final expressions.
Both runtimes expose `completion()` — a single stateless completion against a model tier. It is intentionally minimal: no conversation history, no agent-visible tools, pure text in / text (or object) out. Implemented host-side in `packages/coding-agent/src/eval/completion-bridge.ts` and routed through the existing tool bridge under the reserved name `__completion__`.
### `completion()`
- Signatures:
- JS: `await completion(prompt, { model?, system?, schema? })`
- Python: `completion(prompt, *, model="default", system=None, schema=None)`
- `model` selects a tier (default `"default"`):
- `"smol"` → `@smol` role (fast / cheap)
- `"default"` → the session's active model, falling back to the `@default` role
- `"slow"` → `@slow` role; requests high reasoning effort only on reasoning-capable models
- `system` (optional) supplies a system prompt.
- `schema` (optional) is a plain JSON-Schema object. When present, the model is forced to call a single synthetic `respond` tool with that schema (loose, non-strict), and the helper returns the parsed object. When absent, the helper returns the completion string.
- Errors surface as exceptions: unresolved tier, missing API key, an `error`/`aborted` stop reason, or empty output each raise.
A stateless, tool-free one-shot model call:
### Subagent helper (`agent`)
- JS: `await completion(prompt, { model?, system?, schema? })`
- Python/Ruby/Julia: keyword form with `model`, `system`, and `schema`
- `model`: `"smol"`, `"default"`, or `"slow"` tier; default is the active/default tier.
- `schema`: JSON Schema for a synthetic `respond` tool; successful structured calls return parsed data.
- Unresolved tier, missing credentials, error/abort stop, empty output, and invalid structured output raise into the cell.
Both runtimes expose `agent()` — a single subagent invocation routed through `packages/coding-agent/src/eval/agent-bridge.ts` into the same `runSubprocess(...)` path used by the `task` tool. It uses the current eval session's spawn policy and inherits the parent eval executor id, so parent and subagent code share JS/Python runtime state.
### `agent()`
- Signatures:
- JS: `await agent(prompt, agent?, model?, label?, schema?)` or `await agent(prompt, { agent?, model?, label?, schema?, handle? })`
- Python: `agent(prompt, *, agent="task", model=None, label=None, schema=None, handle=False)`
- `agent` defaults to the bundled `task` agent and resolves through normal agent discovery, so project and user agents work.
- `model` (optional) pins an exact per-call model selector or fallback chain (`string | string[]`) for the subagent, overriding the agent's configured model; omit it to use the agent's default. This route is eval-specific: the `task` tool's own wire schema no longer exposes a per-call `model` (see the 17.1.2 changelog "Removed" note), but the eval `agent()` bridge still accepts and forwards it (`model?` in `agentArgsSchema`, `packages/coding-agent/src/eval/agent-bridge.ts`).
- Shared background is passed via files: write a `local://` file and reference it in the prompt. `label` controls the `agent://<id>` output label prefix.
- `schema` passes a JSON Schema to the subagent structured-output path. When present, the helper parses the final JSON text and returns an object.
- `handle` (default off) returns a DAG node dict — `{ text, output, handle: "agent://<id>", id, agent }`, plus a parsed `data` field when `schema` is set — instead of the bare output, so a downstream stage can reference the transcript by handle.
- Spawn restrictions use `session.getSessionSpawns()` exactly like the `task` tool. Eval-driven subagent recursion is capped at depth 3.
- JS and Python both expose `parallel(thunks)` and `pipeline(items, ...stages)`; both use a bounded async/threaded pool whose width tracks the `task.maxConcurrency` setting (the same ceiling the `task` tool uses; `0` = run every item at once), preserve item order, and propagate rejections. The width is fetched live from the host via the `__concurrency__` bridge, so the helpers no longer take a `concurrency` argument.
- Errors surface as exceptions: unknown or disabled agent, disallowed spawn, recursion cap, subagent failure, or invalid structured output all fail the eval cell.
Runs one subagent through `runStructuredSubagent(...)`:
### Multi-language call behavior
- JS supports the preferred `await agent(prompt, { agent?, model?, label?, schema?, schemaMode?, isolated?, apply?, merge?, handle? })`; legacy positional slots are still implemented.
- Python/Ruby/Julia use keyword arguments (`schema_mode` outside JS).
- `agent` defaults from the current spawn policy. `model` may pin a selector/fallback chain. `schema` overrides agent/session schemas; `schemaMode`/`schema_mode` chooses `permissive` or `strict`.
- `isolated` requests isolation. `apply` controls whether captured changes are integrated; `merge=false` selects patch mode while the normal setting controls branch mode.
- `handle=true` returns `{ text, output, handle, id, agent }`, optional parsed `data`, and isolation metadata instead of only output/data.
- Eval subagents are one-shot (`keepAlive=false`), are unregistered/disposed after completion, and **do not share the caller's eval executor** (`shareEvalSession=false`). Their code mutations therefore do not appear in the caller's retained VM/kernel.
- Spawn policy, discovered-agent availability, task-depth limit 3, hard turn budget, subagent failure, strict schema failure, and isolation-apply failure are enforced as cell errors.
A single tool call can mix Python and JS cells. Persistence is per language runtime:
`parallel(thunks)` runs zero-argument callables in a bounded pool and preserves input order. `pipeline(items, ...stages)` applies each stage as a barriered wave. Pool width is read live from `task.maxConcurrency`; `0` means all items at once. The lowest-index failure is propagated.
- `reset: true` on a Python cell does not touch JS state
- `reset: true` on a JS cell does not touch Python state
- each backend keeps its own retained session keyed from the same session-derived ID
## Side effects and cancellation
## Side Effects
- Prelude helpers may read/write files and call arbitrary registered tools; JS exposes network-capable `fetch`.
- Python, Ruby, and Julia use retained subprocess kernels speaking framed local IPC. JavaScript uses a worker VM.
- Retained runtimes survive calls until reset, owner cleanup, or process exit.
- Cancellation is destructive when needed: JS terminates its worker; managed kernels interrupt and may escalate to shutdown. A reset is likewise destructive to concurrent work sharing that backend session.
- Eval-driven `agent()` may run tools and isolated workspaces, but its child is disposed rather than retained for hub follow-up.
- Filesystem
- JS/Python prelude helpers can read and write filesystem paths under the session cwd or absolute paths.
- JS helper `read()` auto-delegates any non-`local://` scheme URI (`agent://`, `artifact://`, `https://`, ...) to `tool.read(...)` (honoring an `offset`/`limit` line selector), resolves `local://` under its mapped root, reads plain/absolute filesystem paths directly, and rejects directory paths.
- Output may spill to an artifact file via `OutputSink`.
- Network
- Python backend speaks NDJSON to a local `python3` subprocess over stdin/stdout (no network).
- JS runtime exposes `fetch` and `tool.<name>()`; those tools may perform additional network I/O.
- Subprocesses / native bindings
- Python availability check runs `<python> -c ...`.
- Python backend spawns one `python -u runner.py` subprocess per kernel; cancellation sends `SIGINT`. Details in `docs/python-repl.md`.
- `agent()` runs one in-process subagent via the task executor; that subagent may use its configured tools.
- Session state
- `session.assertEvalExecutionAllowed?.()` can block execution.
- `session.trackEvalExecution?.(...)` can register cancellable eval work.
- `session.getSessionFile?.()`, `session.getEvalSessionId?.()`, and `session.getEvalKernelOwnerId?.()` influence VM/kernel reuse and artifact lookup.
- JS VM contexts persist across eval calls until reset/disposal.
- Python retained kernels persist until reset, owner cleanup, or process exit.
- `agent()` allocates `agent://<id>` output artifacts and reuses the parent's eval executor id.
- User-visible prompts / interactive UI
- none; stdin requests are rejected programmatically
- Background work / cancellation
- Python retained kernels have heartbeat and idle cleanup timers.
- Cancellation hard-kills/resets the shared executor for that backend: JS terminates the worker, Python sends SIGINT and may escalate to subprocess shutdown.
## Limits and errors
## Limits & Caps
- Per-cell timeout default: 30s (applied when `timeout` is omitted in `EvalTool.execute()`; clamped through `TOOL_TIMEOUTS.eval.default` in `packages/coding-agent/src/tools/tool-timeouts.ts`)
- Schema-level `timeout` range: integer `1..3600` seconds (enforced by Zod on the cell schema)
- Timeout clamp at runtime: 1s minimum, 3600s maximum (`TOOL_TIMEOUTS.eval` in `packages/coding-agent/src/tools/tool-timeouts.ts`)
- Transcript code/output preview: 10 lines by default (`EVAL_DEFAULT_PREVIEW_LINES` in `packages/coding-agent/src/tools/eval-render.ts`, re-exported from `eval.ts`)
- Output truncation window: 50KB default (`DEFAULT_MAX_BYTES` in `packages/coding-agent/src/session/streaming-output.ts`)
- Output line cap inside truncation helpers: 3000 lines (`DEFAULT_MAX_LINES` in `packages/coding-agent/src/session/streaming-output.ts`)
- Streaming tail buffer for live updates: `DEFAULT_MAX_BYTES * 2` = 100KB (`packages/coding-agent/src/tools/eval.ts`)
- JS/Python `parallel()` / `pipeline()` helper pool width: the `task.maxConcurrency` setting (default 32; `0` = unbounded), resolved live via the `__concurrency__` bridge (`packages/coding-agent/src/eval/concurrency-bridge.ts`)
- Eval-driven `agent()` recursion cap: task depth 3 (`EVAL_AGENT_MAX_DEPTH`)
- Python kernel startup wait: 10s (`STARTUP_TIMEOUT_MS` in `packages/coding-agent/src/eval/py/kernel.ts`)
- Python kernel shutdown grace per escalation step (`exit` request → `SIGTERM` → `SIGKILL`): 1000ms (`SHUTDOWN_GRACE_MS` in `packages/coding-agent/src/eval/py/kernel.ts`)
- Python SIGINT escalation window: 5s without a `done` frame before the subprocess is killed (`INTERRUPT_ESCALATION_MS` in `packages/coding-agent/src/eval/py/kernel.ts`)
- Python auto-restart budget: a dead retained kernel is replaced and the cell retried once per execution (`executeOnSession` in `packages/coding-agent/src/eval/py/executor.ts`)
## Errors
- Zod validation rejects malformed `cells` arrays before `execute()` runs (missing `language`/`code`, out-of-range `timeout`, empty `cells`).
- Missing session without proxy executor throws `ToolError("Eval tool requires a session when not using proxy executor")`.
- Disabled/unavailable backends throw `ToolError` from `resolveBackend()`:
- `eval.py = false` (or `PI_PY=0`) and a `py` cell is requested
- `eval.js = false` (or `PI_JS=0`) and a `js` cell is requested
- Python kernel unavailable and a `py` cell is requested
- JS runtime exceptions are converted into text output plus `exitCode: 1`; cancellations return `cancelled: true` and may append `Command timed out`.
- Python execution errors from the kernel become text output and `exitCode: 1`; later cells are skipped.
- Python stdin requests are treated as errors with the message `Kernel requested stdin; interactive input is not supported.`
- Cancellation is returned, not thrown, once backend execution has started. The tool formats it as a cell failure and sets `details.isError = true`.
- If output truncates, the tool still succeeds; truncation is surfaced through `details.meta` and artifact-backed full output when available.
## Shared executor trade-offs
- Parent agents and subagents share eval state bidirectionally when a subagent inherits the parent's executor id. Mutations in either direction are visible to the other participant.
- Async regions of concurrent runs can interleave. Synchronous JS still blocks the VM event loop; synchronous Python still contends on the GIL.
- Cancelling one run is destructive to the shared backend executor. This is intentional: JS worker termination and Python SIGINT/subprocess shutdown are the only reliable way to interrupt arbitrary user code.
- `reset: true` is destructive for every live run on that backend session id. Concurrent Python resets coalesce — a reset already in flight is awaited rather than duplicated, and runs queued behind it proceed on the freshly-restarted kernel.
- Default timeout: 30 seconds; `0` disables. Nonzero timeouts are clamped through `clampTimeout("eval", ..., tools.maxTimeout)`.
- Output sink default window: 50 KiB (`DEFAULT_MAX_BYTES`); live tail: 100 KiB; truncation helpers cap at 3000 lines.
- Each JSON display value included in model-visible text is capped at 8000 characters; the full structured value remains in `jsonOutputs`.
- Transcript preview defaults to 10 lines.
- Eval subagent recursion cap: 3. Helper fan-out uses `task.maxConcurrency` (default 32, `0` unbounded).
- Malformed params are schema errors; unavailable/disabled backends and missing session are `ToolError`s.
- Runtime exceptions become backend output with nonzero exit. Interactive stdin is an error. Output truncation does not fail the call.
- A dead retained managed kernel may be replaced and the invocation retried once by its executor.
## Notes
- Backend selection is strictly explicit per cell: `language` must be `"py"` or `"js"`. The previous `*** Cell` header parser, the `eval.lark` constrained grammar, and the sniffer-based fallback have all been removed.
- `EvalTool.customFormat` no longer exists. Tool calls flow through the standard JSON schema; there is no Lark-constrained sampling path.
- `tool.<name>()` exists in both JS and Python. Python calls route through a per-run loopback bridge keyed by the current cell id.
- `read()` delegates non-`local://` scheme URIs to `tool.read`, resolves `local://` under its injected root, and resolves plain paths against the session cwd or an absolute filesystem path; `resolveRegularFile()` rejects directory paths. `write()` accepts `local://` and plain paths but rejects any other `scheme://` via `resolveHelperPath()` (`Protocol paths are not supported by write()`).
- Python helper `output(...)` depends on `PI_ARTIFACTS_DIR` or `PI_SESSION_FILE`; it fails outside a session-backed run.
- `display()` can produce text and structured outputs from the same value; the renderer prefers markdown over `text/plain` when both exist.
- JS static imports are rewritten only at top level. Nested imports stay invalid and surface normal JS syntax/runtime errors.
- `EvalTool` is `concurrency = "exclusive"` within one agent session, but parent and subagent sessions can run eval concurrently when they share an inherited executor id.
- The tool description shown to the model is templated by backend availability (`getEvalToolDescription()`); if Python is unavailable, the prompt omits Python-specific instructions.
- One call is one cell. Use separate calls to exploit persistence and rerun only the failed step.
- State is isolated by language; resetting Python does not reset JS, Ruby, or Julia.
- Current schema tokens are only `py`, `js`, `rb`, and `jl`; long language names are renderer/approval formatting aliases, not wire values.
- The former multi-cell `cells` payload, `*** Cell` parser, sniffing fallback, and constrained `eval.lark` grammar are removed.
- Parent and ordinary task subagents may share an inherited eval executor id; children created by eval's own `agent()` explicitly do not.