diff --git a/docs/ai-schema-normalize.md b/docs/ai-schema-normalize.md new file mode 100644 index 000000000..a1a2e1e6b --- /dev/null +++ b/docs/ai-schema-normalize.md @@ -0,0 +1,171 @@ +# AI tool-schema normalization + +`@oh-my-pi/pi-ai` exposes one unified schema normalizer that providers consume +before tools are sent on the wire. All walkers live in +`packages/ai/src/utils/schema/normalize.ts`; the operational contract is +`packages/ai/src/utils/schema/CONSTRAINTS.md`. + +There is no separate `strict-mode.ts` module any more — OpenAI strict-mode +sanitization, OpenAI Responses `oneOf` rewriting, Google/Vertex/Gemini-CLI +sanitization, Cloud Code Assist Claude sanitization, and MCP sanitization all +share the same option-driven walk. + +## Entry points + +All exports live under `@oh-my-pi/pi-ai/utils/schema`: + +- `normalizeSchema(value, options)` — generic option-driven walker. +- `normalizeSchemaForGoogle(value)` — Gemini / Vertex / Gemini CLI. +- `normalizeSchemaForCCA(value)` — Cloud Code Assist Claude (Antigravity + GCA). +- `normalizeSchemaForMCP(value)` — MCP inputSchemas before they enter the + custom-tool registry. `tool-bridge.ts` runs every MCP `inputSchema` through + this dispatcher. +- `normalizeSchemaForOpenAIResponses(schema)` (alias + `sanitizeSchemaForOpenAIResponses`) — rewrites `oneOf` → `anyOf` for the + Responses family. +- `sanitizeSchemaForStrictMode(schema)` and + `enforceStrictSchema(schema)` / `tryEnforceStrictSchema(schema)` — the + OpenAI strict-mode pipeline (sanitize → enforce). All three are exported + from `normalize.ts`. +- `adaptSchemaForStrict(schema, strict)` from `./adapt` — thin composer that + wraps `tryEnforceStrictSchema` for provider call sites and consults + `PI_NO_STRICT` (env `PI_NO_STRICT`) for the global bypass. + +Removed in the unified-flow refactor: + +- `strict-mode.ts` (merged into `normalize.ts`). +- `sanitize-google.ts` and `normalize-cca.ts` (replaced by + `normalizeSchemaFor*` dispatchers). +- `StringEnum` helper — use `z.enum([...])` directly; Zod's emitted JSON + Schema is already wire-compatible with Google and other providers. +- `sanitizeSchemaFor{Google,CCA,MCP}` / `prepareSchemaForCCA` — renamed to + `normalizeSchemaFor{Google,CCA,MCP}`. + +## Dispatcher mapping + +| Provider transport(s) | Dispatcher | +| -------------------------------------------------------------------- | -------------------------------------------- | +| `openai-completions`, `openai-responses`, `openai-codex-responses` | `adaptSchemaForStrict` (sanitize + enforce) | +| `openai-responses` family (`oneOf` → `anyOf` only) | `normalizeSchemaForOpenAIResponses` | +| `google-generative-ai`, `google-vertex`, Gemini CLI | `normalizeSchemaForGoogle` | +| Cloud Code Assist Claude (Antigravity + GCA, `claude-*` model ids) | `normalizeSchemaForCCA` | +| MCP `inputSchema` ingestion | `normalizeSchemaForMCP` | +| `anthropic-messages` (native, not CCA) | per-provider whitelist in `anthropic.ts` | + +Gemini CLI / Antigravity CCA MUST run the full `normalizeSchemaForCCA` +pipeline (not just the first keyword-stripping pass) to keep parity with the +shared Google Claude path. + +## Walk semantics + +`normalizeSchema` first upgrades the input to JSON Schema 2020-12, then +walks the tree with the option set pinned by the dispatcher. Each node: + +1. Inlines `$ref` (see "Edge cases" below). +2. Renames `snake_case` combinator/property keys to camelCase + (`any_of` → `anyOf`, etc.; collisions follow python-genai + `pop(from)`/`set(to)` semantics — snake_case wins). +3. Applies the `handle_null_fields` collapse for nullable unions before + recursing into children. +4. Strips keys the target provider does not support, optionally lifting + human-meaningful keys (`pattern`, `format`, min/max, `default`, + `examples`, ...) into the sibling `description` via the spill formatter + (`spill.ts`). Structural/meta keys (`$ref`, `$defs`, + `additionalProperties`) are not spilled. +5. Normalizes type unions (`type: ["T", "null"]` → `type: "T"` + nullable + marker on Google, plain `type: "T"` on CCA). +6. Collapses object-only / same-type combiners, optionally lossy-collapses + mixed-type combiners (CCA only), and runs the residual-combiner fixpoint. +7. Validates against AJV 2020 when `validateAndFallback` is set (CCA path) + and emits the per-tool fallback `{ "type": "object", "properties": {} }` + on residual incompatibility — `type` array, `type: "null"`, `nullable` + key, or any remaining `anyOf`/`oneOf`/`allOf`. + +## OpenAI strict-mode pipeline + +`adaptSchemaForStrict(schema, strict)` runs `tryEnforceStrictSchema`, +which composes: + +1. **Sanitize** (`sanitizeSchemaForStrictMode`): strips non-structural + keywords (`format`, `pattern`, min/max, `examples`, `default`, + `if`/`then`/`else`, `not`, `unevaluated*`, `patternProperties`, + `dependent*`, `content*`, `min/maxProperties`, `$dynamicRef`, etc.). The + `default` value is inlined into the sibling `description` as + ` (default: X)` before being dropped, unless `description` already + contains `(default:` or no `description` exists. +2. **Enforce** (`enforceStrictSchema`): every object node gets + `additionalProperties: false`, every property goes into `required`, and + optional properties become nullable unions + (`anyOf: [, { "type": "null" }]`). Tuple `prefixItems` are + strictified recursively. + +The two passes share node-level caches and the same epoch-based cycle +guard, so a single walk on the wire path normalizes refs, allOf, and +nullable wrapping consistently. `tryEnforceStrictSchema` is fail-open: +if anything throws, it returns `{ strict: false, schema: original }` so +callers MUST emit `strict: true` only when enforcement actually succeeded. + +### Edge cases the strict-mode normalizer handles + +- **Local `$ref` inlining.** OpenAI strict mode rejects + `{ "$ref": "...", "description": "..." }` with sibling keys. The + sanitizer pre-resolves local `#/...` refs against the root and merges + with **sibling keys winning** over the resolved def — same precedence + as `openai-python`'s `_ensure_strict_json_schema`. Recursive refs are + guarded by the per-walk epoch. +- **Single-item `allOf`.** A `{ "allOf": [X], ...siblings }` collapses to + `{ ...X, ...siblings }` with the inlined entry's keys winning over the + original siblings (matches `openai-python`'s `_pydantic.py:79-83`). Multi- + item `allOf` is left intact for the downstream validator to reject if + needed. +- **Type-array branches and nullable unions.** When a node has + `type: ["T", "U"]`, the sanitizer emits one variant schema per type, + pruning type-specific keywords (e.g. `properties`/`required` only stay on + the `object` variant, `items` only on the `array` variant). The shared + `description` is **hoisted onto the `anyOf` wrapper** instead of being + duplicated on every branch — so a strict nullable union becomes + `{ anyOf: [T, { type: "null" }], description: "..." }`, not + `anyOf: [{ ..., description }, { ..., description }]`. +- **Enum/const without a `type`.** Both sanitize and enforce paths call + `inferStrictPrimitiveTypeFromEnumOrConst` to infer the primitive `type` + from `enum` / `const` values. Mixed-primitive enums (`[1, "two", null]`), + enums containing objects/arrays, and non-primitive `const` values + (`{a:1}`, `[1,2,3]`) cannot be described by a single `type` keyword and + trigger the strict-mode fail-open path — emitting a typeless schema + would just be rejected on the wire by OpenAI. + +## Performance: static fingerprint cache + +`resolveProviderModels` in `packages/ai/src/model-manager.ts` and +`readModelCache`/`writeModelCache` in `model-cache.ts` cooperate via a +schema-v3 `static_fingerprint` column on the `model_cache` SQLite table. + +- `fingerprintStatic(staticModels)` hashes the static catalog slice + (`Bun.hash(JSON.stringify(models))` in base36) and memoizes the result + in a per-process `WeakMap` keyed by the array reference. Multiple + cold-start arms calling `resolveProviderModels` with the same + `staticModels` array pay the JSON+hash cost once. +- On cache read, if the network fetch is being skipped, the cached row is + fresh + authoritative, and the cached `static_fingerprint` matches the + current one, `resolveProviderModels` returns the cached models verbatim + — the cache already incorporates the same static state, so re-running + `mergeDynamicModels(static, cache)` would just rebuild the same objects. +- `mergeModelSources` and `mergeDynamicModels` short-circuit on + empty-source inputs (the common shape after `(static, [])` or for + providers without a static catalog), avoiding Map churn entirely. + +Cache rows written before schema v3 are dropped by the cache-version +check; the column defaults to `''` for any row that survives a version +upgrade so the fingerprint-equality check naturally fails closed and the +full merge re-runs. + +## Related + +- `docs/models.md` — registry, equivalence, compat flags + (`supportsStrictMode`, `toolStrictMode`, `disableStrictTools`). +- `docs/provider-streaming-internals.md` — how the normalized schemas are + used downstream during the provider stream loop. +- `docs/mcp-server-tool-authoring.md` — MCP `inputSchema` ingestion via + `normalizeSchemaForMCP`. +- `packages/ai/src/utils/schema/CONSTRAINTS.md` — operational contract for + every normalization rule. diff --git a/docs/auth-broker-gateway.md b/docs/auth-broker-gateway.md new file mode 100644 index 000000000..07aabda59 --- /dev/null +++ b/docs/auth-broker-gateway.md @@ -0,0 +1,181 @@ +# Auth Broker and Auth Gateway + +The auth broker and auth gateway are two cooperating HTTP services that move OAuth refresh tokens and provider access tokens off developer laptops and into a single broker host. + +- **`omp auth-broker serve`** holds the canonical SQLite credential vault, performs OAuth refreshes, and exposes a small REST API (`/v1/snapshot`, `/v1/credential/:id/refresh`, `/v1/credential/:id/disable`, `/v1/credential`, `/v1/usage`, `/v1/healthz`). +- **`omp auth-gateway serve`** is a forward-proxy. It accepts OpenAI Chat Completions, Anthropic Messages, and OpenAI Responses requests, injects the broker-resolved access token, and forwards the bytes to the real provider. Clients (containerised omp, llm-git, the macOS usage widget, …) never see the access token. + +Transport security between operator, broker, and gateway is delegated to the operator (Tailscale / Wireguard / reverse proxy + TLS). Every endpoint except `/v1/healthz` (broker) and `/healthz` (gateway) requires a bearer token. + +Source: `packages/ai/src/auth-broker/`, `packages/ai/src/auth-gateway/`, `packages/coding-agent/src/cli/auth-broker-cli.ts`, `packages/coding-agent/src/cli/auth-gateway-cli.ts`, `packages/coding-agent/src/session/auth-broker-config.ts`. + +## Data flow + +``` + ┌────────────────────────────────────────────────────────────┐ + │ broker host │ + │ │ + developer ──▶ │ ┌──────────────────────────┐ ┌────────────────────┐ │ + laptop / │ │ omp auth-broker serve │◀──▶│ SQLite agent.db │ │ + CI / robomp │ │ - holds refresh tokens │ │ (canonical writer)│ │ + │ │ - background refresher │ └────────────────────┘ │ + │ │ /v1/{snapshot,refresh,…}│ │ + │ └─────────┬────────────────┘ │ + │ │ bearer ($CONFIG_DIR/auth-broker.token) │ + │ ▼ │ + │ ┌──────────────────────────┐ │ + │ │ omp auth-gateway serve │ RemoteAuthCredentialStore │ + │ │ /v1/{chat,messages,…} │ pulls /v1/snapshot at boot, │ + │ │ /v1/usage, /v1/models │ refreshes credentials by id │ + │ └─────────┬────────────────┘ via the broker on expiry │ + └────────────┼───────────────────────────────────────────────┘ + │ bearer ($CONFIG_DIR/auth-gateway.token) + ▼ + unauthenticated clients + (llm-git, macOS widget, robomp containers, IDE plugins, …) + │ + ▼ same path is forwarded with Authorization + api.anthropic.com / api.openai.com / … +``` + +The broker is the only writer of OAuth refresh tokens. Clients (including the gateway itself) load a redacted snapshot in which every `refresh` field has been replaced with `REMOTE_REFRESH_SENTINEL`; when an access token expires the client calls `POST /v1/credential/:id/refresh` and the broker performs the refresh server-side. `RemoteAuthCredentialStore` rejects any local code path that tries to write through it, with an error pointing at `omp auth-broker login` / `omp auth-broker logout`. + +## auth-broker + +### CLI + +``` +omp auth-broker serve [--bind=host:port] # boot the broker +omp auth-broker token [--regenerate] [--json] # print or rotate the bearer token +omp auth-broker login [--via=user@host] [--dry-run] +omp auth-broker logout +omp auth-broker import [--provider=] [--include-disabled] [--dry-run] [--json] +omp auth-broker migrate --from-local [--dry-run] [--json] +omp auth-broker status [--json] +``` + +- `serve` opens the local SQLite store at `getAgentDbPath()` and binds an HTTP listener (default `127.0.0.1:8765`). On startup a token is ensured at `/auth-broker.token` (mode `0600`, `0700` parent dir). The background refresher refreshes any OAuth credential whose `expires - Date.now() < refreshSkewMs` (default 5 min) every `refreshIntervalMs` (default 60 s). +- `token` prints the cached bearer or generates a new one. `--regenerate` rotates it. +- `login ` runs the per-provider OAuth flow locally, or — with `--via=user@host` — `ssh -L :127.0.0.1: user@host omp auth-broker login ` so the OAuth callback hits the local browser but the credential is written on the broker host. Built-in callback ports: `anthropic:54545`, `openai-codex:1455`, `google-gemini-cli:8085`, `google-antigravity:51121`, `gitlab-duo:8080`. +- `logout ` deletes every credential row for ``. +- `import ` imports CLIProxyAPI-style JSON credentials into the local SQLite store. Maps `type` field → omp provider (`claude → anthropic`, `codex → openai-codex`, `gemini → google-gemini-cli`, `antigravity → google-antigravity`, `gemini-cli → google-gemini-cli`). +- `migrate --from-local` walks the local SQLite store + env-derived credentials and idempotently uploads them to the configured broker (`POST /v1/credential`). +- `status` health-pings the configured remote broker. + +### Endpoints + +| Method | Path | Auth | Purpose | +| ------ | ---- | ---- | ------- | +| `GET` | `/v1/healthz` | none | Liveness + version | +| `GET` | `/v1/snapshot` | bearer | Redacted snapshot (refresh tokens replaced by sentinel) | +| `POST` | `/v1/credential` | bearer | Upsert one OAuth or API-key credential | +| `POST` | `/v1/credential/:id/refresh` | bearer | Force-refresh one OAuth credential | +| `POST` | `/v1/credential/:id/disable` | bearer | Disable one credential with a recorded cause | +| `GET` | `/v1/usage` | bearer | Aggregate `UsageReport[]` across credentials | + +Requests use `Authorization: Bearer `. The server compares against an in-memory token allow-list; the gateway’s implementation uses a timing-safe comparison. + +### Background refresher + +`AuthBrokerRefresher` iterates active OAuth credentials at `refreshIntervalMs` cadence and refreshes any within `refreshSkewMs` of expiry. Refreshes are single-flighted per credential id so a slow refresh cannot be retriggered. The refresher distinguishes: + +- **definitive failures** (`invalid_grant`, `invalid_token`, `revoked`, unauthorized refresh-token, 401/403 not from a network blip) — credentials are passed to `AuthStorage.disableCredentialById(id, cause)` so the next snapshot pull surfaces a clean delete on the client; +- **transient failures** (timeout / ECONNREFUSED / fetch failed) — left in place for the next sweep. + +## auth-gateway + +### CLI + +``` +omp auth-gateway serve [--bind=host:port] [--no-auth] +omp auth-gateway token [--regenerate] [--json] +omp auth-gateway status [--json] +``` + +- `serve` requires `OMP_AUTH_BROKER_URL` (or `auth.broker.url` in `config.yml`) — the gateway is itself a broker client. It calls `AuthBrokerClient.fetchSnapshot()`, wraps it in `RemoteAuthCredentialStore`, and constructs an `AuthStorage` that resolves access tokens through the broker. Default bind is `127.0.0.1:4000`. The gateway token is stored at `/auth-gateway.token` (`0600`); `--no-auth` disables the bearer check entirely (loopback-only use). +- `token` / `status` mirror the broker’s equivalents. + +### Endpoints + +| Method | Path | Auth | Purpose | +| ------ | ---- | ---- | ------- | +| `GET` | `/healthz` | none | Liveness + version | +| `GET` | `/v1/usage` | bearer | Aggregate `UsageReport[]` (proxied through `AuthStorage`) | +| `GET` | `/v1/models` | bearer | Bundled-model catalog filtered to providers with credentials | +| `POST` | `/v1/chat/completions` | bearer | OpenAI Chat Completions wire format | +| `POST` | `/v1/messages` | bearer | Anthropic Messages wire format | +| `POST` | `/v1/responses` | bearer | OpenAI Responses wire format | + +The model id is read from the top-level `model` field. The gateway picks the first bundled `Model` matching that id and: + +- **Passthrough fast-path** — when the inbound wire format matches the model’s native API (`openai-chat → openai-completions`, `anthropic-messages → anthropic-messages`, `openai-responses → openai-responses`), the request body is forwarded byte-for-byte with the client `Authorization`/`x-api-key` stripped and replaced by `Authorization: Bearer `. Provider-specific fields (`cache_control`, `service_tier`, tool-choice extensions, …) flow through unmodified. Hop-by-hop headers (RFC 7230) plus `Content-Encoding`/`Content-Length` are stripped from the upstream response. +- **Translate path** — when the inbound format and the resolved model’s API differ (e.g. `/v1/chat/completions` targeting an Anthropic model, or `/v1/responses` targeting `openai-codex-responses` which runs over a websocket transport), the request is parsed against the wire schema, rebuilt into an omp `Context`, dispatched through `streamSimple()`, and re-encoded back to the inbound format (SSE for streamed responses). + +`idleTimeout` on the underlying `Bun.serve` is set to `255 s` so long thinking-budget calls do not get killed by Bun’s default idle timeout. + +## Usage cache: server-side 5-min jitter + client-side 15 s single-flight + +Two layers cache the aggregate provider-usage report. Both are intentional and stacked. + +### Server-side cache (broker `AuthStorage`) + +`AuthStorage` caches each credential’s `UsageReport` in the broker’s SQLite store at a **5-minute per-credential TTL with ±25 % jitter**. Anthropic and OpenAI rate-limit `/usage` aggressively per source IP, and a synchronized 5-credential fan-out trips 429s every cycle; the jitter decorrelates refresh times within a few cycles. On fetch failure the store keeps the **last-good** report for up to 24 h with a short jittered re-poll window — so a transient upstream blip never blanks out the widget. + +Constants: `USAGE_REPORT_TTL_MS = 5 * 60_000`, `USAGE_LAST_GOOD_RETENTION_MS = 24 * 60 * 60_000` (`packages/ai/src/auth-storage.ts`). + +### Client-side single-flight (`RemoteAuthCredentialStore`) + +When the gateway (or any other broker client) calls `fetchUsageReports()` / `getUsageReport(provider, credential)`, `RemoteAuthCredentialStore` coalesces concurrent calls into a single `GET /v1/usage` round-trip and caches the result for **15 s** in memory. + +- `USAGE_CACHE_TTL_MS = 15_000` (`packages/ai/src/auth-broker/remote-store.ts`). +- A single `#usageInflight` promise is shared across all callers; a per-caller `AbortSignal` is **raced** against the shared promise, not threaded into it, so one caller’s abort never cascades into a peer’s in-flight request. +- On fetch failure the rejected promise is logged and the awaited value is `null` — callers (`AuthStorage.fetchUsageReports`, `#getUsageReport`) treat a `null` report as "no usage signal for this cycle" and proceed without it. **This is the 15 s TTL fallback**: the client absorbs transient broker outages by suppressing the error, returning `null` to ranking, and re-attempting after the 15 s window. + +The 15 s client window deliberately sits below the broker’s 5 min server cache, so almost every client poll is served from the broker’s already-cached value; the client cache exists to absorb the parallel fan-out generated by `AuthStorage.#rankOAuthSelections` into a single broker round-trip. + +## Operator opt-in + +The broker is **off** unless `OMP_AUTH_BROKER_URL` (or `auth.broker.url` in `config.yml`) is set. When set, `discoverAuthStorage` in `packages/coding-agent/src/sdk.ts` swaps the local SQLite credential store for `RemoteAuthCredentialStore` and every API call resolves credentials through the broker. + +### Environment variables + +| Variable | Purpose | Required when | +| -------- | ------- | ------------- | +| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`). Selecting this puts the client in broker mode — local SQLite is bypassed. | Any time the omp client should resolve credentials through a broker (and required by `omp auth-gateway serve`). | +| `OMP_AUTH_BROKER_TOKEN` | Bearer token used for every broker endpoint except `/v1/healthz`. | When `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token`. | + +Resolution order in `resolveAuthBrokerConfig()`: + +1. `OMP_AUTH_BROKER_URL` env (else `auth.broker.url` from `config.yml`, with `$ENV_NAME` resolution); +2. `OMP_AUTH_BROKER_TOKEN` env (else `auth.broker.token` from `config.yml`, else `/auth-broker.token`); +3. URL set but no token resolvable → hard error pointing at the token file path. + +The gateway has no dedicated env vars — it inherits `OMP_AUTH_BROKER_*` because it is itself a broker client. + +### `config.yml` keys + +| Key | Default | Purpose | +| --- | ------- | ------- | +| `auth.broker.url` | unset | Same as `OMP_AUTH_BROKER_URL`; env wins. Hidden from the settings UI. | +| `auth.broker.token` | unset | Same as `OMP_AUTH_BROKER_TOKEN`; env wins. Values may be the literal token or `$ENV_NAME` to indirect through env. | + +### Token files + +| Path | Owner | Mode | +| ---- | ----- | ---- | +| `/auth-broker.token` | `omp auth-broker serve` (created at first start) | `0600` in a `0700` parent dir | +| `/auth-gateway.token` | `omp auth-gateway serve` (skipped under `--no-auth`) | `0600` in a `0700` parent dir | + +`` resolves to `~/.omp/` (respecting `PI_CONFIG_DIR`). + +## Interaction with the local API-key resolution order + +The broker only owns OAuth credentials and provider-API-key credentials that were uploaded to it. The standard credential ladder in `models.md` (`Auth and API key resolution order`) is preserved, with one addition committed alongside the gateway: + +- `AuthStorage.setConfigApiKey / removeConfigApiKey / clearConfigApiKeys` let a `models.yml` `apiKey` beat a stored OAuth token **without** overriding an explicit `--api-key`. This is what allows a broker-resolved OAuth credential to be reliably shadowed by a per-environment `models.yml` config key when both are present. + +## See also + +- [`secrets.md`](./secrets.md) — secret obfuscation around tokens that *do* leak through (e.g. `OMP_AUTH_BROKER_TOKEN` in shell output). +- [`models.md`](./models.md) — provider auth resolution order; the broker plugs in at layers 2–3 (stored credentials). +- [`environment-variables.md`](./environment-variables.md) — full env reference including `OMP_AUTH_BROKER_URL` / `OMP_AUTH_BROKER_TOKEN`. diff --git a/docs/environment-variables.md b/docs/environment-variables.md index 518fcc005..01ec224b7 100644 --- a/docs/environment-variables.md +++ b/docs/environment-variables.md @@ -84,6 +84,17 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not | `GH_TOKEN` | Copilot fallback; GitHub API auth in web scraper | In web scraper: `GITHUB_TOKEN` → `GH_TOKEN` | | `GITHUB_TOKEN` | Copilot fallback; GitHub API auth in web scraper | In web scraper: checked before `GH_TOKEN` | +### Auth broker / auth gateway (remote credential vault) + +When the broker is enabled, the local SQLite credential store is bypassed and all OAuth refresh / access tokens live on the broker host. See [`auth-broker-gateway.md`](./auth-broker-gateway.md) for the full protocol, CLI surface, and 5-min/15-s usage cache layering. + +| Variable | Used for | Required when | Notes / precedence | +| ----------------------- | ------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`); selects broker mode | Resolving credentials through a broker; also required by `omp auth-gateway serve` (the gateway is itself a broker client) | Wins over `auth.broker.url` in `config.yml`. When set with no resolvable token, `resolveAuthBrokerConfig()` hard-errors instead of falling back to local SQLite. | +| `OMP_AUTH_BROKER_TOKEN` | Bearer token sent on every broker endpoint except `/v1/healthz` | `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token` | Resolution: this env → `auth.broker.token` (`$ENV_NAME` indirection supported) → `/auth-broker.token` (mode `0600`). `` is `~/.omp/` (respecting `PI_CONFIG_DIR`). | + +The gateway has no dedicated env vars — it inherits `OMP_AUTH_BROKER_*`. Its own inbound bearer token lives at `/auth-gateway.token` and is managed via `omp auth-gateway token`. + --- ## 2) Provider-specific runtime configuration @@ -277,7 +288,7 @@ Extra conditional behavior: | `PI_SUBPROCESS_CMD` | Overrides subagent spawn command (`omp` / `omp.cmd` resolution bypass) | | `PI_TASK_MAX_OUTPUT_BYTES` | Max captured output bytes per subagent (default `500000`) | | `PI_TASK_MAX_OUTPUT_LINES` | Max captured output lines per subagent (default `5000`) | -| `PI_TIMING` | If `1`, enables startup/tool timing instrumentation logs | +| `PI_TIMING` | If set (any non-empty value), prints a hierarchical timing-span tree to **stderr** via `logger.printTimings()`. In interactive mode the tree prints once the agent is ready (before the TUI starts); in print mode it prints after the whole prompt batch completes. Print-mode prompts are wrapped in `print:prompt:initial` / `print:prompt:next` spans so each user message shows up as its own row. `PI_TIMING=x` exits the process with code 0 right after printing in interactive mode (use to measure cold startup only). `PI_TIMING=full` lists every module-load entry instead of just the top N. | | `PI_PACKAGE_DIR` | Overrides package asset base dir resolution (docs/examples/changelog path lookup) | | `PI_DISABLE_LSPMUX` | If `1`, disables lspmux detection/integration and forces direct LSP server spawning | | `PI_RPC_EMIT_TITLE` | Boolean-like flag enabling title events in RPC mode | diff --git a/docs/install-id.md b/docs/install-id.md new file mode 100644 index 000000000..4c7571132 --- /dev/null +++ b/docs/install-id.md @@ -0,0 +1,41 @@ +# Install ID + +A persistent per-install UUID that identifies a single oh-my-pi installation across sessions. Used as a stable correlation key for server-side dedup of telemetry-style pushes (currently the auto-QA grievance flush from `report_tool_issue`). + +## API + +Exported from `@oh-my-pi/pi-utils` (`packages/utils/src/dirs.ts`): + +| Symbol | Purpose | +| --- | --- | +| `getInstallId(): string` | Returns the install ID, generating and persisting one on first call. Result is cached in-process for the lifetime of the runtime. | +| `__resetInstallIdCacheForTests(): void` | Clears the in-process cache. Test-only — MUST NOT be called from production code. | + +The returned value is a canonical lowercase RFC 4122 UUID matching `^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$`. + +## Storage + +- Path: `/install-id` — i.e. `~/.omp/install-id` by default, respecting `PI_CONFIG_DIR` via `getConfigRootDir()`. +- Format: a single UUID line (trailing `\n`). +- Permissions: file is created with mode `0o600`. +- Lifecycle: independent of `~/.omp/agent/`. Wiping agent state (sessions, settings, DB) does NOT regenerate the install ID; only deleting the `install-id` file itself does. + +## Generation and lifecycle + +1. First call to `getInstallId()` reads the file. If contents parse as a valid UUID, that value is cached and returned. +2. Otherwise the helper calls `crypto.randomUUID()` (Node's CSPRNG-backed UUID v4) to mint a new ID. +3. The new value is written via `open(O_WRONLY | O_CREAT | O_EXCL, 0o600)`. The exclusive-create guard means two processes hitting first-call simultaneously cannot both succeed — the loser sees `EEXIST`, re-reads the winner's file, and adopts that ID. +4. If the existing file contained non-empty garbage (failed UUID regex), it is `unlink`ed before the exclusive create so `O_EXCL` does not trip on stale data. +5. Any other write failure (read-only FS, permission error) is swallowed: the freshly generated UUID is still cached in-memory so the rest of the process sees a stable value, and subsequent process launches will retry persistence. +6. Subsequent in-process calls return the cached value without touching disk. Mutating the file on disk after the first call has no effect until the process restarts (or tests call `__resetInstallIdCacheForTests`). + +## Consumers + +- `packages/coding-agent/src/tools/report-tool-issue.ts` — included as `installId` in the auto-QA grievance push body so the backend can deduplicate repeated reports from the same install. See `dev.autoqaPush.*` settings and `PI_AUTO_QA_PUSH_*` env vars. + +New consumers MUST treat the value as opaque and MUST NOT derive PII from it; the helper does not mix in hostname, username, or any other host-identifying entropy. + +## See also + +- [environment-variables.md](environment-variables.md) — `PI_CONFIG_DIR` controls where `install-id` lives. +- [config-usage.md](config-usage.md) — broader config-root layout. diff --git a/docs/models.md b/docs/models.md index 086fc487c..f8a311697 100644 --- a/docs/models.md +++ b/docs/models.md @@ -148,6 +148,17 @@ ModelRegistry pipeline (on refresh): - otherwise append 6. Load cached/runtime-discovered models (Ollama, llama.cpp, LM Studio, plus built-in provider managers), then re-apply model overrides. +### Provider-model cache and static fingerprint + +Cached per-provider model lists are persisted in the model-cache SQLite +database (schema v3) with a `static_fingerprint` column that hashes the +static catalog slice merged into the row. When `resolveProviderModels` +skips the network fetch and the fingerprint of the in-memory static +catalog matches the cached one, the cached rows are returned verbatim — +the static + dynamic merge is bypassed entirely. The fingerprint is +memoized per process via a WeakMap keyed by the static-models array +reference, so repeated cold-start calls do not re-hash. + ## Canonical model equivalence and coalescing The registry keeps every concrete provider model and then builds a canonical layer above them. @@ -309,6 +320,12 @@ Keyless providers: - Providers marked `auth: none` are treated as available without credentials. - `getApiKey*` returns `kNoAuth` for them. +### Broker mode + +When `OMP_AUTH_BROKER_URL` (or `auth.broker.url`) is set, the local SQLite credential store is replaced by `RemoteAuthCredentialStore`. Layers 2 and 3 above (stored API key / OAuth in `agent.db`) are served from a broker-supplied snapshot whose `refresh` tokens are redacted; expiry triggers `POST /v1/credential/:id/refresh` on the broker rather than a local refresh. + +`AuthStorage.setConfigApiKey` lets a `models.yml` `apiKey` win over a broker-resolved OAuth token without overriding a runtime `--api-key`. See [`auth-broker-gateway.md`](./auth-broker-gateway.md) for the full broker / gateway design and env surface (`OMP_AUTH_BROKER_URL`, `OMP_AUTH_BROKER_TOKEN`, `auth.broker.url`, `auth.broker.token`). + ## Model availability vs all models - `getAll()` returns the loaded model registry (built-in + merged custom + discovered). @@ -530,6 +547,14 @@ providers: ``` `disableStrictTools` is a provider-level flag that applies to all models in the provider. + +Tool schemas going on the wire are normalized by the unified flow in +`packages/ai/src/utils/schema/normalize.ts` (Google/CCA/MCP dispatchers +plus the OpenAI strict-mode sanitize+enforce pipeline). See +[`ai-schema-normalize.md`](./ai-schema-normalize.md) for the strict-mode +edge cases (local `$ref` inlining, single-item `allOf` collapse, +`anyOf`-wrapper description hoist, enum/const primitive-type inference) +and the per-provider dispatcher mapping. ## Practical examples ### Local OpenAI-compatible endpoint (no auth) diff --git a/docs/natives-binding-contract.md b/docs/natives-binding-contract.md index 1821be5b8..f787b58d2 100644 --- a/docs/natives-binding-contract.md +++ b/docs/natives-binding-contract.md @@ -68,7 +68,7 @@ Consumers in `packages/coding-agent` and `packages/tui` import directly from `@o | PTY | `new PtySession()`, `start/write/resize/kill` | `pty.rs` | class / promises | | Process | `killTree(pid, signal)`, `listDescendants(pid)` | `ps.rs` | sync | | Keys | `parseKey`, `matchesKey`, Kitty/legacy helpers | `keys.rs` | sync | -| Text | `wrapTextWithAnsi`, `truncateToWidth`, `sliceWithWidth`, `extractSegments`, `sanitizeText`, `visibleWidth` | `text.rs` | sync | +| Text | `wrapTextWithAnsi`, `truncateToWidth`, `sliceWithWidth`, `extractSegments`, `visibleWidth` | `text.rs` | sync | | Highlight | `highlightCode`, `supportsLanguage`, `getSupportedLanguages` | `highlight.rs` | sync | | HTML | `htmlToMarkdown(html, options?)` | `html.rs` | `Promise` | | Image | `PhotonImage`, `encodeSixel` | `image.rs` | class / sync / promises | diff --git a/docs/natives-build-release-debugging.md b/docs/natives-build-release-debugging.md index 579051d2d..f8a9b2723 100644 --- a/docs/natives-build-release-debugging.md +++ b/docs/natives-build-release-debugging.md @@ -215,3 +215,74 @@ bun --cwd=packages/natives run embed:native # Reset embedded manifest to null stub bun --cwd=packages/natives run embed:native -- --reset ``` + +## Orchestrator-side content-addressed build cache (robomp) + +When `pi-natives` is built inside the robomp orchestrator (`python/robomp/`), workspaces share built artifacts through a content-addressed cache instead of rebuilding from scratch in every per-issue worktree. The cache is **orchestrator-side only** — `bun --cwd=packages/natives run build` itself is unchanged; the cache lives outside the build pipeline and is populated/captured around `ensure_workspace` and post-task success in `python/robomp/src/robomp/natives_cache.py`. + +### What is cached + +The complete set of files in `packages/natives/native/` that are pure functions of the cache-key inputs: + +- `pi_natives.-[-variant].node` (glob `pi_natives.*.node`) +- `index.d.ts` +- `index.js` +- `embedded-addon.js` +- `manifest.json` (cache metadata: key, target triple, capture timestamp, source workspace, commit) + +An entry is only considered a hit when the `.node` glob matches AND every companion plus the manifest is present. Partial entries are evicted on GC. + +### Cache key + +The key is `sha256` over `(path \t git-tree-hash \n)` pairs for the following inputs, in this order (order is significant), followed by the target triple: + +1. `crates` (whole subtree — pi-natives transitively depends on other workspace crates) +2. `Cargo.lock` +3. `Cargo.toml` +4. `rust-toolchain.toml` +5. `packages/natives` (whole subtree — build script, `scripts/*`, package.json with napi config) + +Tree hashes come from one `git cat-file --batch-check` invocation against `HEAD`; paths missing from `HEAD` fold in as a fixed null hash so the key stays deterministic across repos that don't ship every input. The target-triple suffix matches the napi addon basename convention (`-` for non-x64, `--` for x64). When `TARGET_VARIANT` is unset on an x64 host the variant component is `host` rather than autodetected — the key is stable on a given machine but a `modern`/`baseline` build with an explicit `TARGET_VARIANT` gets a different key. + +Anything outside this input set (Rust toolchain auto-installed delta, host glibc, env vars other than `TARGET_VARIANT`) is **not** in the key. If you need to invalidate after such a change, delete the cache directory by hand or bump one of the input files. + +### Layout and ownership + +- Root: `/data/cache/pi-natives` (provisioned by `entrypoint.sh` alongside the cargo caches, owned `root:omp`, mode `02770` setgid so cached files inherit `gid=omp` and stay readable by every slot user). +- Per-repo subdirectory: `//` where the slug is `owner__repo` (mirrors `SandboxManager.pool_path`). +- Per-entry directory: `///` containing the cached files plus `manifest.json`. +- Per-repo lockfile: `//.lock` (advisory `fcntl.flock`, exclusive on capture and GC). +- Staging dirs (`..tmp.`) during capture; renamed atomically into the final entry path. Stale staging dirs from crashed captures are swept on GC. + +### Populate and capture semantics + +- **Populate** (workspace ← cache) runs inside `ensure_workspace`. On a key hit the `.node` is **hardlinked** into the workspace (zero-copy, shared inode); the companion `index.d.ts` / `index.js` / `embedded-addon.js` are **copied** (independent inodes) because the napi build's `installGeneratedBindings` and `gen-enums.ts` rewrite those files via `open(..., 'w')` — an in-place truncate that would otherwise propagate through a hardlink and corrupt the cache. Cross-device hardlink failures (`EXDEV`) fall back to copy. +- **Capture** (cache ← workspace) runs from the post-task success path when the build produced a complete artifact set. Capture uses **copy**, not hardlink: hardlinking a slot-owned workspace file would preserve slot UID ownership on the cached inode and defeat the shared-group model. Copying creates a fresh root-owned, `gid=omp` inode via the setgid cache root. Capture is idempotent under the per-repo flock: a concurrent capture for the same key returns the existing entry. + +### Garbage collection + +A periodic GC loop runs in `WorkerPool` with two caps per repo. When either cap is exceeded, oldest entries (by `manifest.json.captured_at`) are dropped first: + +- entry count cap (`max_entries_per_repo`, default 8) +- byte cap (`max_bytes`, default 4 GiB) + +Workspaces that hardlinked a `.node` before GC retain access via the kernel inode refcount — `rmtree` of the cache entry does not delete the file from the workspace. + +### Configuration (settings on `robomp.config.Settings`) + +| Env var | Default | Effect | +| -------------------------------------------- | ------------------------ | ------------------------------------------------------------- | +| `ROBOMP_NATIVES_CACHE_ENABLED` | `true` | Master switch. When false the populate/capture hooks no-op and every workspace builds from scratch. | +| `ROBOMP_NATIVES_CACHE_ROOT` | `/data/cache/pi-natives` | Cache root directory. Must be `root:omp 02770` for cross-slot reads. | +| `ROBOMP_NATIVES_CACHE_MAX_ENTRIES_PER_REPO` | `8` | LRU entry-count cap, per repo slug. | +| `ROBOMP_NATIVES_CACHE_MAX_BYTES` | `4294967296` (4 GiB) | LRU byte cap, per repo slug. | +| `ROBOMP_NATIVES_CACHE_GC_INTERVAL_SECONDS` | `3600` | Period of the background GC loop in `WorkerPool`. | + +### Manual invalidation + +- One key: `rm -rf /data/cache/pi-natives//`. +- One repo: `rm -rf /data/cache/pi-natives/`. +- Everything: `rm -rf /data/cache/pi-natives/*` (preserve the root so its setgid mode survives). +- Stuck lock: `rm /data/cache/pi-natives//.lock` (only when no orchestrator process is touching the repo). + +Trigger an automatic miss by editing any path in the key set: a single touched byte under `crates/`, `Cargo.lock`, `Cargo.toml`, `rust-toolchain.toml`, or `packages/natives/` shifts the tree hash and forces a fresh build at the next populate. diff --git a/docs/natives-text-search-pipeline.md b/docs/natives-text-search-pipeline.md index bdf460d53..82bccfbb8 100644 --- a/docs/natives-text-search-pipeline.md +++ b/docs/natives-text-search-pipeline.md @@ -37,7 +37,6 @@ Terminology follows `docs/natives-architecture.md`: | `truncateToWidth(text, maxWidth, ellipsis, pad, tabWidth)` | `truncateToWidth` | `text.rs` | | `sliceWithWidth(line, startCol, length, strict, tabWidth)` | `sliceWithWidth` | `text.rs` | | `extractSegments(line, beforeEnd, afterStart, afterLen, strictAfter, tabWidth)` | `extractSegments` | `text.rs` | -| `sanitizeText(text)` | `sanitizeText` | `text.rs` | | `visibleWidth(text, tabWidth)` | `visibleWidth` | `text.rs` | | `highlightCode(code, lang, colors)` | `highlightCode` | `highlight.rs` | | `supportsLanguage(lang)` | `supportsLanguage` | `highlight.rs` | @@ -206,7 +205,7 @@ These are pure, in-memory utilities. - `truncateToWidth`: visible-cell truncation with ellipsis policy (`Unicode`, `Ascii`, `Omit`), optional right padding. - `sliceWithWidth`: column slicing with optional strict width enforcement. - `extractSegments`: extracts before/after segments around an overlay while restoring ANSI state for the `after` segment. -- `sanitizeText`: strips ANSI escapes + control chars, drops lone surrogates, normalizes line endings. +- `sanitizeText` (ANSI/control/surrogate stripping with line-ending normalization) no longer lives in `text.rs`; it moved to `@oh-my-pi/pi-utils` as a pure-JS implementation in `packages/utils/src/sanitize-text.ts`. The native binding was removed in the same change because the JS version was competitive on the benchmarked workloads, and keeping a Rust copy forced every caller (including `pi-utils`) to pull in `@oh-my-pi/pi-natives`. - `visibleWidth`: counts visible terminal cells using caller-supplied tab width. ### Failure behavior diff --git a/docs/sdk.md b/docs/sdk.md index cad7b6cb3..f6ee5d30f 100644 --- a/docs/sdk.md +++ b/docs/sdk.md @@ -308,6 +308,18 @@ type CreateAgentSessionResult = { Use `setToolUIContext(...)` only if your embedder provides UI capabilities that tools/extensions should call into. +## Startup performance + +`createAgentSession()` runs two background optimizations to overlap I/O with the rest of session setup: + +- **Model-host preconnect.** As soon as the model is resolved, the SDK fires a best-effort `fetch.preconnect(model.baseUrl)` so DNS + TCP + TLS + HTTP/2 to the provider's host happens in parallel with extension/skill load, tool registry build, and system-prompt assembly. The first real `fetch(...)` then reuses the warm connection, saving 100–300 ms on transcontinental hops (e.g. residential IP → `api.anthropic.com`). Implementation lives in `preconnectModelHost()` in `packages/coding-agent/src/sdk.ts`. If `fetch.preconnect` is unavailable (non-Bun runtime) or the call throws, the optimization is silently skipped — never a hard dependency. Applies to every mode (interactive, print, RPC, ACP). +- **Conditional LSP warmup.** Startup LSP servers (those returned by `discoverStartupLspServers(cwd)`) are only warmed when **all** of these hold: + - `enableLsp !== false` on the session options, **and** + - `options.hasUI === true` (interactive TUI), **and** + - the `lsp.diagnosticsOnWrite` setting is enabled. + + Print / script / RPC / ACP invocations (`hasUI=false`) skip the warmup entirely: they don't render the warmup status indicator and typically finish before the language servers would stabilize, so warming them just spends CPU parsing big `initialize` responses concurrently with the LLM stream consumer and jitters perceived latency. Tools that actually need an LSP server still spin one up on demand through `getOrCreateClient()` — only the *startup* warmup is skipped. The returned `lspServers` field in `CreateAgentSessionResult` is therefore `undefined` (not an empty array) whenever the warmup branch was bypassed. + ## Minimal controlled embed example ```ts diff --git a/docs/secrets.md b/docs/secrets.md index 2ec81bde6..0b7dcf760 100644 --- a/docs/secrets.md +++ b/docs/secrets.md @@ -106,3 +106,7 @@ Environment variables are collected first, then file-defined entries are appende - `packages/coding-agent/src/secrets/obfuscator.ts` -- `SecretObfuscator` class, placeholder generation, message obfuscation - `packages/coding-agent/src/secrets/regex.ts` -- regex literal parsing and compilation - `packages/coding-agent/src/config/settings-schema.ts` -- `secrets.enabled` setting definition + +## See also + +- [`auth-broker-gateway.md`](./auth-broker-gateway.md) -- remote credential vault and forward-proxy that keep provider OAuth refresh tokens and access tokens off developer hosts entirely (complementary to in-process obfuscation). diff --git a/docs/session-tree-plan.md b/docs/session-tree-plan.md index d5905a3f1..eea5120b2 100644 --- a/docs/session-tree-plan.md +++ b/docs/session-tree-plan.md @@ -181,6 +181,33 @@ Adjacent but related lifecycle hooks: - In-memory sessions never return a branch file path from `createBranchedSession`. - Tree context reconstruction includes service-tier and MCP tool-selection state, but those entries do not become LLM messages. +## Plan approval session naming + +When a user approves a plan from plan mode (`InteractiveMode.#approvePlan`), the approval handler seeds the session name from the plan's title so the resulting (fresh or compacted) session does not stay unnamed. + +Trigger: + +- Plan approval reaches `#approvePlan(...)` with `options.title` populated from the plan-approval details. +- This runs for every approval choice (`Approve and execute`, `Approve and compact context`, plain `Approve`); the synthetic `plan-approved` prompt is what otherwise bypasses the input-controller's title-generation path. + +Naming source: + +- The normalized plan title is humanized via `humanizePlanTitle(title)` (`packages/coding-agent/src/plan-mode/approved-plan.ts`): + - replaces runs of `-`/`_` with a single space + - trims whitespace + - capitalizes the first character + - returns `""` for whitespace-only / separator-only input +- The humanized name is applied with `sessionManager.setSessionName(name, "auto")`. Because `setSessionName` is a no-op when `titleSource === "user"`, the seeded name never overrides a name the user already chose (e.g. on the `preserveContext` path where the session continues with prior naming). +- On successful apply, the terminal title (`setSessionTerminalTitle`) and the editor border color are refreshed to reflect the new name. + +Examples (from `humanizePlanTitle`): + +- `migrate-mcp-loader` → `Migrate mcp loader` +- `fix_session_naming` → `Fix session naming` +- `foo--bar__baz` → `Foo bar baz` +- `RefactorRouter` → `RefactorRouter` (no separators to expand) +- `""` / `"---"` → `""` (no name applied) + ## Legacy compatibility still present Session migrations still run on load: diff --git a/docs/tools/eval.md b/docs/tools/eval.md index da9f37f36..573241f3b 100644 --- a/docs/tools/eval.md +++ b/docs/tools/eval.md @@ -8,8 +8,6 @@ - Entry: `packages/coding-agent/src/tools/eval.ts` - Model-facing prompt: `packages/coding-agent/src/prompts/tools/eval.md` - Key collaborators: - - `packages/coding-agent/src/eval/parse.ts` — lenient cell parser - - `packages/coding-agent/src/eval/sniff.ts` — language sniffing heuristics - `packages/coding-agent/src/eval/backend.ts` — backend execution contract - `packages/coding-agent/src/eval/js/index.ts` — JS backend adapter - `packages/coding-agent/src/eval/js/executor.ts` — JS execution + output sink @@ -24,36 +22,33 @@ ## Inputs +Tool parameters are a JSON object with a single `cells` field — an ordered array of cell objects. Each cell is a structured record; there is no `*** Cell` header parsing, no language sniffing, and no implicit single-cell fallback. Cells run in array order; state persists within each language across cells and across tool calls. + | Field | Type | Required | Description | | --- | --- | --- | --- | -| `input` | `string` | Yes | Cell program text. Parsed by `parseEvalInput()` in `packages/coding-agent/src/eval/parse.ts`, not by JSON subfields. | +| `cells` | `EvalCellInput[]` | Yes | Cells executed in order. At least one cell is required (`.min(1)`). | -`input` syntax accepted at runtime: +Each `EvalCellInput` (from `evalCellSchema` in `packages/coding-agent/src/tools/eval.ts`): -- Cell header: `*** Cell `. Attributes are space-separated tokens with quoted titles (`"..."` or `'...'`). -- Canonical tokens (advertised in the prompt): - - `:""` — language + title shorthand. `lang` is `py` or `js` (lenient: also `ts`, plus the long-form aliases `python`, `javascript`, `typescript`, `ipy`, `ipython`). - - `t:<n>[ms|s|m]` — per-cell timeout (default 30s). - - `rst` — wipe this cell's language kernel before running. -- Lenient additional tokens (accepted by the parser, not advertised): - - bare language token (`py`, `js`) - - `id:"..."` / `title:"..."` / `name:"..."` / `cell:"..."` / `file:"..."` / `label:"..."` — title aliases - - `timeout:` / `duration:` / `time:` — `t:` aliases - - `reset` — `rst` alias - - `rst:true|false|1|0|yes|no|on|off` — explicit boolean form - - a bare positional duration token (`30s`, `2m`, `500ms`) - - any unclassified bare token folds into a positional title fragment -- Cell body: every following line until the next `*** Cell ...`, the optional `*** End`, or `*** Abort`. `*** End` is a quirk fix for GPT-trained models that emit terminators and is not documented in the prompt. +| Field | Type | Required | Description | +| --- | --- | --- | --- | +| `language` | `"py" \| "js"` | Yes | Backend selector. `"py"` maps to the IPython/Jupyter kernel (`python` backend); `"js"` maps to the persistent JavaScript VM. | +| `code` | `string` | Yes | Cell body, verbatim. JSON-encoded — embed newlines, quotes, and indentation directly; no fences, no headers. | +| `title` | `string` | No | Short label rendered in the transcript (e.g. `"imports"`, `"load config"`). | +| `timeout` | `integer` | No | Per-cell timeout in seconds, clamped to `1..600`. Defaults to 30 when omitted. | +| `reset` | `boolean` | No | Wipe this cell's language kernel before running. Reset is per-language: a `py` cell's reset does not touch the JS VM and vice versa. Defaults to `false`. | -Leniencies in `packages/coding-agent/src/eval/parse.ts`: +Minimal example matching the live schema: -- Markers accept two or more leading `*` and flexible whitespace. -- `*** End` is optional everywhere; the parser silently consumes trailing tokens (e.g. `*** End py`). -- Missing terminators between adjacent cells are tolerated; the next `*** Cell` closes the prior cell, and stray non-marker lines between cells fold into the prior cell's body without crashing. -- Bare code or a single markdown fence such as ```` ```py ```` is treated as one implicit cell. -- If `*** Abort` appears, the in-progress cell is dropped and the result carries an abort warning. To preserve a completed cell before `*** Abort`, emit `*** End` first. - -The tool also exposes a custom Lark grammar from `packages/coding-agent/src/eval/eval.lark` for constrained sampling. That grammar is stricter than the runtime parser: it requires the canonical `*** Cell <lang>:"title"` header form with a fixed attribute order, advertises only `py` / `js`, and pins the trailing `*** End` so GPT-trained models' natural terminator habit aligns with the constrained output. +```json +{ + "cells": [ + { "language": "py", "title": "imports", "timeout": 10, "code": "import json\nfrom pathlib import Path" }, + { "language": "py", "title": "load config", "code": "data = json.loads(read('package.json'))\ndisplay(data)" }, + { "language": "js", "title": "summary", "reset": true, "code": "const data = JSON.parse(await read('package.json'));\ndisplay(data);\nreturn data.name;" } + ] +} +``` ## Outputs @@ -69,17 +64,17 @@ Returned shape: - `jsonOutputs`: structured values emitted via `display(...)` - `images`: image payloads emitted by Python rich display or JS `display({ type: "image", ... })` - `statusEvents`: aggregated helper/tool status events - - `notice`: backend fallback notice + - `notice`: backend fallback notice (currently unused; reserved for future per-cell notices) - `meta`: truncation metadata - `isError`: set on cell failure or cancellation Renderer behavior in `packages/coding-agent/src/tools/eval.ts`: -- call preview renders parsed code cells with syntax highlighting +- call preview renders each cell's `code` with syntax highlighting based on its declared `language` - result view renders each cell separately, including status, duration, and output - markdown outputs are rendered with the Markdown component instead of plain text - `jsonOutputs` render as a tree, collapsed or expanded depending on UI state -- timeout / fallback / truncation notices render as dim metadata lines +- timeout / truncation notices render as dim metadata lines - images are carried in `details.images`; generic tool UI image handling renders them outside the text block Side-channel artifacts: @@ -89,54 +84,48 @@ Side-channel artifacts: ## Flow -1. `EvalTool.execute()` in `packages/coding-agent/src/tools/eval.ts` parses `params.input` with `parseEvalInput()`. -2. `parseEvalInput()` normalizes newlines, collects cells, parses attributes, and assigns each cell a language from the header, language sniffing, or the default `python`. -3. Back in `execute()`, each parsed cell is resolved to a backend with `resolveBackend()`: - - explicit `python`/`js` requests are validated against session settings and backend availability - - otherwise `sniffEvalLanguage()` in `packages/coding-agent/src/eval/sniff.ts` tries shebangs and language markers - - if no explicit language was present, later cells prefer the previous runtime language before re-sniffing - - Python is preferred when available; JS is the fallback when Python is unavailable or disabled -4. The tool allocates an `OutputSink`, a `TailBuffer`, per-cell result objects, and a `sessionAbortController`. `session.trackEvalExecution?.(...)` can wrap the whole run for external cancellation tracking. -5. Cells execute sequentially. For each cell, `execute()`: - - clamps the cell timeout through `clampTimeout("eval", ...)` +1. `EvalTool.execute()` in `packages/coding-agent/src/tools/eval.ts` receives `params.cells` already validated by the Zod schema — no string parsing step. +2. For each cell, `execute()` maps `cell.language` to an `EvalLanguage` (`"py"` → `"python"`, `"js"` → `"js"`) and calls `resolveBackend(session, language)`: + - `python` is gated on `eval.py !== false` and `pythonBackend.isAvailable(session)`. + - `js` is gated on `eval.js !== false`. + - A disabled or unavailable requested backend throws `ToolError`; there is no auto-fallback or sniffing. +3. The tool allocates an `OutputSink`, a `TailBuffer`, per-cell result objects, and a `sessionAbortController`. `session.trackEvalExecution?.(...)` can wrap the whole run for external cancellation tracking. +4. Cells execute sequentially. For each cell, `execute()`: + - clamps `(cell.timeout ?? 30) * 1000` ms through `clampTimeout("eval", ...)` - builds a combined abort signal from the tool signal, the timeout, and the session abort controller - marks the cell `running` and emits an update - - calls the backend’s `execute()` with `cwd`, `sessionId`, `sessionFile`, `kernelOwnerId`, `deadlineMs`, `reset`, artifact info, and chunk callback -6. JS cells dispatch through `packages/coding-agent/src/eval/js/index.ts` into `executeJs()`; Python cells dispatch through `packages/coding-agent/src/eval/py/index.ts` into `executePython()`. -7. Backend text chunks stream into the shared `OutputSink`; rich outputs are accumulated separately as JSON, images, markdown markers, and status events. -8. After each cell: + - calls the backend’s `execute()` with `cwd`, `sessionId`, `sessionFile`, `kernelOwnerId`, `deadlineMs`, `reset` (defaults to `false`), artifact info, and chunk callback +5. JS cells dispatch through `packages/coding-agent/src/eval/js/index.ts` into `executeJs()`; Python cells dispatch through `packages/coding-agent/src/eval/py/index.ts` into `executePython()`. +6. Backend text chunks stream into the shared `OutputSink`; rich outputs are accumulated separately as JSON, images, markdown markers, and status events. +7. After each cell: - text output is trimmed and stored on that cell result - multi-cell runs prefix text with `[i/n]` and the optional title - cancellations return early with `isError: true` and a cell-specific abort message - non-zero exit codes return early with `isError: true` and a message naming the failed cell - later cells are skipped after the first error, but earlier cell state persists in the underlying runtime -9. On success, the tool joins all cell outputs, synthesizes `(no text output)` or `(no output)` when needed, and attaches truncation metadata from `summarizeFinal()`. -10. The renderer uses `details.cells`, `details.jsonOutputs`, and `details.statusEvents` to build notebook-style output. `mergeCallAndResult = true` and `inline = true`, so call and result render together in the transcript. +8. On success, the tool joins all cell outputs, synthesizes `(no text output)` or `(no output)` when needed, and attaches truncation metadata from `summarizeFinal()`. +9. The renderer uses `details.cells`, `details.jsonOutputs`, and `details.statusEvents` to build notebook-style output. `mergeCallAndResult = true` and `inline = true`, so call and result render together in the transcript. ## Modes / Variants -### Parsing modes - -- Explicit multi-cell format with `*** Cell ...` headers -- Implicit single-cell fallback for bare code or a single fenced block -- Abort-recovery parse path when `*** Abort` is present - ### Backend selection -- Explicit Python backend -- Explicit JavaScript backend -- Auto-detected backend via `sniffEvalLanguage()` -- Fallback from requested/inferred Python to JS when Python is unavailable -- Fallback notice when JS markers are seen but `eval.js` is disabled and Python is used instead +Backend choice is **explicit per cell** — there is no auto-detection. + +- `language: "py"` → Python (IPython/Jupyter) backend +- `language: "js"` → JavaScript VM backend + +If the requested backend is disabled or unavailable, the tool throws `ToolError` for that cell. The caller chooses; the tool does not silently substitute. ### JavaScript runtime Implemented in `packages/coding-agent/src/eval/js/context-manager.ts` and `packages/coding-agent/src/eval/js/prelude.txt`. - Persistent `vm.Context` instances keyed by `js:${sessionId}` in `vmContexts` -- `rst` calls `resetVmContext(sessionKey)` before the cell executes +- `reset: true` calls `resetVmContext(sessionKey)` before the cell executes - Top-level `await` and bare `return` are supported by wrapping code in an async IIFE when `wrapCode()` sees `await` or `return` - Top-level static `import ... from ...` and dynamic `import(...)` calls are routed through `rewriteImports()`, which sends them via `__omp_import__` so the specifier resolves against the session cwd +- Module cache is busted for **local** imports between cells so edits to source files are picked up without restarting the runtime. `__omp_import__` deletes `require.cache[absPath]` before re-importing whenever the original specifier is a filesystem path: relative (`./x`, `../x`, `.`, `..`), POSIX-absolute (`/...`), home-prefixed (`~/...`), or Windows drive-letter (`C:\...` / `C:/...`). Bare specifiers (`react`, `lodash/x`) and URL/scheme specifiers (`node:fs`, `file://...`, `https://...`) are left in cache so package identity stays stable across cells. The cache-bust only fires when the resolved target is an absolute path — unresolved bare-package fallbacks (`resolveImportSpecifier()` returning the original specifier) skip it. - The prelude installs globals: - `display`, `print` - `read`, `write`, `append`, `sort`, `uniq`, `counter`, `diff`, `tree`, `env`, `output` @@ -155,7 +144,7 @@ Implemented in `packages/coding-agent/src/eval/py/executor.ts`, `packages/coding - Default mode is retained `session` kernels keyed by `python:${sessionId}` - Optional `python.kernelMode = "per-call"` creates a fresh kernel for each cell and shuts it down afterward -- `rst` disposes the retained kernel for that session before the cell runs; later Python cells in the same tool call reuse the fresh kernel +- `reset: true` disposes the retained kernel for that session before the cell runs; later Python cells in the same tool call reuse the fresh kernel - Startup path: - availability check - create/connect kernel @@ -177,8 +166,8 @@ Implemented in `packages/coding-agent/src/eval/py/executor.ts`, `packages/coding A single tool call can mix Python and JS cells. Persistence is per language runtime: -- resetting Python does not touch JS state -- resetting JS does not touch Python state +- `reset: true` on a Python cell does not touch JS state +- `reset: true` on a JS cell does not touch Python state - each backend keeps its own retained session keyed from the same session-derived ID ## Side Effects @@ -206,8 +195,9 @@ A single tool call can mix Python and JS cells. Persistence is per language runt ## Limits & Caps -- Per-cell timeout default: 30s (`DEFAULT_TIMEOUT_MS` in `packages/coding-agent/src/eval/parse.ts`; `TOOL_TIMEOUTS.eval.default` in `packages/coding-agent/src/tools/tool-timeouts.ts`) -- Timeout clamp: 1s minimum, 600s maximum (`TOOL_TIMEOUTS.eval` in `packages/coding-agent/src/tools/tool-timeouts.ts`) +- Per-cell timeout default: 30s (applied when `timeout` is omitted in `EvalTool.execute()`; clamped through `TOOL_TIMEOUTS.eval.default` in `packages/coding-agent/src/tools/tool-timeouts.ts`) +- Schema-level `timeout` range: integer `1..600` seconds (enforced by Zod on the cell schema) +- Timeout clamp at runtime: 1s minimum, 600s maximum (`TOOL_TIMEOUTS.eval` in `packages/coding-agent/src/tools/tool-timeouts.ts`) - Transcript code/output preview: 10 lines by default (`EVAL_DEFAULT_PREVIEW_LINES` in `packages/coding-agent/src/tools/eval.ts`) - Output truncation window: 50KB default (`DEFAULT_MAX_BYTES` in `packages/coding-agent/src/session/streaming-output.ts`) - Output line cap inside truncation helpers: 3000 lines (`DEFAULT_MAX_LINES` in `packages/coding-agent/src/session/streaming-output.ts`) @@ -222,24 +212,22 @@ A single tool call can mix Python and JS cells. Persistence is per language runt ## Errors -- Parse errors from `parseEvalInput()` throw immediately, for example invalid timeout strings. +- Zod validation rejects malformed `cells` arrays before `execute()` runs (missing `language`/`code`, out-of-range `timeout`, empty `cells`). - Missing session without proxy executor throws `ToolError("Eval tool requires a session when not using proxy executor")`. - Disabled/unavailable backends throw `ToolError` from `resolveBackend()`: - - `eval.py = false` - - `eval.js = false` - - Python kernel unavailable - - no backend available + - `eval.py = false` and a `py` cell is requested + - `eval.js = false` and a `js` cell is requested + - Python kernel unavailable and a `py` cell is requested - JS runtime exceptions are converted into text output plus `exitCode: 1`; cancellations return `cancelled: true` and may append `Command timed out`. - Python execution errors from the kernel become text output and `exitCode: 1`; later cells are skipped. - Python stdin requests are treated as errors with the message `Kernel requested stdin; interactive input is not supported.` - Cancellation is returned, not thrown, once backend execution has started. The tool formats it as a cell failure and sets `details.isError = true`. -- If parsing encountered `*** Abort`, the final text appends `ABORT_WARNING`, explicitly telling the model that earlier cells ran and state persists. - If output truncates, the tool still succeeds; truncation is surfaced through `details.meta` and artifact-backed full output when available. ## Notes -- The runtime parser is intentionally more permissive than `packages/coding-agent/src/eval/eval.lark`; maintain both when changing syntax. -- Cell language in `ParsedEvalCell` is not the last word: `EvalTool.execute()` may override backend selection for cells without an explicit header by inheriting the previous runtime language. +- Backend selection is now strictly explicit per cell: `language` must be `"py"` or `"js"`. The previous `*** Cell` header parser, the `eval.lark` constrained grammar, and the sniffer-based fallback have all been removed. +- `EvalTool.customFormat` no longer exists. Tool calls flow through the standard JSON schema; there is no Lark-constrained sampling path. - `tool.<name>()` exists only in JS. Python prelude helpers do not call back into the full tool registry. - JS helper paths reject protocol URIs (`://`) in `resolvePath()`; the JS prelude is filesystem-only unless the code calls `tool.read(...)` or another tool explicitly. - Python helper `output(...)` depends on `PI_SESSION_FILE`; it fails outside a session-backed run. diff --git a/docs/tools/lsp.md b/docs/tools/lsp.md index 7a4d273d8..cfc97551c 100644 --- a/docs/tools/lsp.md +++ b/docs/tools/lsp.md @@ -310,4 +310,5 @@ Same as `definition`, but sends `textDocument/implementation` and reports `imple - `reload` does not recreate a client immediately after killing it; the next request triggers reinitialization. - `workspace/applyEdit` can apply edits initiated by the server outside the direct tool action result path. - `detectLspmux()` can be disabled with `PI_DISABLE_LSPMUX=1`; only `rust-analyzer` is in `DEFAULT_SUPPORTED_SERVERS`. +- Startup LSP warmup (`discoverStartupLspServers(cwd)` in `sdk.ts`) is gated on `enableLsp && options.hasUI && settings.get("lsp.diagnosticsOnWrite")` — print/RPC/ACP/script sessions skip it and let `getOrCreateClient()` cold-start servers on demand. See `docs/sdk.md` § Startup performance. - `configCache` is per-process and never auto-invalidated; config changes require a fresh process to be observed by `getConfig()` callers. \ No newline at end of file diff --git a/docs/ttsr-injection-lifecycle.md b/docs/ttsr-injection-lifecycle.md index e8f25f846..3fa7047bf 100644 --- a/docs/ttsr-injection-lifecycle.md +++ b/docs/ttsr-injection-lifecycle.md @@ -108,7 +108,27 @@ Pending injections are cleared after content generation. ### Non-interrupting matches -If matched rules do not permit interruption (`interruptMode: "never"`, or source-specific `prose-only`/`tool-only` mismatch), they are still queued. After a successful non-error, non-aborted assistant message, `AgentSession` injects the hidden `ttsr-injection` custom message as a follow-up and schedules continuation. +Non-interrupting matches split by `matchContext.source`: + +- **`source === "tool"` (tool-source match).** The rule is bucketed into `#perToolTtsrInjections`, keyed by the matched tool call's `id`. There is **no** deferred follow-up turn and the stream is not aborted. When the tool actually produces a result, the `afterToolCall` hook prepends a rendered `ttsr-tool-reminder.md` block to `ctx.result.content` (a single `text` block inserted ahead of the tool's own content), and persists a `ttsr_injection` entry with the consumed rule names. The template payload is: + + ```xml + <system-reminder reason="rule_violation" rule="{{name}}" path="{{path}}"> + ... + {{content}} + </system-reminder> + ``` + +- **`source === "text"` / `"thinking"` (prose-source match).** Behavior is unchanged: the rule is queued in `#pendingTtsrInjections` and, after a successful non-error, non-aborted assistant message, `AgentSession` injects the hidden `ttsr-injection` custom message as a follow-up and schedules continuation. + +Within a single matching batch, each rule is attached to exactly one sibling tool call — if multiple sibling tool calls would satisfy the same rule, deduplication picks one and the others are left untouched. Multiple distinct rules can still fold onto the same tool call. + +#### Implications for tool authors and transcript readers + +- The tool's own `toolResult` content is preserved verbatim; the reminder is **prepended** as an additional leading text block. Renderers that assume `content[0]` is the tool's primary output must scan past any block whose text begins with `<system-reminder reason="rule_violation"` (or filter on the wrapper tag) to find the real payload. +- The reminder is in-band on the tool result, not a separate `custom_message`/`ttsr-injection` entry. Transcript readers looking for non-interrupting TTSR activity on tool-source rules MUST inspect tool results (and the persisted `ttsr_injection` entry list), not just synthetic injection entries. +- A single tool result may carry reminders for several rules concatenated with a blank line between rendered templates. +- If the assistant message ends with `stopReason === "aborted"` or `"error"` before the matched tools run, the pending per-tool buckets are cleared — those rules are **not** persisted as injected and remain eligible to re-trigger on a future turn (subject to repeat policy). ## 5. Repeat policy and gap logic @@ -169,7 +189,8 @@ Interactive mode uses `session.isTtsrAbortPending` to suppress showing the abort In the current runtime path: - interrupted injections append a hidden `custom_message` with `customType: "ttsr-injection"` and append a `ttsr_injection` entry via `appendTtsrInjection(...)` -- deferred non-interrupting injections are marked/persisted when their queued custom message reaches `message_end` +- deferred non-interrupting prose-source injections are marked/persisted when their queued custom message reaches `message_end` +- non-interrupting tool-source injections are marked at match time and persisted via `appendTtsrInjection(...)` from the `afterToolCall` hook when the matched tool's result is produced - `createAgentSession()` restores `existingSession.injectedTtsrRules` into `ttsrManager` Net effect: injected-rule suppression is persisted/restored across session reload/resume for the current branch path. @@ -196,5 +217,6 @@ During the timer window, state can change (user interruption, mode actions, addi - Duplicate rule names at capability layer: lower-priority duplicates are shadowed before registration. - Duplicate names at manager layer: second registration is ignored. - `contextMode: "keep"`: partial violating output can remain in context before reminder retry. -- `interruptMode: "never"` queues a deferred hidden injection after a successful assistant message rather than aborting mid-stream. +- `interruptMode: "never"`: prose-source matches queue a deferred hidden injection after a successful assistant message; tool-source matches fold an in-band `<system-reminder>` into the matched tool call's `toolResult` content via the `afterToolCall` hook (no mid-stream abort, no separate follow-up turn). +- Tool-source non-interrupting buckets are cleared when the parent assistant message ends with `stopReason === "aborted"` or `"error"`, so rules whose target tool never produced a result remain eligible to re-trigger. - Repeat-after-gap depends on turn count increments at `turn_end`; mid-turn chunks do not advance gap counters. diff --git a/packages/coding-agent/package.json b/packages/coding-agent/package.json index e438be0fb..2abe9a206 100644 --- a/packages/coding-agent/package.json +++ b/packages/coding-agent/package.json @@ -539,7 +539,6 @@ "types": "./src/web/search/providers/*.ts", "import": "./src/web/search/providers/*.ts" }, - "./*.js": "./src/*.ts" } }