feat(cross-cutting): added multi-syntax in-band tool-call support for runtime tool conversion
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls. - Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3. - Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax. - Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
This commit is contained in:
@@ -0,0 +1,630 @@
|
|||||||
|
# Anthropic Claude tool use (Messages API content blocks)
|
||||||
|
|
||||||
|
Anthropic's Claude is a closed, hosted model family; there are no released weights and therefore no `--tool-call-parser` flag to set. The canonical tool-calling convention is the **Messages API** (`POST /v1/messages`, header `anthropic-version: 2023-06-01`): tools are advertised in a top-level `tools` array, the model returns structured `tool_use` **content blocks** with `stop_reason: "tool_use"`, and you feed results back as `tool_result` content blocks inside a `user` message. Tool use is "enabled" simply by including the `tools` parameter (optionally with `tool_choice`); the API then injects a tool-use system prompt and parses the model's output back into JSON blocks for you. This applies to all current models (Claude Opus / Sonnet / Haiku 3.x, 4, 4.x) and is mirrored by gateways such as LiteLLM and by third-party Claude-compatible servers.
|
||||||
|
|
||||||
|
Under the hood the model is trained to emit an **XML** function-call syntax (`<function_calls>` / `<invoke>` / `<parameter>`); the API serializes your JSON-Schema tools into a system prompt and converts the model's XML output into JSON `tool_use` blocks. That underlying format is documented as the *secondary* convention below, together with the older, now-retired prompt-based **legacy XML** format (`<tool_name>` / `<parameters>` / `<function_results>`) that pre-dates the Messages API and still surfaces when you do tool use purely through prompting.
|
||||||
|
|
||||||
|
The primary, authoritative shape for any parser/renderer is the JSON content-block format. The XML is informational (and the only thing visible if you reconstruct prompts at the token level).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Content-block types & stop reasons
|
||||||
|
|
||||||
|
Anthropic has no token-level tool delimiters in the public API. The unit is the **content block**: every `message.content` is an array of typed blocks. Tool calling adds two block types and one stop reason; streaming adds a delta type.
|
||||||
|
|
||||||
|
| Item | Where | Shape / meaning |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `text` block | assistant & user | `{"type":"text","text":"..."}`. Plain prose. Assistant may emit text *before* its tool calls. |
|
||||||
|
| `tool_use` block | assistant | `{"type":"tool_use","id":"toolu_...","name":"<tool>","input":{...}}`. The function call. `input` is a **nested JSON object** (already parsed), conforming to the tool's `input_schema`. |
|
||||||
|
| `tool_result` block | user | `{"type":"tool_result","tool_use_id":"toolu_...","content":<string \| block[]>,"is_error":<bool?>}`. The executed result, sent back in a `user` message. |
|
||||||
|
| `server_tool_use` block | assistant | `{"type":"server_tool_use","id":"srvtoolu_...","name":"web_search","input":{...}}`. Emitted for Anthropic-executed server tools; you do **not** return a `tool_result` for these. |
|
||||||
|
| `web_search_tool_result` (and similar) | assistant | Server-tool output, injected by Anthropic inline in the assistant turn. |
|
||||||
|
| `thinking` / `redacted_thinking` block | assistant | Extended-thinking reasoning blocks; carry a `signature`. Must be preserved verbatim across turns when thinking + tools are combined. |
|
||||||
|
| `stop_reason: "tool_use"` | response top level | The model invoked one or more tools and is waiting for results. Drives the agentic loop. |
|
||||||
|
| `stop_reason: "end_turn"` | response top level | Natural completion (no tool call); the loop exits. |
|
||||||
|
| Other `stop_reason` | response top level | `"max_tokens"`, `"stop_sequence"`, `"pause_turn"` (long server-tool turn, resend as-is to continue), `"refusal"`. |
|
||||||
|
| `id` prefixes | — | Messages `msg_…`; client tool calls `toolu_…`; server tool calls `srvtoolu_…`. |
|
||||||
|
|
||||||
|
Streaming adds these SSE events / delta types (full list under [Roles / channels](#roles--channels--turn-structure) and [Tool-call format](#tool-call-format)):
|
||||||
|
|
||||||
|
| Streaming item | Shape / meaning |
|
||||||
|
| --- | --- |
|
||||||
|
| `message_start` | Carries a `Message` skeleton with empty `content`, `stop_reason: null`. |
|
||||||
|
| `content_block_start` | Opens a block at `index`. For a tool call: `content_block.{type:"tool_use",id,name,input:{}}` — `input` starts as an **empty object**. |
|
||||||
|
| `content_block_delta` / `input_json_delta` | `{"type":"input_json_delta","partial_json":"<chunk>"}` — a **partial JSON string** fragment of `tool_use.input`. |
|
||||||
|
| `content_block_delta` / `text_delta` | `{"type":"text_delta","text":"..."}`. |
|
||||||
|
| `content_block_delta` / `thinking_delta`, `signature_delta` | Extended-thinking content / signature. |
|
||||||
|
| `content_block_stop` | Closes the block at `index`; this is when accumulated `partial_json` is complete and safe to `JSON.parse`. |
|
||||||
|
| `message_delta` | Top-level updates; carries the final `delta.stop_reason` (e.g. `"tool_use"`) and **cumulative** `usage`. |
|
||||||
|
| `message_stop` | End of stream. |
|
||||||
|
| `ping` / `error` | Keep-alive; `error` (e.g. `overloaded_error`) may appear mid-stream. |
|
||||||
|
|
||||||
|
### Legacy XML tags (prompt-based, pre-Messages-API)
|
||||||
|
|
||||||
|
The retired prompt-based format used these tags. They are nested-element tags (no attributes), distinct from the modern attribute form (`<invoke name="…">`). Verified against Anthropic's archived "Legacy tool use" doc (see [Sources](#sources)).
|
||||||
|
|
||||||
|
| Tag | Role | Notes |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `<tools>` … `</tools>` | tool advertising | Container in the system prompt wrapping all `<tool_description>` entries. |
|
||||||
|
| `<tool_description>` | tool advertising | One per tool: holds `<tool_name>`, `<description>`, `<parameters>`. |
|
||||||
|
| `<tool_name>` | both | Function name (used in definitions, calls, and results). |
|
||||||
|
| `<parameters>` / `<parameter>` | definition | `<parameters>` wraps `<parameter>` entries, each with `<name>`, `<type>`, `<description>`. |
|
||||||
|
| `<function_calls>` | model output | Wraps one or more `<invoke>` blocks. |
|
||||||
|
| `<invoke>` | model output | One function call; contains `<tool_name>` + a `<parameters>` block of `<paramName>value</paramName>` child tags. |
|
||||||
|
| `<function_results>` | tool result (fed back) | Wraps `<result>` (success) or `<error>` (failure). |
|
||||||
|
| `<result>` / `<stdout>` | tool result | `<result>` holds `<tool_name>` + `<stdout>`; the output text goes in `<stdout>`. |
|
||||||
|
| `<error>` | tool result | Replaces `<result>` when the function raised. |
|
||||||
|
| `</function_calls>` | stop sequence | Passed as `stop_sequence` so generation halts after a call. |
|
||||||
|
| `<scratchpad>` / `<answer>` | model output | Conventionally used for chain-of-thought and final answer in legacy prompts. |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Roles / channels / turn structure
|
||||||
|
|
||||||
|
The Messages API uses only two conversational roles, `user` and `assistant`, alternating. There is **no** dedicated `tool`/`function` role and **no** top-level `system` role — the system prompt is a separate top-level `system` parameter (string or text-block array). Tool data rides inside the normal roles:
|
||||||
|
|
||||||
|
- `assistant` messages contain AI-generated `text`, `thinking`, and `tool_use` (and `server_tool_use`) blocks.
|
||||||
|
- `user` messages contain your `text`/`image`/`document` content and `tool_result` blocks.
|
||||||
|
|
||||||
|
There are no named "channels". The closest analogue to a reasoning channel is the extended-thinking `thinking` content block (a first-class block with a cryptographic `signature`), kept separate from the user-visible `text` block. When thinking is enabled alongside tools, the `thinking` block(s) from a tool-calling turn must be passed back unmodified in the follow-up request.
|
||||||
|
|
||||||
|
The agentic loop is keyed on `stop_reason`:
|
||||||
|
|
||||||
|
1. Send `tools` + the user message.
|
||||||
|
2. Claude responds with `stop_reason: "tool_use"` and one or more `tool_use` blocks (optionally preceded by a `text` block).
|
||||||
|
3. Execute each tool; build a `tool_result` block per call.
|
||||||
|
4. Append the assistant message **and** a `user` message carrying all `tool_result` blocks; resend.
|
||||||
|
5. Repeat while `stop_reason == "tool_use"`; exit on `end_turn` (or another terminal reason).
|
||||||
|
|
||||||
|
Strict ordering rules (a 400 otherwise):
|
||||||
|
- `tool_result` blocks must come **first** in the `user` message's `content` array (any text after them).
|
||||||
|
- The `tool_result` `user` message must **immediately follow** the assistant `tool_use` message — nothing in between.
|
||||||
|
- Every `tool_use.id` must be answered by a `tool_result.tool_use_id` in that next message.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tool definitions
|
||||||
|
|
||||||
|
Tools are passed in the top-level `tools` array. Each user-defined (client) tool is a **flat** object — no `{"type":"function", "function":{…}}` wrapper (that wrapper is OpenAI's). Fields:
|
||||||
|
|
||||||
|
- `name` — matches `^[a-zA-Z0-9_-]{1,64}$`.
|
||||||
|
- `description` — detailed plaintext (the single biggest driver of tool-call quality).
|
||||||
|
- `input_schema` — a JSON Schema object (**not** `parameters`) describing the input the model must produce.
|
||||||
|
- Optional: `input_examples`, `cache_control`, `strict`, `defer_loading`, `allowed_callers`.
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"name": "get_weather",
|
||||||
|
"description": "Get the current weather in a given location",
|
||||||
|
"input_schema": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"location": {
|
||||||
|
"type": "string",
|
||||||
|
"description": "The city and state, e.g. San Francisco, CA"
|
||||||
|
},
|
||||||
|
"unit": {
|
||||||
|
"type": "string",
|
||||||
|
"enum": ["celsius", "fahrenheit"],
|
||||||
|
"description": "The unit of temperature, either 'celsius' or 'fahrenheit'"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"required": ["location"]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Anthropic-schema client tools (`bash`, `text_editor`, `computer`, `memory`) and server tools (`web_search`, `web_fetch`, `code_execution`, `tool_search`) instead carry a versioned `type`, e.g. `{"type": "web_search_20250305", "name": "web_search"}`.
|
||||||
|
|
||||||
|
`tool_choice` controls invocation (four options):
|
||||||
|
- `{"type":"auto"}` — model decides (default when `tools` present).
|
||||||
|
- `{"type":"any"}` — must call some tool.
|
||||||
|
- `{"type":"tool","name":"get_weather"}` — must call that specific tool.
|
||||||
|
- `{"type":"none"}` — no tools (default when no `tools`).
|
||||||
|
|
||||||
|
With `any` or `tool` the API prefills the assistant turn, so no leading natural-language text precedes the `tool_use` block. Add `"disable_parallel_tool_use": true` inside `tool_choice` to cap at one tool per turn. (Extended thinking only supports `auto`/`none`.)
|
||||||
|
|
||||||
|
### How the API turns this into a prompt (the bridge to XML)
|
||||||
|
|
||||||
|
When `tools` is present, the API constructs a tool-use system prompt with this skeleton (verified from "Define tools"):
|
||||||
|
|
||||||
|
```text
|
||||||
|
In this environment you have access to a set of tools you can use to answer the user's question.
|
||||||
|
{{ FORMATTING INSTRUCTIONS }}
|
||||||
|
String and scalar parameters should be specified as is, while lists and objects should use JSON format. Note that spaces for string values are not stripped. The output is not expected to be valid XML and is parsed with regular expressions.
|
||||||
|
Here are the functions available in JSONSchema format:
|
||||||
|
{{ TOOL DEFINITIONS IN JSON SCHEMA }}
|
||||||
|
{{ USER SYSTEM PROMPT }}
|
||||||
|
{{ TOOL CONFIGURATION }}
|
||||||
|
```
|
||||||
|
|
||||||
|
`{{ TOOL DEFINITIONS IN JSON SCHEMA }}` is your `tools` array serialized to JSON Schema. `{{ FORMATTING INSTRUCTIONS }}` is the (unpublished) block teaching the model the `<function_calls>`/`<invoke name>`/`<parameter name>` syntax shown under [Tool-call format → underlying XML](#underlying-xml-modern-attribute-form). The note "parsed with regular expressions" is why output need not be well-formed XML.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tool-call format
|
||||||
|
|
||||||
|
The wire format your application consumes is JSON. A single call is one `tool_use` content block in the assistant message, with `stop_reason: "tool_use"` at the top level:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "msg_01Aq9w938a90dw8q",
|
||||||
|
"type": "message",
|
||||||
|
"role": "assistant",
|
||||||
|
"model": "claude-opus-4-8",
|
||||||
|
"content": [
|
||||||
|
{
|
||||||
|
"type": "text",
|
||||||
|
"text": "I'll check the current weather in San Francisco for you."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "tool_use",
|
||||||
|
"id": "toolu_01A09q90qw90lq917835lq9",
|
||||||
|
"name": "get_weather",
|
||||||
|
"input": { "location": "San Francisco, CA", "unit": "celsius" }
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"stop_reason": "tool_use",
|
||||||
|
"stop_sequence": null,
|
||||||
|
"usage": { "input_tokens": 472, "output_tokens": 65 }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Key facts for a parser:
|
||||||
|
- `tool_use.input` is an already-parsed **object**, never a JSON string.
|
||||||
|
- A leading `text` block is optional and informational; do not rely on its wording.
|
||||||
|
- Match calls to results by `id` → `tool_use_id`.
|
||||||
|
|
||||||
|
### Underlying XML (modern attribute form)
|
||||||
|
|
||||||
|
Before the API converts it, the model literally emits an XML block. The current (Claude 3+) form is attribute-based:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<function_calls>
|
||||||
|
<invoke name="get_weather">
|
||||||
|
<parameter name="location">San Francisco, CA</parameter>
|
||||||
|
<parameter name="unit">celsius</parameter>
|
||||||
|
</invoke>
|
||||||
|
</function_calls>
|
||||||
|
```
|
||||||
|
|
||||||
|
`[Partially verified]` Anthropic does not publish the literal `{{ FORMATTING INSTRUCTIONS }}`, so the exact tag spelling for current models is reconstructed from the trained format (and matches the task's reference anchor) rather than an official verbatim doc. In production, current Claude models prefix these tags with an `antml:` XML namespace (e.g. `<function_calls>`, `<invoke name="…">`, `<parameter name="…">`); the namespace is widely observed but **not** documented officially — treat it as `[unverified]`. The API strips all of this and exposes only the JSON `tool_use` block; integrators should target the JSON, not the XML.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Multiple / parallel tool calls
|
||||||
|
|
||||||
|
Parallel calls are the default. Claude emits **multiple `tool_use` blocks in a single assistant message**:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"role": "assistant",
|
||||||
|
"content": [
|
||||||
|
{ "type": "text", "text": "Let me check both cities." },
|
||||||
|
{
|
||||||
|
"type": "tool_use",
|
||||||
|
"id": "toolu_01weather_sf",
|
||||||
|
"name": "get_weather",
|
||||||
|
"input": { "location": "San Francisco, CA" }
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "tool_use",
|
||||||
|
"id": "toolu_02weather_nyc",
|
||||||
|
"name": "get_weather",
|
||||||
|
"input": { "location": "New York, NY" }
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
You return **all** results in **one** `user` message, one `tool_result` per call, results first:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"role": "user",
|
||||||
|
"content": [
|
||||||
|
{
|
||||||
|
"type": "tool_result",
|
||||||
|
"tool_use_id": "toolu_01weather_sf",
|
||||||
|
"content": "San Francisco: 68F, partly cloudy"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "tool_result",
|
||||||
|
"tool_use_id": "toolu_02weather_nyc",
|
||||||
|
"content": "New York: 45F, clear skies"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Calls in one turn are **unordered** and may be run concurrently. If two batched calls turn out to depend on each other, return the natural error in a `tool_result` with `"is_error": true`; Claude reissues the dependent call on a later turn. (In the legacy XML format, parallelism is multiple `<invoke>` blocks inside one `<function_calls>`.)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tool-result format
|
||||||
|
|
||||||
|
A result is a `tool_result` block inside a `user` message:
|
||||||
|
|
||||||
|
- `tool_use_id` (required) — the `id` of the `tool_use` it answers.
|
||||||
|
- `content` (optional) — a string, **or** an array of `text`/`image`/`document` blocks. Omit for an empty result.
|
||||||
|
- `is_error` (optional) — `true` for execution failures; put a useful message in `content`.
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"role": "user",
|
||||||
|
"content": [
|
||||||
|
{
|
||||||
|
"type": "tool_result",
|
||||||
|
"tool_use_id": "toolu_01A09q90qw90lq917835lq9",
|
||||||
|
"content": "15 degrees"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Error result:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"role": "user",
|
||||||
|
"content": [
|
||||||
|
{
|
||||||
|
"type": "tool_result",
|
||||||
|
"tool_use_id": "toolu_01A09q90qw90lq917835lq9",
|
||||||
|
"content": "ConnectionError: the weather service API is not available (HTTP 500)",
|
||||||
|
"is_error": true
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Rich result (text + image blocks):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"role": "user",
|
||||||
|
"content": [
|
||||||
|
{
|
||||||
|
"type": "tool_result",
|
||||||
|
"tool_use_id": "toolu_01A09q90qw90lq917835lq9",
|
||||||
|
"content": [
|
||||||
|
{ "type": "text", "text": "15 degrees" },
|
||||||
|
{
|
||||||
|
"type": "image",
|
||||||
|
"source": { "type": "base64", "media_type": "image/jpeg", "data": "/9j/4AAQSkZJRg..." }
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Server tools require **no** `tool_result` from you — Anthropic executes them and injects the result inline in the assistant turn. (Legacy XML feeds results back as `<function_results><result><tool_name>…</tool_name><stdout>…</stdout></result></function_results>`, or `<error>…</error>` on failure.)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## End-to-end example
|
||||||
|
|
||||||
|
A complete multi-turn weather exchange. All JSON is valid.
|
||||||
|
|
||||||
|
**Request 1 — system + tools + user question:**
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"model": "claude-opus-4-8",
|
||||||
|
"max_tokens": 1024,
|
||||||
|
"system": "You are a helpful weather assistant. Use the provided tools to answer.",
|
||||||
|
"tools": [
|
||||||
|
{
|
||||||
|
"name": "get_weather",
|
||||||
|
"description": "Get the current weather in a given location",
|
||||||
|
"input_schema": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"location": { "type": "string", "description": "The city and state, e.g. San Francisco, CA" },
|
||||||
|
"unit": { "type": "string", "enum": ["celsius", "fahrenheit"], "description": "Unit for the temperature" }
|
||||||
|
},
|
||||||
|
"required": ["location"]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"messages": [
|
||||||
|
{ "role": "user", "content": "What's the weather in San Francisco?" }
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response 1 — assistant requests the tool (`stop_reason: "tool_use"`):**
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "msg_01Aq9w938a90dw8q",
|
||||||
|
"type": "message",
|
||||||
|
"role": "assistant",
|
||||||
|
"model": "claude-opus-4-8",
|
||||||
|
"content": [
|
||||||
|
{ "type": "text", "text": "I'll check the current weather in San Francisco for you." },
|
||||||
|
{
|
||||||
|
"type": "tool_use",
|
||||||
|
"id": "toolu_01A09q90qw90lq917835lq9",
|
||||||
|
"name": "get_weather",
|
||||||
|
"input": { "location": "San Francisco, CA", "unit": "celsius" }
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"stop_reason": "tool_use",
|
||||||
|
"stop_sequence": null,
|
||||||
|
"usage": { "input_tokens": 472, "output_tokens": 65 }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Request 2 — replay history, append the assistant turn and the `tool_result`:**
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"model": "claude-opus-4-8",
|
||||||
|
"max_tokens": 1024,
|
||||||
|
"system": "You are a helpful weather assistant. Use the provided tools to answer.",
|
||||||
|
"tools": [
|
||||||
|
{
|
||||||
|
"name": "get_weather",
|
||||||
|
"description": "Get the current weather in a given location",
|
||||||
|
"input_schema": {
|
||||||
|
"type": "object",
|
||||||
|
"properties": {
|
||||||
|
"location": { "type": "string", "description": "The city and state, e.g. San Francisco, CA" },
|
||||||
|
"unit": { "type": "string", "enum": ["celsius", "fahrenheit"], "description": "Unit for the temperature" }
|
||||||
|
},
|
||||||
|
"required": ["location"]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"messages": [
|
||||||
|
{ "role": "user", "content": "What's the weather in San Francisco?" },
|
||||||
|
{
|
||||||
|
"role": "assistant",
|
||||||
|
"content": [
|
||||||
|
{ "type": "text", "text": "I'll check the current weather in San Francisco for you." },
|
||||||
|
{
|
||||||
|
"type": "tool_use",
|
||||||
|
"id": "toolu_01A09q90qw90lq917835lq9",
|
||||||
|
"name": "get_weather",
|
||||||
|
"input": { "location": "San Francisco, CA", "unit": "celsius" }
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"role": "user",
|
||||||
|
"content": [
|
||||||
|
{
|
||||||
|
"type": "tool_result",
|
||||||
|
"tool_use_id": "toolu_01A09q90qw90lq917835lq9",
|
||||||
|
"content": "15 degrees Celsius, partly cloudy"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Response 2 — assistant's final answer (`stop_reason: "end_turn"`):**
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"id": "msg_01EeFG3hijk2lmno4PqrSt",
|
||||||
|
"type": "message",
|
||||||
|
"role": "assistant",
|
||||||
|
"model": "claude-opus-4-8",
|
||||||
|
"content": [
|
||||||
|
{ "type": "text", "text": "It's currently 15 degrees Celsius and partly cloudy in San Francisco." }
|
||||||
|
],
|
||||||
|
"stop_reason": "end_turn",
|
||||||
|
"stop_sequence": null,
|
||||||
|
"usage": { "input_tokens": 530, "output_tokens": 18 }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Streaming (SSE) shape of the tool call
|
||||||
|
|
||||||
|
The same tool call, streamed. Note `tool_use` opens with an empty `input`, the arguments arrive as `input_json_delta.partial_json` fragments, and the final `stop_reason` lands in `message_delta`. This block is reproduced verbatim from Anthropic's streaming docs:
|
||||||
|
|
||||||
|
```text
|
||||||
|
event: message_start
|
||||||
|
data: {"type":"message_start","message":{"id":"msg_014p7gG3wDgGV9EUtLvnow3U","type":"message","role":"assistant","model":"claude-opus-4-8","stop_sequence":null,"usage":{"input_tokens":472,"output_tokens":2},"content":[],"stop_reason":null}}
|
||||||
|
|
||||||
|
event: content_block_start
|
||||||
|
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
|
||||||
|
|
||||||
|
event: ping
|
||||||
|
data: {"type": "ping"}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Okay"}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":" let"}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"'s"}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":" check"}}
|
||||||
|
|
||||||
|
event: content_block_stop
|
||||||
|
data: {"type":"content_block_stop","index":0}
|
||||||
|
|
||||||
|
event: content_block_start
|
||||||
|
data: {"type":"content_block_start","index":1,"content_block":{"type":"tool_use","id":"toolu_01T1x1fJ34qAmk2tNTrN7Up6","name":"get_weather","input":{}}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":""}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":"{\"location\":"}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":" \"San"}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":" Francisc"}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":"o,"}}
|
||||||
|
|
||||||
|
event: content_block_delta
|
||||||
|
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":" CA\"}"}}
|
||||||
|
|
||||||
|
event: content_block_stop
|
||||||
|
data: {"type":"content_block_stop","index":1}
|
||||||
|
|
||||||
|
event: message_delta
|
||||||
|
data: {"type":"message_delta","delta":{"stop_reason":"tool_use","stop_sequence":null},"usage":{"output_tokens":89}}
|
||||||
|
|
||||||
|
event: message_stop
|
||||||
|
data: {"type":"message_stop"}
|
||||||
|
```
|
||||||
|
|
||||||
|
Reassembly: concatenate every `partial_json` for a given `index` (`"" + "{\"location\":" + " \"San" + " Francisc" + "o," + " CA\"}"` → `{"location": "San Francisco, CA"}`), then `JSON.parse` at that block's `content_block_stop`. Tool use also supports fine-grained streaming (`eager_input_streaming` per tool) for finer `partial_json` chunking.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## OpenAI-compatible API mapping
|
||||||
|
|
||||||
|
Anthropic integrates tools into the `user`/`assistant` message structure rather than using OpenAI's separate `tool` role and `function` wrapper. Field-by-field:
|
||||||
|
|
||||||
|
| Concept | Anthropic Messages API | OpenAI Chat Completions |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| Tool definition wrapper | flat `{"name","description","input_schema"}` in `tools[]` | `{"type":"function","function":{"name","description","parameters"}}` in `tools[]` |
|
||||||
|
| Tool schema key | `input_schema` (JSON Schema) | `parameters` (JSON Schema) |
|
||||||
|
| "Must call a tool" | `tool_choice:{"type":"any"}` / `{"type":"tool","name":…}` | `tool_choice:"required"` / `{"type":"function","function":{"name":…}}` |
|
||||||
|
| Disable parallel calls | `tool_choice:{…,"disable_parallel_tool_use":true}` | `parallel_tool_calls:false` (top level) |
|
||||||
|
| Assistant call container | `tool_use` **content block** in `content[]` | `tool_calls[]` on the assistant `message` |
|
||||||
|
| Call id | `tool_use.id` = `toolu_…` | `tool_calls[].id` = `call_…` |
|
||||||
|
| Function name | `tool_use.name` | `tool_calls[].function.name` |
|
||||||
|
| Function arguments | `tool_use.input` = **nested JSON object** (parsed) | `tool_calls[].function.arguments` = **JSON string** (must `JSON.parse`) |
|
||||||
|
| "Tools were called" signal | `stop_reason:"tool_use"` | `finish_reason:"tool_calls"` |
|
||||||
|
| Result message role | `user` message containing `tool_result` block(s) | dedicated `{"role":"tool",…}` message(s) |
|
||||||
|
| Result ↔ call linkage | `tool_result.tool_use_id` | `tool` message `tool_call_id` |
|
||||||
|
| Result payload | `tool_result.content` = string **or** block array (text/image/document) | `tool` message `content` = string |
|
||||||
|
| Error result | `tool_result` with `is_error:true` | no dedicated flag; encode in `content` |
|
||||||
|
| System prompt | top-level `system` param (no `system` role) | `{"role":"system",…}` message |
|
||||||
|
| Streamed args | `input_json_delta.partial_json` fragments | `tool_calls[].function.arguments` string deltas |
|
||||||
|
|
||||||
|
Conversion gotchas:
|
||||||
|
- **Object vs string:** to emit OpenAI shape, `JSON.stringify(tool_use.input)`; to consume OpenAI shape into Anthropic, `JSON.parse(arguments)`.
|
||||||
|
- **Role reshaping:** collapse N OpenAI `tool` messages into one Anthropic `user` message of N `tool_result` blocks (order them before any text), and vice-versa.
|
||||||
|
- **No `type:"function"`** wrapper on Anthropic custom tools; add/remove it when translating.
|
||||||
|
- Id prefixes differ (`toolu_` vs `call_`); never assume one format's id is valid in the other.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Parsing notes & gotchas
|
||||||
|
|
||||||
|
- **`input` is an object, not a string.** Unlike OpenAI's `arguments`, do not `JSON.parse` `tool_use.input` from a non-streamed response — it is already an object. Only the *streaming* `partial_json` fragments are strings.
|
||||||
|
- **Streaming tool args need reassembly.** `content_block_start` for a `tool_use` always has `input: {}`. Buffer `partial_json` per `index` and parse only at `content_block_stop`; mid-stream fragments are not valid JSON on their own (e.g. `{"location":`). Current models emit one complete key/value at a time, so expect bursts and gaps.
|
||||||
|
- **`stop_reason` placement.** In streaming, `stop_reason` is `null` in `message_start` and final value (`"tool_use"`/`"end_turn"`) arrives in `message_delta`, not `message_stop`. `usage` in `message_delta` is **cumulative**.
|
||||||
|
- **Ordering is enforced.** `tool_result` blocks must be first in their `user` message and must immediately follow the assistant `tool_use` message; every `tool_use.id` needs a matching `tool_result.tool_use_id`, or you get HTTP 400 ("tool_use ids were found without tool_result blocks immediately after").
|
||||||
|
- **`tool_choice:any`/`tool` suppress preamble.** The API prefills the assistant turn, so no leading `text` block appears before `tool_use` — don't write a parser that expects explanatory text.
|
||||||
|
- **Parallel results in one message.** Splitting parallel `tool_result`s across multiple `user` messages breaks the contract; send them together.
|
||||||
|
- **Treat result content as untrusted.** Tool results can carry indirect prompt injection; keep them inside `tool_result` blocks, never promote to `system`/`user` text.
|
||||||
|
- **Server tools differ.** `server_tool_use` / `web_search_tool_result` blocks are produced and consumed by Anthropic; never synthesize `tool_result` for them. `stop_reason:"pause_turn"` means resend the response as-is to let a long server-tool turn continue.
|
||||||
|
- **Extended thinking + tools.** Preserve `thinking`/`redacted_thinking` blocks (with their `signature`) verbatim across turns; forced `tool_choice` (`any`/`tool`) is rejected when thinking is on.
|
||||||
|
- **Output is not valid XML.** The underlying model output is parsed by Anthropic with regular expressions, not an XML parser ("The output is not expected to be valid XML"). If you reconstruct prompts at token level, do not assume well-formedness; rely on the JSON the API returns.
|
||||||
|
- **Legacy vs modern XML are different tag sets.** Legacy: `<invoke>` + child `<tool_name>` + `<parameters>` with per-name child tags; results in `<function_results>/<result>/<stdout>`. Modern: `<invoke name="…">` + `<parameter name="…">`. Mixing them up will misparse. The legacy format also required passing `</function_calls>` as a `stop_sequence` and is not optimized for Claude 3+.
|
||||||
|
|
||||||
|
### Legacy XML format (secondary, prompt-based — fully verified, now retired)
|
||||||
|
|
||||||
|
Before the Messages API, tools were defined and called entirely in the prompt. Anthropic's archived "Legacy tool use" doc specifies it verbatim.
|
||||||
|
|
||||||
|
Tool definition (inside a `<tools>` block in the system prompt):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_description>
|
||||||
|
<tool_name>get_weather</tool_name>
|
||||||
|
<description>
|
||||||
|
Retrieves the current weather for a specified location.
|
||||||
|
Returns a dictionary with two fields:
|
||||||
|
- temperature: float, the current temperature in Fahrenheit
|
||||||
|
- conditions: string, a brief description of the current weather conditions
|
||||||
|
Raises ValueError if the provided location cannot be found.
|
||||||
|
</description>
|
||||||
|
<parameters>
|
||||||
|
<parameter>
|
||||||
|
<name>location</name>
|
||||||
|
<type>string</type>
|
||||||
|
<description>The city and state, e.g. San Francisco, CA</description>
|
||||||
|
</parameter>
|
||||||
|
</parameters>
|
||||||
|
</tool_description>
|
||||||
|
```
|
||||||
|
|
||||||
|
Model-emitted call (multiple `<invoke>` for parallel calls; pass `</function_calls>` as a `stop_sequence`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<function_calls>
|
||||||
|
<invoke>
|
||||||
|
<tool_name>get_weather</tool_name>
|
||||||
|
<parameters>
|
||||||
|
<location>San Francisco, CA</location>
|
||||||
|
</parameters>
|
||||||
|
</invoke>
|
||||||
|
</function_calls>
|
||||||
|
```
|
||||||
|
|
||||||
|
Result fed back into the next user turn:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<function_results>
|
||||||
|
<result>
|
||||||
|
<tool_name>get_weather</tool_name>
|
||||||
|
<stdout>
|
||||||
|
59 degrees Fahrenheit, partly cloudy
|
||||||
|
</stdout>
|
||||||
|
</result>
|
||||||
|
</function_results>
|
||||||
|
```
|
||||||
|
|
||||||
|
Error result:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<function_results>
|
||||||
|
<error>
|
||||||
|
error message goes here
|
||||||
|
</error>
|
||||||
|
</function_results>
|
||||||
|
```
|
||||||
|
|
||||||
|
The legacy system-prompt preamble (verbatim from the archived doc) was:
|
||||||
|
|
||||||
|
```text
|
||||||
|
In this environment you have access to a set of tools you can use to answer the user's question.
|
||||||
|
You may call them like this:
|
||||||
|
<function_calls>
|
||||||
|
<invoke>
|
||||||
|
<tool_name>$TOOL_NAME</tool_name>
|
||||||
|
<parameters>
|
||||||
|
<$PARAMETER_NAME>$PARAMETER_VALUE</$PARAMETER_NAME>
|
||||||
|
...
|
||||||
|
</parameters>
|
||||||
|
</invoke>
|
||||||
|
</function_calls>
|
||||||
|
|
||||||
|
Here are the tools available:
|
||||||
|
<tools>
|
||||||
|
...one <tool_description> per tool...
|
||||||
|
</tools>
|
||||||
|
```
|
||||||
|
|
||||||
|
Legacy notes: no built-in tools (everything is prompt-defined); Anthropic recommended ≤3–5 tools; the model conventionally wrapped reasoning in `<scratchpad>` and final output in `<answer>`. This format is "out of date" and "not optimized for Claude 3" — use the JSON Messages API for anything current.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- Tool use overview — https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
|
||||||
|
- How tool use works — https://docs.claude.com/en/docs/agents-and-tools/tool-use/how-tool-use-works
|
||||||
|
- Define tools (tool schema, `input_schema`, `tool_choice`, constructed system prompt) — https://docs.claude.com/en/docs/agents-and-tools/tool-use/define-tools
|
||||||
|
- Handle tool calls (`tool_use`/`tool_result`, `is_error`, ordering rules) — https://docs.claude.com/en/docs/agents-and-tools/tool-use/handle-tool-calls
|
||||||
|
- Parallel tool use — https://docs.claude.com/en/docs/agents-and-tools/tool-use/parallel-tool-use
|
||||||
|
- Streaming messages (SSE events, `input_json_delta`, verbatim tool-use stream) — https://docs.claude.com/en/docs/build-with-claude/streaming
|
||||||
|
- Messages API reference (`stop_reason` enum, response shape, `tools`) — https://docs.claude.com/en/api/messages
|
||||||
|
- Legacy tool use (archived; verbatim XML tags and prompt) — https://web.archive.org/web/20240528231249/https://docs.anthropic.com/en/docs/legacy-tool-use ; also live localized copies, e.g. https://docs.anthropic.com/de/docs/legacy-tool-use (English path now redirects to the tool-use overview)
|
||||||
@@ -0,0 +1,356 @@
|
|||||||
|
# DeepSeek tool-calling wire format
|
||||||
|
|
||||||
|
DeepSeek's chat models (DeepSeek-V3, V3-0324, R1, R1-0528, and DeepSeek-V3.1) share a
|
||||||
|
single tokenizer family and a distinctive envelope built from **fullwidth-pipe** special
|
||||||
|
tokens such as `<|begin▁of▁sentence|>` and `<|User|>`. Tool calling is emitted as a run
|
||||||
|
of dedicated special tokens (`<|tool▁calls▁begin|>` … `<|tool▁calls▁end|>`) rather than
|
||||||
|
JSON-in-text or XML. This document centers on **DeepSeek-V3.1** (the current hybrid
|
||||||
|
thinking/non-thinking model) and documents the older **DeepSeek-V3-0324** and
|
||||||
|
**DeepSeek-R1-0528** format as an explicit version difference, because their on-the-wire
|
||||||
|
tool syntax is *not* the same as V3.1's.
|
||||||
|
|
||||||
|
An inference server enables it with a chat template plus a tool-call parser:
|
||||||
|
|
||||||
|
- vLLM V3.1: `--enable-auto-tool-choice --tool-call-parser deepseek_v31 --chat-template examples/tool_chat_template_deepseekv31.jinja` (optionally `--reasoning-parser deepseek_r1`).
|
||||||
|
- vLLM V3-0324 / R1-0528: `--enable-auto-tool-choice --tool-call-parser deepseek_v3 --chat-template examples/tool_chat_template_deepseekv3.jinja` (V3-0324) or `tool_chat_template_deepseekr1.jinja` (R1-0528).
|
||||||
|
- The model's own `tokenizer_config.json` `chat_template` (and the identical `assets/chat_template.jinja`) renders the V3.1 envelope, tool calls, and tool outputs; it does **not** synthesize the `## Tools` advertisement block, so vLLM ships a template that does (see below).
|
||||||
|
|
||||||
|
> Verified against: the DeepSeek-V3.1 model card "Chat Template" / "ToolCall" sections, the
|
||||||
|
> byte-identical `chat_template` in `tokenizer_config.json` and `assets/chat_template.jinja`,
|
||||||
|
> the `added_tokens` in `tokenizer.json` (token IDs), `config.json` (bos/eos IDs), the
|
||||||
|
> DeepSeek-V3-0324 and DeepSeek-R1-0528 `tokenizer_config.json` chat templates, the vLLM
|
||||||
|
> `tool_chat_template_deepseekv31.jinja`, and the vLLM tool-calling / reasoning-outputs docs.
|
||||||
|
|
||||||
|
## A note on the unusual Unicode (do not substitute ASCII)
|
||||||
|
|
||||||
|
DeepSeek's markers do **not** use the ASCII vertical bar `|` (U+007C) or ASCII underscore
|
||||||
|
`_`. They use:
|
||||||
|
|
||||||
|
- `|` — **U+FF5C FULLWIDTH VERTICAL LINE**, as the delimiter just inside the angle brackets.
|
||||||
|
- `▁` — **U+2581 LOWER ONE EIGHTH BLOCK** (the SentencePiece word-boundary glyph), as the
|
||||||
|
separator *between words* inside a token, e.g. `begin▁of▁sentence`, `tool▁calls▁begin`.
|
||||||
|
|
||||||
|
So `<|tool▁calls▁begin|>` is `<` + `|`(FF5C) + `tool` + `▁`(2581) + `calls` + `▁`(2581) +
|
||||||
|
`begin` + `|`(FF5C) + `>`. Copying these tokens as `<|tool_calls_begin|>` (ASCII pipe +
|
||||||
|
underscore) produces tokens the model never trained on and will silently break parsing and
|
||||||
|
generation. The only DeepSeek markers that use ASCII brackets are the thinking tags
|
||||||
|
`<think>` / `</think>` (plain `<`, `/`, `>`) and the rarely used `<|EOT|>` (ASCII pipes).
|
||||||
|
|
||||||
|
## Special tokens
|
||||||
|
|
||||||
|
Token IDs are from DeepSeek-V3.1 `tokenizer.json` (`added_tokens`); `vocab_size` is 129280.
|
||||||
|
The `special` column reflects the tokenizer's `"special"` flag (it governs
|
||||||
|
`skip_special_tokens`); note that the role/think/tool markers are `special: false`.
|
||||||
|
|
||||||
|
| Token (verbatim) | ID | `special` | Purpose |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `<|begin▁of▁sentence|>` | 0 | true | BOS; prepended once at the very start of the prompt. |
|
||||||
|
| `<|end▁of▁sentence|>` | 1 | true | EOS; ends every assistant/tool turn and is the stop token. |
|
||||||
|
| `<|▁pad▁|>` | 2 | true | Padding (`pad_token`; the model card/config also reuse EOS as pad). |
|
||||||
|
| `<|search▁begin|>` | 128796 | false | Search-agent query open (thinking-mode search tool). |
|
||||||
|
| `<|search▁end|>` | 128797 | false | Search-agent query close. |
|
||||||
|
| `<think>` | 128798 | false | Opens the reasoning/thinking span. ASCII brackets. |
|
||||||
|
| `</think>` | 128799 | false | Closes the reasoning span; **also emitted in non-thinking mode** (see below). |
|
||||||
|
| `<|fim▁hole|>` / `<|fim▁begin|>` / `<|fim▁end|>` | 128800–128802 | false | Fill-in-the-middle (not chat). |
|
||||||
|
| `<|User|>` | 128803 | false | User role marker. |
|
||||||
|
| `<|Assistant|>` | 128804 | false | Assistant role marker. |
|
||||||
|
| `<\|EOT\|>` | 128805 | true | End-of-turn (legacy; ASCII pipes, rarely used in chat). |
|
||||||
|
| `<|tool▁calls▁begin|>` | 128806 | false | Opens the assistant's batch of tool calls. |
|
||||||
|
| `<|tool▁calls▁end|>` | 128807 | false | Closes the batch of tool calls. |
|
||||||
|
| `<|tool▁call▁begin|>` | 128808 | false | Opens a single tool call inside the batch. |
|
||||||
|
| `<|tool▁call▁end|>` | 128809 | false | Closes a single tool call. |
|
||||||
|
| `<|tool▁outputs▁begin|>` | 128810 | false | Opens a batch of tool results (**R1-0528 / V3-0324 only**). |
|
||||||
|
| `<|tool▁outputs▁end|>` | 128811 | false | Closes a batch of tool results (**R1-0528 / V3-0324 only**). |
|
||||||
|
| `<|tool▁output▁begin|>` | 128812 | false | Opens a single tool result. |
|
||||||
|
| `<|tool▁output▁end|>` | 128813 | false | Closes a single tool result. |
|
||||||
|
| `<|tool▁sep|>` | 128814 | false | Separator inside a tool call (between name and arguments). |
|
||||||
|
|
||||||
|
`config.json` confirms `bos_token_id: 0`, `eos_token_id: 1`.
|
||||||
|
|
||||||
|
## Roles / channels / turn structure
|
||||||
|
|
||||||
|
There is no OpenAI-style `system`/`developer` channel token. Roles are inline markers and
|
||||||
|
the prompt is one flat string:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|begin▁of▁sentence|>{system_prompt}<|User|>{query}<|Assistant|>{response}<|end▁of▁sentence|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- **System prompt** has no marker. All `system` messages are concatenated (joined with
|
||||||
|
`\n\n` when there are several) and emitted immediately after `<|begin▁of▁sentence|>`,
|
||||||
|
before the first `<|User|>`. When tools are present the `## Tools` block is appended to
|
||||||
|
this system text (separated by `\n\n`).
|
||||||
|
- **User turn**: `<|User|>` + content. (No EOS after the user text in V3.1; the assistant
|
||||||
|
marker follows directly.)
|
||||||
|
- **Assistant turn**: opens with `<|Assistant|>`, then a thinking tag, then content, then
|
||||||
|
`<|end▁of▁sentence|>`.
|
||||||
|
- **Thinking vs non-thinking (V3.1 hybrid)** — selected by the template, not by the model:
|
||||||
|
- Non-thinking generation prefix: `…<|Assistant|></think>` — the model starts *after* a
|
||||||
|
`</think>` it never had to open. Unlike DeepSeek-V3, V3.1 always injects this `</think>`.
|
||||||
|
- Thinking generation prefix: `…<|Assistant|><think>` — the model emits its chain of
|
||||||
|
thought, closes with `</think>`, then the answer.
|
||||||
|
- In multi-turn context, **every** stored assistant turn keeps a `</think>`; only the last
|
||||||
|
turn's leading thinking tag reflects the requested mode. When rendering a stored
|
||||||
|
assistant message, any text up to and including `</think>` is stripped from `content`
|
||||||
|
before re-emitting (the template does `content.split('</think>', 1)[1]`).
|
||||||
|
- **Tool calling runs in non-thinking mode.** The model card states "Toolcall is supported
|
||||||
|
in non-thinking mode," and the V3.1 tool template opens the tool-call turn with
|
||||||
|
`<|Assistant|></think>`. With vLLM, V3.1 reasoning is disabled by default; enable it via
|
||||||
|
`chat_template_kwargs={"thinking": true}`.
|
||||||
|
- **Search-agent channel**: a separate thinking-mode protocol using `<|search▁begin|>` /
|
||||||
|
`<|search▁end|>` (see the model card's `assets/search_tool_trajectory.html`); out of
|
||||||
|
scope for ordinary function calling.
|
||||||
|
|
||||||
|
## Tool definitions
|
||||||
|
|
||||||
|
Tools are advertised as a **Markdown block injected into the system area** (after the system
|
||||||
|
prompt, before the first `<|User|>`). The chat template in `tokenizer_config.json` does not
|
||||||
|
build this block from a `tools=[…]` argument; the caller (or vLLM's
|
||||||
|
`tool_chat_template_deepseekv31.jinja`) constructs it. Reproduced verbatim from the
|
||||||
|
DeepSeek-V3.1 model card, the full layout is
|
||||||
|
`<|begin▁of▁sentence|>{system prompt}\n\n{tool_description}<|User|>{query}<|Assistant|></think>`
|
||||||
|
where `{tool_description}` is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
## Tools
|
||||||
|
You have access to the following tools:
|
||||||
|
|
||||||
|
### {tool_name1}
|
||||||
|
Description: {description}
|
||||||
|
|
||||||
|
Parameters: {json.dumps(parameters)}
|
||||||
|
|
||||||
|
IMPORTANT: ALWAYS adhere to this exact format for tool use:
|
||||||
|
<|tool▁calls▁begin|><|tool▁call▁begin|>tool_call_name<|tool▁sep|>tool_call_arguments<|tool▁call▁end|>{additional_tool_calls}<|tool▁calls▁end|>
|
||||||
|
|
||||||
|
Where:
|
||||||
|
- `tool_call_name` must be an exact match to one of the available tools
|
||||||
|
- `tool_call_arguments` must be valid JSON that strictly follows the tool's Parameters Schema
|
||||||
|
- For multiple tool calls, chain them directly without separators or spaces
|
||||||
|
```
|
||||||
|
|
||||||
|
Each tool contributes one `### {name}` section with a `Description:` line and a
|
||||||
|
`Parameters: {…}` line whose value is the compact JSON of the JSON-Schema parameters object
|
||||||
|
(`json.dumps(parameters)` in the card, `parameters | tojson` in vLLM's template). The
|
||||||
|
`IMPORTANT:` instruction block is appended once, after the last tool.
|
||||||
|
|
||||||
|
## Tool-call format
|
||||||
|
|
||||||
|
The model emits one batch wrapper containing one or more calls. Each call is
|
||||||
|
`name <|tool▁sep|> arguments`, where **arguments is a raw JSON object string** (no code
|
||||||
|
fence). Minimal single call (what the model generates after the `<|Assistant|></think>`
|
||||||
|
prefix):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool▁calls▁begin|><|tool▁call▁begin|>get_weather<|tool▁sep|>{"location": "San Francisco, CA"}<|tool▁call▁end|><|tool▁calls▁end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Grammar (V3.1):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool▁calls▁begin|><|tool▁call▁begin|>{name}<|tool▁sep|>{json_args}<|tool▁call▁end|>{…more calls…}<|tool▁calls▁end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- `{name}` must exactly match an advertised tool name. It comes **first**, immediately after
|
||||||
|
`<|tool▁call▁begin|>`.
|
||||||
|
- `{json_args}` is valid JSON conforming to the tool's parameter schema, inlined directly.
|
||||||
|
- The whole assistant turn is then closed by the template/server with
|
||||||
|
`<|end▁of▁sentence|>`.
|
||||||
|
|
||||||
|
(V3.1 has **no** `type` field and **no** ` ```json ` fence around arguments — that is the
|
||||||
|
older R1/V3-0324 convention; see Version differences.)
|
||||||
|
|
||||||
|
## Multiple / parallel tool calls
|
||||||
|
|
||||||
|
All calls live inside one `<|tool▁calls▁begin|>…<|tool▁calls▁end|>` wrapper. After the
|
||||||
|
first `<|tool▁call▁begin|>…<|tool▁call▁end|>`, each additional call is **another
|
||||||
|
`<|tool▁call▁begin|>…<|tool▁call▁end|>` chained directly, with no separator, newline, or
|
||||||
|
space between calls** (the card: "chain them directly without separators or spaces"):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool▁calls▁begin|><|tool▁call▁begin|>get_weather<|tool▁sep|>{"location": "San Francisco, CA"}<|tool▁call▁end|><|tool▁call▁begin|>get_weather<|tool▁sep|>{"location": "Seattle, WA"}<|tool▁call▁end|><|tool▁calls▁end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Note that `<|tool▁calls▁begin|>` (plural, id 128806) appears exactly once; each call uses
|
||||||
|
the singular `<|tool▁call▁begin|>` (id 128808) / `<|tool▁call▁end|>` (id 128809).
|
||||||
|
|
||||||
|
## Tool-result format
|
||||||
|
|
||||||
|
Executed results are fed back as `tool`-role messages. In **V3.1** each result is wrapped in
|
||||||
|
the singular output tokens, with **no** plural `<|tool▁outputs▁…|>` wrapper, emitted right
|
||||||
|
after the assistant tool-call turn's `<|end▁of▁sentence|>`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool▁output▁begin|>{result_text}<|tool▁output▁end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
`{result_text}` is the raw tool output (typically a JSON string, but any text). For multiple
|
||||||
|
results, the V3.1 template emits one `<|tool▁output▁begin|>…<|tool▁output▁end|>` per `tool`
|
||||||
|
message, concatenated directly. There is **no tool-call ID in the wire format** — results are
|
||||||
|
matched to calls **positionally** (order of outputs ↔ order of calls).
|
||||||
|
|
||||||
|
The model then produces its final answer **directly after `<|tool▁output▁end|>`** with no
|
||||||
|
`<|Assistant|>` marker and no `</think>` (see Parsing notes — the V3.1 reference template
|
||||||
|
deliberately renders post-tool assistant content as just `content<|end▁of▁sentence|>`).
|
||||||
|
|
||||||
|
> R1-0528 / V3-0324 differ: results are enclosed in a `<|tool▁outputs▁begin|>` …
|
||||||
|
> `<|tool▁outputs▁end|>` batch wrapper, with each result as
|
||||||
|
> `<|tool▁output▁begin|>…<|tool▁output▁end|>` and multiple results newline-separated.
|
||||||
|
|
||||||
|
## End-to-end example
|
||||||
|
|
||||||
|
A complete DeepSeek-V3.1 **non-thinking** multi-turn exchange. Everything is one flat string;
|
||||||
|
inline `←` comments mark where the model's generation begins (they are not part of the
|
||||||
|
stream). Whitespace inside the `## Tools` block is literal newlines.
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|begin▁of▁sentence|>You are a helpful assistant.
|
||||||
|
|
||||||
|
## Tools
|
||||||
|
You have access to the following tools:
|
||||||
|
|
||||||
|
### get_weather
|
||||||
|
Description: Get the current weather for a location
|
||||||
|
|
||||||
|
Parameters: {"type": "object", "properties": {"location": {"type": "string", "description": "City and state, e.g. San Francisco, CA"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}}, "required": ["location"]}
|
||||||
|
|
||||||
|
IMPORTANT: ALWAYS adhere to this exact format for tool use:
|
||||||
|
<|tool▁calls▁begin|><|tool▁call▁begin|>tool_call_name<|tool▁sep|>tool_call_arguments<|tool▁call▁end|>{additional_tool_calls}<|tool▁calls▁end|>
|
||||||
|
|
||||||
|
Where:
|
||||||
|
- `tool_call_name` must be an exact match to one of the available tools
|
||||||
|
- `tool_call_arguments` must be valid JSON that strictly follows the tool's Parameters Schema
|
||||||
|
- For multiple tool calls, chain them directly without separators or spaces
|
||||||
|
<|User|>What's the weather in San Francisco?<|Assistant|></think><|tool▁calls▁begin|><|tool▁call▁begin|>get_weather<|tool▁sep|>{"location": "San Francisco, CA", "unit": "celsius"}<|tool▁call▁end|><|tool▁calls▁end|><|end▁of▁sentence|><|tool▁output▁begin|>{"temperature": 18, "unit": "celsius", "condition": "Foggy"}<|tool▁output▁end|>It's currently 18°C and foggy in San Francisco.<|end▁of▁sentence|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Reading the spans:
|
||||||
|
|
||||||
|
1. `<|begin▁of▁sentence|>` + system text + `\n\n` + `## Tools…` block — prompt prefix.
|
||||||
|
2. `<|User|>What's the weather in San Francisco?` — user turn.
|
||||||
|
3. `<|Assistant|></think>` — non-thinking generation prefix (prompt). **Model generates from here.**
|
||||||
|
4. `<|tool▁calls▁begin|>…<|tool▁calls▁end|>` — the model's tool call; server appends `<|end▁of▁sentence|>` and stops with `finish_reason: "tool_calls"`.
|
||||||
|
5. `<|tool▁output▁begin|>…<|tool▁output▁end|>` — your executed result, appended to the prompt.
|
||||||
|
6. `It's currently 18°C and foggy in San Francisco.<|end▁of▁sentence|>` — **the model generates the final answer directly after the tool output** (no new `<|Assistant|>` marker), ending with EOS.
|
||||||
|
|
||||||
|
## OpenAI-compatible API mapping
|
||||||
|
|
||||||
|
When fronted by an OpenAI-compatible server (e.g. vLLM with `--tool-call-parser
|
||||||
|
deepseek_v31`):
|
||||||
|
|
||||||
|
- **`finish_reason`**: `"tool_calls"` when the model emitted a `<|tool▁calls▁begin|>…`
|
||||||
|
batch; otherwise `"stop"`.
|
||||||
|
- **`message.tool_calls[]`**: one element per `<|tool▁call▁begin|>…<|tool▁call▁end|>`.
|
||||||
|
- `.type` = `"function"`.
|
||||||
|
- `.function.name` = the text between `<|tool▁call▁begin|>` and `<|tool▁sep|>`.
|
||||||
|
- `.function.arguments` = the text between `<|tool▁sep|>` and `<|tool▁call▁end|>`, returned
|
||||||
|
as a **JSON string** (per the OpenAI spec), not a nested object. The model already emits
|
||||||
|
raw JSON there, so it is passed through.
|
||||||
|
- `.id` = **synthesized by the server** (e.g. `chatcmpl-tool-…`). DeepSeek's wire format
|
||||||
|
carries no call ID.
|
||||||
|
- **Tool result messages**: `{"role": "tool", "tool_call_id": "<id>", "content": "<result>"}`.
|
||||||
|
The server renders `content` into `<|tool▁output▁begin|>…<|tool▁output▁end|>`. Because the
|
||||||
|
prompt has no IDs, `tool_call_id` is used only for client-side bookkeeping; **the model
|
||||||
|
relies on ordering**, so preserve the order of results relative to the calls.
|
||||||
|
- **Assistant replay**: when you send a prior assistant turn back with `tool_calls`, the
|
||||||
|
template inlines `function.arguments`. The HF reference template inlines it **verbatim**
|
||||||
|
(assumes it is already a JSON string); vLLM's `tool_chat_template_deepseekv31.jinja` pipes
|
||||||
|
it through `| tojson`. Send `arguments` as a JSON **string** per the OpenAI spec (see the
|
||||||
|
gotcha below about double-encoding).
|
||||||
|
|
||||||
|
## Parsing notes & gotchas
|
||||||
|
|
||||||
|
- **Unicode is load-bearing.** Match `|` = U+FF5C and `▁` = U+2581 exactly. ASCII
|
||||||
|
`<|tool_calls_begin|>` will not tokenize to the special tokens. `<think>`/`</think>` use
|
||||||
|
ASCII brackets; the rare `<|EOT|>` uses ASCII pipes.
|
||||||
|
- **Tool/role markers are `special: false`.** Only `<|begin▁of▁sentence|>`,
|
||||||
|
`<|end▁of▁sentence|>`, `<|▁pad▁|>`, and `<|EOT|>` are flagged `special: true`. So
|
||||||
|
decoding with `skip_special_tokens=True` will **not** strip `<|tool▁calls▁begin|>`,
|
||||||
|
`<|tool▁sep|>`, `<|Assistant|>`, `</think>`, etc. — they remain in the decoded string for
|
||||||
|
the parser to find. (Conversely, do not assume special-token filtering removes them.)
|
||||||
|
- **No code fence / no `type` field in V3.1.** A parser written for R1/V3-0324
|
||||||
|
(`function<|tool▁sep|>name` + ` ```json ` block) will not parse V3.1, and vice-versa.
|
||||||
|
V3.1 is `name<|tool▁sep|>raw_json`.
|
||||||
|
- **Chaining has no delimiter in V3.1.** Calls abut directly:
|
||||||
|
`…<|tool▁call▁end|><|tool▁call▁begin|>…`. Do not split on newlines/whitespace; split on
|
||||||
|
the `<|tool▁call▁begin|>` / `<|tool▁call▁end|>` boundaries. (R1/V3-0324 put a `\n` before
|
||||||
|
each subsequent call.)
|
||||||
|
- **No tool-call IDs on the wire.** Match results to calls by position. A server must
|
||||||
|
generate synthetic `tool_call_id`s for the OpenAI shape.
|
||||||
|
- **`</think>` appears even in non-thinking mode.** Strip the leading `</think>` (and any
|
||||||
|
preceding reasoning) before treating the remainder as the visible answer; the template does
|
||||||
|
`content.split('</think>', 1)[1]` when replaying stored turns.
|
||||||
|
- **Post-tool generation prompt quirk.** The reference V3.1 chat template only appends the
|
||||||
|
`<|Assistant|></think>` generation prefix when the **last message is `user`**. After a
|
||||||
|
`tool` message it appends nothing and the model continues straight after
|
||||||
|
`<|tool▁output▁end|>`. Agent loops that re-template a conversation ending in a tool result
|
||||||
|
must not expect (or double-insert) an assistant marker there.
|
||||||
|
- **`arguments` double-encoding risk.** On replay, vLLM's example template applies
|
||||||
|
`arguments | tojson`. If `arguments` is already a JSON string (the OpenAI convention), that
|
||||||
|
pipe will JSON-encode the string again (wrapping it in quotes and escaping it). Pass an
|
||||||
|
object where the template expects `| tojson`, or a string where the template inlines
|
||||||
|
verbatim — match the template you actually run.
|
||||||
|
- **Streaming.** Tool calls arrive token-by-token; the name is complete only at
|
||||||
|
`<|tool▁sep|>`, and arguments are partial JSON until `<|tool▁call▁end|>`. Buffer per call
|
||||||
|
boundary; do not attempt to `json.loads` arguments before the closing tool-call token.
|
||||||
|
- **Malformed output.** With `tool_choice="auto"` and no structural-tag constraint
|
||||||
|
(`VLLM_ENFORCE_STRICT_TOOL_CALLING=false`), the model can emit invalid JSON in
|
||||||
|
`tool_call_arguments` or a `tool_call_name` that does not match any tool; the parser
|
||||||
|
extracts best-effort. Named/`required` tool choice uses the structured-outputs backend and
|
||||||
|
guarantees schema-valid arguments.
|
||||||
|
|
||||||
|
## Version differences: V3.1 vs V3-0324 / R1-0528
|
||||||
|
|
||||||
|
The pre-V3.1 models (DeepSeek-V3-0324 and DeepSeek-R1-0528) share an older tool-call
|
||||||
|
encoding, served in vLLM with `--tool-call-parser deepseek_v3`. The per-call body is:
|
||||||
|
|
||||||
|
````text
|
||||||
|
<|tool▁call▁begin|>function<|tool▁sep|>{name}
|
||||||
|
```json
|
||||||
|
{json_args}
|
||||||
|
```<|tool▁call▁end|>
|
||||||
|
````
|
||||||
|
|
||||||
|
Differences from V3.1:
|
||||||
|
|
||||||
|
| Aspect | V3.1 (`deepseek_v31`) | V3-0324 / R1-0528 (`deepseek_v3`) |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| Field order in a call | `{name}<|tool▁sep|>{args}` | `function<|tool▁sep|>{name}` (the literal `type`, then name) |
|
||||||
|
| Arguments wrapping | raw JSON, inline | fenced ` ```json … ``` ` block (name and args separated by `\n`) |
|
||||||
|
| Chaining of calls | abut directly, **no separator** | each subsequent call prefixed with `\n` |
|
||||||
|
| Tool results | `<|tool▁output▁begin|>…<|tool▁output▁end|>` per message, no batch wrapper | wrapped in `<|tool▁outputs▁begin|>…<|tool▁outputs▁end|>`, results newline-separated |
|
||||||
|
| User→assistant boundary | user turn = `<|User|>{q}`; `<|Assistant|></think>` added at generation | user turn = `<|User|>{q}<|Assistant|>` (assistant marker appended in the user branch) |
|
||||||
|
| Thinking | hybrid; `thinking` kwarg toggles `<think>` vs `</think>` prefix | R1-0528 always reasoning (bare `<|Assistant|>` generation prefix, model opens `<think>` itself); V3-0324 non-reasoning |
|
||||||
|
| vLLM parser | `--tool-call-parser deepseek_v31` | `--tool-call-parser deepseek_v3` |
|
||||||
|
|
||||||
|
Example R1-0528 / V3-0324 parallel call with its result batch:
|
||||||
|
|
||||||
|
````text
|
||||||
|
<|tool▁calls▁begin|><|tool▁call▁begin|>function<|tool▁sep|>get_weather
|
||||||
|
```json
|
||||||
|
{"location": "San Francisco, CA"}
|
||||||
|
```<|tool▁call▁end|>
|
||||||
|
<|tool▁call▁begin|>function<|tool▁sep|>get_weather
|
||||||
|
```json
|
||||||
|
{"location": "Seattle, WA"}
|
||||||
|
```<|tool▁call▁end|><|tool▁calls▁end|><|end▁of▁sentence|><|tool▁outputs▁begin|><|tool▁output▁begin|>{"temperature": 18}<|tool▁output▁end|>
|
||||||
|
<|tool▁output▁begin|>{"temperature": 14}<|tool▁output▁end|><|tool▁outputs▁end|>
|
||||||
|
````
|
||||||
|
|
||||||
|
The `deepseek_r1` **reasoning** parser (`--reasoning-parser deepseek_r1`) applies to the R1
|
||||||
|
series **and** to DeepSeek-V3.1; it extracts the `<think>…</think>` span into the response's
|
||||||
|
`reasoning` field. It is independent of the tool-call parser.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- DeepSeek-V3.1 model card (Chat Template / ToolCall sections): <https://huggingface.co/deepseek-ai/DeepSeek-V3.1>
|
||||||
|
- DeepSeek-V3.1 `assets/chat_template.jinja`: <https://huggingface.co/deepseek-ai/DeepSeek-V3.1/resolve/main/assets/chat_template.jinja>
|
||||||
|
- DeepSeek-V3.1 `tokenizer_config.json` (`chat_template`, byte-identical to the jinja): <https://huggingface.co/deepseek-ai/DeepSeek-V3.1/resolve/main/tokenizer_config.json>
|
||||||
|
- DeepSeek-V3.1 `tokenizer.json` (`added_tokens` → token IDs and `special` flags): <https://huggingface.co/deepseek-ai/DeepSeek-V3.1/resolve/main/tokenizer.json>
|
||||||
|
- DeepSeek-V3.1 `config.json` (`bos_token_id`, `eos_token_id`, `vocab_size`): <https://huggingface.co/deepseek-ai/DeepSeek-V3.1/resolve/main/config.json>
|
||||||
|
- DeepSeek-R1-0528 model card and `tokenizer_config.json` (older tool format): <https://huggingface.co/deepseek-ai/DeepSeek-R1-0528> · <https://huggingface.co/deepseek-ai/DeepSeek-R1-0528/resolve/main/tokenizer_config.json>
|
||||||
|
- DeepSeek-R1 model card: <https://huggingface.co/deepseek-ai/DeepSeek-R1>
|
||||||
|
- DeepSeek-V3-0324 `tokenizer_config.json` (older tool format): <https://huggingface.co/deepseek-ai/DeepSeek-V3-0324/resolve/main/tokenizer_config.json>
|
||||||
|
- vLLM tool-call template for V3.1 (`## Tools` injection + `| tojson`): <https://github.com/vllm-project/vllm/blob/main/examples/tool_chat_template_deepseekv31.jinja>
|
||||||
|
- vLLM Tool Calling docs (`deepseek_v3`, `deepseek_v31` parser flags): <https://docs.vllm.ai/en/latest/features/tool_calling/>
|
||||||
|
- vLLM Reasoning Outputs docs (`deepseek_r1` reasoning parser; V3.1 thinking default): <https://docs.vllm.ai/en/latest/features/reasoning_outputs/>
|
||||||
@@ -0,0 +1,296 @@
|
|||||||
|
# GLM-4.5 / GLM-4.6 tool-calling format
|
||||||
|
|
||||||
|
Native tool-calling convention of Zhipu AI / Z.ai's **GLM-4.5** family (`zai-org/GLM-4.5` 355B-A32B and `zai-org/GLM-4.5-Air` 106B-A12B, `model_type: "glm4_moe"`), shared byte-for-byte by **GLM-4.6**. Unlike the JSON-in-a-tag conventions used by most families, GLM emits each tool call as an **XML-like** block: `<tool_call>{name}` followed by alternating `<arg_key>`/`<arg_value>` element pairs, closed by `</tool_call>`. The prompt is a GLM-style sequence opened by `[gMASK]<sop>` with turn markers `<|system|>`, `<|user|>`, `<|assistant|>`, `<|observation|>`. An inference server turns the raw stream into OpenAI-style `tool_calls` with a parser plus a reasoning parser: both vLLM and SGLang expose `--tool-call-parser glm45 --reasoning-parser glm45` (vLLM additionally needs `--enable-auto-tool-choice`). Tool calling and reasoning are driven entirely by the bundled `chat_template.jinja`; thinking mode is on by default and is disabled per-request with `chat_template_kwargs={"enable_thinking": false}`.
|
||||||
|
|
||||||
|
This document was verified against the authoritative `chat_template.jinja` from the HF repo (fetched raw and **rendered locally with Jinja2** — `trim_blocks=True, lstrip_blocks=True`, transformers' `tojson` filter — to produce the byte-exact streams below), `tokenizer_config.json` and `generation_config.json` for the exact token IDs and stop tokens, the model card, and the vLLM (`Glm4MoeModelToolParser`) and SGLang (`Glm4MoeDetector`) parser sources. The HF `resolve`/`blob` web paths redirect to the model-card API; the byte-exact source was obtained via the `resolve/main/...:raw` cache (template commit `cbb2c7cfb52fa128a9660cb1a7a78e017899e115`). The GLM-4.5 and GLM-4.6 `chat_template.jinja` files are identical (same content hash `41478957…`).
|
||||||
|
|
||||||
|
## Special tokens
|
||||||
|
|
||||||
|
Token IDs are from `tokenizer_config.json` (`added_tokens_decoder`). Note the split: the turn/role markers are registered as **special** tokens, whereas the structural tool-call and thinking tags are each a single dedicated vocabulary token but flagged **`special: false`** (they are emitted/printed as ordinary text, not stripped as control tokens).
|
||||||
|
|
||||||
|
| Token (verbatim) | ID | `special` | Purpose |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `[gMASK]` | 151331 | true | GLM prefix / blank-infilling sentinel; first token of every prompt |
|
||||||
|
| `<sop>` | 151333 | true | "Start of piece" — immediately follows `[gMASK]` to open the sequence |
|
||||||
|
| `<eop>` | 151334 | true | "End of piece" (not emitted by the chat template) |
|
||||||
|
| `<\|system\|>` | 151335 | true | Opens a system turn (and the injected tools turn) |
|
||||||
|
| `<\|user\|>` | 151336 | true | Opens a user turn (also an EOS id — see below) |
|
||||||
|
| `<\|assistant\|>` | 151337 | true | Opens an assistant turn / generation prompt |
|
||||||
|
| `<\|observation\|>` | 151338 | true | Opens a tool-result (observation) turn (also an EOS id) |
|
||||||
|
| `<\|endoftext\|>` | 151329 | true | End-of-text; `eos_token` and `pad_token` |
|
||||||
|
| `<think>` | 151350 | false | Opens the reasoning span inside an assistant turn |
|
||||||
|
| `</think>` | 151351 | false | Closes the reasoning span |
|
||||||
|
| `<tool_call>` | 151352 | false | Opens one tool call; function name follows on the same line |
|
||||||
|
| `</tool_call>` | 151353 | false | Closes one tool call |
|
||||||
|
| `<arg_key>` | 151356 | false | Opens an argument-name element |
|
||||||
|
| `</arg_key>` | 151357 | false | Closes an argument-name element |
|
||||||
|
| `<arg_value>` | 151358 | false | Opens an argument-value element |
|
||||||
|
| `</arg_value>` | 151359 | false | Closes an argument-value element |
|
||||||
|
| `<tool_response>` | 151354 | false | Wraps one tool result inside an observation turn |
|
||||||
|
| `</tool_response>` | 151355 | false | Closes a tool result |
|
||||||
|
| `/nothink` | 151360 | true | Soft switch appended to user text to suppress thinking |
|
||||||
|
|
||||||
|
Notes on exactness:
|
||||||
|
- All pipes are ASCII `|` (U+007C); GLM uses no fullwidth `|` (U+FF5C) or `▁` (U+2581) variants (unlike DeepSeek). Reproduce `<|system|>`, `<|user|>`, `<|assistant|>`, `<|observation|>` exactly, and `[gMASK]` with literal square brackets.
|
||||||
|
- Because `<tool_call>`, `<arg_key>`, `<arg_value>`, `<tool_response>`, `<think>` (and their closers) each map to exactly **one** token ID, they cost one token apiece in the stream — but being `special: false` they round-trip through detokenization as plain text. Parsers therefore match them as literal substrings in the decoded text, not as control-token ids.
|
||||||
|
- `eos_token_id` is a **list**: `[151329, 151336, 151338]` = `<|endoftext|>`, `<|user|>`, `<|observation|>` (from `generation_config.json`). This is how a tool-call turn ends: after `</tool_call>` the model emits `<|observation|>`, which is an EOS id, so generation halts and the server reports a tool call (see Turn structure).
|
||||||
|
|
||||||
|
## Roles / channels / turn structure
|
||||||
|
|
||||||
|
Every prompt begins with the literal two-token prefix `[gMASK]<sop>` (no following newline). Turns are then concatenated, each introduced by its role marker; there is no per-turn terminator token in rendered history (the next marker, or an EOS id during generation, ends a turn).
|
||||||
|
|
||||||
|
- **System** (`<|system|>`): role marker, newline, then the message text. When `tools` are supplied, a synthetic tools system turn is rendered **first**, before any user-supplied system turn (the two are separate `<|system|>` blocks — see Tool definitions).
|
||||||
|
- **User** (`<|user|>`): role marker, newline, then text. If `enable_thinking` is false, the literal `/nothink` is appended to the user text (unless it already ends with `/nothink`).
|
||||||
|
- **Assistant** (`<|assistant|>`): role marker, then a reasoning span and/or visible content and/or tool calls. The reasoning span is `\n<think>{reasoning}</think>`; visible content follows on its own line; tool calls follow as `<tool_call>…</tool_call>` blocks.
|
||||||
|
- **Tool result** (`<|observation|>`): role marker introducing one or more `<tool_response>…</tool_response>` blocks (see Tool-result format).
|
||||||
|
|
||||||
|
Thinking / reasoning channel:
|
||||||
|
- Reasoning lives in `<think>…</think>` inside the assistant turn. The `--reasoning-parser glm45` extracts it into a separate `reasoning_content` field; the visible answer is whatever follows `</think>`.
|
||||||
|
- **Only the reasoning of assistant turns after the last user message is kept.** The template renders every earlier assistant turn with an empty `<think></think>` and drops its `reasoning_content` (or any inline `<think>…</think>` embedded in `content`). This keeps stale chains of thought out of the context on later turns.
|
||||||
|
- An assistant turn with neither preserved reasoning nor an explicit chain renders `\n<think></think>` (empty), then content/tool calls.
|
||||||
|
|
||||||
|
Generation prompt (`add_generation_prompt=True`):
|
||||||
|
- **Thinking mode (default):** the prompt ends with a bare `<|assistant|>`; the model continues with `\n<think>…</think>` then its answer or tool calls.
|
||||||
|
- **Non-thinking mode** (`enable_thinking=false`): the prompt ends with `<|assistant|>\n<think></think>`, pre-filling an empty reasoning span so the model goes straight to the answer.
|
||||||
|
|
||||||
|
How a tool-call turn terminates: there is no dedicated "stop after tool call" token. The model emits `</tool_call>` and then `<|observation|>` (token 151338), which is one of the three EOS ids, so decoding stops. The server inspects the text, finds `<tool_call>`, and returns `finish_reason: "tool_calls"`.
|
||||||
|
|
||||||
|
## Tool definitions
|
||||||
|
|
||||||
|
When the request carries `tools`, the template prepends one `<|system|>` turn containing a fixed preamble, the tool list wrapped in `<tools>…</tools>`, and a literal description of the output format. Each tool is serialized with `tool | tojson(ensure_ascii=False)` — i.e. the **entire OpenAI tool object verbatim**, including the `{"type": "function", "function": {…}}` wrapper, with default JSON spacing (`", "` / `": "`). One tool per line.
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|system|>
|
||||||
|
# Tools
|
||||||
|
|
||||||
|
You may call one or more functions to assist with the user query.
|
||||||
|
|
||||||
|
You are provided with function signatures within <tools></tools> XML tags:
|
||||||
|
<tools>
|
||||||
|
{"type": "function", "function": {"name": "get_weather", "description": "Get current weather for a city", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "City name"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}}, "required": ["location"]}}}
|
||||||
|
</tools>
|
||||||
|
|
||||||
|
For each function call, output the function name and arguments within the following XML format:
|
||||||
|
<tool_call>{function-name}
|
||||||
|
<arg_key>{arg-key-1}</arg_key>
|
||||||
|
<arg_value>{arg-value-1}</arg_value>
|
||||||
|
<arg_key>{arg-key-2}</arg_key>
|
||||||
|
<arg_value>{arg-value-2}</arg_value>
|
||||||
|
...
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
The `<tool_call>{function-name}` / `<arg_key>` / `<arg_value>` lines above are part of the **prompt text** (the format spec the model is told to follow), not an example call. This tools turn is emitted only when `tools` is non-empty, and it is closed implicitly by the next role marker (e.g. a user-supplied `<|system|>` or the first `<|user|>`), with no blank line between them.
|
||||||
|
|
||||||
|
## Tool-call format
|
||||||
|
|
||||||
|
The model emits a call as an `<tool_call>` block: the function **name on the same line** as the opening tag, a newline, then one `<arg_key>…</arg_key>` + `<arg_value>…</arg_value>` pair per argument, closed by `</tool_call>`. Minimal single call (assistant generation in thinking mode; reasoning shown for realism):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<think>The user wants the weather in Beijing. I'll call get_weather.</think>
|
||||||
|
<tool_call>get_weather
|
||||||
|
<arg_key>location</arg_key>
|
||||||
|
<arg_value>Beijing</arg_value>
|
||||||
|
<arg_key>unit</arg_key>
|
||||||
|
<arg_value>celsius</arg_value>
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
Anatomy and value encoding (this is the single most error-prone part):
|
||||||
|
|
||||||
|
- The function name is the text between `<tool_call>` and the first newline — there is **no** wrapping tag around it and **no** space after `<tool_call>`.
|
||||||
|
- Each argument is two adjacent elements: `<arg_key>name</arg_key>` then `<arg_value>value</arg_value>`, conventionally one pair per line.
|
||||||
|
- **Argument values are NOT uniformly JSON.** The template renders each value as `value | tojson(ensure_ascii=False) if value is not string else value`:
|
||||||
|
- **string** values are emitted **raw, without surrounding quotes** → `<arg_value>Beijing</arg_value>` (not `"Beijing"`).
|
||||||
|
- **non-string** values (number, boolean, null, object, array) are JSON-encoded → `<arg_value>3</arg_value>`, `<arg_value>true</arg_value>`, `<arg_value>{"k": 1}</arg_value>`.
|
||||||
|
- A **zero-argument** call has no pairs: the name is followed by a newline and the closer — `<tool_call>get_time\n</tool_call>`.
|
||||||
|
|
||||||
|
Because string values lose their quotes, a parser must decide per argument whether to JSON-decode or treat the value as a literal string. Both reference parsers do this by consulting the tool's JSON Schema: if the parameter's type is `string`, the raw text is taken as-is; otherwise the value is JSON-decoded (with `ast.literal_eval` and raw-string fallbacks). The model is trained to follow the schema, so it emits a bare string exactly when the parameter is string-typed.
|
||||||
|
|
||||||
|
## Multiple / parallel tool calls
|
||||||
|
|
||||||
|
Two or more calls in one turn are emitted as consecutive `<tool_call>…</tool_call>` blocks separated by a single newline (no wrapper element around the set). Raw assistant emission for two parallel calls with mixed argument types:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<think>Two cities. Call get_weather twice in parallel.</think>
|
||||||
|
<tool_call>get_weather
|
||||||
|
<arg_key>location</arg_key>
|
||||||
|
<arg_value>Beijing</arg_value>
|
||||||
|
<arg_key>unit</arg_key>
|
||||||
|
<arg_value>celsius</arg_value>
|
||||||
|
</tool_call>
|
||||||
|
<tool_call>get_weather
|
||||||
|
<arg_key>location</arg_key>
|
||||||
|
<arg_value>Shanghai</arg_value>
|
||||||
|
<arg_key>days</arg_key>
|
||||||
|
<arg_value>3</arg_value>
|
||||||
|
<arg_key>verbose</arg_key>
|
||||||
|
<arg_value>true</arg_value>
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
Note `Beijing`/`Shanghai`/`celsius` (string) are bare, while `3` (number) and `true` (boolean) are JSON literals. Parsers split on the non-greedy `<tool_call>.*?</tool_call>` regex, so any number of calls is supported; each becomes a separate entry in `tool_calls[]`.
|
||||||
|
|
||||||
|
## Tool-result format
|
||||||
|
|
||||||
|
Results are returned in an **observation** turn. For a single result: the `<|observation|>` marker, a newline, then the result wrapped in `<tool_response>` / `</tool_response>`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|observation|>
|
||||||
|
<tool_response>
|
||||||
|
{"temperature": 26, "unit": "celsius", "condition": "Sunny"}
|
||||||
|
</tool_response>
|
||||||
|
```
|
||||||
|
|
||||||
|
The content between the tags is inserted **verbatim** (callers typically pass a JSON string, but any text is allowed). For **multiple** results from a set of parallel calls, the `<|observation|>` marker appears **once** and each result gets its own `<tool_response>` block (consecutive `tool`-role messages are merged under a single observation turn):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|observation|>
|
||||||
|
<tool_response>
|
||||||
|
{"temperature": 26, "condition": "Sunny"}
|
||||||
|
</tool_response>
|
||||||
|
<tool_response>
|
||||||
|
{"temperature": 30, "condition": "Cloudy"}
|
||||||
|
</tool_response>
|
||||||
|
```
|
||||||
|
|
||||||
|
The chat template reads **only** the tool message's `content` — it does not consult any `tool_call_id`. Results are therefore correlated to calls **positionally / by order**, not by an embedded id (GLM's wire format carries no per-call id; see API mapping).
|
||||||
|
|
||||||
|
## End-to-end example
|
||||||
|
|
||||||
|
A complete multi-turn weather exchange. These are the exact locally rendered streams; newlines inside a turn are literal and turns are otherwise contiguous (no separators between markers).
|
||||||
|
|
||||||
|
**Stage 1 — prompt fed to the model** (`tools` set, one prior system message, `add_generation_prompt=True`, thinking mode):
|
||||||
|
|
||||||
|
```text
|
||||||
|
[gMASK]<sop><|system|>
|
||||||
|
# Tools
|
||||||
|
|
||||||
|
You may call one or more functions to assist with the user query.
|
||||||
|
|
||||||
|
You are provided with function signatures within <tools></tools> XML tags:
|
||||||
|
<tools>
|
||||||
|
{"type": "function", "function": {"name": "get_weather", "description": "Get current weather for a city", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "City name"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}}, "required": ["location"]}}}
|
||||||
|
</tools>
|
||||||
|
|
||||||
|
For each function call, output the function name and arguments within the following XML format:
|
||||||
|
<tool_call>{function-name}
|
||||||
|
<arg_key>{arg-key-1}</arg_key>
|
||||||
|
<arg_value>{arg-value-1}</arg_value>
|
||||||
|
<arg_key>{arg-key-2}</arg_key>
|
||||||
|
<arg_value>{arg-value-2}</arg_value>
|
||||||
|
...
|
||||||
|
</tool_call><|system|>
|
||||||
|
You are a helpful assistant.<|user|>
|
||||||
|
What's the weather in Beijing?<|assistant|>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Assistant generation** (model output; it ends by emitting `<|observation|>`, an EOS id, so decoding stops there; server returns `finish_reason: "tool_calls"`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<think>The user wants the weather in Beijing. I'll call get_weather.</think>
|
||||||
|
<tool_call>get_weather
|
||||||
|
<arg_key>location</arg_key>
|
||||||
|
<arg_value>Beijing</arg_value>
|
||||||
|
<arg_key>unit</arg_key>
|
||||||
|
<arg_value>celsius</arg_value>
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Stage 2 — prompt for the next turn**, after appending the assistant tool-call turn and the tool result, then `add_generation_prompt=True`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
[gMASK]<sop><|system|>
|
||||||
|
# Tools
|
||||||
|
|
||||||
|
You may call one or more functions to assist with the user query.
|
||||||
|
|
||||||
|
You are provided with function signatures within <tools></tools> XML tags:
|
||||||
|
<tools>
|
||||||
|
{"type": "function", "function": {"name": "get_weather", "description": "Get current weather for a city", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "City name"}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}}, "required": ["location"]}}}
|
||||||
|
</tools>
|
||||||
|
|
||||||
|
For each function call, output the function name and arguments within the following XML format:
|
||||||
|
<tool_call>{function-name}
|
||||||
|
<arg_key>{arg-key-1}</arg_key>
|
||||||
|
<arg_value>{arg-value-1}</arg_value>
|
||||||
|
<arg_key>{arg-key-2}</arg_key>
|
||||||
|
<arg_value>{arg-value-2}</arg_value>
|
||||||
|
...
|
||||||
|
</tool_call><|system|>
|
||||||
|
You are a helpful assistant.<|user|>
|
||||||
|
What's the weather in Beijing?<|assistant|>
|
||||||
|
<think>The user wants the weather in Beijing. I'll call get_weather.</think>
|
||||||
|
<tool_call>get_weather
|
||||||
|
<arg_key>location</arg_key>
|
||||||
|
<arg_value>Beijing</arg_value>
|
||||||
|
<arg_key>unit</arg_key>
|
||||||
|
<arg_value>celsius</arg_value>
|
||||||
|
</tool_call><|observation|>
|
||||||
|
<tool_response>
|
||||||
|
{"temperature": 26, "unit": "celsius", "condition": "Sunny"}
|
||||||
|
</tool_response><|assistant|>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Final assistant generation** (natural-language answer, terminated by `<|endoftext|>`; `finish_reason: "stop"`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<think>Got it, 26C and sunny.</think>
|
||||||
|
It's 26°C and sunny in Beijing right now.
|
||||||
|
```
|
||||||
|
|
||||||
|
Two subtleties visible above: (1) the reasoning of the assistant tool-call turn is **preserved** in Stage 2 only because it is the segment after the last user message; with another user turn after it, that `<think>…</think>` would be re-rendered empty. (2) The tool-call turn and the observation turn abut directly (`</tool_call><|observation|>`), and the observation abuts the next assistant marker (`</tool_response><|assistant|>`).
|
||||||
|
|
||||||
|
For **non-thinking** mode the user text carries the soft switch and the generation prompt pre-fills an empty think span:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|user|>
|
||||||
|
Hi there/nothink<|assistant|>
|
||||||
|
<think></think>
|
||||||
|
```
|
||||||
|
|
||||||
|
## OpenAI-compatible API mapping
|
||||||
|
|
||||||
|
With a server parser active (`--tool-call-parser glm45 --reasoning-parser glm45`), the raw stream maps onto Chat Completions as follows:
|
||||||
|
|
||||||
|
- `choices[].finish_reason` = `"tool_calls"` when the output contained at least one `<tool_call>` (otherwise `"stop"`).
|
||||||
|
- `choices[].message.content` = the text **before** the first `<tool_call>` (normalized to `null` if empty/whitespace). The `<think>…</think>` reasoning is removed by the reasoning parser and surfaced separately as `message.reasoning_content`.
|
||||||
|
- `choices[].message.tool_calls[]` — one entry per `<tool_call>…</tool_call>` block:
|
||||||
|
- `.id` = a **server-generated** id (e.g. vLLM's `make_tool_call_id()`), **not** present in the model output. GLM emits no per-call id in the stream.
|
||||||
|
- `.type` = `"function"`.
|
||||||
|
- `.function.name` = the text after `<tool_call>` up to the first newline.
|
||||||
|
- `.function.arguments` = a **JSON string** (an object), reconstructed from the `<arg_key>`/`<arg_value>` pairs with per-argument typing from the tool schema. vLLM returns `json.dumps(arg_dct, ensure_ascii=False)`, e.g. `"{\"location\": \"Beijing\", \"unit\": \"celsius\"}"`. Clients `json.loads()` it before use.
|
||||||
|
- **Request side — tool results** are sent back as `role: "tool"` messages, e.g.:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"role": "tool", "tool_call_id": "call_abc123", "content": "{\"temperature\": 26, \"unit\": \"celsius\", \"condition\": \"Sunny\"}"}
|
||||||
|
```
|
||||||
|
|
||||||
|
The chat template renders only `content` (inside `<tool_response>`); `tool_call_id` is **ignored by the template** and matters only for the client's own bookkeeping. Order results to match the calls.
|
||||||
|
- **Request side — assistant tool-call history**: the OpenAI shape carries `function.arguments` as a JSON **string**, but the chat template iterates `arguments.items()` and therefore needs an **object**. vLLM/SGLang parse the string back into a dict before rendering; if you call `tokenizer.apply_chat_template` directly, pass `arguments` as a dict (and optionally `reasoning_content` as a string) or the template will raise.
|
||||||
|
- Disable thinking via `extra_body={"chat_template_kwargs": {"enable_thinking": false}}` (OpenAI Python client) — this flips the template to the `/nothink` + pre-filled `<think></think>` path.
|
||||||
|
|
||||||
|
## Parsing notes & gotchas
|
||||||
|
|
||||||
|
- **String values are unquoted; typing needs the schema.** The decisive rule: a `<arg_value>` is a literal string iff the parameter is string-typed in the tool's JSON Schema; otherwise it is JSON. vLLM's `_is_string_type` and SGLang's `get_argument_type` both walk `properties[arg].type` (handling `anyOf`/`oneOf`/`enum`/`allOf`/type-arrays). If the schema is missing/loose, they fall back to "try `json.loads`, then `ast.literal_eval`, then treat as string" — so a bare word like `celsius` survives as a string, while `26` becomes a number. A string value that *looks* like JSON (e.g. a parameter typed `string` whose value is `{"a":1}`) is correctly kept as the literal string only because the schema says `string`.
|
||||||
|
- **Extraction regexes (GLM-4.5/4.6).** vLLM: calls via `<tool_call>.*?</tool_call>` (DOTALL); name/body via `<tool_call>([^\n]*)\n(.*)</tool_call>`; pairs via `<arg_key>(.*?)</arg_key>\s*<arg_value>(.*?)</arg_value>`. The name regex **requires a newline** after the name — matching the 4.5/4.6 template. SGLang uses an equivalent `(?:\\n|\n)` form so it also tolerates literal escaped `\n`.
|
||||||
|
- **`</arg_value>` in a value breaks parsing.** Values are captured non-greedily up to the next `</arg_value>`; a value whose text contains `</arg_value>` (or `</tool_call>`) truncates early. There is no escaping mechanism in the wire format.
|
||||||
|
- **Tool calls are parsed from `content` only, not from reasoning.** A `<tool_call>` emitted inside `<think>…</think>` is ignored by the tool parser (vLLM's reasoning/tool parsers cooperate so only post-`</think>` content is scanned). Don't expect calls made "while thinking" to fire.
|
||||||
|
- **Guided decoding is suppressed for GLM.** For `tool_choice: "required"` or a named tool, vLLM deliberately does **not** apply JSON structured-outputs/guided decoding, because that would force JSON output and conflict with GLM's XML syntax; the parser extracts from free-form XML instead.
|
||||||
|
- **`skip_special_tokens` must be off.** Although the tool/think tags are `special: false`, vLLM forces `skip_special_tokens = False` when tools are enabled (defensive against transformers 5.x detokenization changes) so the literal `<tool_call>`/`</tool_call>` text survives for the regex.
|
||||||
|
- **Streaming.** Long string arguments used to be buffered until the closing tag (vLLM issue #32829); the current parser re-parses the accumulated text each delta and emits only the diff, streaming incremental string content with an open-quote-then-fill strategy and holding back any partial trailing tag (`partial_tag_overlap`). The streamed tool name is the text before the first `\n` or `<arg_key>`. SGLang implements the same as an explicit XML→JSON state machine (`INIT → IN_KEY → WAITING_VALUE → IN_VALUE`). Malformed tails (a missing `</arg_value>` before `</tool_call>`) are closed off heuristically.
|
||||||
|
- **Lineage — GLM-4.5 vs GLM-4.6:** identical wire format and identical `chat_template.jinja` (same content hash); the same `glm45` parser serves both.
|
||||||
|
- **Lineage — GLM-4.7 / GLM-5 changed the format.** Newer models drop the structural newlines: the function name may sit **directly** before the first `<arg_key>` (no newline), zero-argument calls may be `<tool_call>func</tool_call>`, and parallel calls may be emitted **back-to-back with no separator** (`…</tool_call><tool_call>…`). These require the distinct `Glm47MoeModelToolParser` (vLLM, `structural_tag_model="glm_4_7"`) / `Glm47MoeDetector` (SGLang), whose `func_detail_regex` makes the newline and the argument section optional (`<tool_call>\s*(\S+?)\s*(<arg_key>.*)?</tool_call>`). Do **not** use a GLM-4.7 stream to validate a GLM-4.5 parser or vice versa.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- Chat template (authoritative; rendered locally for the byte-exact streams), GLM-4.5 commit `cbb2c7c…`: https://huggingface.co/zai-org/GLM-4.5/resolve/main/chat_template.jinja — the `blob`/web path redirects to the model-card API; verified via the raw `resolve/main` cache.
|
||||||
|
- Identical GLM-4.6 template (same content hash, confirming shared format): https://huggingface.co/zai-org/GLM-4.6/resolve/main/chat_template.jinja
|
||||||
|
- Special-token IDs and `special` flags (`added_tokens_decoder`, `additional_special_tokens`): https://huggingface.co/zai-org/GLM-4.5/resolve/main/tokenizer_config.json
|
||||||
|
- Stop tokens (`eos_token_id = [151329, 151336, 151338]`): https://huggingface.co/zai-org/GLM-4.5/resolve/main/generation_config.json
|
||||||
|
- Model card (server flags `--tool-call-parser glm45 --reasoning-parser glm45`, `enable_thinking` switch, parser links): https://huggingface.co/zai-org/GLM-4.5
|
||||||
|
- vLLM GLM-4.5/4.6 tool parser (`Glm4MoeModelToolParser`: regexes, schema-driven string typing, JSON-string `arguments`, streaming, `skip_special_tokens`): https://github.com/vllm-project/vllm/blob/main/vllm/tool_parsers/glm4_moe_tool_parser.py
|
||||||
|
- vLLM GLM-4.7 tool parser (`Glm47MoeModelToolParser`: same-line name, optional/zero args): https://github.com/vllm-project/vllm/blob/main/vllm/tool_parsers/glm47_moe_tool_parser.py
|
||||||
|
- SGLang GLM-4.5/4.6 detector (`Glm4MoeDetector`: format docstring, XML→JSON state machine, argument typing): https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/function_call/glm4_moe_detector.py
|
||||||
|
- SGLang GLM-4.7 detector (`Glm47MoeDetector`: newline-less / back-to-back calls): https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/function_call/glm47_moe_detector.py
|
||||||
|
- vLLM tool-calling docs: https://docs.vllm.ai/en/latest/features/tool_calling/
|
||||||
@@ -0,0 +1,224 @@
|
|||||||
|
# OpenAI Harmony response format
|
||||||
|
|
||||||
|
Harmony is the response format OpenAI trained its open-weight `gpt-oss` models on (`gpt-oss-20b`, `gpt-oss-120b`, released August 2025). It defines the conversation envelope, the multi-channel reasoning/answer separation, and the function-calling wire syntax. The models will not work correctly if prompted without it. The format deliberately mirrors the OpenAI *Responses* API (roles, channels, recipients) rather than the older Chat Completions shape.
|
||||||
|
|
||||||
|
Tokens are produced with the `o200k_harmony` encoding (the `o200k_base` BPE vocab plus a block of Harmony special tokens; see the table below). The reference renderer/parser is the Rust crate `openai-harmony` (Python bindings: `pip install openai-harmony`; encoding name `HarmonyEncodingName.HARMONY_GPT_OSS`).
|
||||||
|
|
||||||
|
You only deal with raw Harmony if you build your own inference loop. Served through an OpenAI-compatible endpoint the server handles it for you:
|
||||||
|
|
||||||
|
- **Ollama / LM Studio / HuggingFace**: Harmony is applied internally; you send normal OpenAI-style JSON.
|
||||||
|
- **vLLM**: `vllm serve openai/gpt-oss-120b --enable-auto-tool-choice --tool-call-parser openai --reasoning-parser openai_gptoss`. Note the tool-call parser flag is `openai` (not `harmony`). vLLM also exposes a Harmony-native path through the `/v1/responses` endpoint.
|
||||||
|
- **SGLang**: `python3 -m sglang.launch_server --model-path openai/gpt-oss-20b --reasoning-parser gpt-oss --tool-call-parser gpt-oss` (in NVIDIA Dynamo disaggregated mode: `--dyn-tool-call-parser harmony --dyn-reasoning-parser gpt_oss`).
|
||||||
|
|
||||||
|
The chat template shipped with the gpt-oss weights renders these same token sequences from the standard `messages`/`tools` arrays.
|
||||||
|
|
||||||
|
## Special tokens
|
||||||
|
|
||||||
|
All Harmony control tokens have the literal form `<|type|>` (ASCII pipes `|`, U+007C — no unicode variants). They are real single tokens in `o200k_harmony`, not text that is BPE-split. The structurally meaningful ones:
|
||||||
|
|
||||||
|
| Token (verbatim) | Token ID | Purpose |
|
||||||
|
| :--------------- | :------- | :------ |
|
||||||
|
| `<\|start\|>` | `200006` | Begins a message; immediately followed by the header (role, optional recipient/channel/content-type). |
|
||||||
|
| `<\|end\|>` | `200007` | Ends a fully-formed message. |
|
||||||
|
| `<\|message\|>` | `200008` | Header → content transition. Everything after it (until a stop/end token) is the message body. |
|
||||||
|
| `<\|channel\|>` | `200005` | Introduces the channel field of the header (`analysis` / `commentary` / `final`). |
|
||||||
|
| `<\|constrain\|>` | `200003` | Marks the content-type / constrained-decoding format in a tool-call header (e.g. `<\|constrain\|>json`). |
|
||||||
|
| `<\|return\|>` | `200002` | Stop token: the model finished its final answer. Decode-time only (see normalization note). |
|
||||||
|
| `<\|call\|>` | `200012` | Stop token: the model is emitting a tool call and wants it executed. |
|
||||||
|
|
||||||
|
`<|return|>` and `<|call|>` are the two valid generation stop tokens — halt inference on either.
|
||||||
|
|
||||||
|
The encoding also defines (same `o200k_harmony` block, IDs `199998`–`200013`) `<|startoftext|>` (199998), `<|endoftext|>` (199999), and reserved slots `<|reserved_200000|>`, `<|reserved_200001|>`, `<|reserved_200004|>`, `<|reserved_200009|>`–`<|reserved_200011|>`, `<|reserved_200013|>`, plus a bulk reserved range `<|reserved_200014|>`…`<|reserved_201088|>`. The renderer additionally knows the names `<|refusal|>`, `<|untrusted|>`, `<|end_untrusted|>`, `<|meta_end|>` but they are not part of the committed gpt-oss vocabulary and do not appear in normal traffic.
|
||||||
|
|
||||||
|
## Roles / channels / turn structure
|
||||||
|
|
||||||
|
**Message envelope.** Every message is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>{header}<|message|>{content}<|end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
`{header}` always begins with the role and may carry an optional recipient (`to=...`), channel, and content-type. A completed message ends with `<|end|>`; an assistant message being generated ends instead with a stop token (`<|return|>` or `<|call|>`).
|
||||||
|
|
||||||
|
**Roles** (five). The instruction hierarchy used to resolve conflicts is `system` > `developer` > `user` > `assistant` > `tool`.
|
||||||
|
|
||||||
|
| Role | Purpose |
|
||||||
|
| :--- | :------ |
|
||||||
|
| `system` | Identity, knowledge cutoff / current date, reasoning effort, valid-channels declaration, built-in tools. NOT the user-facing "system prompt". |
|
||||||
|
| `developer` | The conventional "system prompt": instructions + the `# Tools` function declarations + (optional) structured-output schema. |
|
||||||
|
| `user` | End-user input. |
|
||||||
|
| `assistant` | Model output. Carries a channel and, for tool calls, a recipient. |
|
||||||
|
| `tool` | Output of an executed tool. The message's *author/role is the tool's own name* (e.g. `functions.get_current_weather`), not the literal word `tool`. |
|
||||||
|
|
||||||
|
**Channels** (assistant output only; the channel is mandatory on every assistant message):
|
||||||
|
|
||||||
|
| Channel | Purpose |
|
||||||
|
| :------ | :------ |
|
||||||
|
| `analysis` | Raw chain-of-thought (reasoning). Not held to the same safety bar as `final`; do not show to end users. Built-in `python`/`browser` calls usually go here. |
|
||||||
|
| `commentary` | Function tool calls, and user-visible "preambles" (action plans) before calling multiple tools. |
|
||||||
|
| `final` | The user-facing answer. |
|
||||||
|
|
||||||
|
**Reasoning effort** is set in the system message as `Reasoning: high` (or `medium` / `low`; default is medium). The model emits CoT into `analysis` and the answer into `final`.
|
||||||
|
|
||||||
|
**CoT carry-over rule.** On the next turn, drop prior `analysis` messages *if* the last assistant turn ended in a `final` message. The exception is an in-progress tool-calling turn: the `analysis` that preceded a tool call MUST be fed back in alongside the tool result so the model can continue its reasoning (the `openai-harmony` renderer does this via `RenderConversationConfig { auto_drop_analysis: true }`).
|
||||||
|
|
||||||
|
## Tool definitions
|
||||||
|
|
||||||
|
Function tools are advertised in the **developer** message under a `# Tools` section, inside a TypeScript-style `namespace functions { ... }`. (Built-in `browser`/`python` tools are instead declared in the **system** message under their own `# Tools` / `## browser` / `## python` headings.) The renderer converts each JSON Schema into a TS type with these rules:
|
||||||
|
|
||||||
|
- No-arg function → `type name = () => any;`
|
||||||
|
- With args → the single parameter is named `_` and its object type is inlined: `type name = (_: { ... }) => any;`
|
||||||
|
- Return type is always `any`.
|
||||||
|
- A property `description` becomes a `//` comment on the line *above* the field; a JSON Schema `title` renders as `// TITLE` followed by a `//` blank-comment line; `examples` render as `// Examples:` then `// - "value"` lines.
|
||||||
|
- Optional (non-`required`) fields get a trailing `?`. A `default` renders as a trailing `// default: <value>` comment; an `enum` becomes a `"a" | "b"` union; `oneOf` becomes a multi-line `|` union; JSON `integer` maps to TS `number`.
|
||||||
|
- One blank line separates function definitions; the block closes with `} // namespace functions`.
|
||||||
|
|
||||||
|
If the developer message has no instruction text, the `# Instructions` heading is omitted and the message is just the `# Tools` block. When any function is defined, the system message gains the routing line `Calls to these tools must go to the commentary channel: 'functions'.`
|
||||||
|
|
||||||
|
Verbatim developer-message example (instructions + two functions), exactly as the renderer emits it:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>developer<|message|># Instructions
|
||||||
|
|
||||||
|
Use a friendly tone.
|
||||||
|
|
||||||
|
# Tools
|
||||||
|
|
||||||
|
## functions
|
||||||
|
|
||||||
|
namespace functions {
|
||||||
|
|
||||||
|
// Gets the location of the user.
|
||||||
|
type get_location = () => any;
|
||||||
|
|
||||||
|
// Gets the current weather in the provided location.
|
||||||
|
type get_current_weather = (_: {
|
||||||
|
// The city and state, e.g. San Francisco, CA
|
||||||
|
location: string,
|
||||||
|
format?: "celsius" | "fahrenheit", // default: celsius
|
||||||
|
}) => any;
|
||||||
|
|
||||||
|
// Gets the current weather in the provided list of locations.
|
||||||
|
type get_multiple_weathers = (_: {
|
||||||
|
// List of city and state, e.g. ["San Francisco, CA", "New York, NY"]
|
||||||
|
locations: string[],
|
||||||
|
format?: "celsius" | "fahrenheit", // default: celsius
|
||||||
|
}) => any;
|
||||||
|
|
||||||
|
} // namespace functions<|end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Tool-call format
|
||||||
|
|
||||||
|
A function call is an **assistant** message on the **commentary** channel, addressed to the tool via recipient `to=functions.<name>`, with content-type `<|constrain|>json` and the JSON arguments as the body, terminated by the `<|call|>` stop token.
|
||||||
|
|
||||||
|
The recipient may appear in the *role section* or the *channel section* of the header — both are valid Harmony and the parser accepts either. The model commonly emits it in the channel section:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>assistant<|channel|>commentary to=functions.get_current_weather <|constrain|>json<|message|>{"location":"San Francisco, CA"}<|call|>
|
||||||
|
```
|
||||||
|
|
||||||
|
The `openai-harmony` renderer, when re-serializing a stored call, places the recipient in the role section instead (note the `<|constrain|>` is preceded by a space in both forms):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>assistant to=functions.get_current_weather<|channel|>commentary <|constrain|>json<|message|>{"location":"San Francisco, CA"}<|call|>
|
||||||
|
```
|
||||||
|
|
||||||
|
The arguments body is a raw JSON object. The `<|constrain|>json` content-type signals JSON (and is the hook for constrained/grammar-based decoding); the `<|constrain|>` token is optional, and the content-type may also be a bare word such as `code` (seen with built-in tools). Built-in tools differ only in channel and recipient: they typically render on `analysis`, with recipient `browser.search` / `browser.open` / `browser.find` or always `python`.
|
||||||
|
|
||||||
|
## Multiple / parallel tool calls
|
||||||
|
|
||||||
|
Harmony has no special "parallel" wrapper. Multiple calls are just multiple consecutive messages. The model may first emit an optional **preamble** — a *user-visible* assistant message on the `commentary` channel (unlike `analysis`, this is meant to be shown) — then one tool-call message per function. Each individual call still ends with its own `<|call|>` stop token, so a host that stops on `<|call|>` collects calls one at a time, executes, feeds the result back, and resumes:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|channel|>analysis<|message|>{reasoning}<|end|><|start|>assistant<|channel|>commentary<|message|>**Action plan**:
|
||||||
|
1. Generate an HTML file
|
||||||
|
2. Generate a JavaScript for the Node.js server
|
||||||
|
3. Start the server
|
||||||
|
---
|
||||||
|
Will start executing the plan step by step<|end|><|start|>assistant<|channel|>commentary to=functions.generate_file<|constrain|>json<|message|>{"template": "basic_html", "path": "index.html"}<|call|>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Tool-result format
|
||||||
|
|
||||||
|
The executed tool's output is fed back as a message whose **author/role is the tool's name**, addressed back to the assistant (`to=assistant`), on the **commentary** channel, ending with `<|end|>`. This is the canonical (recommended) form:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>functions.get_current_weather to=assistant<|channel|>commentary<|message|>{"sunny": true, "temperature": 20}<|end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
The header ordering is `{toolname} to=assistant<|channel|>commentary`. Built-in tool results follow the same shape (e.g. `<|start|>browser.search to=assistant<|channel|>commentary<|message|>{"result": "https://openai.com/"}<|end|>`). The minimal form the renderer accepts when channel/recipient are not set on the message is just `<|start|>{toolname}<|message|>{output}<|end|>`, but emitting the full `to=assistant<|channel|>commentary` header is what the reference parser round-trips and is recommended. After appending the result, restart generation by emitting the next `<|start|>assistant`.
|
||||||
|
|
||||||
|
## End-to-end example
|
||||||
|
|
||||||
|
Complete multi-turn weather exchange: system + developer prompt → user question → assistant analysis CoT → assistant commentary tool call → tool result → assistant final answer. This is a single contiguous token stream (newlines inside headers are only between top-level messages for readability; in practice messages are concatenated with no separator).
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>system<|message|>You are ChatGPT, a large language model trained by OpenAI.
|
||||||
|
Knowledge cutoff: 2024-06
|
||||||
|
Current date: 2025-06-28
|
||||||
|
|
||||||
|
Reasoning: high
|
||||||
|
|
||||||
|
# Valid channels: analysis, commentary, final. Channel must be included for every message.
|
||||||
|
Calls to these tools must go to the commentary channel: 'functions'.<|end|><|start|>developer<|message|># Instructions
|
||||||
|
|
||||||
|
Use a friendly tone.
|
||||||
|
|
||||||
|
# Tools
|
||||||
|
|
||||||
|
## functions
|
||||||
|
|
||||||
|
namespace functions {
|
||||||
|
|
||||||
|
// Gets the current weather in the provided location.
|
||||||
|
type get_current_weather = (_: {
|
||||||
|
// The city and state, e.g. San Francisco, CA
|
||||||
|
location: string,
|
||||||
|
format?: "celsius" | "fahrenheit", // default: celsius
|
||||||
|
}) => any;
|
||||||
|
|
||||||
|
} // namespace functions<|end|><|start|>user<|message|>What is the weather like in SF?<|end|><|start|>assistant<|channel|>analysis<|message|>User wants the weather in San Francisco. Use get_current_weather.<|end|><|start|>assistant<|channel|>commentary to=functions.get_current_weather <|constrain|>json<|message|>{"location":"San Francisco, CA"}<|call|><|start|>functions.get_current_weather to=assistant<|channel|>commentary<|message|>{"sunny": true, "temperature": 20}<|end|><|start|>assistant<|channel|>final<|message|>It's sunny and about 20°C in San Francisco right now.<|return|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Turn boundaries:
|
||||||
|
|
||||||
|
- The host stops generation at `<|call|>`, parses the `commentary` call, runs `get_current_weather`, and appends the `functions.get_current_weather to=assistant` result message.
|
||||||
|
- It then appends `<|start|>assistant` and resumes. The preceding `analysis` message is kept (the turn ended in a tool call, not a `final`), so the model can continue its reasoning.
|
||||||
|
- Generation stops at `<|return|>`. When this turn is persisted into history for a *later* turn, normalize the trailing `<|return|>` to `<|end|>` (see next note).
|
||||||
|
|
||||||
|
**`<|return|>` normalization.** `<|return|>` is a decode-time stop token only. When you store the assistant's reply into history for the next turn, replace the trailing `<|return|>` with `<|end|>` so every stored message is a well-formed `<|start|>{header}<|message|>{content}<|end|>`. (For supervised training targets, ending the example with `<|return|>` is correct.)
|
||||||
|
|
||||||
|
## OpenAI-compatible API mapping
|
||||||
|
|
||||||
|
When a server (vLLM/SGLang/Ollama) bridges Harmony to Chat Completions JSON:
|
||||||
|
|
||||||
|
- **`finish_reason`**: `tool_calls` when generation stopped on `<|call|>`; `stop` when it stopped on `<|return|>`.
|
||||||
|
- **`message.tool_calls[]`**: one entry per `commentary` `to=functions.*` call. `function.name` is the recipient with the `functions.` namespace stripped (`get_current_weather`). `function.arguments` is a **JSON string** (the verbatim `<|message|>` body), matching OpenAI semantics — not a parsed object.
|
||||||
|
- **`tool_call_id`**: Harmony has no native call ID. The server synthesizes one (e.g. `call_abc123`) and is responsible for correlating the follow-up `role:"tool"` message back to the Harmony tool-result envelope (recipient `to=functions.<name>` / call order).
|
||||||
|
- **Tool result messages** (`{"role":"tool","tool_call_id":...,"content":...}`) are rendered into `<|start|>{toolname} to=assistant<|channel|>commentary<|message|>{content}<|end|>`. The server maps `tool_call_id` → the original function name to build the `{toolname}` author.
|
||||||
|
- **Reasoning**: `analysis`-channel text is surfaced as `reasoning_content` (vLLM/SGLang) or as a `reasoning`/`thinking` field, and is generally not echoed back on subsequent requests. `final`-channel text is the normal `message.content`. `commentary` preambles, if surfaced, also map to assistant content.
|
||||||
|
- **`tools` / `tool_choice`** request fields are compiled by the chat template into the developer-message `namespace functions { ... }` block; the system message gains the commentary-routing line.
|
||||||
|
|
||||||
|
## Parsing notes & gotchas
|
||||||
|
|
||||||
|
- **Two stop tokens.** Always stop on both `<|return|>` and `<|call|>`. Stopping only on `<|return|>` will run past tool calls; stopping only on `<|end|>` is wrong for assistant generation.
|
||||||
|
- **Recipient position varies.** `to=functions.<name>` may be in the role section (`<|start|>assistant to=...<|channel|>commentary`) or the channel section (`<|channel|>commentary to=... `). A parser must accept both. A space precedes `<|constrain|>` in both renderings.
|
||||||
|
- **Channel is mandatory** on assistant messages; the system message even reminds the model ("Channel must be included for every message."). Missing-channel output is malformed.
|
||||||
|
- **Tool author, not `tool`.** The tool-result message's role is the tool's *name* (`functions.get_current_weather`), not the literal string `tool`. Splitting `functions.x` into namespace + function is the parser's job.
|
||||||
|
- **CoT dropping is conditional.** Drop `analysis` only when the previous assistant turn ended on `final`. Dropping the `analysis` that immediately precedes a `<|call|>` breaks multi-step tool reasoning.
|
||||||
|
- **`arguments` is a string.** Do not double-encode. The body after `<|message|>` is already serialized JSON; pass it through as the `arguments` string.
|
||||||
|
- **Content-type variants.** `<|constrain|>json` is typical, but the content-type can be a bare token (`json`, `code`); treat `<|constrain|>` as optional metadata, not a guarantee of valid JSON. Enforce JSON validity with constrained decoding / your own grammar — the prompt format alone does not guarantee schema adherence (same caveat applies to structured-output `# Response Formats`).
|
||||||
|
- **Streaming.** Use a stateful parser (the library ships `StreamableParser`) so partial UTF-8 and the header/channel/recipient/content-type fields are reconstructed incrementally; a naive substring scan mishandles multi-byte splits and the optional header fields. `parse_messages_from_completion_tokens` takes `strict=True|False` — `strict=False` tolerates some malformed headers. Do not pass the trailing stop token into the parser.
|
||||||
|
- **Encoding.** Use `o200k_harmony` (the `o200k_base` ranks plus the Harmony specials above). Treat the `<|...|>` tokens as atomic special tokens during both encode and decode; encoding them as ordinary text yields different ranks and corrupts the stream.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- OpenAI Cookbook — OpenAI harmony response format: https://cookbook.openai.com/articles/openai-harmony
|
||||||
|
- openai/harmony renderer (README): https://github.com/openai/harmony
|
||||||
|
- openai/harmony canonical format guide: https://raw.githubusercontent.com/openai/harmony/main/docs/format.md
|
||||||
|
- openai/harmony special-token registry (`o200k_harmony` IDs): https://raw.githubusercontent.com/openai/harmony/main/src/tiktoken_ext/public_encodings.rs
|
||||||
|
- openai/harmony renderer/parser tests and schema→TS logic: https://raw.githubusercontent.com/openai/harmony/main/src/tests.rs , https://raw.githubusercontent.com/openai/harmony/main/src/encoding.rs
|
||||||
|
- openai/harmony test fixtures (verbatim rendered streams): `test-data/test_render_functions_with_parameters.txt`, `test-data/test_does_not_drop_if_ongoing_analysis.txt`, `test-data/test_tool_response_parsing.txt`, `test-data/test_streamable_parser.txt`, `test-data/test_browser_and_function_tool.txt` (https://github.com/openai/harmony/tree/main/test-data)
|
||||||
|
- vLLM tool calling / gpt-oss parser flags: https://docs.vllm.ai/en/latest/features/tool_calling/
|
||||||
|
- SGLang gpt-oss usage (`--tool-call-parser gpt-oss`): https://docs.sglang.io/basic_usage/gpt_oss.html
|
||||||
@@ -0,0 +1,182 @@
|
|||||||
|
# Kimi K2 tool-calling format
|
||||||
|
|
||||||
|
Native tool-calling convention of Moonshot AI's **Kimi K2** family (`moonshotai/Kimi-K2-Instruct` and `-Base`, `model_type: "kimi_k2"`, 1T-param MoE). It is a ChatML-like envelope built on a TikToken tokenizer (160K vocab): every turn is `<|im_{class}|>{name}<|im_middle|>{body}<|im_end|>`, and tool calls are emitted inside the assistant turn wrapped by a dedicated `<|tool_calls_section_begin|>…<|tool_calls_section_end|>` block. All control tokens are plain ASCII `<|…|>` forms (no fullwidth/unicode variants, unlike DeepSeek). An inference server turns the raw stream into OpenAI-style `tool_calls` with a parser: vLLM and SGLang both expose `--tool-call-parser kimi_k2` (vLLM additionally requires `--enable-auto-tool-choice`). The chat template (a standalone `chat_template.jinja` since the 2025.8.11 update) injects the tool schemas and renders the per-turn markers.
|
||||||
|
|
||||||
|
This document was verified against the model card, the official `docs/tool_call_guidance.md` and `docs/deploy_guidance.md` (GitHub `MoonshotAI/Kimi-K2`), the raw `chat_template.jinja` and `tokenizer_config.json` from the HF repo (rendered locally for the byte-exact streams below), and the vLLM `kimi_k2` tool parser source.
|
||||||
|
|
||||||
|
## Special tokens
|
||||||
|
|
||||||
|
The five tool-call markers required for manual parsing, plus the ChatML envelope markers. Token IDs are from `tokenizer_config.json` (`added_tokens_decoder`).
|
||||||
|
|
||||||
|
| Token (verbatim) | ID | Purpose |
|
||||||
|
|---|---|---|
|
||||||
|
| `<\|tool_calls_section_begin\|>` | 163595 | Opens the tool-call section inside an assistant turn |
|
||||||
|
| `<\|tool_call_begin\|>` | 163597 | Opens one individual tool call |
|
||||||
|
| `<\|tool_call_argument_begin\|>` | 163598 | Separates the tool-call ID from its JSON arguments |
|
||||||
|
| `<\|tool_call_end\|>` | 163599 | Closes one individual tool call |
|
||||||
|
| `<\|tool_calls_section_end\|>` | 163596 | Closes the tool-call section |
|
||||||
|
| `<\|im_system\|>` | 163594 | Start marker for system-class turns (`system`, `tool`, `tool_declare`) |
|
||||||
|
| `<\|im_user\|>` | 163587 | Start marker for a user turn |
|
||||||
|
| `<\|im_assistant\|>` | 163588 | Start marker for an assistant turn |
|
||||||
|
| `<\|im_middle\|>` | 163601 | Separates the role/name header from the message body |
|
||||||
|
| `<\|im_end\|>` | 163586 | Ends any turn |
|
||||||
|
| `[BOS]` | 163584 | Sequence-begin token (see notes; not emitted by the chat template) |
|
||||||
|
| `[EOS]` | 163585 | Sequence-end token |
|
||||||
|
|
||||||
|
Notes on exactness:
|
||||||
|
- The five tool tokens use ASCII pipe `|` (U+007C) and underscores; reproduce them exactly. There are no fullwidth pipe (`|`) or `▁` variants in Kimi K2.
|
||||||
|
- `<|im_middle|>` is the only envelope token whose ID (163601) is out of sequence with the others (163586–163599); a `163600` slot is unused.
|
||||||
|
- Image inputs render via a content macro as the literal sequence `<|media_start|>image<|media_content|><|media_pad|><|media_end|>`. These media markers appear in the template but are **not** registered in `added_tokens_decoder`, so they tokenize as ordinary text rather than single special tokens. They are irrelevant to text tool calling and are listed here only for completeness.
|
||||||
|
|
||||||
|
## Roles / channels / turn structure
|
||||||
|
|
||||||
|
Kimi K2 uses a ChatML-style envelope. Every message is rendered as:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_{class}|>{name}<|im_middle|>{body}<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- There are exactly **three** start-marker tokens, chosen by `role`:
|
||||||
|
- `user` → `<|im_user|>`
|
||||||
|
- `assistant` → `<|im_assistant|>`
|
||||||
|
- everything else (`system`, `tool`, and the synthetic `tool_declare`) → `<|im_system|>`
|
||||||
|
- The `{name}` segment between the marker and `<|im_middle|>` is `message.name or message.role`. This is the only "channel"/sub-role label Kimi K2 has. For ordinary turns it is literally `system`, `user`, or `assistant`; for a tool-result turn it is the tool's `name` (the function name) when supplied, otherwise `tool`; for the tool-schema turn it is the literal `tool_declare`.
|
||||||
|
- `<|im_end|>` terminates every turn. The chat template does **not** emit `[BOS]`/`[EOS]`; turn boundaries are purely `<|im_*|>` markers (the tokenizer is TikToken-based with `add_bos_token`/`add_eos_token` unset, and the manual-parse flow feeds the rendered template straight to `/completions`).
|
||||||
|
- **Default system prompt:** if the first message is not a `system` message, the template injects `<|im_system|>system<|im_middle|>You are Kimi, an AI assistant created by Moonshot AI.<|im_end|>` before the first turn.
|
||||||
|
- **Generation prompt:** with `add_generation_prompt=True` the template ends with `<|im_assistant|>assistant<|im_middle|>`, and the model generates from there.
|
||||||
|
- **Thinking/reasoning:** `Kimi-K2-Instruct` is a "reflex-grade" model with no long thinking, so there is no reasoning channel in this format. (Thinking variants are handled separately — vLLM ships a distinct `kimi_k2` reasoning parser keyed on a `</think>` token — but that is out of scope for the Instruct tool-call format documented here.)
|
||||||
|
|
||||||
|
## Tool definitions
|
||||||
|
|
||||||
|
Available tools are advertised in a single dedicated turn placed at the very top of the prompt (before any system/user turn), using the synthetic `tool_declare` sub-role under the `<|im_system|>` marker:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_system|>tool_declare<|im_middle|>{TOOLS_JSON}<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
`{TOOLS_JSON}` is the standard OpenAI-style `tools` array serialized to JSON with **compact separators** `(',', ':')` (no spaces). The array elements are passed through verbatim, i.e. each is `{"type":"function","function":{"name":…,"description":…,"parameters":{…}}}` with a JSON-Schema `parameters` object. Example (single tool, exactly as emitted):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_system|>tool_declare<|im_middle|>[{"type":"function","function":{"name":"get_weather","description":"Get weather information. Call this tool when the user needs to get weather information","parameters":{"type":"object","required":["city"],"properties":{"city":{"type":"string","description":"City name"}}}}}]<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
The `tool_declare` turn is rendered only when `tools` is non-empty.
|
||||||
|
|
||||||
|
## Tool-call format
|
||||||
|
|
||||||
|
When the model decides to call a function, it emits — inside the assistant turn, after any natural-language content — a tool-calls section. Minimal single call (this is the assistant generation that follows `<|im_assistant|>assistant<|im_middle|>`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool_calls_section_begin|><|tool_call_begin|>functions.get_weather:0<|tool_call_argument_begin|>{"city": "Beijing"}<|tool_call_end|><|tool_calls_section_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Anatomy of one call:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool_call_begin|> functions.{func_name}:{idx} <|tool_call_argument_begin|> {JSON arguments} <|tool_call_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- The token between `<|tool_call_begin|>` and `<|tool_call_argument_begin|>` is the **tool-call ID**, with the fixed form `functions.{func_name}:{idx}`.
|
||||||
|
- `functions.` is a literal prefix (it is not derived from the tool schema).
|
||||||
|
- `{func_name}` is the called function's name; the function name is recovered by parsing it back out of this ID, not from a separate field.
|
||||||
|
- `{idx}` is the **0-based call index** within the current assistant turn (`0` for the first call, `1` for the second, …).
|
||||||
|
- After `<|tool_call_argument_begin|>` comes the raw JSON arguments object (e.g. `{"city": "Beijing"}`), terminated by `<|tool_call_end|>`.
|
||||||
|
- All calls of the turn live between one `<|tool_calls_section_begin|>` / `<|tool_calls_section_end|>` pair. Any assistant text content precedes `<|tool_calls_section_begin|>`.
|
||||||
|
- The whole assistant turn is still closed by `<|im_end|>` and the completion's `finish_reason` becomes `tool_calls`.
|
||||||
|
|
||||||
|
## Multiple / parallel tool calls
|
||||||
|
|
||||||
|
Two or more calls in one turn are emitted as consecutive `<|tool_call_begin|>…<|tool_call_end|>` blocks inside a single section, with the index incrementing per call. Raw assistant emission for two parallel calls:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool_calls_section_begin|><|tool_call_begin|>functions.get_weather:0<|tool_call_argument_begin|>{"city": "Beijing"}<|tool_call_end|><|tool_call_begin|>functions.get_weather:1<|tool_call_argument_begin|>{"city": "Shanghai"}<|tool_call_end|><|tool_calls_section_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Note the IDs `functions.get_weather:0` and `functions.get_weather:1` — same function, distinct trailing index. The index is per-turn (it resets to `0` in the next assistant turn).
|
||||||
|
|
||||||
|
## Tool-result format
|
||||||
|
|
||||||
|
Tool execution results are fed back as a turn with `role: "tool"`. Because `tool` is not `user`/`assistant`, it renders under the `<|im_system|>` marker; the sub-role label is the message's `name` (the function name) when present, else `tool`. The body is a literal `## Return of {tool_call_id}` header line followed by the result content:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_system|>get_weather<|im_middle|>## Return of functions.get_weather:0
|
||||||
|
{"weather": "Sunny"}<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- `{tool_call_id}` echoes the exact ID from the originating call (`functions.get_weather:0`), which is how the model correlates a result with the call that produced it.
|
||||||
|
- The result `content` is inserted verbatim on the line after the header; callers typically pass a JSON string (e.g. `json.dumps(tool_result)`).
|
||||||
|
- If the `tool` message omits `name`, the envelope becomes `<|im_system|>tool<|im_middle|>## Return of …`.
|
||||||
|
|
||||||
|
## End-to-end example
|
||||||
|
|
||||||
|
A complete multi-turn weather exchange. These are the exact rendered streams (system + user supplied explicitly; line breaks inside a turn are literal, turns are otherwise contiguous).
|
||||||
|
|
||||||
|
**Stage 1 — prompt fed to the model** (`tools` set, `add_generation_prompt=True`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_system|>tool_declare<|im_middle|>[{"type":"function","function":{"name":"get_weather","description":"Get weather information. Call this tool when the user needs to get weather information","parameters":{"type":"object","required":["city"],"properties":{"city":{"type":"string","description":"City name"}}}}}]<|im_end|><|im_system|>system<|im_middle|>You are Kimi, an AI assistant created by Moonshot AI.<|im_end|><|im_user|>user<|im_middle|>What's the weather like in Beijing today? Use the tool to check.<|im_end|><|im_assistant|>assistant<|im_middle|>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Assistant generation** (model output; server reports `finish_reason: "tool_calls"`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool_calls_section_begin|><|tool_call_begin|>functions.get_weather:0<|tool_call_argument_begin|>{"city": "Beijing"}<|tool_call_end|><|tool_calls_section_end|><|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Stage 2 — prompt for the next turn**, after appending the assistant tool-call turn and the tool result turn (`add_generation_prompt=True`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_system|>tool_declare<|im_middle|>[{"type":"function","function":{"name":"get_weather","description":"Get weather information. Call this tool when the user needs to get weather information","parameters":{"type":"object","required":["city"],"properties":{"city":{"type":"string","description":"City name"}}}}}]<|im_end|><|im_system|>system<|im_middle|>You are Kimi, an AI assistant created by Moonshot AI.<|im_end|><|im_user|>user<|im_middle|>What's the weather like in Beijing today? Use the tool to check.<|im_end|><|im_assistant|>assistant<|im_middle|><|tool_calls_section_begin|><|tool_call_begin|>functions.get_weather:0<|tool_call_argument_begin|>{"city": "Beijing"}<|tool_call_end|><|tool_calls_section_end|><|im_end|><|im_system|>get_weather<|im_middle|>## Return of functions.get_weather:0
|
||||||
|
{"weather": "Sunny"}<|im_end|><|im_assistant|>assistant<|im_middle|>
|
||||||
|
```
|
||||||
|
|
||||||
|
**Final assistant generation** (model produces natural-language answer terminated by `<|im_end|>`; `finish_reason: "stop"`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
It's sunny in Beijing today.<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
## OpenAI-compatible API mapping
|
||||||
|
|
||||||
|
With a server parser active (`--tool-call-parser kimi_k2`), the raw stream maps onto the Chat Completions shape as follows:
|
||||||
|
|
||||||
|
- `choices[].finish_reason` = `"tool_calls"` when the turn contained a tool-calls section (otherwise `"stop"`).
|
||||||
|
- `choices[].message.tool_calls[]` — one entry per `<|tool_call_begin|>…<|tool_call_end|>` block:
|
||||||
|
- `.id` = the raw call ID verbatim, e.g. `"functions.get_weather:0"`.
|
||||||
|
- `.type` = `"function"`.
|
||||||
|
- `.function.name` = the function name parsed out of the ID. vLLM computes `id.split(":")[0].split(".")[-1]` → `"get_weather"`.
|
||||||
|
- `.function.arguments` = a **JSON string** (the raw text captured between `<|tool_call_argument_begin|>` and `<|tool_call_end|>`), e.g. `"{\"city\": \"Beijing\"}"`. Clients `json.loads()` it before use.
|
||||||
|
- Tool results are sent back as messages of the form:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"role": "tool", "tool_call_id": "functions.get_weather:0", "name": "get_weather", "content": "{\"weather\": \"Sunny\"}"}
|
||||||
|
```
|
||||||
|
|
||||||
|
`tool_call_id` must equal the `id` returned for the call; `name` becomes the `<|im_system|>{name}<|im_middle|>` sub-role; `content` becomes the body after `## Return of …`.
|
||||||
|
- Streaming: deltas arrive as `choices[].delta.tool_calls[]` with an `index`; the function `name`/`id` stream once the call header is complete, then `function.arguments` streams as incremental string fragments to be concatenated (standard OpenAI tool-call streaming assembly).
|
||||||
|
|
||||||
|
Moonshot's hosted API (`platform.moonshot.ai`) exposes both OpenAI- and Anthropic-compatible endpoints; the Anthropic-compatible one scales temperature as `real_temperature = request_temperature * 0.6`. Recommended sampling temperature for `Kimi-K2-Instruct` is `0.6`.
|
||||||
|
|
||||||
|
## Parsing notes & gotchas
|
||||||
|
|
||||||
|
- **ID → name parsing differs between references.** The official `tool_call_guidance.md` extracts the name with `function_id.split('.')[1].split(':')[0]`, which assumes the ID is exactly `functions.{name}` with no extra dots. vLLM uses the more robust `function_id.split(":")[0].split(".")[-1]` (takes the last dot-segment before `:{idx}`). Prefer the vLLM form so function names containing `.` are handled.
|
||||||
|
- **Extraction regexes differ too.** Guidance: `<\|tool_call_begin\|>\s*(?P<tool_call_id>[\w\.]+:\d+)\s*<\|tool_call_argument_begin\|>\s*(?P<function_arguments>.*?)\s*<\|tool_call_end\|>`. vLLM: ID class is `[^<]+:\d+` and the argument body uses a negative lookahead `(?:(?!<\|tool_call_begin\|>).)*?` so adjacent calls aren't merged. Both run with `DOTALL`.
|
||||||
|
- **`skip_special_tokens` must be False.** The parser depends on the literal marker text surviving detokenization; vLLM forces `skip_special_tokens = False` when tools are enabled and `tool_choice != "none"`. If markers are stripped, no tool call is detected.
|
||||||
|
- **Arguments are unvalidated raw text.** Whatever the model emits between the argument marker and `<|tool_call_end|>` is passed straight through as the `arguments` string; it must be valid JSON for downstream `json.loads`, and the model can emit malformed/truncated JSON. Validate before executing.
|
||||||
|
- **Index semantics.** `{idx}` is the per-turn call counter starting at `0`; it is not a global counter and resets each assistant turn. Do not assume IDs are unique across turns — disambiguate by turn when persisting history.
|
||||||
|
- **Streaming marker splits.** Section and call markers can be split across token boundaries. vLLM holds back any trailing suffix that partially matches a marker (`partial_tag_overlap`) to avoid leaking marker bytes into streamed content, and only streams a call's name once its header is fully received.
|
||||||
|
- **`finish_reason` varies by engine.** The official guide explicitly warns the terminal `finish_reason` for tool calls "may vary across different engines"; loop on `finish_reason == "tool_calls"` but be defensive.
|
||||||
|
- **Engine fallback.** Kimi K2 reuses the DeepSeek-V3 architecture; `config.json` sets `model_type: "kimi_k2"` so engines apply the right parser. If you force `model_type: "deepseek_v3"` as a compatibility workaround, no native Kimi tool parser is available and you must parse the `<|tool_calls_section_*|>` markers manually.
|
||||||
|
- **Parser availability.** vLLM ships both a Python (`KimiK2ToolParser`) and a newer Rust tool parser; SGLang implements its own `kimi_k2` parser. All key off the same five markers and the `functions.{name}:{idx}` ID convention documented here.
|
||||||
|
- **Whitespace artifact.** When no `system` message is supplied, the template injects the default system prompt and a small `\n ` (newline + two spaces) can appear before the first `<|im_user|>` marker. It is harmless (tokenizes around the markers), but supplying an explicit system message yields the clean streams shown above.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- Model card (Tool Calling section, OpenAI-style example, deployment/API notes): https://huggingface.co/moonshotai/Kimi-K2-Instruct
|
||||||
|
- Official tool-call guidance (markers, ID convention, manual parser, `extract_tool_call_info`): https://raw.githubusercontent.com/MoonshotAI/Kimi-K2/main/docs/tool_call_guidance.md (the HF `resolve`/`blob` paths redirected to the model card; verified against this GitHub raw file)
|
||||||
|
- Deployment guide (`--tool-call-parser kimi_k2`, `--enable-auto-tool-choice`, SGLang flag, `model_type` fallback): https://raw.githubusercontent.com/MoonshotAI/Kimi-K2/main/docs/deploy_guidance.md
|
||||||
|
- Chat template (`chat_template.jinja`, rendered locally for byte-exact streams): https://huggingface.co/moonshotai/Kimi-K2-Instruct/resolve/main/chat_template.jinja
|
||||||
|
- Tokenizer config (special-token IDs in `added_tokens_decoder`): https://huggingface.co/moonshotai/Kimi-K2-Instruct/resolve/main/tokenizer_config.json
|
||||||
|
- vLLM `kimi_k2` tool parser (markers, regex, name-parsing, `skip_special_tokens`, streaming): https://github.com/vllm-project/vllm/blob/main/vllm/tool_parsers/kimi_k2_tool_parser.py
|
||||||
|
- vLLM PR adding the parser: https://github.com/vllm-project/vllm/pull/20789
|
||||||
|
- vLLM tool-calling docs: https://docs.vllm.ai/en/latest/features/tool_calling/
|
||||||
@@ -0,0 +1,276 @@
|
|||||||
|
# pi-native tool-call format
|
||||||
|
|
||||||
|
The **pi-native** format is the tool-call serialization used by the omp / pi coding agent. Unlike the JSON-in-a-tag conventions (Hermes/Qwen, Harmony) and unlike the fully separate JSON content-block channel (Anthropic Messages API), pi-native serializes each call as an **XML-flavored block** whose tag carries the tool name — `<call:NAME>…</call:NAME>` — and whose arguments are child elements named after the parameters. It is **schema-driven**: the tool's JSON Schema decides how each value is typed (string vs number vs object vs array) and which compact spellings are legal.
|
||||||
|
|
||||||
|
This document is a **specification** of the format (it is the contract a renderer must emit and a parser must accept), not a reverse-engineering of trained model weights. It is designed around four goals:
|
||||||
|
|
||||||
|
- **Token economy** — the common cases (a single scalar argument; a single string payload) collapse to one short line.
|
||||||
|
- **Verbatim payloads** — a large multi-line string argument (a patch body, a file's contents, a shell script) is carried **raw**, with no JSON string-escaping and no entity-encoding, terminated by the call's own unique closing tag.
|
||||||
|
- **Human legibility** — a call reads like the function it denotes; nesting maps to nesting.
|
||||||
|
- **Lenient parsing** — the tags are plain text matched by a tolerant parser (regex / streaming state machine), not a strict XML parser; output is *not* required to be well-formed XML.
|
||||||
|
|
||||||
|
Scope: pi-native specifies only the **tool-call (and argument) serialization**. It is **envelope-agnostic** — the `<call:…>` blocks are emitted as ordinary assistant text and embed unchanged in any conversation envelope (ChatML, Harmony, the Anthropic two-role shape, …). Reasoning channels, role markers, and result delivery are the host envelope's concern; the one envelope-level requirement pi-native imposes is in [Tool-result correlation](#tool-result-correlation).
|
||||||
|
|
||||||
|
Lineage: the attribute spelling (`<call:read path="…"/>`) follows Anthropic's modern attribute XML (`<invoke name="…">` / `<parameter name="…">`); the schema-driven **unquoted** value rule (a bare string value carries no quotes; non-strings are JSON) follows GLM-4.5's `<arg_value>` convention. pi-native folds both into one recursive, name-on-the-tag grammar and adds the verbatim **inline body** for bulk string arguments. See [`anthropic.md`](./anthropic.md) and [`glm-4.5.md`](./glm-4.5.md) in this folder.
|
||||||
|
|
||||||
|
## Structural tags
|
||||||
|
|
||||||
|
pi-native has **no special tokens**. Every marker is plain UTF-8 text that BPE-splits like any other text and survives detokenization unchanged; a parser matches the tags as literal substrings (and MUST work even when the surrounding stream is not valid XML). All brackets are ASCII `<` `>` `/` and the literal colon `:`. There are no namespaces, no `<?xml?>` prolog, no entity expansion, and no CDATA sections.
|
||||||
|
|
||||||
|
| Tag (verbatim) | Role |
|
||||||
|
|---|---|
|
||||||
|
| `<call:NAME …>` … `</call:NAME>` | One tool call. `NAME` is the tool/recipient name; it is repeated on the closing tag. |
|
||||||
|
| `<call:NAME …/>` | Self-closing tool call (all arguments supplied as attributes). |
|
||||||
|
| `<KEY>` … `</KEY>` | One argument (or one nested field). `KEY` is the parameter name. |
|
||||||
|
| `<KEY …/>` | Self-closing argument: an object-valued field whose scalar sub-fields are attributes, or an empty value. |
|
||||||
|
| `KEY="…"` / `KEY='…'` / `KEY=…` | An attribute: a scalar field given inline on a tag. Quotes are **delimiters, not type markers** (see [value coercion](#value-coercion)). |
|
||||||
|
|
||||||
|
Names (tool names and parameter names) match `^[A-Za-z_][A-Za-z0-9_-]*$`. The `call:` prefix is a literal four-character marker plus the colon; the colon is what distinguishes a call block from a nested argument element of the same name.
|
||||||
|
|
||||||
|
## Tool-call forms
|
||||||
|
|
||||||
|
A single call has three interchangeable surface forms. Which forms are legal for a given tool is decided by its parameter schema; a renderer SHOULD pick the most compact legal form, and a parser MUST accept all three.
|
||||||
|
|
||||||
|
### 1. Element form (canonical, fully general)
|
||||||
|
|
||||||
|
Each top-level argument is a child element named after the parameter; the element body is the value. This form expresses every schema — scalars, strings, arrays, and nested objects:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:read>
|
||||||
|
<path>src/server/auth.ts</path>
|
||||||
|
<offset>50</offset>
|
||||||
|
</call:read>
|
||||||
|
```
|
||||||
|
|
||||||
|
→ `read({ "path": "src/server/auth.ts", "offset": 50 })` (`offset` is JSON because the schema types it as a number; `path` is a verbatim string).
|
||||||
|
|
||||||
|
### 2. Attribute form (compact scalars)
|
||||||
|
|
||||||
|
When the arguments being passed are **top-level scalars** (string, number, integer, boolean, null), they MAY be written as attributes on the call tag. With every argument as an attribute the tag is self-closing:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:read path="src/server/auth.ts"/>
|
||||||
|
```
|
||||||
|
|
||||||
|
→ `read({ "path": "src/server/auth.ts" })`.
|
||||||
|
|
||||||
|
Attributes and child elements MAY be combined on a non-self-closing call tag — attributes carry the scalars, child elements carry anything structured:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:read path="src/server/auth.ts">
|
||||||
|
<offset>50</offset>
|
||||||
|
</call:read>
|
||||||
|
```
|
||||||
|
|
||||||
|
An attribute whose value cannot be represented as a scalar (an object or array argument) MUST use the element form instead — there is no attribute spelling for structured values on a call tag.
|
||||||
|
|
||||||
|
### 3. Inline-body form (verbatim string payload)
|
||||||
|
|
||||||
|
When the tool's parameters are **all strings** — most often a single string parameter — the call body MAY be the argument value written **verbatim**, with no child element tags:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:edit>
|
||||||
|
*** Begin Patch
|
||||||
|
@@ src/server/auth.ts
|
||||||
|
- return user;
|
||||||
|
+ return user ?? null;
|
||||||
|
*** End Patch
|
||||||
|
</call:edit>
|
||||||
|
```
|
||||||
|
|
||||||
|
→ `edit({ "input": "*** Begin Patch\n@@ src/server/auth.ts\n- return user;\n+ return user ?? null;\n*** End Patch" })`.
|
||||||
|
|
||||||
|
Rules for the inline body:
|
||||||
|
|
||||||
|
- The body fills the **first parameter not already supplied by an attribute**. With no attributes that is simply the first parameter (the "first argument verbatim").
|
||||||
|
- It is permitted only when that target parameter is **string**-typed (the enabling condition "the type only contains string arguments"); any *other* parameters set on the same call MUST be scalars given as attributes.
|
||||||
|
- The value is captured **verbatim** up to the call's own closing tag `</call:NAME>`. No JSON escaping, no entity decoding. The body MAY freely contain `<`, `>`, `&`, quotes, JSON, even other `<call:…>`-looking text — the only sequence it MUST NOT contain is the literal closer `</call:NAME>`. Because that closer carries the tool name, collisions are far rarer than with a short generic delimiter.
|
||||||
|
- Whitespace: a single newline immediately after the opening `>` and a single newline immediately before the closing `</` are treated as block delimiters and are **not** part of the value; all other whitespace (indentation, blank lines, trailing spaces) is preserved exactly.
|
||||||
|
|
||||||
|
A multi-string tool can still use the inline body for its bulk argument by passing the others as attributes — the body then fills the first parameter left unset:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:write path="notes/todo.md">
|
||||||
|
# TODO
|
||||||
|
- ship pi-native parser
|
||||||
|
</call:write>
|
||||||
|
```
|
||||||
|
|
||||||
|
→ `write({ "path": "notes/todo.md", "content": "# TODO\n- ship pi-native parser" })` (here `path` is given by attribute, so the body fills the next string parameter, `content`).
|
||||||
|
|
||||||
|
Inline-eligible tools may always fall back to the element form; `<call:edit><input>…</input></call:edit>` and the inline `<call:edit>…</call:edit>` are equivalent.
|
||||||
|
|
||||||
|
## Value model
|
||||||
|
|
||||||
|
The body of a call (and of any nested element) maps to JSON by a single recursive rule set. **Typing is driven by the parameter's JSON Schema**; the parser only falls back to syntactic heuristics when no schema is available.
|
||||||
|
|
||||||
|
### Value coercion
|
||||||
|
|
||||||
|
Let `coerce(text, type)` produce the JSON value for a captured scalar `text`:
|
||||||
|
|
||||||
|
- `type == "string"` → the value is `text`, **verbatim** (never JSON-parsed, never unquoted). This is why `<path>4</path>` for a string parameter is the string `"4"`, and a Windows path `C:\new\tab` survives intact.
|
||||||
|
- `type` is a non-string scalar (`number` / `integer` / `boolean` / `null`) → `JSON.parse(text)` (so `<offset>50</offset>` → `50`, `<recursive>true</recursive>` → `true`).
|
||||||
|
- `type` unknown (no schema) → **best-effort JSON coercion**: try `JSON.parse(text)`; on success use the parsed value (number, boolean, null, quoted-string, object, or array); on failure treat `text` as a literal string. So a bare `4` becomes the number `4`, `foo.ts` (not valid JSON) becomes `"foo.ts"`.
|
||||||
|
|
||||||
|
The same `coerce` applies to **attribute values** after the surrounding quotes (if any) are stripped — the quotes are XML delimiters only. Hence both spellings below are identical, and both yield the **number** `4` (not the string `"4"`) when `y` is untyped/numeric:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<object y=4/> → { "object": { "y": 4 } }
|
||||||
|
<object y="4"/> → { "object": { "y": 4 } }
|
||||||
|
```
|
||||||
|
|
||||||
|
Consequence to internalize: under loose/no schema, quoting does **not** force a string — `"4"` still coerces to `4`. To carry a numeric-looking value *as a string*, the parameter MUST be `string`-typed in the schema (then the verbatim rule keeps `"4"` → `"4"`). Unquoted attribute values run until whitespace or the closing `>` / `/>`; spaces around `=` are tolerated; a bare attribute with no `=value` (e.g. `<call:tool dry_run/>`) denotes boolean `true`.
|
||||||
|
|
||||||
|
### Scalars and strings
|
||||||
|
|
||||||
|
A scalar argument is one element (or one attribute). String values are unquoted and verbatim; non-string scalars are JSON literals:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:bash command="ls -la" timeout=30/>
|
||||||
|
```
|
||||||
|
|
||||||
|
→ `bash({ "command": "ls -la", "timeout": 30 })`.
|
||||||
|
|
||||||
|
### Arrays — repeat the element
|
||||||
|
|
||||||
|
An array-typed field is expressed by **repeating** its element; each occurrence contributes one item, in order:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<list>x</list>
|
||||||
|
<list>y</list>
|
||||||
|
```
|
||||||
|
|
||||||
|
→ `"list": ["x", "y"]`.
|
||||||
|
|
||||||
|
A field the schema types as an **array always yields an array, even for a single occurrence** — so one `<list>x</list>` under an array-typed `list` is `["x"]`, not `"x"`. When no schema is available the parser falls back to a count heuristic: a name appearing **2+ times among its siblings** is an array; a name appearing **once** is a scalar (so schema typing is the only way to express a one-element array with the heuristic alone). Item values coerce by the array's item type (`<ports>80</ports><ports>443</ports>` → `[80, 443]` for a `number[]`); arrays of objects repeat a nested block (see below). There is no attribute spelling for an array (attributes cannot repeat) — arrays require element form.
|
||||||
|
|
||||||
|
### Objects — a nested block
|
||||||
|
|
||||||
|
An object-typed field opens its own block and follows the **same rules recursively**: its child elements become its properties, repeated children become arrays, and nested object children open further blocks.
|
||||||
|
|
||||||
|
```text
|
||||||
|
<object>
|
||||||
|
<list>x</list>
|
||||||
|
</object>
|
||||||
|
```
|
||||||
|
|
||||||
|
→ `"object": { "list": ["x"] }` (with `object` typed object and `list` typed array).
|
||||||
|
|
||||||
|
An object's **scalar** sub-fields MAY instead be written as attributes — `<object y=4/>` is shorthand for `<object><y>4</y></object>`. Attributes and child elements may be combined on the same object element (attributes for scalars, children for structured sub-fields). An empty object is `<object/>` or `<object></object>` → `{}`.
|
||||||
|
|
||||||
|
### Recursion
|
||||||
|
|
||||||
|
The call body, an object element's body, and an array item's body are all parsed by the identical procedure. Parsing element `E` (tag = field name `F`, schema type `T`):
|
||||||
|
|
||||||
|
1. Gather `E`'s attributes → scalar properties via `coerce`.
|
||||||
|
2. Determine `E`'s body shape from `T` (or, with no schema, from whether the body's first non-whitespace content is a child tag):
|
||||||
|
- `T` object → properties from child elements (+ the attributes from step 1).
|
||||||
|
- `T` array (item type `Ti`) → collect **all** siblings named `F`; each occurrence is one item parsed as `Ti`.
|
||||||
|
- `T` scalar/string → the body is captured text; value = `coerce(text, T)`.
|
||||||
|
3. The call itself is element `E` with no enclosing key: its attributes + child elements **are** the arguments object directly (the tool name on `<call:NAME>` is the recipient, not a key).
|
||||||
|
|
||||||
|
## Multiple / parallel tool calls
|
||||||
|
|
||||||
|
There is no wrapper element around a set of calls. Parallel calls are simply **consecutive `<call:…>` blocks** in one assistant turn (separated by whitespace/newlines; interleaved prose is allowed and is ordinary content):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:read path="src/a.ts"/>
|
||||||
|
<call:read path="src/b.ts"/>
|
||||||
|
```
|
||||||
|
|
||||||
|
A parser returns these as `tool_calls[0]`, `tool_calls[1]`, … in emission order. The host executes them and returns one result per call, in the same order (see correlation, next).
|
||||||
|
|
||||||
|
## Tool definitions and schema dependence
|
||||||
|
|
||||||
|
pi-native does not prescribe how tools are advertised; a host typically lists them as JSON Schema, exactly as the OpenAI / Anthropic / Hermes families do. What pi-native **requires** is that the parser have access to each tool's parameter schema, because the schema is what disambiguates:
|
||||||
|
|
||||||
|
- string (verbatim, unquoted) vs other scalar (JSON) values;
|
||||||
|
- a one-element array vs a scalar (a single `<list>…</list>`);
|
||||||
|
- which body shape (text vs nested members) a non-self-closing element carries;
|
||||||
|
- whether the inline-body form is legal (first unset parameter is a string).
|
||||||
|
|
||||||
|
Without a schema the parser MUST degrade gracefully to the syntactic fallbacks named above (JSON-coerce scalars; repetition-counts for arrays; child-tag presence for object bodies). The fallbacks are lossy at exactly the ambiguous points the schema would resolve, so production hosts SHOULD always supply the schema.
|
||||||
|
|
||||||
|
## Tool-result correlation
|
||||||
|
|
||||||
|
pi-native calls carry **no per-call wire id** (like GLM and Qwen, unlike Anthropic's `toolu_…`). Results are therefore correlated to calls **positionally, by emission order**: the host delivers tool outputs in the same order the `<call:…>` blocks appeared, using whatever its envelope provides for tool output (a `tool`/`user` turn, a Harmony tool message, an Anthropic `tool_result` block, …). When a transport requires an id (e.g. an OpenAI-compatible bridge), the host synthesizes one and maintains the call↔result mapping itself; the id never appears in the pi-native text.
|
||||||
|
|
||||||
|
## End-to-end example
|
||||||
|
|
||||||
|
A short agent turn exercising all three forms plus nesting. Schemas in play: `read(path: string, offset?: number)`, `bash(command: string, timeout?: number)`, `edit(input: string)`, and a synthetic `configure(object: { list: string[]; y?: number })`.
|
||||||
|
|
||||||
|
```text
|
||||||
|
I'll inspect the file, run the tests, then apply the fix.
|
||||||
|
|
||||||
|
<call:read path="src/server/auth.ts"/>
|
||||||
|
|
||||||
|
<call:bash command="bun test src/server/auth.test.ts" timeout=120/>
|
||||||
|
|
||||||
|
<call:configure>
|
||||||
|
<object y=4>
|
||||||
|
<list>alpha</list>
|
||||||
|
<list>beta</list>
|
||||||
|
</object>
|
||||||
|
</call:configure>
|
||||||
|
|
||||||
|
<call:edit>
|
||||||
|
*** Begin Patch
|
||||||
|
@@ src/server/auth.ts
|
||||||
|
- return user;
|
||||||
|
+ return user ?? null;
|
||||||
|
*** End Patch
|
||||||
|
</call:edit>
|
||||||
|
```
|
||||||
|
|
||||||
|
Parses to four calls, in order:
|
||||||
|
|
||||||
|
```json
|
||||||
|
[
|
||||||
|
{ "name": "read", "arguments": { "path": "src/server/auth.ts" } },
|
||||||
|
{ "name": "bash", "arguments": { "command": "bun test src/server/auth.test.ts", "timeout": 120 } },
|
||||||
|
{ "name": "configure", "arguments": { "object": { "y": 4, "list": ["alpha", "beta"] } } },
|
||||||
|
{ "name": "edit", "arguments": { "input": "*** Begin Patch\n@@ src/server/auth.ts\n- return user;\n+ return user ?? null;\n*** End Patch" } }
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
Note: `timeout=120` and `y=4` are JSON numbers (numeric/untyped scalars), `path` and the `list` items are verbatim strings (string-typed), `object` opens a nested block whose `y` rides as an attribute while `list` repeats into an array, and the `edit` body is captured verbatim up to `</call:edit>` despite containing `@@`, `-`/`+`, and other non-XML text.
|
||||||
|
|
||||||
|
## Grammar (lenient EBNF)
|
||||||
|
|
||||||
|
This is the shape a tolerant parser accepts; it is intentionally looser than XML (mismatched-but-recoverable tails are closed heuristically — see gotchas).
|
||||||
|
|
||||||
|
```ebnf
|
||||||
|
stream ::= ( text | call )*
|
||||||
|
call ::= self-call | block-call
|
||||||
|
self-call ::= "<call:" Name attr* ws? "/>"
|
||||||
|
block-call ::= "<call:" Name attr* ">" call-body "</call:" Name ">"
|
||||||
|
call-body ::= members | inline-text ; inline-text only if first param is string
|
||||||
|
members ::= ( ws | element )*
|
||||||
|
element ::= self-element | block-element
|
||||||
|
self-element ::= "<" Name attr* ws? "/>" ; object via attrs, or empty value
|
||||||
|
block-element::= "<" Name attr* ">" ( members | scalar-text ) "</" Name ">"
|
||||||
|
attr ::= ws Name ( ws? "=" ws? attr-val )? ; bare Name → boolean true
|
||||||
|
attr-val ::= '"' dq-chars '"' | "'" sq-chars "'" | bareword
|
||||||
|
Name ::= [A-Za-z_] [A-Za-z0-9_-]*
|
||||||
|
scalar-text ::= < any chars up to the matching close tag, verbatim >
|
||||||
|
inline-text ::= < any chars up to "</call:" Name ">", verbatim >
|
||||||
|
```
|
||||||
|
|
||||||
|
## Parsing notes & gotchas
|
||||||
|
|
||||||
|
- **Schema decides string-vs-JSON.** A `string`-typed value is verbatim and unquoted; everything else is JSON. With no schema, scalars best-effort JSON-coerce and fall back to string. This is the single most error-prone rule (identical to GLM-4.5's unquoted strings): emitting `"San Francisco"` for a string parameter yields the literal value *including the quote characters*.
|
||||||
|
- **Quotes are delimiters, not types.** `y="4"` and `y=4` both coerce to the number `4` under loose/no schema. Quoting an attribute never makes it a string; only a `string` schema type does.
|
||||||
|
- **Arrays = repetition; single-element arrays need the schema.** Two same-named siblings is unambiguously an array. One occurrence is a scalar under the count heuristic and an array only because the schema says so — a parser without the schema cannot tell `<list>x</list>` (scalar) from a one-element array.
|
||||||
|
- **Verbatim bodies are delimited by the named closer.** The inline body and any `string`-typed element body are captured up to their matching `</call:NAME>` / `</KEY>`. A body that contains that exact closing sequence truncates early; there is no escaping mechanism. The inline body's risk is minimal because the delimiter includes the tool name (`</call:edit>`), but a short string-typed *element* (e.g. `<note>…</note>`) is more exposed — prefer the inline-body form for any value that might contain markup, or keep such values in the single-string inline payload.
|
||||||
|
- **Element form vs inline body.** A block call whose body's first non-whitespace content is a child tag matching a known parameter is parsed as element form; otherwise (all-string tool) it is the inline body. A string value that legitimately *starts* with a `<param>`-looking token is the one ambiguity — emit such a tool in element form, or rely on the schema (a tool with structured params is never inline-eligible).
|
||||||
|
- **No ids; order is the contract.** Calls carry no id; results MUST be returned in call order. Reordering results silently misattributes them.
|
||||||
|
- **Lenient, not strict XML.** Do not feed pi-native to an XML parser: tag names contain a colon (`call:read`), attribute values may be unquoted, bodies are not entity-encoded, and the stream need not be balanced beyond each call's own open/close. Match the tags as literals (regex / streaming state machine).
|
||||||
|
- **Streaming.** A stateful parser emits the tool name as soon as `<call:NAME` closes, then streams attribute/child deltas; for an inline body it streams body text incrementally and holds back any partial trailing `</call:` until it can decide whether it is the closer. Coercion of a scalar can only finalize at the value's close tag (a partial number/boolean is not yet valid JSON).
|
||||||
|
- **Whitespace.** Element/inline bodies preserve all whitespace except one leading and one trailing newline that delimit the block. Attribute and indentation whitespace between tags is insignificant.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
pi-native is specified here; it is not derived from a published model template. Its two direct influences are documented in this folder:
|
||||||
|
|
||||||
|
- Anthropic attribute XML (`<invoke name="…">` / `<parameter name="…">`, "parsed with regular expressions", not required to be valid XML): [`anthropic.md`](./anthropic.md).
|
||||||
|
- GLM-4.5 schema-driven, **unquoted** string values vs JSON non-strings, and positional (id-less) call↔result correlation: [`glm-4.5.md`](./glm-4.5.md).
|
||||||
@@ -0,0 +1,206 @@
|
|||||||
|
# Qwen3 tool-calling format (Hermes convention)
|
||||||
|
|
||||||
|
Tool-calling convention of Alibaba's **Qwen3** family (`Qwen/Qwen3-*`: dense `0.6B–32B` and MoE `30B-A3B`/`235B-A22B`; same template line as `Qwen2.5-*` and `QwQ-32B`). It is the **Hermes** convention — the XML+JSON format originated by NousResearch's Hermes 2 Pro and adopted verbatim by Qwen, plus a long tail of community fine-tunes. The envelope is **ChatML**: every turn is `<|im_start|>{role}\n{body}<|im_end|>\n`. Available tools are advertised in the system turn inside a `<tools>…</tools>` block (one JSON spec per line); the model emits each call as a `<tool_call>\n{json}\n</tool_call>` block whose `arguments` is a **nested JSON object** (not a stringified JSON); tool results are fed back inside `<tool_response>…</tool_response>`. Hybrid reasoning is carried in `<think>…</think>`. The format ships in the model's own `chat_template`, so an inference server enables it with no extra template: vLLM uses `--enable-auto-tool-choice --tool-call-parser hermes` (pair with `--reasoning-parser deepseek_r1` for the thinking split); SGLang exposes the matching parsers (e.g. `--reasoning-parser qwen3`).
|
||||||
|
|
||||||
|
Verified against: Qwen's canonical function-calling guide (`qwen.readthedocs.io/en/latest/framework/function_call.html`, read in full incl. the Qwen-Agent + vLLM sections), the byte-exact `chat_template` field of `Qwen/Qwen3-8B`'s `tokenizer_config.json` (HF resolve-cache commit `b968826d9c46dd6066d109eabc6255188de91218`, rendered locally with Jinja2 for the raw streams below) and its `added_tokens_decoder` for token IDs, the NousResearch `Hermes-Function-Calling` README, and the vLLM tool-calling docs (`hermes` parser + Qwen models section).
|
||||||
|
|
||||||
|
## Special tokens
|
||||||
|
|
||||||
|
Only the three ChatML markers are "special" control tokens (`special=true`, skipped by `skip_special_tokens`). The reasoning and tool markers are also single vocabulary tokens (one ID each) but are registered with `special=false`, i.e. they render as ordinary text and are **not** stripped by `skip_special_tokens`. The `<tools>`/`</tools>` wrapper has **no** dedicated token at all — it is plain text that BPE-splits into several tokens. IDs are from `Qwen/Qwen3-8B` `added_tokens_decoder`.
|
||||||
|
|
||||||
|
| Token (verbatim) | ID | `special` | Purpose |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `<\|im_start\|>` | 151644 | true | Start of a turn; followed immediately by the role name + `\n` |
|
||||||
|
| `<\|im_end\|>` | 151645 | true | End of a turn; the chat stop token |
|
||||||
|
| `<\|endoftext\|>` | 151643 | true | Base EOS / pad token |
|
||||||
|
| `<think>` | 151667 | false | Opens the reasoning block |
|
||||||
|
| `</think>` | 151668 | false | Closes the reasoning block |
|
||||||
|
| `<tool_call>` | 151657 | false | Opens one tool call |
|
||||||
|
| `</tool_call>` | 151658 | false | Closes one tool call |
|
||||||
|
| `<tool_response>` | 151665 | false | Opens one tool result |
|
||||||
|
| `</tool_response>` | 151666 | false | Closes one tool result |
|
||||||
|
| `<tools>` … `</tools>` | — | — | Plain text wrapper around the tool list in the system turn (not a single token) |
|
||||||
|
|
||||||
|
Notes on exactness:
|
||||||
|
- All markers use the ASCII pipe `|` (U+007C) and ASCII angle brackets. Qwen3 has **no** fullwidth (`|` U+FF5C) or `▁` (U+2581) variants — that is DeepSeek/SentencePiece territory, not Qwen.
|
||||||
|
- `<|im_start|>` and `<|im_end|>` are the only tokens that matter for splitting turns. Because `<tool_call>`, `</tool_call>`, `<tool_response>`, `<think>`, `</think>` are `special=false`, they survive a `skip_special_tokens=True` decode, which is exactly why the regex-based `hermes` parser can recover them from decoded text.
|
||||||
|
- The model card confirms `</think>` = token `151668` (used by the reference parsing snippet `output_ids[::-1].index(151668)`).
|
||||||
|
|
||||||
|
## Roles / channels / turn structure
|
||||||
|
|
||||||
|
ChatML. Each message renders as:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_start|>{role}
|
||||||
|
{body}<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- Roles: `system`, `user`, `assistant`, `tool`. There is no separate "channel" concept; the only sub-stream is the `<think>` reasoning block inside an assistant turn.
|
||||||
|
- `<|im_end|>\n` terminates every turn. With `add_generation_prompt=True` the prompt ends with `<|im_start|>assistant\n` and the model continues from there.
|
||||||
|
- **System turn:** if the caller supplies a `system` message it becomes the first turn. When `tools` are present, the tool advertisement is merged **into** that same system turn (the user's system text first, then `\n\n`, then the `# Tools` block — see below). Qwen3 injects no default system prompt when none is given.
|
||||||
|
- **Tool-result turns use the `user` envelope.** Qwen3's template maps every `role: "tool"` message into a `<|im_start|>user` turn carrying `<tool_response>` blocks (consecutive tool messages are coalesced into one user turn). This differs from classic Hermes 2 Pro, which used a dedicated `<|im_start|>tool` turn for results — Qwen folds them into `user`.
|
||||||
|
- **Thinking/reasoning:** carried in `<think>…</think>` at the start of an assistant turn (see the Parsing notes for the toggle and the history-rerender rule).
|
||||||
|
|
||||||
|
## Tool definitions
|
||||||
|
|
||||||
|
Tools are advertised inside the system turn. The template emits a fixed preamble, then each tool object serialized with `tool | tojson` (`json.dumps(..., ensure_ascii=False)`) on **its own line**, then a fixed trailer. Each list element is the full OpenAI tool object `{"type": "function", "function": {...}}` (with a JSON-Schema `parameters` object). The exact, verbatim wrapper Qwen3 produces:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_start|>system
|
||||||
|
{optional original system content}
|
||||||
|
|
||||||
|
# Tools
|
||||||
|
|
||||||
|
You may call one or more functions to assist with the user query.
|
||||||
|
|
||||||
|
You are provided with function signatures within <tools></tools> XML tags:
|
||||||
|
<tools>
|
||||||
|
{"type": "function", "function": {"name": "get_current_temperature", "description": "Get current temperature at a location.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The location to get the temperature for, in the format \"City, State, Country\"."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "The unit to return the temperature in. Defaults to \"celsius\"."}}, "required": ["location"]}}}
|
||||||
|
{"type": "function", "function": {"name": "get_temperature_date", "description": "Get temperature at a location and date.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The location to get the temperature for, in the format \"City, State, Country\"."}, "date": {"type": "string", "description": "The date to get the temperature for, in the format \"Year-Month-Day\"."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "The unit to return the temperature in. Defaults to \"celsius\"."}}, "required": ["location", "date"]}}}
|
||||||
|
</tools>
|
||||||
|
|
||||||
|
For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:
|
||||||
|
<tool_call>
|
||||||
|
{"name": <function-name>, "arguments": <args-json-object>}
|
||||||
|
</tool_call><|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- If the first message is a `system` message, its content is placed before `# Tools` (separated by a blank line); otherwise the turn opens straight into `# Tools`.
|
||||||
|
- The trailing instruction is a literal part of the prompt, including the placeholder line `{"name": <function-name>, "arguments": <args-json-object>}` (those angle-bracket tokens are instructions, not emitted output).
|
||||||
|
- Version note: the original Hermes 2 Pro system prompt additionally embedded a `FunctionCall` pydantic schema line (`{"title": "FunctionCall", "type": "object", "properties": {"name": …, "arguments": …}}`). Qwen3 dropped that line; the wrapper above is exactly what Qwen3 emits.
|
||||||
|
|
||||||
|
## Tool-call format
|
||||||
|
|
||||||
|
The model emits each call as a `<tool_call>` line, a single-line JSON object, then `</tool_call>`. Minimal single call:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_call>
|
||||||
|
{"name": "get_current_temperature", "arguments": {"location": "San Francisco, CA, USA", "unit": "celsius"}}
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
- `arguments` is a **nested JSON object**, not a JSON-encoded string. On the wire it is `"arguments": {"location": "..."}` — never `"arguments": "{\"location\": ...}"`. (The template renders a dict argument via `tojson`; only if a caller stored `arguments` as a pre-serialized string does it pass through verbatim.)
|
||||||
|
- The call object has exactly two keys, `name` (string) and `arguments` (object). There is no per-call ID on the wire — the OpenAI-style `tool_call_id` is minted by the server, not the model (see API mapping).
|
||||||
|
- A tool-calling assistant turn may also contain natural-language `content` before the first `<tool_call>`; the template inserts a `\n` between that content and the first call.
|
||||||
|
|
||||||
|
## Multiple / parallel tool calls
|
||||||
|
|
||||||
|
Parallel calls are emitted as consecutive `<tool_call>…</tool_call>` blocks within a single assistant turn, each separated by a newline:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_start|>assistant
|
||||||
|
<tool_call>
|
||||||
|
{"name": "get_current_temperature", "arguments": {"location": "San Francisco, CA, USA"}}
|
||||||
|
</tool_call>
|
||||||
|
<tool_call>
|
||||||
|
{"name": "get_temperature_date", "arguments": {"location": "San Francisco, CA, USA", "date": "2024-10-01"}}
|
||||||
|
</tool_call><|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
The parser returns these as `tool_calls[0]`, `tool_calls[1]`, … in emission order. The application must execute them and return one `<tool_response>` per call, in the same order.
|
||||||
|
|
||||||
|
## Tool-result format
|
||||||
|
|
||||||
|
Each executed result is wrapped in `<tool_response>…</tool_response>`. Qwen3 places them inside a **`user`** turn, and **coalesces** consecutive tool results into one turn (one `<tool_response>` block per result, newline-separated, a single closing `<|im_end|>`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_start|>user
|
||||||
|
<tool_response>
|
||||||
|
{"temperature": 26.1, "location": "San Francisco, CA, USA", "unit": "celsius"}
|
||||||
|
</tool_response>
|
||||||
|
<tool_response>
|
||||||
|
{"temperature": 25.9, "location": "San Francisco, CA, USA", "date": "2024-10-01", "unit": "celsius"}
|
||||||
|
</tool_response><|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
- The body between the tags is the tool's return value (typically a JSON string, but any text is allowed). The function name is **not** repeated inside Qwen3's `<tool_response>` — ordering ties results to calls. (Classic Hermes 2 Pro instead nested `{"name": ..., "content": ...}` inside `<tool_response>` under a `tool` turn; Qwen3's template emits the bare content under a `user` turn.)
|
||||||
|
- At the OpenAI API layer a result message is `{"role": "tool", "content": "...", "tool_call_id": "..."}`; the template renders only its `content` into a `<tool_response>` block.
|
||||||
|
|
||||||
|
## End-to-end example
|
||||||
|
|
||||||
|
Complete multi-turn weather exchange in **non-thinking mode** (`enable_thinking=False`), exactly as `apply_chat_template` renders it for the live flow. With thinking disabled, each generation step injects an empty `<think>\n\n</think>\n\n` after `<|im_start|>assistant\n`; the model then emits its tool call / final answer. Copy-pasteable, byte-exact:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_start|>system
|
||||||
|
You are a helpful assistant. Current Date: 2024-09-30.
|
||||||
|
|
||||||
|
# Tools
|
||||||
|
|
||||||
|
You may call one or more functions to assist with the user query.
|
||||||
|
|
||||||
|
You are provided with function signatures within <tools></tools> XML tags:
|
||||||
|
<tools>
|
||||||
|
{"type": "function", "function": {"name": "get_current_temperature", "description": "Get current temperature at a location.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The location to get the temperature for, in the format \"City, State, Country\"."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "The unit to return the temperature in. Defaults to \"celsius\"."}}, "required": ["location"]}}}
|
||||||
|
</tools>
|
||||||
|
|
||||||
|
For each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:
|
||||||
|
<tool_call>
|
||||||
|
{"name": <function-name>, "arguments": <args-json-object>}
|
||||||
|
</tool_call><|im_end|>
|
||||||
|
<|im_start|>user
|
||||||
|
What's the temperature in San Francisco now?<|im_end|>
|
||||||
|
<|im_start|>assistant
|
||||||
|
<think>
|
||||||
|
|
||||||
|
</think>
|
||||||
|
|
||||||
|
<tool_call>
|
||||||
|
{"name": "get_current_temperature", "arguments": {"location": "San Francisco, CA, USA", "unit": "celsius"}}
|
||||||
|
</tool_call><|im_end|>
|
||||||
|
<|im_start|>user
|
||||||
|
<tool_response>
|
||||||
|
{"temperature": 26.1, "location": "San Francisco, CA, USA", "unit": "celsius"}
|
||||||
|
</tool_response><|im_end|>
|
||||||
|
<|im_start|>assistant
|
||||||
|
<think>
|
||||||
|
|
||||||
|
</think>
|
||||||
|
|
||||||
|
The current temperature in San Francisco is 26.1°C.<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
In **thinking mode** (`enable_thinking=True`, the default) the generation prompt instead ends with a bare `<|im_start|>assistant\n` and the model itself produces the `<think>…real reasoning…</think>` block before the `<tool_call>`. (When re-rendering stored history, the template keeps the `<think>` block only for the last assistant message or messages that carry `reasoning_content`, and strips reasoning from earlier turns — see Parsing notes.)
|
||||||
|
|
||||||
|
## OpenAI-compatible API mapping
|
||||||
|
|
||||||
|
With `--enable-auto-tool-choice --tool-call-parser hermes`, vLLM converts the raw stream into a standard Chat Completions response:
|
||||||
|
|
||||||
|
- `finish_reason`: `"tool_calls"` when the turn ended on tool calls (otherwise `"stop"`).
|
||||||
|
- `message.role`: `"assistant"`; `message.content`: `null` for a pure tool-call turn (any pre-call prose becomes `content`).
|
||||||
|
- `message.tool_calls[]`: one entry per `<tool_call>` block, each:
|
||||||
|
- `id`: server-generated, e.g. `"chatcmpl-tool-924d705adb044ff88e0ef3afdd155f15"` (the model emits no ID).
|
||||||
|
- `type`: `"function"`.
|
||||||
|
- `function.name`: the call's `name`.
|
||||||
|
- `function.arguments`: a **JSON string** at the API boundary, e.g. `'{"location": "San Francisco, CA, USA"}'`. The wire format is a nested object, but the server re-serializes it to a string here (`json.loads(...)` it before use), matching OpenAI and Qwen-Agent.
|
||||||
|
- With thinking + `--reasoning-parser deepseek_r1`, the `<think>…</think>` content is split out into `message.reasoning_content` and removed from `content`.
|
||||||
|
- Feeding results back: append `{"role": "tool", "content": <result>, "tool_call_id": <id-from-the-call>}` for each result. `tool_call_id` links a result to its call (Qwen3's template ignores the id when rendering — ordering is what reaches the model — but the API still requires it).
|
||||||
|
|
||||||
|
Example assistant message returned for the two-call query:
|
||||||
|
|
||||||
|
```text
|
||||||
|
finish_reason='tool_calls'
|
||||||
|
message.content = None
|
||||||
|
message.tool_calls = [
|
||||||
|
{id:'chatcmpl-tool-924d…', type:'function', function:{name:'get_current_temperature', arguments:'{"location": "San Francisco, CA, USA"}'}},
|
||||||
|
{id:'chatcmpl-tool-7e30…', type:'function', function:{name:'get_temperature_date', arguments:'{"location": "San Francisco, CA, USA", "date": "2024-10-01"}'}},
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Parsing notes & gotchas
|
||||||
|
|
||||||
|
- **Arguments object vs string:** on the wire `arguments` is a nested JSON object; the OpenAI layer hands it back as a JSON string. Code that reads the raw stream must parse an object; code that reads the API must `json.loads` the string. Do not double-encode.
|
||||||
|
- **`<tools>` is not a token.** Only count on `<|im_start|>`/`<|im_end|>` (and the `*tool_call*`/`*tool_response*`/`*think*` single tokens) being atomic. `<tools>`/`</tools>` are plain text.
|
||||||
|
- **Regex/streaming parse:** the vLLM `hermes` parser (`vllm/tool_parsers/hermes_tool_parser.py`, `Hermes2ProToolParser`) keys on the literal `<tool_call>` / `</tool_call>` substrings and JSON-decodes the body, supporting multiple blocks per turn. In streaming it buffers from `<tool_call>` until it can incrementally parse `name` then `arguments`; partial argument JSON is emitted as argument deltas. Text before the first `<tool_call>` is streamed as ordinary content.
|
||||||
|
- **Thinking toggle:** `enable_thinking=False` (passed via `chat_template_kwargs={"enable_thinking": False}` over the OpenAI API, or `tokenizer.apply_chat_template(..., enable_thinking=False)`) injects an empty `<think>\n\n</think>\n\n` into the generation prompt, hard-suppressing reasoning. Soft switches `/think` and `/no_think` in a user/system message flip it per-turn when thinking is enabled. Greedy decoding is discouraged for Qwen3 (repetition risk).
|
||||||
|
- **History rerender asymmetry:** when `apply_chat_template` re-renders a stored conversation, it emits the `<think>` block only for the final assistant message or messages carrying `reasoning_content`; reasoning from earlier turns is dropped. So a stored intermediate tool-call assistant turn shows no `<think>` block, while the live generation step that produced it was prefixed with one (in non-thinking mode). Reasoning is preserved only within the current multi-step tool sequence (after the last real user query).
|
||||||
|
- **Reasoning models + stopword templates:** Qwen warns against ReAct-style stopword tool templates for Qwen3, since reasoning text may contain the stopwords and corrupt parsing — use this native Hermes template instead.
|
||||||
|
- **Robustness:** the format is prompt/template-driven, so malformed output is possible (truncated JSON, missing `</tool_call>`, prose mixed into a call, an array serialized as a string). Production parsers should tolerate and, on failure, fall back to treating the text as content. Named / `required` tool_choice routes through vLLM's structured-outputs backend for guaranteed-parseable arguments.
|
||||||
|
- **Version/scope:** this `hermes` template covers `Qwen3-*`, `Qwen2.5-*`, and `QwQ-32B`. It does **not** cover `Qwen3-Coder`, which uses a different XML scheme parsed by vLLM's `qwen3_xml` parser — a separate convention.
|
||||||
|
|
||||||
|
## Sources
|
||||||
|
|
||||||
|
- Qwen function-calling guide: https://qwen.readthedocs.io/en/latest/framework/function_call.html
|
||||||
|
- Qwen3-8B chat template + token IDs (`tokenizer_config.json`, `chat_template` + `added_tokens_decoder`): https://huggingface.co/Qwen/Qwen3-8B/resolve/main/tokenizer_config.json (verified via HF resolve-cache commit `b968826d9c46dd6066d109eabc6255188de91218`)
|
||||||
|
- Qwen3-8B model card (thinking modes, `enable_thinking`, `</think>`=151668): https://huggingface.co/Qwen/Qwen3-8B
|
||||||
|
- NousResearch Hermes-Function-Calling (origin of the convention): https://github.com/NousResearch/Hermes-Function-Calling
|
||||||
|
- vLLM tool-calling docs (`hermes` parser, Qwen models, auto tool choice): https://docs.vllm.ai/en/latest/features/tool_calling/
|
||||||
@@ -1,6 +1,24 @@
|
|||||||
# Changelog
|
# Changelog
|
||||||
|
|
||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
### Breaking Changes
|
||||||
|
|
||||||
|
- Removed `harmony-leak` exports from the `@oh-my-pi/pi-agent-core` package entrypoint
|
||||||
|
- Replaced the experimental `promptToolCalls` agent/loop option with `toolCallSyntax`, selecting an explicit in-band tool-call grammar instead of a boolean GLM-only mode.
|
||||||
|
|
||||||
|
### Added
|
||||||
|
|
||||||
|
- Added support for selecting owned in-band tool-call syntax via `PI_OWNED_TOOLS=<syntax>` (for example `hermes` or `qwen3`) while preserving legacy `PI_OWNED_TOOLS=1/true` as GLM mode
|
||||||
|
- Added owned in-band tool calling for multiple syntaxes (`glm`, `hermes`, `kimi`, `xml`, `anthropic`, `deepseek`, `harmony`, `pi-native`, `qwen3`). Owned mode sends no native provider tools, appends a syntax-specific prompt/catalog, re-encodes prior tool calls/results as grammar-owned text, and parses streamed model output back into canonical tool calls.
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
|
||||||
|
- Added owned in-band syntax support to `Agent` loop configuration resolution by selecting syntax from `toolCallSyntax` or `PI_OWNED_TOOLS` when present
|
||||||
|
|
||||||
|
### Fixed
|
||||||
|
|
||||||
|
- Fixed owned in-band tool-calling requests to omit `toolChoice` after stripping native tools, preventing invalid tool-choice requests
|
||||||
|
- Fixed owned tool calling letting the model fabricate tool results by treating grammar-owned tool-result markers in assistant text as a hard turn boundary: calls before the fabrication are kept, fabricated results and dependent calls are dropped, and the real result is fed back on the next turn.
|
||||||
|
|
||||||
## [15.13.1] - 2026-06-15
|
## [15.13.1] - 2026-06-15
|
||||||
|
|
||||||
@@ -779,4 +797,4 @@ Initial release under @oh-my-pi scope. See previous releases at [badlogic/pi-mon
|
|||||||
### Changed
|
### Changed
|
||||||
|
|
||||||
- `Agent` constructor now has all options optional (empty options use defaults).
|
- `Agent` constructor now has all options optional (empty options use defaults).
|
||||||
- `queueMessage()` is now synchronous (no longer returns a Promise).
|
- `queueMessage()` is now synchronous (no longer returns a Promise).
|
||||||
@@ -15,7 +15,12 @@ import {
|
|||||||
validateToolArguments,
|
validateToolArguments,
|
||||||
zodToWireSchema,
|
zodToWireSchema,
|
||||||
} from "@oh-my-pi/pi-ai";
|
} from "@oh-my-pi/pi-ai";
|
||||||
import { logger, sanitizeText } from "@oh-my-pi/pi-utils";
|
import {
|
||||||
|
encodeInbandToolHistory,
|
||||||
|
renderInbandToolPrompt,
|
||||||
|
type ToolCallSyntax,
|
||||||
|
wrapInbandToolStream,
|
||||||
|
} from "@oh-my-pi/pi-ai/grammar";
|
||||||
import {
|
import {
|
||||||
createHarmonyAuditEvent,
|
createHarmonyAuditEvent,
|
||||||
detectHarmonyLeakInAssistantMessage,
|
detectHarmonyLeakInAssistantMessage,
|
||||||
@@ -25,7 +30,8 @@ import {
|
|||||||
isHarmonyLeakMitigationTarget,
|
isHarmonyLeakMitigationTarget,
|
||||||
recoverHarmonyToolCall,
|
recoverHarmonyToolCall,
|
||||||
signalListLabel,
|
signalListLabel,
|
||||||
} from "./harmony-leak";
|
} from "@oh-my-pi/pi-ai/utils/harmony-leak";
|
||||||
|
import { logger, sanitizeText } from "@oh-my-pi/pi-utils";
|
||||||
import { type AgentRunCoverage, type AgentRunSummary, ToolCallBlockedError } from "./run-collector";
|
import { type AgentRunCoverage, type AgentRunSummary, ToolCallBlockedError } from "./run-collector";
|
||||||
import {
|
import {
|
||||||
type AgentTelemetry,
|
type AgentTelemetry,
|
||||||
@@ -76,6 +82,25 @@ class HarmonyLeakInterruption extends Error {
|
|||||||
this.name = "HarmonyLeakInterruption";
|
this.name = "HarmonyLeakInterruption";
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
function resolveOwnedToolSyntaxFromEnv(value: string | undefined): ToolCallSyntax | undefined {
|
||||||
|
switch (value) {
|
||||||
|
case "1":
|
||||||
|
case "true":
|
||||||
|
return "glm";
|
||||||
|
case "glm":
|
||||||
|
case "hermes":
|
||||||
|
case "kimi":
|
||||||
|
case "xml":
|
||||||
|
case "anthropic":
|
||||||
|
case "deepseek":
|
||||||
|
case "harmony":
|
||||||
|
case "pi":
|
||||||
|
case "qwen3":
|
||||||
|
return value;
|
||||||
|
default:
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
type AssistantContentBlock = AssistantMessage["content"][number];
|
type AssistantContentBlock = AssistantMessage["content"][number];
|
||||||
type AssistantToolCallBlock = Extract<AssistantContentBlock, { type: "toolCall" }>;
|
type AssistantToolCallBlock = Extract<AssistantContentBlock, { type: "toolCall" }>;
|
||||||
@@ -896,6 +921,22 @@ async function streamAssistantResponse(
|
|||||||
llmContext = config.transformProviderContext(llmContext, config.model);
|
llmContext = config.transformProviderContext(llmContext, config.model);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// Owned tool calling: take tool calls away from the provider and run them
|
||||||
|
// through the selected in-band prompt syntax. `PI_OWNED_TOOLS=1` still
|
||||||
|
// force-enables GLM; `PI_OWNED_TOOLS=<syntax>` force-enables that syntax.
|
||||||
|
const ownedSyntax: ToolCallSyntax | undefined =
|
||||||
|
config.toolCallSyntax ?? resolveOwnedToolSyntaxFromEnv(Bun.env.PI_OWNED_TOOLS);
|
||||||
|
let promptToolWireTools: Context["tools"];
|
||||||
|
if (ownedSyntax && llmContext.tools && llmContext.tools.length > 0) {
|
||||||
|
promptToolWireTools = llmContext.tools;
|
||||||
|
llmContext = {
|
||||||
|
...llmContext,
|
||||||
|
systemPrompt: [...(llmContext.systemPrompt ?? []), renderInbandToolPrompt(promptToolWireTools, ownedSyntax)],
|
||||||
|
messages: encodeInbandToolHistory(llmContext.messages, ownedSyntax, promptToolWireTools),
|
||||||
|
tools: undefined,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
const streamFunction = streamFn || streamSimple;
|
const streamFunction = streamFn || streamSimple;
|
||||||
|
|
||||||
// Resolve API key (important for expiring tokens) — do this before resolving
|
// Resolve API key (important for expiring tokens) — do this before resolving
|
||||||
@@ -920,12 +961,22 @@ async function streamAssistantResponse(
|
|||||||
: harmonyAbortController.signal
|
: harmonyAbortController.signal
|
||||||
: signal;
|
: signal;
|
||||||
const repetitionAbortController = new AbortController();
|
const repetitionAbortController = new AbortController();
|
||||||
const finalRequestSignal = requestSignal
|
// Owned tool calling: aborted by the stream wrapper when the model starts
|
||||||
? AbortSignal.any([requestSignal, repetitionAbortController.signal])
|
// fabricating a `<tool_response>`, so the provider stops generating the rest of
|
||||||
: repetitionAbortController.signal;
|
// the hallucinated turn. Merged into the provider signal ONLY (not
|
||||||
|
// `requestSignal`), so it cancels the request without tripping the loop's
|
||||||
|
// external-abort handling (`abortRacePromise` / `requestSignal.aborted`).
|
||||||
|
const promptToolAbortController = ownedSyntax ? new AbortController() : undefined;
|
||||||
|
const providerAbortSignals: AbortSignal[] = [];
|
||||||
|
if (requestSignal) providerAbortSignals.push(requestSignal);
|
||||||
|
providerAbortSignals.push(repetitionAbortController.signal);
|
||||||
|
if (promptToolAbortController) providerAbortSignals.push(promptToolAbortController.signal);
|
||||||
|
const finalRequestSignal =
|
||||||
|
providerAbortSignals.length === 1 ? providerAbortSignals[0]! : AbortSignal.any(providerAbortSignals);
|
||||||
const effectiveTemperature =
|
const effectiveTemperature =
|
||||||
harmonyRetryAttempt > 0 && config.temperature !== undefined ? config.temperature + 0.05 : config.temperature;
|
harmonyRetryAttempt > 0 && config.temperature !== undefined ? config.temperature + 0.05 : config.temperature;
|
||||||
const effectiveToolChoice = dynamicToolChoice ?? config.toolChoice;
|
// Owned tool calling sends no native tools, so any tool_choice would error.
|
||||||
|
const effectiveToolChoice = ownedSyntax ? undefined : (dynamicToolChoice ?? config.toolChoice);
|
||||||
const effectiveReasoning = dynamicReasoning ?? config.reasoning;
|
const effectiveReasoning = dynamicReasoning ?? config.reasoning;
|
||||||
const effectiveDisableReasoning = dynamicDisableReasoning ?? config.disableReasoning;
|
const effectiveDisableReasoning = dynamicDisableReasoning ?? config.disableReasoning;
|
||||||
|
|
||||||
@@ -970,7 +1021,7 @@ async function streamAssistantResponse(
|
|||||||
|
|
||||||
try {
|
try {
|
||||||
return await runInActiveSpan(chatSpan, async () => {
|
return await runInActiveSpan(chatSpan, async () => {
|
||||||
const response = await streamFunction(config.model, llmContext, {
|
let response = await streamFunction(config.model, llmContext, {
|
||||||
...config,
|
...config,
|
||||||
// Hand streamSimple a resolver so its central auth-retry policy can
|
// Hand streamSimple a resolver so its central auth-retry policy can
|
||||||
// re-resolve on 401 / usage-limit: the initial step reuses the key
|
// re-resolve on 401 / usage-limit: the initial step reuses the key
|
||||||
@@ -993,6 +1044,14 @@ async function streamAssistantResponse(
|
|||||||
signal: finalRequestSignal,
|
signal: finalRequestSignal,
|
||||||
onResponse: captureOnResponse,
|
onResponse: captureOnResponse,
|
||||||
});
|
});
|
||||||
|
if (promptToolWireTools && ownedSyntax) {
|
||||||
|
// Re-materialize in-band tool-call text as native toolCall content blocks
|
||||||
|
// so the rest of the loop executes them unchanged. The abort callback
|
||||||
|
// cancels the provider when the model starts fabricating tool results.
|
||||||
|
response = wrapInbandToolStream(response, promptToolWireTools, ownedSyntax, () =>
|
||||||
|
promptToolAbortController?.abort(),
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
let partialMessage: AssistantMessage | null = null;
|
let partialMessage: AssistantMessage | null = null;
|
||||||
let addedPartial = false;
|
let addedPartial = false;
|
||||||
|
|||||||
@@ -22,11 +22,12 @@ import {
|
|||||||
type ToolChoice,
|
type ToolChoice,
|
||||||
type ToolResultMessage,
|
type ToolResultMessage,
|
||||||
} from "@oh-my-pi/pi-ai";
|
} from "@oh-my-pi/pi-ai";
|
||||||
|
import type { ToolCallSyntax } from "@oh-my-pi/pi-ai/grammar";
|
||||||
|
import type { HarmonyAuditEvent } from "@oh-my-pi/pi-ai/utils/harmony-leak";
|
||||||
import { getBundledModel } from "@oh-my-pi/pi-catalog/models";
|
import { getBundledModel } from "@oh-my-pi/pi-catalog/models";
|
||||||
import { logger } from "@oh-my-pi/pi-utils";
|
import { logger } from "@oh-my-pi/pi-utils";
|
||||||
import { abortReasonText, agentLoop, agentLoopContinue } from "./agent-loop";
|
import { abortReasonText, agentLoop, agentLoopContinue } from "./agent-loop";
|
||||||
import type { AppendOnlyContextManager } from "./append-only-context";
|
import type { AppendOnlyContextManager } from "./append-only-context";
|
||||||
import type { HarmonyAuditEvent } from "./harmony-leak";
|
|
||||||
import type {
|
import type {
|
||||||
AgentContext,
|
AgentContext,
|
||||||
AgentEvent,
|
AgentEvent,
|
||||||
@@ -220,6 +221,8 @@ export interface AgentOptions {
|
|||||||
|
|
||||||
/** Enable intent tracing schema injection/stripping in the harness. */
|
/** Enable intent tracing schema injection/stripping in the harness. */
|
||||||
intentTracing?: boolean;
|
intentTracing?: boolean;
|
||||||
|
/** Owned tool-calling syntax. Undefined keeps provider-native tool calling. */
|
||||||
|
toolCallSyntax?: ToolCallSyntax;
|
||||||
/** Dynamic tool choice override, resolved per LLM call. */
|
/** Dynamic tool choice override, resolved per LLM call. */
|
||||||
getToolChoice?: () => ToolChoice | undefined;
|
getToolChoice?: () => ToolChoice | undefined;
|
||||||
|
|
||||||
@@ -316,6 +319,7 @@ export class Agent {
|
|||||||
#preferWebsockets?: boolean;
|
#preferWebsockets?: boolean;
|
||||||
#transformToolCallArguments?: (args: Record<string, unknown>, toolName: string) => Record<string, unknown>;
|
#transformToolCallArguments?: (args: Record<string, unknown>, toolName: string) => Record<string, unknown>;
|
||||||
#intentTracing: boolean;
|
#intentTracing: boolean;
|
||||||
|
#toolCallSyntax?: ToolCallSyntax;
|
||||||
#getToolChoice?: () => ToolChoice | undefined;
|
#getToolChoice?: () => ToolChoice | undefined;
|
||||||
#onPayload?: SimpleStreamOptions["onPayload"];
|
#onPayload?: SimpleStreamOptions["onPayload"];
|
||||||
#onResponse?: SimpleStreamOptions["onResponse"];
|
#onResponse?: SimpleStreamOptions["onResponse"];
|
||||||
@@ -378,6 +382,7 @@ export class Agent {
|
|||||||
this.#preferWebsockets = opts.preferWebsockets;
|
this.#preferWebsockets = opts.preferWebsockets;
|
||||||
this.#transformToolCallArguments = opts.transformToolCallArguments;
|
this.#transformToolCallArguments = opts.transformToolCallArguments;
|
||||||
this.#intentTracing = opts.intentTracing === true;
|
this.#intentTracing = opts.intentTracing === true;
|
||||||
|
this.#toolCallSyntax = opts.toolCallSyntax;
|
||||||
this.#getToolChoice = opts.getToolChoice;
|
this.#getToolChoice = opts.getToolChoice;
|
||||||
this.#onAssistantMessageEvent = opts.onAssistantMessageEvent;
|
this.#onAssistantMessageEvent = opts.onAssistantMessageEvent;
|
||||||
this.#onHarmonyLeak = opts.onHarmonyLeak;
|
this.#onHarmonyLeak = opts.onHarmonyLeak;
|
||||||
@@ -1023,6 +1028,7 @@ export class Agent {
|
|||||||
cursorOnToolResult,
|
cursorOnToolResult,
|
||||||
transformToolCallArguments: this.#transformToolCallArguments,
|
transformToolCallArguments: this.#transformToolCallArguments,
|
||||||
intentTracing: this.#intentTracing,
|
intentTracing: this.#intentTracing,
|
||||||
|
toolCallSyntax: this.#toolCallSyntax,
|
||||||
appendOnlyContext: this.#appendOnlyContext,
|
appendOnlyContext: this.#appendOnlyContext,
|
||||||
beforeToolCall: this.beforeToolCall ? (ctx, signal) => this.beforeToolCall?.(ctx, signal) : undefined,
|
beforeToolCall: this.beforeToolCall ? (ctx, signal) => this.beforeToolCall?.(ctx, signal) : undefined,
|
||||||
afterToolCall: this.afterToolCall ? (ctx, signal) => this.afterToolCall?.(ctx, signal) : undefined,
|
afterToolCall: this.afterToolCall ? (ctx, signal) => this.afterToolCall?.(ctx, signal) : undefined,
|
||||||
|
|||||||
@@ -6,7 +6,6 @@ export * from "./agent-loop";
|
|||||||
export * from "./append-only-context";
|
export * from "./append-only-context";
|
||||||
// Compaction
|
// Compaction
|
||||||
export * from "./compaction";
|
export * from "./compaction";
|
||||||
export * from "./harmony-leak";
|
|
||||||
// Proxy utilities
|
// Proxy utilities
|
||||||
export * from "./proxy";
|
export * from "./proxy";
|
||||||
// Run-level telemetry collector + aggregators
|
// Run-level telemetry collector + aggregators
|
||||||
|
|||||||
@@ -17,8 +17,9 @@ import type {
|
|||||||
ToolResultMessage,
|
ToolResultMessage,
|
||||||
TSchema,
|
TSchema,
|
||||||
} from "@oh-my-pi/pi-ai";
|
} from "@oh-my-pi/pi-ai";
|
||||||
|
import type { ToolCallSyntax } from "@oh-my-pi/pi-ai/grammar";
|
||||||
|
import type { HarmonyAuditEvent } from "@oh-my-pi/pi-ai/utils/harmony-leak";
|
||||||
import type { AppendOnlyContextManager } from "./append-only-context";
|
import type { AppendOnlyContextManager } from "./append-only-context";
|
||||||
import type { HarmonyAuditEvent } from "./harmony-leak";
|
|
||||||
import type { AgentRunCoverage, AgentRunSummary } from "./run-collector";
|
import type { AgentRunCoverage, AgentRunSummary } from "./run-collector";
|
||||||
import type { AgentTelemetryConfig } from "./telemetry";
|
import type { AgentTelemetryConfig } from "./telemetry";
|
||||||
|
|
||||||
@@ -199,6 +200,15 @@ export interface AgentLoopConfig extends SimpleStreamOptions {
|
|||||||
* then strips from arguments before executing tools.
|
* then strips from arguments before executing tools.
|
||||||
*/
|
*/
|
||||||
intentTracing?: boolean;
|
intentTracing?: boolean;
|
||||||
|
/**
|
||||||
|
* Owned tool calling syntax.
|
||||||
|
*
|
||||||
|
* Undefined keeps provider-native tool calling. A syntax value sends no
|
||||||
|
* native `tools`, forces `toolChoice` off, appends that syntax's tool catalog
|
||||||
|
* instructions, re-encodes prior tool calls/results as text, and parses the
|
||||||
|
* model's text output back into canonical `toolCall` blocks.
|
||||||
|
*/
|
||||||
|
toolCallSyntax?: ToolCallSyntax;
|
||||||
/**
|
/**
|
||||||
* Append-only context mode — stabilizes system prompt + tool spec bytes
|
* Append-only context mode — stabilizes system prompt + tool spec bytes
|
||||||
* across turns so provider prefix caches hit at maximum rate.
|
* across turns so provider prefix caches hit at maximum rate.
|
||||||
|
|||||||
@@ -0,0 +1,148 @@
|
|||||||
|
import { describe, expect, it } from "bun:test";
|
||||||
|
import { agentLoop } from "@oh-my-pi/pi-agent-core/agent-loop";
|
||||||
|
import type { AgentContext, AgentLoopConfig, AgentMessage, AgentTool } from "@oh-my-pi/pi-agent-core/types";
|
||||||
|
import type { AssistantMessage, Context, Message, TextContent, ToolResultMessage } from "@oh-my-pi/pi-ai";
|
||||||
|
import { createMockModel } from "@oh-my-pi/pi-ai/providers/mock";
|
||||||
|
import { z } from "zod/v4";
|
||||||
|
import { createUserMessage } from "./helpers";
|
||||||
|
|
||||||
|
function identityConverter(messages: AgentMessage[]): Message[] {
|
||||||
|
return messages.filter(m => m.role === "user" || m.role === "assistant" || m.role === "toolResult") as Message[];
|
||||||
|
}
|
||||||
|
|
||||||
|
function wireText(message: Message): string {
|
||||||
|
if (typeof message.content === "string") return message.content;
|
||||||
|
return (message.content as (TextContent | { type: string })[])
|
||||||
|
.map(b => (b.type === "text" ? (b as TextContent).text : ""))
|
||||||
|
.join("");
|
||||||
|
}
|
||||||
|
|
||||||
|
describe("agentLoop with owned in-band tool calls", () => {
|
||||||
|
it("executes <tool_call> text, strips native tools from the wire, and re-encodes history as text", async () => {
|
||||||
|
const echoArgs: Array<{ msg: string }> = [];
|
||||||
|
const toolSchema = z.object({ msg: z.string().describe("message to echo") });
|
||||||
|
const echoTool: AgentTool<typeof toolSchema, { msg: string }> = {
|
||||||
|
name: "echo",
|
||||||
|
label: "Echo",
|
||||||
|
description: "Echo a message back",
|
||||||
|
parameters: toolSchema,
|
||||||
|
async execute(_toolCallId, params) {
|
||||||
|
echoArgs.push(params);
|
||||||
|
return { content: [{ type: "text", text: `echoed:${params.msg}` }], details: params };
|
||||||
|
},
|
||||||
|
};
|
||||||
|
|
||||||
|
const captured: Context[] = [];
|
||||||
|
const mock = createMockModel({
|
||||||
|
responses: [
|
||||||
|
context => {
|
||||||
|
captured.push(context);
|
||||||
|
return {
|
||||||
|
content: [
|
||||||
|
"on it\n<tool_call>echo\n<arg_key>msg</arg_key>\n<arg_value>hello world</arg_value>\n</tool_call>",
|
||||||
|
],
|
||||||
|
};
|
||||||
|
},
|
||||||
|
context => {
|
||||||
|
captured.push(context);
|
||||||
|
return { content: ["all done"] };
|
||||||
|
},
|
||||||
|
],
|
||||||
|
});
|
||||||
|
|
||||||
|
const context: AgentContext = { systemPrompt: ["BASE PROMPT"], messages: [], tools: [echoTool] };
|
||||||
|
const config: AgentLoopConfig = { model: mock.model, convertToLlm: identityConverter, toolCallSyntax: "glm" };
|
||||||
|
|
||||||
|
const messages = await agentLoop([createUserMessage("say hi")], context, config, undefined, mock.stream).result();
|
||||||
|
|
||||||
|
// The tool was actually executed with the parsed (verbatim) argument.
|
||||||
|
expect(echoArgs).toEqual([{ msg: "hello world" }]);
|
||||||
|
expect(captured).toHaveLength(2);
|
||||||
|
|
||||||
|
// First request: no native tools on the wire; catalog + grammar injected.
|
||||||
|
expect(captured[0].tools).toBeUndefined();
|
||||||
|
const sys0 = captured[0].systemPrompt ?? [];
|
||||||
|
expect(sys0[0]).toBe("BASE PROMPT");
|
||||||
|
const promptSection = sys0.join("\n");
|
||||||
|
expect(promptSection).toContain("<tools>");
|
||||||
|
expect(promptSection).toContain('"name":"echo"');
|
||||||
|
expect(promptSection).toContain("YOU MUST EMIT THE STOP SEQUENCE AND HALT");
|
||||||
|
|
||||||
|
// Second request: the wire carries NO native tool blocks — prior call/result
|
||||||
|
// are plain <tool_call> / <tool_response> text, and tools are still stripped.
|
||||||
|
const wire2 = captured[1].messages;
|
||||||
|
expect(captured[1].tools).toBeUndefined();
|
||||||
|
for (const m of wire2) {
|
||||||
|
expect(m.role).not.toBe("toolResult");
|
||||||
|
if (m.role === "assistant") {
|
||||||
|
expect((m.content as { type: string }[]).some(b => b.type === "toolCall")).toBe(false);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const wireAssistant = wire2.find(m => m.role === "assistant");
|
||||||
|
expect(wireAssistant).toBeDefined();
|
||||||
|
const at = wireText(wireAssistant!);
|
||||||
|
expect(at).toContain("on it");
|
||||||
|
expect(at).toContain("<tool_call>echo");
|
||||||
|
expect(at).toContain("<arg_value>hello world</arg_value>");
|
||||||
|
const resultsText = wire2
|
||||||
|
.filter(m => m.role === "user")
|
||||||
|
.map(wireText)
|
||||||
|
.join("\n");
|
||||||
|
expect(resultsText).toContain("<tool_response>");
|
||||||
|
expect(resultsText).toContain("echoed:hello world");
|
||||||
|
|
||||||
|
// The internal store stays canonical: native toolCall block + toolResult message.
|
||||||
|
const internalAssistant = messages.find(
|
||||||
|
(m): m is AssistantMessage => m.role === "assistant" && m.content.some(b => b.type === "toolCall"),
|
||||||
|
);
|
||||||
|
expect(internalAssistant).toBeDefined();
|
||||||
|
const internalResult = messages.find((m): m is ToolResultMessage => m.role === "toolResult");
|
||||||
|
expect(internalResult).toBeDefined();
|
||||||
|
expect(internalResult!.toolName).toBe("echo");
|
||||||
|
expect(wireText(internalResult!)).toBe("echoed:hello world");
|
||||||
|
});
|
||||||
|
|
||||||
|
it("executes Hermes/Qwen JSON tool calls when that syntax is selected", async () => {
|
||||||
|
const echoArgs: Array<{ msg: string }> = [];
|
||||||
|
const toolSchema = z.object({ msg: z.string().describe("message to echo") });
|
||||||
|
const echoTool: AgentTool<typeof toolSchema, { msg: string }> = {
|
||||||
|
name: "echo",
|
||||||
|
label: "Echo",
|
||||||
|
description: "Echo a message back",
|
||||||
|
parameters: toolSchema,
|
||||||
|
async execute(_toolCallId, params) {
|
||||||
|
echoArgs.push(params);
|
||||||
|
return { content: [{ type: "text", text: `echoed:${params.msg}` }], details: params };
|
||||||
|
},
|
||||||
|
};
|
||||||
|
|
||||||
|
const captured: Context[] = [];
|
||||||
|
const mock = createMockModel({
|
||||||
|
responses: [
|
||||||
|
context => {
|
||||||
|
captured.push(context);
|
||||||
|
return { content: ['<tool_call>\n{"name":"echo","arguments":{"msg":"hi"}}\n</tool_call>'] };
|
||||||
|
},
|
||||||
|
context => {
|
||||||
|
captured.push(context);
|
||||||
|
return { content: ["done"] };
|
||||||
|
},
|
||||||
|
],
|
||||||
|
});
|
||||||
|
|
||||||
|
const context: AgentContext = { systemPrompt: ["BASE PROMPT"], messages: [], tools: [echoTool] };
|
||||||
|
const config: AgentLoopConfig = { model: mock.model, convertToLlm: identityConverter, toolCallSyntax: "hermes" };
|
||||||
|
|
||||||
|
await agentLoop([createUserMessage("say hi")], context, config, undefined, mock.stream).result();
|
||||||
|
|
||||||
|
expect(echoArgs).toEqual([{ msg: "hi" }]);
|
||||||
|
expect(captured[0].tools).toBeUndefined();
|
||||||
|
expect((captured[0].systemPrompt ?? []).join("\n")).toContain('"name":"function_name","arguments"');
|
||||||
|
const resultsText = captured[1].messages
|
||||||
|
.filter(m => m.role === "user")
|
||||||
|
.map(wireText)
|
||||||
|
.join("\n");
|
||||||
|
expect(resultsText).toContain("<tool_response>");
|
||||||
|
expect(resultsText).toContain("echoed:hi");
|
||||||
|
});
|
||||||
|
});
|
||||||
@@ -1,6 +1,22 @@
|
|||||||
# Changelog
|
# Changelog
|
||||||
|
|
||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
### Added
|
||||||
|
|
||||||
|
- Added `@oh-my-pi/pi-ai/utils/harmony-leak` export with helpers to detect, audit, and recover GPT-5 Harmony tool-call header leaks
|
||||||
|
- Added the `@oh-my-pi/pi-ai/grammar` public entrypoint for grammar factories, prompt/call rendering, in-band scanning, history encoding, and related typed utilities
|
||||||
|
- Added a unified in-band tool-call grammar engine with syntax-owned scanners, prompts, history rendering, tool-result rendering, and stream adaptation for GLM, Hermes/Qwen, Kimi, XML/Anthropic, DeepSeek, Harmony, and pi-native formats.
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
|
||||||
|
- Added raw in-band tool-call block capture to parsed owned tool calls so debugging can inspect the exact model-emitted call syntax.
|
||||||
|
- Made tool-call argument validation more lenient for schema-directed scalar coercions, including object/array stringification and 0/1 boolean coercion.
|
||||||
|
|
||||||
|
### Fixed
|
||||||
|
|
||||||
|
- Fixed Harmony leak handling support by adding `recoverHarmonyToolCall` plus leak-detection workflows for contaminated assistant messages so recoverable tool-call arguments can be safely truncated and retried
|
||||||
|
- Fixed false-positive gating in Harmony leak heuristics using signal-based checks so unrelated text containing `to=functions...` is not treated as leaked tool-call markup
|
||||||
|
- Routed Kimi, DeepSeek DSML, and plain thinking markup healing through the shared in-band scanners so provider leak repair and owned tool calling parse the same wire formats.
|
||||||
|
|
||||||
## [15.13.1] - 2026-06-15
|
## [15.13.1] - 2026-06-15
|
||||||
|
|
||||||
@@ -3656,4 +3672,4 @@ _Dedicated to Peter's shoulder ([@steipete](https://twitter.com/steipete))_
|
|||||||
|
|
||||||
## [0.9.4] - 2025-11-26
|
## [0.9.4] - 2025-11-26
|
||||||
|
|
||||||
Initial release with multi-provider LLM support.
|
Initial release with multi-provider LLM support.
|
||||||
@@ -91,6 +91,14 @@
|
|||||||
"types": "./src/usage/*.ts",
|
"types": "./src/usage/*.ts",
|
||||||
"import": "./src/usage/*.ts"
|
"import": "./src/usage/*.ts"
|
||||||
},
|
},
|
||||||
|
"./utils/harmony-leak": {
|
||||||
|
"types": "./src/utils/harmony-leak.ts",
|
||||||
|
"import": "./src/utils/harmony-leak.ts"
|
||||||
|
},
|
||||||
|
"./grammar": {
|
||||||
|
"types": "./src/grammar/index.ts",
|
||||||
|
"import": "./src/grammar/index.ts"
|
||||||
|
},
|
||||||
"./utils/*": {
|
"./utils/*": {
|
||||||
"types": "./src/utils/*.ts",
|
"types": "./src/utils/*.ts",
|
||||||
"import": "./src/utils/*.ts"
|
"import": "./src/utils/*.ts"
|
||||||
|
|||||||
@@ -0,0 +1,31 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
A call is a `<function_calls>` block wrapping one or more `<invoke>` blocks, each holding `<parameter>` children:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<function_calls>
|
||||||
|
<invoke name="tool_name"><parameter name="arg_name">arg value</parameter></invoke>
|
||||||
|
</function_calls>
|
||||||
|
```
|
||||||
|
|
||||||
|
Results arrive later in a `<function_results>` block, one `<result>` per call (failures use `<error>` with `<stderr>` in place of `<result>` with `<stdout>`):
|
||||||
|
|
||||||
|
```text
|
||||||
|
<function_results>
|
||||||
|
<result>
|
||||||
|
<tool_name>tool_name</tool_name>
|
||||||
|
<stdout>verbatim tool result</stdout>
|
||||||
|
</result>
|
||||||
|
</function_results>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- `name` MUST match a listed function.
|
||||||
|
- String/scalar parameters: exact text, spaces preserved. Lists/objects: JSON.
|
||||||
|
- Multiple calls: multiple `<invoke>` blocks in one `<function_calls>`.
|
||||||
|
- You MAY write visible text before the calls.
|
||||||
|
- NEVER emit `tool_calls` JSON.
|
||||||
|
- NEVER use the legacy `<tool_name>`/`<parameters>` call syntax.
|
||||||
|
- Read each `<result>`/`<error>` in call order. NEVER emit `<function_results>` yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,520 @@
|
|||||||
|
import { parseJsonWithRepair } from "../utils/json-parse";
|
||||||
|
import grammarPrompt from "./anthropic.md" with { type: "text" };
|
||||||
|
import { buildStringArgsResolver, mintToolCallId } from "./coercion";
|
||||||
|
import { renderAnthropicToolCalls, renderAnthropicToolResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner, InbandScannerOptions } from "./types";
|
||||||
|
|
||||||
|
const MAX_PARTIAL_TAG_LENGTH = 256;
|
||||||
|
const MAX_PARAMETER_VALUE_LENGTH = 1_000_000;
|
||||||
|
|
||||||
|
const WRAPPER_TAGS: Record<string, true> = { function_calls: true, tool_calls: true };
|
||||||
|
const THINKING_TAGS: Record<string, true> = { thinking: true, think: true, scratchpad: true };
|
||||||
|
const BASE_TAG_PREFIXES = [
|
||||||
|
"<function_calls",
|
||||||
|
"</function_calls",
|
||||||
|
"<tool_calls",
|
||||||
|
"</tool_calls",
|
||||||
|
"<invoke",
|
||||||
|
"</invoke",
|
||||||
|
"<parameter",
|
||||||
|
"</parameter",
|
||||||
|
"<antml:function_calls",
|
||||||
|
"</antml:function_calls",
|
||||||
|
"<antml:tool_calls",
|
||||||
|
"</antml:tool_calls",
|
||||||
|
"<antml:invoke",
|
||||||
|
"</antml:invoke",
|
||||||
|
"<antml:parameter",
|
||||||
|
"</antml:parameter",
|
||||||
|
] as const;
|
||||||
|
const THINKING_TAG_PREFIXES = [
|
||||||
|
"<thinking",
|
||||||
|
"</thinking",
|
||||||
|
"<think",
|
||||||
|
"</think",
|
||||||
|
"<scratchpad",
|
||||||
|
"</scratchpad",
|
||||||
|
"<antml:thinking",
|
||||||
|
"</antml:thinking",
|
||||||
|
"<antml:think",
|
||||||
|
"</antml:think",
|
||||||
|
"<antml:scratchpad",
|
||||||
|
"</antml:scratchpad",
|
||||||
|
] as const;
|
||||||
|
|
||||||
|
type ScannerState = "outside" | "section" | "invoke" | "parameter" | "thinking";
|
||||||
|
type ReturnState = "outside" | "section";
|
||||||
|
|
||||||
|
interface ParsedTag {
|
||||||
|
readonly raw: string;
|
||||||
|
readonly localName: string;
|
||||||
|
readonly prefix: string;
|
||||||
|
readonly closing: boolean;
|
||||||
|
readonly selfClosing: boolean;
|
||||||
|
readonly attrs: ReadonlyMap<string, string>;
|
||||||
|
}
|
||||||
|
|
||||||
|
type TagRead = ParsedTag | "partial" | undefined;
|
||||||
|
|
||||||
|
export class AnthropicInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#state: ScannerState = "outside";
|
||||||
|
#returnState: ReturnState = "outside";
|
||||||
|
#afterThinkingState: ReturnState = "outside";
|
||||||
|
#id = "";
|
||||||
|
#name = "";
|
||||||
|
#args: Record<string, unknown> = {};
|
||||||
|
#started = false;
|
||||||
|
#paramName = "";
|
||||||
|
#paramValue = "";
|
||||||
|
#paramString: boolean | undefined;
|
||||||
|
#paramTruncated = false;
|
||||||
|
#paramClosePrefixes: readonly string[] = [];
|
||||||
|
#rawBlock = "";
|
||||||
|
#thinking = "";
|
||||||
|
#thinkingTag = "";
|
||||||
|
#thinkingClosePrefixes: readonly string[] = [];
|
||||||
|
readonly #stringArgs: (toolName: string) => ReadonlySet<string>;
|
||||||
|
readonly #parseThinking: boolean;
|
||||||
|
|
||||||
|
constructor(options: InbandScannerOptions = {}) {
|
||||||
|
this.#stringArgs = options.stringArgs ?? buildStringArgsResolver(options.tools);
|
||||||
|
this.#parseThinking = options.parseThinking === true;
|
||||||
|
}
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
let progressed: boolean;
|
||||||
|
switch (this.#state) {
|
||||||
|
case "outside":
|
||||||
|
progressed = this.#consumeOutside(final, events);
|
||||||
|
break;
|
||||||
|
case "section":
|
||||||
|
progressed = this.#consumeSection(final, events);
|
||||||
|
break;
|
||||||
|
case "invoke":
|
||||||
|
progressed = this.#consumeInvoke(final, events);
|
||||||
|
break;
|
||||||
|
case "parameter":
|
||||||
|
progressed = this.#consumeParameter(final);
|
||||||
|
break;
|
||||||
|
case "thinking":
|
||||||
|
progressed = this.#consumeThinking(final, events);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
if (!progressed) break;
|
||||||
|
}
|
||||||
|
if (final) this.#flushFinal(events);
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeOutside(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const tagStart = this.#buffer.indexOf("<");
|
||||||
|
if (tagStart === -1) {
|
||||||
|
this.#emitText(this.#buffer, events);
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
if (tagStart > 0) {
|
||||||
|
this.#emitText(this.#buffer.slice(0, tagStart), events);
|
||||||
|
this.#buffer = this.#buffer.slice(tagStart);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = this.#peekTag(final, this.#relevantPrefixes());
|
||||||
|
if (tag === "partial") return false;
|
||||||
|
if (!tag) {
|
||||||
|
this.#emitText(this.#buffer[0]!, events);
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!tag.closing && WRAPPER_TAGS[tag.localName] === true) {
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
this.#state = "section";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!tag.closing && tag.localName === "invoke") {
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
this.#startInvoke(tag, "outside", events);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (this.#isThinkingOpen(tag)) {
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
this.#startThinking(tag, "outside", events);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (tag.closing && WRAPPER_TAGS[tag.localName] === true) {
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#emitText(this.#buffer[0]!, events);
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeSection(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const tagStart = this.#buffer.indexOf("<");
|
||||||
|
if (tagStart === -1) {
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
if (tagStart > 0) {
|
||||||
|
this.#buffer = this.#buffer.slice(tagStart);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = this.#peekTag(final, this.#relevantPrefixes());
|
||||||
|
if (tag === "partial") return false;
|
||||||
|
if (!tag) {
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
if (tag.closing && WRAPPER_TAGS[tag.localName] === true) {
|
||||||
|
this.#state = "outside";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!tag.closing && tag.localName === "invoke") {
|
||||||
|
this.#startInvoke(tag, "section", events);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (this.#parseThinking && !tag.closing && THINKING_TAGS[tag.localName] === true) {
|
||||||
|
this.#startThinking(tag, "section", events);
|
||||||
|
}
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeInvoke(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const tagStart = this.#buffer.indexOf("<");
|
||||||
|
if (tagStart === -1) {
|
||||||
|
if (final) this.#resetCall(this.#returnState);
|
||||||
|
else {
|
||||||
|
this.#rawBlock += this.#buffer;
|
||||||
|
this.#buffer = "";
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
if (tagStart > 0) {
|
||||||
|
const consumed = this.#buffer.slice(0, tagStart);
|
||||||
|
this.#rawBlock += consumed;
|
||||||
|
this.#buffer = this.#buffer.slice(tagStart);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = this.#peekTag(final, this.#relevantPrefixes());
|
||||||
|
if (tag === "partial") return false;
|
||||||
|
if (!tag) {
|
||||||
|
const consumed = this.#buffer[0]!;
|
||||||
|
this.#rawBlock += consumed;
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#rawBlock += tag.raw;
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
if (tag.closing && tag.localName === "invoke") {
|
||||||
|
if (this.#started) {
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: this.#name,
|
||||||
|
arguments: this.#args,
|
||||||
|
rawBlock: this.#rawBlock,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
this.#resetCall(this.#returnState);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!tag.closing && tag.localName === "parameter") {
|
||||||
|
this.#startParameter(tag);
|
||||||
|
if (tag.selfClosing) this.#finishParameter();
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeParameter(final: boolean): boolean {
|
||||||
|
const tagStart = this.#buffer.indexOf("<");
|
||||||
|
if (tagStart === -1) {
|
||||||
|
if (final) {
|
||||||
|
this.#resetCall(this.#returnState);
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
this.#appendParameterValue(this.#buffer);
|
||||||
|
this.#rawBlock += this.#buffer;
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
if (tagStart > 0) {
|
||||||
|
const consumed = this.#buffer.slice(0, tagStart);
|
||||||
|
this.#appendParameterValue(consumed);
|
||||||
|
this.#rawBlock += consumed;
|
||||||
|
this.#buffer = this.#buffer.slice(tagStart);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = this.#peekTag(final, this.#paramClosePrefixes);
|
||||||
|
if (tag === "partial") return false;
|
||||||
|
if (tag?.closing && tag.localName === "parameter") {
|
||||||
|
this.#rawBlock += tag.raw;
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
this.#finishParameter();
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (final && !tag) {
|
||||||
|
this.#resetCall(this.#returnState);
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
const consumed = this.#buffer[0]!;
|
||||||
|
this.#appendParameterValue(consumed);
|
||||||
|
this.#rawBlock += consumed;
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeThinking(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const tagStart = this.#buffer.indexOf("<");
|
||||||
|
if (tagStart === -1) {
|
||||||
|
if (final) {
|
||||||
|
this.#appendThinking(this.#buffer, events);
|
||||||
|
this.#buffer = "";
|
||||||
|
this.#finishThinking(events);
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
this.#appendThinking(this.#buffer, events);
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
if (tagStart > 0) {
|
||||||
|
this.#appendThinking(this.#buffer.slice(0, tagStart), events);
|
||||||
|
this.#buffer = this.#buffer.slice(tagStart);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = this.#peekTag(final, this.#thinkingClosePrefixes);
|
||||||
|
if (tag === "partial") return false;
|
||||||
|
if (tag?.closing && tag.localName === this.#thinkingTag) {
|
||||||
|
this.#buffer = this.#buffer.slice(tag.raw.length);
|
||||||
|
this.#finishThinking(events);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (final && !tag) {
|
||||||
|
this.#appendThinking(this.#buffer, events);
|
||||||
|
this.#buffer = "";
|
||||||
|
this.#finishThinking(events);
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
this.#appendThinking(this.#buffer[0]!, events);
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#flushFinal(events: InbandScanEvent[]): void {
|
||||||
|
if (this.#state === "outside") return;
|
||||||
|
if (this.#state === "thinking") this.#finishThinking(events);
|
||||||
|
else this.#resetCall(this.#returnState);
|
||||||
|
this.#state = "outside";
|
||||||
|
this.#buffer = "";
|
||||||
|
}
|
||||||
|
|
||||||
|
#startInvoke(tag: ParsedTag, returnState: ReturnState, events: InbandScanEvent[]): void {
|
||||||
|
this.#returnState = returnState;
|
||||||
|
this.#id = mintToolCallId();
|
||||||
|
this.#name = tag.attrs.get("name")?.trim() ?? "";
|
||||||
|
this.#args = {};
|
||||||
|
this.#rawBlock = tag.raw;
|
||||||
|
this.#started = this.#name.length > 0;
|
||||||
|
this.#state = "invoke";
|
||||||
|
if (this.#started) events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
}
|
||||||
|
|
||||||
|
#startParameter(tag: ParsedTag): void {
|
||||||
|
this.#paramName = tag.attrs.get("name")?.trim() ?? "";
|
||||||
|
this.#paramValue = "";
|
||||||
|
this.#paramTruncated = false;
|
||||||
|
this.#paramString = parseStringAttribute(tag.attrs.get("string"));
|
||||||
|
this.#paramClosePrefixes = closePrefixes("parameter", tag.prefix);
|
||||||
|
this.#state = "parameter";
|
||||||
|
}
|
||||||
|
|
||||||
|
#appendParameterValue(delta: string): void {
|
||||||
|
if (delta.length === 0) return;
|
||||||
|
const remaining = MAX_PARAMETER_VALUE_LENGTH - this.#paramValue.length;
|
||||||
|
if (remaining > 0) this.#paramValue += delta.slice(0, remaining);
|
||||||
|
if (delta.length > remaining) this.#paramTruncated = true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#finishParameter(): void {
|
||||||
|
if (this.#paramName.length > 0) {
|
||||||
|
const value = this.#paramTruncated
|
||||||
|
? `${this.#paramValue}\n…[parameter truncated: exceeded ${MAX_PARAMETER_VALUE_LENGTH} bytes]`
|
||||||
|
: this.#paramValue;
|
||||||
|
this.#args[this.#paramName] = this.#coerceParameterValue(this.#paramName, value, this.#paramString);
|
||||||
|
}
|
||||||
|
this.#paramName = "";
|
||||||
|
this.#paramValue = "";
|
||||||
|
this.#paramString = undefined;
|
||||||
|
this.#paramTruncated = false;
|
||||||
|
this.#paramClosePrefixes = [];
|
||||||
|
this.#state = "invoke";
|
||||||
|
}
|
||||||
|
|
||||||
|
#coerceParameterValue(name: string, raw: string, explicitString: boolean | undefined): unknown {
|
||||||
|
if (explicitString ?? this.#stringArgs(this.#name).has(name)) return raw;
|
||||||
|
const trimmed = raw.trim();
|
||||||
|
if (trimmed.length === 0) return raw;
|
||||||
|
try {
|
||||||
|
return parseJsonWithRepair<unknown>(trimmed);
|
||||||
|
} catch {
|
||||||
|
return raw;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#startThinking(tag: ParsedTag, afterState: ReturnState, events: InbandScanEvent[]): void {
|
||||||
|
this.#afterThinkingState = afterState;
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#thinkingTag = tag.localName;
|
||||||
|
this.#thinkingClosePrefixes = closePrefixes(tag.localName, tag.prefix);
|
||||||
|
this.#state = "thinking";
|
||||||
|
events.push({ type: "thinkingStart" });
|
||||||
|
if (tag.selfClosing) this.#finishThinking(events);
|
||||||
|
}
|
||||||
|
|
||||||
|
#appendThinking(delta: string, events: InbandScanEvent[]): void {
|
||||||
|
if (delta.length === 0) return;
|
||||||
|
this.#thinking += delta;
|
||||||
|
events.push({ type: "thinkingDelta", delta });
|
||||||
|
}
|
||||||
|
|
||||||
|
#finishThinking(events: InbandScanEvent[]): void {
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#thinkingTag = "";
|
||||||
|
this.#thinkingClosePrefixes = [];
|
||||||
|
this.#state = this.#afterThinkingState;
|
||||||
|
this.#afterThinkingState = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#resetCall(nextState: ReturnState): void {
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#args = {};
|
||||||
|
this.#started = false;
|
||||||
|
this.#paramName = "";
|
||||||
|
this.#paramValue = "";
|
||||||
|
this.#paramString = undefined;
|
||||||
|
this.#paramTruncated = false;
|
||||||
|
this.#paramClosePrefixes = [];
|
||||||
|
this.#rawBlock = "";
|
||||||
|
this.#state = nextState;
|
||||||
|
}
|
||||||
|
|
||||||
|
#peekTag(final: boolean, relevantPrefixes: readonly string[]): TagRead {
|
||||||
|
const close = this.#buffer.indexOf(">");
|
||||||
|
if (close === -1) {
|
||||||
|
if (
|
||||||
|
!final &&
|
||||||
|
this.#buffer.length <= MAX_PARTIAL_TAG_LENGTH &&
|
||||||
|
couldBeTagPrefix(this.#buffer, relevantPrefixes)
|
||||||
|
) {
|
||||||
|
return "partial";
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
const raw = this.#buffer.slice(0, close + 1);
|
||||||
|
return parseTag(raw);
|
||||||
|
}
|
||||||
|
|
||||||
|
#isThinkingOpen(tag: ParsedTag): boolean {
|
||||||
|
if (!this.#parseThinking || tag.closing) return false;
|
||||||
|
return THINKING_TAGS[tag.localName] === true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#relevantPrefixes(): readonly string[] {
|
||||||
|
return this.#parseThinking ? ALL_TAG_PREFIXES : BASE_TAG_PREFIXES;
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitText(text: string, events: InbandScanEvent[]): void {
|
||||||
|
if (text.length > 0) events.push({ type: "text", text });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const ALL_TAG_PREFIXES = [...BASE_TAG_PREFIXES, ...THINKING_TAG_PREFIXES] as const;
|
||||||
|
|
||||||
|
function parseTag(raw: string): ParsedTag | undefined {
|
||||||
|
const match = /^<\s*(\/?)\s*(?:(?<prefix>[A-Za-z_][\w.-]*):)?(?<localName>[A-Za-z_][\w.-]*)(?<attrs>[^>]*)>$/s.exec(
|
||||||
|
raw,
|
||||||
|
);
|
||||||
|
const localName = match?.groups?.localName;
|
||||||
|
if (!match || !localName) return undefined;
|
||||||
|
const attrsText = match.groups?.attrs ?? "";
|
||||||
|
return {
|
||||||
|
raw,
|
||||||
|
localName: localName.toLowerCase(),
|
||||||
|
prefix: match.groups?.prefix ?? "",
|
||||||
|
closing: match[1] === "/",
|
||||||
|
selfClosing: match[1] !== "/" && /\/\s*$/.test(attrsText),
|
||||||
|
attrs: parseAttributes(attrsText),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseAttributes(text: string): ReadonlyMap<string, string> {
|
||||||
|
const attrs = new Map<string, string>();
|
||||||
|
const pattern = /([A-Za-z_:][\w:.-]*)\s*=\s*(?:"([^"]*)"|'([^']*)'|([^\s"'<>/=]+))/g;
|
||||||
|
for (const match of text.matchAll(pattern)) {
|
||||||
|
const rawName = match[1];
|
||||||
|
if (!rawName) continue;
|
||||||
|
const colon = rawName.lastIndexOf(":");
|
||||||
|
const name = (colon === -1 ? rawName : rawName.slice(colon + 1)).toLowerCase();
|
||||||
|
attrs.set(name, match[2] ?? match[3] ?? match[4] ?? "");
|
||||||
|
}
|
||||||
|
return attrs;
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseStringAttribute(value: string | undefined): boolean | undefined {
|
||||||
|
if (value === undefined) return undefined;
|
||||||
|
const normalized = value.trim().toLowerCase();
|
||||||
|
if (normalized === "false" || normalized === "0" || normalized === "no") return false;
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
function closePrefixes(localName: string, prefix: string): readonly string[] {
|
||||||
|
const unprefixed = `</${localName}`;
|
||||||
|
const antml = `</antml:${localName}`;
|
||||||
|
if (prefix.length === 0 || prefix === "antml") return [unprefixed, antml];
|
||||||
|
return [`</${prefix}:${localName}`, unprefixed, antml];
|
||||||
|
}
|
||||||
|
|
||||||
|
function couldBeTagPrefix(buffer: string, prefixes: readonly string[]): boolean {
|
||||||
|
if (!buffer.startsWith("<")) return false;
|
||||||
|
for (const prefix of prefixes) {
|
||||||
|
if (prefix.startsWith(buffer) || buffer.startsWith(prefix)) return true;
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "anthropic",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: options => new AnthropicInbandScanner(options),
|
||||||
|
renderAssistantToolCalls: renderAnthropicToolCalls,
|
||||||
|
renderToolResults: renderAnthropicToolResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
import { toolWireSchema } from "../utils/schema";
|
||||||
|
import { getInbandGrammar } from "./factory";
|
||||||
|
import promptTemplate from "./prompt-template.md" with { type: "text" };
|
||||||
|
import type { InbandTool, ToolCallSyntax } from "./types";
|
||||||
|
|
||||||
|
const TOOLS_TOKEN = "{{TOOLS}}";
|
||||||
|
const GRAMMAR_TOKEN = "{{GRAMMAR}}";
|
||||||
|
|
||||||
|
export function renderToolCatalog(tools: readonly InbandTool[]): string {
|
||||||
|
return tools
|
||||||
|
.map(tool =>
|
||||||
|
JSON.stringify({
|
||||||
|
type: "function",
|
||||||
|
function: {
|
||||||
|
name: tool.name,
|
||||||
|
description: tool.description ?? "",
|
||||||
|
parameters: toolWireSchema(tool),
|
||||||
|
},
|
||||||
|
}),
|
||||||
|
)
|
||||||
|
.join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderInbandToolPrompt(tools: readonly InbandTool[], syntax: ToolCallSyntax): string {
|
||||||
|
const prompt = getInbandGrammar(syntax).prompt.trim();
|
||||||
|
return promptTemplate.replace(TOOLS_TOKEN, () => renderToolCatalog(tools)).replace(GRAMMAR_TOKEN, () => prompt);
|
||||||
|
}
|
||||||
@@ -0,0 +1,136 @@
|
|||||||
|
import { toolWireSchema } from "../utils/schema";
|
||||||
|
import type { InbandTool } from "./types";
|
||||||
|
|
||||||
|
export interface ToolArgShape {
|
||||||
|
stringArgs: Set<string>;
|
||||||
|
properties: Record<string, unknown>;
|
||||||
|
parameterOrder: string[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export function buildArgShapes(tools: readonly InbandTool[] = []): Map<string, ToolArgShape> {
|
||||||
|
const shapes = new Map<string, ToolArgShape>();
|
||||||
|
for (const tool of tools) {
|
||||||
|
const schema = resolveToolSchema(tool);
|
||||||
|
const props = schema.properties;
|
||||||
|
const properties =
|
||||||
|
props && typeof props === "object" && !Array.isArray(props) ? (props as Record<string, unknown>) : {};
|
||||||
|
const stringArgs = new Set<string>();
|
||||||
|
const parameterOrder: string[] = [];
|
||||||
|
for (const key in properties) {
|
||||||
|
parameterOrder.push(key);
|
||||||
|
if (isStringOnlySchema(properties[key])) stringArgs.add(key);
|
||||||
|
}
|
||||||
|
shapes.set(tool.name, { stringArgs, properties, parameterOrder });
|
||||||
|
}
|
||||||
|
return shapes;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function buildStringArgsResolver(tools: readonly InbandTool[] = []): (toolName: string) => ReadonlySet<string> {
|
||||||
|
const shapes = buildArgShapes(tools);
|
||||||
|
const empty = new Set<string>();
|
||||||
|
return (toolName: string) => shapes.get(toolName)?.stringArgs ?? empty;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function resolveToolSchema(tool: InbandTool): Record<string, unknown> {
|
||||||
|
try {
|
||||||
|
return toolWireSchema(tool);
|
||||||
|
} catch {
|
||||||
|
const params = tool.parameters;
|
||||||
|
return params && typeof params === "object" && !Array.isArray(params) ? (params as Record<string, unknown>) : {};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export function isStringOnlySchema(schema: unknown): boolean {
|
||||||
|
const types = collectSchemaTypes(schema);
|
||||||
|
types.delete("null");
|
||||||
|
return types.size === 1 && types.has("string");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function collectSchemaTypes(schema: unknown, out: Set<string> = new Set(), depth = 0): Set<string> {
|
||||||
|
if (depth > 8 || !schema || typeof schema !== "object" || Array.isArray(schema)) return out;
|
||||||
|
const node = schema as Record<string, unknown>;
|
||||||
|
const type = node.type;
|
||||||
|
if (typeof type === "string") out.add(type);
|
||||||
|
else if (Array.isArray(type)) for (const t of type) if (typeof t === "string") out.add(t);
|
||||||
|
if (type === undefined && Array.isArray(node.enum)) {
|
||||||
|
for (const value of node.enum) out.add(jsonTypeOf(value));
|
||||||
|
}
|
||||||
|
if (type === undefined && "const" in node) out.add(jsonTypeOf(node.const));
|
||||||
|
for (const key of ["anyOf", "oneOf", "allOf"] as const) {
|
||||||
|
const branch = node[key];
|
||||||
|
if (Array.isArray(branch)) for (const sub of branch) collectSchemaTypes(sub, out, depth + 1);
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function jsonTypeOf(value: unknown): string {
|
||||||
|
const type = typeof value;
|
||||||
|
if (value === null) return "null";
|
||||||
|
if (type === "number" || type === "bigint") return "number";
|
||||||
|
if (type === "boolean") return "boolean";
|
||||||
|
if (type === "string") return "string";
|
||||||
|
return "object";
|
||||||
|
}
|
||||||
|
|
||||||
|
export function decodeValue(raw: string): unknown {
|
||||||
|
const trimmed = raw.trim();
|
||||||
|
if (trimmed.length === 0) return trimmed;
|
||||||
|
try {
|
||||||
|
return JSON.parse(trimmed) as unknown;
|
||||||
|
} catch {
|
||||||
|
return raw;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export function coerceValue(raw: string, schema: unknown): unknown {
|
||||||
|
return isStringOnlySchema(schema) ? raw : decodeValue(raw);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function isArraySchema(schema: unknown): boolean {
|
||||||
|
return collectSchemaTypes(schema).has("array");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function isObjectSchema(schema: unknown): boolean {
|
||||||
|
return collectSchemaTypes(schema).has("object");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getObjectProperties(schema: unknown): Record<string, unknown> {
|
||||||
|
if (!schema || typeof schema !== "object" || Array.isArray(schema)) return {};
|
||||||
|
const props = (schema as Record<string, unknown>).properties;
|
||||||
|
return props && typeof props === "object" && !Array.isArray(props) ? (props as Record<string, unknown>) : {};
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getArrayItemSchema(schema: unknown): unknown {
|
||||||
|
if (!schema || typeof schema !== "object" || Array.isArray(schema)) return undefined;
|
||||||
|
return (schema as Record<string, unknown>).items;
|
||||||
|
}
|
||||||
|
|
||||||
|
let idCounter = 0;
|
||||||
|
export function mintToolCallId(): string {
|
||||||
|
idCounter = (idCounter + 1) % Number.MAX_SAFE_INTEGER;
|
||||||
|
return `ptc_${Date.now().toString(36)}_${idCounter.toString(36)}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function partialSuffixOverlap(text: string, tag: string): number {
|
||||||
|
const max = Math.min(text.length, tag.length - 1);
|
||||||
|
for (let k = max; k > 0; k--) {
|
||||||
|
if (text.endsWith(tag.slice(0, k))) return k;
|
||||||
|
}
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function partialSuffixOverlapAny(text: string, tags: readonly string[]): number {
|
||||||
|
let best = 0;
|
||||||
|
for (const tag of tags) best = Math.max(best, partialSuffixOverlap(text, tag));
|
||||||
|
return best;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function normalizeKimiFunctionName(rawId: string): string {
|
||||||
|
const beforeIndex = rawId.split(":", 1)[0] ?? rawId;
|
||||||
|
const parts = beforeIndex.split(".");
|
||||||
|
return parts[parts.length - 1]?.trim() ?? beforeIndex.trim();
|
||||||
|
}
|
||||||
|
|
||||||
|
export function asRecord(value: unknown): Record<string, unknown> {
|
||||||
|
return value && typeof value === "object" && !Array.isArray(value) ? (value as Record<string, unknown>) : {};
|
||||||
|
}
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
A tool call wraps the function name, a separator, and one JSON object of arguments in fixed tokens. Emit them exactly:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool▁calls▁begin|><|tool▁call▁begin|>tool_name<|tool▁sep|>{"arg":"value"}<|tool▁call▁end|><|tool▁calls▁end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Results arrive as output tokens:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool▁output▁begin|>verbatim tool result<|tool▁output▁end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- Use `|` (U+FF5C) and `▁` (U+2581) exactly.
|
||||||
|
- Tool name MUST match an available function; arguments are one valid JSON object.
|
||||||
|
- NEVER wrap arguments in Markdown fences; NEVER emit a `type` field or `function` prefix.
|
||||||
|
- Multiple calls chain `<|tool▁call▁begin|>...<|tool▁call▁end|>` directly — no separators, spaces, or newlines between them.
|
||||||
|
- Private reasoning, when needed, goes in `<think>...</think>` before the tokens.
|
||||||
|
- Read each output token in call order. NEVER emit output tokens yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,520 @@
|
|||||||
|
import { parseJsonWithRepair } from "../utils/json-parse";
|
||||||
|
import { asRecord, mintToolCallId, partialSuffixOverlapAny } from "./coercion";
|
||||||
|
import grammarPrompt from "./deepseek.md" with { type: "text" };
|
||||||
|
import { renderDeepSeekToolCalls, renderDeepSeekToolResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner, InbandScannerOptions } from "./types";
|
||||||
|
|
||||||
|
export const DEEPSEEK_TOOL_CALLS_BEGIN = "<|tool▁calls▁begin|>";
|
||||||
|
export const DEEPSEEK_TOOL_CALLS_END = "<|tool▁calls▁end|>";
|
||||||
|
export const DEEPSEEK_TOOL_CALL_BEGIN = "<|tool▁call▁begin|>";
|
||||||
|
export const DEEPSEEK_TOOL_CALL_END = "<|tool▁call▁end|>";
|
||||||
|
export const DEEPSEEK_TOOL_SEPARATOR = "<|tool▁sep|>";
|
||||||
|
|
||||||
|
const THINK_OPEN = "<think>";
|
||||||
|
const THINK_CLOSE = "</think>";
|
||||||
|
const LEGACY_TOOL_TYPE = "function";
|
||||||
|
const LEGACY_JSON_FENCE = "```json";
|
||||||
|
const CODE_FENCE = "```";
|
||||||
|
|
||||||
|
const DSML_TOOL_CALLS_OPEN_FULLWIDTH = "<|DSML|tool_calls>";
|
||||||
|
const DSML_TOOL_CALLS_CLOSE_FULLWIDTH = "</|DSML|tool_calls>";
|
||||||
|
const DSML_TOOL_CALLS_OPEN_ASCII = "<|DSML|tool_calls>";
|
||||||
|
const DSML_TOOL_CALLS_CLOSE_ASCII = "</|DSML|tool_calls>";
|
||||||
|
|
||||||
|
const CONTROL_TOKENS = [
|
||||||
|
"<|begin▁of▁sentence|>",
|
||||||
|
"<|end▁of▁sentence|>",
|
||||||
|
"<|▁pad▁|>",
|
||||||
|
"<|User|>",
|
||||||
|
"<|Assistant|>",
|
||||||
|
"<|EOT|>",
|
||||||
|
"<|search▁begin|>",
|
||||||
|
"<|search▁end|>",
|
||||||
|
"<|fim▁hole|>",
|
||||||
|
"<|fim▁begin|>",
|
||||||
|
"<|fim▁end|>",
|
||||||
|
"<|tool▁outputs▁begin|>",
|
||||||
|
"<|tool▁outputs▁end|>",
|
||||||
|
"<|tool▁output▁begin|>",
|
||||||
|
"<|tool▁output▁end|>",
|
||||||
|
] as const;
|
||||||
|
|
||||||
|
const OUTSIDE_TOKENS = [
|
||||||
|
DEEPSEEK_TOOL_CALLS_BEGIN,
|
||||||
|
DEEPSEEK_TOOL_CALLS_END,
|
||||||
|
DEEPSEEK_TOOL_CALL_BEGIN,
|
||||||
|
THINK_OPEN,
|
||||||
|
THINK_CLOSE,
|
||||||
|
DSML_TOOL_CALLS_OPEN_FULLWIDTH,
|
||||||
|
DSML_TOOL_CALLS_OPEN_ASCII,
|
||||||
|
DSML_TOOL_CALLS_CLOSE_FULLWIDTH,
|
||||||
|
DSML_TOOL_CALLS_CLOSE_ASCII,
|
||||||
|
...CONTROL_TOKENS,
|
||||||
|
] as const;
|
||||||
|
|
||||||
|
const SECTION_TOKENS = [DEEPSEEK_TOOL_CALL_BEGIN, DEEPSEEK_TOOL_CALLS_END] as const;
|
||||||
|
const DSML_SECTION_TOKENS = [
|
||||||
|
DSML_TOOL_CALLS_CLOSE_FULLWIDTH,
|
||||||
|
DSML_TOOL_CALLS_CLOSE_ASCII,
|
||||||
|
"<|DSML|invoke",
|
||||||
|
"<|DSML|invoke",
|
||||||
|
] as const;
|
||||||
|
const DSML_INVOKE_TOKENS = ["</|DSML|invoke>", "</|DSML|invoke>", "<|DSML|parameter", "<|DSML|parameter"] as const;
|
||||||
|
const DSML_PARAMETER_CLOSE_TOKENS = ["</|DSML|parameter>", "</|DSML|parameter>"] as const;
|
||||||
|
|
||||||
|
type State =
|
||||||
|
| "outside"
|
||||||
|
| "thinking"
|
||||||
|
| "section"
|
||||||
|
| "header"
|
||||||
|
| "args"
|
||||||
|
| "legacyName"
|
||||||
|
| "legacyArgs"
|
||||||
|
| "dsmlSection"
|
||||||
|
| "dsmlInvoke"
|
||||||
|
| "dsmlParam";
|
||||||
|
|
||||||
|
export class DeepSeekInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#state: State = "outside";
|
||||||
|
#parseThinking: boolean;
|
||||||
|
#inToolSection = false;
|
||||||
|
#id = "";
|
||||||
|
#name = "";
|
||||||
|
#thinking = "";
|
||||||
|
#dsmlArgs: Record<string, unknown> = {};
|
||||||
|
#dsmlParamName = "";
|
||||||
|
#dsmlParamIsString = true;
|
||||||
|
#rawBlock = "";
|
||||||
|
|
||||||
|
constructor(options: InbandScannerOptions = {}) {
|
||||||
|
this.#parseThinking = options.parseThinking ?? true;
|
||||||
|
}
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#state === "outside") {
|
||||||
|
this.#consumeOutside(final, events);
|
||||||
|
if (this.#state !== "outside" && this.#buffer.length > 0) continue;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
if (this.#state === "thinking") {
|
||||||
|
this.#consumeThinking(final, events);
|
||||||
|
if (!final && this.#state === "thinking") break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (this.#state === "section") {
|
||||||
|
if (!this.#consumeSection(final)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (this.#state === "header") {
|
||||||
|
if (!this.#consumeHeader(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (this.#state === "legacyName") {
|
||||||
|
if (!this.#consumeLegacyName(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (this.#state === "args" || this.#state === "legacyArgs") {
|
||||||
|
if (!this.#consumeArgs(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (this.#state === "dsmlSection") {
|
||||||
|
if (!this.#consumeDsmlSection(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (this.#state === "dsmlInvoke") {
|
||||||
|
if (!this.#consumeDsmlInvoke(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (!this.#consumeDsmlParam(final)) break;
|
||||||
|
}
|
||||||
|
if (final && this.#buffer.length === 0 && this.#rawBlock.length > 0) this.#rawBlock = "";
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeOutside(final: boolean, events: InbandScanEvent[]): void {
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
const match = findEarliestToken(this.#buffer, OUTSIDE_TOKENS);
|
||||||
|
if (!match) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, OUTSIDE_TOKENS);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (emit.length > 0) events.push({ type: "text", text: emit });
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (match.index > 0) events.push({ type: "text", text: this.#buffer.slice(0, match.index) });
|
||||||
|
this.#buffer = this.#buffer.slice(match.index);
|
||||||
|
if (this.#buffer.startsWith(DEEPSEEK_TOOL_CALLS_BEGIN)) {
|
||||||
|
this.#buffer = this.#buffer.slice(DEEPSEEK_TOOL_CALLS_BEGIN.length);
|
||||||
|
this.#inToolSection = true;
|
||||||
|
this.#state = "section";
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (this.#buffer.startsWith(DEEPSEEK_TOOL_CALL_BEGIN)) {
|
||||||
|
this.#buffer = this.#buffer.slice(DEEPSEEK_TOOL_CALL_BEGIN.length);
|
||||||
|
this.#rawBlock = DEEPSEEK_TOOL_CALL_BEGIN;
|
||||||
|
this.#inToolSection = false;
|
||||||
|
this.#state = "header";
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (this.#buffer.startsWith(THINK_OPEN)) {
|
||||||
|
this.#buffer = this.#buffer.slice(THINK_OPEN.length);
|
||||||
|
this.#state = "thinking";
|
||||||
|
this.#thinking = "";
|
||||||
|
if (this.#parseThinking) events.push({ type: "thinkingStart" });
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (
|
||||||
|
this.#buffer.startsWith(DSML_TOOL_CALLS_OPEN_FULLWIDTH) ||
|
||||||
|
this.#buffer.startsWith(DSML_TOOL_CALLS_OPEN_ASCII)
|
||||||
|
) {
|
||||||
|
const openToken = this.#buffer.startsWith(DSML_TOOL_CALLS_OPEN_FULLWIDTH)
|
||||||
|
? DSML_TOOL_CALLS_OPEN_FULLWIDTH
|
||||||
|
: DSML_TOOL_CALLS_OPEN_ASCII;
|
||||||
|
this.#buffer = this.#buffer.slice(openToken.length);
|
||||||
|
this.#state = "dsmlSection";
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const control = this.#matchingControlToken();
|
||||||
|
if (control) {
|
||||||
|
this.#buffer = this.#buffer.slice(control.length);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
this.#buffer = this.#buffer.slice(match.token.length);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeThinking(final: boolean, events: InbandScanEvent[]): void {
|
||||||
|
const close = this.#buffer.indexOf(THINK_CLOSE);
|
||||||
|
if (close === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, [THINK_CLOSE]);
|
||||||
|
this.#emitThinking(this.#buffer.slice(0, this.#buffer.length - hold), events);
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
if (final) this.#endThinking(events);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
this.#emitThinking(this.#buffer.slice(0, close), events);
|
||||||
|
this.#buffer = this.#buffer.slice(close + THINK_CLOSE.length);
|
||||||
|
this.#endThinking(events);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeSection(final: boolean): boolean {
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
this.#skipWhitespace();
|
||||||
|
if (this.#buffer.startsWith(DEEPSEEK_TOOL_CALLS_END)) {
|
||||||
|
this.#buffer = this.#buffer.slice(DEEPSEEK_TOOL_CALLS_END.length);
|
||||||
|
this.#inToolSection = false;
|
||||||
|
this.#state = "outside";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (this.#buffer.startsWith(DEEPSEEK_TOOL_CALL_BEGIN)) {
|
||||||
|
this.#buffer = this.#buffer.slice(DEEPSEEK_TOOL_CALL_BEGIN.length);
|
||||||
|
this.#rawBlock = DEEPSEEK_TOOL_CALL_BEGIN;
|
||||||
|
this.#state = "header";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!final && partialSuffixOverlapAny(this.#buffer, SECTION_TOKENS) === this.#buffer.length) return false;
|
||||||
|
if (this.#buffer.length === 0) return false;
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
}
|
||||||
|
return final;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeHeader(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const sep = this.#buffer.indexOf(DEEPSEEK_TOOL_SEPARATOR);
|
||||||
|
if (sep === -1) {
|
||||||
|
if (final) this.#resetTool();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
const rawHead = this.#buffer.slice(0, sep + DEEPSEEK_TOOL_SEPARATOR.length);
|
||||||
|
const head = this.#buffer.slice(0, sep).trim();
|
||||||
|
this.#rawBlock += rawHead;
|
||||||
|
this.#buffer = this.#buffer.slice(rawHead.length);
|
||||||
|
if (head === LEGACY_TOOL_TYPE) {
|
||||||
|
this.#state = "legacyName";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
this.#startTool(head, events);
|
||||||
|
this.#state = "args";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeLegacyName(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const fence = this.#buffer.indexOf(LEGACY_JSON_FENCE);
|
||||||
|
if (fence === -1) {
|
||||||
|
if (final) this.#resetTool();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
const rawName = this.#buffer.slice(0, fence + LEGACY_JSON_FENCE.length);
|
||||||
|
const name = this.#buffer.slice(0, fence).trim();
|
||||||
|
this.#rawBlock += rawName;
|
||||||
|
this.#buffer = this.#buffer.slice(rawName.length);
|
||||||
|
this.#rawBlock += this.#dropOneLineBreak();
|
||||||
|
this.#startTool(name, events);
|
||||||
|
this.#state = "legacyArgs";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeArgs(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const end = this.#buffer.indexOf(DEEPSEEK_TOOL_CALL_END);
|
||||||
|
if (end === -1) {
|
||||||
|
if (final) this.#resetTool();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
let rawArgs = this.#buffer.slice(0, end);
|
||||||
|
if (this.#state === "legacyArgs") {
|
||||||
|
const fence = rawArgs.lastIndexOf(CODE_FENCE);
|
||||||
|
if (fence !== -1) rawArgs = rawArgs.slice(0, fence);
|
||||||
|
}
|
||||||
|
const rawTail = this.#buffer.slice(0, end + DEEPSEEK_TOOL_CALL_END.length);
|
||||||
|
this.#rawBlock += rawTail;
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: this.#name,
|
||||||
|
arguments: this.#parseArgs(rawArgs),
|
||||||
|
rawBlock: this.#rawBlock,
|
||||||
|
});
|
||||||
|
this.#buffer = this.#buffer.slice(rawTail.length);
|
||||||
|
this.#resetTool(this.#inToolSection ? "section" : "outside");
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeDsmlSection(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
this.#skipWhitespace();
|
||||||
|
const close = this.#matchingDsmlClose(DSML_TOOL_CALLS_CLOSE_FULLWIDTH, DSML_TOOL_CALLS_CLOSE_ASCII);
|
||||||
|
if (close) {
|
||||||
|
this.#buffer = this.#buffer.slice(close.length);
|
||||||
|
this.#state = "outside";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
const invoke = this.#matchDsmlOpen("invoke");
|
||||||
|
if (invoke) {
|
||||||
|
this.#rawBlock = invoke.raw;
|
||||||
|
this.#name = invoke.name;
|
||||||
|
this.#id = mintToolCallId();
|
||||||
|
this.#dsmlArgs = {};
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
this.#state = "dsmlInvoke";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!final) {
|
||||||
|
if (
|
||||||
|
(this.#buffer.startsWith("<|DSML|invoke") || this.#buffer.startsWith("<|DSML|invoke")) &&
|
||||||
|
!this.#buffer.includes(">")
|
||||||
|
)
|
||||||
|
return false;
|
||||||
|
if (partialSuffixOverlapAny(this.#buffer, DSML_SECTION_TOKENS) === this.#buffer.length) return false;
|
||||||
|
}
|
||||||
|
if (this.#buffer.length === 0) return false;
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
}
|
||||||
|
return final;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeDsmlInvoke(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
const skipped = this.#skipWhitespace();
|
||||||
|
if (skipped.length > 0) this.#rawBlock += skipped;
|
||||||
|
const close = this.#matchingDsmlClose("</|DSML|invoke>", "</|DSML|invoke>");
|
||||||
|
if (close) {
|
||||||
|
this.#rawBlock += close;
|
||||||
|
this.#buffer = this.#buffer.slice(close.length);
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: this.#name,
|
||||||
|
arguments: this.#dsmlArgs,
|
||||||
|
rawBlock: this.#rawBlock,
|
||||||
|
});
|
||||||
|
this.#resetDsmlTool();
|
||||||
|
this.#state = "dsmlSection";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
const param = this.#matchDsmlOpen("parameter");
|
||||||
|
if (param) {
|
||||||
|
this.#rawBlock += param.raw;
|
||||||
|
this.#dsmlParamName = param.name;
|
||||||
|
this.#dsmlParamIsString = param.stringAttr !== "false";
|
||||||
|
this.#state = "dsmlParam";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!final) {
|
||||||
|
if (
|
||||||
|
(this.#buffer.startsWith("<|DSML|parameter") || this.#buffer.startsWith("<|DSML|parameter")) &&
|
||||||
|
!this.#buffer.includes(">")
|
||||||
|
)
|
||||||
|
return false;
|
||||||
|
if (partialSuffixOverlapAny(this.#buffer, DSML_INVOKE_TOKENS) === this.#buffer.length) return false;
|
||||||
|
}
|
||||||
|
const consumed = this.#buffer[0]!;
|
||||||
|
this.#rawBlock += consumed;
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
}
|
||||||
|
return final;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeDsmlParam(final: boolean): boolean {
|
||||||
|
const close = findEarliestToken(this.#buffer, DSML_PARAMETER_CLOSE_TOKENS);
|
||||||
|
if (!close) {
|
||||||
|
if (final) this.#resetDsmlTool();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
const rawValue = this.#buffer.slice(0, close.index);
|
||||||
|
this.#dsmlArgs[this.#dsmlParamName] = coerceDsmlValue(rawValue, this.#dsmlParamIsString);
|
||||||
|
this.#rawBlock += rawValue + close.token;
|
||||||
|
this.#buffer = this.#buffer.slice(close.index + close.token.length);
|
||||||
|
this.#dsmlParamName = "";
|
||||||
|
this.#dsmlParamIsString = true;
|
||||||
|
this.#state = "dsmlInvoke";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#startTool(name: string, events: InbandScanEvent[]): void {
|
||||||
|
this.#name = name;
|
||||||
|
this.#id = mintToolCallId();
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitThinking(delta: string, events: InbandScanEvent[]): void {
|
||||||
|
if (delta.length === 0) return;
|
||||||
|
if (this.#parseThinking) {
|
||||||
|
this.#thinking += delta;
|
||||||
|
events.push({ type: "thinkingDelta", delta });
|
||||||
|
} else {
|
||||||
|
events.push({ type: "text", text: delta });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#endThinking(events: InbandScanEvent[]): void {
|
||||||
|
if (this.#parseThinking) events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#parseArgs(rawArgs: string): Record<string, unknown> {
|
||||||
|
const trimmed = rawArgs.trim();
|
||||||
|
if (trimmed.length === 0) return {};
|
||||||
|
try {
|
||||||
|
return asRecord(parseJsonWithRepair<unknown>(trimmed));
|
||||||
|
} catch {
|
||||||
|
return {};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#skipWhitespace(): string {
|
||||||
|
let i = 0;
|
||||||
|
while (i < this.#buffer.length && /\s/.test(this.#buffer[i]!)) i++;
|
||||||
|
if (i === 0) return "";
|
||||||
|
const skipped = this.#buffer.slice(0, i);
|
||||||
|
this.#buffer = this.#buffer.slice(i);
|
||||||
|
return skipped;
|
||||||
|
}
|
||||||
|
|
||||||
|
#dropOneLineBreak(): string {
|
||||||
|
if (this.#buffer.startsWith("\r\n")) {
|
||||||
|
this.#buffer = this.#buffer.slice(2);
|
||||||
|
return "\r\n";
|
||||||
|
}
|
||||||
|
if (this.#buffer.startsWith("\n")) {
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return "\n";
|
||||||
|
}
|
||||||
|
return "";
|
||||||
|
}
|
||||||
|
|
||||||
|
#matchingControlToken(): string | undefined {
|
||||||
|
if (this.#buffer.startsWith(DEEPSEEK_TOOL_CALLS_END)) return DEEPSEEK_TOOL_CALLS_END;
|
||||||
|
if (this.#buffer.startsWith(THINK_CLOSE)) return THINK_CLOSE;
|
||||||
|
if (this.#buffer.startsWith(DSML_TOOL_CALLS_CLOSE_FULLWIDTH)) return DSML_TOOL_CALLS_CLOSE_FULLWIDTH;
|
||||||
|
if (this.#buffer.startsWith(DSML_TOOL_CALLS_CLOSE_ASCII)) return DSML_TOOL_CALLS_CLOSE_ASCII;
|
||||||
|
for (const token of CONTROL_TOKENS) {
|
||||||
|
if (this.#buffer.startsWith(token)) return token;
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
#matchingDsmlClose(fullwidth: string, ascii: string): string | undefined {
|
||||||
|
if (this.#buffer.startsWith(fullwidth)) return fullwidth;
|
||||||
|
if (this.#buffer.startsWith(ascii)) return ascii;
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
#matchDsmlOpen(
|
||||||
|
kind: "invoke" | "parameter",
|
||||||
|
): { name: string; stringAttr: string | undefined; raw: string } | undefined {
|
||||||
|
if (!this.#buffer.startsWith(`<|DSML|${kind}`) && !this.#buffer.startsWith(`<|DSML|${kind}`)) return undefined;
|
||||||
|
const end = this.#buffer.indexOf(">");
|
||||||
|
if (end === -1) return undefined;
|
||||||
|
const tag = this.#buffer.slice(0, end + 1);
|
||||||
|
const name = /\sname="([^"]*)"/.exec(tag)?.[1];
|
||||||
|
if (name === undefined) return undefined;
|
||||||
|
const stringAttr = /\sstring="(true|false)"/.exec(tag)?.[1];
|
||||||
|
this.#buffer = this.#buffer.slice(end + 1);
|
||||||
|
return { name, stringAttr, raw: tag };
|
||||||
|
}
|
||||||
|
|
||||||
|
#resetTool(next: State = "outside"): void {
|
||||||
|
this.#state = next;
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#rawBlock = "";
|
||||||
|
}
|
||||||
|
|
||||||
|
#resetDsmlTool(): void {
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#dsmlArgs = {};
|
||||||
|
this.#dsmlParamName = "";
|
||||||
|
this.#dsmlParamIsString = true;
|
||||||
|
this.#rawBlock = "";
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function findEarliestToken(text: string, tokens: readonly string[]): { index: number; token: string } | undefined {
|
||||||
|
let bestIndex = -1;
|
||||||
|
let bestToken = "";
|
||||||
|
for (const token of tokens) {
|
||||||
|
const index = text.indexOf(token);
|
||||||
|
if (index === -1) continue;
|
||||||
|
if (bestIndex === -1 || index < bestIndex || (index === bestIndex && token.length > bestToken.length)) {
|
||||||
|
bestIndex = index;
|
||||||
|
bestToken = token;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return bestIndex === -1 ? undefined : { index: bestIndex, token: bestToken };
|
||||||
|
}
|
||||||
|
|
||||||
|
function coerceDsmlValue(raw: string, isString: boolean): unknown {
|
||||||
|
if (isString) return raw;
|
||||||
|
const trimmed = raw.trim();
|
||||||
|
if (trimmed.length === 0) return raw;
|
||||||
|
try {
|
||||||
|
return parseJsonWithRepair<unknown>(trimmed);
|
||||||
|
} catch {
|
||||||
|
return raw;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "deepseek",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: options => new DeepSeekInbandScanner(options),
|
||||||
|
renderAssistantToolCalls: renderDeepSeekToolCalls,
|
||||||
|
renderToolResults: renderDeepSeekToolResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
import anthropicGrammar from "./anthropic";
|
||||||
|
import deepseekGrammar from "./deepseek";
|
||||||
|
import glmGrammar from "./glm";
|
||||||
|
import harmonyGrammar from "./harmony";
|
||||||
|
import hermesGrammar from "./hermes";
|
||||||
|
import kimiGrammar from "./kimi";
|
||||||
|
import piGrammar from "./pi";
|
||||||
|
import qwen3Grammar from "./qwen3";
|
||||||
|
import type { Grammar, InbandScanner, InbandScannerOptions, ToolCallSyntax } from "./types";
|
||||||
|
import xmlGrammar from "./xml";
|
||||||
|
|
||||||
|
const GRAMMARS: Record<ToolCallSyntax, Grammar> = {
|
||||||
|
glm: glmGrammar,
|
||||||
|
hermes: hermesGrammar,
|
||||||
|
kimi: kimiGrammar,
|
||||||
|
xml: xmlGrammar,
|
||||||
|
anthropic: anthropicGrammar,
|
||||||
|
deepseek: deepseekGrammar,
|
||||||
|
harmony: harmonyGrammar,
|
||||||
|
pi: piGrammar,
|
||||||
|
qwen3: qwen3Grammar,
|
||||||
|
};
|
||||||
|
|
||||||
|
export function getInbandGrammar(syntax: ToolCallSyntax): Grammar {
|
||||||
|
return GRAMMARS[syntax];
|
||||||
|
}
|
||||||
|
|
||||||
|
export function createInbandScanner(syntax: ToolCallSyntax, options: InbandScannerOptions = {}): InbandScanner {
|
||||||
|
return getInbandGrammar(syntax).createScanner(options);
|
||||||
|
}
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
Emit each call as a `<tool_call>` block. The function name goes on the same line as the opening tag, followed by one `<arg_key>`/`<arg_value>` pair per argument, closed by `</tool_call>`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_call>get_weather
|
||||||
|
<arg_key>location</arg_key>
|
||||||
|
<arg_value>Beijing</arg_value>
|
||||||
|
<arg_key>days</arg_key>
|
||||||
|
<arg_value>3</arg_value>
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
Tool results return in an observation block:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<observation>
|
||||||
|
<tool_response>
|
||||||
|
verbatim tool result
|
||||||
|
</tool_response>
|
||||||
|
</observation>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- The name after `<tool_call>` must match a listed function and sit on the same line.
|
||||||
|
- Emit one `<arg_key>name</arg_key>` + `<arg_value>value</arg_value>` pair per argument; omit unset optional args.
|
||||||
|
- String values are raw text (no quotes, no escaping); non-string values are valid JSON.
|
||||||
|
- Multiple calls are consecutive `<tool_call>…</tool_call>` blocks.
|
||||||
|
- Private reasoning goes in `<think>…</think>`; NEVER put tool calls inside `<think>`.
|
||||||
|
- Read each `<tool_response>` in call order. NEVER emit `<tool_response>` yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,383 @@
|
|||||||
|
import {
|
||||||
|
buildStringArgsResolver,
|
||||||
|
decodeValue,
|
||||||
|
mintToolCallId,
|
||||||
|
partialSuffixOverlap,
|
||||||
|
partialSuffixOverlapAny,
|
||||||
|
} from "./coercion";
|
||||||
|
import grammarPrompt from "./glm.md" with { type: "text" };
|
||||||
|
import { renderGlmToolCalls, renderGlmToolResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner, InbandScannerOptions } from "./types";
|
||||||
|
|
||||||
|
const TOOL_OPEN = "<tool_call>";
|
||||||
|
const TOOL_CLOSE = "</tool_call>";
|
||||||
|
const ARG_KEY_OPEN = "<arg_key>";
|
||||||
|
const ARG_KEY_CLOSE = "</arg_key>";
|
||||||
|
const ARG_VALUE_OPEN = "<arg_value>";
|
||||||
|
const ARG_VALUE_CLOSE = "</arg_value>";
|
||||||
|
const RESPONSE_OPEN = "<tool_response>";
|
||||||
|
const RESPONSE_CLOSE = "</tool_response>";
|
||||||
|
const THINK_OPEN = "<think>";
|
||||||
|
const THINK_CLOSE = "</think>";
|
||||||
|
|
||||||
|
const OUTSIDE_TAGS = [
|
||||||
|
TOOL_OPEN,
|
||||||
|
ARG_KEY_OPEN,
|
||||||
|
ARG_KEY_CLOSE,
|
||||||
|
ARG_VALUE_OPEN,
|
||||||
|
ARG_VALUE_CLOSE,
|
||||||
|
RESPONSE_OPEN,
|
||||||
|
RESPONSE_CLOSE,
|
||||||
|
THINK_OPEN,
|
||||||
|
THINK_CLOSE,
|
||||||
|
] as const;
|
||||||
|
const OUTSIDE_TAGS_NO_THINK = [
|
||||||
|
TOOL_OPEN,
|
||||||
|
ARG_KEY_OPEN,
|
||||||
|
ARG_KEY_CLOSE,
|
||||||
|
ARG_VALUE_OPEN,
|
||||||
|
ARG_VALUE_CLOSE,
|
||||||
|
RESPONSE_OPEN,
|
||||||
|
RESPONSE_CLOSE,
|
||||||
|
] as const;
|
||||||
|
const BODY_TAGS = [ARG_KEY_OPEN, TOOL_CLOSE] as const;
|
||||||
|
|
||||||
|
type State = "outside" | "thinking" | "name" | "body" | "key" | "afterkey" | "value";
|
||||||
|
|
||||||
|
interface OpenCall {
|
||||||
|
id: string;
|
||||||
|
name: string;
|
||||||
|
stringArgs: ReadonlySet<string>;
|
||||||
|
arguments: Record<string, unknown>;
|
||||||
|
key: string | null;
|
||||||
|
valueRaw: string;
|
||||||
|
rawBlock: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface TagMatch {
|
||||||
|
index: number;
|
||||||
|
tag: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export class GLMInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#state: State = "outside";
|
||||||
|
#call: OpenCall | null = null;
|
||||||
|
#thinking = "";
|
||||||
|
#parseThinking: boolean;
|
||||||
|
#stringArgs: (toolName: string) => ReadonlySet<string>;
|
||||||
|
|
||||||
|
constructor(options: InbandScannerOptions = {}) {
|
||||||
|
this.#parseThinking = options.parseThinking !== false;
|
||||||
|
this.#stringArgs = options.stringArgs ?? buildStringArgsResolver(options.tools);
|
||||||
|
}
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#state === "outside") {
|
||||||
|
if (!this.#consumeOutside(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "thinking") {
|
||||||
|
this.#consumeThinking(final, events);
|
||||||
|
if (this.#state === "thinking") break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "name") {
|
||||||
|
if (!this.#consumeName(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "body") {
|
||||||
|
if (!this.#consumeBody(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "key") {
|
||||||
|
if (!this.#consumeKey(final)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "afterkey") {
|
||||||
|
if (!this.#consumeAfterKey(final)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!this.#consumeValue(final, events)) break;
|
||||||
|
}
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeOutside(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const tags = this.#parseThinking ? OUTSIDE_TAGS : OUTSIDE_TAGS_NO_THINK;
|
||||||
|
const match = findFirstTag(this.#buffer, tags);
|
||||||
|
if (!match) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, tags);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (emit.length > 0) events.push({ type: "text", text: emit });
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (match.index > 0) events.push({ type: "text", text: this.#buffer.slice(0, match.index) });
|
||||||
|
this.#buffer = this.#buffer.slice(match.index + match.tag.length);
|
||||||
|
|
||||||
|
if (match.tag === TOOL_OPEN) {
|
||||||
|
this.#state = "name";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (match.tag === THINK_OPEN && this.#parseThinking) {
|
||||||
|
this.#thinking = "";
|
||||||
|
events.push({ type: "thinkingStart" });
|
||||||
|
this.#state = "thinking";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (match.tag === RESPONSE_OPEN) {
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeThinking(final: boolean, events: InbandScanEvent[]): void {
|
||||||
|
const close = this.#buffer.indexOf(THINK_CLOSE);
|
||||||
|
if (close === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlap(this.#buffer, THINK_CLOSE);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
this.#emitThinking(emit, events);
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
if (final) this.#endThinking(events);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
this.#emitThinking(this.#buffer.slice(0, close), events);
|
||||||
|
this.#buffer = this.#buffer.slice(close + THINK_CLOSE.length);
|
||||||
|
this.#endThinking(events);
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeName(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const newline = this.#buffer.indexOf("\n");
|
||||||
|
const key = this.#buffer.indexOf(ARG_KEY_OPEN);
|
||||||
|
const close = this.#buffer.indexOf(TOOL_CLOSE);
|
||||||
|
const delimiter = minFound(newline, key, close);
|
||||||
|
if (delimiter === -1) {
|
||||||
|
if (!final) return false;
|
||||||
|
this.#beginCall(this.#buffer, events);
|
||||||
|
this.#buffer = "";
|
||||||
|
this.#endCall(events);
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const rawName = this.#buffer.slice(0, delimiter);
|
||||||
|
this.#beginCall(rawName, events);
|
||||||
|
if (delimiter === newline) {
|
||||||
|
this.#appendCallRaw("\n");
|
||||||
|
this.#buffer = this.#buffer.slice(delimiter + 1);
|
||||||
|
this.#state = "body";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (delimiter === key) {
|
||||||
|
this.#appendCallRaw(ARG_KEY_OPEN);
|
||||||
|
this.#buffer = this.#buffer.slice(delimiter + ARG_KEY_OPEN.length);
|
||||||
|
this.#state = "key";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
this.#appendCallRaw(TOOL_CLOSE);
|
||||||
|
this.#buffer = this.#buffer.slice(delimiter + TOOL_CLOSE.length);
|
||||||
|
this.#endCall(events);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeBody(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
this.#appendCallRaw(this.#skipWhitespace());
|
||||||
|
if (this.#buffer.length === 0) return false;
|
||||||
|
if (this.#buffer.startsWith(ARG_KEY_OPEN)) {
|
||||||
|
this.#appendCallRaw(ARG_KEY_OPEN);
|
||||||
|
this.#buffer = this.#buffer.slice(ARG_KEY_OPEN.length);
|
||||||
|
this.#state = "key";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (this.#buffer.startsWith(TOOL_CLOSE)) {
|
||||||
|
this.#appendCallRaw(TOOL_CLOSE);
|
||||||
|
this.#buffer = this.#buffer.slice(TOOL_CLOSE.length);
|
||||||
|
this.#endCall(events);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!final && partialSuffixOverlapAny(this.#buffer, BODY_TAGS) === this.#buffer.length) return false;
|
||||||
|
this.#appendCallRaw(this.#buffer[0] ?? "");
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeKey(final: boolean): boolean {
|
||||||
|
const close = this.#buffer.indexOf(ARG_KEY_CLOSE);
|
||||||
|
if (close === -1) {
|
||||||
|
if (final) this.#dropCall();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
if (this.#call) {
|
||||||
|
this.#call.key = this.#buffer.slice(0, close).trim();
|
||||||
|
this.#appendCallRaw(this.#buffer.slice(0, close + ARG_KEY_CLOSE.length));
|
||||||
|
}
|
||||||
|
this.#buffer = this.#buffer.slice(close + ARG_KEY_CLOSE.length);
|
||||||
|
this.#state = "afterkey";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeAfterKey(final: boolean): boolean {
|
||||||
|
this.#appendCallRaw(this.#skipWhitespace());
|
||||||
|
if (this.#buffer.length === 0) return false;
|
||||||
|
if (this.#buffer.startsWith(ARG_VALUE_OPEN)) {
|
||||||
|
this.#appendCallRaw(ARG_VALUE_OPEN);
|
||||||
|
this.#buffer = this.#buffer.slice(ARG_VALUE_OPEN.length);
|
||||||
|
if (this.#call) this.#call.valueRaw = "";
|
||||||
|
this.#state = "value";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (!final && ARG_VALUE_OPEN.startsWith(this.#buffer)) return false;
|
||||||
|
this.#appendCallRaw(this.#buffer[0] ?? "");
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeValue(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const close = this.#buffer.indexOf(ARG_VALUE_CLOSE);
|
||||||
|
if (close === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlap(this.#buffer, ARG_VALUE_CLOSE);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
this.#streamValue(emit, events);
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
if (final) this.#dropCall();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
this.#streamValue(this.#buffer.slice(0, close), events);
|
||||||
|
this.#appendCallRaw(ARG_VALUE_CLOSE);
|
||||||
|
this.#buffer = this.#buffer.slice(close + ARG_VALUE_CLOSE.length);
|
||||||
|
this.#endValue();
|
||||||
|
this.#state = "body";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#beginCall(rawName: string, events: InbandScanEvent[]): void {
|
||||||
|
const name = rawName.trim();
|
||||||
|
if (name.length === 0) {
|
||||||
|
this.#dropCall();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
const id = mintToolCallId();
|
||||||
|
this.#call = {
|
||||||
|
id,
|
||||||
|
name,
|
||||||
|
stringArgs: this.#stringArgs(name),
|
||||||
|
arguments: {},
|
||||||
|
key: null,
|
||||||
|
valueRaw: "",
|
||||||
|
rawBlock: `${TOOL_OPEN}${rawName}`,
|
||||||
|
};
|
||||||
|
events.push({ type: "toolStart", id, name });
|
||||||
|
}
|
||||||
|
|
||||||
|
#streamValue(chunk: string, events: InbandScanEvent[]): void {
|
||||||
|
const call = this.#call;
|
||||||
|
if (!call || call.key === null || chunk.length === 0) return;
|
||||||
|
call.valueRaw += chunk;
|
||||||
|
call.rawBlock += chunk;
|
||||||
|
events.push({ type: "toolArgDelta", id: call.id, name: call.name, key: call.key, delta: chunk });
|
||||||
|
}
|
||||||
|
|
||||||
|
#endValue(): void {
|
||||||
|
const call = this.#call;
|
||||||
|
if (!call || call.key === null) return;
|
||||||
|
call.arguments[call.key] = call.stringArgs.has(call.key) ? call.valueRaw : decodeValue(call.valueRaw);
|
||||||
|
call.key = null;
|
||||||
|
call.valueRaw = "";
|
||||||
|
}
|
||||||
|
|
||||||
|
#endCall(events: InbandScanEvent[]): void {
|
||||||
|
const call = this.#call;
|
||||||
|
if (!call) {
|
||||||
|
this.#state = "outside";
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: call.id,
|
||||||
|
name: call.name,
|
||||||
|
arguments: call.arguments,
|
||||||
|
rawBlock: call.rawBlock,
|
||||||
|
});
|
||||||
|
this.#call = null;
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#dropCall(): void {
|
||||||
|
this.#call = null;
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#appendCallRaw(text: string): void {
|
||||||
|
if (this.#call && text.length > 0) this.#call.rawBlock += text;
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitThinking(delta: string, events: InbandScanEvent[]): void {
|
||||||
|
if (delta.length === 0) return;
|
||||||
|
this.#thinking += delta;
|
||||||
|
events.push({ type: "thinkingDelta", delta });
|
||||||
|
}
|
||||||
|
|
||||||
|
#endThinking(events: InbandScanEvent[]): void {
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#skipWhitespace(): string {
|
||||||
|
let i = 0;
|
||||||
|
while (i < this.#buffer.length && " \n\t\r".includes(this.#buffer[i]!)) i++;
|
||||||
|
const skipped = this.#buffer.slice(0, i);
|
||||||
|
if (i > 0) this.#buffer = this.#buffer.slice(i);
|
||||||
|
return skipped;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function findFirstTag(text: string, tags: readonly string[]): TagMatch | null {
|
||||||
|
let best: TagMatch | null = null;
|
||||||
|
for (const tag of tags) {
|
||||||
|
const index = text.indexOf(tag);
|
||||||
|
if (index === -1) continue;
|
||||||
|
if (!best || index < best.index) best = { index, tag };
|
||||||
|
}
|
||||||
|
return best;
|
||||||
|
}
|
||||||
|
|
||||||
|
function minFound(...values: readonly number[]): number {
|
||||||
|
let best = -1;
|
||||||
|
for (const value of values) {
|
||||||
|
if (value === -1) continue;
|
||||||
|
if (best === -1 || value < best) best = value;
|
||||||
|
}
|
||||||
|
return best;
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "glm",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: options => new GLMInbandScanner(options),
|
||||||
|
renderAssistantToolCalls: renderGlmToolCalls,
|
||||||
|
renderToolResults: renderGlmToolResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
Each function call is one assistant message on the `commentary` channel addressed to the function, emitted as text:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>assistant<|channel|>commentary to=functions.function_name <|constrain|>json<|message|>{"arg":"value"}<|call|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Put private reasoning in an `analysis` message:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>assistant<|channel|>analysis<|message|>private reasoning<|end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Tool results arrive as messages authored by the function, addressed back to the assistant:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|start|>functions.function_name to=assistant<|channel|>commentary<|message|>verbatim tool result<|end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- Recipient is `functions.` + a listed function name.
|
||||||
|
- Body is one JSON object matching the schema; omit optional arguments you are not setting.
|
||||||
|
- Multiple calls = consecutive call messages.
|
||||||
|
- An optional visible preamble is a `commentary` message ending `<|end|>`.
|
||||||
|
- NEVER put tool calls in `analysis`.
|
||||||
|
- NEVER wrap calls in Markdown/code fences.
|
||||||
|
- Read each tool-result message in call order. NEVER emit tool-result messages yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,271 @@
|
|||||||
|
import { parseJsonWithRepair } from "../utils/json-parse";
|
||||||
|
import { asRecord, mintToolCallId, partialSuffixOverlapAny } from "./coercion";
|
||||||
|
import grammarPrompt from "./harmony.md" with { type: "text" };
|
||||||
|
import { renderHarmonyToolCalls, renderHarmonyToolResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner } from "./types";
|
||||||
|
|
||||||
|
const START = "<|start|>";
|
||||||
|
const END = "<|end|>";
|
||||||
|
const MESSAGE = "<|message|>";
|
||||||
|
const CHANNEL = "<|channel|>";
|
||||||
|
const CONSTRAIN = "<|constrain|>";
|
||||||
|
const RETURN = "<|return|>";
|
||||||
|
const CALL = "<|call|>";
|
||||||
|
|
||||||
|
const ALL_TOKENS = [START, END, MESSAGE, CHANNEL, CONSTRAIN, RETURN, CALL] as const;
|
||||||
|
const BODY_TOKENS = [END, CALL, RETURN, START, CHANNEL, MESSAGE, CONSTRAIN] as const;
|
||||||
|
|
||||||
|
type State = "outside" | "header" | "body";
|
||||||
|
type BodyMode = "text" | "thinking" | "tool" | "skip";
|
||||||
|
|
||||||
|
interface HeaderFields {
|
||||||
|
role: string;
|
||||||
|
channel: string;
|
||||||
|
recipient: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface TokenMatch {
|
||||||
|
index: number;
|
||||||
|
token: string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export class HarmonyInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#state: State = "outside";
|
||||||
|
#mode: BodyMode = "skip";
|
||||||
|
#id = "";
|
||||||
|
#name = "";
|
||||||
|
#toolArgs = "";
|
||||||
|
#thinking = "";
|
||||||
|
#rawBlock = "";
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#state === "outside") {
|
||||||
|
const next = findNextToken(this.#buffer, ALL_TOKENS);
|
||||||
|
if (!next) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, ALL_TOKENS);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (emit.length > 0) events.push({ type: "text", text: emit });
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (next.index > 0) events.push({ type: "text", text: this.#buffer.slice(0, next.index) });
|
||||||
|
if (next.token === START) {
|
||||||
|
this.#rawBlock = START;
|
||||||
|
this.#buffer = this.#buffer.slice(next.index + START.length);
|
||||||
|
this.#state = "header";
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (next.token === CHANNEL) {
|
||||||
|
this.#rawBlock = "";
|
||||||
|
this.#buffer = this.#buffer.slice(next.index);
|
||||||
|
this.#state = "header";
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#buffer = this.#buffer.slice(next.index + next.token.length);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "header") {
|
||||||
|
const message = this.#buffer.indexOf(MESSAGE);
|
||||||
|
if (message === -1) {
|
||||||
|
if (final) this.#resetAll();
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
const rawHeader = this.#buffer.slice(0, message);
|
||||||
|
this.#rawBlock += this.#buffer.slice(0, message + MESSAGE.length);
|
||||||
|
const header = this.#parseHeader(rawHeader);
|
||||||
|
this.#buffer = this.#buffer.slice(message + MESSAGE.length);
|
||||||
|
this.#enterBody(header, events);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
const next = findNextToken(this.#buffer, BODY_TOKENS);
|
||||||
|
if (!next) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, BODY_TOKENS);
|
||||||
|
this.#emitBody(this.#buffer.slice(0, this.#buffer.length - hold), events);
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
if (final && this.#buffer.length === 0) {
|
||||||
|
this.#finishBody(events);
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#emitBody(this.#buffer.slice(0, next.index), events);
|
||||||
|
if (next.token === END || next.token === CALL || next.token === RETURN) {
|
||||||
|
if (this.#mode === "tool") this.#rawBlock += next.token;
|
||||||
|
this.#buffer = this.#buffer.slice(next.index + next.token.length);
|
||||||
|
this.#finishBody(events);
|
||||||
|
this.#state = "outside";
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (next.token === START) {
|
||||||
|
this.#finishBody(events);
|
||||||
|
this.#rawBlock = START;
|
||||||
|
this.#buffer = this.#buffer.slice(next.index + START.length);
|
||||||
|
this.#state = "header";
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (next.token === CHANNEL) {
|
||||||
|
this.#finishBody(events);
|
||||||
|
this.#rawBlock = "";
|
||||||
|
this.#buffer = this.#buffer.slice(next.index);
|
||||||
|
this.#state = "header";
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#mode === "tool") this.#rawBlock += next.token;
|
||||||
|
this.#buffer = this.#buffer.slice(next.index + next.token.length);
|
||||||
|
}
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#enterBody(header: HeaderFields, events: InbandScanEvent[]): void {
|
||||||
|
this.#clearBody(false);
|
||||||
|
this.#state = "body";
|
||||||
|
|
||||||
|
const assistantMessage = header.role === "" || header.role === "assistant";
|
||||||
|
if (!assistantMessage) {
|
||||||
|
this.#mode = "skip";
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (header.recipient.length > 0 && header.recipient !== "assistant") {
|
||||||
|
this.#mode = "tool";
|
||||||
|
this.#id = mintToolCallId();
|
||||||
|
this.#name = header.recipient.startsWith("functions.")
|
||||||
|
? header.recipient.slice("functions.".length)
|
||||||
|
: header.recipient;
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (header.channel === "analysis") {
|
||||||
|
this.#mode = "thinking";
|
||||||
|
events.push({ type: "thinkingStart" });
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#mode = "text";
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitBody(chunk: string, events: InbandScanEvent[]): void {
|
||||||
|
if (chunk.length === 0) return;
|
||||||
|
if (this.#mode === "text") {
|
||||||
|
events.push({ type: "text", text: chunk });
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (this.#mode === "thinking") {
|
||||||
|
this.#thinking += chunk;
|
||||||
|
events.push({ type: "thinkingDelta", delta: chunk });
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (this.#mode === "tool") {
|
||||||
|
this.#rawBlock += chunk;
|
||||||
|
this.#toolArgs += chunk;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#finishBody(events: InbandScanEvent[]): void {
|
||||||
|
if (this.#mode === "thinking") {
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
} else if (this.#mode === "tool" && this.#name.length > 0) {
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: this.#name,
|
||||||
|
arguments: this.#parseArgs(),
|
||||||
|
rawBlock: this.#rawBlock,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
this.#clearBody();
|
||||||
|
}
|
||||||
|
|
||||||
|
#parseHeader(rawHeader: string): HeaderFields {
|
||||||
|
const channelIndex = rawHeader.indexOf(CHANNEL);
|
||||||
|
const rolePart = channelIndex === -1 ? rawHeader : rawHeader.slice(0, channelIndex);
|
||||||
|
const channelPart = channelIndex === -1 ? "" : rawHeader.slice(channelIndex + CHANNEL.length);
|
||||||
|
return {
|
||||||
|
role: firstWord(rolePart),
|
||||||
|
channel: firstWord(channelPart),
|
||||||
|
recipient: parseRecipient(rawHeader),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
#parseArgs(): Record<string, unknown> {
|
||||||
|
const raw = this.#toolArgs.trim();
|
||||||
|
if (raw.length === 0) return {};
|
||||||
|
try {
|
||||||
|
return asRecord(parseJsonWithRepair<unknown>(raw));
|
||||||
|
} catch {
|
||||||
|
return {};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#clearBody(resetRawBlock = true): void {
|
||||||
|
this.#mode = "skip";
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#toolArgs = "";
|
||||||
|
this.#thinking = "";
|
||||||
|
if (resetRawBlock) this.#rawBlock = "";
|
||||||
|
}
|
||||||
|
|
||||||
|
#resetAll(): void {
|
||||||
|
this.#buffer = "";
|
||||||
|
this.#state = "outside";
|
||||||
|
this.#clearBody();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function findNextToken(text: string, tokens: readonly string[]): TokenMatch | undefined {
|
||||||
|
let match: TokenMatch | undefined;
|
||||||
|
for (const token of tokens) {
|
||||||
|
const index = text.indexOf(token);
|
||||||
|
if (index !== -1 && (!match || index < match.index)) match = { index, token };
|
||||||
|
}
|
||||||
|
return match;
|
||||||
|
}
|
||||||
|
|
||||||
|
function firstWord(text: string): string {
|
||||||
|
const trimmed = text.trimStart();
|
||||||
|
let end = 0;
|
||||||
|
while (end < trimmed.length) {
|
||||||
|
const ch = trimmed[end]!;
|
||||||
|
if (ch === "<" || /\s/.test(ch)) break;
|
||||||
|
end++;
|
||||||
|
}
|
||||||
|
return trimmed.slice(0, end);
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseRecipient(header: string): string {
|
||||||
|
const match = /(?:^|\s)to=([^\s<]+)/.exec(header);
|
||||||
|
return match?.[1] ?? "";
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "harmony",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: () => new HarmonyInbandScanner(),
|
||||||
|
renderAssistantToolCalls: renderHarmonyToolCalls,
|
||||||
|
renderToolResults: renderHarmonyToolResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
Emit each call as a `<tool_call>` block wrapping a single-line JSON object with `name` and `arguments`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_call>
|
||||||
|
{"name":"function_name","arguments":{"arg":"value"}}
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
Results arrive later as `<tool_response>` blocks:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_response>
|
||||||
|
verbatim tool result
|
||||||
|
</tool_response>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- `name` MUST match a listed function; `arguments` is a JSON object, never a stringified JSON.
|
||||||
|
- Emit multiple calls as consecutive `<tool_call>` blocks; keep any prose outside them.
|
||||||
|
- Read each `<tool_response>` in call order. NEVER emit `<tool_response>` yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,170 @@
|
|||||||
|
import { parseJsonWithRepair, parseStreamingJson } from "../utils/json-parse";
|
||||||
|
import { asRecord, mintToolCallId, partialSuffixOverlapAny } from "./coercion";
|
||||||
|
import grammarPrompt from "./hermes.md" with { type: "text" };
|
||||||
|
import { renderHermesToolCalls, renderToolResponseResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner, InbandScannerOptions } from "./types";
|
||||||
|
|
||||||
|
const TOOL_OPEN = "<tool_call>";
|
||||||
|
const TOOL_CLOSE = "</tool_call>";
|
||||||
|
const THINK_OPEN = "<think>";
|
||||||
|
const THINK_CLOSE = "</think>";
|
||||||
|
const HOLD_TAGS = [TOOL_OPEN, TOOL_CLOSE, THINK_OPEN, THINK_CLOSE] as const;
|
||||||
|
|
||||||
|
export class HermesInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#inside = false;
|
||||||
|
#id = "";
|
||||||
|
#name = "";
|
||||||
|
#started = false;
|
||||||
|
#parseThinking: boolean;
|
||||||
|
#inThinking = false;
|
||||||
|
#thinking = "";
|
||||||
|
|
||||||
|
constructor(options: InbandScannerOptions = {}) {
|
||||||
|
this.#parseThinking = options.parseThinking === true;
|
||||||
|
}
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#inThinking) {
|
||||||
|
const closeThink = this.#buffer.indexOf(THINK_CLOSE);
|
||||||
|
if (closeThink === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, [THINK_CLOSE]);
|
||||||
|
const thinking = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (thinking.length > 0) {
|
||||||
|
this.#thinking += thinking;
|
||||||
|
events.push({ type: "thinkingDelta", delta: thinking });
|
||||||
|
}
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
if (final) {
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#inThinking = false;
|
||||||
|
}
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
const thinking = this.#buffer.slice(0, closeThink);
|
||||||
|
if (thinking.length > 0) {
|
||||||
|
this.#thinking += thinking;
|
||||||
|
events.push({ type: "thinkingDelta", delta: thinking });
|
||||||
|
}
|
||||||
|
this.#buffer = this.#buffer.slice(closeThink + THINK_CLOSE.length);
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#inThinking = false;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!this.#inside) {
|
||||||
|
const open = this.#buffer.indexOf(TOOL_OPEN);
|
||||||
|
const think = this.#parseThinking ? this.#buffer.indexOf(THINK_OPEN) : -1;
|
||||||
|
const start = open === -1 ? think : think === -1 ? open : Math.min(open, think);
|
||||||
|
if (start === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, HOLD_TAGS);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (emit.length > 0) events.push({ type: "text", text: emit });
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
if (start > 0) events.push({ type: "text", text: this.#buffer.slice(0, start) });
|
||||||
|
if (start === think) {
|
||||||
|
this.#buffer = this.#buffer.slice(start + THINK_OPEN.length);
|
||||||
|
this.#inThinking = true;
|
||||||
|
this.#thinking = "";
|
||||||
|
events.push({ type: "thinkingStart" });
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
this.#buffer = this.#buffer.slice(start + TOOL_OPEN.length);
|
||||||
|
this.#inside = true;
|
||||||
|
this.#id = mintToolCallId();
|
||||||
|
this.#name = "";
|
||||||
|
this.#started = false;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
const close = this.#buffer.indexOf(TOOL_CLOSE);
|
||||||
|
const body = close === -1 ? this.#buffer : this.#buffer.slice(0, close);
|
||||||
|
if (!this.#started) this.#tryStart(body, events);
|
||||||
|
if (close === -1) {
|
||||||
|
if (final) this.#reset();
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
const parsed = this.#parseCall(body);
|
||||||
|
if (parsed) {
|
||||||
|
if (!this.#started) {
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: parsed.name });
|
||||||
|
this.#started = true;
|
||||||
|
}
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: parsed.name,
|
||||||
|
arguments: parsed.arguments,
|
||||||
|
rawBlock: `${TOOL_OPEN}${body}${TOOL_CLOSE}`,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
this.#buffer = this.#buffer.slice(close + TOOL_CLOSE.length);
|
||||||
|
this.#reset();
|
||||||
|
}
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#tryStart(body: string, events: InbandScanEvent[]): void {
|
||||||
|
try {
|
||||||
|
const partial = parseStreamingJson<{ name?: unknown }>(body);
|
||||||
|
if (typeof partial.name !== "string" || partial.name.length === 0) return;
|
||||||
|
this.#name = partial.name;
|
||||||
|
this.#started = true;
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
} catch {
|
||||||
|
// Partial JSON is allowed until the closing tag arrives.
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#parseCall(body: string): { name: string; arguments: Record<string, unknown> } | undefined {
|
||||||
|
try {
|
||||||
|
const parsed = parseJsonWithRepair<{ name?: unknown; arguments?: unknown }>(body.trim());
|
||||||
|
if (typeof parsed.name !== "string" || parsed.name.length === 0) return undefined;
|
||||||
|
let args = parsed.arguments;
|
||||||
|
if (typeof args === "string") {
|
||||||
|
try {
|
||||||
|
args = parseJsonWithRepair<unknown>(args);
|
||||||
|
} catch {
|
||||||
|
args = {};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return { name: parsed.name, arguments: asRecord(args) };
|
||||||
|
} catch {
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#reset(): void {
|
||||||
|
this.#inside = false;
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#started = false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "hermes",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: options => new HermesInbandScanner(options),
|
||||||
|
renderAssistantToolCalls: renderHermesToolCalls,
|
||||||
|
renderToolResults: renderToolResponseResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,81 @@
|
|||||||
|
import type {
|
||||||
|
AssistantMessage,
|
||||||
|
Context,
|
||||||
|
ImageContent,
|
||||||
|
Message,
|
||||||
|
TextContent,
|
||||||
|
ToolCall,
|
||||||
|
ToolResultMessage,
|
||||||
|
} from "../types";
|
||||||
|
import { getInbandGrammar } from "./factory";
|
||||||
|
import type { Grammar, GrammarToolResult, InbandTool, ToolCallSyntax } from "./types";
|
||||||
|
|
||||||
|
export function encodeInbandToolHistory(
|
||||||
|
messages: Context["messages"],
|
||||||
|
syntax: ToolCallSyntax,
|
||||||
|
tools: readonly InbandTool[] = [],
|
||||||
|
): Context["messages"] {
|
||||||
|
const grammar = getInbandGrammar(syntax);
|
||||||
|
const out: Message[] = [];
|
||||||
|
for (let i = 0; i < messages.length; i++) {
|
||||||
|
const message = messages[i]!;
|
||||||
|
if (message.role === "assistant") {
|
||||||
|
out.push(encodeAssistantMessage(message, grammar, tools));
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (message.role === "toolResult") {
|
||||||
|
const run: ToolResultMessage[] = [];
|
||||||
|
let j = i;
|
||||||
|
while (j < messages.length && messages[j]!.role === "toolResult") {
|
||||||
|
run.push(messages[j] as ToolResultMessage);
|
||||||
|
j++;
|
||||||
|
}
|
||||||
|
out.push(encodeToolResults(run, grammar));
|
||||||
|
i = j - 1;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
out.push(message);
|
||||||
|
}
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
function encodeAssistantMessage(
|
||||||
|
message: AssistantMessage,
|
||||||
|
grammar: Grammar,
|
||||||
|
tools: readonly InbandTool[],
|
||||||
|
): AssistantMessage {
|
||||||
|
const toolCalls = message.content.filter((block): block is ToolCall => block.type === "toolCall");
|
||||||
|
if (toolCalls.length === 0) return message;
|
||||||
|
const prose = message.content
|
||||||
|
.filter((block): block is TextContent => block.type === "text")
|
||||||
|
.map(block => block.text)
|
||||||
|
.join("\n");
|
||||||
|
const rendered = grammar.renderAssistantToolCalls(toolCalls, { tools });
|
||||||
|
const text = prose.trim().length > 0 ? `${prose.trimEnd()}\n${rendered}` : rendered;
|
||||||
|
return { ...message, content: [{ type: "text", text }] };
|
||||||
|
}
|
||||||
|
|
||||||
|
function encodeToolResults(results: readonly ToolResultMessage[], grammar: Grammar): Message {
|
||||||
|
const grammarResults: GrammarToolResult[] = [];
|
||||||
|
const images: ImageContent[] = [];
|
||||||
|
for (let index = 0; index < results.length; index++) {
|
||||||
|
const result = results[index]!;
|
||||||
|
let text = "";
|
||||||
|
for (const block of result.content) {
|
||||||
|
if (block.type === "text") text += block.text;
|
||||||
|
else if (block.type === "image") images.push(block);
|
||||||
|
}
|
||||||
|
grammarResults.push({
|
||||||
|
id: result.toolCallId,
|
||||||
|
name: result.toolName,
|
||||||
|
index,
|
||||||
|
text,
|
||||||
|
isError: result.isError,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
const content: (TextContent | ImageContent)[] = [
|
||||||
|
{ type: "text", text: grammar.renderToolResults(grammarResults) },
|
||||||
|
...images,
|
||||||
|
];
|
||||||
|
return { role: "user", content, timestamp: results[0]?.timestamp ?? Date.now() };
|
||||||
|
}
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
export * from "./catalog";
|
||||||
|
export * from "./coercion";
|
||||||
|
export * from "./factory";
|
||||||
|
export * from "./history";
|
||||||
|
export * from "./owned-stream";
|
||||||
|
export * from "./types";
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
Emit every call of a turn inside one section. Each call is an id of the fixed form `functions.NAME:INDEX` followed by one JSON arguments object:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|tool_calls_section_begin|><|tool_call_begin|>functions.NAME:INDEX<|tool_call_argument_begin|>{"arg":"value"}<|tool_call_end|><|tool_calls_section_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
Tool results arrive later as turns whose body is a `## Return of functions.NAME:INDEX` header then the verbatim result:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<|im_system|>NAME<|im_middle|>## Return of functions.NAME:INDEX
|
||||||
|
verbatim tool result<|im_end|>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- `NAME` MUST match a listed function exactly.
|
||||||
|
- Arguments MUST be one JSON object with double-quoted keys.
|
||||||
|
- Multiple calls = consecutive `<|tool_call_begin|>…<|tool_call_end|>` blocks in the same section; `INDEX` increments from `0`.
|
||||||
|
- This format has no thinking channel; NEVER emit `<think>` tags.
|
||||||
|
- Read each result turn in call order. NEVER emit result turns yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,197 @@
|
|||||||
|
import { parseJsonWithRepair } from "../utils/json-parse";
|
||||||
|
import { asRecord, normalizeKimiFunctionName, partialSuffixOverlapAny } from "./coercion";
|
||||||
|
import grammarPrompt from "./kimi.md" with { type: "text" };
|
||||||
|
import { renderKimiToolCalls, renderKimiToolResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner } from "./types";
|
||||||
|
|
||||||
|
export const KIMI_SECTION_BEGIN = "<|tool_calls_section_begin|>";
|
||||||
|
export const KIMI_SECTION_END = "<|tool_calls_section_end|>";
|
||||||
|
export const KIMI_CALL_BEGIN = "<|tool_call_begin|>";
|
||||||
|
export const KIMI_CALL_END = "<|tool_call_end|>";
|
||||||
|
export const KIMI_ARG_BEGIN = "<|tool_call_argument_begin|>";
|
||||||
|
|
||||||
|
const TOKENS = [KIMI_SECTION_BEGIN, KIMI_SECTION_END, KIMI_CALL_BEGIN, KIMI_CALL_END, KIMI_ARG_BEGIN] as const;
|
||||||
|
|
||||||
|
type State = "outside" | "section" | "header" | "args";
|
||||||
|
|
||||||
|
export class KimiInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#state: State = "outside";
|
||||||
|
#id = "";
|
||||||
|
#name = "";
|
||||||
|
#rawBlock = "";
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#state === "outside") {
|
||||||
|
if (!this.#consumeOutside(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "section") {
|
||||||
|
if (!this.#consumeSection(final)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "header") {
|
||||||
|
if (!this.#consumeHeader(final, events)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!this.#consumeArgs(final, events)) break;
|
||||||
|
}
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeOutside(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const tokenStart = this.#nextTokenIndex();
|
||||||
|
if (tokenStart === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, TOKENS);
|
||||||
|
const emitEnd = this.#buffer.length - hold;
|
||||||
|
if (emitEnd > 0) events.push({ type: "text", text: this.#buffer.slice(0, emitEnd) });
|
||||||
|
this.#buffer = this.#buffer.slice(emitEnd);
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (tokenStart > 0) events.push({ type: "text", text: this.#buffer.slice(0, tokenStart) });
|
||||||
|
this.#buffer = this.#buffer.slice(tokenStart);
|
||||||
|
const token = this.#tokenAtStart();
|
||||||
|
if (!token) return false;
|
||||||
|
this.#buffer = this.#buffer.slice(token.length);
|
||||||
|
if (token === KIMI_SECTION_BEGIN) this.#state = "section";
|
||||||
|
else events.push({ type: "text", text: token });
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeSection(final: boolean): boolean {
|
||||||
|
this.#skipWhitespace();
|
||||||
|
if (this.#buffer.length === 0) return false;
|
||||||
|
|
||||||
|
const token = this.#tokenAtStart();
|
||||||
|
if (token === KIMI_SECTION_END) {
|
||||||
|
this.#buffer = this.#buffer.slice(KIMI_SECTION_END.length);
|
||||||
|
this.#state = "outside";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (token === KIMI_CALL_BEGIN) {
|
||||||
|
this.#buffer = this.#buffer.slice(KIMI_CALL_BEGIN.length);
|
||||||
|
this.#state = "header";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
if (token) {
|
||||||
|
this.#buffer = this.#buffer.slice(token.length);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!final && partialSuffixOverlapAny(this.#buffer, TOKENS) === this.#buffer.length) return false;
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeHeader(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const sep = this.#buffer.indexOf(KIMI_ARG_BEGIN);
|
||||||
|
if (sep === -1) {
|
||||||
|
if (final) this.#dropBufferedCall();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const rawHeader = this.#buffer.slice(0, sep);
|
||||||
|
this.#id = rawHeader.trim();
|
||||||
|
this.#name = normalizeKimiFunctionName(this.#id);
|
||||||
|
this.#rawBlock = `${KIMI_CALL_BEGIN}${rawHeader}${KIMI_ARG_BEGIN}`;
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
this.#buffer = this.#buffer.slice(sep + KIMI_ARG_BEGIN.length);
|
||||||
|
this.#state = "args";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeArgs(final: boolean, events: InbandScanEvent[]): boolean {
|
||||||
|
const end = this.#buffer.indexOf(KIMI_CALL_END);
|
||||||
|
if (end === -1) {
|
||||||
|
if (final) this.#dropBufferedCall();
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const rawArgsBlock = this.#buffer.slice(0, end);
|
||||||
|
const rawArgs = rawArgsBlock.trim();
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: this.#name,
|
||||||
|
arguments: this.#parseArgs(rawArgs),
|
||||||
|
rawBlock: `${this.#rawBlock}${rawArgsBlock}${KIMI_CALL_END}`,
|
||||||
|
});
|
||||||
|
this.#buffer = this.#buffer.slice(end + KIMI_CALL_END.length);
|
||||||
|
this.#resetCall();
|
||||||
|
this.#state = "section";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#parseArgs(rawArgs: string): Record<string, unknown> {
|
||||||
|
if (rawArgs.length === 0) return {};
|
||||||
|
try {
|
||||||
|
return asRecord(parseJsonWithRepair<unknown>(rawArgs));
|
||||||
|
} catch {
|
||||||
|
return {};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#nextTokenIndex(): number {
|
||||||
|
let best = -1;
|
||||||
|
for (const token of TOKENS) {
|
||||||
|
const index = this.#buffer.indexOf(token);
|
||||||
|
if (index !== -1 && (best === -1 || index < best)) best = index;
|
||||||
|
}
|
||||||
|
return best;
|
||||||
|
}
|
||||||
|
|
||||||
|
#tokenAtStart(): string | undefined {
|
||||||
|
for (const token of TOKENS) {
|
||||||
|
if (this.#buffer.startsWith(token)) return token;
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
#skipWhitespace(): void {
|
||||||
|
let i = 0;
|
||||||
|
while (i < this.#buffer.length && isWhitespace(this.#buffer.charCodeAt(i))) i++;
|
||||||
|
if (i > 0) this.#buffer = this.#buffer.slice(i);
|
||||||
|
}
|
||||||
|
|
||||||
|
#dropBufferedCall(): void {
|
||||||
|
this.#buffer = "";
|
||||||
|
this.#resetCall();
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#resetCall(): void {
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#rawBlock = "";
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function isWhitespace(cp: number): boolean {
|
||||||
|
return cp === 0x20 || cp === 0x09 || cp === 0x0a || cp === 0x0d || cp === 0x0b || cp === 0x0c;
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "kimi",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: () => new KimiInbandScanner(),
|
||||||
|
renderAssistantToolCalls: renderKimiToolCalls,
|
||||||
|
renderToolResults: renderKimiToolResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,329 @@
|
|||||||
|
import type {
|
||||||
|
AssistantMessage,
|
||||||
|
AssistantMessageEventStream as AssistantMessageEventStreamType,
|
||||||
|
TextContent,
|
||||||
|
ThinkingContent,
|
||||||
|
ToolCall,
|
||||||
|
} from "../types";
|
||||||
|
import { AssistantMessageEventStream } from "../utils/event-stream";
|
||||||
|
import { buildStringArgsResolver } from "./coercion";
|
||||||
|
import { createInbandScanner } from "./factory";
|
||||||
|
import type { InbandScanEvent, InbandScanner, InbandTool, ToolCallSyntax } from "./types";
|
||||||
|
|
||||||
|
const RESPONSE_OPEN_TOKENS: Record<ToolCallSyntax, readonly string[]> = {
|
||||||
|
glm: ["<tool_response>"],
|
||||||
|
hermes: ["<tool_response>"],
|
||||||
|
kimi: ["<|im_system|>"],
|
||||||
|
xml: ["<tool_response>"],
|
||||||
|
anthropic: ["<function_results>", "<tool_response>"],
|
||||||
|
deepseek: ["<|tool▁outputs▁begin|>", "<|tool▁output▁begin|>"],
|
||||||
|
harmony: ["<|start|>functions."],
|
||||||
|
pi: ["<tool_response>"],
|
||||||
|
qwen3: ["<tool_response>"],
|
||||||
|
};
|
||||||
|
|
||||||
|
function firstTokenIndex(text: string, tokens: readonly string[]): number {
|
||||||
|
let best = -1;
|
||||||
|
for (const token of tokens) {
|
||||||
|
const index = text.indexOf(token);
|
||||||
|
if (index !== -1 && (best === -1 || index < best)) best = index;
|
||||||
|
}
|
||||||
|
return best;
|
||||||
|
}
|
||||||
|
|
||||||
|
type OpenText = { index: number } | undefined;
|
||||||
|
type OpenThinking = { index: number; text: string } | undefined;
|
||||||
|
|
||||||
|
export function parseInbandToolMessage(
|
||||||
|
message: AssistantMessage,
|
||||||
|
syntax: ToolCallSyntax,
|
||||||
|
tools: readonly InbandTool[],
|
||||||
|
): AssistantMessage {
|
||||||
|
const projector = new InbandStreamProjector(new AssistantMessageEventStream(), tools, syntax, message, false);
|
||||||
|
for (const block of message.content) {
|
||||||
|
if (block.type === "text") projector.text(block.text);
|
||||||
|
else projector.keep(block);
|
||||||
|
}
|
||||||
|
return projector.finish(message, false);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function wrapInbandToolStream(
|
||||||
|
inner: AssistantMessageEventStreamType,
|
||||||
|
tools: readonly InbandTool[],
|
||||||
|
syntax: ToolCallSyntax,
|
||||||
|
onAbort?: () => void,
|
||||||
|
): AssistantMessageEventStreamType {
|
||||||
|
const out = new AssistantMessageEventStream();
|
||||||
|
void (async () => {
|
||||||
|
try {
|
||||||
|
let projector: InbandStreamProjector | undefined;
|
||||||
|
for await (const event of inner) {
|
||||||
|
switch (event.type) {
|
||||||
|
case "start":
|
||||||
|
projector = new InbandStreamProjector(out, tools, syntax, event.partial, true);
|
||||||
|
break;
|
||||||
|
case "thinking_start":
|
||||||
|
projector?.thinkingStart();
|
||||||
|
break;
|
||||||
|
case "thinking_delta":
|
||||||
|
projector?.thinkingDelta(event.delta);
|
||||||
|
break;
|
||||||
|
case "thinking_end":
|
||||||
|
projector?.thinkingEnd();
|
||||||
|
break;
|
||||||
|
case "text_delta":
|
||||||
|
if (projector?.text(event.delta)) {
|
||||||
|
projector.finish(event.partial, true);
|
||||||
|
onAbort?.();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
break;
|
||||||
|
case "done":
|
||||||
|
projector ??= new InbandStreamProjector(out, tools, syntax, event.message, true);
|
||||||
|
projector.finish(event.message, true);
|
||||||
|
return;
|
||||||
|
case "error":
|
||||||
|
out.push(event);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
} catch (err) {
|
||||||
|
out.fail(err);
|
||||||
|
}
|
||||||
|
})();
|
||||||
|
return out;
|
||||||
|
}
|
||||||
|
|
||||||
|
class InbandStreamProjector {
|
||||||
|
readonly #out: AssistantMessageEventStream;
|
||||||
|
readonly #scanner: InbandScanner;
|
||||||
|
readonly #emitEvents: boolean;
|
||||||
|
readonly #responseOpenTokens: readonly string[];
|
||||||
|
readonly #responseOverlapLength: number;
|
||||||
|
#partial: AssistantMessage;
|
||||||
|
#text: OpenText;
|
||||||
|
#thinking: OpenThinking;
|
||||||
|
#toolBlocks = new Map<string, { index: number; block: ToolCall; currentKey?: string; rawValue: string }>();
|
||||||
|
#fedLen = 0;
|
||||||
|
#stopped = false;
|
||||||
|
#responsePending = "";
|
||||||
|
|
||||||
|
constructor(
|
||||||
|
out: AssistantMessageEventStream,
|
||||||
|
tools: readonly InbandTool[],
|
||||||
|
syntax: ToolCallSyntax,
|
||||||
|
seed: AssistantMessage,
|
||||||
|
emitEvents: boolean,
|
||||||
|
) {
|
||||||
|
this.#out = out;
|
||||||
|
this.#emitEvents = emitEvents;
|
||||||
|
this.#scanner = createInbandScanner(syntax, {
|
||||||
|
tools,
|
||||||
|
stringArgs: buildStringArgsResolver(tools),
|
||||||
|
parseThinking: true,
|
||||||
|
});
|
||||||
|
this.#responseOpenTokens = RESPONSE_OPEN_TOKENS[syntax];
|
||||||
|
this.#responseOverlapLength = Math.max(0, ...this.#responseOpenTokens.map(token => token.length - 1));
|
||||||
|
this.#partial = { ...seed, content: [] };
|
||||||
|
if (emitEvents) this.#out.push({ type: "start", partial: this.#partial });
|
||||||
|
}
|
||||||
|
|
||||||
|
keep(block: AssistantMessage["content"][number]): void {
|
||||||
|
this.#closeText();
|
||||||
|
this.#closeThinking();
|
||||||
|
this.#partial.content.push(block);
|
||||||
|
}
|
||||||
|
|
||||||
|
text(delta: string): boolean {
|
||||||
|
if (this.#stopped) return true;
|
||||||
|
this.#fedLen += delta.length;
|
||||||
|
const combined = this.#responsePending + delta;
|
||||||
|
const responseIndex = firstTokenIndex(combined, this.#responseOpenTokens);
|
||||||
|
if (responseIndex !== -1) {
|
||||||
|
this.#responsePending = "";
|
||||||
|
this.#apply(this.#scanner.feed(combined.slice(0, responseIndex)));
|
||||||
|
this.#stopped = true;
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (combined.length <= this.#responseOverlapLength) {
|
||||||
|
this.#responsePending = combined;
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const emitLength = combined.length - this.#responseOverlapLength;
|
||||||
|
this.#responsePending = combined.slice(emitLength);
|
||||||
|
this.#apply(this.#scanner.feed(combined.slice(0, emitLength)));
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
thinkingStart(): void {
|
||||||
|
this.#closeText();
|
||||||
|
if (this.#thinking) return;
|
||||||
|
const block: ThinkingContent = { type: "thinking", thinking: "" };
|
||||||
|
this.#partial.content.push(block);
|
||||||
|
this.#thinking = { index: this.#partial.content.length - 1, text: "" };
|
||||||
|
if (this.#emitEvents)
|
||||||
|
this.#out.push({ type: "thinking_start", contentIndex: this.#thinking.index, partial: this.#partial });
|
||||||
|
}
|
||||||
|
|
||||||
|
thinkingDelta(delta: string): void {
|
||||||
|
if (!this.#thinking) this.thinkingStart();
|
||||||
|
const thinking = this.#thinking;
|
||||||
|
if (!thinking) return;
|
||||||
|
const block = this.#partial.content[thinking.index] as ThinkingContent;
|
||||||
|
block.thinking += delta;
|
||||||
|
thinking.text += delta;
|
||||||
|
if (this.#emitEvents)
|
||||||
|
this.#out.push({ type: "thinking_delta", contentIndex: thinking.index, delta, partial: this.#partial });
|
||||||
|
}
|
||||||
|
|
||||||
|
thinkingEnd(): void {
|
||||||
|
this.#closeThinking();
|
||||||
|
}
|
||||||
|
|
||||||
|
finish(message: AssistantMessage, emitDone: boolean): AssistantMessage {
|
||||||
|
let fullText = "";
|
||||||
|
for (const block of message.content) if (block.type === "text") fullText += block.text;
|
||||||
|
if (!this.#stopped && fullText.length > this.#fedLen) this.text(fullText.slice(this.#fedLen));
|
||||||
|
if (!this.#stopped && this.#responsePending.length > 0) {
|
||||||
|
this.#apply(this.#scanner.feed(this.#responsePending));
|
||||||
|
this.#responsePending = "";
|
||||||
|
}
|
||||||
|
this.#apply(this.#scanner.flush());
|
||||||
|
this.#closeText();
|
||||||
|
this.#closeThinking();
|
||||||
|
const hasTools = this.#partial.content.some(block => block.type === "toolCall");
|
||||||
|
const reason =
|
||||||
|
hasTools && message.stopReason !== "length" ? "toolUse" : message.stopReason === "length" ? "length" : "stop";
|
||||||
|
const finalMessage: AssistantMessage = { ...message, content: this.#partial.content, stopReason: reason };
|
||||||
|
if (emitDone) this.#out.push({ type: "done", reason, message: finalMessage });
|
||||||
|
return finalMessage;
|
||||||
|
}
|
||||||
|
|
||||||
|
#apply(events: InbandScanEvent[]): void {
|
||||||
|
for (const event of events) {
|
||||||
|
switch (event.type) {
|
||||||
|
case "text":
|
||||||
|
this.#emitText(event.text);
|
||||||
|
break;
|
||||||
|
case "thinkingStart":
|
||||||
|
this.thinkingStart();
|
||||||
|
break;
|
||||||
|
case "thinkingDelta":
|
||||||
|
this.thinkingDelta(event.delta);
|
||||||
|
break;
|
||||||
|
case "thinkingEnd":
|
||||||
|
this.thinkingEnd();
|
||||||
|
break;
|
||||||
|
case "toolStart":
|
||||||
|
this.#beginTool(event);
|
||||||
|
break;
|
||||||
|
case "toolArgDelta":
|
||||||
|
this.#deltaTool(event);
|
||||||
|
break;
|
||||||
|
case "toolEnd":
|
||||||
|
this.#endTool(event);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitText(text: string): void {
|
||||||
|
if (text.length === 0) return;
|
||||||
|
this.#closeThinking();
|
||||||
|
if (!this.#text) {
|
||||||
|
this.#partial.content.push({ type: "text", text: "" });
|
||||||
|
this.#text = { index: this.#partial.content.length - 1 };
|
||||||
|
if (this.#emitEvents)
|
||||||
|
this.#out.push({ type: "text_start", contentIndex: this.#text.index, partial: this.#partial });
|
||||||
|
}
|
||||||
|
const block = this.#partial.content[this.#text.index] as TextContent;
|
||||||
|
block.text += text;
|
||||||
|
if (this.#emitEvents)
|
||||||
|
this.#out.push({ type: "text_delta", contentIndex: this.#text.index, delta: text, partial: this.#partial });
|
||||||
|
}
|
||||||
|
|
||||||
|
#closeText(): void {
|
||||||
|
if (!this.#text) return;
|
||||||
|
const block = this.#partial.content[this.#text.index] as TextContent;
|
||||||
|
if (this.#emitEvents) {
|
||||||
|
this.#out.push({
|
||||||
|
type: "text_end",
|
||||||
|
contentIndex: this.#text.index,
|
||||||
|
content: block.text,
|
||||||
|
partial: this.#partial,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
this.#text = undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
#closeThinking(): void {
|
||||||
|
if (!this.#thinking) return;
|
||||||
|
const block = this.#partial.content[this.#thinking.index] as ThinkingContent;
|
||||||
|
if (this.#emitEvents) {
|
||||||
|
this.#out.push({
|
||||||
|
type: "thinking_end",
|
||||||
|
contentIndex: this.#thinking.index,
|
||||||
|
content: block.thinking,
|
||||||
|
partial: this.#partial,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
this.#thinking = undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
#beginTool(event: Extract<InbandScanEvent, { type: "toolStart" }>): void {
|
||||||
|
this.#closeText();
|
||||||
|
this.#closeThinking();
|
||||||
|
if (this.#toolBlocks.has(event.id)) return;
|
||||||
|
const block: ToolCall = { type: "toolCall", id: event.id, name: event.name, arguments: {} };
|
||||||
|
this.#partial.content.push(block);
|
||||||
|
const entry = { index: this.#partial.content.length - 1, block, rawValue: "" };
|
||||||
|
this.#toolBlocks.set(event.id, entry);
|
||||||
|
if (this.#emitEvents)
|
||||||
|
this.#out.push({ type: "toolcall_start", contentIndex: entry.index, partial: this.#partial });
|
||||||
|
}
|
||||||
|
|
||||||
|
#deltaTool(event: Extract<InbandScanEvent, { type: "toolArgDelta" }>): void {
|
||||||
|
let entry = this.#toolBlocks.get(event.id);
|
||||||
|
if (!entry) {
|
||||||
|
this.#beginTool({ type: "toolStart", id: event.id, name: event.name });
|
||||||
|
entry = this.#toolBlocks.get(event.id);
|
||||||
|
}
|
||||||
|
if (!entry) return;
|
||||||
|
if (entry.currentKey !== event.key) {
|
||||||
|
entry.currentKey = event.key;
|
||||||
|
entry.rawValue =
|
||||||
|
typeof entry.block.arguments[event.key] === "string" ? String(entry.block.arguments[event.key]) : "";
|
||||||
|
}
|
||||||
|
entry.rawValue += event.delta;
|
||||||
|
entry.block.arguments[event.key] = entry.rawValue;
|
||||||
|
if (this.#emitEvents)
|
||||||
|
this.#out.push({
|
||||||
|
type: "toolcall_delta",
|
||||||
|
contentIndex: entry.index,
|
||||||
|
delta: event.delta,
|
||||||
|
partial: this.#partial,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
#endTool(event: Extract<InbandScanEvent, { type: "toolEnd" }>): void {
|
||||||
|
let entry = this.#toolBlocks.get(event.id);
|
||||||
|
if (!entry) {
|
||||||
|
this.#beginTool({ type: "toolStart", id: event.id, name: event.name });
|
||||||
|
entry = this.#toolBlocks.get(event.id);
|
||||||
|
}
|
||||||
|
if (!entry) return;
|
||||||
|
entry.block.name = event.name;
|
||||||
|
entry.block.arguments = event.arguments;
|
||||||
|
if (event.rawBlock !== undefined) entry.block.rawBlock = event.rawBlock;
|
||||||
|
if (this.#emitEvents)
|
||||||
|
this.#out.push({
|
||||||
|
type: "toolcall_end",
|
||||||
|
contentIndex: entry.index,
|
||||||
|
toolCall: entry.block,
|
||||||
|
partial: this.#partial,
|
||||||
|
});
|
||||||
|
this.#toolBlocks.delete(event.id);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
A tool call is a `<call:NAME>…</call:NAME>` block (or self-closing `<call:NAME …/>`) written as plain assistant text; arguments are given as tag attributes, child elements, or a verbatim inline body.
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:read path="src/a.ts" offset=50/>
|
||||||
|
```
|
||||||
|
|
||||||
|
Objects and arrays use child elements, repeating an element for each array item:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:configure>
|
||||||
|
<object>
|
||||||
|
<y>4</y>
|
||||||
|
<list>alpha</list>
|
||||||
|
<list>beta</list>
|
||||||
|
</object>
|
||||||
|
</call:configure>
|
||||||
|
```
|
||||||
|
|
||||||
|
A single string argument can fill the body directly:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<call:edit>
|
||||||
|
*** Begin Patch
|
||||||
|
...
|
||||||
|
*** End Patch
|
||||||
|
</call:edit>
|
||||||
|
```
|
||||||
|
|
||||||
|
Tool results arrive as response blocks, read in call order:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_response>
|
||||||
|
verbatim tool result
|
||||||
|
</tool_response>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- `NAME` must match a listed function; never wrap calls in JSON or fences.
|
||||||
|
- Use attributes only for top-level scalars; put objects, arrays, and long strings in child elements.
|
||||||
|
- Strings are verbatim (no quotes, no entity escaping); numbers, booleans, and null are JSON literals.
|
||||||
|
- An object opens a child block whose scalar subfields may also be attributes; an array repeats its element once per item.
|
||||||
|
- The inline body fills the first unset string-typed parameter and may contain any raw text except `</call:NAME>`.
|
||||||
|
- Emit parallel calls as consecutive blocks. NEVER invent call ids; results are positional.
|
||||||
|
- This format defines no thinking channel; never emit `<think>`.
|
||||||
|
- Read each `<tool_response>` in call order. NEVER emit `<tool_response>` yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,584 @@
|
|||||||
|
import type { ToolArgShape } from "./coercion";
|
||||||
|
import {
|
||||||
|
buildArgShapes,
|
||||||
|
coerceValue,
|
||||||
|
collectSchemaTypes,
|
||||||
|
getArrayItemSchema,
|
||||||
|
getObjectProperties,
|
||||||
|
isArraySchema,
|
||||||
|
isObjectSchema,
|
||||||
|
isStringOnlySchema,
|
||||||
|
mintToolCallId,
|
||||||
|
partialSuffixOverlapAny,
|
||||||
|
} from "./coercion";
|
||||||
|
import grammarPrompt from "./pi.md" with { type: "text" };
|
||||||
|
import { renderPiNativeToolCalls, renderToolResponseResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner, InbandScannerOptions } from "./types";
|
||||||
|
|
||||||
|
const CALL_PREFIX = "<call:";
|
||||||
|
const NAME_START = /[A-Za-z_]/;
|
||||||
|
const NAME_CHAR = /[A-Za-z0-9_-]/;
|
||||||
|
const EMPTY_STRING_ARGS: ReadonlySet<string> = new Set<string>();
|
||||||
|
|
||||||
|
type ScannerState = "outside" | "body";
|
||||||
|
type BodyMode = "undecided" | "inline" | "members";
|
||||||
|
|
||||||
|
type RawAttribute = { name: string; value: string | true };
|
||||||
|
|
||||||
|
type OpenTag = {
|
||||||
|
name: string;
|
||||||
|
rawAttrs: RawAttribute[];
|
||||||
|
selfClosing: boolean;
|
||||||
|
end: number;
|
||||||
|
};
|
||||||
|
|
||||||
|
type MembersResult = {
|
||||||
|
ok: boolean;
|
||||||
|
value: Record<string, unknown>;
|
||||||
|
next: number;
|
||||||
|
};
|
||||||
|
|
||||||
|
type ValueResult = {
|
||||||
|
ok: boolean;
|
||||||
|
value: unknown;
|
||||||
|
next: number;
|
||||||
|
};
|
||||||
|
|
||||||
|
export class PiNativeInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#state: ScannerState = "outside";
|
||||||
|
#bodyMode: BodyMode = "undecided";
|
||||||
|
#id = "";
|
||||||
|
#name = "";
|
||||||
|
#args: Record<string, unknown> = {};
|
||||||
|
#inlineKey = "";
|
||||||
|
#inlineValue = "";
|
||||||
|
#inlineLeading = false;
|
||||||
|
#rawBlock = "";
|
||||||
|
readonly #argShapes: Map<string, ToolArgShape>;
|
||||||
|
readonly #stringArgs: (toolName: string) => ReadonlySet<string>;
|
||||||
|
|
||||||
|
constructor(options: InbandScannerOptions = {}) {
|
||||||
|
this.#argShapes = buildArgShapes(options.tools);
|
||||||
|
this.#stringArgs =
|
||||||
|
options.stringArgs ?? (toolName => this.#argShapes.get(toolName)?.stringArgs ?? EMPTY_STRING_ARGS);
|
||||||
|
}
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#state === "outside") {
|
||||||
|
if (!this.#consumeOutside(events, final)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#bodyMode === "undecided") {
|
||||||
|
const mode = this.#classifyBody(final);
|
||||||
|
if (!mode) break;
|
||||||
|
this.#bodyMode = mode;
|
||||||
|
if (mode === "inline") {
|
||||||
|
this.#inlineKey = this.#inlineTargetKey() ?? "input";
|
||||||
|
this.#inlineValue = "";
|
||||||
|
this.#inlineLeading = true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#bodyMode === "inline") {
|
||||||
|
if (!this.#consumeInline(events, final)) break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!this.#consumeMembers(events, final)) break;
|
||||||
|
}
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeOutside(events: InbandScanEvent[], final: boolean): boolean {
|
||||||
|
const open = this.#buffer.indexOf(CALL_PREFIX);
|
||||||
|
if (open === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, [CALL_PREFIX]);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (emit.length > 0) events.push({ type: "text", text: emit });
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (open > 0) {
|
||||||
|
events.push({ type: "text", text: this.#buffer.slice(0, open) });
|
||||||
|
this.#buffer = this.#buffer.slice(open);
|
||||||
|
}
|
||||||
|
|
||||||
|
const tagEnd = findTagEnd(this.#buffer, 0);
|
||||||
|
if (tagEnd === -1) {
|
||||||
|
if (final) this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = parseCallOpenTag(this.#buffer);
|
||||||
|
if (!tag) {
|
||||||
|
events.push({ type: "text", text: this.#buffer[0] ?? "" });
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#beginCall(tag, events);
|
||||||
|
this.#buffer = this.#buffer.slice(tag.end);
|
||||||
|
if (tag.selfClosing) {
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: this.#name,
|
||||||
|
arguments: this.#args,
|
||||||
|
rawBlock: this.#rawBlock,
|
||||||
|
});
|
||||||
|
this.#reset();
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#state = "body";
|
||||||
|
this.#bodyMode = "undecided";
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#beginCall(tag: OpenTag, events: InbandScanEvent[]): void {
|
||||||
|
this.#id = mintToolCallId();
|
||||||
|
this.#name = tag.name;
|
||||||
|
this.#args = coerceAttributes(tag.rawAttrs, this.#shape()?.properties ?? {});
|
||||||
|
this.#inlineKey = "";
|
||||||
|
this.#inlineValue = "";
|
||||||
|
this.#inlineLeading = false;
|
||||||
|
this.#rawBlock = this.#buffer.slice(0, tag.end);
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
}
|
||||||
|
|
||||||
|
#classifyBody(final: boolean): BodyMode | undefined {
|
||||||
|
const closeTag = this.#closeTag();
|
||||||
|
const first = skipWhitespace(this.#buffer, 0);
|
||||||
|
const close = this.#buffer.indexOf(closeTag);
|
||||||
|
const inlineKey = this.#inlineTargetKey();
|
||||||
|
|
||||||
|
if (close !== -1 && first >= close) return inlineKey ? "inline" : "members";
|
||||||
|
if (first >= this.#buffer.length) return undefined;
|
||||||
|
|
||||||
|
const fromFirst = this.#buffer.slice(first);
|
||||||
|
if (!final && closeTag.startsWith(fromFirst)) return undefined;
|
||||||
|
|
||||||
|
if (this.#buffer[first] !== "<") return inlineKey ? "inline" : "members";
|
||||||
|
if (this.#buffer.startsWith(closeTag, first)) return inlineKey ? "inline" : "members";
|
||||||
|
if (this.#buffer.startsWith("</", first)) return "members";
|
||||||
|
|
||||||
|
const elementName = readElementNamePrefix(this.#buffer, first);
|
||||||
|
if (elementName === undefined) return final ? (inlineKey ? "inline" : "members") : undefined;
|
||||||
|
if (elementName.length === 0) return inlineKey ? "inline" : "members";
|
||||||
|
|
||||||
|
const shape = this.#shape();
|
||||||
|
if (!shape) return "members";
|
||||||
|
return Object.hasOwn(shape.properties, elementName) ? "members" : inlineKey ? "inline" : "members";
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeInline(events: InbandScanEvent[], final: boolean): boolean {
|
||||||
|
this.#stripInlineLeadingDelimiter(final);
|
||||||
|
const closeTag = this.#closeTag();
|
||||||
|
const close = this.#buffer.indexOf(closeTag);
|
||||||
|
if (close === -1) {
|
||||||
|
if (final) {
|
||||||
|
this.#reset();
|
||||||
|
this.#buffer = "";
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
const overlap = partialSuffixOverlapAny(this.#buffer, [closeTag]);
|
||||||
|
let hold = Math.max(1, overlap);
|
||||||
|
if (overlap > 0) {
|
||||||
|
const beforeOverlap = this.#buffer.length - overlap - 1;
|
||||||
|
if (this.#buffer[beforeOverlap] === "\n") {
|
||||||
|
hold = Math.max(hold, overlap + 1);
|
||||||
|
if (this.#buffer[beforeOverlap - 1] === "\r") hold = Math.max(hold, overlap + 2);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
const emitLength = this.#buffer.length - hold;
|
||||||
|
if (emitLength > 0) {
|
||||||
|
const delta = this.#buffer.slice(0, emitLength);
|
||||||
|
this.#rawBlock += delta;
|
||||||
|
this.#emitInlineDelta(delta, events);
|
||||||
|
this.#buffer = this.#buffer.slice(emitLength);
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const rawDelta = this.#buffer.slice(0, close);
|
||||||
|
this.#rawBlock += rawDelta + closeTag;
|
||||||
|
let delta = rawDelta;
|
||||||
|
if (delta.endsWith("\r\n")) delta = delta.slice(0, -2);
|
||||||
|
else if (delta.endsWith("\n")) delta = delta.slice(0, -1);
|
||||||
|
this.#emitInlineDelta(delta, events);
|
||||||
|
this.#args[this.#inlineKey] = this.#inlineValue;
|
||||||
|
events.push({ type: "toolEnd", id: this.#id, name: this.#name, arguments: this.#args, rawBlock: this.#rawBlock });
|
||||||
|
this.#buffer = this.#buffer.slice(close + closeTag.length);
|
||||||
|
this.#reset();
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeMembers(events: InbandScanEvent[], final: boolean): boolean {
|
||||||
|
const closeTag = this.#closeTag();
|
||||||
|
let searchFrom = 0;
|
||||||
|
while (true) {
|
||||||
|
const close = this.#buffer.indexOf(closeTag, searchFrom);
|
||||||
|
if (close === -1) {
|
||||||
|
if (final) {
|
||||||
|
this.#reset();
|
||||||
|
this.#buffer = "";
|
||||||
|
}
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
|
||||||
|
const body = this.#buffer.slice(0, close);
|
||||||
|
const parsed = parseMembers(body, 0, undefined, this.#shape()?.properties ?? {});
|
||||||
|
if (!parsed.ok || skipWhitespace(body, parsed.next) !== body.length) {
|
||||||
|
searchFrom = close + closeTag.length;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
const bodyArgs = parsed.value;
|
||||||
|
const args = { ...this.#args, ...bodyArgs };
|
||||||
|
this.#rawBlock += body + closeTag;
|
||||||
|
this.#emitCompletedStringDeltas(bodyArgs, events);
|
||||||
|
events.push({ type: "toolEnd", id: this.#id, name: this.#name, arguments: args, rawBlock: this.#rawBlock });
|
||||||
|
this.#buffer = this.#buffer.slice(close + closeTag.length);
|
||||||
|
this.#reset();
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitCompletedStringDeltas(args: Record<string, unknown>, events: InbandScanEvent[]): void {
|
||||||
|
for (const key in args) {
|
||||||
|
const value = args[key];
|
||||||
|
if (typeof value === "string" && value.length > 0) {
|
||||||
|
events.push({ type: "toolArgDelta", id: this.#id, name: this.#name, key, delta: value });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#stripInlineLeadingDelimiter(final: boolean): void {
|
||||||
|
if (!this.#inlineLeading) return;
|
||||||
|
if (this.#buffer.length === 0) return;
|
||||||
|
if (this.#buffer[0] === "\r") {
|
||||||
|
if (this.#buffer.length === 1 && !final) return;
|
||||||
|
if (this.#buffer[1] === "\n") {
|
||||||
|
this.#rawBlock += this.#buffer.slice(0, 2);
|
||||||
|
this.#buffer = this.#buffer.slice(2);
|
||||||
|
}
|
||||||
|
this.#inlineLeading = false;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (this.#buffer[0] === "\n") {
|
||||||
|
this.#rawBlock += this.#buffer[0];
|
||||||
|
this.#buffer = this.#buffer.slice(1);
|
||||||
|
}
|
||||||
|
this.#inlineLeading = false;
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitInlineDelta(delta: string, events: InbandScanEvent[]): void {
|
||||||
|
if (delta.length === 0) return;
|
||||||
|
this.#inlineValue += delta;
|
||||||
|
events.push({ type: "toolArgDelta", id: this.#id, name: this.#name, key: this.#inlineKey, delta });
|
||||||
|
}
|
||||||
|
|
||||||
|
#inlineTargetKey(): string | undefined {
|
||||||
|
const shape = this.#shape();
|
||||||
|
if (shape) {
|
||||||
|
for (const key of shape.parameterOrder) {
|
||||||
|
if (Object.hasOwn(this.#args, key)) continue;
|
||||||
|
return isStringOnlySchema(shape.properties[key]) ? key : undefined;
|
||||||
|
}
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
|
||||||
|
for (const key of this.#stringArgs(this.#name)) {
|
||||||
|
if (!Object.hasOwn(this.#args, key)) return key;
|
||||||
|
}
|
||||||
|
return "input";
|
||||||
|
}
|
||||||
|
|
||||||
|
#shape(): ToolArgShape | undefined {
|
||||||
|
return this.#argShapes.get(this.#name);
|
||||||
|
}
|
||||||
|
|
||||||
|
#closeTag(): string {
|
||||||
|
return `</call:${this.#name}>`;
|
||||||
|
}
|
||||||
|
|
||||||
|
#reset(): void {
|
||||||
|
this.#state = "outside";
|
||||||
|
this.#bodyMode = "undecided";
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#args = {};
|
||||||
|
this.#inlineKey = "";
|
||||||
|
this.#inlineValue = "";
|
||||||
|
this.#inlineLeading = false;
|
||||||
|
this.#rawBlock = "";
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseMembers(
|
||||||
|
text: string,
|
||||||
|
position: number,
|
||||||
|
endTag: string | undefined,
|
||||||
|
properties: Record<string, unknown>,
|
||||||
|
): MembersResult {
|
||||||
|
const value: Record<string, unknown> = {};
|
||||||
|
let index = position;
|
||||||
|
while (index < text.length) {
|
||||||
|
index = skipWhitespace(text, index);
|
||||||
|
if (endTag && text.startsWith(endTag, index)) return { ok: true, value, next: index + endTag.length };
|
||||||
|
if (index >= text.length) break;
|
||||||
|
if (text[index] !== "<" || text.startsWith("</", index) || text.startsWith(CALL_PREFIX, index)) {
|
||||||
|
return { ok: false, value, next: index };
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = parseElementOpenTag(text, index);
|
||||||
|
if (!tag) return { ok: false, value, next: index };
|
||||||
|
const propertySchema = properties[tag.name];
|
||||||
|
const schemaArray = isArraySchema(propertySchema);
|
||||||
|
const itemSchema = schemaArray ? getArrayItemSchema(propertySchema) : propertySchema;
|
||||||
|
const parsed = parseElementValue(text, tag, itemSchema);
|
||||||
|
if (!parsed.ok) return { ok: false, value, next: index };
|
||||||
|
addMember(value, tag.name, parsed.value, schemaArray);
|
||||||
|
index = parsed.next;
|
||||||
|
}
|
||||||
|
|
||||||
|
return endTag ? { ok: false, value, next: index } : { ok: true, value, next: index };
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseElementValue(text: string, tag: OpenTag, schema: unknown): ValueResult {
|
||||||
|
const attrProperties = getObjectProperties(schema);
|
||||||
|
const attrs = coerceAttributes(tag.rawAttrs, attrProperties);
|
||||||
|
if (tag.selfClosing) {
|
||||||
|
if (isObjectSchema(schema) || tag.rawAttrs.length > 0) return { ok: true, value: attrs, next: tag.end };
|
||||||
|
return { ok: true, value: coerceValue("", schema), next: tag.end };
|
||||||
|
}
|
||||||
|
|
||||||
|
const bodyStart = tag.end;
|
||||||
|
const closeTag = `</${tag.name}>`;
|
||||||
|
if (shouldParseObjectBody(text, bodyStart, closeTag, schema, tag.rawAttrs.length > 0)) {
|
||||||
|
const parsed = parseMembers(text, bodyStart, closeTag, attrProperties);
|
||||||
|
if (!parsed.ok) return { ok: false, value: undefined, next: bodyStart };
|
||||||
|
return { ok: true, value: { ...attrs, ...parsed.value }, next: parsed.next };
|
||||||
|
}
|
||||||
|
|
||||||
|
const close = text.indexOf(closeTag, bodyStart);
|
||||||
|
if (close === -1) return { ok: false, value: undefined, next: bodyStart };
|
||||||
|
const raw = stripBlockDelimiters(text.slice(bodyStart, close));
|
||||||
|
return { ok: true, value: coerceValue(raw, schema), next: close + closeTag.length };
|
||||||
|
}
|
||||||
|
|
||||||
|
function shouldParseObjectBody(
|
||||||
|
text: string,
|
||||||
|
bodyStart: number,
|
||||||
|
closeTag: string,
|
||||||
|
schema: unknown,
|
||||||
|
hasAttrs: boolean,
|
||||||
|
): boolean {
|
||||||
|
if (isObjectSchema(schema)) return true;
|
||||||
|
if (isTypedScalarSchema(schema)) return false;
|
||||||
|
if (hasAttrs) return true;
|
||||||
|
const first = skipWhitespace(text, bodyStart);
|
||||||
|
if (text.startsWith(closeTag, first)) return false;
|
||||||
|
return text[first] === "<" && !text.startsWith("</", first) && !text.startsWith(CALL_PREFIX, first);
|
||||||
|
}
|
||||||
|
|
||||||
|
function isTypedScalarSchema(schema: unknown): boolean {
|
||||||
|
const types = collectSchemaTypes(schema);
|
||||||
|
if (types.size === 0) return false;
|
||||||
|
return !types.has("object") && !types.has("array");
|
||||||
|
}
|
||||||
|
|
||||||
|
function addMember(target: Record<string, unknown>, key: string, value: unknown, schemaArray: boolean): void {
|
||||||
|
if (schemaArray) {
|
||||||
|
const existing = target[key];
|
||||||
|
if (Array.isArray(existing)) existing.push(value);
|
||||||
|
else target[key] = [value];
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!Object.hasOwn(target, key)) {
|
||||||
|
target[key] = value;
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const existing = target[key];
|
||||||
|
if (Array.isArray(existing)) existing.push(value);
|
||||||
|
else target[key] = [existing, value];
|
||||||
|
}
|
||||||
|
|
||||||
|
function coerceAttributes(
|
||||||
|
rawAttrs: readonly RawAttribute[],
|
||||||
|
properties: Record<string, unknown>,
|
||||||
|
): Record<string, unknown> {
|
||||||
|
const attrs: Record<string, unknown> = {};
|
||||||
|
for (const attr of rawAttrs) {
|
||||||
|
attrs[attr.name] = attr.value === true ? true : coerceValue(attr.value, properties[attr.name]);
|
||||||
|
}
|
||||||
|
return attrs;
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseCallOpenTag(text: string): OpenTag | undefined {
|
||||||
|
if (!text.startsWith(CALL_PREFIX)) return undefined;
|
||||||
|
const tagEnd = findTagEnd(text, 0);
|
||||||
|
if (tagEnd === -1) return undefined;
|
||||||
|
return parseOpenTagContent(text, CALL_PREFIX.length, tagEnd);
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseElementOpenTag(text: string, start: number): OpenTag | undefined {
|
||||||
|
if (text[start] !== "<" || text.startsWith("</", start) || text.startsWith(CALL_PREFIX, start)) return undefined;
|
||||||
|
const tagEnd = findTagEnd(text, start);
|
||||||
|
if (tagEnd === -1) return undefined;
|
||||||
|
return parseOpenTagContent(text, start + 1, tagEnd);
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseOpenTagContent(text: string, contentStart: number, tagEnd: number): OpenTag | undefined {
|
||||||
|
let contentEnd = tagEnd;
|
||||||
|
let cursor = skipWhitespace(text, contentStart);
|
||||||
|
const nameStart = cursor;
|
||||||
|
if (!isNameStart(text[cursor])) return undefined;
|
||||||
|
cursor++;
|
||||||
|
while (cursor < contentEnd && isNameChar(text[cursor])) cursor++;
|
||||||
|
const name = text.slice(nameStart, cursor);
|
||||||
|
|
||||||
|
let selfClosing = false;
|
||||||
|
let last = contentEnd - 1;
|
||||||
|
while (last >= cursor && isWhitespace(text[last])) last--;
|
||||||
|
if (text[last] === "/") {
|
||||||
|
selfClosing = true;
|
||||||
|
contentEnd = last;
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
name,
|
||||||
|
rawAttrs: parseRawAttributes(text.slice(cursor, contentEnd)),
|
||||||
|
selfClosing,
|
||||||
|
end: tagEnd + 1,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseRawAttributes(text: string): RawAttribute[] {
|
||||||
|
const attrs: RawAttribute[] = [];
|
||||||
|
let index = 0;
|
||||||
|
while (index < text.length) {
|
||||||
|
index = skipWhitespace(text, index);
|
||||||
|
if (index >= text.length) break;
|
||||||
|
if (!isNameStart(text[index])) {
|
||||||
|
index++;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
const nameStart = index;
|
||||||
|
index++;
|
||||||
|
while (index < text.length && isNameChar(text[index])) index++;
|
||||||
|
const name = text.slice(nameStart, index);
|
||||||
|
index = skipWhitespace(text, index);
|
||||||
|
if (text[index] !== "=") {
|
||||||
|
attrs.push({ name, value: true });
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
index++;
|
||||||
|
index = skipWhitespace(text, index);
|
||||||
|
if (index >= text.length) {
|
||||||
|
attrs.push({ name, value: "" });
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
|
||||||
|
const quote = text[index];
|
||||||
|
if (quote === '"' || quote === "'") {
|
||||||
|
const valueStart = ++index;
|
||||||
|
while (index < text.length && text[index] !== quote) index++;
|
||||||
|
attrs.push({ name, value: text.slice(valueStart, index) });
|
||||||
|
if (index < text.length) index++;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
const valueStart = index;
|
||||||
|
while (index < text.length && !isWhitespace(text[index])) index++;
|
||||||
|
attrs.push({ name, value: text.slice(valueStart, index) });
|
||||||
|
}
|
||||||
|
return attrs;
|
||||||
|
}
|
||||||
|
|
||||||
|
function readElementNamePrefix(text: string, ltIndex: number): string | undefined {
|
||||||
|
let index = ltIndex + 1;
|
||||||
|
if (index >= text.length) return undefined;
|
||||||
|
if (!isNameStart(text[index])) return "";
|
||||||
|
const start = index;
|
||||||
|
index++;
|
||||||
|
while (index < text.length && isNameChar(text[index])) index++;
|
||||||
|
if (index >= text.length) return undefined;
|
||||||
|
const next = text[index];
|
||||||
|
return isWhitespace(next) || next === "/" || next === ">" ? text.slice(start, index) : "";
|
||||||
|
}
|
||||||
|
|
||||||
|
function findTagEnd(text: string, start: number): number {
|
||||||
|
let quote = "";
|
||||||
|
for (let index = start; index < text.length; index++) {
|
||||||
|
const ch = text[index];
|
||||||
|
if (quote) {
|
||||||
|
if (ch === quote) quote = "";
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (ch === '"' || ch === "'") {
|
||||||
|
quote = ch;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
if (ch === ">") return index;
|
||||||
|
}
|
||||||
|
return -1;
|
||||||
|
}
|
||||||
|
|
||||||
|
function stripBlockDelimiters(raw: string): string {
|
||||||
|
let start = 0;
|
||||||
|
let end = raw.length;
|
||||||
|
if (raw.startsWith("\r\n")) start = 2;
|
||||||
|
else if (raw.startsWith("\n")) start = 1;
|
||||||
|
if (end > start) {
|
||||||
|
if (raw.endsWith("\r\n")) end -= 2;
|
||||||
|
else if (raw.endsWith("\n")) end -= 1;
|
||||||
|
}
|
||||||
|
return raw.slice(start, end);
|
||||||
|
}
|
||||||
|
|
||||||
|
function skipWhitespace(text: string, index: number): number {
|
||||||
|
while (index < text.length && isWhitespace(text[index])) index++;
|
||||||
|
return index;
|
||||||
|
}
|
||||||
|
|
||||||
|
function isWhitespace(ch: string | undefined): boolean {
|
||||||
|
return ch === " " || ch === "\n" || ch === "\r" || ch === "\t" || ch === "\f";
|
||||||
|
}
|
||||||
|
|
||||||
|
function isNameStart(ch: string | undefined): boolean {
|
||||||
|
return ch !== undefined && NAME_START.test(ch);
|
||||||
|
}
|
||||||
|
|
||||||
|
function isNameChar(ch: string | undefined): boolean {
|
||||||
|
return ch !== undefined && NAME_CHAR.test(ch);
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "pi",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: options => new PiNativeInbandScanner(options),
|
||||||
|
renderAssistantToolCalls: renderPiNativeToolCalls,
|
||||||
|
renderToolResults: renderToolResponseResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
# Tools
|
||||||
|
|
||||||
|
You may call one or more functions to assist with the user query.
|
||||||
|
Tool calls are emitted as text using the exact syntax below, not as native provider tool messages.
|
||||||
|
|
||||||
|
Available functions are listed inside `<tools></tools>` as one JSON object per line:
|
||||||
|
|
||||||
|
<tools>
|
||||||
|
{{TOOLS}}
|
||||||
|
</tools>
|
||||||
|
|
||||||
|
{{GRAMMAR}}
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
Emit each tool call as one `<tool_call>` block wrapping a single-line JSON object with `name` and a nested `arguments` object:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_call>
|
||||||
|
{"name":"function_name","arguments":{"arg":"value"}}
|
||||||
|
</tool_call>
|
||||||
|
```
|
||||||
|
|
||||||
|
Do any private reasoning in `<think>...</think>` before your tool calls.
|
||||||
|
|
||||||
|
Tool results arrive later in a user turn:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_response>
|
||||||
|
verbatim tool result
|
||||||
|
</tool_response>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- `name` MUST match a listed function; `arguments` is a JSON object, never a JSON string.
|
||||||
|
- Multiple calls = consecutive `<tool_call>...</tool_call>` blocks; keep prose outside them.
|
||||||
|
- NEVER put tool calls inside `<think>`.
|
||||||
|
- Read each `<tool_response>` in call order. NEVER emit `<tool_response>` yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,202 @@
|
|||||||
|
import { parseJsonWithRepair } from "../utils/json-parse";
|
||||||
|
import { asRecord, mintToolCallId, partialSuffixOverlapAny } from "./coercion";
|
||||||
|
import grammarPrompt from "./qwen3.md" with { type: "text" };
|
||||||
|
import { renderHermesToolCalls, renderToolResponseResults } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner, InbandScannerOptions } from "./types";
|
||||||
|
|
||||||
|
const TOOL_OPEN = "<tool_call>";
|
||||||
|
const TOOL_CLOSE = "</tool_call>";
|
||||||
|
const THINK_OPEN = "<think>";
|
||||||
|
const THINK_CLOSE = "</think>";
|
||||||
|
|
||||||
|
const TOOL_START_TAGS = [TOOL_OPEN] as const;
|
||||||
|
const START_TAGS = [TOOL_OPEN, THINK_OPEN] as const;
|
||||||
|
const THINK_CLOSE_TAGS = [THINK_CLOSE] as const;
|
||||||
|
const COMPLETE_NAME = /^\s*\{\s*"name"\s*:\s*("(?:\\.|[^"\\])*")/;
|
||||||
|
|
||||||
|
type State = "outside" | "thinking" | "tool";
|
||||||
|
|
||||||
|
export class Qwen3InbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#state: State = "outside";
|
||||||
|
#id = "";
|
||||||
|
#name = "";
|
||||||
|
#started = false;
|
||||||
|
#thinking = "";
|
||||||
|
readonly #parseThinking: boolean;
|
||||||
|
|
||||||
|
constructor(options: InbandScannerOptions = {}) {
|
||||||
|
this.#parseThinking = options.parseThinking !== false;
|
||||||
|
}
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#consume(true);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#state === "outside") {
|
||||||
|
this.#consumeOutside(final, events);
|
||||||
|
if (this.#state === "outside") break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (this.#state === "thinking") {
|
||||||
|
this.#consumeThinking(final, events);
|
||||||
|
if (this.#state === "thinking") break;
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#consumeTool(final, events);
|
||||||
|
if (this.#state === "tool") break;
|
||||||
|
}
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeOutside(final: boolean, events: InbandScanEvent[]): void {
|
||||||
|
const tool = this.#buffer.indexOf(TOOL_OPEN);
|
||||||
|
const think = this.#parseThinking ? this.#buffer.indexOf(THINK_OPEN) : -1;
|
||||||
|
let start = tool;
|
||||||
|
let isThink = false;
|
||||||
|
if (think !== -1 && (start === -1 || think < start)) {
|
||||||
|
start = think;
|
||||||
|
isThink = true;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (start === -1) {
|
||||||
|
const tags = this.#parseThinking ? START_TAGS : TOOL_START_TAGS;
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, tags);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (emit.length > 0) events.push({ type: "text", text: emit });
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
if (start > 0) events.push({ type: "text", text: this.#buffer.slice(0, start) });
|
||||||
|
if (isThink) {
|
||||||
|
this.#buffer = this.#buffer.slice(start + THINK_OPEN.length);
|
||||||
|
this.#state = "thinking";
|
||||||
|
this.#thinking = "";
|
||||||
|
events.push({ type: "thinkingStart" });
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#buffer = this.#buffer.slice(start + TOOL_OPEN.length);
|
||||||
|
this.#state = "tool";
|
||||||
|
this.#id = mintToolCallId();
|
||||||
|
this.#name = "";
|
||||||
|
this.#started = false;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeThinking(final: boolean, events: InbandScanEvent[]): void {
|
||||||
|
const close = this.#buffer.indexOf(THINK_CLOSE);
|
||||||
|
if (close === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, THINK_CLOSE_TAGS);
|
||||||
|
const delta = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
this.#emitThinkingDelta(delta, events);
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
if (final) this.#endThinking(events);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
this.#emitThinkingDelta(this.#buffer.slice(0, close), events);
|
||||||
|
this.#buffer = this.#buffer.slice(close + THINK_CLOSE.length);
|
||||||
|
this.#endThinking(events);
|
||||||
|
}
|
||||||
|
|
||||||
|
#consumeTool(final: boolean, events: InbandScanEvent[]): void {
|
||||||
|
const close = this.#buffer.indexOf(TOOL_CLOSE);
|
||||||
|
const body = close === -1 ? this.#buffer : this.#buffer.slice(0, close);
|
||||||
|
if (!this.#started) this.#tryStart(body, events);
|
||||||
|
if (close === -1) {
|
||||||
|
if (final) this.#resetTool();
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
const parsed = this.#parseCall(body);
|
||||||
|
if (parsed) {
|
||||||
|
if (!this.#started) {
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: parsed.name });
|
||||||
|
this.#started = true;
|
||||||
|
}
|
||||||
|
events.push({
|
||||||
|
type: "toolEnd",
|
||||||
|
id: this.#id,
|
||||||
|
name: parsed.name,
|
||||||
|
arguments: parsed.arguments,
|
||||||
|
rawBlock: `${TOOL_OPEN}${body}${TOOL_CLOSE}`,
|
||||||
|
});
|
||||||
|
}
|
||||||
|
this.#buffer = this.#buffer.slice(close + TOOL_CLOSE.length);
|
||||||
|
this.#resetTool();
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitThinkingDelta(delta: string, events: InbandScanEvent[]): void {
|
||||||
|
if (delta.length === 0) return;
|
||||||
|
this.#thinking += delta;
|
||||||
|
events.push({ type: "thinkingDelta", delta });
|
||||||
|
}
|
||||||
|
|
||||||
|
#endThinking(events: InbandScanEvent[]): void {
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#state = "outside";
|
||||||
|
}
|
||||||
|
|
||||||
|
#tryStart(body: string, events: InbandScanEvent[]): void {
|
||||||
|
const nameMatch = COMPLETE_NAME.exec(body);
|
||||||
|
if (!nameMatch) return;
|
||||||
|
let name: unknown;
|
||||||
|
try {
|
||||||
|
name = JSON.parse(nameMatch[1]!);
|
||||||
|
} catch {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (typeof name !== "string" || name.length === 0) return;
|
||||||
|
this.#name = name;
|
||||||
|
this.#started = true;
|
||||||
|
events.push({ type: "toolStart", id: this.#id, name: this.#name });
|
||||||
|
}
|
||||||
|
|
||||||
|
#parseCall(body: string): { name: string; arguments: Record<string, unknown> } | undefined {
|
||||||
|
try {
|
||||||
|
const parsed = parseJsonWithRepair<{ name?: unknown; arguments?: unknown }>(body.trim());
|
||||||
|
if (typeof parsed.name !== "string" || parsed.name.length === 0) return undefined;
|
||||||
|
let args = parsed.arguments;
|
||||||
|
if (typeof args === "string") {
|
||||||
|
try {
|
||||||
|
args = parseJsonWithRepair<unknown>(args);
|
||||||
|
} catch {
|
||||||
|
args = {};
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return { name: parsed.name, arguments: asRecord(args) };
|
||||||
|
} catch {
|
||||||
|
return undefined;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
#resetTool(): void {
|
||||||
|
this.#state = "outside";
|
||||||
|
this.#id = "";
|
||||||
|
this.#name = "";
|
||||||
|
this.#started = false;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "qwen3",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: options => new Qwen3InbandScanner(options),
|
||||||
|
renderAssistantToolCalls: renderHermesToolCalls,
|
||||||
|
renderToolResults: renderToolResponseResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -0,0 +1,190 @@
|
|||||||
|
import type { ToolCall } from "../types";
|
||||||
|
import { buildArgShapes, getArrayItemSchema, getObjectProperties, isStringOnlySchema } from "./coercion";
|
||||||
|
import type { GrammarRenderOptions, GrammarToolResult, InbandTool } from "./types";
|
||||||
|
|
||||||
|
const DEEPSEEK_TOOL_CALLS_BEGIN = "<|tool▁calls▁begin|>";
|
||||||
|
const DEEPSEEK_TOOL_CALLS_END = "<|tool▁calls▁end|>";
|
||||||
|
const DEEPSEEK_TOOL_CALL_BEGIN = "<|tool▁call▁begin|>";
|
||||||
|
const DEEPSEEK_TOOL_CALL_END = "<|tool▁call▁end|>";
|
||||||
|
const DEEPSEEK_TOOL_SEPARATOR = "<|tool▁sep|>";
|
||||||
|
const DEEPSEEK_TOOL_OUTPUT_BEGIN = "<|tool▁output▁begin|>";
|
||||||
|
const DEEPSEEK_TOOL_OUTPUT_END = "<|tool▁output▁end|>";
|
||||||
|
|
||||||
|
export function renderGlmToolCalls(calls: readonly ToolCall[], options: GrammarRenderOptions = {}): string {
|
||||||
|
const shapes = buildArgShapes(options.tools);
|
||||||
|
return calls
|
||||||
|
.map(call => {
|
||||||
|
const shape = shapes.get(call.name);
|
||||||
|
let body = `<tool_call>${call.name}`;
|
||||||
|
for (const key in call.arguments) {
|
||||||
|
const value = call.arguments[key];
|
||||||
|
const rendered = shape?.stringArgs.has(key) && typeof value === "string" ? value : stringifyJson(value);
|
||||||
|
body += `\n<arg_key>${key}</arg_key>\n<arg_value>${rendered}</arg_value>`;
|
||||||
|
}
|
||||||
|
return `${body}\n</tool_call>`;
|
||||||
|
})
|
||||||
|
.join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderGlmToolResults(results: readonly GrammarToolResult[]): string {
|
||||||
|
return `<observation>\n${renderToolResponseResults(results)}\n</observation>`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderHermesToolCalls(calls: readonly ToolCall[]): string {
|
||||||
|
return calls
|
||||||
|
.map(call => `<tool_call>\n${stringifyJson({ name: call.name, arguments: call.arguments })}\n</tool_call>`)
|
||||||
|
.join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderKimiToolCalls(calls: readonly ToolCall[]): string {
|
||||||
|
if (calls.length === 0) return "";
|
||||||
|
const body = calls
|
||||||
|
.map(
|
||||||
|
(call, index) =>
|
||||||
|
`<|tool_call_begin|>${kimiCallId(call.name, call.id, index)}<|tool_call_argument_begin|>${stringifyJson(call.arguments)}<|tool_call_end|>`,
|
||||||
|
)
|
||||||
|
.join("");
|
||||||
|
return `<|tool_calls_section_begin|>${body}<|tool_calls_section_end|>`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderKimiToolResults(results: readonly GrammarToolResult[]): string {
|
||||||
|
return results
|
||||||
|
.map(
|
||||||
|
result =>
|
||||||
|
`<|im_system|>${result.name}<|im_middle|>## Return of ${kimiCallId(result.name, result.id, result.index)}\n${result.text}<|im_end|>`,
|
||||||
|
)
|
||||||
|
.join("");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderDeepSeekToolCalls(calls: readonly ToolCall[]): string {
|
||||||
|
if (calls.length === 0) return "";
|
||||||
|
const body = calls
|
||||||
|
.map(
|
||||||
|
call =>
|
||||||
|
`${DEEPSEEK_TOOL_CALL_BEGIN}${call.name}${DEEPSEEK_TOOL_SEPARATOR}${stringifyJson(call.arguments)}${DEEPSEEK_TOOL_CALL_END}`,
|
||||||
|
)
|
||||||
|
.join("");
|
||||||
|
return `${DEEPSEEK_TOOL_CALLS_BEGIN}${body}${DEEPSEEK_TOOL_CALLS_END}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderDeepSeekToolResults(results: readonly GrammarToolResult[]): string {
|
||||||
|
return results.map(result => `${DEEPSEEK_TOOL_OUTPUT_BEGIN}${result.text}${DEEPSEEK_TOOL_OUTPUT_END}`).join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderHarmonyToolCalls(calls: readonly ToolCall[]): string {
|
||||||
|
return calls
|
||||||
|
.map(
|
||||||
|
call =>
|
||||||
|
`<|start|>assistant<|channel|>commentary to=${harmonyRecipient(call.name)} <|constrain|>json<|message|>${stringifyJson(call.arguments)}<|call|>`,
|
||||||
|
)
|
||||||
|
.join("");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderHarmonyToolResults(results: readonly GrammarToolResult[]): string {
|
||||||
|
return results
|
||||||
|
.map(
|
||||||
|
result =>
|
||||||
|
`<|start|>${harmonyRecipient(result.name)} to=assistant<|channel|>commentary<|message|>${result.text}<|end|>`,
|
||||||
|
)
|
||||||
|
.join("");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderAnthropicToolCalls(calls: readonly ToolCall[], options: GrammarRenderOptions = {}): string {
|
||||||
|
if (calls.length === 0) return "";
|
||||||
|
return `<function_calls>\n${renderXmlInvokes(calls, options.tools ?? [])}\n</function_calls>`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderAnthropicToolResults(results: readonly GrammarToolResult[]): string {
|
||||||
|
const body = results
|
||||||
|
.map(result => {
|
||||||
|
const tag = result.isError ? "error" : "result";
|
||||||
|
const streamTag = result.isError ? "stderr" : "stdout";
|
||||||
|
return `<${tag}>\n<tool_name>${escapeXmlText(result.name)}</tool_name>\n<${streamTag}>${result.text}</${streamTag}>\n</${tag}>`;
|
||||||
|
})
|
||||||
|
.join("\n");
|
||||||
|
return `<function_results>\n${body}\n</function_results>`;
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderXmlToolCalls(calls: readonly ToolCall[], options: GrammarRenderOptions = {}): string {
|
||||||
|
return renderXmlInvokes(calls, options.tools ?? []);
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderPiNativeToolCalls(calls: readonly ToolCall[], options: GrammarRenderOptions = {}): string {
|
||||||
|
const shapes = buildArgShapes(options.tools);
|
||||||
|
return calls
|
||||||
|
.map(call => {
|
||||||
|
const shape = shapes.get(call.name);
|
||||||
|
let body = `<call:${call.name}>`;
|
||||||
|
for (const key in call.arguments) {
|
||||||
|
body += `\n${renderPiNativeElement(key, call.arguments[key], shape?.properties[key])}`;
|
||||||
|
}
|
||||||
|
return `${body}\n</call:${call.name}>`;
|
||||||
|
})
|
||||||
|
.join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
export function renderToolResponseResults(results: readonly GrammarToolResult[]): string {
|
||||||
|
return results.map(result => `<tool_response>\n${result.text}\n</tool_response>`).join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderXmlInvokes(calls: readonly ToolCall[], tools: readonly InbandTool[]): string {
|
||||||
|
const shapes = buildArgShapes(tools);
|
||||||
|
return calls
|
||||||
|
.map(call => {
|
||||||
|
const shape = shapes.get(call.name);
|
||||||
|
let body = `<invoke name="${escapeXmlAttr(call.name)}">`;
|
||||||
|
for (const key in call.arguments) {
|
||||||
|
const value = call.arguments[key];
|
||||||
|
const isString = shape?.stringArgs.has(key) === true;
|
||||||
|
const stringAttr = isString ? ' string="true"' : ' string="false"';
|
||||||
|
const rendered = isString && typeof value === "string" ? value : stringifyJson(value);
|
||||||
|
body += `<parameter name="${escapeXmlAttr(key)}"${stringAttr}>${rendered}</parameter>`;
|
||||||
|
}
|
||||||
|
return `${body}</invoke>`;
|
||||||
|
})
|
||||||
|
.join("\n");
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderPiNativeElement(key: string, value: unknown, schema: unknown): string {
|
||||||
|
if (Array.isArray(value)) {
|
||||||
|
const itemSchema = getArrayItemSchema(schema);
|
||||||
|
return value.map(item => renderPiNativeElement(key, item, itemSchema)).join("\n");
|
||||||
|
}
|
||||||
|
if (value && typeof value === "object") {
|
||||||
|
const record = value as Record<string, unknown>;
|
||||||
|
const properties = getObjectProperties(schema);
|
||||||
|
let body = `<${key}>`;
|
||||||
|
for (const childKey in record) {
|
||||||
|
body += `\n${renderPiNativeElement(childKey, record[childKey], properties[childKey])}`;
|
||||||
|
}
|
||||||
|
return `${body}\n</${key}>`;
|
||||||
|
}
|
||||||
|
return `<${key}>${renderPiNativeScalar(value, schema)}</${key}>`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function renderPiNativeScalar(value: unknown, schema: unknown): string {
|
||||||
|
if (typeof value === "string") return value;
|
||||||
|
if (isStringOnlySchema(schema) && value === null) return "";
|
||||||
|
return stringifyJson(value);
|
||||||
|
}
|
||||||
|
|
||||||
|
function kimiCallId(name: string, id: string, index: number): string {
|
||||||
|
const trimmed = id.trim();
|
||||||
|
return trimmed.startsWith("functions.") ? trimmed : `functions.${name}:${index}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function harmonyRecipient(name: string): string {
|
||||||
|
return name.startsWith("functions.") ? name : `functions.${name}`;
|
||||||
|
}
|
||||||
|
|
||||||
|
function stringifyJson(value: unknown): string {
|
||||||
|
return JSON.stringify(value) ?? "null";
|
||||||
|
}
|
||||||
|
|
||||||
|
function escapeXmlAttr(value: string): string {
|
||||||
|
return value.replaceAll("&", "&").replaceAll('"', """).replaceAll("<", "<").replaceAll(">", ">");
|
||||||
|
}
|
||||||
|
|
||||||
|
function escapeXmlText(value: string): string {
|
||||||
|
return value.replaceAll("&", "&").replaceAll("<", "<").replaceAll(">", ">");
|
||||||
|
}
|
||||||
@@ -0,0 +1,91 @@
|
|||||||
|
import { partialSuffixOverlapAny } from "./coercion";
|
||||||
|
import type { InbandScanEvent, InbandScanner } from "./types";
|
||||||
|
|
||||||
|
const THINK_OPEN = "<think>";
|
||||||
|
const THINK_CLOSE = "</think>";
|
||||||
|
const THINKING_OPEN = "<thinking>";
|
||||||
|
const THINKING_CLOSE = "</thinking>";
|
||||||
|
const TAGS = [
|
||||||
|
{ open: THINK_OPEN, close: THINK_CLOSE },
|
||||||
|
{ open: THINKING_OPEN, close: THINKING_CLOSE },
|
||||||
|
] as const;
|
||||||
|
const OPENS = [THINK_OPEN, THINKING_OPEN] as const;
|
||||||
|
|
||||||
|
type Tag = { readonly open: string; readonly close: string };
|
||||||
|
|
||||||
|
export class ThinkingInbandScanner implements InbandScanner {
|
||||||
|
#buffer = "";
|
||||||
|
#closeTag = "";
|
||||||
|
#thinking = "";
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
if (text.length === 0) return [];
|
||||||
|
this.#buffer += text;
|
||||||
|
return this.#consume(false);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
const events = this.#consume(true);
|
||||||
|
if (this.#buffer.length === 0) return events;
|
||||||
|
if (this.#closeTag) {
|
||||||
|
this.#emitThinking(this.#buffer, events);
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
} else {
|
||||||
|
events.push({ type: "text", text: this.#buffer });
|
||||||
|
}
|
||||||
|
this.#buffer = "";
|
||||||
|
this.#closeTag = "";
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#consume(final: boolean): InbandScanEvent[] {
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
while (this.#buffer.length > 0) {
|
||||||
|
if (this.#closeTag) {
|
||||||
|
const close = this.#buffer.indexOf(this.#closeTag);
|
||||||
|
if (close === -1) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, [this.#closeTag]);
|
||||||
|
this.#emitThinking(this.#buffer.slice(0, this.#buffer.length - hold), events);
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
this.#emitThinking(this.#buffer.slice(0, close), events);
|
||||||
|
this.#buffer = this.#buffer.slice(close + this.#closeTag.length);
|
||||||
|
events.push({ type: "thinkingEnd", thinking: this.#thinking });
|
||||||
|
this.#thinking = "";
|
||||||
|
this.#closeTag = "";
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
|
||||||
|
const tag = findEarliestOpen(this.#buffer);
|
||||||
|
if (!tag) {
|
||||||
|
const hold = final ? 0 : partialSuffixOverlapAny(this.#buffer, OPENS);
|
||||||
|
const emit = this.#buffer.slice(0, this.#buffer.length - hold);
|
||||||
|
if (emit.length > 0) events.push({ type: "text", text: emit });
|
||||||
|
this.#buffer = this.#buffer.slice(this.#buffer.length - hold);
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
if (tag.index > 0) events.push({ type: "text", text: this.#buffer.slice(0, tag.index) });
|
||||||
|
this.#buffer = this.#buffer.slice(tag.index + tag.open.length);
|
||||||
|
this.#closeTag = tag.close;
|
||||||
|
this.#thinking = "";
|
||||||
|
events.push({ type: "thinkingStart" });
|
||||||
|
}
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
#emitThinking(delta: string, events: InbandScanEvent[]): void {
|
||||||
|
if (delta.length === 0) return;
|
||||||
|
this.#thinking += delta;
|
||||||
|
events.push({ type: "thinkingDelta", delta });
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function findEarliestOpen(buffer: string): (Tag & { index: number }) | undefined {
|
||||||
|
let best: (Tag & { index: number }) | undefined;
|
||||||
|
for (const tag of TAGS) {
|
||||||
|
const index = buffer.indexOf(tag.open);
|
||||||
|
if (index !== -1 && (!best || index < best.index)) best = { ...tag, index };
|
||||||
|
}
|
||||||
|
return best;
|
||||||
|
}
|
||||||
@@ -0,0 +1,50 @@
|
|||||||
|
import type { Context, ToolCall } from "../types";
|
||||||
|
|
||||||
|
export type ToolCallSyntax = "glm" | "hermes" | "kimi" | "xml" | "anthropic" | "deepseek" | "harmony" | "pi" | "qwen3";
|
||||||
|
|
||||||
|
export type InbandScanEvent =
|
||||||
|
| { type: "text"; text: string }
|
||||||
|
| { type: "thinkingStart" }
|
||||||
|
| { type: "thinkingDelta"; delta: string }
|
||||||
|
| { type: "thinkingEnd"; thinking: string }
|
||||||
|
| { type: "toolStart"; id: string; name: string }
|
||||||
|
| { type: "toolArgDelta"; id: string; name: string; key: string; delta: string }
|
||||||
|
| { type: "toolEnd"; id: string; name: string; arguments: Record<string, unknown>; rawBlock?: string };
|
||||||
|
|
||||||
|
export interface InbandScanner {
|
||||||
|
feed(text: string): InbandScanEvent[];
|
||||||
|
flush(): InbandScanEvent[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface GrammarToolResult {
|
||||||
|
readonly id: string;
|
||||||
|
readonly name: string;
|
||||||
|
readonly index: number;
|
||||||
|
readonly text: string;
|
||||||
|
readonly isError: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface GrammarRenderOptions {
|
||||||
|
readonly tools?: readonly InbandTool[];
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface Grammar {
|
||||||
|
readonly syntax: ToolCallSyntax;
|
||||||
|
readonly prompt: string;
|
||||||
|
createScanner(options?: InbandScannerOptions): InbandScanner;
|
||||||
|
renderAssistantToolCalls(calls: readonly ToolCall[], options?: GrammarRenderOptions): string;
|
||||||
|
renderToolResults(results: readonly GrammarToolResult[], options?: GrammarRenderOptions): string;
|
||||||
|
}
|
||||||
|
|
||||||
|
export interface InbandScannerOptions {
|
||||||
|
/** string-typed arg names for a tool → read verbatim. Ignored by JSON-carrying syntaxes. */
|
||||||
|
stringArgs?: (toolName: string) => ReadonlySet<string>;
|
||||||
|
/** Full tool schemas for schema-driven syntaxes such as GLM XML and pi-native. */
|
||||||
|
tools?: readonly InbandTool[];
|
||||||
|
/** XML only: parse pipe-wrapped DeepSeek DSML tags vs plain Anthropic invoke/parameter tags. */
|
||||||
|
xmlTagset?: "anthropic" | "dsml";
|
||||||
|
/** Emit thinking markers as thinking events instead of visible text when the syntax defines them. */
|
||||||
|
parseThinking?: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
export type InbandTool = NonNullable<Context["tools"]>[number];
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
## Format guide
|
||||||
|
|
||||||
|
A call is one `<invoke>` element whose `<parameter>` children carry its arguments:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<invoke name="fn"><parameter name="arg">value</parameter></invoke>
|
||||||
|
```
|
||||||
|
|
||||||
|
Emit consecutive `<invoke>…</invoke>` blocks for multiple calls; you MAY wrap them in `<tool_calls>…</tool_calls>`. Each call's result arrives as a response block:
|
||||||
|
|
||||||
|
```text
|
||||||
|
<tool_response>
|
||||||
|
verbatim tool result
|
||||||
|
</tool_response>
|
||||||
|
```
|
||||||
|
|
||||||
|
## Rules
|
||||||
|
|
||||||
|
- `name` MUST match a listed function.
|
||||||
|
- String values are literal text (no JSON quotes or escaping); non-string values are JSON. Add `string="false"` to a parameter only to force JSON parsing of a value the schema treats as a string.
|
||||||
|
- Read each `<tool_response>` in call order. NEVER emit `<tool_response>` yourself.
|
||||||
|
- After emitting your tool calls, YOU MUST EMIT THE STOP SEQUENCE AND HALT.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
import { AnthropicInbandScanner } from "./anthropic";
|
||||||
|
import { DeepSeekInbandScanner } from "./deepseek";
|
||||||
|
import { renderToolResponseResults, renderXmlToolCalls } from "./rendering";
|
||||||
|
import type { Grammar, InbandScanEvent, InbandScanner, InbandScannerOptions } from "./types";
|
||||||
|
import grammarPrompt from "./xml.md" with { type: "text" };
|
||||||
|
|
||||||
|
export class XmlInbandScanner implements InbandScanner {
|
||||||
|
readonly #inner: InbandScanner;
|
||||||
|
|
||||||
|
constructor(options: InbandScannerOptions = {}) {
|
||||||
|
this.#inner =
|
||||||
|
options.xmlTagset === "dsml" ? new DeepSeekInbandScanner(options) : new AnthropicInbandScanner(options);
|
||||||
|
}
|
||||||
|
|
||||||
|
feed(text: string): InbandScanEvent[] {
|
||||||
|
return this.#inner.feed(text);
|
||||||
|
}
|
||||||
|
|
||||||
|
flush(): InbandScanEvent[] {
|
||||||
|
return this.#inner.flush();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
const grammar: Grammar = {
|
||||||
|
syntax: "xml",
|
||||||
|
prompt: grammarPrompt,
|
||||||
|
createScanner: options => new XmlInbandScanner(options),
|
||||||
|
renderAssistantToolCalls: renderXmlToolCalls,
|
||||||
|
renderToolResults: renderToolResponseResults,
|
||||||
|
};
|
||||||
|
|
||||||
|
export default grammar;
|
||||||
@@ -426,6 +426,11 @@ export interface ToolCall {
|
|||||||
arguments: Record<string, any>;
|
arguments: Record<string, any>;
|
||||||
thoughtSignature?: string; // Google-specific: opaque signature for reusing thought context
|
thoughtSignature?: string; // Google-specific: opaque signature for reusing thought context
|
||||||
intent?: string; // Harness-level intent metadata extracted from traced tool arguments
|
intent?: string; // Harness-level intent metadata extracted from traced tool arguments
|
||||||
|
/**
|
||||||
|
* Verbatim in-band syntax block that produced this synthetic `ptc_*` call.
|
||||||
|
* Present only for owned prompt/tool-call formats; provider-native calls omit it.
|
||||||
|
*/
|
||||||
|
rawBlock?: string;
|
||||||
/**
|
/**
|
||||||
* Original wire-level name when the tool was invoked via OpenAI's custom-tool
|
* Original wire-level name when the tool was invoked via OpenAI's custom-tool
|
||||||
* mechanism (e.g., `apply_patch`). Set by `openai-responses` on receive so
|
* mechanism (e.g., `apply_patch`). Set by `openai-responses` on receive so
|
||||||
|
|||||||
@@ -7,7 +7,7 @@
|
|||||||
* hashline DSL form. Other tools and surfaces fall through to
|
* hashline DSL form. Other tools and surfaces fall through to
|
||||||
* abort-and-retry handled by the agent loop.
|
* abort-and-retry handled by the agent loop.
|
||||||
*/
|
*/
|
||||||
import type { AssistantMessage, Model, ToolCall } from "@oh-my-pi/pi-ai";
|
import type { AssistantMessage, Model, ToolCall } from "../types";
|
||||||
|
|
||||||
// Single source of truth for the marker pattern. `M` in the errata.
|
// Single source of truth for the marker pattern. `M` in the errata.
|
||||||
// Use a fresh non-global instance for `.test()` to avoid lastIndex pitfalls.
|
// Use a fresh non-global instance for `.test()` to avoid lastIndex pitfalls.
|
||||||
@@ -2,60 +2,20 @@
|
|||||||
* Streaming-safe filters for leaked chat-template tool-call and thinking markup.
|
* Streaming-safe filters for leaked chat-template tool-call and thinking markup.
|
||||||
*
|
*
|
||||||
* Hosted models sometimes leak raw template markup into visible `content` instead
|
* Hosted models sometimes leak raw template markup into visible `content` instead
|
||||||
* of returning structured events. One `StreamMarkupHealing` instance owns one stream
|
* of returning structured events. Tool-call healing delegates to the same
|
||||||
* and one grammar selected by options:
|
* grammar scanners used by owned in-band tool calling; this file keeps the
|
||||||
*
|
* provider-facing compatibility wrapper and model/provider gating.
|
||||||
* - `kimi`: Kimi K2 `<|tool_calls_section_begin|>` sections.
|
|
||||||
* - `dsml`: DeepSeek `<|DSML|tool_calls>` envelopes.
|
|
||||||
* - `thinking`: plain `<think>` / `<thinking>` blocks used by MiniMax-style streams.
|
|
||||||
*
|
|
||||||
* The parser strips marker bytes, reconstructs embedded calls, emits thinking
|
|
||||||
* deltas for thinking blocks, and holds partial tags across chunk boundaries.
|
|
||||||
*/
|
*/
|
||||||
|
|
||||||
import { isDeepseekModelIdOrName } from "@oh-my-pi/pi-catalog/identity";
|
import { isDeepseekModelIdOrName } from "@oh-my-pi/pi-catalog/identity";
|
||||||
|
|
||||||
import { parseJsonWithRepair } from "./json-parse";
|
import { createInbandScanner } from "../grammar/factory";
|
||||||
|
import { ThinkingInbandScanner } from "../grammar/thinking";
|
||||||
|
import type { InbandScanEvent, InbandScanner } from "../grammar/types";
|
||||||
|
|
||||||
const KIMI_SECTION_BEGIN = "<|tool_calls_section_begin|>";
|
|
||||||
const KIMI_SECTION_END = "<|tool_calls_section_end|>";
|
const KIMI_SECTION_END = "<|tool_calls_section_end|>";
|
||||||
const KIMI_CALL_BEGIN = "<|tool_call_begin|>";
|
const DSML_TOOL_CALLS_CLOSE_FULLWIDTH = "</|DSML|tool_calls>";
|
||||||
const KIMI_CALL_END = "<|tool_call_end|>";
|
const DSML_TOOL_CALLS_CLOSE_ASCII = "</|DSML|tool_calls>";
|
||||||
const KIMI_ARG_BEGIN = "<|tool_call_argument_begin|>";
|
|
||||||
const KIMI_TOKENS = [KIMI_SECTION_BEGIN, KIMI_SECTION_END, KIMI_CALL_BEGIN, KIMI_CALL_END, KIMI_ARG_BEGIN] as const;
|
|
||||||
|
|
||||||
/** Maximum buffered Kimi partial-token length before giving up holdback. */
|
|
||||||
const MAX_KIMI_PARTIAL_HOLD = 64;
|
|
||||||
|
|
||||||
/** Both fullwidth (U+FF5C) and ASCII pipes are observed in DeepSeek DSML leaks. */
|
|
||||||
const DSML_PIPE = "[||]";
|
|
||||||
const DSML_TOOL_CALLS_OPEN_RE = new RegExp(`<${DSML_PIPE}DSML${DSML_PIPE}tool_calls>`, "y");
|
|
||||||
const DSML_TOOL_CALLS_CLOSE_RE = new RegExp(`</${DSML_PIPE}DSML${DSML_PIPE}tool_calls>`, "y");
|
|
||||||
const DSML_INVOKE_OPEN_RE = new RegExp(`<${DSML_PIPE}DSML${DSML_PIPE}invoke\\s+name="([^"]*)"\\s*>`, "y");
|
|
||||||
const DSML_INVOKE_CLOSE_RE = new RegExp(`</${DSML_PIPE}DSML${DSML_PIPE}invoke>`, "y");
|
|
||||||
const DSML_PARAMETER_OPEN_RE = new RegExp(
|
|
||||||
`<${DSML_PIPE}DSML${DSML_PIPE}parameter\\s+name="([^"]*)"(?:\\s+string="(true|false)")?\\s*>`,
|
|
||||||
"y",
|
|
||||||
);
|
|
||||||
const DSML_PARAMETER_CLOSE_RE = new RegExp(`</${DSML_PIPE}DSML${DSML_PIPE}parameter>`, "y");
|
|
||||||
/** Canonical DSML section-open shape; `|` positions accept either pipe variant. */
|
|
||||||
const DSML_SECTION_OPEN_TEMPLATE = "<|DSML|tool_calls>";
|
|
||||||
|
|
||||||
const THINK_OPEN = "<think>";
|
|
||||||
const THINK_CLOSE = "</think>";
|
|
||||||
const THINKING_OPEN = "<thinking>";
|
|
||||||
const THINKING_CLOSE = "</thinking>";
|
|
||||||
|
|
||||||
const PLAIN_THINKING_TAGS = [
|
|
||||||
{ open: THINK_OPEN, close: THINK_CLOSE },
|
|
||||||
{ open: THINKING_OPEN, close: THINKING_CLOSE },
|
|
||||||
] as const;
|
|
||||||
|
|
||||||
/** Cap held-back XML tag bytes so a stray `<` in prose cannot grow unboundedly. */
|
|
||||||
const MAX_XML_PARTIAL_HOLD = 256;
|
|
||||||
|
|
||||||
/** Maximum parameter bytes to accumulate before abandoning a pathological XML call. */
|
|
||||||
const MAX_XML_PARAM_VALUE_LENGTH = 1_000_000;
|
|
||||||
|
|
||||||
export interface HealedToolCall {
|
export interface HealedToolCall {
|
||||||
readonly id: string;
|
readonly id: string;
|
||||||
@@ -74,48 +34,28 @@ export type StreamMarkupHealingEvent =
|
|||||||
| { readonly type: "thinking"; readonly thinking: string }
|
| { readonly type: "thinking"; readonly thinking: string }
|
||||||
| { readonly type: "toolCall"; readonly call: HealedToolCall };
|
| { readonly type: "toolCall"; readonly call: HealedToolCall };
|
||||||
|
|
||||||
type XmlToolState =
|
|
||||||
| { readonly kind: "idle" }
|
|
||||||
| { readonly kind: "section" }
|
|
||||||
| { readonly kind: "invoke"; readonly name: string; readonly args: Record<string, unknown> }
|
|
||||||
| {
|
|
||||||
readonly kind: "parameter";
|
|
||||||
readonly invokeName: string;
|
|
||||||
readonly args: Record<string, unknown>;
|
|
||||||
readonly paramName: string;
|
|
||||||
readonly isString: boolean;
|
|
||||||
value: string;
|
|
||||||
truncated?: boolean;
|
|
||||||
};
|
|
||||||
|
|
||||||
type ThinkingTag = { readonly open: string; readonly close: string };
|
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* State machine that consumes streamed visible text and emits cleaned text,
|
* State machine that consumes streamed visible text and emits cleaned text,
|
||||||
* thinking deltas, and reconstructed tool calls.
|
* thinking deltas, and reconstructed tool calls.
|
||||||
*
|
*
|
||||||
* Feed only one stream channel (usually `delta.content` / `message.content`).
|
* Feed only one stream channel (usually `delta.content` / `message.content`).
|
||||||
* Mixing reasoning and visible text into the same instance can corrupt the
|
* Mixing reasoning and visible text into the same instance can corrupt held-back
|
||||||
* held-back partial tag buffer.
|
* partial tag buffers.
|
||||||
*/
|
*/
|
||||||
export class StreamMarkupHealing {
|
export class StreamMarkupHealing {
|
||||||
readonly #pattern: StreamMarkupHealingPattern;
|
readonly #pattern: StreamMarkupHealingPattern;
|
||||||
#buffer = "";
|
readonly #scanner: InbandScanner;
|
||||||
#offset = 0;
|
|
||||||
|
|
||||||
#kimiInSection = false;
|
|
||||||
#kimiInCall = false;
|
|
||||||
#kimiInArgs = false;
|
|
||||||
#kimiPendingId = "";
|
|
||||||
#kimiPendingArgs = "";
|
|
||||||
|
|
||||||
#xmlState: XmlToolState = { kind: "idle" };
|
|
||||||
#thinkingCloseTag = "";
|
|
||||||
#sectionTerminated = false;
|
#sectionTerminated = false;
|
||||||
readonly #completed: HealedToolCall[] = [];
|
readonly #completed: HealedToolCall[] = [];
|
||||||
|
|
||||||
constructor(options: StreamMarkupHealingOptions) {
|
constructor(options: StreamMarkupHealingOptions) {
|
||||||
this.#pattern = options.pattern;
|
this.#pattern = options.pattern;
|
||||||
|
this.#scanner =
|
||||||
|
options.pattern === "kimi"
|
||||||
|
? createInbandScanner("kimi")
|
||||||
|
: options.pattern === "dsml"
|
||||||
|
? createInbandScanner("xml", { xmlTagset: "dsml" })
|
||||||
|
: new ThinkingInbandScanner();
|
||||||
}
|
}
|
||||||
|
|
||||||
get pattern(): StreamMarkupHealingPattern {
|
get pattern(): StreamMarkupHealingPattern {
|
||||||
@@ -125,8 +65,8 @@ export class StreamMarkupHealing {
|
|||||||
/**
|
/**
|
||||||
* Feed a chunk and return visible text only. Reconstructed tool calls are
|
* Feed a chunk and return visible text only. Reconstructed tool calls are
|
||||||
* stored for {@link drainCompleted}; thinking blocks are intentionally not
|
* stored for {@link drainCompleted}; thinking blocks are intentionally not
|
||||||
* returned by this compatibility helper. Use {@link feedEvents} when the
|
* returned by this compatibility helper. Use {@link feedEvents} when the caller
|
||||||
* caller needs ordered text/thinking/tool-call events.
|
* needs ordered text/thinking/tool-call events.
|
||||||
*/
|
*/
|
||||||
feed(text: string): string {
|
feed(text: string): string {
|
||||||
let clean = "";
|
let clean = "";
|
||||||
@@ -143,16 +83,8 @@ export class StreamMarkupHealing {
|
|||||||
/** Feed a chunk and return cleaned text/thinking/tool-call events in stream order. */
|
/** Feed a chunk and return cleaned text/thinking/tool-call events in stream order. */
|
||||||
feedEvents(text: string): StreamMarkupHealingEvent[] {
|
feedEvents(text: string): StreamMarkupHealingEvent[] {
|
||||||
if (text.length === 0) return [];
|
if (text.length === 0) return [];
|
||||||
this.#compact();
|
this.#markSectionClosed(text);
|
||||||
this.#buffer += text;
|
return this.#convertScannerEvents(this.#scanner.feed(text));
|
||||||
switch (this.#pattern) {
|
|
||||||
case "kimi":
|
|
||||||
return this.#consumeKimiEvents();
|
|
||||||
case "dsml":
|
|
||||||
return this.#consumeDsmlEvents();
|
|
||||||
case "thinking":
|
|
||||||
return this.#consumePlainThinkingEvents();
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
@@ -176,32 +108,12 @@ export class StreamMarkupHealing {
|
|||||||
|
|
||||||
/**
|
/**
|
||||||
* Flush held-back stream-end fragments as ordered events. Partial tool-call
|
* Flush held-back stream-end fragments as ordered events. Partial tool-call
|
||||||
* sections/envelopes are dropped; unterminated thinking blocks are emitted as
|
* sections/envelopes are dropped by the delegated scanners; unterminated
|
||||||
* thinking, matching the previous MiniMax parser behavior.
|
* thinking blocks are emitted as thinking, matching the previous MiniMax parser
|
||||||
|
* behavior.
|
||||||
*/
|
*/
|
||||||
flushEvents(): StreamMarkupHealingEvent[] {
|
flushEvents(): StreamMarkupHealingEvent[] {
|
||||||
const tail = this.#remaining();
|
return this.#convertScannerEvents(this.#scanner.flush());
|
||||||
this.#buffer = "";
|
|
||||||
this.#offset = 0;
|
|
||||||
|
|
||||||
switch (this.#pattern) {
|
|
||||||
case "kimi": {
|
|
||||||
const inTemplate = this.#kimiInCall || this.#kimiInSection;
|
|
||||||
this.#resetKimi();
|
|
||||||
return inTemplate || tail.length === 0 ? [] : [{ type: "text", text: tail }];
|
|
||||||
}
|
|
||||||
case "dsml": {
|
|
||||||
const state = this.#xmlState;
|
|
||||||
this.#xmlState = { kind: "idle" };
|
|
||||||
return state.kind !== "idle" || tail.length === 0 ? [] : [{ type: "text", text: tail }];
|
|
||||||
}
|
|
||||||
case "thinking": {
|
|
||||||
const closeTag = this.#thinkingCloseTag;
|
|
||||||
this.#thinkingCloseTag = "";
|
|
||||||
if (tail.length === 0) return [];
|
|
||||||
return closeTag ? [{ type: "thinking", thinking: tail }] : [{ type: "text", text: tail }];
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
|
|
||||||
/** Flush held-back text only. Reconstructed calls are retained for {@link drainCompleted}. */
|
/** Flush held-back text only. Reconstructed calls are retained for {@link drainCompleted}. */
|
||||||
@@ -222,393 +134,44 @@ export class StreamMarkupHealing {
|
|||||||
return this.#sectionTerminated;
|
return this.#sectionTerminated;
|
||||||
}
|
}
|
||||||
|
|
||||||
#remaining(): string {
|
#markSectionClosed(text: string): void {
|
||||||
return this.#offset === 0 ? this.#buffer : this.#buffer.slice(this.#offset);
|
if (this.#sectionTerminated) return;
|
||||||
}
|
if (this.#pattern === "kimi") {
|
||||||
|
this.#sectionTerminated = text.includes(KIMI_SECTION_END);
|
||||||
#compact(): void {
|
return;
|
||||||
if (this.#offset === 0) return;
|
|
||||||
this.#buffer = this.#buffer.slice(this.#offset);
|
|
||||||
this.#offset = 0;
|
|
||||||
}
|
|
||||||
|
|
||||||
#consumeKimiEvents(): StreamMarkupHealingEvent[] {
|
|
||||||
const events: StreamMarkupHealingEvent[] = [];
|
|
||||||
let clean = "";
|
|
||||||
const flushClean = (): void => {
|
|
||||||
if (clean.length === 0) return;
|
|
||||||
events.push({ type: "text", text: clean });
|
|
||||||
clean = "";
|
|
||||||
};
|
|
||||||
|
|
||||||
while (this.#offset < this.#buffer.length) {
|
|
||||||
if (this.#startsWithPartialToken(KIMI_TOKENS, MAX_KIMI_PARTIAL_HOLD)) break;
|
|
||||||
|
|
||||||
if (this.#matchesToken(KIMI_SECTION_BEGIN)) {
|
|
||||||
this.#kimiInSection = true;
|
|
||||||
this.#offset += KIMI_SECTION_BEGIN.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (this.#matchesToken(KIMI_SECTION_END)) {
|
|
||||||
this.#kimiInSection = false;
|
|
||||||
this.#sectionTerminated = true;
|
|
||||||
this.#offset += KIMI_SECTION_END.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (this.#matchesToken(KIMI_CALL_BEGIN)) {
|
|
||||||
if (!this.#kimiInSection) {
|
|
||||||
clean += KIMI_CALL_BEGIN;
|
|
||||||
this.#offset += KIMI_CALL_BEGIN.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
this.#kimiInCall = true;
|
|
||||||
this.#kimiInArgs = false;
|
|
||||||
this.#kimiPendingId = "";
|
|
||||||
this.#kimiPendingArgs = "";
|
|
||||||
this.#offset += KIMI_CALL_BEGIN.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (this.#matchesToken(KIMI_ARG_BEGIN)) {
|
|
||||||
if (!this.#kimiInSection) {
|
|
||||||
clean += KIMI_ARG_BEGIN;
|
|
||||||
this.#offset += KIMI_ARG_BEGIN.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
this.#kimiInArgs = true;
|
|
||||||
this.#offset += KIMI_ARG_BEGIN.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (this.#matchesToken(KIMI_CALL_END)) {
|
|
||||||
if (!this.#kimiInSection || !this.#kimiInCall) {
|
|
||||||
clean += KIMI_CALL_END;
|
|
||||||
this.#offset += KIMI_CALL_END.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
const call = this.#finalizeKimiCall();
|
|
||||||
flushClean();
|
|
||||||
events.push({ type: "toolCall", call });
|
|
||||||
this.#offset += KIMI_CALL_END.length;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
const ch = this.#buffer[this.#offset]!;
|
|
||||||
this.#offset += 1;
|
|
||||||
|
|
||||||
if (this.#kimiInCall) {
|
|
||||||
if (this.#kimiInArgs) {
|
|
||||||
this.#kimiPendingArgs += ch;
|
|
||||||
} else {
|
|
||||||
this.#kimiPendingId += ch;
|
|
||||||
}
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
if (!this.#kimiInSection) clean += ch;
|
|
||||||
}
|
}
|
||||||
|
this.#sectionTerminated =
|
||||||
flushClean();
|
text.includes(DSML_TOOL_CALLS_CLOSE_FULLWIDTH) || text.includes(DSML_TOOL_CALLS_CLOSE_ASCII);
|
||||||
return events;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
#consumeDsmlEvents(): StreamMarkupHealingEvent[] {
|
#convertScannerEvents(events: readonly InbandScanEvent[]): StreamMarkupHealingEvent[] {
|
||||||
return this.#consumeXmlToolEvents({
|
const out: StreamMarkupHealingEvent[] = [];
|
||||||
getState: () => this.#xmlState,
|
for (const event of events) {
|
||||||
setState: state => {
|
switch (event.type) {
|
||||||
this.#xmlState = state;
|
case "text":
|
||||||
},
|
out.push({ type: "text", text: event.text });
|
||||||
sectionOpen: DSML_TOOL_CALLS_OPEN_RE,
|
break;
|
||||||
sectionClose: DSML_TOOL_CALLS_CLOSE_RE,
|
case "thinkingDelta":
|
||||||
invokeOpen: DSML_INVOKE_OPEN_RE,
|
if (event.delta.length > 0) out.push({ type: "thinking", thinking: event.delta });
|
||||||
invokeClose: DSML_INVOKE_CLOSE_RE,
|
break;
|
||||||
parameterOpen: DSML_PARAMETER_OPEN_RE,
|
case "toolEnd":
|
||||||
parameterClose: DSML_PARAMETER_CLOSE_RE,
|
out.push({
|
||||||
coerceStringByDefault: true,
|
type: "toolCall",
|
||||||
});
|
call: {
|
||||||
}
|
id: generateHealedToolCallId(),
|
||||||
|
name: event.name,
|
||||||
#consumePlainThinkingEvents(): StreamMarkupHealingEvent[] {
|
arguments: JSON.stringify(event.arguments),
|
||||||
const events: StreamMarkupHealingEvent[] = [];
|
},
|
||||||
let clean = "";
|
|
||||||
let thinking = "";
|
|
||||||
const flushClean = (): void => {
|
|
||||||
if (clean.length === 0) return;
|
|
||||||
events.push({ type: "text", text: clean });
|
|
||||||
clean = "";
|
|
||||||
};
|
|
||||||
const flushThinking = (): void => {
|
|
||||||
if (thinking.length === 0) return;
|
|
||||||
events.push({ type: "thinking", thinking });
|
|
||||||
thinking = "";
|
|
||||||
};
|
|
||||||
|
|
||||||
while (this.#offset < this.#buffer.length) {
|
|
||||||
if (this.#thinkingCloseTag) {
|
|
||||||
if (this.#matchesToken(this.#thinkingCloseTag)) {
|
|
||||||
flushThinking();
|
|
||||||
this.#offset += this.#thinkingCloseTag.length;
|
|
||||||
this.#thinkingCloseTag = "";
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (this.#startsWithPartialToken([this.#thinkingCloseTag], MAX_XML_PARTIAL_HOLD)) break;
|
|
||||||
const ch = this.#buffer[this.#offset]!;
|
|
||||||
this.#offset += 1;
|
|
||||||
thinking += ch;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
const thinkingTag = this.#tryMatchThinkingOpen(PLAIN_THINKING_TAGS);
|
|
||||||
if (thinkingTag) {
|
|
||||||
flushClean();
|
|
||||||
this.#thinkingCloseTag = thinkingTag.close;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (this.#startsWithPartialThinkingOpen(PLAIN_THINKING_TAGS)) break;
|
|
||||||
|
|
||||||
const ch = this.#buffer[this.#offset]!;
|
|
||||||
this.#offset += 1;
|
|
||||||
clean += ch;
|
|
||||||
}
|
|
||||||
|
|
||||||
flushClean();
|
|
||||||
flushThinking();
|
|
||||||
return events;
|
|
||||||
}
|
|
||||||
|
|
||||||
#consumeXmlToolEvents(config: {
|
|
||||||
readonly getState: () => XmlToolState;
|
|
||||||
readonly setState: (state: XmlToolState) => void;
|
|
||||||
readonly sectionOpen: RegExp;
|
|
||||||
readonly sectionClose: RegExp;
|
|
||||||
readonly invokeOpen: RegExp;
|
|
||||||
readonly invokeClose: RegExp;
|
|
||||||
readonly parameterOpen: RegExp;
|
|
||||||
readonly parameterClose: RegExp;
|
|
||||||
readonly coerceStringByDefault: boolean;
|
|
||||||
}): StreamMarkupHealingEvent[] {
|
|
||||||
const events: StreamMarkupHealingEvent[] = [];
|
|
||||||
let clean = "";
|
|
||||||
const flushClean = (): void => {
|
|
||||||
if (clean.length === 0) return;
|
|
||||||
events.push({ type: "text", text: clean });
|
|
||||||
clean = "";
|
|
||||||
};
|
|
||||||
|
|
||||||
while (this.#offset < this.#buffer.length) {
|
|
||||||
const state = config.getState();
|
|
||||||
|
|
||||||
if (state.kind === "idle") {
|
|
||||||
if (this.#tryMatch(config.sectionOpen)) {
|
|
||||||
config.setState({ kind: "section" });
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
} else if (state.kind === "section") {
|
|
||||||
if (this.#tryMatch(config.sectionClose)) {
|
|
||||||
config.setState({ kind: "idle" });
|
|
||||||
this.#sectionTerminated = true;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
const invokeMatch = this.#tryMatchCapture(config.invokeOpen);
|
|
||||||
if (invokeMatch) {
|
|
||||||
config.setState({ kind: "invoke", name: invokeMatch[1] ?? "", args: {} });
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
} else if (state.kind === "invoke") {
|
|
||||||
if (this.#tryMatch(config.invokeClose)) {
|
|
||||||
const call = finalizeXmlToolCall(state.name, state.args);
|
|
||||||
flushClean();
|
|
||||||
events.push({ type: "toolCall", call });
|
|
||||||
config.setState({ kind: "section" });
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
const paramMatch = this.#tryMatchCapture(config.parameterOpen);
|
|
||||||
if (paramMatch) {
|
|
||||||
const stringAttr = paramMatch[2];
|
|
||||||
config.setState({
|
|
||||||
kind: "parameter",
|
|
||||||
invokeName: state.name,
|
|
||||||
args: state.args,
|
|
||||||
paramName: paramMatch[1] ?? "",
|
|
||||||
isString: config.coerceStringByDefault ? stringAttr !== "false" : false,
|
|
||||||
value: "",
|
|
||||||
});
|
});
|
||||||
continue;
|
break;
|
||||||
}
|
case "thinkingStart":
|
||||||
} else if (this.#tryMatch(config.parameterClose)) {
|
case "thinkingEnd":
|
||||||
// A capped value executes with silently corrupted input unless the
|
case "toolStart":
|
||||||
// truncation is made explicit — the marker fails JSON params loudly
|
case "toolArgDelta":
|
||||||
// and tells the model/tool what happened to string params.
|
break;
|
||||||
const paramValue = state.truncated
|
|
||||||
? `${state.value}\n…[parameter truncated: exceeded ${MAX_XML_PARAM_VALUE_LENGTH} bytes]`
|
|
||||||
: state.value;
|
|
||||||
state.args[state.paramName] = coerceXmlParamValue(paramValue, state.isString);
|
|
||||||
config.setState({ kind: "invoke", name: state.invokeName, args: state.args });
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
|
|
||||||
if (state.kind === "idle") {
|
|
||||||
// In idle, a bare `<` is legitimate output (`a < b`, generics, JSX).
|
|
||||||
// Only hold back tails that could still grow into the DSML
|
|
||||||
// section-open tag; everything else flows through immediately.
|
|
||||||
if (this.#startsWithPartialDsmlSectionOpen()) break;
|
|
||||||
} else if (this.#startsWithPartialXmlTag()) {
|
|
||||||
break;
|
|
||||||
}
|
|
||||||
|
|
||||||
const ch = this.#buffer[this.#offset]!;
|
|
||||||
this.#offset += 1;
|
|
||||||
if (state.kind === "idle") {
|
|
||||||
clean += ch;
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (state.kind === "parameter") {
|
|
||||||
if (state.value.length < MAX_XML_PARAM_VALUE_LENGTH) {
|
|
||||||
state.value += ch;
|
|
||||||
} else {
|
|
||||||
// Beyond the cap the value stops growing, but we stay in
|
|
||||||
// `parameter` state so the rest of the envelope — including its
|
|
||||||
// close tags — is still swallowed instead of leaking into
|
|
||||||
// visible text. The close handler appends an explicit marker.
|
|
||||||
state.truncated = true;
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
return out;
|
||||||
flushClean();
|
|
||||||
return events;
|
|
||||||
}
|
|
||||||
|
|
||||||
#tryMatch(pattern: RegExp): boolean {
|
|
||||||
pattern.lastIndex = this.#offset;
|
|
||||||
const match = pattern.exec(this.#buffer);
|
|
||||||
if (!match) return false;
|
|
||||||
this.#offset += match[0].length;
|
|
||||||
return true;
|
|
||||||
}
|
|
||||||
|
|
||||||
#tryMatchCapture(pattern: RegExp): RegExpExecArray | undefined {
|
|
||||||
pattern.lastIndex = this.#offset;
|
|
||||||
const match = pattern.exec(this.#buffer);
|
|
||||||
if (!match) return undefined;
|
|
||||||
this.#offset += match[0].length;
|
|
||||||
return match;
|
|
||||||
}
|
|
||||||
|
|
||||||
#tryMatchThinkingOpen(tags: readonly ThinkingTag[]): ThinkingTag | undefined {
|
|
||||||
for (const tag of tags) {
|
|
||||||
if (!this.#matchesToken(tag.open)) continue;
|
|
||||||
this.#offset += tag.open.length;
|
|
||||||
return tag;
|
|
||||||
}
|
|
||||||
return undefined;
|
|
||||||
}
|
|
||||||
|
|
||||||
#matchesToken(token: string): boolean {
|
|
||||||
return this.#buffer.startsWith(token, this.#offset);
|
|
||||||
}
|
|
||||||
|
|
||||||
#startsWithPartialThinkingOpen(tags: readonly ThinkingTag[]): boolean {
|
|
||||||
for (const tag of tags) {
|
|
||||||
if (this.#startsWithPartialToken([tag.open], MAX_XML_PARTIAL_HOLD)) return true;
|
|
||||||
}
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
|
|
||||||
#startsWithPartialToken(tokens: readonly string[], maxHold: number): boolean {
|
|
||||||
const remainingLength = this.#buffer.length - this.#offset;
|
|
||||||
if (remainingLength === 0 || remainingLength > maxHold) return false;
|
|
||||||
for (const token of tokens) {
|
|
||||||
if (token.length <= remainingLength) continue;
|
|
||||||
if (this.#bufferIsPrefixOf(token, remainingLength)) return true;
|
|
||||||
}
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
|
|
||||||
#startsWithPartialXmlTag(): boolean {
|
|
||||||
if (this.#buffer[this.#offset] !== "<") return false;
|
|
||||||
const tailLength = this.#buffer.length - this.#offset;
|
|
||||||
if (tailLength > MAX_XML_PARTIAL_HOLD) return false;
|
|
||||||
for (let i = this.#offset + 1; i < this.#buffer.length; i++) {
|
|
||||||
if (this.#buffer[i] === ">") return false;
|
|
||||||
}
|
|
||||||
return true;
|
|
||||||
}
|
|
||||||
|
|
||||||
#startsWithPartialDsmlSectionOpen(): boolean {
|
|
||||||
const tailLength = this.#buffer.length - this.#offset;
|
|
||||||
if (tailLength === 0 || tailLength >= DSML_SECTION_OPEN_TEMPLATE.length) return false;
|
|
||||||
for (let i = 0; i < tailLength; i++) {
|
|
||||||
const ch = this.#buffer[this.#offset + i]!;
|
|
||||||
const expected = DSML_SECTION_OPEN_TEMPLATE[i]!;
|
|
||||||
if (expected === "|") {
|
|
||||||
if (ch !== "|" && ch !== "|") return false;
|
|
||||||
} else if (ch !== expected) {
|
|
||||||
return false;
|
|
||||||
}
|
|
||||||
}
|
|
||||||
return true;
|
|
||||||
}
|
|
||||||
|
|
||||||
#bufferIsPrefixOf(token: string, remainingLength: number): boolean {
|
|
||||||
for (let i = 0; i < remainingLength; i++) {
|
|
||||||
if (this.#buffer[this.#offset + i] !== token[i]) return false;
|
|
||||||
}
|
|
||||||
return true;
|
|
||||||
}
|
|
||||||
|
|
||||||
#finalizeKimiCall(): HealedToolCall {
|
|
||||||
const rawId = this.#kimiPendingId.trim();
|
|
||||||
const rawArgs = this.#kimiPendingArgs.trim();
|
|
||||||
const name = normalizeKimiFunctionName(rawId);
|
|
||||||
|
|
||||||
let argsJson = rawArgs;
|
|
||||||
if (rawArgs.length > 0) {
|
|
||||||
try {
|
|
||||||
argsJson = JSON.stringify(parseJsonWithRepair<unknown>(rawArgs));
|
|
||||||
} catch {
|
|
||||||
// Leave raw; downstream parseStreamingJson absorbs the failure.
|
|
||||||
}
|
|
||||||
} else {
|
|
||||||
argsJson = "{}";
|
|
||||||
}
|
|
||||||
|
|
||||||
this.#kimiInCall = false;
|
|
||||||
this.#kimiInArgs = false;
|
|
||||||
this.#kimiPendingId = "";
|
|
||||||
this.#kimiPendingArgs = "";
|
|
||||||
return { id: generateHealedToolCallId(), name, arguments: argsJson };
|
|
||||||
}
|
|
||||||
|
|
||||||
#resetKimi(): void {
|
|
||||||
this.#kimiInSection = false;
|
|
||||||
this.#kimiInCall = false;
|
|
||||||
this.#kimiInArgs = false;
|
|
||||||
this.#kimiPendingId = "";
|
|
||||||
this.#kimiPendingArgs = "";
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
function normalizeKimiFunctionName(rawId: string): string {
|
|
||||||
const stripped = rawId.startsWith("functions.") ? rawId.slice("functions.".length) : rawId;
|
|
||||||
const colon = stripped.indexOf(":");
|
|
||||||
return colon >= 0 ? stripped.slice(0, colon) : stripped;
|
|
||||||
}
|
|
||||||
|
|
||||||
function finalizeXmlToolCall(name: string, args: Record<string, unknown>): HealedToolCall {
|
|
||||||
return {
|
|
||||||
id: generateHealedToolCallId(),
|
|
||||||
name: name.trim(),
|
|
||||||
arguments: JSON.stringify(args),
|
|
||||||
};
|
|
||||||
}
|
|
||||||
|
|
||||||
function coerceXmlParamValue(raw: string, isString: boolean): unknown {
|
|
||||||
if (isString) return raw;
|
|
||||||
const trimmed = raw.trim();
|
|
||||||
if (trimmed.length === 0) return raw;
|
|
||||||
try {
|
|
||||||
return parseJsonWithRepair<unknown>(trimmed);
|
|
||||||
} catch {
|
|
||||||
return raw;
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -11,9 +11,9 @@
|
|||||||
* 2. Normalizes LLM quirks (null / "null" → omit-or-default substitution)
|
* 2. Normalizes LLM quirks (null / "null" → omit-or-default substitution)
|
||||||
* against the JSON Schema before validation.
|
* against the JSON Schema before validation.
|
||||||
* 3. Validates with the Zod or JSON-Schema validator.
|
* 3. Validates with the Zod or JSON-Schema validator.
|
||||||
* 4. On failure, walks the resulting issues and coerces JSON-stringified
|
* 4. On failure, walks the resulting issues and coerces common LLM type
|
||||||
* values (`"[1,2]"` → `[1,2]`), drops unrecognized keys, and retries up
|
* drift (JSON-stringified values, boolean/number/string scalar drift),
|
||||||
* to `MAX_COERCION_PASSES` times.
|
* drops unrecognized keys, and retries up to `MAX_COERCION_PASSES` times.
|
||||||
* 5. Throws a formatted error if reconciliation fails; otherwise returns
|
* 5. Throws a formatted error if reconciliation fails; otherwise returns
|
||||||
* the parsed arguments with original unknown root fields preserved (so
|
* the parsed arguments with original unknown root fields preserved (so
|
||||||
* hallucinated top-level keys still surface to the caller).
|
* hallucinated top-level keys still surface to the caller).
|
||||||
@@ -38,20 +38,20 @@ import { isZodSchema, zodToWireSchema } from "./schema/wire";
|
|||||||
// Type Coercion Utilities
|
// Type Coercion Utilities
|
||||||
// ============================================================================
|
// ============================================================================
|
||||||
//
|
//
|
||||||
// LLMs sometimes produce tool arguments where a value that should be a number,
|
// LLMs sometimes produce tool arguments where a value has the right meaning but
|
||||||
// boolean, array, or object is instead passed as a JSON-encoded string. For
|
// the wrong JSON type. For example, an array parameter might arrive as
|
||||||
// example, an array parameter might arrive as `"[1, 2, 3]"` instead of `[1, 2, 3]`.
|
// `"[1, 2, 3]"`, a boolean as `"yes"` or `1`, or a string field as a structured
|
||||||
|
// object that should be embedded verbatim.
|
||||||
//
|
//
|
||||||
// Rather than rejecting these outright, we attempt automatic coercion:
|
// Rather than rejecting these outright, we attempt automatic coercion:
|
||||||
// 1. Validate against the tool's schema (Zod, derived from TypeBox when the
|
// 1. Validate against the tool's schema (Zod, derived from TypeBox when the
|
||||||
// tool was authored with TypeBox).
|
// tool was authored with TypeBox).
|
||||||
// 2. For each type error where the actual value is a string, we check if
|
// 2. For each type error, perform only the schema-directed rewrite that
|
||||||
// parsing it as JSON yields a value matching the expected type.
|
// matches the expected type.
|
||||||
// 3. If so, we replace the string with the parsed value and re-validate.
|
// 3. Re-validate the full argument object after each coercion pass.
|
||||||
//
|
//
|
||||||
// This is intentionally conservative: we only parse strings that look like
|
// This is intentionally conservative: each rewrite is small and validation
|
||||||
// valid JSON literals (objects, arrays, booleans, null, numbers) and only
|
// remains the source of truth for whether the result is accepted.
|
||||||
// accept the result if it matches the schema's expected type.
|
|
||||||
// ============================================================================
|
// ============================================================================
|
||||||
|
|
||||||
/** Regex matching valid JSON number literals (integers, decimals, scientific notation) */
|
/** Regex matching valid JSON number literals (integers, decimals, scientific notation) */
|
||||||
@@ -109,6 +109,85 @@ function tryParseNumberString(value: string, expectedTypes: string[]): { value:
|
|||||||
return { value: parsed, changed: true };
|
return { value: parsed, changed: true };
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function tryCoerceBoolean(value: unknown, expectedTypes: string[]): { value: unknown; changed: boolean } {
|
||||||
|
if (!expectedTypes.includes("boolean")) {
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
|
||||||
|
if (typeof value === "number") {
|
||||||
|
if (value === 0) return { value: false, changed: true };
|
||||||
|
if (value === 1) return { value: true, changed: true };
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
|
||||||
|
if (typeof value !== "string") {
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
|
||||||
|
switch (value.trim().toLowerCase()) {
|
||||||
|
case "true":
|
||||||
|
case "1":
|
||||||
|
case "yes":
|
||||||
|
case "on":
|
||||||
|
return { value: true, changed: true };
|
||||||
|
case "false":
|
||||||
|
case "0":
|
||||||
|
case "no":
|
||||||
|
case "off":
|
||||||
|
return { value: false, changed: true };
|
||||||
|
default:
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function tryCoerceBooleanToNumber(value: unknown, expectedTypes: string[]): { value: unknown; changed: boolean } {
|
||||||
|
if (!expectedTypes.includes("number") && !expectedTypes.includes("integer")) {
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
if (typeof value !== "boolean") {
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
return { value: value ? 1 : 0, changed: true };
|
||||||
|
}
|
||||||
|
|
||||||
|
function tryCoerceString(value: unknown, expectedTypes: string[]): { value: unknown; changed: boolean } {
|
||||||
|
if (!expectedTypes.includes("string") || typeof value === "string" || value === null || value === undefined) {
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
|
||||||
|
if (Array.isArray(value) || typeof value === "object") {
|
||||||
|
try {
|
||||||
|
const stringified = JSON.stringify(value);
|
||||||
|
if (stringified === undefined) return { value, changed: false };
|
||||||
|
return { value: stringified, changed: true };
|
||||||
|
} catch {
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (typeof value === "function") {
|
||||||
|
return { value, changed: false };
|
||||||
|
}
|
||||||
|
|
||||||
|
return { value: String(value), changed: true };
|
||||||
|
}
|
||||||
|
|
||||||
|
function tryCoerceForExpectedTypes(value: unknown, expectedTypes: string[]): { value: unknown; changed: boolean } {
|
||||||
|
if (typeof value === "string") {
|
||||||
|
const parsed = tryParseJsonForTypes(value, expectedTypes);
|
||||||
|
if (parsed.changed) return parsed;
|
||||||
|
return tryCoerceBoolean(value, expectedTypes);
|
||||||
|
}
|
||||||
|
|
||||||
|
const booleanCoercion = tryCoerceBoolean(value, expectedTypes);
|
||||||
|
if (booleanCoercion.changed) return booleanCoercion;
|
||||||
|
|
||||||
|
const numericCoercion = tryCoerceBooleanToNumber(value, expectedTypes);
|
||||||
|
if (numericCoercion.changed) return numericCoercion;
|
||||||
|
|
||||||
|
return tryCoerceString(value, expectedTypes);
|
||||||
|
}
|
||||||
|
|
||||||
function tryParseLeadingJsonContainer(value: string): unknown | undefined {
|
function tryParseLeadingJsonContainer(value: string): unknown | undefined {
|
||||||
const firstChar = value[0];
|
const firstChar = value[0];
|
||||||
const closingChar = firstChar === "{" ? "}" : firstChar === "[" ? "]" : undefined;
|
const closingChar = firstChar === "{" ? "}" : firstChar === "[" ? "]" : undefined;
|
||||||
@@ -806,10 +885,11 @@ function flattenIssues(issues: ReadonlyArray<ZodIssue>): FlatIssue[] {
|
|||||||
* Repair issues raised by the validator before we surface them to the caller.
|
* Repair issues raised by the validator before we surface them to the caller.
|
||||||
*
|
*
|
||||||
* Two kinds of repair are applied:
|
* Two kinds of repair are applied:
|
||||||
* - **type**: when a value is a JSON-encoded string and the schema wants
|
* - **type**: when a value has a common LLM-produced shape mismatch, rewrite
|
||||||
* something else, parse it and substitute the parsed value. When a
|
* it only in the direction requested by the schema: parse JSON strings,
|
||||||
* non-union schema wants an array but receives a singleton value, wrap that
|
* accept boolean spellings, stringify non-null values for string fields,
|
||||||
* value in a one-element array.
|
* map booleans to numeric 0/1, and wrap singleton array values for non-union
|
||||||
|
* array expectations.
|
||||||
* - **unrecognized**: when a strict object received an extra key (Zod's
|
* - **unrecognized**: when a strict object received an extra key (Zod's
|
||||||
* `unrecognized_keys` or JSON Schema's `additionalProperties: false`),
|
* `unrecognized_keys` or JSON Schema's `additionalProperties: false`),
|
||||||
* drop that key so re-validation succeeds. This effectively coerces every
|
* drop that key so re-validation succeeds. This effectively coerces every
|
||||||
@@ -818,9 +898,8 @@ function flattenIssues(issues: ReadonlyArray<ZodIssue>): FlatIssue[] {
|
|||||||
*
|
*
|
||||||
* The function is safe and conservative:
|
* The function is safe and conservative:
|
||||||
* - Only processes "type" and "unrecognized" issues
|
* - Only processes "type" and "unrecognized" issues
|
||||||
* - Only attempts JSON coercion on string values
|
* - Only attempts schema-directed coercions for the expected type
|
||||||
* - Only wraps singleton array values for non-union type expectations
|
* - Only wraps singleton array values for non-union type expectations
|
||||||
* - Only accepts parsed results that match the expected type
|
|
||||||
* - Clones the args object before mutation (copy-on-write)
|
* - Clones the args object before mutation (copy-on-write)
|
||||||
*/
|
*/
|
||||||
function coerceArgsFromIssues(args: unknown, issues: FlatIssue[]): { value: unknown; changed: boolean } {
|
function coerceArgsFromIssues(args: unknown, issues: FlatIssue[]): { value: unknown; changed: boolean } {
|
||||||
@@ -845,10 +924,7 @@ function coerceArgsFromIssues(args: unknown, issues: FlatIssue[]): { value: unkn
|
|||||||
if (issue.expectedTypes.length === 0) continue;
|
if (issue.expectedTypes.length === 0) continue;
|
||||||
|
|
||||||
const currentValue = getValueAtPointer(nextArgs, issue.instancePath);
|
const currentValue = getValueAtPointer(nextArgs, issue.instancePath);
|
||||||
const result =
|
const result = tryCoerceForExpectedTypes(currentValue, issue.expectedTypes);
|
||||||
typeof currentValue === "string"
|
|
||||||
? tryParseJsonForTypes(currentValue, issue.expectedTypes)
|
|
||||||
: { value: currentValue, changed: false };
|
|
||||||
const coercedValue = result.changed
|
const coercedValue = result.changed
|
||||||
? result.value
|
? result.value
|
||||||
: issue.expectedTypes.includes("array") &&
|
: issue.expectedTypes.includes("array") &&
|
||||||
|
|||||||
@@ -1,4 +1,5 @@
|
|||||||
import { describe, expect, it } from "bun:test";
|
import { describe, expect, it } from "bun:test";
|
||||||
|
import type { AssistantMessage, Model, ToolCall, Usage } from "@oh-my-pi/pi-ai";
|
||||||
import {
|
import {
|
||||||
createHarmonyAuditEvent,
|
createHarmonyAuditEvent,
|
||||||
detectHarmonyLeak,
|
detectHarmonyLeak,
|
||||||
@@ -7,11 +8,9 @@ import {
|
|||||||
isHarmonyLeakMitigationTarget,
|
isHarmonyLeakMitigationTarget,
|
||||||
recoverHarmonyToolCall,
|
recoverHarmonyToolCall,
|
||||||
signalListLabel,
|
signalListLabel,
|
||||||
} from "@oh-my-pi/pi-agent-core/harmony-leak";
|
} from "@oh-my-pi/pi-ai/utils/harmony-leak";
|
||||||
import type { AssistantMessage, Model, ToolCall } from "@oh-my-pi/pi-ai";
|
|
||||||
import { getBundledModel } from "@oh-my-pi/pi-catalog/models";
|
import { getBundledModel } from "@oh-my-pi/pi-catalog/models";
|
||||||
import corpus from "./fixtures/harmony-leak-corpus.json" with { type: "json" };
|
import corpus from "./fixtures/harmony-leak-corpus.json" with { type: "json" };
|
||||||
import { createAssistantMessage } from "./helpers";
|
|
||||||
|
|
||||||
interface CorpusPositive {
|
interface CorpusPositive {
|
||||||
id: string;
|
id: string;
|
||||||
@@ -30,6 +29,33 @@ const negatives = corpus.negatives as CorpusNegative[];
|
|||||||
const codexModel: Model = getBundledModel("openai-codex", "gpt-5.4");
|
const codexModel: Model = getBundledModel("openai-codex", "gpt-5.4");
|
||||||
const anthropicModel: Model = getBundledModel("anthropic", "claude-sonnet-4-5");
|
const anthropicModel: Model = getBundledModel("anthropic", "claude-sonnet-4-5");
|
||||||
|
|
||||||
|
function createAssistantMessage(
|
||||||
|
content: AssistantMessage["content"],
|
||||||
|
stopReason: AssistantMessage["stopReason"] = "stop",
|
||||||
|
): AssistantMessage {
|
||||||
|
return {
|
||||||
|
role: "assistant",
|
||||||
|
content,
|
||||||
|
api: "mock",
|
||||||
|
provider: "mock",
|
||||||
|
model: "mock-model",
|
||||||
|
usage: createUsage(),
|
||||||
|
stopReason,
|
||||||
|
timestamp: Date.now(),
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function createUsage(): Usage {
|
||||||
|
return {
|
||||||
|
input: 0,
|
||||||
|
output: 0,
|
||||||
|
cacheRead: 0,
|
||||||
|
cacheWrite: 0,
|
||||||
|
totalTokens: 0,
|
||||||
|
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 },
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
function makeToolCallMessage(toolName: string, input: string | null, argJson: string | null): AssistantMessage {
|
function makeToolCallMessage(toolName: string, input: string | null, argJson: string | null): AssistantMessage {
|
||||||
const callArgs: Record<string, unknown> =
|
const callArgs: Record<string, unknown> =
|
||||||
input !== null ? { input } : argJson !== null ? (JSON.parse(argJson) as Record<string, unknown>) : {};
|
input !== null ? { input } : argJson !== null ? (JSON.parse(argJson) as Record<string, unknown>) : {};
|
||||||
@@ -0,0 +1,254 @@
|
|||||||
|
import { describe, expect, it } from "bun:test";
|
||||||
|
import type { AssistantMessage, Context, ToolCall, ToolResultMessage, Usage } from "@oh-my-pi/pi-ai";
|
||||||
|
import {
|
||||||
|
createInbandScanner,
|
||||||
|
encodeInbandToolHistory,
|
||||||
|
type GrammarToolResult,
|
||||||
|
getInbandGrammar,
|
||||||
|
type InbandScanEvent,
|
||||||
|
parseInbandToolMessage,
|
||||||
|
renderInbandToolPrompt,
|
||||||
|
type ToolCallSyntax,
|
||||||
|
} from "@oh-my-pi/pi-ai/grammar";
|
||||||
|
|
||||||
|
const TOOLS = [
|
||||||
|
{
|
||||||
|
name: "read",
|
||||||
|
description: "Read a file",
|
||||||
|
parameters: {
|
||||||
|
type: "object",
|
||||||
|
properties: { path: { type: "string" }, count: { type: "number" } },
|
||||||
|
required: ["path"],
|
||||||
|
},
|
||||||
|
},
|
||||||
|
{
|
||||||
|
name: "write",
|
||||||
|
description: "Write a file",
|
||||||
|
parameters: {
|
||||||
|
type: "object",
|
||||||
|
properties: { path: { type: "string" }, content: { type: "string" } },
|
||||||
|
required: ["path", "content"],
|
||||||
|
},
|
||||||
|
},
|
||||||
|
] as unknown as NonNullable<Context["tools"]>;
|
||||||
|
|
||||||
|
const SYNTAXES: readonly ToolCallSyntax[] = [
|
||||||
|
"glm",
|
||||||
|
"hermes",
|
||||||
|
"kimi",
|
||||||
|
"xml",
|
||||||
|
"anthropic",
|
||||||
|
"deepseek",
|
||||||
|
"harmony",
|
||||||
|
"pi",
|
||||||
|
"qwen3",
|
||||||
|
];
|
||||||
|
|
||||||
|
function usage(): Usage {
|
||||||
|
return {
|
||||||
|
input: 0,
|
||||||
|
output: 0,
|
||||||
|
cacheRead: 0,
|
||||||
|
cacheWrite: 0,
|
||||||
|
totalTokens: 0,
|
||||||
|
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 },
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function assistant(content: AssistantMessage["content"]): AssistantMessage {
|
||||||
|
return {
|
||||||
|
role: "assistant",
|
||||||
|
content,
|
||||||
|
api: "mock",
|
||||||
|
provider: "mock",
|
||||||
|
model: "mock-model",
|
||||||
|
usage: usage(),
|
||||||
|
stopReason: "toolUse",
|
||||||
|
timestamp: 0,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function result(toolCallId: string, toolName: string, text: string, isError = false): ToolResultMessage {
|
||||||
|
return { role: "toolResult", toolCallId, toolName, content: [{ type: "text", text }], isError, timestamp: 0 };
|
||||||
|
}
|
||||||
|
|
||||||
|
function feedText(syntax: ToolCallSyntax, text: string): InbandScanEvent[] {
|
||||||
|
const scanner = createInbandScanner(syntax, { tools: TOOLS, parseThinking: true });
|
||||||
|
const events: InbandScanEvent[] = [];
|
||||||
|
for (const char of text) events.push(...scanner.feed(char));
|
||||||
|
events.push(...scanner.flush());
|
||||||
|
return events;
|
||||||
|
}
|
||||||
|
|
||||||
|
function toolEnds(events: readonly InbandScanEvent[]): Extract<InbandScanEvent, { type: "toolEnd" }>[] {
|
||||||
|
return events.filter((event): event is Extract<InbandScanEvent, { type: "toolEnd" }> => event.type === "toolEnd");
|
||||||
|
}
|
||||||
|
|
||||||
|
function firstRawBlock(syntax: ToolCallSyntax, text: string): string | undefined {
|
||||||
|
return toolEnds(feedText(syntax, text))[0]?.rawBlock;
|
||||||
|
}
|
||||||
|
|
||||||
|
function expectRawBlock(syntax: ToolCallSyntax, text: string, expected: string): void {
|
||||||
|
expect(firstRawBlock(syntax, text), syntax).toBe(expected);
|
||||||
|
}
|
||||||
|
|
||||||
|
describe("in-band tool grammars", () => {
|
||||||
|
it("renders a tool prompt for every syntax", () => {
|
||||||
|
for (const syntax of SYNTAXES) {
|
||||||
|
const prompt = renderInbandToolPrompt(TOOLS, syntax);
|
||||||
|
expect(prompt).toContain("<tools>");
|
||||||
|
expect(prompt).toContain("</tools>");
|
||||||
|
expect(prompt).toContain('"name":"read"');
|
||||||
|
expect(prompt).toContain(getInbandGrammar(syntax).prompt.trim().split("\n", 1)[0]!);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
it("each grammar renders calls that its scanner parses back", () => {
|
||||||
|
const call: ToolCall = {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "functions.read:0",
|
||||||
|
name: "read",
|
||||||
|
arguments: { path: "src/a.ts", count: 2 },
|
||||||
|
};
|
||||||
|
for (const syntax of SYNTAXES) {
|
||||||
|
const grammar = getInbandGrammar(syntax);
|
||||||
|
const rendered = grammar.renderAssistantToolCalls([call], { tools: TOOLS });
|
||||||
|
const calls = toolEnds(feedText(syntax, rendered));
|
||||||
|
expect(calls, syntax).toHaveLength(1);
|
||||||
|
expect(calls[0]!.name).toBe("read");
|
||||||
|
expect(calls[0]!.arguments).toEqual({ path: "src/a.ts", count: 2 });
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|
||||||
|
it("captures exact raw tool call blocks for debugging", () => {
|
||||||
|
expectRawBlock(
|
||||||
|
"glm",
|
||||||
|
"<tool_call>read\n<arg_key>path</arg_key>\n<arg_value>src/a.ts</arg_value>\n</tool_call>",
|
||||||
|
"<tool_call>read\n<arg_key>path</arg_key>\n<arg_value>src/a.ts</arg_value>\n</tool_call>",
|
||||||
|
);
|
||||||
|
expectRawBlock(
|
||||||
|
"kimi",
|
||||||
|
'<|tool_calls_section_begin|><|tool_call_begin|>functions.read:0<|tool_call_argument_begin|> {"path":"src/a.ts"}\n<|tool_call_end|><|tool_calls_section_end|>',
|
||||||
|
'<|tool_call_begin|>functions.read:0<|tool_call_argument_begin|> {"path":"src/a.ts"}\n<|tool_call_end|>',
|
||||||
|
);
|
||||||
|
expectRawBlock(
|
||||||
|
"deepseek",
|
||||||
|
'<|DSML|tool_calls>\n<|DSML|invoke name="read">\n <|DSML|parameter name="path" string="true">src/a.ts</|DSML|parameter>\n</|DSML|invoke>\n</|DSML|tool_calls>',
|
||||||
|
'<|DSML|invoke name="read">\n <|DSML|parameter name="path" string="true">src/a.ts</|DSML|parameter>\n</|DSML|invoke>',
|
||||||
|
);
|
||||||
|
expectRawBlock(
|
||||||
|
"xml",
|
||||||
|
'<function_calls>\n<invoke name="read"><parameter name="path" string="true">src/a.ts</parameter></invoke>\n</function_calls>',
|
||||||
|
'<invoke name="read"><parameter name="path" string="true">src/a.ts</parameter></invoke>',
|
||||||
|
);
|
||||||
|
expectRawBlock(
|
||||||
|
"harmony",
|
||||||
|
'<|start|>assistant<|channel|>commentary to=functions.read <|constrain|>json<|message|>{"path":"src/a.ts"}<|call|>',
|
||||||
|
'<|start|>assistant<|channel|>commentary to=functions.read <|constrain|>json<|message|>{"path":"src/a.ts"}<|call|>',
|
||||||
|
);
|
||||||
|
expectRawBlock(
|
||||||
|
"pi",
|
||||||
|
'<call:write path="out.ts">\nhello\n</call:write>',
|
||||||
|
'<call:write path="out.ts">\nhello\n</call:write>',
|
||||||
|
);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("projects raw tool blocks onto parsed ToolCall content", () => {
|
||||||
|
const raw =
|
||||||
|
'<|start|>assistant<|channel|>commentary to=functions.read <|constrain|>json<|message|>{"path":"src/a.ts"}<|call|>';
|
||||||
|
const parsed = parseInbandToolMessage(assistant([{ type: "text", text: raw }]), "harmony", TOOLS);
|
||||||
|
const call = parsed.content.find((block): block is ToolCall => block.type === "toolCall");
|
||||||
|
|
||||||
|
expect(call?.rawBlock).toBe(raw);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("stops before hallucinated Anthropic function results", () => {
|
||||||
|
const parsed = parseInbandToolMessage(
|
||||||
|
assistant([
|
||||||
|
{
|
||||||
|
type: "text",
|
||||||
|
text: '<invoke name="read"><parameter name="path">rubygems.ts:85-93</parameter></invoke>\n<function_results>\n<result>\n<tool_name>read</tool_name>\n<stdout>[rubygems.ts#A1B2]</stdout>\n</result>\n</function_results>\n<invoke name="edit"><parameter name="input">[rubygems.ts#A1B2]\nXCHG 89..89:\n+ fake</parameter></invoke>',
|
||||||
|
},
|
||||||
|
]),
|
||||||
|
"anthropic",
|
||||||
|
TOOLS,
|
||||||
|
);
|
||||||
|
const calls = parsed.content.filter((block): block is ToolCall => block.type === "toolCall");
|
||||||
|
|
||||||
|
expect(calls.map(call => call.name)).toEqual(["read"]);
|
||||||
|
expect(calls[0]?.arguments).toEqual({ path: "rubygems.ts:85-93" });
|
||||||
|
});
|
||||||
|
|
||||||
|
it("keeps result rendering in the owning grammar", () => {
|
||||||
|
const resultBlock: GrammarToolResult = {
|
||||||
|
id: "functions.read:0",
|
||||||
|
name: "read",
|
||||||
|
index: 0,
|
||||||
|
text: "FILE",
|
||||||
|
isError: false,
|
||||||
|
};
|
||||||
|
expect(getInbandGrammar("glm").renderToolResults([resultBlock])).toBe(
|
||||||
|
"<observation>\n<tool_response>\nFILE\n</tool_response>\n</observation>",
|
||||||
|
);
|
||||||
|
expect(getInbandGrammar("deepseek").renderToolResults([resultBlock])).toBe(
|
||||||
|
"<|tool▁output▁begin|>FILE<|tool▁output▁end|>",
|
||||||
|
);
|
||||||
|
expect(getInbandGrammar("kimi").renderToolResults([resultBlock])).toBe(
|
||||||
|
"<|im_system|>read<|im_middle|>## Return of functions.read:0\nFILE<|im_end|>",
|
||||||
|
);
|
||||||
|
expect(getInbandGrammar("harmony").renderToolResults([resultBlock])).toBe(
|
||||||
|
"<|start|>functions.read to=assistant<|channel|>commentary<|message|>FILE<|end|>",
|
||||||
|
);
|
||||||
|
expect(getInbandGrammar("anthropic").renderToolResults([resultBlock])).toBe(
|
||||||
|
"<function_results>\n<result>\n<tool_name>read</tool_name>\n<stdout>FILE</stdout>\n</result>\n</function_results>",
|
||||||
|
);
|
||||||
|
expect(getInbandGrammar("qwen3").renderToolResults([resultBlock])).toBe(
|
||||||
|
"<tool_response>\nFILE\n</tool_response>",
|
||||||
|
);
|
||||||
|
expect(getInbandGrammar("pi").renderToolResults([resultBlock])).toBe("<tool_response>\nFILE\n</tool_response>");
|
||||||
|
});
|
||||||
|
|
||||||
|
it("encodes assistant calls and tool results through the selected grammar", () => {
|
||||||
|
const history: Context["messages"] = [
|
||||||
|
{ role: "user", content: "hi", timestamp: 0 },
|
||||||
|
assistant([
|
||||||
|
{ type: "text", text: "let me read" },
|
||||||
|
{ type: "toolCall", id: "functions.read:0", name: "read", arguments: { path: "a.ts" } },
|
||||||
|
]),
|
||||||
|
result("functions.read:0", "read", "FILE A"),
|
||||||
|
];
|
||||||
|
const enc = encodeInbandToolHistory(history, "kimi", TOOLS);
|
||||||
|
expect(enc[0]).toBe(history[0]);
|
||||||
|
expect(enc[1]!.role).toBe("assistant");
|
||||||
|
expect(enc[2]!.role).toBe("user");
|
||||||
|
const assistantBlock = (enc[1] as AssistantMessage).content[0]!;
|
||||||
|
const assistantText = assistantBlock.type === "text" ? assistantBlock.text : "";
|
||||||
|
expect(assistantText).toContain("<|tool_calls_section_begin|>");
|
||||||
|
expect(assistantText).toContain("functions.read:0");
|
||||||
|
const resultText =
|
||||||
|
Array.isArray(enc[2]!.content) && enc[2]!.content[0]!.type === "text" ? enc[2]!.content[0]!.text : "";
|
||||||
|
expect(resultText).toBe("<|im_system|>read<|im_middle|>## Return of functions.read:0\nFILE A<|im_end|>");
|
||||||
|
});
|
||||||
|
|
||||||
|
it("streams string arguments incrementally for GLM", () => {
|
||||||
|
const text = getInbandGrammar("glm").renderAssistantToolCalls(
|
||||||
|
[
|
||||||
|
{
|
||||||
|
type: "toolCall",
|
||||||
|
id: "c1",
|
||||||
|
name: "write",
|
||||||
|
arguments: { path: "out.ts", content: "line1\nconst x = `a`;" },
|
||||||
|
},
|
||||||
|
],
|
||||||
|
{ tools: TOOLS },
|
||||||
|
);
|
||||||
|
const deltas = feedText("glm", text)
|
||||||
|
.filter(
|
||||||
|
(event): event is Extract<InbandScanEvent, { type: "toolArgDelta" }> =>
|
||||||
|
event.type === "toolArgDelta" && event.key === "content",
|
||||||
|
)
|
||||||
|
.map(event => event.delta)
|
||||||
|
.join("");
|
||||||
|
expect(deltas).toBe("line1\nconst x = `a`;");
|
||||||
|
});
|
||||||
|
});
|
||||||
@@ -42,6 +42,163 @@ describe("Tool argument coercion", () => {
|
|||||||
expect(typeof result.label).toBe("string");
|
expect(typeof result.label).toBe("string");
|
||||||
});
|
});
|
||||||
|
|
||||||
|
it("stringifies object values when schema expects string", () => {
|
||||||
|
const tool: Tool = {
|
||||||
|
name: "object-string",
|
||||||
|
description: "",
|
||||||
|
parameters: z.object({ payload: z.string() }),
|
||||||
|
};
|
||||||
|
|
||||||
|
const result = validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-object-string",
|
||||||
|
name: "object-string",
|
||||||
|
arguments: { payload: { a: 1, nested: ["x"] } },
|
||||||
|
}) as { payload: string };
|
||||||
|
|
||||||
|
expect(result.payload).toBe('{"a":1,"nested":["x"]}');
|
||||||
|
});
|
||||||
|
|
||||||
|
it("stringifies array values when schema expects string", () => {
|
||||||
|
const tool: Tool = {
|
||||||
|
name: "array-string",
|
||||||
|
description: "",
|
||||||
|
parameters: z.object({ payload: z.string() }),
|
||||||
|
};
|
||||||
|
|
||||||
|
const result = validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-array-string",
|
||||||
|
name: "array-string",
|
||||||
|
arguments: { payload: ["a", 2, true] },
|
||||||
|
}) as { payload: string };
|
||||||
|
|
||||||
|
expect(result.payload).toBe('["a",2,true]');
|
||||||
|
});
|
||||||
|
|
||||||
|
it("coerces numeric 0 and 1 to booleans", () => {
|
||||||
|
const tool: Tool = {
|
||||||
|
name: "numeric-booleans",
|
||||||
|
description: "",
|
||||||
|
parameters: z.object({ enabled: z.boolean(), disabled: z.boolean() }),
|
||||||
|
};
|
||||||
|
|
||||||
|
const result = validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-numeric-booleans",
|
||||||
|
name: "numeric-booleans",
|
||||||
|
arguments: { enabled: 1, disabled: 0 },
|
||||||
|
}) as { enabled: boolean; disabled: boolean };
|
||||||
|
|
||||||
|
expect(result).toEqual({ enabled: true, disabled: false });
|
||||||
|
});
|
||||||
|
|
||||||
|
it("coerces booleans to numeric 0 and 1", () => {
|
||||||
|
const tool: Tool = {
|
||||||
|
name: "boolean-numbers",
|
||||||
|
description: "",
|
||||||
|
parameters: z.object({ enabled: z.number(), disabled: z.number().int() }),
|
||||||
|
};
|
||||||
|
|
||||||
|
const result = validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-boolean-numbers",
|
||||||
|
name: "boolean-numbers",
|
||||||
|
arguments: { enabled: true, disabled: false },
|
||||||
|
}) as { enabled: number; disabled: number };
|
||||||
|
|
||||||
|
expect(result).toEqual({ enabled: 1, disabled: 0 });
|
||||||
|
});
|
||||||
|
|
||||||
|
it("rejects numeric boolean values other than 0 or 1", () => {
|
||||||
|
const tool: Tool = {
|
||||||
|
name: "invalid-numeric-boolean",
|
||||||
|
description: "",
|
||||||
|
parameters: z.object({ enabled: z.boolean() }),
|
||||||
|
};
|
||||||
|
|
||||||
|
expect(() =>
|
||||||
|
validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-invalid-numeric-boolean",
|
||||||
|
name: "invalid-numeric-boolean",
|
||||||
|
arguments: { enabled: 2 },
|
||||||
|
}),
|
||||||
|
).toThrow('Validation failed for tool "invalid-numeric-boolean"');
|
||||||
|
});
|
||||||
|
|
||||||
|
it("keeps raw in-band blocks out of validation errors", () => {
|
||||||
|
const tool: Tool = {
|
||||||
|
name: "raw-debug",
|
||||||
|
description: "",
|
||||||
|
parameters: z.object({ input: z.string() }),
|
||||||
|
};
|
||||||
|
const rawBlock = '<|start|>assistant<|channel|>commentary to=functions.edit <|message|>{"input":"x"}}<|call|>';
|
||||||
|
|
||||||
|
expect(() =>
|
||||||
|
validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-raw-debug",
|
||||||
|
name: "raw-debug",
|
||||||
|
arguments: {},
|
||||||
|
rawBlock,
|
||||||
|
}),
|
||||||
|
).toThrow('Validation failed for tool "raw-debug"');
|
||||||
|
expect(() =>
|
||||||
|
validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-raw-debug",
|
||||||
|
name: "raw-debug",
|
||||||
|
arguments: {},
|
||||||
|
rawBlock,
|
||||||
|
}),
|
||||||
|
).not.toThrow(rawBlock);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("coerces common string boolean forms", () => {
|
||||||
|
const tool: Tool = {
|
||||||
|
name: "string-booleans",
|
||||||
|
description: "",
|
||||||
|
parameters: z.object({
|
||||||
|
t: z.boolean(),
|
||||||
|
f: z.boolean(),
|
||||||
|
one: z.boolean(),
|
||||||
|
zero: z.boolean(),
|
||||||
|
yes: z.boolean(),
|
||||||
|
no: z.boolean(),
|
||||||
|
on: z.boolean(),
|
||||||
|
off: z.boolean(),
|
||||||
|
}),
|
||||||
|
};
|
||||||
|
|
||||||
|
const result = validateToolArguments(tool, {
|
||||||
|
type: "toolCall",
|
||||||
|
id: "call-string-booleans",
|
||||||
|
name: "string-booleans",
|
||||||
|
arguments: {
|
||||||
|
t: "TRUE",
|
||||||
|
f: "false",
|
||||||
|
one: "1",
|
||||||
|
zero: "0",
|
||||||
|
yes: "yes",
|
||||||
|
no: "NO",
|
||||||
|
on: "on",
|
||||||
|
off: "OFF",
|
||||||
|
},
|
||||||
|
}) as Record<string, boolean>;
|
||||||
|
|
||||||
|
expect(result).toEqual({
|
||||||
|
t: true,
|
||||||
|
f: false,
|
||||||
|
one: true,
|
||||||
|
zero: false,
|
||||||
|
yes: true,
|
||||||
|
no: false,
|
||||||
|
on: true,
|
||||||
|
off: false,
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
it("parses JSON arrays in string values when schema expects array", () => {
|
it("parses JSON arrays in string values when schema expects array", () => {
|
||||||
const tool: Tool = {
|
const tool: Tool = {
|
||||||
name: "t3",
|
name: "t3",
|
||||||
|
|||||||
@@ -1,6 +1,19 @@
|
|||||||
# Changelog
|
# Changelog
|
||||||
|
|
||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
### Added
|
||||||
|
|
||||||
|
- Added `flux-1-schnell-fp8` to the Fireworks serverless model catalog
|
||||||
|
- Added `gpt-oss-20b` to the Fireworks model catalog
|
||||||
|
- Added `qwen3-embedding-8b` to the Fireworks model catalog
|
||||||
|
- Added `qwen3-reranker-8b` to the Fireworks model catalog
|
||||||
|
- Added `Gemma 4 E2B IT` and `Gemma 4 E4B IT` to the Google model catalog
|
||||||
|
- Added `qwen/qwen3-asr-flash` to the Zenmux model catalog
|
||||||
|
- Added sparse `supportsTools` model metadata so providers can mark models that require in-band tool-call formatting.
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
|
||||||
|
- Kept non-tool-capable Fireworks serverless models in discovery results and marked them with `supportsTools: false` for fallback-aware handling
|
||||||
|
|
||||||
## [15.13.1] - 2026-06-15
|
## [15.13.1] - 2026-06-15
|
||||||
|
|
||||||
@@ -167,4 +180,4 @@
|
|||||||
|
|
||||||
### Removed
|
### Removed
|
||||||
|
|
||||||
- Removed the runtime enrichment layer: `enrichModelThinking` (and its non-enumerable memo-slot cache), `refreshModelThinking`, `modelOmitsReasoningEffort`, and the `model-thinking` re-exports of generator-only policies. Thinking metadata is resolved exactly once inside `buildModel`; runtime helpers (`getSupportedEfforts`, `clampThinkingLevelForModel`, `requireSupportedEffort`, the effort mappers) are pure field reads.
|
- Removed the runtime enrichment layer: `enrichModelThinking` (and its non-enumerable memo-slot cache), `refreshModelThinking`, `modelOmitsReasoningEffort`, and the `model-thinking` re-exports of generator-only policies. Thinking metadata is resolved exactly once inside `buildModel`; runtime helpers (`getSupportedEfforts`, `clampThinkingLevelForModel`, `requireSupportedEffort`, the effort mappers) are pure field reads.
|
||||||
@@ -13692,6 +13692,39 @@
|
|||||||
"requiresAssistantContentForToolCalls": true
|
"requiresAssistantContentForToolCalls": true
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
"flux-1-schnell-fp8": {
|
||||||
|
"id": "flux-1-schnell-fp8",
|
||||||
|
"name": "FLUX.1 [schnell] FP8",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"provider": "fireworks",
|
||||||
|
"baseUrl": "https://api.fireworks.ai/inference/v1",
|
||||||
|
"reasoning": true,
|
||||||
|
"input": [
|
||||||
|
"text"
|
||||||
|
],
|
||||||
|
"cost": {
|
||||||
|
"input": 0,
|
||||||
|
"output": 0,
|
||||||
|
"cacheRead": 0,
|
||||||
|
"cacheWrite": 0
|
||||||
|
},
|
||||||
|
"contextWindow": null,
|
||||||
|
"maxTokens": null,
|
||||||
|
"supportsTools": false,
|
||||||
|
"thinking": {
|
||||||
|
"mode": "effort",
|
||||||
|
"efforts": [
|
||||||
|
"minimal",
|
||||||
|
"low",
|
||||||
|
"medium",
|
||||||
|
"high",
|
||||||
|
"xhigh"
|
||||||
|
],
|
||||||
|
"effortMap": {
|
||||||
|
"minimal": "none"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
"glm-5": {
|
"glm-5": {
|
||||||
"id": "glm-5",
|
"id": "glm-5",
|
||||||
"name": "GLM-5",
|
"name": "GLM-5",
|
||||||
@@ -13783,6 +13816,26 @@
|
|||||||
]
|
]
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
"gpt-oss-20b": {
|
||||||
|
"id": "gpt-oss-20b",
|
||||||
|
"name": "OpenAI gpt-oss-20b",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"provider": "fireworks",
|
||||||
|
"baseUrl": "https://api.fireworks.ai/inference/v1",
|
||||||
|
"reasoning": false,
|
||||||
|
"input": [
|
||||||
|
"text"
|
||||||
|
],
|
||||||
|
"cost": {
|
||||||
|
"input": 0,
|
||||||
|
"output": 0,
|
||||||
|
"cacheRead": 0,
|
||||||
|
"cacheWrite": 0
|
||||||
|
},
|
||||||
|
"contextWindow": 131072,
|
||||||
|
"maxTokens": 65536,
|
||||||
|
"supportsTools": false
|
||||||
|
},
|
||||||
"kimi-k2.5": {
|
"kimi-k2.5": {
|
||||||
"id": "kimi-k2.5",
|
"id": "kimi-k2.5",
|
||||||
"name": "Kimi K2.5",
|
"name": "Kimi K2.5",
|
||||||
@@ -14003,6 +14056,70 @@
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
"qwen3-embedding-8b": {
|
||||||
|
"id": "qwen3-embedding-8b",
|
||||||
|
"name": "Qwen3 Embedding 8B",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"provider": "fireworks",
|
||||||
|
"baseUrl": "https://api.fireworks.ai/inference/v1",
|
||||||
|
"reasoning": true,
|
||||||
|
"input": [
|
||||||
|
"text"
|
||||||
|
],
|
||||||
|
"cost": {
|
||||||
|
"input": 0,
|
||||||
|
"output": 0,
|
||||||
|
"cacheRead": 0,
|
||||||
|
"cacheWrite": 0
|
||||||
|
},
|
||||||
|
"contextWindow": 40960,
|
||||||
|
"maxTokens": null,
|
||||||
|
"supportsTools": false,
|
||||||
|
"thinking": {
|
||||||
|
"mode": "effort",
|
||||||
|
"efforts": [
|
||||||
|
"minimal",
|
||||||
|
"low",
|
||||||
|
"medium",
|
||||||
|
"high"
|
||||||
|
],
|
||||||
|
"effortMap": {
|
||||||
|
"minimal": "none"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"qwen3-reranker-8b": {
|
||||||
|
"id": "qwen3-reranker-8b",
|
||||||
|
"name": "Qwen3 Reranker 8B",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"provider": "fireworks",
|
||||||
|
"baseUrl": "https://api.fireworks.ai/inference/v1",
|
||||||
|
"reasoning": true,
|
||||||
|
"input": [
|
||||||
|
"text"
|
||||||
|
],
|
||||||
|
"cost": {
|
||||||
|
"input": 0,
|
||||||
|
"output": 0,
|
||||||
|
"cacheRead": 0,
|
||||||
|
"cacheWrite": 0
|
||||||
|
},
|
||||||
|
"contextWindow": 40960,
|
||||||
|
"maxTokens": null,
|
||||||
|
"supportsTools": false,
|
||||||
|
"thinking": {
|
||||||
|
"mode": "effort",
|
||||||
|
"efforts": [
|
||||||
|
"minimal",
|
||||||
|
"low",
|
||||||
|
"medium",
|
||||||
|
"high"
|
||||||
|
],
|
||||||
|
"effortMap": {
|
||||||
|
"minimal": "none"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
"qwen3.6-plus": {
|
"qwen3.6-plus": {
|
||||||
"id": "qwen3.6-plus",
|
"id": "qwen3.6-plus",
|
||||||
"name": "Qwen3.6 Plus",
|
"name": "Qwen3.6 Plus",
|
||||||
@@ -16509,6 +16626,64 @@
|
|||||||
"high"
|
"high"
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
|
},
|
||||||
|
"gemma-4-E2B-it": {
|
||||||
|
"id": "gemma-4-E2B-it",
|
||||||
|
"name": "Gemma 4 E2B IT",
|
||||||
|
"api": "google-generative-ai",
|
||||||
|
"provider": "google",
|
||||||
|
"baseUrl": "https://generativelanguage.googleapis.com/v1beta",
|
||||||
|
"reasoning": true,
|
||||||
|
"input": [
|
||||||
|
"text",
|
||||||
|
"image"
|
||||||
|
],
|
||||||
|
"cost": {
|
||||||
|
"input": 0,
|
||||||
|
"output": 0,
|
||||||
|
"cacheRead": 0,
|
||||||
|
"cacheWrite": 0
|
||||||
|
},
|
||||||
|
"contextWindow": 131072,
|
||||||
|
"maxTokens": 8192,
|
||||||
|
"thinking": {
|
||||||
|
"mode": "budget",
|
||||||
|
"efforts": [
|
||||||
|
"minimal",
|
||||||
|
"low",
|
||||||
|
"medium",
|
||||||
|
"high"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"gemma-4-E4B-it": {
|
||||||
|
"id": "gemma-4-E4B-it",
|
||||||
|
"name": "Gemma 4 E4B IT",
|
||||||
|
"api": "google-generative-ai",
|
||||||
|
"provider": "google",
|
||||||
|
"baseUrl": "https://generativelanguage.googleapis.com/v1beta",
|
||||||
|
"reasoning": true,
|
||||||
|
"input": [
|
||||||
|
"text",
|
||||||
|
"image"
|
||||||
|
],
|
||||||
|
"cost": {
|
||||||
|
"input": 0,
|
||||||
|
"output": 0,
|
||||||
|
"cacheRead": 0,
|
||||||
|
"cacheWrite": 0
|
||||||
|
},
|
||||||
|
"contextWindow": 131072,
|
||||||
|
"maxTokens": 8192,
|
||||||
|
"thinking": {
|
||||||
|
"mode": "budget",
|
||||||
|
"efforts": [
|
||||||
|
"minimal",
|
||||||
|
"low",
|
||||||
|
"medium",
|
||||||
|
"high"
|
||||||
|
]
|
||||||
|
}
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
"google-antigravity": {
|
"google-antigravity": {
|
||||||
@@ -75419,6 +75594,34 @@
|
|||||||
"requiresEffort": true
|
"requiresEffort": true
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
|
"qwen/qwen3-asr-flash": {
|
||||||
|
"id": "qwen/qwen3-asr-flash",
|
||||||
|
"name": "Qwen3-ASR-Flash",
|
||||||
|
"api": "openai-completions",
|
||||||
|
"provider": "zenmux",
|
||||||
|
"baseUrl": "https://zenmux.ai/api/v1",
|
||||||
|
"reasoning": true,
|
||||||
|
"input": [
|
||||||
|
"text"
|
||||||
|
],
|
||||||
|
"cost": {
|
||||||
|
"input": 0,
|
||||||
|
"output": 0,
|
||||||
|
"cacheRead": 0,
|
||||||
|
"cacheWrite": 0
|
||||||
|
},
|
||||||
|
"contextWindow": 1000000,
|
||||||
|
"maxTokens": null,
|
||||||
|
"thinking": {
|
||||||
|
"mode": "effort",
|
||||||
|
"efforts": [
|
||||||
|
"minimal",
|
||||||
|
"low",
|
||||||
|
"medium",
|
||||||
|
"high"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
},
|
||||||
"qwen/qwen3-coder": {
|
"qwen/qwen3-coder": {
|
||||||
"id": "qwen/qwen3-coder",
|
"id": "qwen/qwen3-coder",
|
||||||
"name": "Qwen3-Coder",
|
"name": "Qwen3-Coder",
|
||||||
|
|||||||
@@ -1160,6 +1160,7 @@ function mapFireworksControlPlaneModel(
|
|||||||
): ModelSpec<"openai-completions"> {
|
): ModelSpec<"openai-completions"> {
|
||||||
const name = toModelName(record.displayName, reference?.name ?? publicModelId);
|
const name = toModelName(record.displayName, reference?.name ?? publicModelId);
|
||||||
const supportsImage = toBoolean(record.supportsImageInput) === true;
|
const supportsImage = toBoolean(record.supportsImageInput) === true;
|
||||||
|
const supportsTools = toBoolean(record.supportsTools);
|
||||||
const contextWindow = toPositiveNumber(record.contextLength, reference?.contextWindow ?? null);
|
const contextWindow = toPositiveNumber(record.contextLength, reference?.contextWindow ?? null);
|
||||||
// The control plane reports no max-output budget; default the Kimi family to
|
// The control plane reports no max-output budget; default the Kimi family to
|
||||||
// its published cap, everyone else to the discovery fallback, then clamp.
|
// its published cap, everyone else to the discovery fallback, then clamp.
|
||||||
@@ -1192,6 +1193,7 @@ function mapFireworksControlPlaneModel(
|
|||||||
input: supportsImage ? ["text", "image"] : (reference?.input ?? ["text"]),
|
input: supportsImage ? ["text", "image"] : (reference?.input ?? ["text"]),
|
||||||
contextWindow,
|
contextWindow,
|
||||||
maxTokens,
|
maxTokens,
|
||||||
|
...(supportsTools === false ? { supportsTools: false } : {}),
|
||||||
};
|
};
|
||||||
return stripFireworksDeepSeekThinkingToggle(model, publicModelId);
|
return stripFireworksDeepSeekThinkingToggle(model, publicModelId);
|
||||||
}
|
}
|
||||||
@@ -1240,7 +1242,6 @@ async function fetchFireworksServerlessModels(options: {
|
|||||||
if (!isRecord(entry)) continue;
|
if (!isRecord(entry)) continue;
|
||||||
const record = entry as FireworksControlPlaneModel;
|
const record = entry as FireworksControlPlaneModel;
|
||||||
if (toBoolean(record.supportsServerless) !== true) continue;
|
if (toBoolean(record.supportsServerless) !== true) continue;
|
||||||
if (toBoolean(record.supportsTools) !== true) continue;
|
|
||||||
if (typeof record.state === "string" && record.state !== "READY") continue;
|
if (typeof record.state === "string" && record.state !== "READY") continue;
|
||||||
const wireName = typeof record.name === "string" ? record.name : "";
|
const wireName = typeof record.name === "string" ? record.name : "";
|
||||||
if (!wireName) continue;
|
if (!wireName) continue;
|
||||||
@@ -1396,6 +1397,7 @@ function mapWaferModel(
|
|||||||
const capabilities = wafer?.capabilities ?? {};
|
const capabilities = wafer?.capabilities ?? {};
|
||||||
const reasoning = capabilities.reasoning === true;
|
const reasoning = capabilities.reasoning === true;
|
||||||
const vision = capabilities.vision === true;
|
const vision = capabilities.vision === true;
|
||||||
|
const supportsTools = toBoolean(capabilities.tools) === false ? false : undefined;
|
||||||
const contextWindow = toPositiveNumber(
|
const contextWindow = toPositiveNumber(
|
||||||
wafer?.context_length,
|
wafer?.context_length,
|
||||||
toPositiveNumber((entry as { max_model_len?: unknown }).max_model_len, defaults.contextWindow),
|
toPositiveNumber((entry as { max_model_len?: unknown }).max_model_len, defaults.contextWindow),
|
||||||
@@ -1434,6 +1436,7 @@ function mapWaferModel(
|
|||||||
cost,
|
cost,
|
||||||
contextWindow,
|
contextWindow,
|
||||||
maxTokens,
|
maxTokens,
|
||||||
|
...(supportsTools === false ? { supportsTools } : {}),
|
||||||
};
|
};
|
||||||
if (reasoning) {
|
if (reasoning) {
|
||||||
// Wafer's `wafer.provider` envelope tells us which upstream backend serves
|
// Wafer's `wafer.provider` envelope tells us which upstream backend serves
|
||||||
@@ -2928,6 +2931,7 @@ export function mapModelsDevToModels(
|
|||||||
},
|
},
|
||||||
contextWindow: toPositiveNumber(m.limit?.context, desc.defaultContextWindow ?? null),
|
contextWindow: toPositiveNumber(m.limit?.context, desc.defaultContextWindow ?? null),
|
||||||
maxTokens: toPositiveNumber(m.limit?.output, desc.defaultMaxTokens ?? null),
|
maxTokens: toPositiveNumber(m.limit?.output, desc.defaultMaxTokens ?? null),
|
||||||
|
...(m.tool_call === false ? { supportsTools: false } : {}),
|
||||||
...(desc.compat && { compat: desc.compat }),
|
...(desc.compat && { compat: desc.compat }),
|
||||||
...(desc.headers && { headers: { ...desc.headers } }),
|
...(desc.headers && { headers: { ...desc.headers } }),
|
||||||
};
|
};
|
||||||
|
|||||||
@@ -436,6 +436,13 @@ export interface Model<TApi extends Api = Api> {
|
|||||||
baseUrl: string;
|
baseUrl: string;
|
||||||
reasoning: boolean;
|
reasoning: boolean;
|
||||||
input: ("text" | "image")[];
|
input: ("text" | "image")[];
|
||||||
|
/**
|
||||||
|
* Native provider tool-call support. `false` is the only unsupported signal:
|
||||||
|
* `true` and `undefined` both mean callers may use native tools. Catalog and
|
||||||
|
* discovery sources should set this sparsely when an upstream explicitly
|
||||||
|
* reports that native tool calling is unsupported.
|
||||||
|
*/
|
||||||
|
supportsTools?: boolean;
|
||||||
cost: {
|
cost: {
|
||||||
input: number; // $/million tokens
|
input: number; // $/million tokens
|
||||||
output: number; // $/million tokens
|
output: number; // $/million tokens
|
||||||
|
|||||||
@@ -30,7 +30,7 @@ const PAGE_1 = [
|
|||||||
supportsServerless: true,
|
supportsServerless: true,
|
||||||
state: "READY",
|
state: "READY",
|
||||||
},
|
},
|
||||||
// Image model: serverless but not tool-capable — must be filtered out.
|
// Serverless but not tool-capable — still listed, but marked for owned-tool fallback.
|
||||||
{
|
{
|
||||||
name: "accounts/fireworks/models/flux-1-schnell-fp8",
|
name: "accounts/fireworks/models/flux-1-schnell-fp8",
|
||||||
displayName: "FLUX.1 [schnell]",
|
displayName: "FLUX.1 [schnell]",
|
||||||
@@ -139,10 +139,11 @@ describe("Fireworks control-plane serverless discovery", () => {
|
|||||||
expect(kimi.reasoning).toBe(true);
|
expect(kimi.reasoning).toBe(true);
|
||||||
});
|
});
|
||||||
|
|
||||||
it("filters non-serverless, non-tool, and not-ready records", async () => {
|
it("keeps non-tool serverless records flagged and filters unavailable records", async () => {
|
||||||
const { models } = await discover();
|
const { models } = await discover();
|
||||||
const ids = models.map(m => m.id);
|
const ids = models.map(m => m.id);
|
||||||
expect(ids).not.toContain("flux-1-schnell-fp8"); // tools: false
|
const flux = models.find(m => m.id === "flux-1-schnell-fp8");
|
||||||
|
expect(flux?.supportsTools).toBe(false);
|
||||||
expect(ids).not.toContain("kimi-k2-instruct-0905"); // serverless: false
|
expect(ids).not.toContain("kimi-k2-instruct-0905"); // serverless: false
|
||||||
expect(ids).not.toContain("some-pending-model"); // state: DEPLOYING
|
expect(ids).not.toContain("some-pending-model"); // state: DEPLOYING
|
||||||
});
|
});
|
||||||
|
|||||||
@@ -1,11 +1,17 @@
|
|||||||
# Changelog
|
# Changelog
|
||||||
|
|
||||||
## [Unreleased]
|
## [Unreleased]
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
|
|
||||||
|
- Added `supportsTools` to model definitions and overrides so custom model configs can declare whether a model supports native tool calls
|
||||||
|
- Added `tools.format` for choosing native tool calling or a specific owned in-band format (`glm`, `hermes`, `kimi`, `xml`), with `auto` falling back to GLM only for models marked as not supporting native tools.
|
||||||
- Added a conditional easter-egg tip recommending nerd fonts when using the unicode symbol preset.
|
- Added a conditional easter-egg tip recommending nerd fonts when using the unicode symbol preset.
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
|
||||||
|
- Expanded `tools.format` to support additional in-band tool-call syntaxes, including `anthropic`, `deepseek`, `harmony`, `pi`, and `qwen3`
|
||||||
|
- Changed the experimental owned tool-calling prompt from a GLM-only toggle to syntax-specific grammar prompts and result formats. `PI_OWNED_TOOLS=1` still forces GLM; `PI_OWNED_TOOLS=<syntax>` forces that syntax.
|
||||||
|
|
||||||
### Fixed
|
### Fixed
|
||||||
|
|
||||||
- Fixed Auto-Promote Context being pre-empted by compaction: the pre-prompt context check ran compaction directly, so snapcompact (or any strategy) fired before promotion ever got a chance. It now tries promotion to a larger-context model first — mirroring the post-turn threshold path — and only compacts when no larger-context target is available. Snapcompact (auto and manual) also falls back to a context-full LLM summary when its frame archive plus kept history would still overflow the model's usable window, instead of leaving the session over the limit.
|
- Fixed Auto-Promote Context being pre-empted by compaction: the pre-prompt context check ran compaction directly, so snapcompact (or any strategy) fired before promotion ever got a chance. It now tries promotion to a larger-context model first — mirroring the post-turn threshold path — and only compacts when no larger-context target is available. Snapcompact (auto and manual) also falls back to a context-full LLM summary when its frame archive plus kept history would still overflow the model's usable window, instead of leaving the session over the limit.
|
||||||
@@ -11703,4 +11709,4 @@ Initial public release.
|
|||||||
|
|
||||||
## [0.7.6] - 2025-11-13
|
## [0.7.6] - 2025-11-13
|
||||||
|
|
||||||
Previous releases did not maintain a changelog.
|
Previous releases did not maintain a changelog.
|
||||||
@@ -130,11 +130,13 @@ export function mergeDiscoveredModel<TApi extends Api>(
|
|||||||
providerOverride?: Pick<ProviderOverride, "baseUrl" | "headers" | "transport">,
|
providerOverride?: Pick<ProviderOverride, "baseUrl" | "headers" | "transport">,
|
||||||
): Model<TApi> {
|
): Model<TApi> {
|
||||||
if (existing) {
|
if (existing) {
|
||||||
|
const supportsTools = model.supportsTools ?? existing.supportsTools;
|
||||||
return buildModel({
|
return buildModel({
|
||||||
...model,
|
...model,
|
||||||
baseUrl: providerOverride?.baseUrl ?? model.baseUrl ?? existing.baseUrl,
|
baseUrl: providerOverride?.baseUrl ?? model.baseUrl ?? existing.baseUrl,
|
||||||
headers: existing.headers ? { ...existing.headers, ...model.headers } : model.headers,
|
headers: existing.headers ? { ...existing.headers, ...model.headers } : model.headers,
|
||||||
transport: providerOverride?.transport ?? existing.transport ?? model.transport,
|
transport: providerOverride?.transport ?? existing.transport ?? model.transport,
|
||||||
|
...(supportsTools !== undefined ? { supportsTools } : {}),
|
||||||
compat: model.compatConfig,
|
compat: model.compatConfig,
|
||||||
} as ModelSpec<TApi>);
|
} as ModelSpec<TApi>);
|
||||||
}
|
}
|
||||||
@@ -370,6 +372,7 @@ interface ModelPatch {
|
|||||||
reasoning?: boolean;
|
reasoning?: boolean;
|
||||||
thinking?: ThinkingConfig;
|
thinking?: ThinkingConfig;
|
||||||
input?: ("text" | "image")[];
|
input?: ("text" | "image")[];
|
||||||
|
supportsTools?: boolean;
|
||||||
cost?: Partial<Model<Api>["cost"]>;
|
cost?: Partial<Model<Api>["cost"]>;
|
||||||
contextWindow?: number;
|
contextWindow?: number;
|
||||||
maxTokens?: number;
|
maxTokens?: number;
|
||||||
@@ -395,6 +398,7 @@ function applyModelPatch(base: Model<Api>, patch: ModelPatch, transport: ModelTr
|
|||||||
if (patch.reasoning !== undefined) result.reasoning = patch.reasoning;
|
if (patch.reasoning !== undefined) result.reasoning = patch.reasoning;
|
||||||
if (patch.thinking !== undefined) result.thinking = patch.thinking;
|
if (patch.thinking !== undefined) result.thinking = patch.thinking;
|
||||||
if (patch.input !== undefined) result.input = patch.input;
|
if (patch.input !== undefined) result.input = patch.input;
|
||||||
|
if (patch.supportsTools !== undefined) result.supportsTools = patch.supportsTools;
|
||||||
if (patch.contextWindow !== undefined) result.contextWindow = patch.contextWindow;
|
if (patch.contextWindow !== undefined) result.contextWindow = patch.contextWindow;
|
||||||
if (patch.maxTokens !== undefined) result.maxTokens = patch.maxTokens;
|
if (patch.maxTokens !== undefined) result.maxTokens = patch.maxTokens;
|
||||||
if (patch.omitMaxOutputTokens !== undefined) result.omitMaxOutputTokens = patch.omitMaxOutputTokens;
|
if (patch.omitMaxOutputTokens !== undefined) result.omitMaxOutputTokens = patch.omitMaxOutputTokens;
|
||||||
@@ -506,6 +510,7 @@ function buildCustomModelOverlay(
|
|||||||
reasoning: modelDef.reasoning,
|
reasoning: modelDef.reasoning,
|
||||||
thinking: modelDef.thinking,
|
thinking: modelDef.thinking,
|
||||||
input: modelDef.input,
|
input: modelDef.input,
|
||||||
|
supportsTools: modelDef.supportsTools,
|
||||||
cost: modelDef.cost,
|
cost: modelDef.cost,
|
||||||
contextWindow: modelDef.contextWindow,
|
contextWindow: modelDef.contextWindow,
|
||||||
maxTokens: modelDef.maxTokens,
|
maxTokens: modelDef.maxTokens,
|
||||||
@@ -535,6 +540,7 @@ function finalizeCustomModel(model: CustomModelOverlay, options: CustomModelBuil
|
|||||||
reference?.cost ??
|
reference?.cost ??
|
||||||
(options.useDefaults ? { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 } : undefined);
|
(options.useDefaults ? { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 } : undefined);
|
||||||
const input = resolvedModel.input ?? reference?.input ?? (options.useDefaults ? ["text"] : undefined);
|
const input = resolvedModel.input ?? reference?.input ?? (options.useDefaults ? ["text"] : undefined);
|
||||||
|
const supportsTools = resolvedModel.supportsTools ?? reference?.supportsTools;
|
||||||
return buildModel({
|
return buildModel({
|
||||||
id: resolvedModel.id,
|
id: resolvedModel.id,
|
||||||
name: resolvedModel.name ?? (options.useDefaults ? resolvedModel.id : undefined),
|
name: resolvedModel.name ?? (options.useDefaults ? resolvedModel.id : undefined),
|
||||||
@@ -544,6 +550,7 @@ function finalizeCustomModel(model: CustomModelOverlay, options: CustomModelBuil
|
|||||||
reasoning: resolvedModel.reasoning ?? reference?.reasoning ?? (options.useDefaults ? false : undefined),
|
reasoning: resolvedModel.reasoning ?? reference?.reasoning ?? (options.useDefaults ? false : undefined),
|
||||||
thinking: resolvedModel.thinking ?? reference?.thinking,
|
thinking: resolvedModel.thinking ?? reference?.thinking,
|
||||||
input: input as ("text" | "image")[],
|
input: input as ("text" | "image")[],
|
||||||
|
...(supportsTools !== undefined ? { supportsTools } : {}),
|
||||||
cost,
|
cost,
|
||||||
contextWindow: resolvedModel.contextWindow ?? reference?.contextWindow ?? (options.useDefaults ? 128000 : null),
|
contextWindow: resolvedModel.contextWindow ?? reference?.contextWindow ?? (options.useDefaults ? 128000 : null),
|
||||||
maxTokens: resolvedModel.maxTokens ?? reference?.maxTokens ?? (options.useDefaults ? 16384 : null),
|
maxTokens: resolvedModel.maxTokens ?? reference?.maxTokens ?? (options.useDefaults ? 16384 : null),
|
||||||
@@ -878,10 +885,12 @@ export class ModelRegistry {
|
|||||||
#mergeResolvedModels(baseModels: Model<Api>[], replacementModels: Model<Api>[]): Model<Api>[] {
|
#mergeResolvedModels(baseModels: Model<Api>[], replacementModels: Model<Api>[]): Model<Api>[] {
|
||||||
return mergeByModelKey(baseModels, replacementModels, (existing, replacementModel) => {
|
return mergeByModelKey(baseModels, replacementModels, (existing, replacementModel) => {
|
||||||
if (!existing) return replacementModel;
|
if (!existing) return replacementModel;
|
||||||
|
const supportsTools = replacementModel.supportsTools ?? existing.supportsTools;
|
||||||
return {
|
return {
|
||||||
...replacementModel,
|
...replacementModel,
|
||||||
contextWindow: replacementModel.contextWindow ?? existing.contextWindow,
|
contextWindow: replacementModel.contextWindow ?? existing.contextWindow,
|
||||||
maxTokens: replacementModel.maxTokens ?? existing.maxTokens,
|
maxTokens: replacementModel.maxTokens ?? existing.maxTokens,
|
||||||
|
...(supportsTools !== undefined ? { supportsTools } : {}),
|
||||||
};
|
};
|
||||||
});
|
});
|
||||||
}
|
}
|
||||||
@@ -2205,6 +2214,7 @@ export interface ProviderConfigInput {
|
|||||||
reasoning: boolean;
|
reasoning: boolean;
|
||||||
thinking?: ThinkingConfig;
|
thinking?: ThinkingConfig;
|
||||||
input: ("text" | "image")[];
|
input: ("text" | "image")[];
|
||||||
|
supportsTools?: boolean;
|
||||||
cost: { input: number; output: number; cacheRead: number; cacheWrite: number };
|
cost: { input: number; output: number; cacheRead: number; cacheWrite: number };
|
||||||
contextWindow: number;
|
contextWindow: number;
|
||||||
maxTokens: number;
|
maxTokens: number;
|
||||||
|
|||||||
@@ -133,6 +133,7 @@ const ModelDefinitionSchema = z.object({
|
|||||||
reasoning: z.boolean().optional(),
|
reasoning: z.boolean().optional(),
|
||||||
thinking: ModelThinkingSchema.optional(),
|
thinking: ModelThinkingSchema.optional(),
|
||||||
input: z.array(z.enum(["text", "image"])).optional(),
|
input: z.array(z.enum(["text", "image"])).optional(),
|
||||||
|
supportsTools: z.boolean().optional(),
|
||||||
cost: z
|
cost: z
|
||||||
.object({
|
.object({
|
||||||
input: z.number(),
|
input: z.number(),
|
||||||
@@ -155,6 +156,7 @@ export const ModelOverrideSchema = z.object({
|
|||||||
reasoning: z.boolean().optional(),
|
reasoning: z.boolean().optional(),
|
||||||
thinking: ModelThinkingSchema.optional(),
|
thinking: ModelThinkingSchema.optional(),
|
||||||
input: z.array(z.enum(["text", "image"])).optional(),
|
input: z.array(z.enum(["text", "image"])).optional(),
|
||||||
|
supportsTools: z.boolean().optional(),
|
||||||
cost: z
|
cost: z
|
||||||
.object({
|
.object({
|
||||||
input: z.number().optional(),
|
input: z.number().optional(),
|
||||||
|
|||||||
@@ -17,6 +17,7 @@ export interface ProviderValidationModel {
|
|||||||
id: string;
|
id: string;
|
||||||
api?: Api;
|
api?: Api;
|
||||||
contextWindow?: number;
|
contextWindow?: number;
|
||||||
|
supportsTools?: boolean;
|
||||||
maxTokens?: number;
|
maxTokens?: number;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -1719,6 +1719,48 @@ export const SETTINGS_SCHEMA = {
|
|||||||
},
|
},
|
||||||
},
|
},
|
||||||
|
|
||||||
|
"tools.format": {
|
||||||
|
type: "enum",
|
||||||
|
values: [
|
||||||
|
"auto",
|
||||||
|
"native",
|
||||||
|
"glm",
|
||||||
|
"hermes",
|
||||||
|
"kimi",
|
||||||
|
"xml",
|
||||||
|
"anthropic",
|
||||||
|
"deepseek",
|
||||||
|
"harmony",
|
||||||
|
"pi",
|
||||||
|
"qwen3",
|
||||||
|
] as const,
|
||||||
|
default: "auto",
|
||||||
|
ui: {
|
||||||
|
tab: "context",
|
||||||
|
group: "Experimental",
|
||||||
|
label: "Tool Call Format",
|
||||||
|
description:
|
||||||
|
"Controls how tools are exposed to the model. Auto uses native tool calls unless the selected model is marked as not supporting tools, then falls back to GLM-style in-band tool calls. Native forces provider-native tools; the other values force the named in-band syntax. Applies on session start.",
|
||||||
|
options: [
|
||||||
|
{
|
||||||
|
value: "auto",
|
||||||
|
label: "Auto",
|
||||||
|
description: "Use native tool calls unless the model is known not to support them.",
|
||||||
|
},
|
||||||
|
{ value: "native", label: "Native", description: "Use provider-native tool calls." },
|
||||||
|
{ value: "glm", label: "GLM", description: "Use GLM-style in-band tool calls." },
|
||||||
|
{ value: "hermes", label: "Hermes", description: "Use Hermes-style in-band tool calls." },
|
||||||
|
{ value: "kimi", label: "Kimi", description: "Use Kimi-style in-band tool calls." },
|
||||||
|
{ value: "xml", label: "XML", description: "Use generic XML in-band tool calls." },
|
||||||
|
{ value: "anthropic", label: "Anthropic", description: "Use Anthropic-style in-band tool calls." },
|
||||||
|
{ value: "deepseek", label: "DeepSeek", description: "Use DeepSeek-style in-band tool calls." },
|
||||||
|
{ value: "harmony", label: "Harmony", description: "Use Harmony-style in-band tool calls." },
|
||||||
|
{ value: "pi", label: "Pi", description: "Use Pi-style in-band tool calls." },
|
||||||
|
{ value: "qwen3", label: "Qwen3", description: "Use Qwen3-style in-band tool calls." },
|
||||||
|
],
|
||||||
|
},
|
||||||
|
},
|
||||||
|
|
||||||
"snapcompact.shape": {
|
"snapcompact.shape": {
|
||||||
type: "enum",
|
type: "enum",
|
||||||
values: ["auto", ...SHAPE_VARIANT_NAMES] as const,
|
values: ["auto", ...SHAPE_VARIANT_NAMES] as const,
|
||||||
|
|||||||
@@ -16,6 +16,7 @@ import {
|
|||||||
type SimpleStreamOptions,
|
type SimpleStreamOptions,
|
||||||
streamSimple,
|
streamSimple,
|
||||||
} from "@oh-my-pi/pi-ai";
|
} from "@oh-my-pi/pi-ai";
|
||||||
|
import type { ToolCallSyntax } from "@oh-my-pi/pi-ai/grammar";
|
||||||
import {
|
import {
|
||||||
getOpenAICodexTransportDetails,
|
getOpenAICodexTransportDetails,
|
||||||
prewarmOpenAICodexResponses,
|
prewarmOpenAICodexResponses,
|
||||||
@@ -550,6 +551,17 @@ export interface CreateAgentSessionResult {
|
|||||||
eventBus: EventBus;
|
eventBus: EventBus;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
export type ToolCallFormat = "auto" | "native" | ToolCallSyntax;
|
||||||
|
|
||||||
|
export function resolveToolCallSyntax(
|
||||||
|
format: ToolCallFormat,
|
||||||
|
model: Pick<Model, "supportsTools"> | undefined,
|
||||||
|
): ToolCallSyntax | undefined {
|
||||||
|
if (format === "native") return undefined;
|
||||||
|
if (format === "auto") return model?.supportsTools === false ? "glm" : undefined;
|
||||||
|
return format;
|
||||||
|
}
|
||||||
|
|
||||||
// Re-exports
|
// Re-exports
|
||||||
|
|
||||||
export type { PromptTemplate } from "./config/prompt-templates";
|
export type { PromptTemplate } from "./config/prompt-templates";
|
||||||
@@ -2489,6 +2501,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
|
|||||||
return result;
|
return result;
|
||||||
},
|
},
|
||||||
intentTracing: !!intentField,
|
intentTracing: !!intentField,
|
||||||
|
toolCallSyntax: resolveToolCallSyntax(settings.get("tools.format"), model),
|
||||||
getToolChoice: () => session?.nextToolChoice(),
|
getToolChoice: () => session?.nextToolChoice(),
|
||||||
telemetry: options.telemetry,
|
telemetry: options.telemetry,
|
||||||
appendOnlyContext: model
|
appendOnlyContext: model
|
||||||
|
|||||||
@@ -0,0 +1,16 @@
|
|||||||
|
import { describe, expect, it } from "bun:test";
|
||||||
|
import { resolveToolCallSyntax } from "@oh-my-pi/pi-coding-agent/sdk";
|
||||||
|
|
||||||
|
describe("resolveToolCallSyntax", () => {
|
||||||
|
it("uses GLM in auto mode only for models known not to support native tools", () => {
|
||||||
|
expect(resolveToolCallSyntax("auto", { supportsTools: false })).toBe("glm");
|
||||||
|
expect(resolveToolCallSyntax("auto", { supportsTools: true })).toBeUndefined();
|
||||||
|
expect(resolveToolCallSyntax("auto", {})).toBeUndefined();
|
||||||
|
expect(resolveToolCallSyntax("auto", undefined)).toBeUndefined();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("keeps native unset and passes explicit in-band syntaxes through", () => {
|
||||||
|
expect(resolveToolCallSyntax("native", { supportsTools: false })).toBeUndefined();
|
||||||
|
expect(resolveToolCallSyntax("qwen3", undefined)).toBe("qwen3");
|
||||||
|
});
|
||||||
|
});
|
||||||
@@ -5,6 +5,7 @@ import * as path from "node:path";
|
|||||||
import { Effort } from "@oh-my-pi/pi-ai";
|
import { Effort } from "@oh-my-pi/pi-ai";
|
||||||
import {
|
import {
|
||||||
getDefault,
|
getDefault,
|
||||||
|
getEnumValues,
|
||||||
onAppendOnlyModeChanged,
|
onAppendOnlyModeChanged,
|
||||||
onStatusLineSessionAccentChanged,
|
onStatusLineSessionAccentChanged,
|
||||||
resetSettingsForTest,
|
resetSettingsForTest,
|
||||||
@@ -64,6 +65,23 @@ describe("Settings", () => {
|
|||||||
const settings = await Settings.init({ cwd: projectDir, agentDir });
|
const settings = await Settings.init({ cwd: projectDir, agentDir });
|
||||||
expect(settings.get("tui.maxInlineImages")).toBe(8);
|
expect(settings.get("tui.maxInlineImages")).toBe(8);
|
||||||
});
|
});
|
||||||
|
|
||||||
|
it("exposes all tool call format options", () => {
|
||||||
|
const values = getEnumValues("tools.format");
|
||||||
|
expect(values).toEqual([
|
||||||
|
"auto",
|
||||||
|
"native",
|
||||||
|
"glm",
|
||||||
|
"hermes",
|
||||||
|
"kimi",
|
||||||
|
"xml",
|
||||||
|
"anthropic",
|
||||||
|
"deepseek",
|
||||||
|
"harmony",
|
||||||
|
"pi-native",
|
||||||
|
"qwen3",
|
||||||
|
]);
|
||||||
|
});
|
||||||
});
|
});
|
||||||
|
|
||||||
describe("get()", () => {
|
describe("get()", () => {
|
||||||
|
|||||||
@@ -639,6 +639,11 @@ class LiveProgress {
|
|||||||
const metaParts = [op, target].filter((v): v is string => Boolean(v));
|
const metaParts = [op, target].filter((v): v is string => Boolean(v));
|
||||||
const meta = metaParts.length > 0 ? paint(ANSI.dim, metaParts.join(" ")) : "";
|
const meta = metaParts.length > 0 ? paint(ANSI.dim, metaParts.join(" ")) : "";
|
||||||
console.log(` ${tag}${meta ? ` ${meta}` : ""} ${clipped}`);
|
console.log(` ${tag}${meta ? ` ${meta}` : ""} ${clipped}`);
|
||||||
|
if (failure.rawBlock) {
|
||||||
|
const rawLine = failure.rawBlock.replace(/\s+/g, " ").trim();
|
||||||
|
const clippedRaw = rawLine.length > 240 ? `${rawLine.slice(0, 237)}...` : rawLine;
|
||||||
|
console.log(` ${paint(ANSI.dim, "raw")} ${clippedRaw}`);
|
||||||
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -780,10 +780,16 @@ export interface ToolCallStats {
|
|||||||
totalInputChars: number;
|
totalInputChars: number;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
interface PendingEditCall {
|
||||||
|
args: unknown;
|
||||||
|
rawBlock?: string;
|
||||||
|
}
|
||||||
|
|
||||||
export interface EditFailure {
|
export interface EditFailure {
|
||||||
toolCallId: string;
|
toolCallId: string;
|
||||||
args: unknown;
|
args: unknown;
|
||||||
error: string;
|
error: string;
|
||||||
|
rawBlock?: string;
|
||||||
category?: EditFailureCategory;
|
category?: EditFailureCategory;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1196,9 +1202,16 @@ async function runSingleTask(
|
|||||||
});
|
});
|
||||||
break;
|
break;
|
||||||
}
|
}
|
||||||
const pendingEdits = new Map<string, unknown>();
|
const pendingEdits = new Map<string, PendingEditCall>();
|
||||||
|
const rawToolBlocks = new Map<string, string>();
|
||||||
for (const event of events) {
|
for (const event of events) {
|
||||||
|
if (event.type === "message_end") {
|
||||||
|
for (const raw of extractAssistantToolRawBlocks(event)) {
|
||||||
|
rawToolBlocks.set(raw.id, raw.rawBlock);
|
||||||
|
const pending = pendingEdits.get(raw.id);
|
||||||
|
if (pending) pending.rawBlock = raw.rawBlock;
|
||||||
|
}
|
||||||
|
}
|
||||||
if (event.type === "tool_execution_start") {
|
if (event.type === "tool_execution_start") {
|
||||||
const e = event as { toolName?: string; toolCallId?: string; args?: unknown };
|
const e = event as { toolName?: string; toolCallId?: string; args?: unknown };
|
||||||
const toolName = e.toolName;
|
const toolName = e.toolName;
|
||||||
@@ -1206,7 +1219,8 @@ async function runSingleTask(
|
|||||||
toolStats.read++;
|
toolStats.read++;
|
||||||
} else if (isEditTool(toolName)) {
|
} else if (isEditTool(toolName)) {
|
||||||
toolStats.edit++;
|
toolStats.edit++;
|
||||||
if (e.toolCallId) pendingEdits.set(e.toolCallId, e.args);
|
if (e.toolCallId)
|
||||||
|
pendingEdits.set(e.toolCallId, { args: e.args, rawBlock: rawToolBlocks.get(e.toolCallId) });
|
||||||
} else if (toolName === "write") {
|
} else if (toolName === "write") {
|
||||||
toolStats.write++;
|
toolStats.write++;
|
||||||
}
|
}
|
||||||
@@ -1218,7 +1232,8 @@ async function runSingleTask(
|
|||||||
} else if (event.type === "tool_execution_end") {
|
} else if (event.type === "tool_execution_end") {
|
||||||
const e = event as { toolName?: string; toolCallId?: string; isError?: boolean; result?: unknown };
|
const e = event as { toolName?: string; toolCallId?: string; isError?: boolean; result?: unknown };
|
||||||
if (isEditTool(e.toolName) && e.toolCallId && pendingEdits.has(e.toolCallId)) {
|
if (isEditTool(e.toolName) && e.toolCallId && pendingEdits.has(e.toolCallId)) {
|
||||||
const args = pendingEdits.get(e.toolCallId) ?? null;
|
const pendingEdit = pendingEdits.get(e.toolCallId) ?? { args: null };
|
||||||
|
const args = pendingEdit.args;
|
||||||
pendingEdits.delete(e.toolCallId);
|
pendingEdits.delete(e.toolCallId);
|
||||||
if (config.editVariant === "hashline" && args) {
|
if (config.editVariant === "hashline" && args) {
|
||||||
const counts = countHashlineEditSubtypes(args);
|
const counts = countHashlineEditSubtypes(args);
|
||||||
@@ -1238,6 +1253,7 @@ async function runSingleTask(
|
|||||||
toolCallId: e.toolCallId,
|
toolCallId: e.toolCallId,
|
||||||
args,
|
args,
|
||||||
error,
|
error,
|
||||||
|
rawBlock: pendingEdit.rawBlock,
|
||||||
category: categorizeEditFailure(error, args),
|
category: categorizeEditFailure(error, args),
|
||||||
});
|
});
|
||||||
} else {
|
} else {
|
||||||
@@ -1410,6 +1426,27 @@ function extractToolErrorMessage(result: unknown): string {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function extractAssistantToolRawBlocks(event: {
|
||||||
|
type: string;
|
||||||
|
[key: string]: unknown;
|
||||||
|
}): Array<{ id: string; rawBlock: string }> {
|
||||||
|
const message = event.message;
|
||||||
|
if (message === null || typeof message !== "object") return [];
|
||||||
|
const role = (message as { role?: unknown }).role;
|
||||||
|
if (role !== "assistant") return [];
|
||||||
|
const content = (message as { content?: unknown }).content;
|
||||||
|
if (!Array.isArray(content)) return [];
|
||||||
|
const rawBlocks: Array<{ id: string; rawBlock: string }> = [];
|
||||||
|
for (const block of content) {
|
||||||
|
if (block === null || typeof block !== "object") continue;
|
||||||
|
const typedBlock = block as { type?: unknown; id?: unknown; rawBlock?: unknown };
|
||||||
|
if (typedBlock.type !== "toolCall") continue;
|
||||||
|
if (typeof typedBlock.id !== "string" || typeof typedBlock.rawBlock !== "string") continue;
|
||||||
|
rawBlocks.push({ id: typedBlock.id, rawBlock: typedBlock.rawBlock });
|
||||||
|
}
|
||||||
|
return rawBlocks;
|
||||||
|
}
|
||||||
|
|
||||||
function shuffle<T>(items: T[]): T[] {
|
function shuffle<T>(items: T[]): T[] {
|
||||||
const copy = items.slice();
|
const copy = items.slice();
|
||||||
for (let i = copy.length - 1; i > 0; i--) {
|
for (let i = copy.length - 1; i > 0; i--) {
|
||||||
|
|||||||
Reference in New Issue
Block a user