chore: update stale docs

This commit is contained in:
can1357
2026-08-03 16:37:05 +02:00
parent fc04aa6fa7
commit ebd5e3f86f
120 changed files with 5246 additions and 4691 deletions
+23
View File
@@ -190,6 +190,12 @@ Before the API converts it, the model literally emits an XML block. The current
Current Claude models prefix these tags with an `antml:` XML namespace prefix (e.g. `antml:function_calls`, `antml:invoke name="…"`, `antml:parameter name="…"`). The API strips all of this and exposes only the JSON `tool_use` block; integrators should target the JSON, not the XML.
### OMP `anthropic` dialect
OMP operates on the underlying prompt-driven XML rather than Messages API content blocks. Its renderer always emits the unprefixed attribute form above, wraps multiple calls in one `<function_calls>` block, and renders each argument as a `<parameter name="…">` child. With a tool schema, declared string arguments are inserted as literal text; other values are JSON-serialized. The streaming scanner also accepts `antml:`-prefixed tags, `<tool_calls>` as a wrapper alias, and a bare `<invoke>` outside either wrapper.
The scanner mints call ids because this XML has none. It scans streamed text statefully, emits `toolArgDelta` events while each parameter body arrives, and publishes the coerced argument object with `toolEnd` after `</invoke>`. A parameter value is capped at 1,000,000 JavaScript string code units; overflow gains an explicit truncation suffix. JSON-like values are parsed with repair, while schema-declared strings stay strings. With `parseThinking: true`, `<thinking>`, `<think>`, and `<scratchpad>` (prefixed or unprefixed) become thinking events; otherwise those tags remain visible text.
---
## Multiple / parallel tool calls
@@ -301,6 +307,23 @@ Rich result (text + image blocks):
Server tools require **no** `tool_result` from you — Anthropic executes them and injects the result inline in the assistant turn. (Legacy XML feeds results back as `<function_results><result><tool_name>…</tool_name><stdout>…</stdout></result></function_results>`, or `<error>…</error>` on failure.)
OMP's prompt-driven dialect is different from Anthropic's server-tool behavior. It renders client results as:
```text
<function_results>
<result>
<tool_name>get_weather</tool_name>
<stdout>15 degrees</stdout>
</result>
<error>
<tool_name>other_tool</tool_name>
<stderr>execution failed</stderr>
</error>
</function_results>
```
There is no result id in this XML, so results are correlated by call order. OMP includes the tool name in both success and error entries and does not encode `isError` anywhere else.
---
## End-to-end example
+31
View File
@@ -368,6 +368,37 @@ reuse the same fullwidth pipe (`|`, U+FF5C), but the body is an Anthropic-styl
structured `tool_calls`; a parser must heal it back into tool calls and strip the markers from
user-visible text.
## omp / pi converter behavior
The repository's `deepseek` dialect is an **owned in-band converter**, not a
vLLM parser wrapper. Select it with `PI_DIALECT=deepseek` (or the equivalent
agent configuration). When tools are present, the agent appends the dialect
guide and compact tool catalog to the system prompt, removes native provider
tools from the request, re-encodes prior calls/results with this syntax, and
scans streamed assistant text back into canonical pi tool-call events.
The current scanner accepts all three forms described above:
- V3.1 `name<|tool▁sep|>{json}` calls;
- legacy `function<|tool▁sep|>name` plus a fenced JSON body; and
- fullwidth or ASCII DSML `invoke` / `parameter` blocks.
For V3.1 and legacy calls, omp emits `toolStart` after the header is complete
but buffers arguments until `</|tool▁call▁end|>`; it then uses the shared
repairing JSON parser. A missing/invalid argument object becomes `{}`, and an
unfinished call is discarded on flush. DSML is genuinely incremental:
`toolArgDelta` events are emitted while each parameter body arrives. A DSML
parameter is a raw string unless `string="false"`; the latter is
repairing-JSON-decoded and falls back to the raw text if decoding fails. Call
IDs for id-less DeepSeek forms are synthesized as `ptc_…`.
The scanner also removes leaked DeepSeek chat-template control tokens from
visible text and, by default, maps `<think>…</think>` to thinking events. Its
renderer emits V3.1 calls, joins parallel calls without separators, and renders
multiple results as singular output blocks separated by newlines. The DSML
syntax is accepted for healing leaked provider output but is not the owned
dialect's emitted history format.
## Sources
- DeepSeek-V3.1 model card (Chat Template / ToolCall sections): <https://huggingface.co/deepseek-ai/DeepSeek-V3.1>
+10 -8
View File
@@ -39,11 +39,11 @@ Tools are advertised in the prompt as a JSON-Schema catalog. Gemma 3's official
2. **JSON** (the sibling convention — see `qwen3.md` for the closely related Hermes shape):
> … you MUST put it in the format of `{"name": function name, "parameters": dictionary of argument name and its value}`
Hosted Gemini wraps the same idea in markdown fences and the `default_api` namespace. The function signatures themselves are passed as OpenAI-style tool JSON (`{"type":"function","function":{name,description,parameters}}`).
Hosted Gemini wraps the same idea in markdown fences and the `default_api` namespace. The function signatures themselves are passed as OpenAI-style tool JSON (`{"type":"function","function":{name,description,parameters}}`). OMP's renderer emits `default_api.NAME(...)` without `print`; its scanner also accepts the wrapped and bare variants below.
## Tool-call format
One call is a Python call expression. The hosted-Gemini canonical form is a `print()` of a `default_api` method, fenced:
One call is a Python call expression. Hosted Gemini commonly emits a `print()` of a `default_api` method:
````text
```tool_code
@@ -73,15 +73,15 @@ Strings use Python escaping (`\n`, `\t`, `\\`, `\'`, `\"`); hosted Gemini emits
## Multiple / parallel tool calls
Two encodings exist, both inside a single `tool_code` block:
Two encodings occur inside a single `tool_code` block:
- **Gemma 3 Pythonic prompt** — a Python **list** of call expressions:
- **OMP / Gemma 3 Pythonic form** — a Python **list** of call expressions. OMP renders this form for two or more calls:
````text
```tool_code
[get_current_temperature(location="London"), get_temperature_date(location="London", date="2024-10-01")]
[default_api.get_current_temperature(location="London"), default_api.get_temperature_date(location="London", date="2024-10-01")]
```
````
- **Hosted Gemini** — one `print(default_api...)` **statement per line**:
- **Hosted Gemini variant** — one `print(default_api...)` **statement per line**:
````text
```tool_code
print(default_api.get_current_temperature(location="London"))
@@ -89,11 +89,11 @@ Two encodings exist, both inside a single `tool_code` block:
```
````
Either way the calls are returned in source order; the application executes them and returns one result per call in the same order.
The OMP scanner extracts top-level call expressions in source order from either form. It mints a tool-call id for each parsed call; the text convention itself has no id.
## Tool-result format
Executed results are returned to the model in a ```` ```tool_outputs ```` block. Gemma 3 docs use assignment-style values (`result = 92.3`); for opaque tool output the block simply carries the returned text/JSON:
Executed results are returned to the model in ```` ```tool_outputs ```` blocks. OMP renders one complete block per result, in call order; it does not encode `isError` separately. Gemma 3 docs also show assignment-style values (`result = 92.3`), while opaque output can be returned as text/JSON:
````text
```tool_outputs
@@ -136,6 +136,8 @@ It's currently 11.4°C in London.
- **Skip string contents when scanning.** A call like `search(pattern="foo(")` contains a `(` inside a string; a naive `\w+\(` scan mis-detects `foo` as a callee. Track string state and only treat top-level `(` as a call opener.
- **Fence ambiguity.** The body terminates at the first bare ` ``` `; a string argument literally containing ` ``` ` will truncate the block early (rare, accepted limitation).
- **It leaks.** Because nothing is a special token, the format appears verbatim in normal responses when the model "decides" to call a tool but the structured decoder misfires. Production code reading raw text should detect ` ```tool_code ` and parse it; production code on the structured API should retry on `MALFORMED_FUNCTION_CALL`.
- **OMP streaming behavior.** The scanner buffers the entire `tool_code` body and emits tool events only after the closing fence; it does not stream partial arguments. An unterminated block is discarded on flush rather than exposed as text. Positional arguments and malformed keyword segments are skipped. In addition to ordinary quoted strings, the literal decoder accepts Python raw/byte/unicode prefixes, triple quotes, octal escapes, and `\x`/`\u`/`\U` escapes.
- **Transcript rendering.** OMP wraps the transcript in `<bos>` and Gemma-style `<start_of_turn>user|model` turns. `developer` text is prepended to the next user turn (or emitted as its own user turn if no user message follows); consecutive tool results become one user turn containing their separate `tool_outputs` blocks.
- **Variant divergence.** Gemma **4** abandoned this Pythonic form for a token-delimited brace syntax (`<|tool_call>call:NAME{…}<tool_call|>`) — a different convention documented in `gemma.md`. This spec covers hosted Gemini and Gemma 3.
## Sources
+4
View File
@@ -62,6 +62,7 @@ The OMP parser is the streaming `GemmaInbandScanner` (`packages/ai/src/dialect/g
1. finds the matching `<tool_call|>` close, skipping any `<|"|>…<|"|>` string span so a `<tool_call|>` sequence that appears inside a string value does not end the block early;
2. matches the `call:NAME{` head, then takes the brace body up to its depth-matched `}`;
3. splits that body into `key:value` pairs at top-level commas — bracket depth (`[]`, `{}`) and `<|"|>` string spans are skipped — and decodes each value per the grammar above, so nested lists and objects parse correctly (a single-level regex would not).
Calls are emitted only after the complete close marker arrives; there are no partial-argument events. If the stream is flushed with an unterminated tool block, OMP drops that incomplete block. A syntactically closed block with a missing final argument brace is still parsed from the available body.
## Multiple / parallel tool calls
@@ -76,6 +77,8 @@ Each result is `<|tool_response>response:NAME{output:VALUE}<tool_response|>`. `r
<|tool_response>response:read{output:<|"|>FILE<|"|>}<tool_response|>
```
The Gemma wire form has no dedicated success/error field. OMP renders `isError` results in the same `response:NAME{output:…}` shape as successful results, so any failure indication must be present in the result text itself.
## End-to-end example
`renderTranscript` output for a weather query. The system turn also carries the `<tools>` catalog and format guide (see *Tool definitions*, abbreviated here); the model's call merges with its tool response into one `model` turn (response right after the call), and the final answer is the next `model` turn. Turns are emitted back-to-back with no separator — only the `\n` after each role is literal:
@@ -94,6 +97,7 @@ The current weather in Tokyo is 15 degrees Celsius and sunny.<turn|>
- **Asymmetric pipes.** The closer is `<tool_call|>`, not `</tool_call>` or `<|tool_call>`. Matching the wrong pipe side will never close the block.
- **One call per block.** Unlike a JSON `tool_calls[]` array, parallelism is "more blocks", not "more entries in one block".
- **Bare scalars.** A value not wrapped in `<|"|>` is `true`/`false` → bool, `null`/`none` → null, numeric → number, otherwise a bare string (e.g. an unquoted enum or type name like `STRING`).
- **Tool-call ids are synthesized.** The format carries no id; OMP mints one when the scanner opens each call block and correlates rendered responses by the surrounding message order/name.
- **Not Gemma 3 / hosted Gemini.** Those use the Pythonic `tool_code` / `default_api` form in `gemini.md`. Gemma 4 replaced it with this token syntax; the two are not interchangeable.
## Sources
+38 -1
View File
@@ -280,7 +280,44 @@ With a server parser active (`--tool-call-parser glm45 --reasoning-parser glm45`
- **`skip_special_tokens` must be off.** Although the tool/think tags are `special: false`, vLLM forces `skip_special_tokens = False` when tools are enabled (defensive against transformers 5.x detokenization changes) so the literal `<tool_call>`/`</tool_call>` text survives for the regex.
- **Streaming.** Long string arguments used to be buffered until the closing tag (vLLM issue #32829); the current parser re-parses the accumulated text each delta and emits only the diff, streaming incremental string content with an open-quote-then-fill strategy and holding back any partial trailing tag (`partial_tag_overlap`). The streamed tool name is the text before the first `\n` or `<arg_key>`. SGLang implements the same as an explicit XML→JSON state machine (`INIT → IN_KEY → WAITING_VALUE → IN_VALUE`). Malformed tails (a missing `</arg_value>` before `</tool_call>`) are closed off heuristically.
- **Lineage — GLM-4.5 vs GLM-4.6:** identical wire format and identical `chat_template.jinja` (same content hash); the same `glm45` parser serves both.
- **Lineage — GLM-4.7 / GLM-5 changed the format.** Newer models drop the structural newlines: the function name may sit **directly** before the first `<arg_key>` (no newline), zero-argument calls may be `<tool_call>func</tool_call>`, and parallel calls may be emitted **back-to-back with no separator** (`…</tool_call><tool_call>…`). These require the distinct `Glm47MoeModelToolParser` (vLLM, `structural_tag_model="glm_4_7"`) / `Glm47MoeDetector` (SGLang), whose `func_detail_regex` makes the newline and the argument section optional (`<tool_call>\s*(\S+?)\s*(<arg_key>.*)?</tool_call>`). Do **not** use a GLM-4.7 stream to validate a GLM-4.5 parser or vice versa.
- **Lineage — GLM-4.7 / GLM-5 changed the format.** Newer models may omit
structural newlines: the function name can sit directly before the first
`<arg_key>`, zero-argument calls can be `<tool_call>func</tool_call>`, and
parallel calls can abut. vLLM/SGLang require their distinct GLM-4.7
parsers for this variant. omp's repository scanner is intentionally broader:
it accepts newline, `<arg_key>`, or `</tool_call>` as the name delimiter, so
the same `glm` dialect scanner handles both layouts.
## omp / pi converter behavior
The repository's `glm` dialect is an **owned in-band converter**. Select it
with `PI_DIALECT=glm`; legacy `PI_DIALECT=1` and `PI_DIALECT=true` also resolve
to GLM. With tools present, the agent appends the GLM format guide and compact
tool catalog to the system prompt, removes native provider tools, rewrites
prior calls/results into grammar-owned text, and scans assistant text back
into canonical pi events. GLM-family model affinity resolves to this dialect.
The owned renderer always emits the GLM-4.5 newline layout. It consults each
tool's normalized schema: string-only properties are emitted raw, while all
other values are JSON-serialized. Parallel calls are newline-separated. In
owned history, result batches become a synthetic user message containing
`<observation>` with one `<tool_response>` per result; the lower-level GLM
transcript renderer uses the model-native `<|observation|>` role marker
instead.
The scanner synthesizes `ptc_…` ids, emits `toolStart` once the name delimiter
arrives, and streams each argument body as keyed `toolArgDelta` events.
String-only schema properties stay verbatim; every other property is parsed
with strict `JSON.parse` after trimming and falls back to the original raw text
on failure. An unfinished key/value drops the open call on flush. The scanner
also heals narrowly recognizable model mistakes: `</arg_key>` used in place
of `</arg_value>`, a stray wrong closer before the real closer, and a missing
value closer immediately before the next argument or call close.
Thinking parsing is enabled by default and excludes `<think>…</think>` from
visible text. If `<tool_response>` appears in assistant output, the scanner
drops that tag and the remainder of the currently buffered chunk rather than
treating the hallucinated result as assistant content.
## Sources
+9
View File
@@ -126,6 +126,12 @@ Some Harmony serializers include an explicit JSON content type and place the rec
The arguments body is a raw JSON object. The optional `<|constrain|>json` content-type signals JSON (and is the hook for constrained/grammar-based decoding); the content-type may also be a bare word such as `code` (seen with built-in tools). Built-in tools differ only in channel and recipient: they typically render on `analysis`, with recipient `browser.search` / `browser.open` / `browser.find` or always `python`.
### OMP `harmony` dialect behavior
OMP emits the first form above: no `<|constrain|>` marker, recipient in the channel section, and compact JSON arguments. It synthesizes a call id on receipt because Harmony carries none. The stateful scanner accepts the recipient in either header section, strips a leading `functions.` from the exposed tool name, and treats any nonempty recipient other than `assistant` as a tool call (including built-ins such as `browser.search`).
Arguments are accumulated until `<|call|>`, `<|end|>`, or `<|return|>` and parsed with JSON repair. Empty arguments, or input that still cannot be parsed after repair, become `{}` rather than a scanner error. The scanner emits `toolStart` when the header completes and `toolEnd` only at the message terminator; `analysis` body chunks stream as thinking deltas, while ordinary assistant `commentary`/`final` bodies stream as text. Non-assistant messages, including tool-result envelopes, are skipped by this output scanner.
## Multiple / parallel tool calls
Harmony has no special "parallel" wrapper. Multiple calls are just multiple consecutive messages. The model may first emit an optional **preamble** — a *user-visible* assistant message on the `commentary` channel (unlike `analysis`, this is meant to be shown) — then one tool-call message per function. Each individual call still ends with its own `<|call|>` stop token, so a host that stops on `<|call|>` collects calls one at a time, executes, feeds the result back, and resumes:
@@ -149,6 +155,8 @@ The executed tool's output is fed back as a message whose **author/role is the t
The header ordering is `{toolname} to=assistant<|channel|>commentary`. Built-in tool results follow the same shape (e.g. `<|start|>browser.search to=assistant<|channel|>commentary<|message|>{"result": "https://openai.com/"}<|end|>`). The minimal form the renderer accepts when channel/recipient are not set on the message is just `<|start|>{toolname}<|message|>{output}<|end|>`, but emitting the full `to=assistant<|channel|>commentary` header is what the reference parser round-trips and is recommended. After appending the result, restart generation by emitting the next `<|start|>assistant`.
OMP always renders the full canonical result header shown above and passes `result.text` through verbatim. Harmony has no dedicated error bit, so `isError` is not represented separately; a failure must be described in the result payload.
## End-to-end example
Complete multi-turn weather exchange: system + developer prompt → user question → assistant analysis CoT → assistant commentary tool call → tool result → assistant final answer. This is a single contiguous token stream (newlines inside headers are only between top-level messages for readability; in practice messages are concatenated with no separator).
@@ -198,6 +206,7 @@ When a server (vLLM/SGLang/Ollama) bridges Harmony to Chat Completions JSON:
- **`tool_call_id`**: Harmony has no native call ID. The server synthesizes one (e.g. `call_abc123`) and is responsible for correlating the follow-up `role:"tool"` message back to the Harmony tool-result envelope (recipient `to=functions.<name>` / call order).
- **Tool result messages** (`{"role":"tool","tool_call_id":...,"content":...}`) are rendered into `<|start|>{toolname} to=assistant<|channel|>commentary<|message|>{content}<|end|>`. The server maps `tool_call_id` → the original function name to build the `{toolname}` author.
- **Reasoning**: `analysis`-channel text is surfaced as `reasoning_content` (vLLM/SGLang) or as a `reasoning`/`thinking` field, and is generally not echoed back on subsequent requests. `final`-channel text is the normal `message.content`. `commentary` preambles, if surfaced, also map to assistant content.
- **OMP transcript rendering:** `developer`, `user`, and other non-assistant roles map directly to Harmony envelopes. Assistant messages emit, in order, a complete `analysis` message for thinking, a complete `final` message for visible text, then one `commentary` call message per tool call. Thus visible text accompanying a tool call is rendered as `final`, not as a commentary preamble. Tool-result runs become consecutive canonical tool-author envelopes.
- **`tools` / `tool_choice`** request fields are compiled by the chat template into the developer-message `namespace functions { ... }` block; the system message gains the commentary-routing line.
## Parsing notes & gotchas
+41 -2
View File
@@ -159,17 +159,56 @@ Moonshot's hosted API (`platform.moonshot.ai`) exposes both OpenAI- and Anthropi
## Parsing notes & gotchas
- **ID → name parsing differs between references.** The official `tool_call_guidance.md` extracts the name with `function_id.split('.')[1].split(':')[0]`, which assumes the ID is exactly `functions.{name}` with no extra dots. vLLM uses the more robust `function_id.split(":")[0].split(".")[-1]` (takes the last dot-segment before `:{idx}`). Prefer the vLLM form so function names containing `.` are handled.
- **ID → name parsing differs between references.** The official
`tool_call_guidance.md` uses `function_id.split('.')[1].split(':')[0]`,
while vLLM and omp take the last dot-separated segment before the colon.
The latter tolerates additional namespace segments, but neither convention
can preserve a literal dot as part of the function name; tool names SHOULD
follow the documented `functions.{name}:{idx}` shape without dots in
`{name}`.
- **Extraction regexes differ too.** Guidance: `<\|tool_call_begin\|>\s*(?P<tool_call_id>[\w\.]+:\d+)\s*<\|tool_call_argument_begin\|>\s*(?P<function_arguments>.*?)\s*<\|tool_call_end\|>`. vLLM: ID class is `[^<]+:\d+` and the argument body uses a negative lookahead `(?:(?!<\|tool_call_begin\|>).)*?` so adjacent calls aren't merged. Both run with `DOTALL`.
- **`skip_special_tokens` must be False.** The parser depends on the literal marker text surviving detokenization; vLLM forces `skip_special_tokens = False` when tools are enabled and `tool_choice != "none"`. If markers are stripped, no tool call is detected.
- **Arguments are unvalidated raw text.** Whatever the model emits between the argument marker and `<|tool_call_end|>` is passed straight through as the `arguments` string; it must be valid JSON for downstream `json.loads`, and the model can emit malformed/truncated JSON. Validate before executing.
- **Index semantics.** `{idx}` is the per-turn call counter starting at `0`; it is not a global counter and resets each assistant turn. Do not assume IDs are unique across turns — disambiguate by turn when persisting history.
- **Streaming marker splits.** Section and call markers can be split across token boundaries. vLLM holds back any trailing suffix that partially matches a marker (`partial_tag_overlap`) to avoid leaking marker bytes into streamed content, and only streams a call's name once its header is fully received.
- **Streaming marker splits.** Section and call markers can be split across token boundaries.
vLLM holds back any trailing suffix that partially matches a marker and streams
argument fragments. omp's owned scanner also holds partial markers, but buffers a
call's arguments until `<|tool_call_end|>` and emits no `toolArgDelta` events.
- **`finish_reason` varies by engine.** The official guide explicitly warns the terminal `finish_reason` for tool calls "may vary across different engines"; loop on `finish_reason == "tool_calls"` but be defensive.
- **Engine fallback.** Kimi K2 reuses the DeepSeek-V3 architecture; `config.json` sets `model_type: "kimi_k2"` so engines apply the right parser. If you force `model_type: "deepseek_v3"` as a compatibility workaround, no native Kimi tool parser is available and you must parse the `<|tool_calls_section_*|>` markers manually.
- **Parser availability.** vLLM ships both a Python (`KimiK2ToolParser`) and a newer Rust tool parser; SGLang implements its own `kimi_k2` parser. All key off the same five markers and the `functions.{name}:{idx}` ID convention documented here.
- **Whitespace artifact.** When no `system` message is supplied, the template injects the default system prompt and a small `\n ` (newline + two spaces) can appear before the first `<|im_user|>` marker. It is harmless (tokenizes around the markers), but supplying an explicit system message yields the clean streams shown above.
## omp / pi converter behavior
The repository's `kimi` dialect is an **owned in-band converter**. Select it
with `PI_DIALECT=kimi` (or the equivalent agent configuration). With tools
present, the agent appends the Kimi guide and compact tool catalog to the
system prompt, removes native provider tools, rewrites prior calls/results in
Kimi text form, and converts streamed output back into canonical pi events.
Kimi-family model affinity resolves to this dialect.
The renderer emits one section per assistant call batch. It preserves a
pre-existing id that already begins with `functions.`; otherwise it generates
`functions.{name}:{batchIndex}`. Tool results are rendered as consecutive
`<|im_system|>{name}<|im_middle|>## Return of …<|im_end|>` turns, and canonical
tool-result messages are collapsed into one synthetic user message containing
that text.
The scanner recognizes only calls inside a section. Once the argument marker
arrives it preserves the raw header as the call id and derives the name from
the last dot-separated segment before the first colon. It emits `toolStart` at
that point, buffers the argument body until `<|tool_call_end|>`, then applies
the shared repairing JSON parser and emits `toolEnd`; it does **not** emit
incremental argument deltas. Invalid/non-object arguments normalize to `{}`,
and an unfinished call is discarded on flush. Section markers are suppressed
from visible text, while an isolated call marker outside a section remains
ordinary text.
Thinking parsing is enabled by default and maps `<think>…</think>` to thinking
events. `parseThinking: false` leaves those tags and their contents in visible
text.
## Sources
- Model card (Tool Calling section, OpenAI-style example, deployment/API notes): https://huggingface.co/moonshotai/Kimi-K2-Instruct
+134 -246
View File
@@ -1,276 +1,164 @@
# pi-native tool-call format
# pi-native auth-gateway transport
The **pi-native** format is the tool-call serialization used by the omp / pi coding agent. Unlike the JSON-in-a-tag conventions (Hermes/Qwen, Harmony) and unlike the fully separate JSON content-block channel (Anthropic Messages API), pi-native serializes each call as an **XML-flavored block** whose tag carries the tool name — `<call:NAME>…</call:NAME>` — and whose arguments are child elements named after the parameters. It is **schema-driven**: the tool's JSON Schema decides how each value is typed (string vs number vs object vs array) and which compact spellings are legal.
`pi-native` is the lossless transport between a pi-ai client and an
`omp auth-gateway`. It is **not a textual tool-call dialect**: there is no
`<call:NAME>` grammar, parser, renderer, or `PI_DIALECT=pi-native` value in the
current implementation. Tool calls remain canonical pi-ai `ToolCall` content
blocks inside `Context` and `AssistantMessageEvent`.
This document is a **specification** of the format (it is the contract a renderer must emit and a parser must accept), not a reverse-engineering of trained model weights. It is designed around four goals:
Use this transport when the client already speaks pi-ai and the gateway owns
provider credentials—for example, a containerized omp talking to a host
gateway or a robomp slot talking to its sidecar. OpenAI/Anthropic-compatible
routes translate and can lose pi-specific fields; pi-native sends the
canonical types directly, preserving service tier, cache markers, thinking
budgets, tool-choice variants, images, and tool-call IDs.
- **Token economy** — the common cases (a single scalar argument; a single string payload) collapse to one short line.
- **Verbatim payloads** — a large multi-line string argument (a patch body, a file's contents, a shell script) is carried **raw**, with no JSON string-escaping and no entity-encoding, terminated by the call's own unique closing tag.
- **Human legibility** — a call reads like the function it denotes; nesting maps to nesting.
- **Lenient parsing** — the tags are plain text matched by a tolerant parser (regex / streaming state machine), not a strict XML parser; output is *not* required to be well-formed XML.
## Configuration and dispatch
Scope: pi-native specifies only the **tool-call (and argument) serialization**. It is **envelope-agnostic** — the `<call:…>` blocks are emitted as ordinary assistant text and embed unchanged in any conversation envelope (ChatML, Harmony, the Anthropic two-role shape, …). Reasoning channels, role markers, and result delivery are the host envelope's concern; the one envelope-level requirement pi-native imposes is in [Tool-result correlation](#tool-result-correlation).
A model opts in with:
Lineage: the attribute spelling (`<call:read path="…"/>`) follows Anthropic's modern attribute XML (`<invoke name="…">` / `<parameter name="…">`); the schema-driven **unquoted** value rule (a bare string value carries no quotes; non-strings are JSON) follows GLM-4.5's `<arg_value>` convention. pi-native folds both into one recursive, name-on-the-tag grammar and adds the verbatim **inline body** for bulk string arguments. See [`anthropic.md`](./anthropic.md) and [`glm-4.5.md`](./glm-4.5.md) in this folder.
## Structural tags
pi-native has **no special tokens**. Every marker is plain UTF-8 text that BPE-splits like any other text and survives detokenization unchanged; a parser matches the tags as literal substrings (and MUST work even when the surrounding stream is not valid XML). All brackets are ASCII `<` `>` `/` and the literal colon `:`. There are no namespaces, no `<?xml?>` prolog, no entity expansion, and no CDATA sections.
| Tag (verbatim) | Role |
|---|---|
| `<call:NAME …>` … `</call:NAME>` | One tool call. `NAME` is the tool/recipient name; it is repeated on the closing tag. |
| `<call:NAME …/>` | Self-closing tool call (all arguments supplied as attributes). |
| `<KEY>` … `</KEY>` | One argument (or one nested field). `KEY` is the parameter name. |
| `<KEY …/>` | Self-closing argument: an object-valued field whose scalar sub-fields are attributes, or an empty value. |
| `KEY="…"` / `KEY='…'` / `KEY=…` | An attribute: a scalar field given inline on a tag. Quotes are **delimiters, not type markers** (see [value coercion](#value-coercion)). |
Names (tool names and parameter names) match `^[A-Za-z_][A-Za-z0-9_-]*$`. The `call:` prefix is a literal four-character marker plus the colon; the colon is what distinguishes a call block from a nested argument element of the same name.
## Tool-call forms
A single call has three interchangeable surface forms. Which forms are legal for a given tool is decided by its parameter schema; a renderer SHOULD pick the most compact legal form, and a parser MUST accept all three.
### 1. Element form (canonical, fully general)
Each top-level argument is a child element named after the parameter; the element body is the value. This form expresses every schema — scalars, strings, arrays, and nested objects:
```text
<call:read>
<path>src/server/auth.ts</path>
<offset>50</offset>
</call:read>
```yaml
transport: pi-native
baseUrl: http://gateway.internal:4000
```
→ `read({ "path": "src/server/auth.ts", "offset": 50 })` (`offset` is JSON because the schema types it as a number; `path` is a verbatim string).
### 2. Attribute form (compact scalars)
When the arguments being passed are **top-level scalars** (string, number, integer, boolean, null), they MAY be written as attributes on the call tag. With every argument as an attribute the tag is self-closing:
`baseUrl` MUST identify an `omp auth-gateway` (or compatible service). Missing
`baseUrl` fails with:
```text
<call:read path="src/server/auth.ts"/>
pi-native transport requires `baseUrl` on model MODEL_ID (set it on the provider config in models.yml)
```
→ `read({ "path": "src/server/auth.ts" })`.
When `model.transport === "pi-native"`, `streamSimple` bypasses the normal
per-API provider implementation and calls `streamPiNative`. The client removes
trailing slashes from `baseUrl` and posts to `/v1/pi/stream`.
Attributes and child elements MAY be combined on a non-self-closing call tag — attributes carry the scalars, child elements carry anything structured:
The gateway bearer is the resolved model/API key. It is sent as
`Authorization: Bearer …`, never in the JSON options. Model headers are also
forwarded; an explicit `model.headers.Authorization` takes precedence over the
resolved key.
`transport` changes only dispatch. Pricing, context window, maximum-token and
thinking metadata still resolve locally from the model catalog.
## Request
```http
POST /v1/pi/stream
Content-Type: application/json
Accept: text/event-stream
{
"modelId": "provider/model-id",
"context": {
"systemPrompt": ["..."],
"messages": [],
"tools": []
},
"options": {},
"stream": true
}
```
The client always qualifies `modelId` as `${provider}/${id}` and always
requests streaming. The server also accepts `modelId`, a string `model`, or
`model.id`; its lower-level request parser defaults `stream` to `true`.
Validation at the gateway boundary is intentionally shallow:
- the body MUST be an object;
- a non-empty model identifier MUST be present;
- `context` MUST be an object with a `messages` array;
- when present, `context.systemPrompt` and `context.tools` MUST be arrays.
Invalid shapes produce validation errors. Canonical message/tool internals are
not revalidated at this boundary; downstream failures surface as gateway
upstream errors.
## Options crossing the wire
The server accepts this `SimpleStreamOptions` subset:
`temperature`, `topP`, `topK`, `minP`, `presencePenalty`,
`frequencyPenalty`, `repetitionPenalty`, `stopSequences`, `maxTokens`,
`cacheRetention`, `cachedContent`, `headers`, `initiatorOverride`,
`maxRetryDelayMs`, `metadata`, `sessionId`, `promptCacheKey`, `promptCache`,
`statefulResponses`, `streamFirstEventTimeoutMs`, `streamIdleTimeoutMs`,
`reasoning`, `disableReasoning`, `hideThinkingSummary`, `thinkingBudgets`,
`toolChoice`, `serviceTier`, `kimiApiFormat`, `syntheticApiFormat`,
`preferWebsockets`, `openrouterVariant`, and `loopGuard`.
Unknown, `null`, and `undefined` option values are silently dropped by the
server. The client additionally strips runtime/server-owned fields:
`signal`, `apiKey`, `fetch`, `onPayload`, `onResponse`, `onSseEvent`,
`execHandlers`, `cursorExecHandlers`, `cursorOnToolResult`, and
`providerSessionState`. `onResponse` still runs locally against the gateway's
HTTP response; callbacks and runtime handles themselves never cross the wire.
## Streaming response
Each canonical `AssistantMessageEvent` is JSON-serialized without reshaping
and SSE-framed:
```text
<call:read path="src/server/auth.ts">
<offset>50</offset>
</call:read>
data: {"type":"start",...}
data: {"type":"text_delta",...}
data: {"type":"done","reason":"stop","message":{...}}
data: [DONE]
```
An attribute whose value cannot be represented as a scalar (an object or array argument) MUST use the element form instead — there is no attribute spelling for structured values on a call tag.
The server stops after a canonical `done` or `error` event and then writes
`[DONE]`. If its event iterator throws first, it best-effort emits
`{"type":"error","reason":"error","errorMessage":"..."}` followed by `[DONE]`.
Cancelling the HTTP body propagates cancellation to the gateway request.
### 3. Inline-body form (verbatim string payload)
The client parses every event and pushes it verbatim into
`AssistantMessageEventStream`; there is no partial-content reconstruction or
tool conversion. A caller abort cancels the response body. First-event and
idle watchdogs use request options when supplied, otherwise the standard
`PI_STREAM_FIRST_EVENT_TIMEOUT_MS` / `PI_STREAM_IDLE_TIMEOUT_MS` policy.
The initial `start` event is not considered progress for the idle watchdog.
When the tool's parameters are **all strings** — most often a single string parameter — the call body MAY be the argument value written **verbatim**, with no child element tags:
If the SSE connection closes without a terminal event, the client synthesizes
a terminal assistant boundary so `.result()` cannot hang: caller cancellation
becomes an `error` event with `stopReason: "aborted"`; otherwise it becomes an
empty `done` event with `stopReason: "stop"`.
```text
<call:edit>
*** Begin Patch
@@ src/server/auth.ts
- return user;
+ return user ?? null;
*** End Patch
</call:edit>
```
→ `edit({ "input": "*** Begin Patch\n@@ src/server/auth.ts\n- return user;\n+ return user ?? null;\n*** End Patch" })`.
Rules for the inline body:
- The body fills the **first parameter not already supplied by an attribute**. With no attributes that is simply the first parameter (the "first argument verbatim").
- It is permitted only when that target parameter is **string**-typed (the enabling condition "the type only contains string arguments"); any *other* parameters set on the same call MUST be scalars given as attributes.
- The value is captured **verbatim** up to the call's own closing tag `</call:NAME>`. No JSON escaping, no entity decoding. The body MAY freely contain `<`, `>`, `&`, quotes, JSON, even other `<call:…>`-looking text — the only sequence it MUST NOT contain is the literal closer `</call:NAME>`. Because that closer carries the tool name, collisions are far rarer than with a short generic delimiter.
- Whitespace: a single newline immediately after the opening `>` and a single newline immediately before the closing `</` are treated as block delimiters and are **not** part of the value; all other whitespace (indentation, blank lines, trailing spaces) is preserved exactly.
A multi-string tool can still use the inline body for its bulk argument by passing the others as attributes — the body then fills the first parameter left unset:
```text
<call:write path="notes/todo.md">
# TODO
- ship pi-native parser
</call:write>
```
→ `write({ "path": "notes/todo.md", "content": "# TODO\n- ship pi-native parser" })` (here `path` is given by attribute, so the body fills the next string parameter, `content`).
Inline-eligible tools may always fall back to the element form; `<call:edit><input>…</input></call:edit>` and the inline `<call:edit>…</call:edit>` are equivalent.
## Value model
The body of a call (and of any nested element) maps to JSON by a single recursive rule set. **Typing is driven by the parameter's JSON Schema**; the parser only falls back to syntactic heuristics when no schema is available.
### Value coercion
Let `coerce(text, type)` produce the JSON value for a captured scalar `text`:
- `type == "string"` → the value is `text`, **verbatim** (never JSON-parsed, never unquoted). This is why `<path>4</path>` for a string parameter is the string `"4"`, and a Windows path `C:\new\tab` survives intact.
- `type` is a non-string scalar (`number` / `integer` / `boolean` / `null`) → `JSON.parse(text)` (so `<offset>50</offset>` → `50`, `<recursive>true</recursive>` → `true`).
- `type` unknown (no schema) → **best-effort JSON coercion**: try `JSON.parse(text)`; on success use the parsed value (number, boolean, null, quoted-string, object, or array); on failure treat `text` as a literal string. So a bare `4` becomes the number `4`, `foo.ts` (not valid JSON) becomes `"foo.ts"`.
The same `coerce` applies to **attribute values** after the surrounding quotes (if any) are stripped — the quotes are XML delimiters only. Hence both spellings below are identical, and both yield the **number** `4` (not the string `"4"`) when `y` is untyped/numeric:
```text
<object y=4/> → { "object": { "y": 4 } }
<object y="4"/> → { "object": { "y": 4 } }
```
Consequence to internalize: under loose/no schema, quoting does **not** force a string — `"4"` still coerces to `4`. To carry a numeric-looking value *as a string*, the parameter MUST be `string`-typed in the schema (then the verbatim rule keeps `"4"` → `"4"`). Unquoted attribute values run until whitespace or the closing `>` / `/>`; spaces around `=` are tolerated; a bare attribute with no `=value` (e.g. `<call:tool dry_run/>`) denotes boolean `true`.
### Scalars and strings
A scalar argument is one element (or one attribute). String values are unquoted and verbatim; non-string scalars are JSON literals:
```text
<call:bash command="ls -la" timeout=30/>
```
→ `bash({ "command": "ls -la", "timeout": 30 })`.
### Arrays — repeat the element
An array-typed field is expressed by **repeating** its element; each occurrence contributes one item, in order:
```text
<list>x</list>
<list>y</list>
```
→ `"list": ["x", "y"]`.
A field the schema types as an **array always yields an array, even for a single occurrence** — so one `<list>x</list>` under an array-typed `list` is `["x"]`, not `"x"`. When no schema is available the parser falls back to a count heuristic: a name appearing **2+ times among its siblings** is an array; a name appearing **once** is a scalar (so schema typing is the only way to express a one-element array with the heuristic alone). Item values coerce by the array's item type (`<ports>80</ports><ports>443</ports>` → `[80, 443]` for a `number[]`); arrays of objects repeat a nested block (see below). There is no attribute spelling for an array (attributes cannot repeat) — arrays require element form.
### Objects — a nested block
An object-typed field opens its own block and follows the **same rules recursively**: its child elements become its properties, repeated children become arrays, and nested object children open further blocks.
```text
<object>
<list>x</list>
</object>
```
→ `"object": { "list": ["x"] }` (with `object` typed object and `list` typed array).
An object's **scalar** sub-fields MAY instead be written as attributes — `<object y=4/>` is shorthand for `<object><y>4</y></object>`. Attributes and child elements may be combined on the same object element (attributes for scalars, children for structured sub-fields). An empty object is `<object/>` or `<object></object>` → `{}`.
### Recursion
The call body, an object element's body, and an array item's body are all parsed by the identical procedure. Parsing element `E` (tag = field name `F`, schema type `T`):
1. Gather `E`'s attributes → scalar properties via `coerce`.
2. Determine `E`'s body shape from `T` (or, with no schema, from whether the body's first non-whitespace content is a child tag):
- `T` object → properties from child elements (+ the attributes from step 1).
- `T` array (item type `Ti`) → collect **all** siblings named `F`; each occurrence is one item parsed as `Ti`.
- `T` scalar/string → the body is captured text; value = `coerce(text, T)`.
3. The call itself is element `E` with no enclosing key: its attributes + child elements **are** the arguments object directly (the tool name on `<call:NAME>` is the recipient, not a key).
## Multiple / parallel tool calls
There is no wrapper element around a set of calls. Parallel calls are simply **consecutive `<call:…>` blocks** in one assistant turn (separated by whitespace/newlines; interleaved prose is allowed and is ordinary content):
```text
<call:read path="src/a.ts"/>
<call:read path="src/b.ts"/>
```
A parser returns these as `tool_calls[0]`, `tool_calls[1]`, … in emission order. The host executes them and returns one result per call, in the same order (see correlation, next).
## Tool definitions and schema dependence
pi-native does not prescribe how tools are advertised; a host typically lists them as JSON Schema, exactly as the OpenAI / Anthropic / Hermes families do. What pi-native **requires** is that the parser have access to each tool's parameter schema, because the schema is what disambiguates:
- string (verbatim, unquoted) vs other scalar (JSON) values;
- a one-element array vs a scalar (a single `<list>…</list>`);
- which body shape (text vs nested members) a non-self-closing element carries;
- whether the inline-body form is legal (first unset parameter is a string).
Without a schema the parser MUST degrade gracefully to the syntactic fallbacks named above (JSON-coerce scalars; repetition-counts for arrays; child-tag presence for object bodies). The fallbacks are lossy at exactly the ambiguous points the schema would resolve, so production hosts SHOULD always supply the schema.
## Tool-result correlation
pi-native calls carry **no per-call wire id** (like GLM and Qwen, unlike Anthropic's `toolu_…`). Results are therefore correlated to calls **positionally, by emission order**: the host delivers tool outputs in the same order the `<call:…>` blocks appeared, using whatever its envelope provides for tool output (a `tool`/`user` turn, a Harmony tool message, an Anthropic `tool_result` block, …). When a transport requires an id (e.g. an OpenAI-compatible bridge), the host synthesizes one and maintains the call↔result mapping itself; the id never appears in the pi-native text.
## End-to-end example
A short agent turn exercising all three forms plus nesting. Schemas in play: `read(path: string, offset?: number)`, `bash(command: string, timeout?: number)`, `edit(input: string)`, and a synthetic `configure(object: { list: string[]; y?: number })`.
```text
I'll inspect the file, run the tests, then apply the fix.
<call:read path="src/server/auth.ts"/>
<call:bash command="bun test src/server/auth.test.ts" timeout=120/>
<call:configure>
<object y=4>
<list>alpha</list>
<list>beta</list>
</object>
</call:configure>
<call:edit>
*** Begin Patch
@@ src/server/auth.ts
- return user;
+ return user ?? null;
*** End Patch
</call:edit>
```
Parses to four calls, in order:
The client consumes streaming responses only. The server endpoint also
supports `stream: false`, returning:
```json
[
{ "name": "read", "arguments": { "path": "src/server/auth.ts" } },
{ "name": "bash", "arguments": { "command": "bun test src/server/auth.test.ts", "timeout": 120 } },
{ "name": "configure", "arguments": { "object": { "y": 4, "list": ["alpha", "beta"] } } },
{ "name": "edit", "arguments": { "input": "*** Begin Patch\n@@ src/server/auth.ts\n- return user;\n+ return user ?? null;\n*** End Patch" } }
]
{ "message": { "role": "assistant", "content": [] } }
```
Note: `timeout=120` and `y=4` are JSON numbers (numeric/untyped scalars), `path` and the `list` items are verbatim strings (string-typed), `object` opens a nested block whose `y` rides as an attribute while `list` repeats into an array, and the `edit` body is captured verbatim up to `</call:edit>` despite containing `@@`, `-`/`+`, and other non-XML text.
with the full canonical `AssistantMessage` in `message`.
## Grammar (lenient EBNF)
## Errors
This is the shape a tolerant parser accepts; it is intentionally looser than XML (mismatched-but-recoverable tails are closed heuristically — see gotchas).
Pre-stream HTTP failures use:
```ebnf
stream ::= ( text | call )*
call ::= self-call | block-call
self-call ::= "<call:" Name attr* ws? "/>"
block-call ::= "<call:" Name attr* ">" call-body "</call:" Name ">"
call-body ::= members | inline-text ; inline-text only if first param is string
members ::= ( ws | element )*
element ::= self-element | block-element
self-element ::= "<" Name attr* ws? "/>" ; object via attrs, or empty value
block-element::= "<" Name attr* ">" ( members | scalar-text ) "</" Name ">"
attr ::= ws Name ( ws? "=" ws? attr-val )? ; bare Name → boolean true
attr-val ::= '"' dq-chars '"' | "'" sq-chars "'" | bareword
Name ::= [A-Za-z_] [A-Za-z0-9_-]*
scalar-text ::= < any chars up to the matching close tag, verbatim >
inline-text ::= < any chars up to "</call:" Name ">", verbatim >
```json
{ "error": { "type": "rate_limit_error", "message": "..." } }
```
## Parsing notes & gotchas
with the appropriate HTTP status, `Content-Type: application/json`, and
`Cache-Control: no-store`. The client converts this shape into
`AuthGatewayError`, preserving status, response headers, and `type`. A
nonconforming error body falls back to
`auth-gateway STATUS: BODY_OR_STATUS_TEXT`. A successful response with no body
is also an `AuthGatewayError`.
- **Schema decides string-vs-JSON.** A `string`-typed value is verbatim and unquoted; everything else is JSON. With no schema, scalars best-effort JSON-coerce and fall back to string. This is the single most error-prone rule (identical to GLM-4.5's unquoted strings): emitting `"San Francisco"` for a string parameter yields the literal value *including the quote characters*.
- **Quotes are delimiters, not types.** `y="4"` and `y=4` both coerce to the number `4` under loose/no schema. Quoting an attribute never makes it a string; only a `string` schema type does.
- **Arrays = repetition; single-element arrays need the schema.** Two same-named siblings is unambiguously an array. One occurrence is a scalar under the count heuristic and an array only because the schema says so — a parser without the schema cannot tell `<list>x</list>` (scalar) from a one-element array.
- **Verbatim bodies are delimited by the named closer.** The inline body and any `string`-typed element body are captured up to their matching `</call:NAME>` / `</KEY>`. A body that contains that exact closing sequence truncates early; there is no escaping mechanism. The inline body's risk is minimal because the delimiter includes the tool name (`</call:edit>`), but a short string-typed *element* (e.g. `<note>…</note>`) is more exposed — prefer the inline-body form for any value that might contain markup, or keep such values in the single-string inline payload.
- **Element form vs inline body.** A block call whose body's first non-whitespace content is a child tag matching a known parameter is parsed as element form; otherwise (all-string tool) it is the inline body. A string value that legitimately *starts* with a `<param>`-looking token is the one ambiguity — emit such a tool in element form, or rely on the schema (a tool with structured params is never inline-eligible).
- **No ids; order is the contract.** Calls carry no id; results MUST be returned in call order. Reordering results silently misattributes them.
- **Lenient, not strict XML.** Do not feed pi-native to an XML parser: tag names contain a colon (`call:read`), attribute values may be unquoted, bodies are not entity-encoded, and the stream need not be balanced beyond each call's own open/close. Match the tags as literals (regex / streaming state machine).
- **Streaming.** A stateful parser emits the tool name as soon as `<call:NAME` closes, then streams attribute/child deltas; for an inline body it streams body text incrementally and holds back any partial trailing `</call:` until it can decide whether it is the closer. Coercion of a scalar can only finalize at the value's close tag (a partial number/boolean is not yet valid JSON).
- **Whitespace.** Element/inline bodies preserve all whitespace except one leading and one trailing newline that delimit the block. Attribute and indentation whitespace between tags is insignificant.
## Source of truth
## Sources
pi-native is specified here; it is not derived from a published model template. Its two direct influences are documented in this folder:
- Anthropic attribute XML (`<invoke name="…">` / `<parameter name="…">`, "parsed with regular expressions", not required to be valid XML): [`anthropic.md`](./anthropic.md).
- GLM-4.5 schema-driven, **unquoted** string values vs JSON non-strings, and positional (id-less) call↔result correlation: [`glm-4.5.md`](./glm-4.5.md).
- `packages/catalog/src/types.ts` — `Model.transport`
- `packages/ai/src/stream.ts` — pi-native dispatch
- `packages/ai/src/providers/pi-native-client.ts` — request, auth, SSE and
timeout behavior
- `packages/ai/src/providers/pi-native-server.ts` — request validation,
option allow-list, SSE and error envelopes
- `packages/ai/src/auth-gateway/server.ts` — `/v1/pi/stream` route and gateway
model/credential resolution
+41 -1
View File
@@ -186,6 +186,39 @@ message.tool_calls = [
]
```
## omp / pi converter behavior
The repository's `qwen3` dialect is an **owned in-band converter**. Select it
with `PI_DIALECT=qwen3` (or the equivalent agent configuration). With tools
present, the agent appends the Qwen3 format guide and compact tool catalog to
the system prompt, removes native provider tools, rewrites earlier calls and
results as text in this syntax, and scans streamed output back into canonical
pi tool-call events. `hermes` remains a separate selectable dialect even
though both emit the same basic JSON-in-`<tool_call>` convention.
The catalog's current family-affinity helper maps every model id containing
`qwen` to `qwen3`, including Qwen3-Coder. That broad affinity does not change
the format distinction described below, so callers must explicitly select the
appropriate dialect for Coder endpoints.
The omp renderer always writes a nested `arguments` object and renders
parallel calls newline-separated. Results become newline-delimited
`<tool_response>` blocks inside the synthetic user history message. The
scanner mints an id (`ptc_…`), emits `toolStart` as soon as the leading JSON
contains a complete string `name`, and waits for `</tool_call>` before emitting
`toolEnd`; it does not stream argument deltas. At close it uses the shared
repairing JSON parser. For compatibility it also accepts a stringified
`arguments` value and parses it once more, although the owned renderer never
emits that shape. A malformed completed block, a non-object argument value, or
an unfinished block does not become visible fallback prose: malformed calls
are consumed, with a string parse failure/non-object normalized to `{}` or the
whole call omitted when its outer object/name cannot be recovered.
Thinking parsing is enabled by default: `<think>…</think>` becomes thinking
events and is excluded from visible text. Callers creating the scanner can set
`parseThinking: false`, in which case thinking markup is left as ordinary
text.
## Parsing notes & gotchas
- **Arguments object vs string:** on the wire `arguments` is a nested JSON object; the OpenAI layer hands it back as a JSON string. Code that reads the raw stream must parse an object; code that reads the API must `json.loads` the string. Do not double-encode.
@@ -194,7 +227,14 @@ message.tool_calls = [
- **Thinking toggle:** `enable_thinking=False` (passed via `chat_template_kwargs={"enable_thinking": False}` over the OpenAI API, or `tokenizer.apply_chat_template(..., enable_thinking=False)`) injects an empty `<think>\n\n</think>\n\n` into the generation prompt, hard-suppressing reasoning. Soft switches `/think` and `/no_think` in a user/system message flip it per-turn when thinking is enabled. Greedy decoding is discouraged for Qwen3 (repetition risk).
- **History rerender asymmetry:** when `apply_chat_template` re-renders a stored conversation, it emits the `<think>` block only for the final assistant message or messages carrying `reasoning_content`; reasoning from earlier turns is dropped. So a stored intermediate tool-call assistant turn shows no `<think>` block, while the live generation step that produced it was prefixed with one (in non-thinking mode). Reasoning is preserved only within the current multi-step tool sequence (after the last real user query).
- **Reasoning models + stopword templates:** Qwen warns against ReAct-style stopword tool templates for Qwen3, since reasoning text may contain the stopwords and corrupt parsing — use this native Hermes template instead.
- **Robustness:** the format is prompt/template-driven, so malformed output is possible (truncated JSON, missing `</tool_call>`, prose mixed into a call, an array serialized as a string). Production parsers should tolerate and, on failure, fall back to treating the text as content. Named / `required` tool_choice routes through vLLM's structured-outputs backend for guaranteed-parseable arguments.
- **Robustness:** the format is prompt/template-driven, so malformed output is possible
(truncated JSON, missing `</tool_call>`, prose mixed into a call, or stringified
arguments). vLLM may fall back to content depending on its parser path; omp's
owned scanner instead consumes a recognized block and emits no call when the
outer JSON/name cannot be recovered. Named / `required` tool choice can route
through vLLM's structured-outputs backend when using vLLM native tools, but
owned mode sends no native provider tool definition and therefore cannot rely
on that backend.
- **Version/scope:** this `hermes` template covers `Qwen3-*`, `Qwen2.5-*`, and `QwQ-32B`. It does **not** cover `Qwen3-Coder`, which uses a different XML scheme parsed by vLLM's `qwen3_xml` parser — a separate convention.
## Sources