feat: added native tool inventory rendering with TypeScript signatures

- Added `jsonSchemaToTypeScript` and `renderToolInventory` to generate tool blocks with TypeScript signatures.
- Added `examples` and `TSchema` fields to dump-tool metadata and passed them through prompt rendering.
- Changed Harmony invocation rendering to omit `<|constrain|>json` markers in tool call payloads.
- Added compact native tool list-mode inventory rendering with full `# Tool:` output elsewhere.
This commit is contained in:
can1357
2026-06-15 08:12:58 +02:00
parent ba82ed6e59
commit 8b7dd10a8a
23 changed files with 603 additions and 100 deletions
+3 -3
View File
@@ -138,7 +138,7 @@ Here are the functions available in JSONSchema format:
{{ TOOL CONFIGURATION }}
```
`{{ TOOL DEFINITIONS IN JSON SCHEMA }}` is your `tools` array serialized to JSON Schema. `{{ FORMATTING INSTRUCTIONS }}` is the (unpublished) block teaching the model the `<function_calls>`/`<invoke name>`/`<parameter name>` syntax shown under [Tool-call format → underlying XML](#underlying-xml-modern-attribute-form). The note "parsed with regular expressions" is why output need not be well-formed XML.
`{{ TOOL DEFINITIONS IN JSON SCHEMA }}` is your `tools` array serialized to JSON Schema. `{{ FORMATTING INSTRUCTIONS }}` is the (unpublished) block teaching the model the XML syntax with `antml:` namespace prefixes (shown under [Tool-call format → underlying XML](#underlying-xml-modern-attribute-form-with-antml-namespace)). The note "parsed with regular expressions" is why output need not be well-formed XML.
---
@@ -175,7 +175,7 @@ Key facts for a parser:
- A leading `text` block is optional and informational; do not rely on its wording.
- Match calls to results by `id` → `tool_use_id`.
### Underlying XML (modern attribute form)
### Underlying XML (modern attribute form with antml: namespace)
Before the API converts it, the model literally emits an XML block. The current (Claude 3+) form is attribute-based:
@@ -188,7 +188,7 @@ Before the API converts it, the model literally emits an XML block. The current
</function_calls>
```
`[Partially verified]` Anthropic does not publish the literal `{{ FORMATTING INSTRUCTIONS }}`, so the exact tag spelling for current models is reconstructed from the trained format (and matches the task's reference anchor) rather than an official verbatim doc. In production, current Claude models prefix these tags with an `antml:` XML namespace (e.g. `<function_calls>`, `<invoke name="…">`, `<parameter name="…">`); the namespace is widely observed but **not** documented officially — treat it as `[unverified]`. The API strips all of this and exposes only the JSON `tool_use` block; integrators should target the JSON, not the XML.
Current Claude models prefix these tags with an `antml:` XML namespace prefix (e.g. `antml:function_calls`, `antml:invoke name="…"`, `antml:parameter name="…"`). The API strips all of this and exposes only the JSON `tool_use` block; integrators should target the JSON, not the XML.
---
+9 -9
View File
@@ -110,21 +110,21 @@ format?: "celsius" | "fahrenheit", // default: celsius
## Tool-call format
A function call is an **assistant** message on the **commentary** channel, addressed to the tool via recipient `to=functions.<name>`, with content-type `<|constrain|>json` and the JSON arguments as the body, terminated by the `<|call|>` stop token.
A function call is an **assistant** message on the **commentary** channel, addressed to the tool via recipient `to=functions.<name>`, with the JSON arguments as the body, terminated by the `<|call|>` stop token.
The recipient may appear in the *role section* or the *channel section* of the header — both are valid Harmony and the parser accepts either. The model commonly emits it in the channel section:
The recipient may appear in the *role section* or the *channel section* of the header — both are valid Harmony and the parser accepts either. The model commonly emits it in the channel section. The pi renderer omits the optional content-type marker:
```text
<|start|>assistant<|channel|>commentary to=functions.get_current_weather <|constrain|>json<|message|>{"location":"San Francisco, CA"}<|call|>
<|start|>assistant<|channel|>commentary to=functions.get_current_weather<|message|>{"location":"San Francisco, CA"}<|call|>
```
The `openai-harmony` renderer, when re-serializing a stored call, places the recipient in the role section instead (note the `<|constrain|>` is preceded by a space in both forms):
Some Harmony serializers include an explicit JSON content type and place the recipient in the role section instead:
```text
<|start|>assistant to=functions.get_current_weather<|channel|>commentary <|constrain|>json<|message|>{"location":"San Francisco, CA"}<|call|>
```
The arguments body is a raw JSON object. The `<|constrain|>json` content-type signals JSON (and is the hook for constrained/grammar-based decoding); the `<|constrain|>` token is optional, and the content-type may also be a bare word such as `code` (seen with built-in tools). Built-in tools differ only in channel and recipient: they typically render on `analysis`, with recipient `browser.search` / `browser.open` / `browser.find` or always `python`.
The arguments body is a raw JSON object. The optional `<|constrain|>json` content-type signals JSON (and is the hook for constrained/grammar-based decoding); the content-type may also be a bare word such as `code` (seen with built-in tools). Built-in tools differ only in channel and recipient: they typically render on `analysis`, with recipient `browser.search` / `browser.open` / `browser.find` or always `python`.
## Multiple / parallel tool calls
@@ -136,7 +136,7 @@ Harmony has no special "parallel" wrapper. Multiple calls are just multiple cons
2. Generate a JavaScript for the Node.js server
3. Start the server
---
Will start executing the plan step by step<|end|><|start|>assistant<|channel|>commentary to=functions.generate_file<|constrain|>json<|message|>{"template": "basic_html", "path": "index.html"}<|call|>
Will start executing the plan step by step<|end|><|start|>assistant<|channel|>commentary to=functions.generate_file<|message|>{"template": "basic_html", "path": "index.html"}<|call|>
```
## Tool-result format
@@ -178,7 +178,7 @@ location: string,
format?: "celsius" | "fahrenheit", // default: celsius
}) => any;
} // namespace functions<|end|><|start|>user<|message|>What is the weather like in SF?<|end|><|start|>assistant<|channel|>analysis<|message|>User wants the weather in San Francisco. Use get_current_weather.<|end|><|start|>assistant<|channel|>commentary to=functions.get_current_weather <|constrain|>json<|message|>{"location":"San Francisco, CA"}<|call|><|start|>functions.get_current_weather to=assistant<|channel|>commentary<|message|>{"sunny": true, "temperature": 20}<|end|><|start|>assistant<|channel|>final<|message|>It's sunny and about 20°C in San Francisco right now.<|return|>
} // namespace functions<|end|><|start|>user<|message|>What is the weather like in SF?<|end|><|start|>assistant<|channel|>analysis<|message|>User wants the weather in San Francisco. Use get_current_weather.<|end|><|start|>assistant<|channel|>commentary to=functions.get_current_weather<|message|>{"location":"San Francisco, CA"}<|call|><|start|>functions.get_current_weather to=assistant<|channel|>commentary<|message|>{"sunny": true, "temperature": 20}<|end|><|start|>assistant<|channel|>final<|message|>It's sunny and about 20°C in San Francisco right now.<|return|>
```
Turn boundaries:
@@ -203,12 +203,12 @@ When a server (vLLM/SGLang/Ollama) bridges Harmony to Chat Completions JSON:
## Parsing notes & gotchas
- **Two stop tokens.** Always stop on both `<|return|>` and `<|call|>`. Stopping only on `<|return|>` will run past tool calls; stopping only on `<|end|>` is wrong for assistant generation.
- **Recipient position varies.** `to=functions.<name>` may be in the role section (`<|start|>assistant to=...<|channel|>commentary`) or the channel section (`<|channel|>commentary to=... `). A parser must accept both. A space precedes `<|constrain|>` in both renderings.
- **Recipient position varies.** `to=functions.<name>` may be in the role section (`<|start|>assistant to=...<|channel|>commentary`) or the channel section (`<|channel|>commentary to=...`). A parser must accept both.
- **Channel is mandatory** on assistant messages; the system message even reminds the model ("Channel must be included for every message."). Missing-channel output is malformed.
- **Tool author, not `tool`.** The tool-result message's role is the tool's *name* (`functions.get_current_weather`), not the literal string `tool`. Splitting `functions.x` into namespace + function is the parser's job.
- **CoT dropping is conditional.** Drop `analysis` only when the previous assistant turn ended on `final`. Dropping the `analysis` that immediately precedes a `<|call|>` breaks multi-step tool reasoning.
- **`arguments` is a string.** Do not double-encode. The body after `<|message|>` is already serialized JSON; pass it through as the `arguments` string.
- **Content-type variants.** `<|constrain|>json` is typical, but the content-type can be a bare token (`json`, `code`); treat `<|constrain|>` as optional metadata, not a guarantee of valid JSON. Enforce JSON validity with constrained decoding / your own grammar — the prompt format alone does not guarantee schema adherence (same caveat applies to structured-output `# Response Formats`).
- **Content-type variants.** `<|constrain|>json` is optional. If present, it is metadata, not a guarantee of valid JSON. Enforce JSON validity with constrained decoding / your own grammar — the prompt format alone does not guarantee schema adherence (same caveat applies to structured-output `# Response Formats`).
- **Streaming.** Use a stateful parser (the library ships `StreamableParser`) so partial UTF-8 and the header/channel/recipient/content-type fields are reconstructed incrementally; a naive substring scan mishandles multi-byte splits and the optional header fields. `parse_messages_from_completion_tokens` takes `strict=True|False` — `strict=False` tolerates some malformed headers. Do not pass the trailing stop token into the parser.
- **Encoding.** Use `o200k_harmony` (the `o200k_base` ranks plus the Harmony specials above). Treat the `<|...|>` tokens as atomic special tokens during both encode and decode; encoding them as ordinary text yields different ranks and corrupts the stream.