feat(coding-agent/web): introduced structured web-search query parsing module

- Implemented a structured web-search query parsing module supporting directives, tokenization, date parsing, and syntax serialization.
- Updated search providers to map query directives and date bounds to native provider parameters and filters.
- Added lenient result constraint post-filtering and configuration settings for enhanced engine routing.
- Added comprehensive unit and integration tests covering query parsing, constraint filtering, and provider-specific request mapping.
This commit is contained in:
can1357
2026-07-24 16:03:28 +02:00
parent db937bd149
commit 5acdefc7b3
46 changed files with 2758 additions and 61 deletions
+5
View File
@@ -7,15 +7,20 @@
- Added in-process moreutils-style shell builtins to the bash tool's embedded shell: `ts` (timestamp lines; `-i`/`-s` elapsed modes, `-m` monotonic clock, `-r` relative rewriting of RFC3339/syslog timestamps, `%.S`/`%.s`/`%.T` subsecond extensions), `sponge` (soak stdin fully before atomically writing the target, so `foo file | ... | sponge file` works; `-a` appends), `ifne` (run a command only when stdin is non-empty; `-n` inverts and passes non-empty stdin through), `isutf8` (streaming UTF-8 validation with line/char/byte diagnostics; `-q`, `-l`, `-i`), `combine` (boolean `and`/`not`/`or`/`xor` on the lines of two files, `-` for stdin), and `errno` (errno name/number/description lookup with `-l` list and `-s` search; unix only). Like the uutils-backed builtins, they run in-process against the command's own stdio, resolve paths against the shell working directory, honor cancellation, and are disabled by `PI_DISABLE_UUTILS_BUILTINS`.
- The bash tool prompt now lists the available shell builtins (`mkdir` through `jq`, `rm`/`mv`/`ln`, and the moreutils set) so the model relies on them without existence checks; the line is dropped when `PI_DISABLE_UUTILS_BUILTINS` disables the builtins and omits unix-only `errno` on Windows.
- Sessions using a broker-backed auth store now report each completed request's token usage and cost to the auth broker (batched, 10s cadence) so the broker can track actual token burn per client install
- Added a per-spawn `effort` parameter to the `task` tool (`"lo"` | `"med"` | `"hi"`): each selector maps onto the resolved model's supported thinking range (lowest, middle, and highest level — whatever the model tops out at, e.g. `high`, `xhigh`, or `max`) and overrides the agent's default selector, including `auto`. Omitting `effort` keeps the existing automatic per-prompt thinking classification.
- Added `searxng.engines` setting for the SearXNG web search provider: a comma-separated list of engine names or bang shortcuts (e.g. `ddg, br, startpage`) sent as the API's `engines=` parameter. Shortcuts are resolved to canonical engine names via the instance's `/config` endpoint (cached per endpoint; entries pass through verbatim if `/config` is unreachable). Bang syntax in queries (`!ddg foo`) continues to pass through to the instance, and external bangs (`!!g`) are now stripped client-side since SearXNG answers them with an HTTP redirect even for JSON requests.
- Web search queries now understand Google-style directives on every provider: `site:`/`-site:` (plus `domain:`/`host:` aliases), `after:`/`before:`/`since:`/`until:` date bounds, `inurl:`/`intitle:`/`intext:`/`allin*:`, `filetype:`/`ext:`, `lang:`, quoted phrases (including smart quotes), `+term`, `-`/`NOT` exclusions, and `OR`/`|` groups. A shared parser (`web/search/query.ts`) structures the query once per request; each provider maps constraints onto native API filters where the upstream supports them (Perplexity domain/date/language filters on both the API-key and ask paths, Tavily `include_domains`/`exclude_domains` + `start_date`/`end_date`, Exa domain lists + published-date bounds on API and MCP paths, Anthropic `allowed_domains`/`blocked_domains`, xAI `filters.allowed_domains`/`excluded_domains`, Parallel `source_policy`, Brave absolute `freshness` ranges, Firecrawl `tbs=cdr` date ranges, Jina `X-Site`, SearXNG `language`) and otherwise re-emits only the operator syntax its engine parses (full Google syntax for Gemini grounding, OpenAI, Kagi, and the credential-free scrapers — with scraper-hostile path-`site:`/`inurl:` operators demoted to plain keywords; conservative subsets for DuckDuckGo, Mojeek, Kimi, Z.AI, TinyFish, and Synthetic). The pipeline then applies a lenient post-filter to every response: constraints the engine ignored are enforced on the returned sources, and any constraint dimension that would eliminate every result is relaxed and reported to the model (`Note: no results matched \`site:...\`; the constraint was relaxed`) instead of returning nothing. Directive-free queries are passed through byte-identical everywhere.
### Changed
- Large pastes saved via the large-paste menu now insert `local://paste-N.md` references (previously `local://attachment-N`), so the saved paste carries a markdown extension and a clearer name.
- Raw SSE debug capture now trims over-budget events smartly instead of chopping off the tail: tool definitions inside `data:` payloads are compacted first (name kept, schema/description elided — often enough to keep the whole payload as valid JSON), and anything still over the 64k cap keeps its head and tail with a `: omp-debug-elided chars=N` comment marking the removed middle, so trailing fields like `usage` stay visible.
- The `web_search` tool prompt now tells the model to never search for content that is programmatically accessible or has a known URL (GitHub, known arXiv papers, Wikipedia pages, official docs) and to `read` the URL directly instead.
### Fixed
- Fixed `todo` calls that omit `op` hard-failing validation ("op must be operation to apply (was missing)"): the tool now validates leniently and infers the op for unambiguous payloads (`list` → `init`, `phase`+`items` → `append`, bare `items` on an empty list → `init`); `op` stays required in the schema, and ambiguous op-less calls surface the schema error as a retryable tool error.
- Fixed credential-free web search engines (SearXNG, DuckDuckGo, Google, Startpage, Ecosia, Mojeek, and the Public Web fan-out) returning zero results for queries with `site:` paths (e.g. `site:github.com/owner/repo`) or `inurl:` operators: scraper engines only match `site:` against a bare domain and DuckDuckGo ignores `inurl:` entirely, so such queries silently emptied the result set and fell through to the next provider in the chain. A shared `formatScraperQuery` formatter now structurally demotes path-carrying `site:` and all `inurl:` values to plain search terms (covering OR-grouped and quoted directives) while preserving bare-domain `site:` filters, negated operators, and each engine's supported syntax; the pipeline post-filter still enforces the demoted constraints on returned sources.
- Fixed `ast_edit` previews reading like applied edits to the model: the `⟨proposed⟩` badge was TUI-only, so the model-visible result (hashline header + `-`/`+` rows, identical to applied edit output) carried no staged-proposal signal. The preview result now leads with a "Staged as a proposal — files NOT modified yet" notice naming `xd://resolve`/`xd://reject`, the injected resolve reminder names the source tool, and the `ast_edit` tool prompt documents the two-phase flow.
- Fixed the `hub` launch `ps`/`list` response burying the active process behind every exited one and growing without bound in long-lived projects: the broker now lists non-terminal daemons first (oldest to newest) and caps exited/failed history at the 10 most recently exited, so the active launch is immediately visible and the response stays bounded. Broker recovery also preserves each already-terminal daemon's real exit time instead of overwriting it with the restart timestamp, so the history cap keeps the genuinely most-recently-exited processes after an idle-broker restart ([#6517](https://github.com/can1357/oh-my-pi/issues/6517)).
@@ -130,8 +130,15 @@ ${chalk.bold("Options:")}
--compact Render condensed output
-h, --help Show this help
${chalk.bold("Query directives:")}
site:/-site: after:/before: (YYYY-MM-DD) inurl: intitle: filetype:
"exact phrase" -term OR
Mapped to native provider filters where available, otherwise applied as a
lenient post-filter (a constraint matching nothing is relaxed, not fatal).
${chalk.bold("Examples:")}
${APP_NAME} q --provider=exa "what's the color of the sky"
${APP_NAME} q --provider=brave --recency=week "latest TypeScript 5.7 changes"
${APP_NAME} q 'transformer scaling site:arxiv.org after:2024 -site:reddit.com'
`);
}
@@ -5266,6 +5266,11 @@ export const SETTINGS_SCHEMA = {
default: undefined,
},
"searxng.engines": {
type: "string",
default: undefined,
},
"searxng.language": {
type: "string",
default: undefined,
@@ -3,4 +3,6 @@ Searches the web for up-to-date information beyond knowledge cutoff.
<instruction>
- You SHOULD prefer primary sources (papers, official docs) and corroborate key claims with multiple sources
- You MUST include links for cited sources in the final response
- NEVER use for content that is programmatically accessible or whose URL you already know (GitHub repos/issues, a known arXiv paper, a Wikipedia page, official docs) — `read` the URL directly instead
- `query` supports Google-style directives on every provider: `site:`/`-site:`, `after:`/`before:` (`YYYY-MM-DD`), `inurl:`, `intitle:`, `filetype:`, `"exact phrase"`, `-term`, `OR`. Constraints map to native provider filters where available; otherwise results are filtered leniently — a constraint matching nothing is relaxed and reported instead of returning zero results.
</instruction>
+28 -5
View File
@@ -27,6 +27,7 @@ import {
type SearchProvider,
type SearchProviderCandidate,
} from "./provider";
import { applyQueryConstraints, parseSearchQuery } from "./query";
import { renderSearchCall, renderSearchResult, type SearchRenderDetails } from "./render";
import type { SearchProviderId, SearchResponse } from "./types";
import { SearchProviderError } from "./types";
@@ -57,9 +58,12 @@ function formatCount(label: string, count: number): string {
return `${count} ${label}${count === 1 ? "" : "s"}`;
}
/** Format response for LLM consumption */
function formatForLLM(response: SearchResponse): string {
/** Format response for LLM consumption. `notes` lead the output (e.g. relaxed-constraint warnings). */
function formatForLLM(response: SearchResponse, notes: readonly string[] = []): string {
const parts: string[] = [];
for (const note of notes) {
parts.push(`Note: ${note}`);
}
if (response.answer) {
parts.push(response.answer);
@@ -141,6 +145,8 @@ async function executeSearch(
candidates = resolveProviderCandidates();
}
const parsedQuery = parseSearchQuery(params.query);
// Invariant across providers; read once and tolerate an uninitialized
// Settings singleton (e.g. `omp q ...` CLI path, unit tests) so the
// provider-fallback loop never aborts before any provider runs.
@@ -182,6 +188,7 @@ async function executeSearch(
const response = await provider.search({
query: params.query,
parsedQuery,
limit: params.limit,
recency: params.recency,
systemPrompt: webSearchSystemPrompt,
@@ -196,15 +203,31 @@ async function executeSearch(
geminiModel,
});
if (!hasRenderableSearchContent(response)) {
// Lenient constraint pass over whatever the provider returned: enforce
// site:/inurl:/intitle:/filetype:/date directives the provider could
// not (or only partially) honor natively, relaxing any dimension that
// would wipe out every result. Citations/answer text stay untouched.
let finalResponse = response;
const constraintNotes: string[] = [];
if (parsedQuery.hasConstraints && response.sources.length > 0) {
const filtered = applyQueryConstraints(response.sources, parsedQuery);
if (filtered.sources.length !== response.sources.length) {
finalResponse = { ...response, sources: filtered.sources };
}
for (const label of filtered.dropped) {
constraintNotes.push(`no results matched \`${label}\`; the constraint was relaxed`);
}
}
if (!hasRenderableSearchContent(finalResponse)) {
throw new SearchProviderError(provider.id, `${provider.label} returned no renderable search content.`, 204);
}
const text = formatForLLM(response);
const text = formatForLLM(finalResponse, constraintNotes);
return {
content: [{ type: "text" as const, text }],
details: { response },
details: { response: finalResponse },
};
} catch (error) {
// Surface user-initiated cancellation immediately so the session sees
@@ -28,6 +28,7 @@ import type {
SearchSource,
} from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery, type QuerySyntax, type StructuredQuery } from "../query";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
import { classifyProviderHttpError, withHardTimeout } from "./utils";
@@ -36,6 +37,59 @@ const DEFAULT_MODEL = "claude-haiku-4-5";
const DEFAULT_MAX_TOKENS = 4096;
const WEB_SEARCH_TOOL_NAME = "web_search";
const WEB_SEARCH_TOOL_TYPE = "web_search_20250305";
/**
* Claude's search backend understands common Google-style operators, so most
* directives are re-emitted as query text. `site:` is intentionally absent:
* site includes/excludes map onto the web_search tool's native
* `allowed_domains`/`blocked_domains` parameters instead.
*/
const ANTHROPIC_QUERY_SYNTAX: QuerySyntax = {
phrases: true,
negation: true,
or: true,
inUrl: true,
inTitle: true,
filetype: true,
dateRange: true,
};
/** Upstream request shape derived from the parsed query. */
interface AnthropicQueryPlan {
query: string;
allowedDomains?: string[];
blockedDomains?: string[];
}
/**
* Map parsed directives onto the request: `site:` includes become
* `allowed_domains`, `-site:` exclusions become `blocked_domains` (the two are
* mutually exclusive on the API, so exclusions are only sent when there are no
* includes), and remaining directives are re-emitted as query syntax.
* Directive-free queries pass through byte-identical. Anthropic domain
* filters take bare hosts (subdomains included automatically); any path part
* of a `site:` value is enforced by the central constraint filter.
*/
function planQuery(rawQuery: string, parsed: StructuredQuery): AnthropicQueryPlan {
if (!parsed.hasDirectives) return { query: rawQuery };
const hosts = (sites: readonly string[]) => {
const unique = new Set<string>();
for (const site of sites) {
const slash = site.indexOf("/");
const host = slash === -1 ? site : site.slice(0, slash);
if (host.length > 0) unique.add(host);
}
return [...unique];
};
const allowed = hosts(parsed.sites);
const blocked = allowed.length === 0 ? hosts(parsed.excludedSites) : [];
return {
query: formatQuery(parsed, ANTHROPIC_QUERY_SYNTAX),
allowedDomains: allowed.length > 0 ? allowed : undefined,
blockedDomains: blocked.length > 0 ? blocked : undefined,
};
}
export interface AnthropicSearchParams {
query: string;
system_prompt?: string;
@@ -82,7 +136,7 @@ function buildSystemBlocks(
* Calls the Anthropic API with web search tool enabled.
* @param auth - Authentication configuration (API key or OAuth)
* @param model - Model identifier to use
* @param query - Search query from the user
* @param plan - Query text plus native domain filters derived from parsed directives
* @param metadataUserId - Optional Anthropic Messages metadata.user_id (already shaped for OAuth)
* @param systemPrompt - Optional system prompt for guiding response style
* @returns Raw API response from Anthropic
@@ -91,7 +145,7 @@ function buildSystemBlocks(
async function callSearch(
auth: AnthropicAuthConfig,
model: string,
query: string,
plan: AnthropicQueryPlan,
metadataUserId?: string,
systemPrompt?: string,
maxTokens?: number,
@@ -107,11 +161,13 @@ async function callSearch(
const body: Record<string, unknown> = {
model,
max_tokens: maxTokens ?? DEFAULT_MAX_TOKENS,
messages: [{ role: "user", content: query }],
messages: [{ role: "user", content: plan.query }],
tools: [
{
type: WEB_SEARCH_TOOL_TYPE,
name: WEB_SEARCH_TOOL_NAME,
...(plan.allowedDomains ? { allowed_domains: plan.allowedDomains } : {}),
...(plan.blockedDomains ? { blocked_domains: plan.blockedDomains } : {}),
},
],
};
@@ -283,6 +339,8 @@ export async function searchAnthropic(
const callerSessionId = "authStorage" in params ? params.sessionId : undefined;
const accountId =
"authStorage" in params ? params.authStorage.getOAuthAccountId("anthropic", params.sessionId) : undefined;
const parsed = ("parsedQuery" in params ? params.parsedQuery : undefined) ?? parseSearchQuery(params.query);
const plan = planQuery(params.query, parsed);
const response = await withAuth(
keyOrResolver,
key => {
@@ -302,7 +360,7 @@ export async function searchAnthropic(
return callSearch(
auth,
model,
params.query,
plan,
metadataUserId,
systemPrompt,
maxTokens,
@@ -1,5 +1,6 @@
import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai";
import type { ModelRegistry } from "../../../config/model-registry";
import type { StructuredQuery } from "../query";
import type { SearchProviderId, SearchResponse } from "../types";
/**
@@ -14,6 +15,21 @@ import type { SearchProviderId, SearchResponse } from "../types";
*/
export interface SearchParams {
query: string;
/**
* Structured view of `query`, parsed once by the search pipeline:
* Google-style directives (`site:`, `before:`/`after:`, `inurl:`,
* `intitle:`, `filetype:`, quoted phrases, `OR` groups, `-exclusions`)
* extracted into fields.
*
* Providers SHOULD map constraints onto native API parameters
* (domain/date filters) or engine query syntax (`formatQuery`) where the
* upstream supports them, and lean lenient otherwise: the pipeline
* post-filters every response with `applyQueryConstraints`, which
* relaxes any constraint that would eliminate all results — so a
* best-effort search always beats an empty one. When absent (direct
* provider calls), parse with `parseSearchQuery(params.query)`.
*/
parsedQuery?: StructuredQuery;
limit?: number;
/**
* Temporal filter narrowing results to the specified time window.
@@ -7,6 +7,8 @@
import { type AuthStorage, type FetchImpl, getEnvApiKey } from "@oh-my-pi/pi-ai";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import type { QuerySyntax, StructuredQuery } from "../query";
import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery } from "../query";
import { clampNumResults, dateToAgeSeconds } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -23,10 +25,32 @@ const RECENCY_MAP: Record<"day" | "week" | "month" | "year", "pd" | "pw" | "pm"
year: "py",
};
/**
* Brave parses the classic operator set inline (site:, quotes, -, OR…) but
* date bounds map onto the native `freshness` param, so `before:`/`after:`
* tokens are stripped from the rebuilt query string.
*/
const BRAVE_QUERY_SYNTAX: QuerySyntax = { ...GOOGLE_QUERY_SYNTAX, dateRange: false };
/**
* Freshness param: explicit `after:`/`before:` bounds win over the
* recency-derived period, rendered as Brave's absolute range
* `YYYY-MM-DDtoYYYY-MM-DD` with sensible open ends.
*/
function braveFreshness(parsed: StructuredQuery, recency?: keyof typeof RECENCY_MAP): string | undefined {
if (parsed.after || parsed.before) {
const start = parsed.after ?? "1970-01-01";
const end = parsed.before ?? new Date().toISOString().slice(0, 10);
return `${start}to${end}`;
}
return recency ? RECENCY_MAP[recency] : undefined;
}
export interface BraveSearchParams {
query: string;
num_results?: number;
recency?: "day" | "week" | "month" | "year";
parsedQuery?: StructuredQuery;
signal?: AbortSignal;
fetch?: FetchImpl;
}
@@ -73,12 +97,14 @@ async function callBraveSearch(
params: BraveSearchParams,
): Promise<{ response: BraveSearchResponse; requestId?: string }> {
const numResults = clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS);
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const url = new URL(BRAVE_SEARCH_URL);
url.searchParams.set("q", params.query);
url.searchParams.set("q", parsed.hasDirectives ? formatQuery(parsed, BRAVE_QUERY_SYNTAX) : params.query);
url.searchParams.set("count", String(numResults));
url.searchParams.set("extra_snippets", "true");
if (params.recency) {
url.searchParams.set("freshness", RECENCY_MAP[params.recency]);
const freshness = braveFreshness(parsed, params.recency);
if (freshness) {
url.searchParams.set("freshness", freshness);
}
const fetchImpl = params.fetch ?? fetch;
@@ -145,6 +171,7 @@ export class BraveProvider extends SearchProvider {
query: params.query,
num_results: params.numSearchResults ?? params.limit,
recency: params.recency,
parsedQuery: params.parsedQuery,
signal: params.signal,
fetch: params.fetch,
});
@@ -31,6 +31,7 @@ import packageJson from "../../../../package.json" with { type: "json" };
import type { ModelRegistry } from "../../../config/model-registry";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery } from "../query";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
import { classifyProviderHttpError, withHardTimeout } from "./utils";
@@ -572,6 +573,7 @@ async function callCodexSearch(
async function runCodexSearchCandidates(options: {
auth: { accessToken: string; accountId?: string };
params: SearchParams;
query: string;
modelCandidates: CodexModelCandidate[];
modelWasConfigured: boolean;
transport: CodexSearchTransport;
@@ -582,7 +584,7 @@ async function runCodexSearchCandidates(options: {
if (!candidate) continue;
try {
return await callCodexSearch(options.auth, options.params.query, {
return await callCodexSearch(options.auth, options.query, {
signal: options.params.signal,
systemPrompt: options.params.systemPrompt,
searchContextSize: "high",
@@ -622,6 +624,15 @@ export async function searchCodex(params: SearchParams): Promise<SearchResponse>
throw new SearchProviderError("codex", "No Codex web search model is configured.");
}
const transport = resolveCodexSearchTransport(params.modelRegistry, firstCandidate.modelId);
// The ChatGPT-backend Codex endpoint speaks the undocumented codex-rs
// request shape (responses-lite moves tools into an `additional_tools`
// developer item), so the documented `web_search.filters.allowed_domains`
// parameter cannot be assumed to survive it. Instead, re-emit directive
// queries with the full Google-style operator syntax — the backing index
// parses the classic operator set — and leave directive-free queries
// byte-identical.
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const query = parsed.hasDirectives ? formatQuery(parsed, GOOGLE_QUERY_SYNTAX) : params.query;
let result: CodexSearchResult;
if (transport.customEndpoint) {
@@ -652,6 +663,7 @@ export async function searchCodex(params: SearchParams): Promise<SearchResponse>
runCodexSearchCandidates({
auth: { accessToken },
params,
query,
modelCandidates,
modelWasConfigured: configuredModel !== undefined,
transport,
@@ -682,6 +694,7 @@ export async function searchCodex(params: SearchParams): Promise<SearchResponse>
return runCodexSearchCandidates({
auth: { accessToken: access.accessToken, accountId },
params,
query,
modelCandidates,
modelWasConfigured: configuredModel !== undefined,
transport,
@@ -1,6 +1,8 @@
import type { AuthStorage } from "@oh-my-pi/pi-ai";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import type { QuerySyntax } from "../query";
import { formatScraperQuery } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -118,8 +120,28 @@ function isAnomalyResponse(html: string): boolean {
return html.includes("anomaly-modal") || html.includes("anomaly.js");
}
/**
* Query syntax the DDG HTML frontend parses: quotes, `-`, OR, site:,
* filetype:, intitle:, inurl:, intext:. Date bounds (`before:`/`after:`) are
* deliberately off — DDG does not parse them, so they are stripped from the
* query and enforced by the pipeline's lenient post-filter instead.
*/
const DDG_QUERY_SYNTAX: QuerySyntax = {
phrases: true,
negation: true,
or: true,
site: true,
inUrl: true,
inTitle: true,
inText: true,
filetype: true,
};
async function callDuckDuckGoHtml(params: SearchParams): Promise<string> {
const form = new URLSearchParams({ q: params.query, kl: "us-en" });
const form = new URLSearchParams({
q: formatScraperQuery(params.query, params.parsedQuery, DDG_QUERY_SYNTAX),
kl: "us-en",
});
const df = params.recency ? RECENCY_TO_DDG_DF[params.recency] : undefined;
if (df) form.set("df", df);
// Add b: "" parameter as specified in the browser fetch template to match real browser form submission
@@ -2,6 +2,7 @@ import type { AuthStorage } from "@oh-my-pi/pi-ai";
import { parseHTML } from "linkedom";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatScraperQuery } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -100,7 +101,10 @@ function isBlockedPage(page: LoadedHtmlPage): boolean {
async function callEcosiaHtml(params: SearchParams): Promise<string> {
const signal = withHardTimeout(params.signal);
const url = new URL(ECOSIA_SEARCH_URL);
url.searchParams.set("q", params.query);
// Ecosia serves Google-backed results, so classic operators pass through
// inline; canonicalize aliases (domain: -> site:, since: -> after:) and
// demote scraper-hostile operators via the shared scraper formatter.
url.searchParams.set("q", formatScraperQuery(params.query, params.parsedQuery));
let page: LoadedHtmlPage;
try {
@@ -12,6 +12,7 @@ import { findApiKey, isSearchResponse } from "../../../exa/mcp-client";
import { parseSSE } from "../../../mcp/json-rpc";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery, type StructuredQuery } from "../query";
import { dateToAgeSeconds } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -472,8 +473,9 @@ export class ExaProvider extends SearchProvider {
}
search(params: SearchParams): Promise<SearchResponse> {
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
return searchExa({
query: params.query,
...directiveParams(parsed),
num_results: params.numSearchResults ?? params.limit,
signal: params.signal,
authStorage: params.authStorage,
@@ -482,3 +484,28 @@ export class ExaProvider extends SearchProvider {
});
}
}
/**
* Map parsed query directives onto Exa's native request parameters:
* `site:` → includeDomains, `-site:` → excludeDomains (bare hosts; path parts
* are enforced by the central constraint filter), `after:`/`before:` →
* start/endPublishedDate (ISO 8601). Exa's neural search prefers natural
* language, so the query itself is re-emitted with quoted phrases only.
* Directive-free queries pass through byte-identical.
*/
function directiveParams(
parsed: StructuredQuery,
): Pick<
ExaSearchParams,
"query" | "include_domains" | "exclude_domains" | "start_published_date" | "end_published_date"
> {
if (!parsed.hasDirectives) return { query: parsed.raw };
const hosts = (sites: readonly string[]) => [...new Set(sites.map(site => site.split("/", 1)[0]))];
return {
query: formatQuery(parsed, { phrases: true }),
include_domains: parsed.sites.length ? hosts(parsed.sites) : undefined,
exclude_domains: parsed.excludedSites.length ? hosts(parsed.excludedSites) : undefined,
start_published_date: parsed.after,
end_published_date: parsed.before,
};
}
@@ -14,6 +14,7 @@ import {
} from "@oh-my-pi/pi-ai";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery, type StructuredQuery } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -34,6 +35,8 @@ export interface FirecrawlSearchParams {
query: string;
num_results?: number;
recency?: SearchParams["recency"];
/** Explicit `tbs` (custom date range); takes precedence over `recency`. */
tbs?: string;
signal?: AbortSignal;
fetch?: FetchImpl;
}
@@ -67,8 +70,9 @@ function buildRequestBody(params: FirecrawlSearchParams): Record<string, unknown
limit: clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS),
sources: [{ type: "web" }],
};
if (params.recency) {
body.tbs = RECENCY_TBS[params.recency];
const tbs = params.tbs ?? (params.recency ? RECENCY_TBS[params.recency] : undefined);
if (tbs) {
body.tbs = tbs;
}
return body;
}
@@ -104,12 +108,42 @@ async function callFirecrawlSearch(
return (await response.json()) as FirecrawlSearchResponse;
}
/** ISO `YYYY-MM-DD` to Google `MM/DD/YYYY` for `tbs=cdr` custom date ranges. */
function toGoogleDate(iso: string): string {
const [year, month, day] = iso.split("-");
return `${month}/${day}/${year}`;
}
/**
* Map explicit `before:`/`after:` bounds to a Firecrawl `tbs` custom date
* range (`cdr:1,cd_min:MM/DD/YYYY,cd_max:MM/DD/YYYY`), or undefined when the
* query carries no absolute date bounds.
*/
function buildDateTbs(parsed: StructuredQuery): string | undefined {
if (!parsed.after && !parsed.before) return undefined;
const parts = ["cdr:1"];
if (parsed.after) parts.push(`cd_min:${toGoogleDate(parsed.after)}`);
if (parsed.before) parts.push(`cd_max:${toGoogleDate(parsed.before)}`);
return parts.join(",");
}
/** Execute Firecrawl web search. */
export async function searchFirecrawl(params: SearchParams): Promise<SearchResponse> {
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
let query = params.query;
let tbs: string | undefined;
if (parsed.hasDirectives) {
// Firecrawl search is SERP-backed: the query supports Google operators
// (site:, inurl:, intitle:, quotes, -, OR). Absolute date bounds move to
// the native tbs param and are stripped from the query string.
tbs = buildDateTbs(parsed);
query = formatQuery(parsed, tbs ? { ...GOOGLE_QUERY_SYNTAX, dateRange: false } : GOOGLE_QUERY_SYNTAX);
}
const firecrawlParams: FirecrawlSearchParams = {
query: params.query,
query,
num_results: params.numSearchResults ?? params.limit,
recency: params.recency,
tbs,
signal: params.signal,
fetch: params.fetch,
};
@@ -18,6 +18,7 @@ import { fetchWithRetry } from "@oh-my-pi/pi-utils";
import type { SearchCitation, SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery, type StructuredQuery } from "../query";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
import { classifyProviderHttpError, withHardTimeout } from "./utils";
@@ -51,6 +52,8 @@ interface GeminiToolParams {
export interface GeminiSearchParams extends GeminiToolParams {
query: string;
/** Pre-parsed structured query; falls back to parsing `query` when omitted. */
parsedQuery?: StructuredQuery;
system_prompt?: string;
num_results?: number;
/** Maximum output tokens. */
@@ -508,6 +511,12 @@ async function callGeminiDeveloperSearch(
*/
export async function searchGemini(params: GeminiSearchParams): Promise<SearchResponse> {
const selectedModel = resolveGeminiSearchModel(params.geminiModel);
// Gemini's googleSearch grounding forwards the query to Google Search, which
// understands the classic operator set natively. Normalize directive aliases
// (domain: → site:, since: → after:, …) to canonical Google forms; leave
// directive-free queries byte-identical.
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const searchQuery = parsed.hasDirectives ? formatQuery(parsed, GOOGLE_QUERY_SYNTAX) : params.query;
const seed = await findGeminiAuth(params.authStorage, params.sessionId, params.signal);
let result: GeminiSearchResult;
@@ -528,7 +537,7 @@ export async function searchGemini(params: GeminiSearchParams): Promise<SearchRe
isAntigravity,
},
selectedModel,
params.query,
searchQuery,
params.system_prompt,
params.max_output_tokens,
params.temperature,
@@ -555,7 +564,7 @@ export async function searchGemini(params: GeminiSearchParams): Promise<SearchRe
result = await callGeminiDeveloperSearch(
apiKey,
selectedModel,
params.query,
searchQuery,
params.system_prompt,
params.max_output_tokens,
params.temperature,
@@ -601,6 +610,7 @@ export class GeminiProvider extends SearchProvider {
search(params: SearchParams): Promise<SearchResponse> {
return searchGemini({
query: params.query,
parsedQuery: params.parsedQuery,
system_prompt: params.systemPrompt,
num_results: params.numSearchResults ?? params.limit,
max_output_tokens: params.maxOutputTokens,
@@ -2,6 +2,7 @@ import type { AuthStorage } from "@oh-my-pi/pi-ai";
import { parseHTML } from "linkedom";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatScraperQuery } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -91,7 +92,7 @@ function parseHtmlResults(html: string): ParsedResult[] {
function buildSearchUrl(params: SearchParams, numResults: number): string {
const url = new URL(GOOGLE_SEARCH_URL);
url.searchParams.set("q", params.query);
url.searchParams.set("q", formatScraperQuery(params.query, params.parsedQuery));
url.searchParams.set("num", String(numResults));
url.searchParams.set("hl", "en");
url.searchParams.set("gl", "us");
@@ -8,6 +8,7 @@
import { type AuthStorage, type FetchImpl, getEnvApiKey } from "@oh-my-pi/pi-ai";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery } from "../query";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
import { classifyProviderHttpError, withHardTimeout } from "./utils";
@@ -18,6 +19,8 @@ type SearchParamsWithFetch = SearchParams & { fetch?: FetchImpl };
export interface JinaSearchParams {
query: string;
num_results?: number;
/** Single bare host for Jina's `X-Site` in-site search header. */
site?: string;
signal?: AbortSignal;
fetch?: FetchImpl;
}
@@ -39,15 +42,18 @@ export function findApiKey(): string | null {
async function callJinaSearch(
apiKey: string,
query: string,
site?: string,
signal?: AbortSignal,
fetchImpl: FetchImpl = fetch,
): Promise<JinaSearchResponse> {
const requestUrl = `${JINA_SEARCH_URL}/${encodeURIComponent(query)}`;
const headers: Record<string, string> = {
Accept: "application/json",
Authorization: `Bearer ${apiKey}`,
};
if (site) headers["X-Site"] = site;
const response = await fetchImpl(requestUrl, {
headers: {
Accept: "application/json",
Authorization: `Bearer ${apiKey}`,
},
headers,
signal: withHardTimeout(signal),
});
@@ -69,7 +75,7 @@ export async function searchJina(params: JinaSearchParams): Promise<SearchRespon
throw new Error("JINA_API_KEY not found. Set it in environment or .env file.");
}
const response = await callJinaSearch(apiKey, params.query, params.signal, params.fetch);
const response = await callJinaSearch(apiKey, params.query, params.site, params.signal, params.fetch);
const sources: SearchSource[] = [];
for (const result of response) {
@@ -99,13 +105,30 @@ export class JinaProvider extends SearchProvider {
}
search(params: SearchParamsWithFetch): Promise<SearchResponse> {
const fetchImpl = params.fetch;
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
let query = params.query;
let site: string | undefined;
if (parsed.hasDirectives) {
// Jina's X-Site header takes a single domain; with exactly one
// include site, send its host there and strip site: tokens from
// the query. Multiple sites stay inline (Bing-backed, parses them).
if (parsed.sites.length === 1) site = parsed.sites[0]!.split("/")[0];
query = formatQuery(parsed, {
phrases: true,
negation: true,
site: !site,
inTitle: true,
inUrl: true,
filetype: true,
});
}
return searchJina({
query: params.query,
query,
num_results: params.numSearchResults ?? params.limit,
site,
signal: params.signal,
fetch: fetchImpl,
fetch: params.fetch,
});
}
}
@@ -7,6 +7,8 @@ import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai";
import type { SearchResponse } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { KagiApiError, searchWithKagi } from "../../kagi";
import type { StructuredQuery } from "../query";
import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -22,16 +24,22 @@ export async function searchKagi(params: {
query: string;
num_results?: number;
recency?: SearchParams["recency"];
parsedQuery?: StructuredQuery;
signal?: AbortSignal;
authStorage: AuthStorage;
sessionId?: string;
fetch?: FetchImpl;
}): Promise<SearchResponse> {
const numResults = clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS);
// Kagi's index understands the classic Google operator set: canonicalize
// directives (domain: -> site:, until: -> before:YYYY-MM-DD, ...) and pass
// them through in the query string. Directive-free queries stay untouched.
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const query = parsed.hasDirectives ? formatQuery(parsed, GOOGLE_QUERY_SYNTAX) : params.query;
try {
const result = await searchWithKagi(
params.query,
query,
{
limit: numResults,
recency: params.recency,
@@ -75,6 +83,7 @@ export class KagiProvider extends SearchProvider {
return searchKagi({
query: params.query,
parsedQuery: params.parsedQuery,
num_results: params.numSearchResults ?? params.limit,
recency: params.recency,
signal: params.signal,
@@ -12,6 +12,7 @@ import { $env } from "@oh-my-pi/pi-utils";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery, type QuerySyntax, type StructuredQuery } from "../query";
import { clampNumResults, dateToAgeSeconds } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -25,8 +26,19 @@ const DEFAULT_NUM_RESULTS = 10;
const MAX_NUM_RESULTS = 20;
const DEFAULT_TIMEOUT_SECONDS = 30;
/** Kimi Code search is Bing-flavored: re-emit the operators Bing parses; dates/lang stay with the central filter. */
const KIMI_QUERY_SYNTAX: QuerySyntax = {
phrases: true,
negation: true,
site: true,
inTitle: true,
inUrl: true,
filetype: true,
};
export interface KimiSearchParams {
query: string;
parsedQuery?: StructuredQuery;
num_results?: number;
include_content?: boolean;
signal?: AbortSignal;
@@ -138,12 +150,14 @@ export async function searchKimi(params: KimiSearchParams): Promise<SearchRespon
);
}
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const query = parsed.hasDirectives ? formatQuery(parsed, KIMI_QUERY_SYNTAX) : params.query;
const limit = clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS);
const { response, requestId } = await withAuth(
keyOrResolver,
key =>
callKimiSearch(key, {
query: params.query,
query,
limit,
includeContent: params.include_content ?? false,
signal: params.signal,
@@ -192,6 +206,7 @@ export class KimiProvider extends SearchProvider {
return searchKimi({
query: params.query,
parsedQuery: params.parsedQuery,
num_results: params.numSearchResults ?? params.limit,
signal: params.signal,
authStorage: params.authStorage,
@@ -4,6 +4,7 @@ import { parseHTML } from "linkedom";
import type { Page } from "puppeteer-core";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatScraperQuery, type QuerySyntax } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -81,9 +82,21 @@ function parseHtmlResults(html: string): ParsedResult[] {
return results;
}
/**
* Syntax re-emitted to Mojeek for directive-carrying queries. Mojeek's
* support page (mojeek.com/support/search-operators.html) confirms `site:`,
* and the community docs confirm quoted phrases and `-` exclusions. Mojeek
* also parses `in*:` operators and its own date syntax (`since:`/`before:`
* with YYYYMMDD), but the latter differs from Google's `after:`/`before:`
* ISO form and `since` is already claimed by `recency`, so date bounds and
* `in*` constraints are conservatively left to the pipeline's lenient
* post-filter instead.
*/
const MOJEEK_QUERY_SYNTAX: QuerySyntax = { phrases: true, negation: true, site: true };
function buildSearchUrl(params: SearchParams, numResults: number): string {
const url = new URL(MOJEEK_SEARCH_URL);
url.searchParams.set("q", params.query);
url.searchParams.set("q", formatScraperQuery(params.query, params.parsedQuery, MOJEEK_QUERY_SYNTAX));
url.searchParams.set("t", String(numResults));
url.searchParams.set("arc", "none");
url.searchParams.set("lang", "en");
@@ -9,6 +9,7 @@ import {
parseParallelErrorResponse,
parseParallelSearchPayload,
} from "../../parallel";
import { formatQuery, parseSearchQuery, type StructuredQuery } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -17,6 +18,42 @@ import { classifyProviderHttpError, toSearchSources, withHardTimeout } from "./u
const DEFAULT_NUM_RESULTS = 10;
const MAX_NUM_RESULTS = 40;
/** Query-string caps for Parallel: natural-language objective, no field operators. */
const PARALLEL_QUERY_SYNTAX = { phrases: true, negation: true, or: true } as const;
/** Parallel `source_policy` (beta Search API): bare-host allow/deny lists + freshness floor. */
interface ParallelSourcePolicy {
include_domains?: string[];
exclude_domains?: string[];
after_date?: string;
}
/** Site values may carry paths (`github.com/anthropics`); Parallel takes bare hosts. */
function toHosts(sites: readonly string[]): string[] {
const hosts = new Set<string>();
for (const site of sites) {
const host = site.split("/", 1)[0];
if (host) hosts.add(host);
}
return [...hosts];
}
/**
* Map parsed `site:`/`-site:`/`after:` directives onto Parallel's
* `source_policy`. Per Parallel docs, `exclude_domains` is ignored when
* `include_domains` is set, so exclusions are only sent without an allow
* list (the central lenient filter enforces them regardless).
*/
function toSourcePolicy(parsed: StructuredQuery): ParallelSourcePolicy | undefined {
const policy: ParallelSourcePolicy = {};
const include = toHosts(parsed.sites);
const exclude = toHosts(parsed.excludedSites);
if (include.length) policy.include_domains = include;
else if (exclude.length) policy.exclude_domains = exclude;
if (parsed.after) policy.after_date = parsed.after;
return Object.keys(policy).length ? policy : undefined;
}
async function searchWithAuthStorage(
objective: string,
queries: string[],
@@ -26,6 +63,7 @@ async function searchWithAuthStorage(
},
authStorage: AuthStorage,
sessionId?: string,
sourcePolicy?: ParallelSourcePolicy,
): Promise<ParallelSearchResult> {
const apiKey = await authStorage.getApiKey("parallel", sessionId, { signal: params.signal });
if (!apiKey) {
@@ -57,6 +95,7 @@ async function searchWithAuthStorage(
excerpts: {
max_chars_per_result: 10_000,
},
...(sourcePolicy && { source_policy: sourcePolicy }),
}),
signal: withHardTimeout(params.signal),
});
@@ -78,22 +117,28 @@ export async function searchParallel(
num_results?: number;
signal?: AbortSignal;
fetch?: FetchImpl;
parsedQuery?: StructuredQuery;
},
authStorage: AuthStorage,
sessionId?: string,
): Promise<SearchResponse> {
const numResults = clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS);
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
// Back-compat: without directives the upstream request is byte-identical.
const query = parsed.hasDirectives ? formatQuery(parsed, PARALLEL_QUERY_SYNTAX) : params.query;
const sourcePolicy = parsed.hasDirectives ? toSourcePolicy(parsed) : undefined;
try {
const result = await searchWithAuthStorage(
params.query,
[params.query],
query,
[query],
{
signal: params.signal,
fetch: params.fetch,
},
authStorage,
sessionId,
sourcePolicy,
);
return {
@@ -128,6 +173,7 @@ export class ParallelProvider extends SearchProvider {
num_results: params.numSearchResults ?? params.limit,
signal: params.signal,
fetch: params.fetch,
parsedQuery: params.parsedQuery,
},
params.authStorage,
params.sessionId,
@@ -30,6 +30,7 @@ import type {
SearchSource,
} from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery, type QuerySyntax, type StructuredQuery } from "../query";
import { dateToAgeSeconds } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -46,6 +47,71 @@ const OAUTH_USER_AGENT = "Perplexity/641 CFNetwork/1568 Darwin/25.2.0";
const ANONYMOUS_USER_AGENT =
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/149.0.0.0 Safari/537.36";
/**
* Query-string operators Perplexity's search backend tolerates as text signal.
* `site:`/date/`lang:` directives are excluded: they map onto native request
* fields (`search_domain_filter`, `search_*_date_filter`,
* `search_language_filter`) and must be stripped from the query so the engine
* is not double-constrained.
*/
const PERPLEXITY_QUERY_SYNTAX: QuerySyntax = {
phrases: true,
negation: true,
or: true,
inUrl: true,
inTitle: true,
filetype: true,
};
/** Native Perplexity search filters derived from parsed query directives. */
interface PerplexityNativeFilters {
/** Query rebuilt without natively-mapped directives. */
query: string;
/** `search_domain_filter`: allow entries as bare hosts, deny entries as `-host`. */
domainFilter?: string[];
/** `search_after_date_filter`, `%m/%d/%Y`. */
afterDate?: string;
/** `search_before_date_filter`, `%m/%d/%Y`. */
beforeDate?: string;
/** `search_language_filter`: ISO 639-1 two-letter codes. */
languageFilter?: string[];
}
/**
* Bare host of a `site:` value (`github.com/anthropics` → `github.com`);
* Perplexity's domain filter takes hosts only, the path part is enforced by
* the central lenient post-filter.
*/
function siteHost(site: string): string {
const slash = site.indexOf("/");
return slash === -1 ? site : site.slice(0, slash);
}
/** ISO `YYYY-MM-DD` → Perplexity's documented `%m/%d/%Y` date-filter format (e.g. `3/1/2025`). */
function toPerplexityDate(iso: string): string {
const [year, month, day] = iso.split("-");
return `${Number(month)}/${Number(day)}/${year}`;
}
/** Map parsed query directives onto native Perplexity search filters. */
function buildNativeFilters(parsed: StructuredQuery, rawQuery: string): PerplexityNativeFilters {
if (!parsed.hasDirectives) return { query: rawQuery };
// Allow + deny share one array; the API caps it at 20 entries.
const domains = [
...new Set([...parsed.sites.map(siteHost), ...parsed.excludedSites.map(site => `-${siteHost(site)}`)]),
].slice(0, 20);
// search_language_filter takes ISO 639-1 two-letter codes; pass `en-us` as
// `en`, and leave anything else to the central post-filter.
const langCode = parsed.lang ? /^([a-z]{2})(?:[-_]|$)/.exec(parsed.lang)?.[1] : undefined;
return {
query: formatQuery(parsed, PERPLEXITY_QUERY_SYNTAX),
domainFilter: domains.length > 0 ? domains : undefined,
afterDate: parsed.after ? toPerplexityDate(parsed.after) : undefined,
beforeDate: parsed.before ? toPerplexityDate(parsed.before) : undefined,
languageFilter: langCode ? [langCode] : undefined,
};
}
interface PerplexityOAuthStreamMarkdownBlock {
answer?: string;
chunks?: string[];
@@ -258,6 +324,8 @@ export interface PerplexitySearchParams {
signal?: AbortSignal;
query: string;
system_prompt?: string;
/** Pre-parsed view of `query` from the search pipeline; parsed locally when absent. */
parsedQuery?: StructuredQuery;
search_recency_filter?: "hour" | "day" | "week" | "month" | "year";
num_results?: number;
/** Maximum output tokens. Defaults to 8192. */
@@ -352,6 +420,10 @@ function buildPerplexityExtraBody(request: PerplexityRequest): Record<string, un
language_preference: request.language_preference,
return_related_questions: request.return_related_questions,
search_recency_filter: request.search_recency_filter,
search_domain_filter: request.search_domain_filter,
search_after_date_filter: request.search_after_date_filter,
search_before_date_filter: request.search_before_date_filter,
search_language_filter: request.search_language_filter,
};
}
@@ -531,6 +603,7 @@ function buildOAuthAnswer(event: PerplexityOAuthStreamEvent): string {
async function callPerplexityAsk(
auth: { type: "oauth"; token: string } | { type: "cookies"; cookies: string } | { type: "anonymous" },
params: PerplexitySearchParams,
filters: PerplexityNativeFilters,
): Promise<{ answer: string; sources: SearchSource[]; model?: string; requestId?: string }> {
const requestId = crypto.randomUUID();
// The consumer `perplexity_ask` endpoint is itself a research assistant and
@@ -539,7 +612,7 @@ async function callPerplexityAsk(
// "I don't have access to web-search tools in this turn", so ask-endpoint
// searches send the bare query. (The API-key path still uses system_prompt
// as a proper `system` message.)
const effectiveQuery = params.query;
const effectiveQuery = filters.query;
const headers: Record<string, string> = {
"Content-Type": "application/json",
@@ -578,7 +651,9 @@ async function callPerplexityAsk(
version: OAUTH_API_VERSION,
language: "en-US",
timezone: Intl.DateTimeFormat().resolvedOptions().timeZone,
search_recency_filter: params.search_recency_filter ?? null,
// Recency cannot be combined with absolute date filters; explicit
// before:/after: bounds take precedence.
search_recency_filter: filters.afterDate || filters.beforeDate ? null : (params.search_recency_filter ?? null),
is_incognito: true,
use_schematized_api: true,
// `true` (the native app's default) lets the backend classifier skip
@@ -603,6 +678,10 @@ async function callPerplexityAsk(
if (auth.type === "anonymous") {
requestParams.send_back_text_in_streaming_api = true;
}
if (filters.domainFilter) requestParams.search_domain_filter = filters.domainFilter;
if (filters.afterDate) requestParams.search_after_date_filter = filters.afterDate;
if (filters.beforeDate) requestParams.search_before_date_filter = filters.beforeDate;
if (filters.languageFilter) requestParams.search_language_filter = filters.languageFilter;
const requestInit = {
method: "POST",
@@ -781,12 +860,14 @@ function applySourceLimit(result: SearchResponse, limit?: number): SearchRespons
/** Execute Perplexity web search */
export async function searchPerplexity(params: PerplexitySearchParams): Promise<SearchResponse> {
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const filters = buildNativeFilters(parsed, params.query);
const systemPrompt = params.system_prompt;
const messages: PerplexityRequest["messages"] = [];
if (systemPrompt) {
messages.push({ role: "system", content: systemPrompt });
}
messages.push({ role: "user", content: params.query });
messages.push({ role: "user", content: filters.query });
const request: PerplexityRequest = {
model: "sonar-pro",
@@ -805,7 +886,13 @@ export async function searchPerplexity(params: PerplexitySearchParams): Promise<
return_related_questions: true,
};
if (params.search_recency_filter) {
if (filters.domainFilter) request.search_domain_filter = filters.domainFilter;
if (filters.afterDate) request.search_after_date_filter = filters.afterDate;
if (filters.beforeDate) request.search_before_date_filter = filters.beforeDate;
if (filters.languageFilter) request.search_language_filter = filters.languageFilter;
// The API rejects search_recency_filter combined with absolute date
// filters; explicit before:/after: bounds take precedence.
if (params.search_recency_filter && !filters.afterDate && !filters.beforeDate) {
request.search_recency_filter = params.search_recency_filter;
}
@@ -830,10 +917,10 @@ export async function searchPerplexity(params: PerplexitySearchParams): Promise<
? await withOAuthAccess(
params.authStorage,
"perplexity",
access => callPerplexityAsk({ type: "oauth", token: access.accessToken }, params),
access => callPerplexityAsk({ type: "oauth", token: access.accessToken }, params, filters),
{ sessionId: params.sessionId, signal: params.signal, seed: auth.access },
)
: await callPerplexityAsk(auth, params);
: await callPerplexityAsk(auth, params, filters);
return applySourceLimit(
{
provider: "perplexity",
@@ -892,6 +979,7 @@ export class PerplexityProvider extends SearchProvider {
return searchPerplexity({
signal: params.signal,
query: params.query,
parsedQuery: params.parsedQuery,
temperature: params.temperature,
max_tokens: params.maxOutputTokens,
num_search_results: params.numSearchResults,
@@ -14,6 +14,9 @@
* searxng.basicUsername - Optional RFC 7617 Basic auth username
* searxng.basicPassword - Optional RFC 7617 Basic auth password
* searxng.categories - Optional comma-separated categories filter
* searxng.engines - Optional comma-separated engine names or shortcuts
* (e.g. "duckduckgo, br, sp"); shortcuts resolve via
* the instance's /config endpoint
* searxng.language - Optional language code (e.g. en, zh-CN)
*
* Environment variable fallbacks:
@@ -22,6 +25,11 @@
* SEARXNG_BASIC_USERNAME - Optional RFC 7617 Basic auth username
* SEARXNG_BASIC_PASSWORD - Optional RFC 7617 Basic auth password
*
* Bang syntax in queries is passed through: `!ddg foo` selects an engine or
* category server-side and the bang token is stripped from the upstream query.
* External bangs (`!!g`) are removed client-side because SearXNG answers them
* with an HTTP redirect even for JSON requests.
*
* Reference: https://docs.searxng.org/dev/search_api.html
*/
@@ -30,6 +38,8 @@ import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai";
import { settings } from "../../../config/settings";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import type { StructuredQuery } from "../query";
import { formatScraperQuery, parseSearchQuery } from "../query";
import { clampNumResults, dateToAgeSeconds } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -73,6 +83,11 @@ interface SearXNGAuth {
value: string;
}
/** Subset of the SearXNG /config payload used for engine shortcut resolution. */
interface SearXNGConfig {
engines?: Array<{ name?: string; shortcut?: string }>;
}
/** Find SearXNG endpoint from settings or environment. */
function findEndpoint(): string | null {
try {
@@ -150,6 +165,110 @@ function findAuth(): SearXNGAuth | null {
return token ? { type: "bearer", value: token } : null;
}
/** Find configured engine names/shortcuts from settings. */
function findEngines(): string | null {
try {
const engines = settings.get("searxng.engines");
if (engines) return engines;
} catch {
// Settings not initialized yet
}
return null;
}
/** Build request headers including authentication. */
function buildHeaders(auth: SearXNGAuth | null): Record<string, string> {
const headers: Record<string, string> = { Accept: "application/json" };
if (auth?.type === "basic") {
headers.Authorization = `Basic ${auth.value}`;
} else if (auth?.type === "bearer") {
headers.Authorization = `Bearer ${auth.value}`;
}
return headers;
}
/** Per-endpoint cache of shortcut/name → canonical engine name maps. */
const engineNameMapCache = new Map<string, Promise<Map<string, string> | null>>();
/** Fetch the instance's /config and build a lookup of lowercased engine names
* and shortcuts to canonical engine names. Returns null on any failure. */
async function fetchEngineNameMap(
base: string,
auth: SearXNGAuth | null,
fetchImpl: FetchImpl | undefined,
signal: AbortSignal | undefined,
): Promise<Map<string, string> | null> {
try {
const response = await (fetchImpl ?? fetch)(`${base}/config`, {
headers: buildHeaders(auth),
signal: withHardTimeout(signal),
});
if (!response.ok) return null;
const config = (await response.json()) as SearXNGConfig;
const map = new Map<string, string>();
for (const engine of config.engines ?? []) {
if (!engine.name) continue;
map.set(engine.name.toLowerCase(), engine.name);
if (engine.shortcut) map.set(engine.shortcut.toLowerCase(), engine.name);
}
return map.size ? map : null;
} catch {
return null;
}
}
/** Get the engine name map for an endpoint, cached for the process lifetime.
* Failures are not cached so a transient error retries on the next search. */
function getEngineNameMap(
endpoint: string,
auth: SearXNGAuth | null,
fetchImpl: FetchImpl | undefined,
signal: AbortSignal | undefined,
): Promise<Map<string, string> | null> {
const base = endpoint.replace(/\/+$/, "");
let cached = engineNameMapCache.get(base);
if (!cached) {
cached = fetchEngineNameMap(base, auth, fetchImpl, signal).then(map => {
if (!map) engineNameMapCache.delete(base);
return map;
});
engineNameMapCache.set(base, cached);
}
return cached;
}
/** Resolve configured engine entries (canonical names or shortcuts like `ddg`)
* to canonical names for SearXNG's `engines=` parameter, which accepts names
* only — shortcuts resolve exclusively through bang syntax. Unknown entries
* pass through verbatim; the server drops them and falls back to categories. */
async function resolveEngineNames(
raw: string,
endpoint: string,
auth: SearXNGAuth | null,
fetchImpl: FetchImpl | undefined,
signal: AbortSignal | undefined,
): Promise<string | undefined> {
const entries = raw
.split(",")
.map(entry => entry.trim())
.filter(Boolean);
if (!entries.length) return undefined;
const map = await getEngineNameMap(endpoint, auth, fetchImpl, signal);
if (!map) return entries.join(",");
return entries.map(entry => map.get(entry.toLowerCase()) ?? entry).join(",");
}
/** Strip external bang tokens (`!!g`, bare `!!`): SearXNG answers them with an
* HTTP redirect even for JSON requests, which breaks response parsing.
* Single-bang engine/category selectors (`!ddg`, `!images`) are kept — the
* instance resolves and removes them server-side. */
function stripExternalBangs(query: string): string {
return query
.split(/\s+/)
.filter(part => !part.startsWith("!!"))
.join(" ");
}
/** Build the search URL and headers for a SearXNG request */
function buildRequest(
endpoint: string,
@@ -158,6 +277,7 @@ function buildRequest(
num_results?: number;
recency?: "day" | "week" | "month" | "year";
categories?: string;
engines?: string;
language?: string;
signal?: AbortSignal;
},
@@ -181,19 +301,15 @@ function buildRequest(
url.searchParams.set("categories", params.categories);
}
if (params.engines) {
url.searchParams.set("engines", params.engines);
}
if (params.language) {
url.searchParams.set("language", params.language);
}
const headers: Record<string, string> = {
Accept: "application/json",
};
if (auth?.type === "basic") {
headers.Authorization = `Basic ${auth.value}`;
} else if (auth?.type === "bearer") {
headers.Authorization = `Bearer ${auth.value}`;
}
const headers = buildHeaders(auth);
return { url, headers };
}
@@ -205,6 +321,7 @@ async function callSearXNGSearch(
num_results?: number;
recency?: "day" | "week" | "month" | "year";
categories?: string;
engines?: string;
language?: string;
signal?: AbortSignal;
fetch?: FetchImpl;
@@ -231,6 +348,7 @@ async function callSearXNGSearch(
/** Execute SearXNG web search. */
export async function searchSearXNG(params: {
query: string;
parsedQuery?: StructuredQuery;
num_results?: number;
recency?: "day" | "week" | "month" | "year";
signal?: AbortSignal;
@@ -255,12 +373,29 @@ export async function searchSearXNG(params: {
} catch {
// Settings not initialized yet
}
const configuredEngines = findEngines();
// SearXNG forwards `q` to downstream engines, so build it with the shared
// scraper formatter: operators are canonicalized and scraper-hostile ones
// (path-carrying `site:`, `inurl:`) are structurally demoted to plain
// terms before formatting, so paren-grouped `site:` filters are covered
// too. `lang:` maps onto the native `language` param (overriding the
// configured default).
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const query = formatScraperQuery(params.query, parsed);
if (parsed.lang) language = parsed.lang;
const engines = configuredEngines
? await resolveEngineNames(configuredEngines, endpoint, auth, params.fetch, params.signal)
: undefined;
const response = await callSearXNGSearch(
endpoint,
{
...params,
query: stripExternalBangs(query),
categories,
engines,
language,
fetch: params.fetch,
},
@@ -315,6 +450,7 @@ export class SearXNGProvider extends SearchProvider {
search(params: SearchParams): Promise<SearchResponse> {
return searchSearXNG({
parsedQuery: params.parsedQuery,
query: params.query,
num_results: params.numSearchResults ?? params.limit,
recency: params.recency,
@@ -2,6 +2,7 @@ import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai";
import { parseHTML } from "linkedom";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatScraperQuery } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -136,12 +137,17 @@ async function callStartpageHtml(params: SearchParams): Promise<string> {
const fetchImpl = params.fetch ?? fetch;
const signal = withHardTimeout(params.signal);
const withDate = params.recency ? RECENCY_TO_STARTPAGE_WITH_DATE[params.recency] : undefined;
// Startpage proxies Google, so the operator set works inline; rebuild via
// the shared scraper formatter to canonicalize aliases (domain: → site:,
// since: → after:, …) and demote scraper-hostile operators. Directive-free
// queries pass through byte-identical.
const query = formatScraperQuery(params.query, params.parsedQuery);
const formInputs = await fetchFormInputs(fetchImpl, signal);
let page: LoadedHtmlPage;
if (formInputs) {
const form = new URLSearchParams(formInputs);
form.set("query", params.query);
form.set("query", query);
if (withDate) form.set("with_date", withDate);
page = await browserFetch(STARTPAGE_SEARCH_URL, {
fetch: fetchImpl,
@@ -152,7 +158,7 @@ async function callStartpageHtml(params: SearchParams): Promise<string> {
});
} else {
const url = new URL(STARTPAGE_SEARCH_URL);
url.searchParams.set("query", params.query);
url.searchParams.set("query", query);
if (withDate) url.searchParams.set("with_date", withDate);
page = await browserFetch(url.href, {
fetch: fetchImpl,
@@ -8,6 +8,7 @@
import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth } from "@oh-my-pi/pi-ai";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery } from "../query";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
import { classifyProviderHttpError, withHardTimeout } from "./utils";
@@ -73,8 +74,13 @@ export async function searchSynthetic(params: SearchParamsWithFetch): Promise<Se
sessionId: params.sessionId,
});
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const query = parsed.hasDirectives
? formatQuery(parsed, { phrases: true, negation: true, site: true })
: params.query;
const fetchImpl = params.fetch;
const data = await withAuth(keyOrResolver, key => callSyntheticSearch(key, params.query, params.signal, fetchImpl), {
const data = await withAuth(keyOrResolver, key => callSyntheticSearch(key, query, params.signal, fetchImpl), {
signal: params.signal,
missingKeyMessage: "Synthetic credentials not found. Set SYNTHETIC_API_KEY or login with 'omp /login synthetic'.",
});
@@ -7,6 +7,7 @@
import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth } from "@oh-my-pi/pi-ai";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery } from "../query";
import { clampNumResults, dateToAgeSeconds } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -20,6 +21,14 @@ export interface TavilySearchParams {
query: string;
num_results?: number;
recency?: "day" | "week" | "month" | "year";
/** `site:` hosts mapped to Tavily's `include_domains`. */
include_domains?: string[];
/** `-site:` hosts mapped to Tavily's `exclude_domains`. */
exclude_domains?: string[];
/** `after:` inclusive lower bound, ISO `YYYY-MM-DD`, mapped to `start_date`. */
start_date?: string;
/** `before:` upper bound, ISO `YYYY-MM-DD`, mapped to `end_date`. */
end_date?: string;
signal?: AbortSignal;
fetch?: FetchImpl;
}
@@ -83,7 +92,21 @@ export function buildRequestBody(params: TavilySearchParams): Record<string, unk
include_answer: "advanced",
include_raw_content: false,
};
if (params.recency) {
if (params.include_domains?.length) {
body.include_domains = params.include_domains;
}
if (params.exclude_domains?.length) {
body.exclude_domains = params.exclude_domains;
}
if (params.start_date) {
body.start_date = params.start_date;
}
if (params.end_date) {
body.end_date = params.end_date;
}
// Explicit before:/after: bounds take precedence over the relative recency
// window; sending both would over-restrict.
if (params.recency && !params.start_date && !params.end_date) {
body.time_range = params.recency;
}
return body;
@@ -148,8 +171,19 @@ function hasRenderableResponse(response: SearchResponse): boolean {
return response.sources.length > 0;
}
/** Bare hosts from `site:` values (path parts are enforced by the central lenient filter). */
function siteHosts(sites: readonly string[]): string[] {
const hosts = new Set<string>();
for (const site of sites) {
const host = site.split("/", 1)[0];
if (host) hosts.add(host);
}
return [...hosts];
}
/** Execute Tavily web search. */
export async function searchTavily(params: SearchParams): Promise<SearchResponse> {
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const tavilyParams: TavilySearchParams = {
query: params.query,
num_results: params.numSearchResults ?? params.limit,
@@ -157,6 +191,16 @@ export async function searchTavily(params: SearchParams): Promise<SearchResponse
signal: params.signal,
fetch: params.fetch,
};
if (parsed.hasDirectives) {
// Tavily prefers clean natural text; re-emit only phrases and -exclusions.
tavilyParams.query = formatQuery(parsed, { phrases: true, negation: true });
const include = siteHosts(parsed.sites);
const exclude = siteHosts(parsed.excludedSites);
if (include.length > 0) tavilyParams.include_domains = include;
if (exclude.length > 0) tavilyParams.exclude_domains = exclude;
if (parsed.after) tavilyParams.start_date = parsed.after;
if (parsed.before) tavilyParams.end_date = parsed.before;
}
const keyOrResolver: ApiKey = params.authStorage.resolver("tavily", {
sessionId: params.sessionId,
});
@@ -171,11 +215,16 @@ export async function searchTavily(params: SearchParams): Promise<SearchResponse
withAuth(keyOrResolver, key => callTavilySearch(key, searchParams), authOptions);
const response = toSearchResponse(await callWithAuth(tavilyParams), numResults);
if (!tavilyParams.recency || hasRenderableResponse(response)) {
const hasTimeFilter = Boolean(tavilyParams.recency || tavilyParams.start_date || tavilyParams.end_date);
if (!hasTimeFilter || hasRenderableResponse(response)) {
return response;
}
return toSearchResponse(await callWithAuth({ ...tavilyParams, recency: undefined }), numResults);
// Time filters commonly zero out results; retry once without them.
return toSearchResponse(
await callWithAuth({ ...tavilyParams, recency: undefined, start_date: undefined, end_date: undefined }),
numResults,
);
}
/** Search provider for Tavily web search. */
@@ -7,6 +7,7 @@
import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth } from "@oh-my-pi/pi-ai";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery, type QuerySyntax } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -17,6 +18,9 @@ const DEFAULT_NUM_RESULTS = 10;
const MAX_NUM_RESULTS = 20;
const MAX_PAGE = 10;
/** TinyFish is SERP-backed: common Google-style operators pass through. */
const TINYFISH_QUERY_SYNTAX: QuerySyntax = { phrases: true, negation: true, site: true, filetype: true };
const RECENCY_MINUTES: Record<NonNullable<SearchParams["recency"]>, number> = {
day: 1440,
week: 10080,
@@ -107,8 +111,9 @@ function appendTinyFishSources(sources: SearchSource[], results: readonly TinyFi
export async function searchTinyFish(params: SearchParams): Promise<SearchResponse> {
const numResults = clampNumResults(params.numSearchResults ?? params.limit, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS);
const pageSize = Math.min(numResults, DEFAULT_NUM_RESULTS);
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const tinyFishParams: TinyFishSearchParams = {
query: params.query,
query: parsed.hasDirectives ? formatQuery(parsed, TINYFISH_QUERY_SYNTAX) : params.query,
num_results: pageSize,
recency: params.recency,
signal: params.signal,
@@ -3,6 +3,7 @@ import { $env } from "@oh-my-pi/pi-utils";
import { resolveXAIHttpTransport, type XAIHttpProvider, type XAIHttpTransport } from "../../../lib/xai-http";
import type { SearchCitation, SearchResponse, SearchSource, SearchUsage } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery, type QuerySyntax } from "../query";
import { clampNumResults } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -57,14 +58,60 @@ interface XAIResponsesResponse {
usage?: XAIResponsesUsage | null;
}
/**
* Query syntax re-emitted for the Grok search agent. `site:`/`-site:` are
* stripped because hosts map natively onto the web_search domain filters;
* `before:`/`after:` stay in the query text — the Responses web_search tool
* has no date parameters (`from_date`/`to_date` exist only on `x_search` and
* the deprecated Live Search `search_parameters`, which now returns 410) and
* the agent honors the tokens as natural-language hints.
*/
const XAI_QUERY_SYNTAX: QuerySyntax = {
phrases: true,
negation: true,
or: true,
inUrl: true,
inTitle: true,
filetype: true,
dateRange: true,
};
/** xAI web_search accepts at most 5 allowed or excluded domains per request. */
const MAX_DOMAIN_FILTERS = 5;
/** Bare hosts of `site:` values (`github.com/anthropics` → `github.com`), deduped, capped at 5; path parts are enforced by the central constraint filter. */
function domainFilterList(sites: readonly string[]): string[] {
const hosts = new Set<string>();
for (const site of sites) {
const slash = site.indexOf("/");
hosts.add(slash === -1 ? site : site.slice(0, slash));
if (hosts.size === MAX_DOMAIN_FILTERS) break;
}
return [...hosts];
}
function buildRequestBody(params: SearchParams): Record<string, unknown> {
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
const webSearchTool: Record<string, unknown> = { type: "web_search" };
let query = params.query;
if (parsed.hasDirectives) {
query = formatQuery(parsed, XAI_QUERY_SYNTAX);
// allowed_domains and excluded_domains are mutually exclusive per
// request; prefer the allow list, the central filter enforces exclusions.
if (parsed.sites.length > 0) {
webSearchTool.filters = { allowed_domains: domainFilterList(parsed.sites) };
} else if (parsed.excludedSites.length > 0) {
webSearchTool.filters = { excluded_domains: domainFilterList(parsed.excludedSites) };
}
}
const body: Record<string, unknown> = {
model: XAI_WEB_SEARCH_MODEL,
input: [
{ role: "system", content: params.systemPrompt },
{ role: "user", content: params.query },
{ role: "user", content: query },
],
tools: [{ type: "web_search" }],
tools: [webSearchTool],
reasoning: { effort: XAI_WEB_SEARCH_REASONING_EFFORT },
};
@@ -8,6 +8,7 @@ import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth }
import { isRecord } from "@oh-my-pi/pi-utils";
import type { SearchResponse, SearchSource } from "../../../web/search/types";
import { SearchProviderError } from "../../../web/search/types";
import { formatQuery, parseSearchQuery, type QuerySyntax } from "../query";
import { dateToAgeSeconds } from "../utils";
import type { SearchParams } from "./base";
import { SearchProvider } from "./base";
@@ -17,6 +18,22 @@ const ZAI_MCP_URL = "https://api.z.ai/api/mcp/web_search_prime/mcp";
const ZAI_TOOL_NAME = "web_search_prime";
const DEFAULT_NUM_RESULTS = 10;
/**
* webSearchPrime exposes no native filter args (the Web Search API's
* `search_domain_filter`/`search_recency_filter` are scoped to the
* `search_pro_jina` engine, not `search-prime`), but its Bing-flavored
* backend parses the common inline operators. Dates and language are left
* to the central lenient post-filter.
*/
const ZAI_QUERY_SYNTAX: QuerySyntax = {
phrases: true,
negation: true,
site: true,
inTitle: true,
inUrl: true,
filetype: true,
};
export interface ZaiSearchParams {
query: string;
num_results?: number;
@@ -407,8 +424,9 @@ export class ZaiProvider extends SearchProvider {
search(params: SearchParams): Promise<SearchResponse> {
const { fetch: fetchOverride } = params as ZaiProviderSearchParams;
const parsed = params.parsedQuery ?? parseSearchQuery(params.query);
return searchZai({
query: params.query,
query: parsed.hasDirectives ? formatQuery(parsed, ZAI_QUERY_SYNTAX) : params.query,
num_results: params.numSearchResults ?? params.limit,
signal: params.signal,
authStorage: params.authStorage,
@@ -0,0 +1,850 @@
/**
* Structured web-search query parsing.
*
* Agents habitually embed Google-style directives in search queries —
* `site:`, `before:`/`after:`, `inurl:`, `filetype:`, quoted phrases, `OR`
* groups, `-exclusions` — regardless of whether the backing engine parses
* them. This module turns a raw query into a {@link StructuredQuery} so each
* provider can:
*
* 1. map constraints onto native API parameters where they exist (Perplexity
* `search_domain_filter`, Tavily `include_domains`, Exa date bounds, …),
* 2. rebuild a query string containing only the syntax the target engine
* understands ({@link formatQuery}), and
* 3. post-filter returned sources leniently ({@link applyQueryConstraints}):
* a constraint dimension that would eliminate every result is dropped and
* reported rather than returning nothing.
*/
import type { SearchSource } from "./types";
/** One free-text token of the query (everything that is not a recognized directive). */
export interface QueryTerm {
/** Term text without quotes or operator prefixes. */
text: string;
/** Quoted exact phrase (`"like this"`) or verbatim-required (`+term`). */
phrase?: boolean;
/** Excluded via `-term` or `NOT term`. */
negated?: boolean;
/**
* OR-group id. Terms sharing an id are alternatives (`a OR b`); terms
* without a group are implicitly AND-ed. Groups are always contiguous
* runs in {@link StructuredQuery.terms}.
*/
group?: number;
}
/**
* A raw query decomposed into free text plus every recognized constraint.
*
* All list fields are always present (possibly empty) so consumers can map
* over them without null checks. Values are stored as typed by the user
* except for normalization noted per field.
*/
export interface StructuredQuery {
/** Original query string, verbatim. */
raw: string;
/**
* Free-text remainder with all recognized directives removed; phrases
* stay quoted, exclusions keep `-`, OR groups keep `OR`. Empty when the
* query was directives only — use {@link formatQuery} for a never-empty
* engine query.
*/
text: string;
/** Ordered free-text terms (phrases, exclusions, OR groups). */
terms: QueryTerm[];
/** `site:`/`domain:`/`host:` includes — any-of. Lowercased, scheme stripped, may carry a path (`github.com/anthropics`). */
sites: string[];
/** `-site:` exclusions, same normalization as {@link sites}. */
excludedSites: string[];
/** `inurl:`/`url:`/`allinurl:` substrings — all must appear in the URL. */
inUrl: string[];
/** `-inurl:` substrings — none may appear in the URL. */
excludedInUrl: string[];
/** `intitle:`/`title:`/`allintitle:` substrings — all must appear in the title. */
inTitle: string[];
/** `-intitle:` substrings — none may appear in the title. */
excludedInTitle: string[];
/** `intext:`/`inbody:`/`inanchor:`/`allintext:` body substrings. Not post-filterable (snippets are partial); query-building only. */
inText: string[];
/** `-intext:` body exclusions. Query-building only. */
excludedInText: string[];
/** `filetype:`/`ext:` extensions — any-of. Lowercased, no leading dot. */
filetypes: string[];
/** `-filetype:`/`-ext:` extensions — none may match. */
excludedFiletypes: string[];
/** Inclusive lower publish-date bound from `after:`/`since:`, ISO `YYYY-MM-DD`. */
after?: string;
/** Exclusive upper publish-date bound from `before:`/`until:`, ISO `YYYY-MM-DD`. */
before?: string;
/** Language code from `lang:`/`language:`, lowercased (e.g. `en`, `en-us`). */
lang?: string;
/** True when any directive or boolean operator was recognized. */
hasDirectives: boolean;
/** True when any post-filterable constraint is set (sites, url/title terms, filetypes, date bounds). */
hasConstraints: boolean;
}
/**
* Query-syntax capabilities of a target engine, used by {@link formatQuery}
* to decide which parsed features are re-emitted as query text. Everything
* defaults to `false`: the zero-value produces plain keywords suitable for
* natural-language APIs.
*/
export interface QuerySyntax {
/** Emit `"quoted phrases"`. */
phrases?: boolean;
/** Emit `-term` exclusions (negated terms are dropped otherwise). */
negation?: boolean;
/** Emit `OR` between alternatives (groups are flattened to keywords otherwise). */
or?: boolean;
/** Emit `site:`/`-site:`. */
site?: boolean;
/** Emit `inurl:`/`-inurl:`. */
inUrl?: boolean;
/** Emit `intitle:`/`-intitle:`. */
inTitle?: boolean;
/** Emit `intext:`/`-intext:`. */
inText?: boolean;
/** Emit `filetype:`/`-filetype:`. */
filetype?: boolean;
/** Emit `before:`/`after:` ISO date bounds. */
dateRange?: boolean;
}
/** Full Google-style syntax: engines that parse the classic operator set (Google, Startpage, Ecosia, Brave, Kagi, Mojeek, SearXNG…). */
export const GOOGLE_QUERY_SYNTAX: QuerySyntax = {
phrases: true,
negation: true,
or: true,
site: true,
inUrl: true,
inTitle: true,
inText: true,
filetype: true,
dateRange: true,
};
/** Result of {@link applyQueryConstraints}. */
export interface ConstraintFilterResult {
/** Sources surviving the lenient filter — never empty when the input was non-empty. */
sources: SearchSource[];
/**
* Directive renderings (`site:arxiv.org`, `before:2024-01-01`, …) of the
* constraint dimensions that matched zero sources and were therefore
* relaxed instead of enforced.
*/
dropped: string[];
}
const DIRECTIVE_PATTERN = /^([+-]?)([a-z][a-z-]*):(.*)$/i;
type AllMode = "inTitle" | "inUrl" | "inText";
interface RawToken {
text: string;
/** Entire token was a quoted phrase. */
quoted: boolean;
/** Directive value was quoted (`intitle:"a b"`). */
quotedValue?: boolean;
}
function isQuote(ch: string): boolean {
return ch === '"' || ch === "\u201c" || ch === "\u201d";
}
/** Unicode-aware whitespace (agents paste NBSP and friends). */
const WHITESPACE = /\s/;
/** Split a raw query into whitespace-delimited tokens, honoring quoted spans and standalone parens. */
function tokenize(raw: string): RawToken[] {
const tokens: RawToken[] = [];
const n = raw.length;
let i = 0;
while (i < n) {
const ch = raw[i];
if (WHITESPACE.test(ch)) {
i++;
continue;
}
if (isQuote(ch)) {
let j = i + 1;
let buf = "";
while (j < n && !isQuote(raw[j])) {
buf += raw[j];
j++;
}
if (buf.trim().length > 0) tokens.push({ text: buf.trim(), quoted: true });
i = j + 1;
continue;
}
// Bare word; a quote directly after `name:` swallows the quoted span
// into the same token (`intitle:"budget tips"`).
let buf = "";
let quotedValue = false;
while (i < n && !WHITESPACE.test(raw[i])) {
const c = raw[i];
if (isQuote(c) && buf.endsWith(":")) {
let j = i + 1;
while (j < n && !isQuote(raw[j])) {
buf += raw[j];
j++;
}
quotedValue = true;
i = j + 1;
continue;
}
if (isQuote(c)) break; // `foo"bar` — stop the word, let the quote start a phrase
buf += c;
i++;
}
if (buf.length > 0) tokens.push({ text: buf, quoted: false, quotedValue });
}
return splitParens(tokens);
}
/**
* Split leading `(` and unbalanced trailing `)` into standalone tokens so
* `(react OR vue)` parses while `site:wikipedia.org/Foo_(bar)` stays whole.
*/
function splitParens(tokens: RawToken[]): RawToken[] {
const out: RawToken[] = [];
for (const tok of tokens) {
if (tok.quoted || tok.quotedValue) {
out.push(tok);
continue;
}
let text = tok.text;
while (text.startsWith("(")) {
out.push({ text: "(", quoted: false });
text = text.slice(1);
}
let trailing = 0;
while (text.endsWith(")")) {
// Only strip parens that do not close an opener inside the word.
const body = text.slice(0, -1);
let depth = 0;
for (const c of body) {
if (c === "(") depth++;
else if (c === ")") depth--;
}
if (depth > 0) break;
text = body;
trailing++;
}
if (text.length > 0) out.push({ text, quoted: false });
for (let k = 0; k < trailing; k++) out.push({ text: ")", quoted: false });
}
return out;
}
/** Convert year/month/day parts to a validated ISO date, or undefined. */
function isoDate(year: number, month: number, day: number): string | undefined {
if (year < 1000 || year > 9999 || month < 1 || month > 12 || day < 1 || day > 31) return undefined;
return `${year}-${String(month).padStart(2, "0")}-${String(day).padStart(2, "0")}`;
}
/**
* Parse a `before:`/`after:` value into ISO `YYYY-MM-DD`.
* Accepts `YYYY`, `YYYY-MM`, `YYYY-MM-DD` (also `/` and `.` separators) and
* `MM/DD/YYYY` (day-first assumed when the first field exceeds 12).
* Bare years/months resolve to the first day of the period, matching
* Google's `after:2024` ≙ `after:2024-01-01` semantics.
*/
export function parseDateValue(value: string): string | undefined {
const t = value.trim();
let m = /^(\d{4})(?:[-/.](\d{1,2})(?:[-/.](\d{1,2}))?)?$/.exec(t);
if (m) return isoDate(Number(m[1]), m[2] ? Number(m[2]) : 1, m[3] ? Number(m[3]) : 1);
m = /^(\d{1,2})[-/.](\d{1,2})[-/.](\d{4})$/.exec(t);
if (m) {
let month = Number(m[1]);
let day = Number(m[2]);
if (month > 12 && day <= 12) [month, day] = [day, month];
return isoDate(Number(m[3]), month, day);
}
return undefined;
}
/** Lowercase a `site:` value and strip scheme, `*.` wildcard, and trailing slash/dot. */
function normalizeSite(value: string): string {
let site = value.trim().toLowerCase();
site = site.replace(/^[a-z][a-z0-9+.-]*:\/\//, "");
if (site.startsWith("*.")) site = site.slice(2);
site = site.replace(/[/.]+$/, "");
return site;
}
/** Directive names mapped to their canonical field. */
const DIRECTIVE_FIELDS: Record<
string,
"site" | "inUrl" | "inTitle" | "inText" | "filetype" | "before" | "after" | "lang"
> = {
site: "site",
domain: "site",
host: "site",
inurl: "inUrl",
url: "inUrl",
intitle: "inTitle",
title: "inTitle",
intext: "inText",
inbody: "inText",
inanchor: "inText",
filetype: "filetype",
ext: "filetype",
before: "before",
until: "before",
after: "after",
since: "after",
lang: "lang",
language: "lang",
};
/** `allin*:` directives that capture every following plain term. */
const ALL_MODES: Record<string, AllMode> = {
allintitle: "inTitle",
allinurl: "inUrl",
allintext: "inText",
};
/** True for operator/paren tokens and recognized directives — anything a bare `name:` must not adopt as its value. */
function isReservedToken(text: string): boolean {
if (
text === "(" ||
text === ")" ||
text === "OR" ||
text === "AND" ||
text === "NOT" ||
text === "|" ||
text === "||" ||
text === "&&" ||
text === "!"
) {
return true;
}
const m = DIRECTIVE_PATTERN.exec(text);
if (!m) return false;
const name = m[2].toLowerCase();
return DIRECTIVE_FIELDS[name] !== undefined || ALL_MODES[name] !== undefined;
}
/**
* Parse a raw query into a {@link StructuredQuery}.
*
* Lenient by construction: unknown `name:value` tokens (URLs, `C:\paths`,
* `TS2345:`, jargon) stay in the free text verbatim, and a directive with an
* unparseable value (`before:someday`) degrades to a plain term instead of
* being dropped.
*/
export function parseSearchQuery(raw: string): StructuredQuery {
const q: StructuredQuery = {
raw,
text: "",
terms: [],
sites: [],
excludedSites: [],
inUrl: [],
excludedInUrl: [],
inTitle: [],
excludedInTitle: [],
inText: [],
excludedInText: [],
filetypes: [],
excludedFiletypes: [],
hasDirectives: false,
hasConstraints: false,
};
const tokens = tokenize(raw);
let negateNext = false;
let orPending = false;
let lastWasTerm = false;
let groupSeq = 0;
let allMode: AllMode | undefined;
const pushConstraint = (
field: "site" | "inUrl" | "inTitle" | "inText" | "filetype",
value: string,
negated: boolean,
): void => {
q.hasDirectives = true;
orPending = false;
lastWasTerm = false;
const v = value.trim();
if (!v) return;
switch (field) {
case "site": {
const site = normalizeSite(v);
if (site) (negated ? q.excludedSites : q.sites).push(site);
break;
}
case "inUrl":
(negated ? q.excludedInUrl : q.inUrl).push(v);
break;
case "inTitle":
(negated ? q.excludedInTitle : q.inTitle).push(v);
break;
case "inText":
(negated ? q.excludedInText : q.inText).push(v);
break;
case "filetype": {
const ext = v.toLowerCase().replace(/^\.+/, "");
if (ext) (negated ? q.excludedFiletypes : q.filetypes).push(ext);
break;
}
}
};
const pushTerm = (text: string, phrase: boolean): void => {
const negated = negateNext;
negateNext = false;
if (allMode && !negated) {
pushConstraint(allMode, text, false);
return;
}
if (allMode && negated) {
pushConstraint(allMode, text, true);
return;
}
const term: QueryTerm = { text };
if (phrase) term.phrase = true;
if (negated) term.negated = true;
if (orPending && lastWasTerm) {
const prev = q.terms[q.terms.length - 1];
if (prev) {
prev.group ??= ++groupSeq;
term.group = prev.group;
}
}
orPending = false;
lastWasTerm = true;
q.terms.push(term);
};
for (let idx = 0; idx < tokens.length; idx++) {
const tok = tokens[idx];
if (tok.quoted) {
pushTerm(tok.text, true);
continue;
}
// Boolean operators and grouping parens.
if (tok.text === "(" || tok.text === ")") continue;
if (tok.text === "OR" || tok.text === "|" || tok.text === "||") {
orPending = true;
q.hasDirectives = true;
continue;
}
if (tok.text === "AND" || tok.text === "&&") {
q.hasDirectives = true;
continue;
}
if (tok.text === "NOT" || tok.text === "!") {
negateNext = true;
q.hasDirectives = true;
continue;
}
if (tok.text === "-" || tok.text === "+") {
// The tokenizer splits `-"exact phrase"` into `-` + phrase; carry the negation over.
if (tok.text === "-" && tokens[idx + 1]?.quoted) negateNext = true;
continue;
}
const match = DIRECTIVE_PATTERN.exec(tok.text);
const name = match?.[2].toLowerCase();
const allMatch = name ? ALL_MODES[name] : undefined;
const field = name ? DIRECTIVE_FIELDS[name] : undefined;
if (match && allMatch) {
allMode = allMatch;
q.hasDirectives = true;
// `allintitle:budget tips` — inline value plus every following term.
const inline = match[3].trim();
if (inline) pushConstraint(allMatch, inline, match[1] === "-");
orPending = false;
lastWasTerm = false;
continue;
}
if (match && field) {
let value = match[3].trim();
// `site: example.com` — lenient: adopt the next plain token as the value.
if (!value) {
const next = tokens[idx + 1];
if (next && (next.quoted || !isReservedToken(next.text))) {
value = next.text.trim();
idx++;
}
}
if (!value) {
q.hasDirectives = true;
continue;
}
const negated = match[1] === "-" || negateNext;
negateNext = false;
switch (field) {
case "before":
case "after": {
const iso = parseDateValue(value);
if (!iso) {
pushTerm(tok.text, false);
continue;
}
if (field === "before") q.before = iso;
else q.after = iso;
q.hasDirectives = true;
orPending = false;
lastWasTerm = false;
break;
}
case "lang":
q.lang = value.toLowerCase();
q.hasDirectives = true;
orPending = false;
lastWasTerm = false;
break;
default:
pushConstraint(field, value, negated);
}
continue;
}
// Plain term with optional +/- prefix.
let text = tok.text;
if (text.startsWith("-") && text.length > 1) {
negateNext = true;
q.hasDirectives = true;
text = text.replace(/^-+/, "");
if (!text) continue;
// `-site:x` arrives pre-split only when written `- site:x`; re-check directive.
const negMatch = DIRECTIVE_PATTERN.exec(text);
const negName = negMatch?.[2].toLowerCase();
const negField = negName ? DIRECTIVE_FIELDS[negName] : undefined;
if (negMatch && negField && negField !== "before" && negField !== "after" && negField !== "lang") {
negateNext = false;
pushConstraint(negField, negMatch[3].trim(), true);
continue;
}
pushTerm(text, false);
continue;
}
if (text.startsWith("+") && text.length > 1) {
// Legacy Google `+term`: verbatim/required — treat as an exact phrase.
pushTerm(text.slice(1), true);
q.hasDirectives = true;
continue;
}
pushTerm(text, false);
}
q.text = renderTerms(q.terms, { phrases: true, negation: true, or: true });
q.hasConstraints =
q.sites.length > 0 ||
q.excludedSites.length > 0 ||
q.inUrl.length > 0 ||
q.excludedInUrl.length > 0 ||
q.inTitle.length > 0 ||
q.excludedInTitle.length > 0 ||
q.filetypes.length > 0 ||
q.excludedFiletypes.length > 0 ||
q.before !== undefined ||
q.after !== undefined;
return q;
}
/** Quote a directive value when it contains whitespace. */
function quoteValue(value: string): string {
return /\s/.test(value) ? `"${value}"` : value;
}
/** Render the free-text terms per the target syntax. */
function renderTerms(terms: readonly QueryTerm[], syntax: QuerySyntax): string {
const parts: string[] = [];
for (let i = 0; i < terms.length; i++) {
const term = terms[i];
if (term.group !== undefined && syntax.or) {
const members: string[] = [];
let j = i;
for (; j < terms.length && terms[j].group === term.group; j++) {
const rendered = renderTerm(terms[j], syntax);
if (rendered) members.push(rendered);
}
i = j - 1;
if (members.length > 1) parts.push(`(${members.join(" OR ")})`);
else if (members.length === 1) parts.push(members[0]);
continue;
}
const rendered = renderTerm(term, syntax);
if (rendered) parts.push(rendered);
}
return parts.join(" ");
}
function renderTerm(term: QueryTerm, syntax: QuerySyntax): string | undefined {
if (term.negated && !syntax.negation) return undefined;
const body = term.phrase && syntax.phrases ? `"${term.text}"` : term.text;
return term.negated ? `-${body}` : body;
}
/**
* Rebuild a query string for an engine with the given {@link QuerySyntax}.
*
* Constraints whose syntax the engine lacks are omitted (the caller maps
* them onto API parameters or relies on {@link applyQueryConstraints}).
* Never returns an empty string for a non-empty input: a directives-only
* query falls back to the constraint values as keywords, then to `raw` — an
* engine searching *something* beats an empty-query error.
*/
export function formatQuery(q: StructuredQuery, syntax: QuerySyntax = {}): string {
const parts: string[] = [];
const text = renderTerms(q.terms, syntax);
if (text) parts.push(text);
if (syntax.site) {
if (q.sites.length > 1 && syntax.or) parts.push(`(${q.sites.map(s => `site:${s}`).join(" OR ")})`);
else parts.push(...q.sites.map(s => `site:${s}`));
parts.push(...q.excludedSites.map(s => `-site:${s}`));
}
if (syntax.inUrl) {
parts.push(...q.inUrl.map(v => `inurl:${quoteValue(v)}`));
parts.push(...q.excludedInUrl.map(v => `-inurl:${quoteValue(v)}`));
}
if (syntax.inTitle) {
parts.push(...q.inTitle.map(v => `intitle:${quoteValue(v)}`));
parts.push(...q.excludedInTitle.map(v => `-intitle:${quoteValue(v)}`));
}
if (syntax.inText) {
parts.push(...q.inText.map(v => `intext:${quoteValue(v)}`));
parts.push(...q.excludedInText.map(v => `-intext:${quoteValue(v)}`));
}
if (syntax.filetype) {
if (q.filetypes.length > 1 && syntax.or) parts.push(`(${q.filetypes.map(f => `filetype:${f}`).join(" OR ")})`);
else parts.push(...q.filetypes.map(f => `filetype:${f}`));
parts.push(...q.excludedFiletypes.map(f => `-filetype:${f}`));
}
if (syntax.dateRange) {
if (q.after) parts.push(`after:${q.after}`);
if (q.before) parts.push(`before:${q.before}`);
}
let result = parts.join(" ").trim();
if (!result) {
// Directives-only query and no directive syntax: search the constraint
// values as plain keywords so the engine still gets a meaningful query.
const fallback = [...q.sites, ...q.inTitle, ...q.inUrl, ...q.inText, ...q.filetypes];
result = fallback.join(" ").trim();
}
return result || q.raw.trim();
}
/**
* Build the engine query for a credential-free HTML engine (Google,
* Startpage, DuckDuckGo, Ecosia, Mojeek, SearXNG, and the Public Web
* fan-out over them).
*
* Canonicalizes directives via {@link formatQuery} with the engine's
* {@link QuerySyntax} (default: full Google syntax), after demoting the
* operators that zero-match across the scraper set: engines only match
* `site:` against a bare domain (a path yields zero results everywhere),
* and DuckDuckGo ignores `inurl:` entirely — so either operator silently
* empties the result set. The raw URL as a plain term matches fine, so
* bare-domain `site:` filters are kept while path-carrying `site:` and all
* `inurl:` values become plain keywords; the demotion is structural (before
* formatting), so OR-grouped and quoted directives are covered. Negated
* forms (`-site:`, `-inurl:`) pass through untouched — demoting them would
* invert an exclusion into a search term; the pipeline post-filter
* ({@link applyQueryConstraints}) enforces every demoted or unsupported
* constraint on the returned sources. Directive-free queries pass through
* byte-identical.
*/
export function formatScraperQuery(
query: string,
parsedQuery?: StructuredQuery,
syntax: QuerySyntax = GOOGLE_QUERY_SYNTAX,
): string {
const parsed = parsedQuery ?? parseSearchQuery(query);
if (!parsed.hasDirectives) return query;
const demoted = [...parsed.sites.filter(site => site.includes("/")), ...parsed.inUrl];
const downgraded: StructuredQuery = {
...parsed,
sites: parsed.sites.filter(site => !site.includes("/")),
inUrl: [],
terms: [...parsed.terms, ...demoted.map(text => ({ text }))],
};
return formatQuery(downgraded, syntax);
}
/** Hostname (lowercased) and pathname of a URL, or undefined when unparsable. */
function hostAndPath(url: string): { host: string; path: string } | undefined {
try {
const u = new URL(url);
return { host: u.hostname.toLowerCase(), path: u.pathname };
} catch {
return undefined;
}
}
/**
* `site:` matcher: exact host or subdomain of `site`; when `site` carries a
* path (`github.com/anthropics`), the URL path must start with it.
*/
export function matchesSite(url: string, site: string): boolean {
const parsed = hostAndPath(url);
if (!parsed) return false;
const slash = site.indexOf("/");
const siteHost = slash === -1 ? site : site.slice(0, slash);
const sitePath = slash === -1 ? "" : site.slice(slash);
if (parsed.host !== siteHost && !parsed.host.endsWith(`.${siteHost}`)) return false;
if (sitePath && !parsed.path.toLowerCase().startsWith(sitePath.toLowerCase())) return false;
return true;
}
/** `filetype:` matcher: URL pathname ends with `.ext`. */
function matchesFiletype(url: string, ext: string): boolean {
const parsed = hostAndPath(url);
if (!parsed) return false;
return parsed.path.toLowerCase().endsWith(`.${ext}`);
}
const RELATIVE_AGE_PATTERN = /^(\d+)\s*(minute|min|hour|hr|day|week|month|mo|year|yr|[mhdwy])s?\s+ago$/i;
const RELATIVE_UNIT_SECONDS: Record<string, number> = {
m: 60,
min: 60,
minute: 60,
h: 3600,
hr: 3600,
hour: 3600,
d: 86_400,
day: 86_400,
w: 604_800,
week: 604_800,
mo: 2_592_000,
month: 2_592_000,
y: 31_536_000,
yr: 31_536_000,
year: 31_536_000,
};
/** Best-effort publish time (ms epoch) of a source from `ageSeconds`, ISO, or relative dates. */
function sourceTime(source: SearchSource): number | undefined {
if (typeof source.ageSeconds === "number" && Number.isFinite(source.ageSeconds)) {
return Date.now() - source.ageSeconds * 1000;
}
if (!source.publishedDate) return undefined;
const rel = RELATIVE_AGE_PATTERN.exec(source.publishedDate.trim());
if (rel) {
const seconds = Number(rel[1]) * (RELATIVE_UNIT_SECONDS[rel[2].toLowerCase()] ?? 0);
return seconds > 0 ? Date.now() - seconds * 1000 : undefined;
}
const parsed = Date.parse(source.publishedDate);
return Number.isNaN(parsed) ? undefined : parsed;
}
/**
* Strict per-source constraint check: every filterable dimension of `q` must
* pass. Sources without a resolvable date pass date bounds (a missing date
* is not proof of violation). For custom provider flows; the standard path
* is {@link applyQueryConstraints}.
*/
export function matchesQueryConstraints(source: SearchSource, q: StructuredQuery): boolean {
for (const dim of constraintDimensions(q)) {
if (!dim.pred(source)) return false;
}
return true;
}
interface ConstraintDimension {
/** Directive rendering for relaxation notes (`site:arxiv.org`). */
label: string;
pred: (source: SearchSource) => boolean;
}
function constraintDimensions(q: StructuredQuery): ConstraintDimension[] {
const dims: ConstraintDimension[] = [];
const lower = (s: string | undefined): string => (s ?? "").toLowerCase();
if (q.sites.length > 0) {
dims.push({
label: q.sites.map(s => `site:${s}`).join(" OR "),
pred: src => q.sites.some(site => matchesSite(src.url, site)),
});
}
if (q.excludedSites.length > 0) {
dims.push({
label: q.excludedSites.map(s => `-site:${s}`).join(" "),
pred: src => !q.excludedSites.some(site => matchesSite(src.url, site)),
});
}
if (q.inUrl.length > 0) {
dims.push({
label: q.inUrl.map(v => `inurl:${v}`).join(" "),
pred: src => q.inUrl.every(v => lower(src.url).includes(v.toLowerCase())),
});
}
if (q.excludedInUrl.length > 0) {
dims.push({
label: q.excludedInUrl.map(v => `-inurl:${v}`).join(" "),
pred: src => !q.excludedInUrl.some(v => lower(src.url).includes(v.toLowerCase())),
});
}
if (q.inTitle.length > 0) {
dims.push({
label: q.inTitle.map(v => `intitle:${v}`).join(" "),
pred: src => q.inTitle.every(v => lower(src.title).includes(v.toLowerCase())),
});
}
if (q.excludedInTitle.length > 0) {
dims.push({
label: q.excludedInTitle.map(v => `-intitle:${v}`).join(" "),
pred: src => !q.excludedInTitle.some(v => lower(src.title).includes(v.toLowerCase())),
});
}
if (q.filetypes.length > 0) {
dims.push({
label: q.filetypes.map(f => `filetype:${f}`).join(" OR "),
pred: src => q.filetypes.some(ext => matchesFiletype(src.url, ext)),
});
}
if (q.excludedFiletypes.length > 0) {
dims.push({
label: q.excludedFiletypes.map(f => `-filetype:${f}`).join(" "),
pred: src => !q.excludedFiletypes.some(ext => matchesFiletype(src.url, ext)),
});
}
if (q.after !== undefined || q.before !== undefined) {
const afterMs = q.after !== undefined ? Date.parse(q.after) : undefined;
const beforeMs = q.before !== undefined ? Date.parse(q.before) : undefined;
const label = [q.after ? `after:${q.after}` : "", q.before ? `before:${q.before}` : ""].filter(Boolean).join(" ");
dims.push({
label,
pred: src => {
const time = sourceTime(src);
if (time === undefined) return true; // undated → cannot prove violation
if (afterMs !== undefined && time < afterMs) return false;
if (beforeMs !== undefined && time >= beforeMs) return false;
return true;
},
});
}
return dims;
}
/**
* Lenient post-filter: applies each constraint dimension of `q` in turn,
* skipping (and reporting) any dimension that would eliminate every
* remaining source. Guarantees a non-empty result for a non-empty input, so
* a mis-scoped directive degrades to unfiltered results plus a note instead
* of a dead search.
*/
export function applyQueryConstraints(sources: readonly SearchSource[], q: StructuredQuery): ConstraintFilterResult {
let current = [...sources];
const dropped: string[] = [];
if (current.length === 0) return { sources: current, dropped };
for (const dim of constraintDimensions(q)) {
const kept = current.filter(dim.pred);
if (kept.length > 0) current = kept;
else dropped.push(dim.label);
}
return { sources: current, dropped };
}
@@ -275,6 +275,39 @@ describe("searchCodex model selection", () => {
expect(result.sources).toEqual([{ title: "Example Article", url: "https://example.com/article" }]);
});
function sentUserText(): string | undefined {
const input = capturedRequest?.body?.input as Array<Record<string, unknown>> | undefined;
const userItem = input?.find(item => item.role === "user");
const content = userItem?.content as Array<Record<string, unknown>> | undefined;
return content?.[0]?.text as string | undefined;
}
it("re-emits directive queries with normalized Google-style operators", async () => {
delete process.env.PI_CODEX_WEB_SEARCH_MODEL;
await searchCodex(
makeSearchParams(
'bun runtime site:bun.sh -site:reddit.com after:2024-01-01 "exact phrase"',
mockCodexFetch("gpt-5.6-luna"),
),
);
expect(capturedRequest).not.toBeNull();
expect(sentUserText()).toBe('bun runtime "exact phrase" site:bun.sh -site:reddit.com after:2024-01-01');
// Tool config stays untouched: the ChatGPT backend's filter support is
// unverified, so no `filters` field is added to the web_search tool.
const input = capturedRequest?.body?.input as Array<Record<string, unknown>>;
const additionalTools = input.find(item => item.type === "additional_tools");
expect(additionalTools?.tools).toEqual([{ type: "web_search", search_context_size: "high" }]);
});
it("sends directive-free queries byte-identical", async () => {
delete process.env.PI_CODEX_WEB_SEARCH_MODEL;
const query = "how does the bun runtime schedule timers?";
await searchCodex(makeSearchParams(query, mockCodexFetch("gpt-5.6-luna")));
expect(sentUserText()).toBe(query);
});
it("uses configured Codex endpoint, API key, and headers without OAuth", async () => {
process.env.PI_CODEX_WEB_SEARCH_MODEL = "gpt-5.4";
const result = await searchCodex({
@@ -297,6 +297,38 @@ describe("searchExa", () => {
contents: { summary: { query: "shape test" } },
});
});
it("maps site:/before: directives to native Exa params with an operator-free query", async () => {
await withLocalAuthStorage(authStorage =>
new ExaProvider().search({
query: "vector db benchmarks site:qdrant.tech before:2025-01-01",
systemPrompt: "",
authStorage,
fetch: mockFetch(makeMockExaResponse()),
}),
);
expect(capturedRequestBody!.query).toBe("vector db benchmarks");
expect(capturedRequestBody!.includeDomains).toEqual(["qdrant.tech"]);
expect(capturedRequestBody!.endPublishedDate).toBe("2025-01-01");
expect(capturedRequestBody!.startPublishedDate).toBeUndefined();
expect(capturedRequestBody!.excludeDomains).toBeUndefined();
});
it("sends directive-free queries byte-identical with no domain/date params", async () => {
await withLocalAuthStorage(authStorage =>
new ExaProvider().search({
query: "plain natural language question",
systemPrompt: "",
authStorage,
fetch: mockFetch(makeMockExaResponse()),
}),
);
expect(capturedRequestBody).toEqual({
query: "plain natural language question",
numResults: 10,
type: "auto",
contents: { summary: { query: "plain natural language question" } },
});
});
it("paces consecutive Exa API requests by the configured delay", async () => {
resetSettingsForTest();
@@ -104,6 +104,54 @@ describe("Firecrawl web search provider", () => {
authMode: "api_key",
});
});
it("maps before:/after: to a cdr tbs and strips dates from the operator query", async () => {
const captured: { body?: unknown } = {};
const fetchMock: FetchImpl = async (_input, init) => {
captured.body = JSON.parse(String(init?.body ?? "null")) as unknown;
return new Response(JSON.stringify({ data: { web: [] } }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
};
await searchFirecrawl({
...makeParams("bun runtime site:github.com/oven-sh intitle:install after:2024-01-01 before:2024-06-30"),
recency: "month",
fetch: fetchMock,
});
expect(captured.body).toEqual({
query: "bun runtime site:github.com/oven-sh intitle:install",
limit: 10,
sources: [{ type: "web" }],
// Explicit absolute bounds take precedence over the qdr:m recency window.
tbs: "cdr:1,cd_min:01/01/2024,cd_max:06/30/2024",
});
});
it("re-emits non-date operators in the query while keeping recency tbs", async () => {
const captured: { body?: unknown } = {};
const fetchMock: FetchImpl = async (_input, init) => {
captured.body = JSON.parse(String(init?.body ?? "null")) as unknown;
return new Response(JSON.stringify({ data: { web: [] } }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
};
await searchFirecrawl({
...makeParams('"exact phrase" -site:reddit.com filetype:pdf'),
recency: "week",
fetch: fetchMock,
});
expect(captured.body).toEqual({
query: '"exact phrase" -site:reddit.com filetype:pdf',
limit: 10,
sources: [{ type: "web" }],
tbs: "qdr:w",
});
});
it("uses the initially resolved credential for the first authenticated request", async () => {
let resolutionCount = 0;
@@ -111,6 +111,34 @@ describe("searchGemini tools serialization", () => {
});
});
it("normalizes query directive aliases to canonical Google forms in the grounding request", async () => {
const fetchMock = mockGeminiFetch();
await searchGemini({
...makeParams("k8s domain:kubernetes.io since:2024"),
fetch: fetchMock,
});
expect(capturedRequest).not.toBeNull();
const request = capturedRequest?.body?.request as Record<string, unknown>;
expect(request).toMatchObject({
contents: [{ role: "user", parts: [{ text: "k8s site:kubernetes.io after:2024-01-01" }] }],
});
});
it("leaves directive-free queries untouched in the developer API request", async () => {
const fetchMock = mockGeminiFetch(DEVELOPER_SSE_RESPONSE);
await searchGemini({
...makeParams("plain query with no operators"),
authStorage: apiKeyAuthStorage,
fetch: fetchMock,
});
expect(capturedRequest).not.toBeNull();
expect(capturedRequest?.body).toMatchObject({
contents: [{ role: "user", parts: [{ text: "plain query with no operators" }] }],
});
});
it("uses configured developer API model and reports it when modelVersion is absent", async () => {
const fetchMock = mockGeminiFetch(DEVELOPER_SSE_RESPONSE_WITHOUT_MODEL);
const response = await searchGemini({
@@ -198,6 +198,32 @@ describe("Kagi search result parsing", () => {
expect(requestBody?.filters?.after).toBe(expected);
});
it("canonicalizes Google-style directives into the upstream query", async () => {
let requestBody: KagiSearchRequest | undefined;
const fetchMock: FetchImpl = async (input, init) => {
const urlStr = typeof input === "string" ? input : input instanceof URL ? input.toString() : input.url;
if (urlStr === "https://kagi.com/api/v1/search") {
requestBody = JSON.parse(init?.body as string) as KagiSearchRequest;
return new Response(JSON.stringify({ meta: { trace: "req-directives" }, data: { search: [] } }), {
status: 200,
headers: { "Content-Type": "application/json" },
});
}
return new Response("not mocked", { status: 500 });
};
await searchKagi({
query: "x domain:kagi.com until:2025",
authStorage: fakeAuthStorage,
fetch: fetchMock,
});
expect(requestBody?.query).toBe("x site:kagi.com before:2025-01-01");
await searchKagi({ query: "plain kagi query", authStorage: fakeAuthStorage, fetch: fetchMock });
expect(requestBody?.query).toBe("plain kagi query");
});
it("uses a Bearer authorization header", async () => {
let capturedAuth: string | null = null;
const fetchMock: FetchImpl = async (input, init) => {
@@ -110,3 +110,50 @@ describe("searchKimi credential resolution", () => {
});
});
});
describe("searchKimi query directives", () => {
afterEach(() => {
restoreSearchApiKeyEnv();
vi.restoreAllMocks();
});
function captureBodyFetch(capture: { body?: { text_query?: string } }): FetchImpl {
return (_url, init) => {
capture.body = JSON.parse(String(init?.body)) as { text_query?: string };
return Promise.resolve(
new Response(JSON.stringify({ search_results: [] }), {
status: 200,
headers: { "Content-Type": "application/json" },
}),
);
};
}
it("sends directive-free queries upstream unchanged", async () => {
process.env.KIMI_SEARCH_API_KEY = "test-key";
const capture: { body?: { text_query?: string } } = {};
await withLocalAuthStorage(async authStorage => {
await searchKimi({ query: "bun test runner docs", authStorage, fetch: captureBodyFetch(capture) });
});
expect(capture.body?.text_query).toBe("bun test runner docs");
});
it("rebuilds directive queries with Bing-style operators and drops date directives", async () => {
process.env.KIMI_SEARCH_API_KEY = "test-key";
const capture: { body?: { text_query?: string } } = {};
await withLocalAuthStorage(async authStorage => {
await searchKimi({
query: 'site:github.com intitle:changelog filetype:md after:2024-01-01 "bun runtime" -deprecated',
authStorage,
fetch: captureBodyFetch(capture),
});
});
const sent = capture.body?.text_query ?? "";
expect(sent).toContain("site:github.com");
expect(sent).toContain("intitle:changelog");
expect(sent).toContain("filetype:md");
expect(sent).toContain('"bun runtime"');
expect(sent).toContain("-deprecated");
expect(sent).not.toContain("after:");
});
});
@@ -91,6 +91,19 @@ describe("Mojeek web search provider", () => {
expect(url.searchParams.get("since")).toBeNull();
});
it("re-emits supported operators (phrases, -, site:) and strips unsupported date bounds from the query", async () => {
let capturedUrl = "";
const fetchMock: FetchImpl = input => {
capturedUrl = typeof input === "string" ? input : input.toString();
return Promise.resolve(new Response(resultsPage(""), { status: 200 }));
};
await searchMojeek(makeParams("independent index site:mojeek.com after:2024", fetchMock));
const url = new URL(capturedUrl);
expect(url.searchParams.get("q")).toBe("independent index site:mojeek.com");
});
it("parses result rows, deduplicates targets, and skips junk and intra-Mojeek rows", async () => {
const html = resultsPage(
[
@@ -131,6 +131,45 @@ describe("Parallel web search", () => {
]);
});
it("maps site: directives onto source_policy.include_domains and strips them from the query", async () => {
const fetchMock = mockFetch({
search_id: "search-parallel-3",
results: [],
warnings: null,
usage: null,
});
await searchParallel({ query: "web api site:parallel.ai", fetch: fetchMock }, fakeAuthStorage);
expect(capturedRequestBody).toEqual({
objective: "web api",
search_queries: ["web api"],
mode: "fast",
excerpts: { max_chars_per_result: 10_000 },
source_policy: { include_domains: ["parallel.ai"] },
});
});
it("maps -site: and after: onto exclude_domains/after_date, keeping phrases and negation", async () => {
const fetchMock = mockFetch({
search_id: "search-parallel-4",
results: [],
warnings: null,
usage: null,
});
await searchParallel(
{ query: '"web api" -legacy -site:reddit.com/r/node after:2025-06-01', fetch: fetchMock },
fakeAuthStorage,
);
expect(capturedRequestBody).toEqual({
objective: '"web api" -legacy',
search_queries: ['"web api" -legacy'],
mode: "fast",
excerpts: { max_chars_per_result: 10_000 },
source_policy: { exclude_domains: ["reddit.com"], after_date: "2025-06-01" },
});
});
it("surfaces plain-text Parallel API errors", async () => {
const fetchMock: FetchImpl = () => Promise.resolve(new Response("upstream unavailable", { status: 503 }));
await expect(searchParallel({ query: "broken", fetch: fetchMock }, fakeAuthStorage)).rejects.toMatchObject({
@@ -59,6 +59,68 @@ describe("SearXNG web search provider", () => {
});
});
it("demotes engine-hostile operators while keeping bare-domain site: filters", async () => {
process.env.SEARXNG_ENDPOINT = "https://searx.example.org";
const captured: { q?: string | null } = {};
const fetchMock: FetchImpl = input => {
captured.q = new URL(input.toString()).searchParams.get("q");
return Promise.resolve(
new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com" }] }), {
status: 200,
headers: { "Content-Type": "application/json" },
}),
);
};
await searchSearXNG({
query: "site:github.com/can1357/oh-my-pi inurl:releases site:github.com 17.1.1 release",
fetch: fetchMock,
});
expect(captured.q).toBe("17.1.1 release github.com/can1357/oh-my-pi releases site:github.com");
});
it("maps lang: to the language param and re-emits remaining directives in q", async () => {
process.env.SEARXNG_ENDPOINT = "https://searx.example.org";
const captured: { url?: URL } = {};
const fetchMock: FetchImpl = input => {
captured.url = new URL(input.toString());
return Promise.resolve(
new Response(JSON.stringify({ results: [{ title: "r", url: "https://searxng.org/docs" }] }), {
status: 200,
headers: { "Content-Type": "application/json" },
}),
);
};
await searchSearXNG({ query: "docs lang:de site:searxng.org", fetch: fetchMock });
expect(captured.url?.searchParams.get("q")).toBe("docs site:searxng.org");
expect(captured.url?.searchParams.get("language")).toBe("de");
});
it("sends directive-free queries verbatim without a language param", async () => {
process.env.SEARXNG_ENDPOINT = "https://searx.example.org";
const captured: { url?: URL } = {};
const fetchMock: FetchImpl = input => {
captured.url = new URL(input.toString());
return Promise.resolve(
new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com" }] }), {
status: 200,
headers: { "Content-Type": "application/json" },
}),
);
};
await searchSearXNG({ query: "plain metasearch query", fetch: fetchMock });
expect(captured.url?.searchParams.get("q")).toBe("plain metasearch query");
expect(captured.url?.searchParams.get("language")).toBeNull();
});
it("reads Basic auth credentials from nested config.yml settings", async () => {
const agentDir = await fs.mkdtemp(path.join(os.tmpdir(), "searxng-settings-"));
try {
@@ -231,6 +293,107 @@ describe("SearXNG web search provider", () => {
expect(captured.headers?.get("Authorization")).toBe("Bearer bearer-token");
});
it("resolves engine shortcuts via /config into canonical names for the engines parameter", async () => {
const agentDir = await fs.mkdtemp(path.join(os.tmpdir(), "searxng-engines-"));
try {
await Bun.write(
path.join(agentDir, "config.yml"),
[
"searxng:",
" endpoint: https://searx-shortcuts.example.org",
' engines: "ddg, Brave, unknown"',
"",
].join("\n"),
);
await Settings.init({ agentDir });
const requested: URL[] = [];
const fetchMock: FetchImpl = input => {
const url = new URL(input.toString());
requested.push(url);
if (url.pathname === "/config") {
return Promise.resolve(
new Response(
JSON.stringify({
engines: [
{ name: "duckduckgo", shortcut: "ddg" },
{ name: "brave", shortcut: "br" },
],
}),
{ status: 200, headers: { "Content-Type": "application/json" } },
),
);
}
return Promise.resolve(
new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com/r" }] }), {
status: 200,
headers: { "Content-Type": "application/json" },
}),
);
};
await searchSearXNG({ query: "engine selection", fetch: fetchMock });
const searchUrl = requested.find(url => url.pathname === "/search");
expect(searchUrl?.searchParams.get("engines")).toBe("duckduckgo,brave,unknown");
} finally {
await removeWithRetries(agentDir);
}
});
it("passes configured engines verbatim when /config is unavailable", async () => {
const agentDir = await fs.mkdtemp(path.join(os.tmpdir(), "searxng-engines-fallback-"));
try {
await Bun.write(
path.join(agentDir, "config.yml"),
["searxng:", " endpoint: https://searx-noconfig.example.org", ' engines: "ddg,brave"', ""].join("\n"),
);
await Settings.init({ agentDir });
const requested: URL[] = [];
const fetchMock: FetchImpl = input => {
const url = new URL(input.toString());
requested.push(url);
if (url.pathname === "/config") {
return Promise.resolve(new Response("forbidden", { status: 403 }));
}
return Promise.resolve(
new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com/r" }] }), {
status: 200,
headers: { "Content-Type": "application/json" },
}),
);
};
await searchSearXNG({ query: "fallback engines", fetch: fetchMock });
const searchUrl = requested.find(url => url.pathname === "/search");
expect(searchUrl?.searchParams.get("engines")).toBe("ddg,brave");
} finally {
await removeWithRetries(agentDir);
}
});
it("strips external bang tokens but keeps engine bangs in the query", async () => {
process.env.SEARXNG_ENDPOINT = "https://searx-bangs.example.org";
const captured: { url?: URL } = {};
const fetchMock: FetchImpl = input => {
captured.url = new URL(input.toString());
return Promise.resolve(
new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com/r" }] }), {
status: 200,
headers: { "Content-Type": "application/json" },
}),
);
};
await searchSearXNG({ query: "!!g rust !ddg lifetimes", fetch: fetchMock });
expect(captured.url?.pathname).toBe("/search");
expect(captured.url?.searchParams.get("q")).toBe("rust !ddg lifetimes");
});
it("treats empty SearXNG results with upstream failures as a provider error", async () => {
process.env.SEARXNG_ENDPOINT = "https://searx.example.org";
@@ -65,6 +65,48 @@ function expectTinyFishParams(url: URL, expectedParams: readonly string[]): void
}
describe("TinyFish web search provider", () => {
it("rewrites directive queries with supported operators only", async () => {
const captured: URL[] = [];
const fetchMock: FetchImpl = async input => {
const url = input instanceof URL ? input : new URL(typeof input === "string" ? input : input.url);
captured.push(url);
return new Response(JSON.stringify(tinyFishPage(tinyFishResults("tinyfish", 3))), {
status: 200,
headers: { "Content-Type": "application/json" },
});
};
await searchTinyFish({
...makeParams(
'"error handling" rust site:github.com -site:gitlab.com filetype:pdf intitle:tokio after:2024-01-01',
),
fetch: fetchMock,
});
expect(captured).toHaveLength(1);
expect(captured[0].searchParams.get("query")).toBe(
'"error handling" rust site:github.com -site:gitlab.com filetype:pdf',
);
expectTinyFishParams(captured[0], ["query", "num_results", "page"]);
});
it("sends directive-free queries verbatim", async () => {
const captured: URL[] = [];
const fetchMock: FetchImpl = async input => {
const url = input instanceof URL ? input : new URL(typeof input === "string" ? input : input.url);
captured.push(url);
return new Response(JSON.stringify(tinyFishPage(tinyFishResults("tinyfish", 3))), {
status: 200,
headers: { "Content-Type": "application/json" },
});
};
await searchTinyFish({ ...makeParams("plain query with ordinary words"), fetch: fetchMock });
expect(captured).toHaveLength(1);
expect(captured[0].searchParams.get("query")).toBe("plain query with ordinary words");
});
it("passes TinyFish num_results and applies numSearchResults across pages", async () => {
const captured: { url: URL; init?: RequestInit }[] = [];
const pages = new Map([
@@ -163,6 +163,48 @@ describe("xAI web search provider", () => {
expect(capture.capturedRequest?.body).not.toHaveProperty("search_parameters");
});
it("maps site: onto web_search allowed_domains and strips it from the query", async () => {
const capture = captureFetch({ id: "resp_directives", model: "grok-4.3", output_text: "directive answer" });
await searchXAI({
...makeParams(capture.fetchMock),
query: "grok api site:docs.x.ai after:2025-01-01",
});
const body = capture.capturedRequest?.body;
expect(body?.tools).toEqual([{ type: "web_search", filters: { allowed_domains: ["docs.x.ai"] } }]);
// The Responses web_search tool has no from_date/to_date, so the date
// bound stays in the query text for the agent while site: is stripped.
const input = body?.input as { role: string; content: string }[];
expect(input[1]?.content).toBe("grok api after:2025-01-01");
});
it("maps -site: onto excluded_domains as bare hosts only when no allow list is present", async () => {
const capture = captureFetch({ id: "resp_excludes", model: "grok-4.3", output_text: "exclude answer" });
await searchXAI({
...makeParams(capture.fetchMock),
query: "grok changelog -site:reddit.com/r/grok -site:news.ycombinator.com",
});
const body = capture.capturedRequest?.body;
expect(body?.tools).toEqual([
{ type: "web_search", filters: { excluded_domains: ["reddit.com", "news.ycombinator.com"] } },
]);
const input = body?.input as { role: string; content: string }[];
expect(input[1]?.content).toBe("grok changelog");
await searchXAI({
...makeParams(capture.fetchMock),
query: "grok changelog site:docs.x.ai -site:reddit.com",
});
// allowed_domains and excluded_domains are mutually exclusive per
// request: the allow list wins, exclusions fall to the central filter.
expect(capture.capturedRequest?.body?.tools).toEqual([
{ type: "web_search", filters: { allowed_domains: ["docs.x.ai"] } },
]);
});
it("uses dedicated xAI OAuth credentials for Responses API bearer auth", async () => {
const capture = captureFetch({ id: "resp_xai_oauth", model: "grok-4.3", output_text: "xAI OAuth answer" });
@@ -75,4 +75,56 @@ describe("Anthropic search request body", () => {
expect(userId.account_uuid).toBe(accountUuid);
expect(userId.device_id).toMatch(/^[0-9a-f]{64}$/);
});
it("maps site: to allowed_domains and strips the directive from the query", async () => {
using tempDir = TempDir.createSync("@pi-anthropic-search-sites-");
const authStorage = await CodingAuthStorage.create(path.join(tempDir.path(), "auth.db"));
try {
authStorage.setRuntimeApiKey("anthropic", "test-key");
const cap = makeCaptureFetch();
await searchAnthropic({
query: "sdk docs site:docs.anthropic.com",
systemPrompt: "Use web search.",
sessionId: "session-2295",
authStorage,
fetch: cap.fetch,
});
const body = cap.body();
const tool = (body?.tools as Record<string, unknown>[] | undefined)?.[0];
expect(tool?.allowed_domains).toEqual(["docs.anthropic.com"]);
expect(tool).not.toHaveProperty("blocked_domains");
const messages = body?.messages as { content: string }[];
expect(messages[0]?.content).toBe("sdk docs");
} finally {
authStorage.close();
}
});
it("maps -site: to blocked_domains when there are no site includes", async () => {
using tempDir = TempDir.createSync("@pi-anthropic-search-blocked-");
const authStorage = await CodingAuthStorage.create(path.join(tempDir.path(), "auth.db"));
try {
authStorage.setRuntimeApiKey("anthropic", "test-key");
const cap = makeCaptureFetch();
await searchAnthropic({
query: "rust async runtime -site:reddit.com",
systemPrompt: "Use web search.",
sessionId: "session-2295",
authStorage,
fetch: cap.fetch,
});
const body = cap.body();
const tool = (body?.tools as Record<string, unknown>[] | undefined)?.[0];
expect(tool?.blocked_domains).toEqual(["reddit.com"]);
expect(tool).not.toHaveProperty("allowed_domains");
const messages = body?.messages as { content: string }[];
expect(messages[0]?.content).toBe("rust async runtime");
} finally {
authStorage.close();
}
});
});
@@ -118,6 +118,22 @@ describe("Perplexity API-key request shape", () => {
expect(body?.num_search_results).toBe(5);
});
it("maps site:/-site:/after: directives onto native filters and strips them from the query", async () => {
let body: Record<string, unknown> | undefined;
const fetchMock = mockApi(b => (body = b), baseResponse());
await searchPerplexity({
query: "rust site:docs.rs -site:reddit.com after:2024-06-01",
authStorage: apiKeyAuthStorage,
fetch: fetchMock,
});
expect(body?.search_domain_filter).toEqual(["docs.rs", "-reddit.com"]);
expect(body?.search_after_date_filter).toBe("6/1/2024");
const messages = body?.messages as { role: string; content: string }[];
expect(messages.at(-1)?.content).toBe("rust");
});
it("parses related_questions into relatedQuestions, preserving order and dropping blanks", async () => {
const fetchMock = mockApi(
() => {},
@@ -380,6 +396,26 @@ describe("Perplexity OAuth request shape", () => {
expect(response.authMode).toBe("oauth");
expect(response.answer).toBe("OAuth answer");
});
it("maps directives onto ask-endpoint native filters and rewrites query_str", async () => {
let body: Record<string, unknown> | undefined;
const fetchMock = mockOAuth(b => (body = b));
await searchPerplexity({
query: "rust site:docs.rs -site:reddit.com after:2024-06-01",
search_recency_filter: "month",
authStorage: oauthAuthStorage,
fetch: fetchMock,
});
expect(body!.query_str).toBe("rust");
const params = body!.params as Record<string, unknown>;
expect(params.query_str).toBe("rust");
expect(params.search_domain_filter).toEqual(["docs.rs", "-reddit.com"]);
expect(params.search_after_date_filter).toBe("6/1/2024");
// Absolute date bounds take precedence over recency.
expect(params.search_recency_filter).toBeNull();
});
});
describe("Perplexity OAuth transport failure (issue #5315)", () => {
@@ -0,0 +1,70 @@
/**
* Central directive pipeline: executeSearch parses the query once, hands the
* StructuredQuery to the provider, then lenient-filters the returned sources
* — enforcing constraints the provider ignored and relaxing (with a note)
* any dimension that would eliminate every result.
*/
import { afterEach, describe, expect, it, vi } from "bun:test";
import type { AuthStorage } from "@oh-my-pi/pi-ai";
import { runSearchQuery } from "@oh-my-pi/pi-coding-agent/web/search";
import type { SearchParams } from "@oh-my-pi/pi-coding-agent/web/search/provider";
import * as provider from "@oh-my-pi/pi-coding-agent/web/search/provider";
import type { SearchProviderId, SearchResponse, SearchSource } from "@oh-my-pi/pi-coding-agent/web/search/types";
const SOURCES: SearchSource[] = [
{ title: "Docs page", url: "https://docs.example.com/guide" },
{ title: "Blog post", url: "https://blog.other.com/post" },
];
function stubProvider(id: SearchProviderId, behaviour: (params: SearchParams) => Promise<SearchResponse>) {
const stub: provider.SearchProvider = {
id,
label: id,
isAvailable: () => true,
isExplicitlyAvailable: () => true,
search: behaviour,
};
vi.spyOn(provider, "resolveProviderCandidates").mockReturnValue([{ id, explicit: true }]);
vi.spyOn(provider, "getSearchProvider").mockImplementation(async requested => {
if (requested !== id) throw new Error(`Unexpected provider: ${requested}`);
return stub;
});
}
describe("web search directive pipeline", () => {
afterEach(() => vi.restoreAllMocks());
it("passes the parsed query to the provider and post-filters sources it did not constrain", async () => {
let seen: SearchParams | undefined;
stubProvider("brave", async params => {
seen = params;
return { provider: "brave", sources: SOURCES };
});
const result = await runSearchQuery(
{ query: "guide site:docs.example.com", provider: "brave" },
{ authStorage: {} as AuthStorage },
);
expect(seen?.parsedQuery?.sites).toEqual(["docs.example.com"]);
expect(seen?.parsedQuery?.text).toBe("guide");
expect(result.details.response.sources.map(s => s.url)).toEqual(["https://docs.example.com/guide"]);
expect(result.content[0]?.text).not.toContain("Note:");
});
it("relaxes a constraint that matches nothing and leads the LLM text with a note", async () => {
stubProvider("brave", async () => ({ provider: "brave", sources: SOURCES }));
const result = await runSearchQuery(
{ query: "guide site:nowhere.example", provider: "brave" },
{ authStorage: {} as AuthStorage },
);
// Leniency: nothing matched site:nowhere.example, so all sources survive
// and the model is told the constraint was relaxed.
expect(result.details.response.sources).toHaveLength(SOURCES.length);
expect(result.content[0]?.text).toStartWith(
"Note: no results matched `site:nowhere.example`; the constraint was relaxed",
);
});
});
@@ -0,0 +1,298 @@
import { describe, expect, it } from "bun:test";
import {
applyQueryConstraints,
formatQuery,
formatScraperQuery,
GOOGLE_QUERY_SYNTAX,
matchesQueryConstraints,
matchesSite,
parseDateValue,
parseSearchQuery,
} from "@oh-my-pi/pi-coding-agent/web/search/query";
import type { SearchSource } from "@oh-my-pi/pi-coding-agent/web/search/types";
describe("parseSearchQuery", () => {
it("leaves plain queries untouched", () => {
const q = parseSearchQuery("rust async runtime comparison");
expect(q.text).toBe("rust async runtime comparison");
expect(q.hasDirectives).toBe(false);
expect(q.hasConstraints).toBe(false);
expect(q.terms.map(t => t.text)).toEqual(["rust", "async", "runtime", "comparison"]);
});
it("keeps unknown colon tokens (URLs, paths, error codes) as text", () => {
const q = parseSearchQuery("error TS2345: https://example.com/a?b=c C:\\Users\\me");
expect(q.hasConstraints).toBe(false);
expect(q.terms.map(t => t.text)).toEqual(["error", "TS2345:", "https://example.com/a?b=c", "C:\\Users\\me"]);
});
it("parses site: aliases and exclusions with normalization", () => {
const q = parseSearchQuery(
"kubernetes site:HTTPS://Docs.K8s.IO/ domain:cncf.io -site:*.reddit.com host:github.com/kubernetes",
);
expect(q.sites).toEqual(["docs.k8s.io", "cncf.io", "github.com/kubernetes"]);
expect(q.excludedSites).toEqual(["reddit.com"]);
expect(q.text).toBe("kubernetes");
expect(q.hasConstraints).toBe(true);
});
it("adopts the next token when a directive value is space-separated", () => {
const q = parseSearchQuery("site: arxiv.org transformer scaling");
expect(q.sites).toEqual(["arxiv.org"]);
expect(q.text).toBe("transformer scaling");
});
it("does not adopt operators or directives as space-separated values", () => {
const q = parseSearchQuery("site: OR site:example.com");
expect(q.sites).toEqual(["example.com"]);
});
it("parses inurl/intitle/intext variants including quoted values", () => {
const q = parseSearchQuery('inurl:docs intitle:"getting started" -intitle:deprecated inbody:websocket handshake');
expect(q.inUrl).toEqual(["docs"]);
expect(q.inTitle).toEqual(["getting started"]);
expect(q.excludedInTitle).toEqual(["deprecated"]);
expect(q.inText).toEqual(["websocket"]);
expect(q.text).toBe("handshake");
});
it("routes every following term for allintitle:", () => {
const q = parseSearchQuery("allintitle: budget planning tips");
expect(q.inTitle).toEqual(["budget", "planning", "tips"]);
expect(q.text).toBe("");
});
it("parses filetype/ext with normalization", () => {
const q = parseSearchQuery("quarterly report filetype:PDF ext:.xlsx -filetype:doc");
expect(q.filetypes).toEqual(["pdf", "xlsx"]);
expect(q.excludedFiletypes).toEqual(["doc"]);
});
it("parses date bounds in common forms", () => {
const q = parseSearchQuery("llm evals after:2024 before:2025-06-15");
expect(q.after).toBe("2024-01-01");
expect(q.before).toBe("2025-06-15");
expect(q.text).toBe("llm evals");
expect(parseSearchQuery("x since:2023/07/01").after).toBe("2023-07-01");
expect(parseSearchQuery("x until:2024-02").before).toBe("2024-02-01");
expect(parseSearchQuery("x after:6/15/2024").after).toBe("2024-06-15");
expect(parseSearchQuery("x after:15/6/2024").after).toBe("2024-06-15");
});
it("degrades unparseable date values to plain terms", () => {
const q = parseSearchQuery("release before:soon");
expect(q.before).toBeUndefined();
expect(q.terms.map(t => t.text)).toEqual(["release", "before:soon"]);
});
it("parses quoted phrases including smart quotes and negated phrases", () => {
const q = parseSearchQuery('"exact phrase" \u201csmart quoted\u201d -"not this"');
expect(q.terms).toEqual([
{ text: "exact phrase", phrase: true },
{ text: "smart quoted", phrase: true },
{ text: "not this", phrase: true, negated: true },
]);
expect(q.text).toBe('"exact phrase" "smart quoted" -"not this"');
});
it("groups OR alternatives and treats AND as default conjunction", () => {
const q = parseSearchQuery("(react OR vue OR svelte) AND hooks");
const groups = q.terms.map(t => t.group);
expect(q.terms.map(t => t.text)).toEqual(["react", "vue", "svelte", "hooks"]);
expect(groups[0]).toBeDefined();
expect(groups[0]).toBe(groups[1]);
expect(groups[1]).toBe(groups[2]);
expect(groups[3]).toBeUndefined();
expect(q.text).toBe("(react OR vue OR svelte) hooks");
});
it("supports pipe as OR and NOT as negation", () => {
const q = parseSearchQuery("deno | bun NOT node");
expect(q.terms[0].group).toBe(q.terms[1].group);
expect(q.terms[2]).toMatchObject({ text: "node", negated: true });
});
it("swallows OR between directives instead of grouping terms", () => {
const q = parseSearchQuery("caching site:redis.io OR site:memcached.org");
expect(q.sites).toEqual(["redis.io", "memcached.org"]);
expect(q.terms).toEqual([{ text: "caching" }]);
});
it("keeps wikipedia-style parens inside directive values", () => {
const q = parseSearchQuery("site:en.wikipedia.org/wiki/Rust_(programming_language) borrow checker");
expect(q.sites).toEqual(["en.wikipedia.org/wiki/rust_(programming_language)"]);
expect(q.text).toBe("borrow checker");
});
it("treats legacy +term as an exact phrase", () => {
const q = parseSearchQuery("+immutable data");
expect(q.terms[0]).toEqual({ text: "immutable", phrase: true });
});
it("parses lang: into a language code", () => {
const q = parseSearchQuery("documentation lang:EN-us");
expect(q.lang).toBe("en-us");
expect(q.hasConstraints).toBe(false);
});
});
describe("parseDateValue", () => {
it("rejects invalid components", () => {
expect(parseDateValue("2024-13-01")).toBeUndefined();
expect(parseDateValue("2024-00-10")).toBeUndefined();
expect(parseDateValue("notadate")).toBeUndefined();
expect(parseDateValue("24-01-01")).toBeUndefined();
});
});
describe("formatQuery", () => {
it("re-emits full Google syntax", () => {
const q = parseSearchQuery(
'release notes site:github.com -site:gist.github.com filetype:md after:2024-05-01 intitle:"v2"',
);
expect(formatQuery(q, GOOGLE_QUERY_SYNTAX)).toBe(
"release notes site:github.com -site:gist.github.com intitle:v2 filetype:md after:2024-05-01",
);
});
it("emits OR-grouped sites for multi-site queries", () => {
const q = parseSearchQuery("cve site:nvd.nist.gov site:mitre.org");
expect(formatQuery(q, GOOGLE_QUERY_SYNTAX)).toBe("cve (site:nvd.nist.gov OR site:mitre.org)");
});
it("produces plain keywords when the engine supports no syntax", () => {
const q = parseSearchQuery('(react OR vue) "state management" -redux site:dev.to filetype:pdf');
expect(formatQuery(q, {})).toBe("react vue state management");
});
it("falls back to constraint values for directive-only queries", () => {
const q = parseSearchQuery("site:kubernetes.io filetype:yaml");
expect(formatQuery(q, {})).toBe("kubernetes.io yaml");
});
});
describe("formatScraperQuery", () => {
it("demotes path-carrying site: and inurl: to plain terms, keeping bare-domain site:", () => {
expect(formatScraperQuery("site:github.com/can1357/oh-my-pi inurl:releases site:github.com 17.1.1 release")).toBe(
"17.1.1 release github.com/can1357/oh-my-pi releases site:github.com",
);
});
it("demotes every site in an OR-groupable multi-site query when all carry paths", () => {
expect(formatScraperQuery("cve site:nvd.nist.gov/vuln site:mitre.org/cgi-bin")).toBe(
"cve nvd.nist.gov/vuln mitre.org/cgi-bin",
);
});
it("passes negated site:/inurl: through as operators", () => {
expect(formatScraperQuery("foo -site:github.com/x -inurl:bar")).toBe("foo -site:github.com/x -inurl:bar");
});
it("passes directive-free queries through byte-identical", () => {
expect(formatScraperQuery("github.com/can1357/oh-my-pi 17.1.1 release")).toBe(
"github.com/can1357/oh-my-pi 17.1.1 release",
);
});
it("respects a narrower engine syntax while still demoting hostile operators", () => {
expect(
formatScraperQuery("a site:x.com site:y.com/z inurl:w", undefined, {
phrases: true,
negation: true,
site: true,
}),
).toBe("a y.com/z w site:x.com");
});
});
describe("matchesSite", () => {
it("matches exact hosts, subdomains, and path prefixes", () => {
expect(matchesSite("https://docs.k8s.io/setup", "k8s.io")).toBe(true);
expect(matchesSite("https://k8s.io/", "k8s.io")).toBe(true);
expect(matchesSite("https://notk8s.io/", "k8s.io")).toBe(false);
expect(matchesSite("https://github.com/anthropics/sdk", "github.com/anthropics")).toBe(true);
expect(matchesSite("https://github.com/other/sdk", "github.com/anthropics")).toBe(false);
expect(matchesSite("not a url", "k8s.io")).toBe(false);
});
});
describe("applyQueryConstraints", () => {
const sources: SearchSource[] = [
{ title: "K8s docs — Install", url: "https://kubernetes.io/docs/setup/install.pdf" },
{ title: "Random blog", url: "https://blog.example.com/k8s", publishedDate: "2023-01-15" },
{ title: "Reddit thread", url: "https://www.reddit.com/r/kubernetes/post" },
{ title: "K8s blog", url: "https://kubernetes.io/blog/2024", ageSeconds: 3600 },
];
it("filters by included site", () => {
const q = parseSearchQuery("install site:kubernetes.io");
const { sources: out, dropped } = applyQueryConstraints(sources, q);
expect(out.map(s => s.url)).toEqual([
"https://kubernetes.io/docs/setup/install.pdf",
"https://kubernetes.io/blog/2024",
]);
expect(dropped).toEqual([]);
});
it("drops a constraint that would eliminate every result and reports it", () => {
const q = parseSearchQuery("install site:nonexistent.example");
const { sources: out, dropped } = applyQueryConstraints(sources, q);
expect(out).toHaveLength(sources.length);
expect(dropped).toEqual(["site:nonexistent.example"]);
});
it("relaxes dimensions independently", () => {
// site matches two sources, filetype:docx matches none of those → only filetype relaxed.
const q = parseSearchQuery("install site:kubernetes.io filetype:docx");
const { sources: out, dropped } = applyQueryConstraints(sources, q);
expect(out.map(s => s.url)).toEqual([
"https://kubernetes.io/docs/setup/install.pdf",
"https://kubernetes.io/blog/2024",
]);
expect(dropped).toEqual(["filetype:docx"]);
});
it("applies exclusions and filetype filters", () => {
const q = parseSearchQuery("k8s -site:reddit.com filetype:pdf");
const { sources: out, dropped } = applyQueryConstraints(sources, q);
expect(out.map(s => s.url)).toEqual(["https://kubernetes.io/docs/setup/install.pdf"]);
expect(dropped).toEqual([]);
});
it("filters by date bounds while letting undated sources pass", () => {
const q = parseSearchQuery("k8s after:2024-01-01");
const { sources: out } = applyQueryConstraints(sources, q);
// 2023 blog post is provably too old; undated sources survive.
expect(out.map(s => s.url)).toEqual([
"https://kubernetes.io/docs/setup/install.pdf",
"https://www.reddit.com/r/kubernetes/post",
"https://kubernetes.io/blog/2024",
]);
});
it("parses relative published dates", () => {
const q = parseSearchQuery("news before:2020");
const relative: SearchSource[] = [
{ title: "old", url: "https://a.example/1", publishedDate: "2 days ago" },
{ title: "undated", url: "https://a.example/2" },
];
const { sources: out, dropped } = applyQueryConstraints(relative, q);
// "2 days ago" is provably after 2020 → violates before:2020 → only undated survives.
expect(out.map(s => s.url)).toEqual(["https://a.example/2"]);
expect(dropped).toEqual([]);
});
it("returns empty input untouched", () => {
const q = parseSearchQuery("x site:a.com");
expect(applyQueryConstraints([], q)).toEqual({ sources: [], dropped: [] });
});
});
describe("matchesQueryConstraints", () => {
it("checks all dimensions strictly", () => {
const q = parseSearchQuery("site:kubernetes.io inurl:docs");
expect(matchesQueryConstraints({ title: "t", url: "https://kubernetes.io/docs/setup" }, q)).toBe(true);
expect(matchesQueryConstraints({ title: "t", url: "https://kubernetes.io/blog/x" }, q)).toBe(false);
});
});
@@ -45,6 +45,13 @@ describe("Tavily buildRequestBody", () => {
expect(body.include_answer).toBe("advanced");
expect(body.include_raw_content).toBe(false);
});
it("prefers explicit start_date/end_date over time_range", () => {
const body = buildRequestBody({ query: "q", recency: "week", start_date: "2026-01-01", end_date: "2026-02-01" });
expect(body.start_date).toBe("2026-01-01");
expect(body.end_date).toBe("2026-02-01");
expect(body).not.toHaveProperty("time_range");
});
});
describe("Tavily searchTavily request shape (integration)", () => {
@@ -139,4 +146,72 @@ describe("Tavily searchTavily request shape (integration)", () => {
expect(capturedBody).not.toHaveProperty("topic");
expect(capturedBody).not.toHaveProperty("time_range");
});
it("maps site: directives to include/exclude_domains and strips them from the query", async () => {
process.env.TAVILY_API_KEY = "test-key";
let capturedBody: Record<string, unknown> | undefined;
const fetchMock: FetchImpl = async (input, init) => {
const url =
typeof input === "string" ? input : input instanceof URL ? input.toString() : (input as Request).url;
if (url === "https://api.tavily.com/search") {
capturedBody = JSON.parse(init?.body as string);
return new Response(
JSON.stringify({
answer: "test answer",
results: [{ title: "Pricing", url: "https://tavily.com/pricing", content: "plans" }],
request_id: "req-1",
}),
{ status: 200, headers: { "Content-Type": "application/json" } },
);
}
return new Response("not mocked", { status: 500 });
};
await searchTavily({
...makeParams("pricing site:tavily.com -site:reddit.com"),
fetch: fetchMock,
});
expect(capturedBody).toBeDefined();
expect(capturedBody?.include_domains).toEqual(["tavily.com"]);
expect(capturedBody?.exclude_domains).toEqual(["reddit.com"]);
expect(capturedBody?.query).toBe("pricing");
});
it("maps after:/before: to start_date/end_date and retries without them on empty results", async () => {
process.env.TAVILY_API_KEY = "test-key";
const capturedBodies: Record<string, unknown>[] = [];
const fetchMock: FetchImpl = async (input, init) => {
const url =
typeof input === "string" ? input : input instanceof URL ? input.toString() : (input as Request).url;
if (url === "https://api.tavily.com/search") {
capturedBodies.push(JSON.parse(init?.body as string));
const empty = capturedBodies.length === 1;
return new Response(
JSON.stringify({
answer: "",
results: empty ? [] : [{ title: "Post", url: "https://example.com/post", content: "text" }],
request_id: `req-${capturedBodies.length}`,
}),
{ status: 200, headers: { "Content-Type": "application/json" } },
);
}
return new Response("not mocked", { status: 500 });
};
const response = await searchTavily({
...makeParams('"llm agents" after:2026-01-01 before:2026-06-01'),
fetch: fetchMock,
});
expect(capturedBodies).toHaveLength(2);
expect(capturedBodies[0]?.start_date).toBe("2026-01-01");
expect(capturedBodies[0]?.end_date).toBe("2026-06-01");
expect(capturedBodies[0]?.query).toBe('"llm agents"');
expect(capturedBodies[1]).not.toHaveProperty("start_date");
expect(capturedBodies[1]).not.toHaveProperty("end_date");
expect(response.sources).toHaveLength(1);
});
});
@@ -1,6 +1,6 @@
import { describe, expect, it } from "bun:test";
import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai";
import { searchZai } from "@oh-my-pi/pi-coding-agent/web/search/providers/zai";
import { searchZai, ZaiProvider } from "@oh-my-pi/pi-coding-agent/web/search/providers/zai";
interface CapturedRequest {
method: string | undefined;
@@ -108,4 +108,92 @@ describe("Z.AI web search provider", () => {
},
]);
});
function createMcpFetch(): { fetchImpl: FetchImpl; capturedRequests: CapturedRequest[] } {
const capturedRequests: CapturedRequest[] = [];
const fetchImpl: FetchImpl = (_input, init) => {
const request = {
method: init?.method,
headers: new Headers(init?.headers),
body: JSON.parse(String(init?.body)) as Record<string, unknown>,
};
capturedRequests.push(request);
if (request.body.method === "initialize") {
return Promise.resolve(
new Response(
JSON.stringify({
jsonrpc: "2.0",
id: request.body.id,
result: {
protocolVersion: "2025-03-26",
capabilities: { tools: {} },
serverInfo: { name: "zai-web-search", version: "test" },
},
}),
{
status: 200,
headers: { "Content-Type": "application/json", "Mcp-Session-Id": "zai-session-1" },
},
),
);
}
if (request.body.method === "notifications/initialized") {
return Promise.resolve(new Response(null, { status: 202 }));
}
return Promise.resolve(
new Response(
JSON.stringify({
jsonrpc: "2.0",
id: request.body.id,
result: {
content: [{ type: "text", text: JSON.stringify({ search_result: [] }) }],
},
}),
{ status: 200, headers: { "Content-Type": "application/json" } },
),
);
};
return { fetchImpl, capturedRequests };
}
const authStorage = {
resolver() {
return async () => "zai-test-key";
},
hasAuth(provider: string) {
return provider === "zai";
},
} as unknown as AuthStorage;
function toolCallQuery(capturedRequests: CapturedRequest[]): unknown {
const toolCall = capturedRequests.find(request => request.body.method === "tools/call");
const params = toolCall?.body.params as { arguments?: { query?: unknown } } | undefined;
return params?.arguments?.query;
}
it("rewrites directive queries into Bing-flavored operator syntax, dropping date bounds", async () => {
const { fetchImpl, capturedRequests } = createMcpFetch();
await new ZaiProvider().search({
query: 'pytest "fixture scope" site:docs.pytest.org -inurl:changelog filetype:html after:2024-01-01',
systemPrompt: "",
authStorage,
fetch: fetchImpl,
});
expect(toolCallQuery(capturedRequests)).toBe(
'pytest "fixture scope" site:docs.pytest.org -inurl:changelog filetype:html',
);
});
it("sends directive-free queries upstream byte-identical", async () => {
const { fetchImpl, capturedRequests } = createMcpFetch();
await new ZaiProvider().search({
query: "latest bun release notes",
systemPrompt: "",
authStorage,
fetch: fetchImpl,
});
expect(toolCallQuery(capturedRequests)).toBe("latest bun release notes");
});
});