diff --git a/packages/coding-agent/CHANGELOG.md b/packages/coding-agent/CHANGELOG.md index 20eeacc3c..ee7fad936 100644 --- a/packages/coding-agent/CHANGELOG.md +++ b/packages/coding-agent/CHANGELOG.md @@ -7,15 +7,20 @@ - Added in-process moreutils-style shell builtins to the bash tool's embedded shell: `ts` (timestamp lines; `-i`/`-s` elapsed modes, `-m` monotonic clock, `-r` relative rewriting of RFC3339/syslog timestamps, `%.S`/`%.s`/`%.T` subsecond extensions), `sponge` (soak stdin fully before atomically writing the target, so `foo file | ... | sponge file` works; `-a` appends), `ifne` (run a command only when stdin is non-empty; `-n` inverts and passes non-empty stdin through), `isutf8` (streaming UTF-8 validation with line/char/byte diagnostics; `-q`, `-l`, `-i`), `combine` (boolean `and`/`not`/`or`/`xor` on the lines of two files, `-` for stdin), and `errno` (errno name/number/description lookup with `-l` list and `-s` search; unix only). Like the uutils-backed builtins, they run in-process against the command's own stdio, resolve paths against the shell working directory, honor cancellation, and are disabled by `PI_DISABLE_UUTILS_BUILTINS`. - The bash tool prompt now lists the available shell builtins (`mkdir` through `jq`, `rm`/`mv`/`ln`, and the moreutils set) so the model relies on them without existence checks; the line is dropped when `PI_DISABLE_UUTILS_BUILTINS` disables the builtins and omits unix-only `errno` on Windows. - Sessions using a broker-backed auth store now report each completed request's token usage and cost to the auth broker (batched, 10s cadence) so the broker can track actual token burn per client install +- Added a per-spawn `effort` parameter to the `task` tool (`"lo"` | `"med"` | `"hi"`): each selector maps onto the resolved model's supported thinking range (lowest, middle, and highest level — whatever the model tops out at, e.g. `high`, `xhigh`, or `max`) and overrides the agent's default selector, including `auto`. Omitting `effort` keeps the existing automatic per-prompt thinking classification. +- Added `searxng.engines` setting for the SearXNG web search provider: a comma-separated list of engine names or bang shortcuts (e.g. `ddg, br, startpage`) sent as the API's `engines=` parameter. Shortcuts are resolved to canonical engine names via the instance's `/config` endpoint (cached per endpoint; entries pass through verbatim if `/config` is unreachable). Bang syntax in queries (`!ddg foo`) continues to pass through to the instance, and external bangs (`!!g`) are now stripped client-side since SearXNG answers them with an HTTP redirect even for JSON requests. +- Web search queries now understand Google-style directives on every provider: `site:`/`-site:` (plus `domain:`/`host:` aliases), `after:`/`before:`/`since:`/`until:` date bounds, `inurl:`/`intitle:`/`intext:`/`allin*:`, `filetype:`/`ext:`, `lang:`, quoted phrases (including smart quotes), `+term`, `-`/`NOT` exclusions, and `OR`/`|` groups. A shared parser (`web/search/query.ts`) structures the query once per request; each provider maps constraints onto native API filters where the upstream supports them (Perplexity domain/date/language filters on both the API-key and ask paths, Tavily `include_domains`/`exclude_domains` + `start_date`/`end_date`, Exa domain lists + published-date bounds on API and MCP paths, Anthropic `allowed_domains`/`blocked_domains`, xAI `filters.allowed_domains`/`excluded_domains`, Parallel `source_policy`, Brave absolute `freshness` ranges, Firecrawl `tbs=cdr` date ranges, Jina `X-Site`, SearXNG `language`) and otherwise re-emits only the operator syntax its engine parses (full Google syntax for Gemini grounding, OpenAI, Kagi, and the credential-free scrapers — with scraper-hostile path-`site:`/`inurl:` operators demoted to plain keywords; conservative subsets for DuckDuckGo, Mojeek, Kimi, Z.AI, TinyFish, and Synthetic). The pipeline then applies a lenient post-filter to every response: constraints the engine ignored are enforced on the returned sources, and any constraint dimension that would eliminate every result is relaxed and reported to the model (`Note: no results matched \`site:...\`; the constraint was relaxed`) instead of returning nothing. Directive-free queries are passed through byte-identical everywhere. ### Changed - Large pastes saved via the large-paste menu now insert `local://paste-N.md` references (previously `local://attachment-N`), so the saved paste carries a markdown extension and a clearer name. - Raw SSE debug capture now trims over-budget events smartly instead of chopping off the tail: tool definitions inside `data:` payloads are compacted first (name kept, schema/description elided — often enough to keep the whole payload as valid JSON), and anything still over the 64k cap keeps its head and tail with a `: omp-debug-elided chars=N` comment marking the removed middle, so trailing fields like `usage` stay visible. +- The `web_search` tool prompt now tells the model to never search for content that is programmatically accessible or has a known URL (GitHub, known arXiv papers, Wikipedia pages, official docs) and to `read` the URL directly instead. ### Fixed - Fixed `todo` calls that omit `op` hard-failing validation ("op must be operation to apply (was missing)"): the tool now validates leniently and infers the op for unambiguous payloads (`list` → `init`, `phase`+`items` → `append`, bare `items` on an empty list → `init`); `op` stays required in the schema, and ambiguous op-less calls surface the schema error as a retryable tool error. +- Fixed credential-free web search engines (SearXNG, DuckDuckGo, Google, Startpage, Ecosia, Mojeek, and the Public Web fan-out) returning zero results for queries with `site:` paths (e.g. `site:github.com/owner/repo`) or `inurl:` operators: scraper engines only match `site:` against a bare domain and DuckDuckGo ignores `inurl:` entirely, so such queries silently emptied the result set and fell through to the next provider in the chain. A shared `formatScraperQuery` formatter now structurally demotes path-carrying `site:` and all `inurl:` values to plain search terms (covering OR-grouped and quoted directives) while preserving bare-domain `site:` filters, negated operators, and each engine's supported syntax; the pipeline post-filter still enforces the demoted constraints on returned sources. - Fixed `ast_edit` previews reading like applied edits to the model: the `⟨proposed⟩` badge was TUI-only, so the model-visible result (hashline header + `-`/`+` rows, identical to applied edit output) carried no staged-proposal signal. The preview result now leads with a "Staged as a proposal — files NOT modified yet" notice naming `xd://resolve`/`xd://reject`, the injected resolve reminder names the source tool, and the `ast_edit` tool prompt documents the two-phase flow. - Fixed the `hub` launch `ps`/`list` response burying the active process behind every exited one and growing without bound in long-lived projects: the broker now lists non-terminal daemons first (oldest to newest) and caps exited/failed history at the 10 most recently exited, so the active launch is immediately visible and the response stays bounded. Broker recovery also preserves each already-terminal daemon's real exit time instead of overwriting it with the restart timestamp, so the history cap keeps the genuinely most-recently-exited processes after an idle-broker restart ([#6517](https://github.com/can1357/oh-my-pi/issues/6517)). diff --git a/packages/coding-agent/src/cli/web-search-cli.ts b/packages/coding-agent/src/cli/web-search-cli.ts index b642cda67..cd7f12564 100644 --- a/packages/coding-agent/src/cli/web-search-cli.ts +++ b/packages/coding-agent/src/cli/web-search-cli.ts @@ -130,8 +130,15 @@ ${chalk.bold("Options:")} --compact Render condensed output -h, --help Show this help +${chalk.bold("Query directives:")} + site:/-site: after:/before: (YYYY-MM-DD) inurl: intitle: filetype: + "exact phrase" -term OR + Mapped to native provider filters where available, otherwise applied as a + lenient post-filter (a constraint matching nothing is relaxed, not fatal). + ${chalk.bold("Examples:")} ${APP_NAME} q --provider=exa "what's the color of the sky" ${APP_NAME} q --provider=brave --recency=week "latest TypeScript 5.7 changes" + ${APP_NAME} q 'transformer scaling site:arxiv.org after:2024 -site:reddit.com' `); } diff --git a/packages/coding-agent/src/config/settings-schema.ts b/packages/coding-agent/src/config/settings-schema.ts index b1204abe4..8d808ad55 100644 --- a/packages/coding-agent/src/config/settings-schema.ts +++ b/packages/coding-agent/src/config/settings-schema.ts @@ -5266,6 +5266,11 @@ export const SETTINGS_SCHEMA = { default: undefined, }, + "searxng.engines": { + type: "string", + default: undefined, + }, + "searxng.language": { type: "string", default: undefined, diff --git a/packages/coding-agent/src/prompts/tools/web-search.md b/packages/coding-agent/src/prompts/tools/web-search.md index a5dcc4004..a48555a07 100644 --- a/packages/coding-agent/src/prompts/tools/web-search.md +++ b/packages/coding-agent/src/prompts/tools/web-search.md @@ -3,4 +3,6 @@ Searches the web for up-to-date information beyond knowledge cutoff. - You SHOULD prefer primary sources (papers, official docs) and corroborate key claims with multiple sources - You MUST include links for cited sources in the final response +- NEVER use for content that is programmatically accessible or whose URL you already know (GitHub repos/issues, a known arXiv paper, a Wikipedia page, official docs) — `read` the URL directly instead +- `query` supports Google-style directives on every provider: `site:`/`-site:`, `after:`/`before:` (`YYYY-MM-DD`), `inurl:`, `intitle:`, `filetype:`, `"exact phrase"`, `-term`, `OR`. Constraints map to native provider filters where available; otherwise results are filtered leniently — a constraint matching nothing is relaxed and reported instead of returning zero results. diff --git a/packages/coding-agent/src/web/search/index.ts b/packages/coding-agent/src/web/search/index.ts index 83cf2c5dc..41c1413a0 100644 --- a/packages/coding-agent/src/web/search/index.ts +++ b/packages/coding-agent/src/web/search/index.ts @@ -27,6 +27,7 @@ import { type SearchProvider, type SearchProviderCandidate, } from "./provider"; +import { applyQueryConstraints, parseSearchQuery } from "./query"; import { renderSearchCall, renderSearchResult, type SearchRenderDetails } from "./render"; import type { SearchProviderId, SearchResponse } from "./types"; import { SearchProviderError } from "./types"; @@ -57,9 +58,12 @@ function formatCount(label: string, count: number): string { return `${count} ${label}${count === 1 ? "" : "s"}`; } -/** Format response for LLM consumption */ -function formatForLLM(response: SearchResponse): string { +/** Format response for LLM consumption. `notes` lead the output (e.g. relaxed-constraint warnings). */ +function formatForLLM(response: SearchResponse, notes: readonly string[] = []): string { const parts: string[] = []; + for (const note of notes) { + parts.push(`Note: ${note}`); + } if (response.answer) { parts.push(response.answer); @@ -141,6 +145,8 @@ async function executeSearch( candidates = resolveProviderCandidates(); } + const parsedQuery = parseSearchQuery(params.query); + // Invariant across providers; read once and tolerate an uninitialized // Settings singleton (e.g. `omp q ...` CLI path, unit tests) so the // provider-fallback loop never aborts before any provider runs. @@ -182,6 +188,7 @@ async function executeSearch( const response = await provider.search({ query: params.query, + parsedQuery, limit: params.limit, recency: params.recency, systemPrompt: webSearchSystemPrompt, @@ -196,15 +203,31 @@ async function executeSearch( geminiModel, }); - if (!hasRenderableSearchContent(response)) { + // Lenient constraint pass over whatever the provider returned: enforce + // site:/inurl:/intitle:/filetype:/date directives the provider could + // not (or only partially) honor natively, relaxing any dimension that + // would wipe out every result. Citations/answer text stay untouched. + let finalResponse = response; + const constraintNotes: string[] = []; + if (parsedQuery.hasConstraints && response.sources.length > 0) { + const filtered = applyQueryConstraints(response.sources, parsedQuery); + if (filtered.sources.length !== response.sources.length) { + finalResponse = { ...response, sources: filtered.sources }; + } + for (const label of filtered.dropped) { + constraintNotes.push(`no results matched \`${label}\`; the constraint was relaxed`); + } + } + + if (!hasRenderableSearchContent(finalResponse)) { throw new SearchProviderError(provider.id, `${provider.label} returned no renderable search content.`, 204); } - const text = formatForLLM(response); + const text = formatForLLM(finalResponse, constraintNotes); return { content: [{ type: "text" as const, text }], - details: { response }, + details: { response: finalResponse }, }; } catch (error) { // Surface user-initiated cancellation immediately so the session sees diff --git a/packages/coding-agent/src/web/search/providers/anthropic.ts b/packages/coding-agent/src/web/search/providers/anthropic.ts index 7f25f7b80..d34661d85 100644 --- a/packages/coding-agent/src/web/search/providers/anthropic.ts +++ b/packages/coding-agent/src/web/search/providers/anthropic.ts @@ -28,6 +28,7 @@ import type { SearchSource, } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery, type QuerySyntax, type StructuredQuery } from "../query"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; import { classifyProviderHttpError, withHardTimeout } from "./utils"; @@ -36,6 +37,59 @@ const DEFAULT_MODEL = "claude-haiku-4-5"; const DEFAULT_MAX_TOKENS = 4096; const WEB_SEARCH_TOOL_NAME = "web_search"; const WEB_SEARCH_TOOL_TYPE = "web_search_20250305"; + +/** + * Claude's search backend understands common Google-style operators, so most + * directives are re-emitted as query text. `site:` is intentionally absent: + * site includes/excludes map onto the web_search tool's native + * `allowed_domains`/`blocked_domains` parameters instead. + */ +const ANTHROPIC_QUERY_SYNTAX: QuerySyntax = { + phrases: true, + negation: true, + or: true, + inUrl: true, + inTitle: true, + filetype: true, + dateRange: true, +}; + +/** Upstream request shape derived from the parsed query. */ +interface AnthropicQueryPlan { + query: string; + allowedDomains?: string[]; + blockedDomains?: string[]; +} + +/** + * Map parsed directives onto the request: `site:` includes become + * `allowed_domains`, `-site:` exclusions become `blocked_domains` (the two are + * mutually exclusive on the API, so exclusions are only sent when there are no + * includes), and remaining directives are re-emitted as query syntax. + * Directive-free queries pass through byte-identical. Anthropic domain + * filters take bare hosts (subdomains included automatically); any path part + * of a `site:` value is enforced by the central constraint filter. + */ +function planQuery(rawQuery: string, parsed: StructuredQuery): AnthropicQueryPlan { + if (!parsed.hasDirectives) return { query: rawQuery }; + const hosts = (sites: readonly string[]) => { + const unique = new Set(); + for (const site of sites) { + const slash = site.indexOf("/"); + const host = slash === -1 ? site : site.slice(0, slash); + if (host.length > 0) unique.add(host); + } + return [...unique]; + }; + const allowed = hosts(parsed.sites); + const blocked = allowed.length === 0 ? hosts(parsed.excludedSites) : []; + return { + query: formatQuery(parsed, ANTHROPIC_QUERY_SYNTAX), + allowedDomains: allowed.length > 0 ? allowed : undefined, + blockedDomains: blocked.length > 0 ? blocked : undefined, + }; +} + export interface AnthropicSearchParams { query: string; system_prompt?: string; @@ -82,7 +136,7 @@ function buildSystemBlocks( * Calls the Anthropic API with web search tool enabled. * @param auth - Authentication configuration (API key or OAuth) * @param model - Model identifier to use - * @param query - Search query from the user + * @param plan - Query text plus native domain filters derived from parsed directives * @param metadataUserId - Optional Anthropic Messages metadata.user_id (already shaped for OAuth) * @param systemPrompt - Optional system prompt for guiding response style * @returns Raw API response from Anthropic @@ -91,7 +145,7 @@ function buildSystemBlocks( async function callSearch( auth: AnthropicAuthConfig, model: string, - query: string, + plan: AnthropicQueryPlan, metadataUserId?: string, systemPrompt?: string, maxTokens?: number, @@ -107,11 +161,13 @@ async function callSearch( const body: Record = { model, max_tokens: maxTokens ?? DEFAULT_MAX_TOKENS, - messages: [{ role: "user", content: query }], + messages: [{ role: "user", content: plan.query }], tools: [ { type: WEB_SEARCH_TOOL_TYPE, name: WEB_SEARCH_TOOL_NAME, + ...(plan.allowedDomains ? { allowed_domains: plan.allowedDomains } : {}), + ...(plan.blockedDomains ? { blocked_domains: plan.blockedDomains } : {}), }, ], }; @@ -283,6 +339,8 @@ export async function searchAnthropic( const callerSessionId = "authStorage" in params ? params.sessionId : undefined; const accountId = "authStorage" in params ? params.authStorage.getOAuthAccountId("anthropic", params.sessionId) : undefined; + const parsed = ("parsedQuery" in params ? params.parsedQuery : undefined) ?? parseSearchQuery(params.query); + const plan = planQuery(params.query, parsed); const response = await withAuth( keyOrResolver, key => { @@ -302,7 +360,7 @@ export async function searchAnthropic( return callSearch( auth, model, - params.query, + plan, metadataUserId, systemPrompt, maxTokens, diff --git a/packages/coding-agent/src/web/search/providers/base.ts b/packages/coding-agent/src/web/search/providers/base.ts index 6841cf253..a123331b4 100644 --- a/packages/coding-agent/src/web/search/providers/base.ts +++ b/packages/coding-agent/src/web/search/providers/base.ts @@ -1,5 +1,6 @@ import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai"; import type { ModelRegistry } from "../../../config/model-registry"; +import type { StructuredQuery } from "../query"; import type { SearchProviderId, SearchResponse } from "../types"; /** @@ -14,6 +15,21 @@ import type { SearchProviderId, SearchResponse } from "../types"; */ export interface SearchParams { query: string; + /** + * Structured view of `query`, parsed once by the search pipeline: + * Google-style directives (`site:`, `before:`/`after:`, `inurl:`, + * `intitle:`, `filetype:`, quoted phrases, `OR` groups, `-exclusions`) + * extracted into fields. + * + * Providers SHOULD map constraints onto native API parameters + * (domain/date filters) or engine query syntax (`formatQuery`) where the + * upstream supports them, and lean lenient otherwise: the pipeline + * post-filters every response with `applyQueryConstraints`, which + * relaxes any constraint that would eliminate all results — so a + * best-effort search always beats an empty one. When absent (direct + * provider calls), parse with `parseSearchQuery(params.query)`. + */ + parsedQuery?: StructuredQuery; limit?: number; /** * Temporal filter narrowing results to the specified time window. diff --git a/packages/coding-agent/src/web/search/providers/brave.ts b/packages/coding-agent/src/web/search/providers/brave.ts index 5fdbc2c02..57ac1f480 100644 --- a/packages/coding-agent/src/web/search/providers/brave.ts +++ b/packages/coding-agent/src/web/search/providers/brave.ts @@ -7,6 +7,8 @@ import { type AuthStorage, type FetchImpl, getEnvApiKey } from "@oh-my-pi/pi-ai"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import type { QuerySyntax, StructuredQuery } from "../query"; +import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery } from "../query"; import { clampNumResults, dateToAgeSeconds } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -23,10 +25,32 @@ const RECENCY_MAP: Record<"day" | "week" | "month" | "year", "pd" | "pw" | "pm" year: "py", }; +/** + * Brave parses the classic operator set inline (site:, quotes, -, OR…) but + * date bounds map onto the native `freshness` param, so `before:`/`after:` + * tokens are stripped from the rebuilt query string. + */ +const BRAVE_QUERY_SYNTAX: QuerySyntax = { ...GOOGLE_QUERY_SYNTAX, dateRange: false }; + +/** + * Freshness param: explicit `after:`/`before:` bounds win over the + * recency-derived period, rendered as Brave's absolute range + * `YYYY-MM-DDtoYYYY-MM-DD` with sensible open ends. + */ +function braveFreshness(parsed: StructuredQuery, recency?: keyof typeof RECENCY_MAP): string | undefined { + if (parsed.after || parsed.before) { + const start = parsed.after ?? "1970-01-01"; + const end = parsed.before ?? new Date().toISOString().slice(0, 10); + return `${start}to${end}`; + } + return recency ? RECENCY_MAP[recency] : undefined; +} + export interface BraveSearchParams { query: string; num_results?: number; recency?: "day" | "week" | "month" | "year"; + parsedQuery?: StructuredQuery; signal?: AbortSignal; fetch?: FetchImpl; } @@ -73,12 +97,14 @@ async function callBraveSearch( params: BraveSearchParams, ): Promise<{ response: BraveSearchResponse; requestId?: string }> { const numResults = clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS); + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); const url = new URL(BRAVE_SEARCH_URL); - url.searchParams.set("q", params.query); + url.searchParams.set("q", parsed.hasDirectives ? formatQuery(parsed, BRAVE_QUERY_SYNTAX) : params.query); url.searchParams.set("count", String(numResults)); url.searchParams.set("extra_snippets", "true"); - if (params.recency) { - url.searchParams.set("freshness", RECENCY_MAP[params.recency]); + const freshness = braveFreshness(parsed, params.recency); + if (freshness) { + url.searchParams.set("freshness", freshness); } const fetchImpl = params.fetch ?? fetch; @@ -145,6 +171,7 @@ export class BraveProvider extends SearchProvider { query: params.query, num_results: params.numSearchResults ?? params.limit, recency: params.recency, + parsedQuery: params.parsedQuery, signal: params.signal, fetch: params.fetch, }); diff --git a/packages/coding-agent/src/web/search/providers/codex.ts b/packages/coding-agent/src/web/search/providers/codex.ts index 9b543f0f7..4599b3e5e 100644 --- a/packages/coding-agent/src/web/search/providers/codex.ts +++ b/packages/coding-agent/src/web/search/providers/codex.ts @@ -31,6 +31,7 @@ import packageJson from "../../../../package.json" with { type: "json" }; import type { ModelRegistry } from "../../../config/model-registry"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery } from "../query"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; import { classifyProviderHttpError, withHardTimeout } from "./utils"; @@ -572,6 +573,7 @@ async function callCodexSearch( async function runCodexSearchCandidates(options: { auth: { accessToken: string; accountId?: string }; params: SearchParams; + query: string; modelCandidates: CodexModelCandidate[]; modelWasConfigured: boolean; transport: CodexSearchTransport; @@ -582,7 +584,7 @@ async function runCodexSearchCandidates(options: { if (!candidate) continue; try { - return await callCodexSearch(options.auth, options.params.query, { + return await callCodexSearch(options.auth, options.query, { signal: options.params.signal, systemPrompt: options.params.systemPrompt, searchContextSize: "high", @@ -622,6 +624,15 @@ export async function searchCodex(params: SearchParams): Promise throw new SearchProviderError("codex", "No Codex web search model is configured."); } const transport = resolveCodexSearchTransport(params.modelRegistry, firstCandidate.modelId); + // The ChatGPT-backend Codex endpoint speaks the undocumented codex-rs + // request shape (responses-lite moves tools into an `additional_tools` + // developer item), so the documented `web_search.filters.allowed_domains` + // parameter cannot be assumed to survive it. Instead, re-emit directive + // queries with the full Google-style operator syntax — the backing index + // parses the classic operator set — and leave directive-free queries + // byte-identical. + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + const query = parsed.hasDirectives ? formatQuery(parsed, GOOGLE_QUERY_SYNTAX) : params.query; let result: CodexSearchResult; if (transport.customEndpoint) { @@ -652,6 +663,7 @@ export async function searchCodex(params: SearchParams): Promise runCodexSearchCandidates({ auth: { accessToken }, params, + query, modelCandidates, modelWasConfigured: configuredModel !== undefined, transport, @@ -682,6 +694,7 @@ export async function searchCodex(params: SearchParams): Promise return runCodexSearchCandidates({ auth: { accessToken: access.accessToken, accountId }, params, + query, modelCandidates, modelWasConfigured: configuredModel !== undefined, transport, diff --git a/packages/coding-agent/src/web/search/providers/duckduckgo.ts b/packages/coding-agent/src/web/search/providers/duckduckgo.ts index c77972375..f8343d7fb 100644 --- a/packages/coding-agent/src/web/search/providers/duckduckgo.ts +++ b/packages/coding-agent/src/web/search/providers/duckduckgo.ts @@ -1,6 +1,8 @@ import type { AuthStorage } from "@oh-my-pi/pi-ai"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import type { QuerySyntax } from "../query"; +import { formatScraperQuery } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -118,8 +120,28 @@ function isAnomalyResponse(html: string): boolean { return html.includes("anomaly-modal") || html.includes("anomaly.js"); } +/** + * Query syntax the DDG HTML frontend parses: quotes, `-`, OR, site:, + * filetype:, intitle:, inurl:, intext:. Date bounds (`before:`/`after:`) are + * deliberately off — DDG does not parse them, so they are stripped from the + * query and enforced by the pipeline's lenient post-filter instead. + */ +const DDG_QUERY_SYNTAX: QuerySyntax = { + phrases: true, + negation: true, + or: true, + site: true, + inUrl: true, + inTitle: true, + inText: true, + filetype: true, +}; + async function callDuckDuckGoHtml(params: SearchParams): Promise { - const form = new URLSearchParams({ q: params.query, kl: "us-en" }); + const form = new URLSearchParams({ + q: formatScraperQuery(params.query, params.parsedQuery, DDG_QUERY_SYNTAX), + kl: "us-en", + }); const df = params.recency ? RECENCY_TO_DDG_DF[params.recency] : undefined; if (df) form.set("df", df); // Add b: "" parameter as specified in the browser fetch template to match real browser form submission diff --git a/packages/coding-agent/src/web/search/providers/ecosia.ts b/packages/coding-agent/src/web/search/providers/ecosia.ts index a41cca27a..fcbb37163 100644 --- a/packages/coding-agent/src/web/search/providers/ecosia.ts +++ b/packages/coding-agent/src/web/search/providers/ecosia.ts @@ -2,6 +2,7 @@ import type { AuthStorage } from "@oh-my-pi/pi-ai"; import { parseHTML } from "linkedom"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatScraperQuery } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -100,7 +101,10 @@ function isBlockedPage(page: LoadedHtmlPage): boolean { async function callEcosiaHtml(params: SearchParams): Promise { const signal = withHardTimeout(params.signal); const url = new URL(ECOSIA_SEARCH_URL); - url.searchParams.set("q", params.query); + // Ecosia serves Google-backed results, so classic operators pass through + // inline; canonicalize aliases (domain: -> site:, since: -> after:) and + // demote scraper-hostile operators via the shared scraper formatter. + url.searchParams.set("q", formatScraperQuery(params.query, params.parsedQuery)); let page: LoadedHtmlPage; try { diff --git a/packages/coding-agent/src/web/search/providers/exa.ts b/packages/coding-agent/src/web/search/providers/exa.ts index a26143d67..1e71ee610 100644 --- a/packages/coding-agent/src/web/search/providers/exa.ts +++ b/packages/coding-agent/src/web/search/providers/exa.ts @@ -12,6 +12,7 @@ import { findApiKey, isSearchResponse } from "../../../exa/mcp-client"; import { parseSSE } from "../../../mcp/json-rpc"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery, type StructuredQuery } from "../query"; import { dateToAgeSeconds } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -472,8 +473,9 @@ export class ExaProvider extends SearchProvider { } search(params: SearchParams): Promise { + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); return searchExa({ - query: params.query, + ...directiveParams(parsed), num_results: params.numSearchResults ?? params.limit, signal: params.signal, authStorage: params.authStorage, @@ -482,3 +484,28 @@ export class ExaProvider extends SearchProvider { }); } } + +/** + * Map parsed query directives onto Exa's native request parameters: + * `site:` → includeDomains, `-site:` → excludeDomains (bare hosts; path parts + * are enforced by the central constraint filter), `after:`/`before:` → + * start/endPublishedDate (ISO 8601). Exa's neural search prefers natural + * language, so the query itself is re-emitted with quoted phrases only. + * Directive-free queries pass through byte-identical. + */ +function directiveParams( + parsed: StructuredQuery, +): Pick< + ExaSearchParams, + "query" | "include_domains" | "exclude_domains" | "start_published_date" | "end_published_date" +> { + if (!parsed.hasDirectives) return { query: parsed.raw }; + const hosts = (sites: readonly string[]) => [...new Set(sites.map(site => site.split("/", 1)[0]))]; + return { + query: formatQuery(parsed, { phrases: true }), + include_domains: parsed.sites.length ? hosts(parsed.sites) : undefined, + exclude_domains: parsed.excludedSites.length ? hosts(parsed.excludedSites) : undefined, + start_published_date: parsed.after, + end_published_date: parsed.before, + }; +} diff --git a/packages/coding-agent/src/web/search/providers/firecrawl.ts b/packages/coding-agent/src/web/search/providers/firecrawl.ts index cd8ad00e0..ff87a465b 100644 --- a/packages/coding-agent/src/web/search/providers/firecrawl.ts +++ b/packages/coding-agent/src/web/search/providers/firecrawl.ts @@ -14,6 +14,7 @@ import { } from "@oh-my-pi/pi-ai"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery, type StructuredQuery } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -34,6 +35,8 @@ export interface FirecrawlSearchParams { query: string; num_results?: number; recency?: SearchParams["recency"]; + /** Explicit `tbs` (custom date range); takes precedence over `recency`. */ + tbs?: string; signal?: AbortSignal; fetch?: FetchImpl; } @@ -67,8 +70,9 @@ function buildRequestBody(params: FirecrawlSearchParams): Record { + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + let query = params.query; + let tbs: string | undefined; + if (parsed.hasDirectives) { + // Firecrawl search is SERP-backed: the query supports Google operators + // (site:, inurl:, intitle:, quotes, -, OR). Absolute date bounds move to + // the native tbs param and are stripped from the query string. + tbs = buildDateTbs(parsed); + query = formatQuery(parsed, tbs ? { ...GOOGLE_QUERY_SYNTAX, dateRange: false } : GOOGLE_QUERY_SYNTAX); + } const firecrawlParams: FirecrawlSearchParams = { - query: params.query, + query, num_results: params.numSearchResults ?? params.limit, recency: params.recency, + tbs, signal: params.signal, fetch: params.fetch, }; diff --git a/packages/coding-agent/src/web/search/providers/gemini.ts b/packages/coding-agent/src/web/search/providers/gemini.ts index 27612ae77..ccd36fd8a 100644 --- a/packages/coding-agent/src/web/search/providers/gemini.ts +++ b/packages/coding-agent/src/web/search/providers/gemini.ts @@ -18,6 +18,7 @@ import { fetchWithRetry } from "@oh-my-pi/pi-utils"; import type { SearchCitation, SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery, type StructuredQuery } from "../query"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; import { classifyProviderHttpError, withHardTimeout } from "./utils"; @@ -51,6 +52,8 @@ interface GeminiToolParams { export interface GeminiSearchParams extends GeminiToolParams { query: string; + /** Pre-parsed structured query; falls back to parsing `query` when omitted. */ + parsedQuery?: StructuredQuery; system_prompt?: string; num_results?: number; /** Maximum output tokens. */ @@ -508,6 +511,12 @@ async function callGeminiDeveloperSearch( */ export async function searchGemini(params: GeminiSearchParams): Promise { const selectedModel = resolveGeminiSearchModel(params.geminiModel); + // Gemini's googleSearch grounding forwards the query to Google Search, which + // understands the classic operator set natively. Normalize directive aliases + // (domain: → site:, since: → after:, …) to canonical Google forms; leave + // directive-free queries byte-identical. + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + const searchQuery = parsed.hasDirectives ? formatQuery(parsed, GOOGLE_QUERY_SYNTAX) : params.query; const seed = await findGeminiAuth(params.authStorage, params.sessionId, params.signal); let result: GeminiSearchResult; @@ -528,7 +537,7 @@ export async function searchGemini(params: GeminiSearchParams): Promise { return searchGemini({ query: params.query, + parsedQuery: params.parsedQuery, system_prompt: params.systemPrompt, num_results: params.numSearchResults ?? params.limit, max_output_tokens: params.maxOutputTokens, diff --git a/packages/coding-agent/src/web/search/providers/google.ts b/packages/coding-agent/src/web/search/providers/google.ts index 574f2dcaf..3cad6b51c 100644 --- a/packages/coding-agent/src/web/search/providers/google.ts +++ b/packages/coding-agent/src/web/search/providers/google.ts @@ -2,6 +2,7 @@ import type { AuthStorage } from "@oh-my-pi/pi-ai"; import { parseHTML } from "linkedom"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatScraperQuery } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -91,7 +92,7 @@ function parseHtmlResults(html: string): ParsedResult[] { function buildSearchUrl(params: SearchParams, numResults: number): string { const url = new URL(GOOGLE_SEARCH_URL); - url.searchParams.set("q", params.query); + url.searchParams.set("q", formatScraperQuery(params.query, params.parsedQuery)); url.searchParams.set("num", String(numResults)); url.searchParams.set("hl", "en"); url.searchParams.set("gl", "us"); diff --git a/packages/coding-agent/src/web/search/providers/jina.ts b/packages/coding-agent/src/web/search/providers/jina.ts index 5b31e0630..a46ea5834 100644 --- a/packages/coding-agent/src/web/search/providers/jina.ts +++ b/packages/coding-agent/src/web/search/providers/jina.ts @@ -8,6 +8,7 @@ import { type AuthStorage, type FetchImpl, getEnvApiKey } from "@oh-my-pi/pi-ai"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery } from "../query"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; import { classifyProviderHttpError, withHardTimeout } from "./utils"; @@ -18,6 +19,8 @@ type SearchParamsWithFetch = SearchParams & { fetch?: FetchImpl }; export interface JinaSearchParams { query: string; num_results?: number; + /** Single bare host for Jina's `X-Site` in-site search header. */ + site?: string; signal?: AbortSignal; fetch?: FetchImpl; } @@ -39,15 +42,18 @@ export function findApiKey(): string | null { async function callJinaSearch( apiKey: string, query: string, + site?: string, signal?: AbortSignal, fetchImpl: FetchImpl = fetch, ): Promise { const requestUrl = `${JINA_SEARCH_URL}/${encodeURIComponent(query)}`; + const headers: Record = { + Accept: "application/json", + Authorization: `Bearer ${apiKey}`, + }; + if (site) headers["X-Site"] = site; const response = await fetchImpl(requestUrl, { - headers: { - Accept: "application/json", - Authorization: `Bearer ${apiKey}`, - }, + headers, signal: withHardTimeout(signal), }); @@ -69,7 +75,7 @@ export async function searchJina(params: JinaSearchParams): Promise { - const fetchImpl = params.fetch; + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + let query = params.query; + let site: string | undefined; + if (parsed.hasDirectives) { + // Jina's X-Site header takes a single domain; with exactly one + // include site, send its host there and strip site: tokens from + // the query. Multiple sites stay inline (Bing-backed, parses them). + if (parsed.sites.length === 1) site = parsed.sites[0]!.split("/")[0]; + query = formatQuery(parsed, { + phrases: true, + negation: true, + site: !site, + inTitle: true, + inUrl: true, + filetype: true, + }); + } return searchJina({ - query: params.query, + query, num_results: params.numSearchResults ?? params.limit, + site, signal: params.signal, - fetch: fetchImpl, + fetch: params.fetch, }); } } diff --git a/packages/coding-agent/src/web/search/providers/kagi.ts b/packages/coding-agent/src/web/search/providers/kagi.ts index 081022304..8de2ee349 100644 --- a/packages/coding-agent/src/web/search/providers/kagi.ts +++ b/packages/coding-agent/src/web/search/providers/kagi.ts @@ -7,6 +7,8 @@ import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai"; import type { SearchResponse } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; import { KagiApiError, searchWithKagi } from "../../kagi"; +import type { StructuredQuery } from "../query"; +import { formatQuery, GOOGLE_QUERY_SYNTAX, parseSearchQuery } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -22,16 +24,22 @@ export async function searchKagi(params: { query: string; num_results?: number; recency?: SearchParams["recency"]; + parsedQuery?: StructuredQuery; signal?: AbortSignal; authStorage: AuthStorage; sessionId?: string; fetch?: FetchImpl; }): Promise { const numResults = clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS); + // Kagi's index understands the classic Google operator set: canonicalize + // directives (domain: -> site:, until: -> before:YYYY-MM-DD, ...) and pass + // them through in the query string. Directive-free queries stay untouched. + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + const query = parsed.hasDirectives ? formatQuery(parsed, GOOGLE_QUERY_SYNTAX) : params.query; try { const result = await searchWithKagi( - params.query, + query, { limit: numResults, recency: params.recency, @@ -75,6 +83,7 @@ export class KagiProvider extends SearchProvider { return searchKagi({ query: params.query, + parsedQuery: params.parsedQuery, num_results: params.numSearchResults ?? params.limit, recency: params.recency, signal: params.signal, diff --git a/packages/coding-agent/src/web/search/providers/kimi.ts b/packages/coding-agent/src/web/search/providers/kimi.ts index c265747a1..969960005 100644 --- a/packages/coding-agent/src/web/search/providers/kimi.ts +++ b/packages/coding-agent/src/web/search/providers/kimi.ts @@ -12,6 +12,7 @@ import { $env } from "@oh-my-pi/pi-utils"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery, type QuerySyntax, type StructuredQuery } from "../query"; import { clampNumResults, dateToAgeSeconds } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -25,8 +26,19 @@ const DEFAULT_NUM_RESULTS = 10; const MAX_NUM_RESULTS = 20; const DEFAULT_TIMEOUT_SECONDS = 30; +/** Kimi Code search is Bing-flavored: re-emit the operators Bing parses; dates/lang stay with the central filter. */ +const KIMI_QUERY_SYNTAX: QuerySyntax = { + phrases: true, + negation: true, + site: true, + inTitle: true, + inUrl: true, + filetype: true, +}; + export interface KimiSearchParams { query: string; + parsedQuery?: StructuredQuery; num_results?: number; include_content?: boolean; signal?: AbortSignal; @@ -138,12 +150,14 @@ export async function searchKimi(params: KimiSearchParams): Promise callKimiSearch(key, { - query: params.query, + query, limit, includeContent: params.include_content ?? false, signal: params.signal, @@ -192,6 +206,7 @@ export class KimiProvider extends SearchProvider { return searchKimi({ query: params.query, + parsedQuery: params.parsedQuery, num_results: params.numSearchResults ?? params.limit, signal: params.signal, authStorage: params.authStorage, diff --git a/packages/coding-agent/src/web/search/providers/mojeek.ts b/packages/coding-agent/src/web/search/providers/mojeek.ts index 4951803bc..ed06c2fa3 100644 --- a/packages/coding-agent/src/web/search/providers/mojeek.ts +++ b/packages/coding-agent/src/web/search/providers/mojeek.ts @@ -4,6 +4,7 @@ import { parseHTML } from "linkedom"; import type { Page } from "puppeteer-core"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatScraperQuery, type QuerySyntax } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -81,9 +82,21 @@ function parseHtmlResults(html: string): ParsedResult[] { return results; } +/** + * Syntax re-emitted to Mojeek for directive-carrying queries. Mojeek's + * support page (mojeek.com/support/search-operators.html) confirms `site:`, + * and the community docs confirm quoted phrases and `-` exclusions. Mojeek + * also parses `in*:` operators and its own date syntax (`since:`/`before:` + * with YYYYMMDD), but the latter differs from Google's `after:`/`before:` + * ISO form and `since` is already claimed by `recency`, so date bounds and + * `in*` constraints are conservatively left to the pipeline's lenient + * post-filter instead. + */ +const MOJEEK_QUERY_SYNTAX: QuerySyntax = { phrases: true, negation: true, site: true }; + function buildSearchUrl(params: SearchParams, numResults: number): string { const url = new URL(MOJEEK_SEARCH_URL); - url.searchParams.set("q", params.query); + url.searchParams.set("q", formatScraperQuery(params.query, params.parsedQuery, MOJEEK_QUERY_SYNTAX)); url.searchParams.set("t", String(numResults)); url.searchParams.set("arc", "none"); url.searchParams.set("lang", "en"); diff --git a/packages/coding-agent/src/web/search/providers/parallel.ts b/packages/coding-agent/src/web/search/providers/parallel.ts index 5c22d03c3..957efde5f 100644 --- a/packages/coding-agent/src/web/search/providers/parallel.ts +++ b/packages/coding-agent/src/web/search/providers/parallel.ts @@ -9,6 +9,7 @@ import { parseParallelErrorResponse, parseParallelSearchPayload, } from "../../parallel"; +import { formatQuery, parseSearchQuery, type StructuredQuery } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -17,6 +18,42 @@ import { classifyProviderHttpError, toSearchSources, withHardTimeout } from "./u const DEFAULT_NUM_RESULTS = 10; const MAX_NUM_RESULTS = 40; +/** Query-string caps for Parallel: natural-language objective, no field operators. */ +const PARALLEL_QUERY_SYNTAX = { phrases: true, negation: true, or: true } as const; + +/** Parallel `source_policy` (beta Search API): bare-host allow/deny lists + freshness floor. */ +interface ParallelSourcePolicy { + include_domains?: string[]; + exclude_domains?: string[]; + after_date?: string; +} + +/** Site values may carry paths (`github.com/anthropics`); Parallel takes bare hosts. */ +function toHosts(sites: readonly string[]): string[] { + const hosts = new Set(); + for (const site of sites) { + const host = site.split("/", 1)[0]; + if (host) hosts.add(host); + } + return [...hosts]; +} + +/** + * Map parsed `site:`/`-site:`/`after:` directives onto Parallel's + * `source_policy`. Per Parallel docs, `exclude_domains` is ignored when + * `include_domains` is set, so exclusions are only sent without an allow + * list (the central lenient filter enforces them regardless). + */ +function toSourcePolicy(parsed: StructuredQuery): ParallelSourcePolicy | undefined { + const policy: ParallelSourcePolicy = {}; + const include = toHosts(parsed.sites); + const exclude = toHosts(parsed.excludedSites); + if (include.length) policy.include_domains = include; + else if (exclude.length) policy.exclude_domains = exclude; + if (parsed.after) policy.after_date = parsed.after; + return Object.keys(policy).length ? policy : undefined; +} + async function searchWithAuthStorage( objective: string, queries: string[], @@ -26,6 +63,7 @@ async function searchWithAuthStorage( }, authStorage: AuthStorage, sessionId?: string, + sourcePolicy?: ParallelSourcePolicy, ): Promise { const apiKey = await authStorage.getApiKey("parallel", sessionId, { signal: params.signal }); if (!apiKey) { @@ -57,6 +95,7 @@ async function searchWithAuthStorage( excerpts: { max_chars_per_result: 10_000, }, + ...(sourcePolicy && { source_policy: sourcePolicy }), }), signal: withHardTimeout(params.signal), }); @@ -78,22 +117,28 @@ export async function searchParallel( num_results?: number; signal?: AbortSignal; fetch?: FetchImpl; + parsedQuery?: StructuredQuery; }, authStorage: AuthStorage, sessionId?: string, ): Promise { const numResults = clampNumResults(params.num_results, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS); + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + // Back-compat: without directives the upstream request is byte-identical. + const query = parsed.hasDirectives ? formatQuery(parsed, PARALLEL_QUERY_SYNTAX) : params.query; + const sourcePolicy = parsed.hasDirectives ? toSourcePolicy(parsed) : undefined; try { const result = await searchWithAuthStorage( - params.query, - [params.query], + query, + [query], { signal: params.signal, fetch: params.fetch, }, authStorage, sessionId, + sourcePolicy, ); return { @@ -128,6 +173,7 @@ export class ParallelProvider extends SearchProvider { num_results: params.numSearchResults ?? params.limit, signal: params.signal, fetch: params.fetch, + parsedQuery: params.parsedQuery, }, params.authStorage, params.sessionId, diff --git a/packages/coding-agent/src/web/search/providers/perplexity.ts b/packages/coding-agent/src/web/search/providers/perplexity.ts index 548c8dbc8..c9c40c0a6 100644 --- a/packages/coding-agent/src/web/search/providers/perplexity.ts +++ b/packages/coding-agent/src/web/search/providers/perplexity.ts @@ -30,6 +30,7 @@ import type { SearchSource, } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery, type QuerySyntax, type StructuredQuery } from "../query"; import { dateToAgeSeconds } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -46,6 +47,71 @@ const OAUTH_USER_AGENT = "Perplexity/641 CFNetwork/1568 Darwin/25.2.0"; const ANONYMOUS_USER_AGENT = "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/149.0.0.0 Safari/537.36"; +/** + * Query-string operators Perplexity's search backend tolerates as text signal. + * `site:`/date/`lang:` directives are excluded: they map onto native request + * fields (`search_domain_filter`, `search_*_date_filter`, + * `search_language_filter`) and must be stripped from the query so the engine + * is not double-constrained. + */ +const PERPLEXITY_QUERY_SYNTAX: QuerySyntax = { + phrases: true, + negation: true, + or: true, + inUrl: true, + inTitle: true, + filetype: true, +}; + +/** Native Perplexity search filters derived from parsed query directives. */ +interface PerplexityNativeFilters { + /** Query rebuilt without natively-mapped directives. */ + query: string; + /** `search_domain_filter`: allow entries as bare hosts, deny entries as `-host`. */ + domainFilter?: string[]; + /** `search_after_date_filter`, `%m/%d/%Y`. */ + afterDate?: string; + /** `search_before_date_filter`, `%m/%d/%Y`. */ + beforeDate?: string; + /** `search_language_filter`: ISO 639-1 two-letter codes. */ + languageFilter?: string[]; +} + +/** + * Bare host of a `site:` value (`github.com/anthropics` → `github.com`); + * Perplexity's domain filter takes hosts only, the path part is enforced by + * the central lenient post-filter. + */ +function siteHost(site: string): string { + const slash = site.indexOf("/"); + return slash === -1 ? site : site.slice(0, slash); +} + +/** ISO `YYYY-MM-DD` → Perplexity's documented `%m/%d/%Y` date-filter format (e.g. `3/1/2025`). */ +function toPerplexityDate(iso: string): string { + const [year, month, day] = iso.split("-"); + return `${Number(month)}/${Number(day)}/${year}`; +} + +/** Map parsed query directives onto native Perplexity search filters. */ +function buildNativeFilters(parsed: StructuredQuery, rawQuery: string): PerplexityNativeFilters { + if (!parsed.hasDirectives) return { query: rawQuery }; + // Allow + deny share one array; the API caps it at 20 entries. + const domains = [ + ...new Set([...parsed.sites.map(siteHost), ...parsed.excludedSites.map(site => `-${siteHost(site)}`)]), + ].slice(0, 20); + // search_language_filter takes ISO 639-1 two-letter codes; pass `en-us` as + // `en`, and leave anything else to the central post-filter. + const langCode = parsed.lang ? /^([a-z]{2})(?:[-_]|$)/.exec(parsed.lang)?.[1] : undefined; + return { + query: formatQuery(parsed, PERPLEXITY_QUERY_SYNTAX), + domainFilter: domains.length > 0 ? domains : undefined, + afterDate: parsed.after ? toPerplexityDate(parsed.after) : undefined, + beforeDate: parsed.before ? toPerplexityDate(parsed.before) : undefined, + languageFilter: langCode ? [langCode] : undefined, + }; +} + interface PerplexityOAuthStreamMarkdownBlock { answer?: string; chunks?: string[]; @@ -258,6 +324,8 @@ export interface PerplexitySearchParams { signal?: AbortSignal; query: string; system_prompt?: string; + /** Pre-parsed view of `query` from the search pipeline; parsed locally when absent. */ + parsedQuery?: StructuredQuery; search_recency_filter?: "hour" | "day" | "week" | "month" | "year"; num_results?: number; /** Maximum output tokens. Defaults to 8192. */ @@ -352,6 +420,10 @@ function buildPerplexityExtraBody(request: PerplexityRequest): Record { const requestId = crypto.randomUUID(); // The consumer `perplexity_ask` endpoint is itself a research assistant and @@ -539,7 +612,7 @@ async function callPerplexityAsk( // "I don't have access to web-search tools in this turn", so ask-endpoint // searches send the bare query. (The API-key path still uses system_prompt // as a proper `system` message.) - const effectiveQuery = params.query; + const effectiveQuery = filters.query; const headers: Record = { "Content-Type": "application/json", @@ -578,7 +651,9 @@ async function callPerplexityAsk( version: OAUTH_API_VERSION, language: "en-US", timezone: Intl.DateTimeFormat().resolvedOptions().timeZone, - search_recency_filter: params.search_recency_filter ?? null, + // Recency cannot be combined with absolute date filters; explicit + // before:/after: bounds take precedence. + search_recency_filter: filters.afterDate || filters.beforeDate ? null : (params.search_recency_filter ?? null), is_incognito: true, use_schematized_api: true, // `true` (the native app's default) lets the backend classifier skip @@ -603,6 +678,10 @@ async function callPerplexityAsk( if (auth.type === "anonymous") { requestParams.send_back_text_in_streaming_api = true; } + if (filters.domainFilter) requestParams.search_domain_filter = filters.domainFilter; + if (filters.afterDate) requestParams.search_after_date_filter = filters.afterDate; + if (filters.beforeDate) requestParams.search_before_date_filter = filters.beforeDate; + if (filters.languageFilter) requestParams.search_language_filter = filters.languageFilter; const requestInit = { method: "POST", @@ -781,12 +860,14 @@ function applySourceLimit(result: SearchResponse, limit?: number): SearchRespons /** Execute Perplexity web search */ export async function searchPerplexity(params: PerplexitySearchParams): Promise { + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + const filters = buildNativeFilters(parsed, params.query); const systemPrompt = params.system_prompt; const messages: PerplexityRequest["messages"] = []; if (systemPrompt) { messages.push({ role: "system", content: systemPrompt }); } - messages.push({ role: "user", content: params.query }); + messages.push({ role: "user", content: filters.query }); const request: PerplexityRequest = { model: "sonar-pro", @@ -805,7 +886,13 @@ export async function searchPerplexity(params: PerplexitySearchParams): Promise< return_related_questions: true, }; - if (params.search_recency_filter) { + if (filters.domainFilter) request.search_domain_filter = filters.domainFilter; + if (filters.afterDate) request.search_after_date_filter = filters.afterDate; + if (filters.beforeDate) request.search_before_date_filter = filters.beforeDate; + if (filters.languageFilter) request.search_language_filter = filters.languageFilter; + // The API rejects search_recency_filter combined with absolute date + // filters; explicit before:/after: bounds take precedence. + if (params.search_recency_filter && !filters.afterDate && !filters.beforeDate) { request.search_recency_filter = params.search_recency_filter; } @@ -830,10 +917,10 @@ export async function searchPerplexity(params: PerplexitySearchParams): Promise< ? await withOAuthAccess( params.authStorage, "perplexity", - access => callPerplexityAsk({ type: "oauth", token: access.accessToken }, params), + access => callPerplexityAsk({ type: "oauth", token: access.accessToken }, params, filters), { sessionId: params.sessionId, signal: params.signal, seed: auth.access }, ) - : await callPerplexityAsk(auth, params); + : await callPerplexityAsk(auth, params, filters); return applySourceLimit( { provider: "perplexity", @@ -892,6 +979,7 @@ export class PerplexityProvider extends SearchProvider { return searchPerplexity({ signal: params.signal, query: params.query, + parsedQuery: params.parsedQuery, temperature: params.temperature, max_tokens: params.maxOutputTokens, num_search_results: params.numSearchResults, diff --git a/packages/coding-agent/src/web/search/providers/searxng.ts b/packages/coding-agent/src/web/search/providers/searxng.ts index 9decfe9c7..b5e4119e5 100644 --- a/packages/coding-agent/src/web/search/providers/searxng.ts +++ b/packages/coding-agent/src/web/search/providers/searxng.ts @@ -14,6 +14,9 @@ * searxng.basicUsername - Optional RFC 7617 Basic auth username * searxng.basicPassword - Optional RFC 7617 Basic auth password * searxng.categories - Optional comma-separated categories filter + * searxng.engines - Optional comma-separated engine names or shortcuts + * (e.g. "duckduckgo, br, sp"); shortcuts resolve via + * the instance's /config endpoint * searxng.language - Optional language code (e.g. en, zh-CN) * * Environment variable fallbacks: @@ -22,6 +25,11 @@ * SEARXNG_BASIC_USERNAME - Optional RFC 7617 Basic auth username * SEARXNG_BASIC_PASSWORD - Optional RFC 7617 Basic auth password * + * Bang syntax in queries is passed through: `!ddg foo` selects an engine or + * category server-side and the bang token is stripped from the upstream query. + * External bangs (`!!g`) are removed client-side because SearXNG answers them + * with an HTTP redirect even for JSON requests. + * * Reference: https://docs.searxng.org/dev/search_api.html */ @@ -30,6 +38,8 @@ import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai"; import { settings } from "../../../config/settings"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import type { StructuredQuery } from "../query"; +import { formatScraperQuery, parseSearchQuery } from "../query"; import { clampNumResults, dateToAgeSeconds } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -73,6 +83,11 @@ interface SearXNGAuth { value: string; } +/** Subset of the SearXNG /config payload used for engine shortcut resolution. */ +interface SearXNGConfig { + engines?: Array<{ name?: string; shortcut?: string }>; +} + /** Find SearXNG endpoint from settings or environment. */ function findEndpoint(): string | null { try { @@ -150,6 +165,110 @@ function findAuth(): SearXNGAuth | null { return token ? { type: "bearer", value: token } : null; } +/** Find configured engine names/shortcuts from settings. */ +function findEngines(): string | null { + try { + const engines = settings.get("searxng.engines"); + if (engines) return engines; + } catch { + // Settings not initialized yet + } + return null; +} + +/** Build request headers including authentication. */ +function buildHeaders(auth: SearXNGAuth | null): Record { + const headers: Record = { Accept: "application/json" }; + if (auth?.type === "basic") { + headers.Authorization = `Basic ${auth.value}`; + } else if (auth?.type === "bearer") { + headers.Authorization = `Bearer ${auth.value}`; + } + return headers; +} + +/** Per-endpoint cache of shortcut/name → canonical engine name maps. */ +const engineNameMapCache = new Map | null>>(); + +/** Fetch the instance's /config and build a lookup of lowercased engine names + * and shortcuts to canonical engine names. Returns null on any failure. */ +async function fetchEngineNameMap( + base: string, + auth: SearXNGAuth | null, + fetchImpl: FetchImpl | undefined, + signal: AbortSignal | undefined, +): Promise | null> { + try { + const response = await (fetchImpl ?? fetch)(`${base}/config`, { + headers: buildHeaders(auth), + signal: withHardTimeout(signal), + }); + if (!response.ok) return null; + const config = (await response.json()) as SearXNGConfig; + const map = new Map(); + for (const engine of config.engines ?? []) { + if (!engine.name) continue; + map.set(engine.name.toLowerCase(), engine.name); + if (engine.shortcut) map.set(engine.shortcut.toLowerCase(), engine.name); + } + return map.size ? map : null; + } catch { + return null; + } +} + +/** Get the engine name map for an endpoint, cached for the process lifetime. + * Failures are not cached so a transient error retries on the next search. */ +function getEngineNameMap( + endpoint: string, + auth: SearXNGAuth | null, + fetchImpl: FetchImpl | undefined, + signal: AbortSignal | undefined, +): Promise | null> { + const base = endpoint.replace(/\/+$/, ""); + let cached = engineNameMapCache.get(base); + if (!cached) { + cached = fetchEngineNameMap(base, auth, fetchImpl, signal).then(map => { + if (!map) engineNameMapCache.delete(base); + return map; + }); + engineNameMapCache.set(base, cached); + } + return cached; +} + +/** Resolve configured engine entries (canonical names or shortcuts like `ddg`) + * to canonical names for SearXNG's `engines=` parameter, which accepts names + * only — shortcuts resolve exclusively through bang syntax. Unknown entries + * pass through verbatim; the server drops them and falls back to categories. */ +async function resolveEngineNames( + raw: string, + endpoint: string, + auth: SearXNGAuth | null, + fetchImpl: FetchImpl | undefined, + signal: AbortSignal | undefined, +): Promise { + const entries = raw + .split(",") + .map(entry => entry.trim()) + .filter(Boolean); + if (!entries.length) return undefined; + const map = await getEngineNameMap(endpoint, auth, fetchImpl, signal); + if (!map) return entries.join(","); + return entries.map(entry => map.get(entry.toLowerCase()) ?? entry).join(","); +} + +/** Strip external bang tokens (`!!g`, bare `!!`): SearXNG answers them with an + * HTTP redirect even for JSON requests, which breaks response parsing. + * Single-bang engine/category selectors (`!ddg`, `!images`) are kept — the + * instance resolves and removes them server-side. */ +function stripExternalBangs(query: string): string { + return query + .split(/\s+/) + .filter(part => !part.startsWith("!!")) + .join(" "); +} + /** Build the search URL and headers for a SearXNG request */ function buildRequest( endpoint: string, @@ -158,6 +277,7 @@ function buildRequest( num_results?: number; recency?: "day" | "week" | "month" | "year"; categories?: string; + engines?: string; language?: string; signal?: AbortSignal; }, @@ -181,19 +301,15 @@ function buildRequest( url.searchParams.set("categories", params.categories); } + if (params.engines) { + url.searchParams.set("engines", params.engines); + } + if (params.language) { url.searchParams.set("language", params.language); } - const headers: Record = { - Accept: "application/json", - }; - - if (auth?.type === "basic") { - headers.Authorization = `Basic ${auth.value}`; - } else if (auth?.type === "bearer") { - headers.Authorization = `Bearer ${auth.value}`; - } + const headers = buildHeaders(auth); return { url, headers }; } @@ -205,6 +321,7 @@ async function callSearXNGSearch( num_results?: number; recency?: "day" | "week" | "month" | "year"; categories?: string; + engines?: string; language?: string; signal?: AbortSignal; fetch?: FetchImpl; @@ -231,6 +348,7 @@ async function callSearXNGSearch( /** Execute SearXNG web search. */ export async function searchSearXNG(params: { query: string; + parsedQuery?: StructuredQuery; num_results?: number; recency?: "day" | "week" | "month" | "year"; signal?: AbortSignal; @@ -255,12 +373,29 @@ export async function searchSearXNG(params: { } catch { // Settings not initialized yet } + const configuredEngines = findEngines(); + + // SearXNG forwards `q` to downstream engines, so build it with the shared + // scraper formatter: operators are canonicalized and scraper-hostile ones + // (path-carrying `site:`, `inurl:`) are structurally demoted to plain + // terms before formatting, so paren-grouped `site:` filters are covered + // too. `lang:` maps onto the native `language` param (overriding the + // configured default). + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + const query = formatScraperQuery(params.query, parsed); + if (parsed.lang) language = parsed.lang; + + const engines = configuredEngines + ? await resolveEngineNames(configuredEngines, endpoint, auth, params.fetch, params.signal) + : undefined; const response = await callSearXNGSearch( endpoint, { ...params, + query: stripExternalBangs(query), categories, + engines, language, fetch: params.fetch, }, @@ -315,6 +450,7 @@ export class SearXNGProvider extends SearchProvider { search(params: SearchParams): Promise { return searchSearXNG({ + parsedQuery: params.parsedQuery, query: params.query, num_results: params.numSearchResults ?? params.limit, recency: params.recency, diff --git a/packages/coding-agent/src/web/search/providers/startpage.ts b/packages/coding-agent/src/web/search/providers/startpage.ts index be8cacb5f..5076f6612 100644 --- a/packages/coding-agent/src/web/search/providers/startpage.ts +++ b/packages/coding-agent/src/web/search/providers/startpage.ts @@ -2,6 +2,7 @@ import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai"; import { parseHTML } from "linkedom"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatScraperQuery } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -136,12 +137,17 @@ async function callStartpageHtml(params: SearchParams): Promise { const fetchImpl = params.fetch ?? fetch; const signal = withHardTimeout(params.signal); const withDate = params.recency ? RECENCY_TO_STARTPAGE_WITH_DATE[params.recency] : undefined; + // Startpage proxies Google, so the operator set works inline; rebuild via + // the shared scraper formatter to canonicalize aliases (domain: → site:, + // since: → after:, …) and demote scraper-hostile operators. Directive-free + // queries pass through byte-identical. + const query = formatScraperQuery(params.query, params.parsedQuery); const formInputs = await fetchFormInputs(fetchImpl, signal); let page: LoadedHtmlPage; if (formInputs) { const form = new URLSearchParams(formInputs); - form.set("query", params.query); + form.set("query", query); if (withDate) form.set("with_date", withDate); page = await browserFetch(STARTPAGE_SEARCH_URL, { fetch: fetchImpl, @@ -152,7 +158,7 @@ async function callStartpageHtml(params: SearchParams): Promise { }); } else { const url = new URL(STARTPAGE_SEARCH_URL); - url.searchParams.set("query", params.query); + url.searchParams.set("query", query); if (withDate) url.searchParams.set("with_date", withDate); page = await browserFetch(url.href, { fetch: fetchImpl, diff --git a/packages/coding-agent/src/web/search/providers/synthetic.ts b/packages/coding-agent/src/web/search/providers/synthetic.ts index 4a84c3547..ad4397eb2 100644 --- a/packages/coding-agent/src/web/search/providers/synthetic.ts +++ b/packages/coding-agent/src/web/search/providers/synthetic.ts @@ -8,6 +8,7 @@ import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth } from "@oh-my-pi/pi-ai"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery } from "../query"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; import { classifyProviderHttpError, withHardTimeout } from "./utils"; @@ -73,8 +74,13 @@ export async function searchSynthetic(params: SearchParamsWithFetch): Promise callSyntheticSearch(key, params.query, params.signal, fetchImpl), { + const data = await withAuth(keyOrResolver, key => callSyntheticSearch(key, query, params.signal, fetchImpl), { signal: params.signal, missingKeyMessage: "Synthetic credentials not found. Set SYNTHETIC_API_KEY or login with 'omp /login synthetic'.", }); diff --git a/packages/coding-agent/src/web/search/providers/tavily.ts b/packages/coding-agent/src/web/search/providers/tavily.ts index b4cd5258f..fd53f9fc9 100644 --- a/packages/coding-agent/src/web/search/providers/tavily.ts +++ b/packages/coding-agent/src/web/search/providers/tavily.ts @@ -7,6 +7,7 @@ import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth } from "@oh-my-pi/pi-ai"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery } from "../query"; import { clampNumResults, dateToAgeSeconds } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -20,6 +21,14 @@ export interface TavilySearchParams { query: string; num_results?: number; recency?: "day" | "week" | "month" | "year"; + /** `site:` hosts mapped to Tavily's `include_domains`. */ + include_domains?: string[]; + /** `-site:` hosts mapped to Tavily's `exclude_domains`. */ + exclude_domains?: string[]; + /** `after:` inclusive lower bound, ISO `YYYY-MM-DD`, mapped to `start_date`. */ + start_date?: string; + /** `before:` upper bound, ISO `YYYY-MM-DD`, mapped to `end_date`. */ + end_date?: string; signal?: AbortSignal; fetch?: FetchImpl; } @@ -83,7 +92,21 @@ export function buildRequestBody(params: TavilySearchParams): Record 0; } +/** Bare hosts from `site:` values (path parts are enforced by the central lenient filter). */ +function siteHosts(sites: readonly string[]): string[] { + const hosts = new Set(); + for (const site of sites) { + const host = site.split("/", 1)[0]; + if (host) hosts.add(host); + } + return [...hosts]; +} + /** Execute Tavily web search. */ export async function searchTavily(params: SearchParams): Promise { + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); const tavilyParams: TavilySearchParams = { query: params.query, num_results: params.numSearchResults ?? params.limit, @@ -157,6 +191,16 @@ export async function searchTavily(params: SearchParams): Promise 0) tavilyParams.include_domains = include; + if (exclude.length > 0) tavilyParams.exclude_domains = exclude; + if (parsed.after) tavilyParams.start_date = parsed.after; + if (parsed.before) tavilyParams.end_date = parsed.before; + } const keyOrResolver: ApiKey = params.authStorage.resolver("tavily", { sessionId: params.sessionId, }); @@ -171,11 +215,16 @@ export async function searchTavily(params: SearchParams): Promise callTavilySearch(key, searchParams), authOptions); const response = toSearchResponse(await callWithAuth(tavilyParams), numResults); - if (!tavilyParams.recency || hasRenderableResponse(response)) { + const hasTimeFilter = Boolean(tavilyParams.recency || tavilyParams.start_date || tavilyParams.end_date); + if (!hasTimeFilter || hasRenderableResponse(response)) { return response; } - return toSearchResponse(await callWithAuth({ ...tavilyParams, recency: undefined }), numResults); + // Time filters commonly zero out results; retry once without them. + return toSearchResponse( + await callWithAuth({ ...tavilyParams, recency: undefined, start_date: undefined, end_date: undefined }), + numResults, + ); } /** Search provider for Tavily web search. */ diff --git a/packages/coding-agent/src/web/search/providers/tinyfish.ts b/packages/coding-agent/src/web/search/providers/tinyfish.ts index ba1b968fa..e96d42012 100644 --- a/packages/coding-agent/src/web/search/providers/tinyfish.ts +++ b/packages/coding-agent/src/web/search/providers/tinyfish.ts @@ -7,6 +7,7 @@ import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth } from "@oh-my-pi/pi-ai"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery, type QuerySyntax } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -17,6 +18,9 @@ const DEFAULT_NUM_RESULTS = 10; const MAX_NUM_RESULTS = 20; const MAX_PAGE = 10; +/** TinyFish is SERP-backed: common Google-style operators pass through. */ +const TINYFISH_QUERY_SYNTAX: QuerySyntax = { phrases: true, negation: true, site: true, filetype: true }; + const RECENCY_MINUTES: Record, number> = { day: 1440, week: 10080, @@ -107,8 +111,9 @@ function appendTinyFishSources(sources: SearchSource[], results: readonly TinyFi export async function searchTinyFish(params: SearchParams): Promise { const numResults = clampNumResults(params.numSearchResults ?? params.limit, DEFAULT_NUM_RESULTS, MAX_NUM_RESULTS); const pageSize = Math.min(numResults, DEFAULT_NUM_RESULTS); + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); const tinyFishParams: TinyFishSearchParams = { - query: params.query, + query: parsed.hasDirectives ? formatQuery(parsed, TINYFISH_QUERY_SYNTAX) : params.query, num_results: pageSize, recency: params.recency, signal: params.signal, diff --git a/packages/coding-agent/src/web/search/providers/xai.ts b/packages/coding-agent/src/web/search/providers/xai.ts index 5dd5e9c6d..8ef2d6f44 100644 --- a/packages/coding-agent/src/web/search/providers/xai.ts +++ b/packages/coding-agent/src/web/search/providers/xai.ts @@ -3,6 +3,7 @@ import { $env } from "@oh-my-pi/pi-utils"; import { resolveXAIHttpTransport, type XAIHttpProvider, type XAIHttpTransport } from "../../../lib/xai-http"; import type { SearchCitation, SearchResponse, SearchSource, SearchUsage } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery, type QuerySyntax } from "../query"; import { clampNumResults } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -57,14 +58,60 @@ interface XAIResponsesResponse { usage?: XAIResponsesUsage | null; } +/** + * Query syntax re-emitted for the Grok search agent. `site:`/`-site:` are + * stripped because hosts map natively onto the web_search domain filters; + * `before:`/`after:` stay in the query text — the Responses web_search tool + * has no date parameters (`from_date`/`to_date` exist only on `x_search` and + * the deprecated Live Search `search_parameters`, which now returns 410) and + * the agent honors the tokens as natural-language hints. + */ +const XAI_QUERY_SYNTAX: QuerySyntax = { + phrases: true, + negation: true, + or: true, + inUrl: true, + inTitle: true, + filetype: true, + dateRange: true, +}; + +/** xAI web_search accepts at most 5 allowed or excluded domains per request. */ +const MAX_DOMAIN_FILTERS = 5; + +/** Bare hosts of `site:` values (`github.com/anthropics` → `github.com`), deduped, capped at 5; path parts are enforced by the central constraint filter. */ +function domainFilterList(sites: readonly string[]): string[] { + const hosts = new Set(); + for (const site of sites) { + const slash = site.indexOf("/"); + hosts.add(slash === -1 ? site : site.slice(0, slash)); + if (hosts.size === MAX_DOMAIN_FILTERS) break; + } + return [...hosts]; +} + function buildRequestBody(params: SearchParams): Record { + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); + const webSearchTool: Record = { type: "web_search" }; + let query = params.query; + if (parsed.hasDirectives) { + query = formatQuery(parsed, XAI_QUERY_SYNTAX); + // allowed_domains and excluded_domains are mutually exclusive per + // request; prefer the allow list, the central filter enforces exclusions. + if (parsed.sites.length > 0) { + webSearchTool.filters = { allowed_domains: domainFilterList(parsed.sites) }; + } else if (parsed.excludedSites.length > 0) { + webSearchTool.filters = { excluded_domains: domainFilterList(parsed.excludedSites) }; + } + } + const body: Record = { model: XAI_WEB_SEARCH_MODEL, input: [ { role: "system", content: params.systemPrompt }, - { role: "user", content: params.query }, + { role: "user", content: query }, ], - tools: [{ type: "web_search" }], + tools: [webSearchTool], reasoning: { effort: XAI_WEB_SEARCH_REASONING_EFFORT }, }; diff --git a/packages/coding-agent/src/web/search/providers/zai.ts b/packages/coding-agent/src/web/search/providers/zai.ts index 9b264109d..9fbd92303 100644 --- a/packages/coding-agent/src/web/search/providers/zai.ts +++ b/packages/coding-agent/src/web/search/providers/zai.ts @@ -8,6 +8,7 @@ import { type ApiKey, type AuthStorage, type FetchImpl, getEnvApiKey, withAuth } import { isRecord } from "@oh-my-pi/pi-utils"; import type { SearchResponse, SearchSource } from "../../../web/search/types"; import { SearchProviderError } from "../../../web/search/types"; +import { formatQuery, parseSearchQuery, type QuerySyntax } from "../query"; import { dateToAgeSeconds } from "../utils"; import type { SearchParams } from "./base"; import { SearchProvider } from "./base"; @@ -17,6 +18,22 @@ const ZAI_MCP_URL = "https://api.z.ai/api/mcp/web_search_prime/mcp"; const ZAI_TOOL_NAME = "web_search_prime"; const DEFAULT_NUM_RESULTS = 10; +/** + * webSearchPrime exposes no native filter args (the Web Search API's + * `search_domain_filter`/`search_recency_filter` are scoped to the + * `search_pro_jina` engine, not `search-prime`), but its Bing-flavored + * backend parses the common inline operators. Dates and language are left + * to the central lenient post-filter. + */ +const ZAI_QUERY_SYNTAX: QuerySyntax = { + phrases: true, + negation: true, + site: true, + inTitle: true, + inUrl: true, + filetype: true, +}; + export interface ZaiSearchParams { query: string; num_results?: number; @@ -407,8 +424,9 @@ export class ZaiProvider extends SearchProvider { search(params: SearchParams): Promise { const { fetch: fetchOverride } = params as ZaiProviderSearchParams; + const parsed = params.parsedQuery ?? parseSearchQuery(params.query); return searchZai({ - query: params.query, + query: parsed.hasDirectives ? formatQuery(parsed, ZAI_QUERY_SYNTAX) : params.query, num_results: params.numSearchResults ?? params.limit, signal: params.signal, authStorage: params.authStorage, diff --git a/packages/coding-agent/src/web/search/query.ts b/packages/coding-agent/src/web/search/query.ts new file mode 100644 index 000000000..984efd5c0 --- /dev/null +++ b/packages/coding-agent/src/web/search/query.ts @@ -0,0 +1,850 @@ +/** + * Structured web-search query parsing. + * + * Agents habitually embed Google-style directives in search queries — + * `site:`, `before:`/`after:`, `inurl:`, `filetype:`, quoted phrases, `OR` + * groups, `-exclusions` — regardless of whether the backing engine parses + * them. This module turns a raw query into a {@link StructuredQuery} so each + * provider can: + * + * 1. map constraints onto native API parameters where they exist (Perplexity + * `search_domain_filter`, Tavily `include_domains`, Exa date bounds, …), + * 2. rebuild a query string containing only the syntax the target engine + * understands ({@link formatQuery}), and + * 3. post-filter returned sources leniently ({@link applyQueryConstraints}): + * a constraint dimension that would eliminate every result is dropped and + * reported rather than returning nothing. + */ + +import type { SearchSource } from "./types"; + +/** One free-text token of the query (everything that is not a recognized directive). */ +export interface QueryTerm { + /** Term text without quotes or operator prefixes. */ + text: string; + /** Quoted exact phrase (`"like this"`) or verbatim-required (`+term`). */ + phrase?: boolean; + /** Excluded via `-term` or `NOT term`. */ + negated?: boolean; + /** + * OR-group id. Terms sharing an id are alternatives (`a OR b`); terms + * without a group are implicitly AND-ed. Groups are always contiguous + * runs in {@link StructuredQuery.terms}. + */ + group?: number; +} + +/** + * A raw query decomposed into free text plus every recognized constraint. + * + * All list fields are always present (possibly empty) so consumers can map + * over them without null checks. Values are stored as typed by the user + * except for normalization noted per field. + */ +export interface StructuredQuery { + /** Original query string, verbatim. */ + raw: string; + /** + * Free-text remainder with all recognized directives removed; phrases + * stay quoted, exclusions keep `-`, OR groups keep `OR`. Empty when the + * query was directives only — use {@link formatQuery} for a never-empty + * engine query. + */ + text: string; + /** Ordered free-text terms (phrases, exclusions, OR groups). */ + terms: QueryTerm[]; + /** `site:`/`domain:`/`host:` includes — any-of. Lowercased, scheme stripped, may carry a path (`github.com/anthropics`). */ + sites: string[]; + /** `-site:` exclusions, same normalization as {@link sites}. */ + excludedSites: string[]; + /** `inurl:`/`url:`/`allinurl:` substrings — all must appear in the URL. */ + inUrl: string[]; + /** `-inurl:` substrings — none may appear in the URL. */ + excludedInUrl: string[]; + /** `intitle:`/`title:`/`allintitle:` substrings — all must appear in the title. */ + inTitle: string[]; + /** `-intitle:` substrings — none may appear in the title. */ + excludedInTitle: string[]; + /** `intext:`/`inbody:`/`inanchor:`/`allintext:` body substrings. Not post-filterable (snippets are partial); query-building only. */ + inText: string[]; + /** `-intext:` body exclusions. Query-building only. */ + excludedInText: string[]; + /** `filetype:`/`ext:` extensions — any-of. Lowercased, no leading dot. */ + filetypes: string[]; + /** `-filetype:`/`-ext:` extensions — none may match. */ + excludedFiletypes: string[]; + /** Inclusive lower publish-date bound from `after:`/`since:`, ISO `YYYY-MM-DD`. */ + after?: string; + /** Exclusive upper publish-date bound from `before:`/`until:`, ISO `YYYY-MM-DD`. */ + before?: string; + /** Language code from `lang:`/`language:`, lowercased (e.g. `en`, `en-us`). */ + lang?: string; + /** True when any directive or boolean operator was recognized. */ + hasDirectives: boolean; + /** True when any post-filterable constraint is set (sites, url/title terms, filetypes, date bounds). */ + hasConstraints: boolean; +} + +/** + * Query-syntax capabilities of a target engine, used by {@link formatQuery} + * to decide which parsed features are re-emitted as query text. Everything + * defaults to `false`: the zero-value produces plain keywords suitable for + * natural-language APIs. + */ +export interface QuerySyntax { + /** Emit `"quoted phrases"`. */ + phrases?: boolean; + /** Emit `-term` exclusions (negated terms are dropped otherwise). */ + negation?: boolean; + /** Emit `OR` between alternatives (groups are flattened to keywords otherwise). */ + or?: boolean; + /** Emit `site:`/`-site:`. */ + site?: boolean; + /** Emit `inurl:`/`-inurl:`. */ + inUrl?: boolean; + /** Emit `intitle:`/`-intitle:`. */ + inTitle?: boolean; + /** Emit `intext:`/`-intext:`. */ + inText?: boolean; + /** Emit `filetype:`/`-filetype:`. */ + filetype?: boolean; + /** Emit `before:`/`after:` ISO date bounds. */ + dateRange?: boolean; +} + +/** Full Google-style syntax: engines that parse the classic operator set (Google, Startpage, Ecosia, Brave, Kagi, Mojeek, SearXNG…). */ +export const GOOGLE_QUERY_SYNTAX: QuerySyntax = { + phrases: true, + negation: true, + or: true, + site: true, + inUrl: true, + inTitle: true, + inText: true, + filetype: true, + dateRange: true, +}; + +/** Result of {@link applyQueryConstraints}. */ +export interface ConstraintFilterResult { + /** Sources surviving the lenient filter — never empty when the input was non-empty. */ + sources: SearchSource[]; + /** + * Directive renderings (`site:arxiv.org`, `before:2024-01-01`, …) of the + * constraint dimensions that matched zero sources and were therefore + * relaxed instead of enforced. + */ + dropped: string[]; +} + +const DIRECTIVE_PATTERN = /^([+-]?)([a-z][a-z-]*):(.*)$/i; + +type AllMode = "inTitle" | "inUrl" | "inText"; + +interface RawToken { + text: string; + /** Entire token was a quoted phrase. */ + quoted: boolean; + /** Directive value was quoted (`intitle:"a b"`). */ + quotedValue?: boolean; +} + +function isQuote(ch: string): boolean { + return ch === '"' || ch === "\u201c" || ch === "\u201d"; +} + +/** Unicode-aware whitespace (agents paste NBSP and friends). */ +const WHITESPACE = /\s/; + +/** Split a raw query into whitespace-delimited tokens, honoring quoted spans and standalone parens. */ +function tokenize(raw: string): RawToken[] { + const tokens: RawToken[] = []; + const n = raw.length; + let i = 0; + while (i < n) { + const ch = raw[i]; + if (WHITESPACE.test(ch)) { + i++; + continue; + } + if (isQuote(ch)) { + let j = i + 1; + let buf = ""; + while (j < n && !isQuote(raw[j])) { + buf += raw[j]; + j++; + } + if (buf.trim().length > 0) tokens.push({ text: buf.trim(), quoted: true }); + i = j + 1; + continue; + } + // Bare word; a quote directly after `name:` swallows the quoted span + // into the same token (`intitle:"budget tips"`). + let buf = ""; + let quotedValue = false; + while (i < n && !WHITESPACE.test(raw[i])) { + const c = raw[i]; + if (isQuote(c) && buf.endsWith(":")) { + let j = i + 1; + while (j < n && !isQuote(raw[j])) { + buf += raw[j]; + j++; + } + quotedValue = true; + i = j + 1; + continue; + } + if (isQuote(c)) break; // `foo"bar` — stop the word, let the quote start a phrase + buf += c; + i++; + } + if (buf.length > 0) tokens.push({ text: buf, quoted: false, quotedValue }); + } + return splitParens(tokens); +} + +/** + * Split leading `(` and unbalanced trailing `)` into standalone tokens so + * `(react OR vue)` parses while `site:wikipedia.org/Foo_(bar)` stays whole. + */ +function splitParens(tokens: RawToken[]): RawToken[] { + const out: RawToken[] = []; + for (const tok of tokens) { + if (tok.quoted || tok.quotedValue) { + out.push(tok); + continue; + } + let text = tok.text; + while (text.startsWith("(")) { + out.push({ text: "(", quoted: false }); + text = text.slice(1); + } + let trailing = 0; + while (text.endsWith(")")) { + // Only strip parens that do not close an opener inside the word. + const body = text.slice(0, -1); + let depth = 0; + for (const c of body) { + if (c === "(") depth++; + else if (c === ")") depth--; + } + if (depth > 0) break; + text = body; + trailing++; + } + if (text.length > 0) out.push({ text, quoted: false }); + for (let k = 0; k < trailing; k++) out.push({ text: ")", quoted: false }); + } + return out; +} + +/** Convert year/month/day parts to a validated ISO date, or undefined. */ +function isoDate(year: number, month: number, day: number): string | undefined { + if (year < 1000 || year > 9999 || month < 1 || month > 12 || day < 1 || day > 31) return undefined; + return `${year}-${String(month).padStart(2, "0")}-${String(day).padStart(2, "0")}`; +} + +/** + * Parse a `before:`/`after:` value into ISO `YYYY-MM-DD`. + * Accepts `YYYY`, `YYYY-MM`, `YYYY-MM-DD` (also `/` and `.` separators) and + * `MM/DD/YYYY` (day-first assumed when the first field exceeds 12). + * Bare years/months resolve to the first day of the period, matching + * Google's `after:2024` ≙ `after:2024-01-01` semantics. + */ +export function parseDateValue(value: string): string | undefined { + const t = value.trim(); + let m = /^(\d{4})(?:[-/.](\d{1,2})(?:[-/.](\d{1,2}))?)?$/.exec(t); + if (m) return isoDate(Number(m[1]), m[2] ? Number(m[2]) : 1, m[3] ? Number(m[3]) : 1); + m = /^(\d{1,2})[-/.](\d{1,2})[-/.](\d{4})$/.exec(t); + if (m) { + let month = Number(m[1]); + let day = Number(m[2]); + if (month > 12 && day <= 12) [month, day] = [day, month]; + return isoDate(Number(m[3]), month, day); + } + return undefined; +} + +/** Lowercase a `site:` value and strip scheme, `*.` wildcard, and trailing slash/dot. */ +function normalizeSite(value: string): string { + let site = value.trim().toLowerCase(); + site = site.replace(/^[a-z][a-z0-9+.-]*:\/\//, ""); + if (site.startsWith("*.")) site = site.slice(2); + site = site.replace(/[/.]+$/, ""); + return site; +} + +/** Directive names mapped to their canonical field. */ +const DIRECTIVE_FIELDS: Record< + string, + "site" | "inUrl" | "inTitle" | "inText" | "filetype" | "before" | "after" | "lang" +> = { + site: "site", + domain: "site", + host: "site", + inurl: "inUrl", + url: "inUrl", + intitle: "inTitle", + title: "inTitle", + intext: "inText", + inbody: "inText", + inanchor: "inText", + filetype: "filetype", + ext: "filetype", + before: "before", + until: "before", + after: "after", + since: "after", + lang: "lang", + language: "lang", +}; + +/** `allin*:` directives that capture every following plain term. */ +const ALL_MODES: Record = { + allintitle: "inTitle", + allinurl: "inUrl", + allintext: "inText", +}; + +/** True for operator/paren tokens and recognized directives — anything a bare `name:` must not adopt as its value. */ +function isReservedToken(text: string): boolean { + if ( + text === "(" || + text === ")" || + text === "OR" || + text === "AND" || + text === "NOT" || + text === "|" || + text === "||" || + text === "&&" || + text === "!" + ) { + return true; + } + const m = DIRECTIVE_PATTERN.exec(text); + if (!m) return false; + const name = m[2].toLowerCase(); + return DIRECTIVE_FIELDS[name] !== undefined || ALL_MODES[name] !== undefined; +} + +/** + * Parse a raw query into a {@link StructuredQuery}. + * + * Lenient by construction: unknown `name:value` tokens (URLs, `C:\paths`, + * `TS2345:`, jargon) stay in the free text verbatim, and a directive with an + * unparseable value (`before:someday`) degrades to a plain term instead of + * being dropped. + */ +export function parseSearchQuery(raw: string): StructuredQuery { + const q: StructuredQuery = { + raw, + text: "", + terms: [], + sites: [], + excludedSites: [], + inUrl: [], + excludedInUrl: [], + inTitle: [], + excludedInTitle: [], + inText: [], + excludedInText: [], + filetypes: [], + excludedFiletypes: [], + hasDirectives: false, + hasConstraints: false, + }; + + const tokens = tokenize(raw); + let negateNext = false; + let orPending = false; + let lastWasTerm = false; + let groupSeq = 0; + let allMode: AllMode | undefined; + + const pushConstraint = ( + field: "site" | "inUrl" | "inTitle" | "inText" | "filetype", + value: string, + negated: boolean, + ): void => { + q.hasDirectives = true; + orPending = false; + lastWasTerm = false; + const v = value.trim(); + if (!v) return; + switch (field) { + case "site": { + const site = normalizeSite(v); + if (site) (negated ? q.excludedSites : q.sites).push(site); + break; + } + case "inUrl": + (negated ? q.excludedInUrl : q.inUrl).push(v); + break; + case "inTitle": + (negated ? q.excludedInTitle : q.inTitle).push(v); + break; + case "inText": + (negated ? q.excludedInText : q.inText).push(v); + break; + case "filetype": { + const ext = v.toLowerCase().replace(/^\.+/, ""); + if (ext) (negated ? q.excludedFiletypes : q.filetypes).push(ext); + break; + } + } + }; + + const pushTerm = (text: string, phrase: boolean): void => { + const negated = negateNext; + negateNext = false; + if (allMode && !negated) { + pushConstraint(allMode, text, false); + return; + } + if (allMode && negated) { + pushConstraint(allMode, text, true); + return; + } + const term: QueryTerm = { text }; + if (phrase) term.phrase = true; + if (negated) term.negated = true; + if (orPending && lastWasTerm) { + const prev = q.terms[q.terms.length - 1]; + if (prev) { + prev.group ??= ++groupSeq; + term.group = prev.group; + } + } + orPending = false; + lastWasTerm = true; + q.terms.push(term); + }; + + for (let idx = 0; idx < tokens.length; idx++) { + const tok = tokens[idx]; + + if (tok.quoted) { + pushTerm(tok.text, true); + continue; + } + + // Boolean operators and grouping parens. + if (tok.text === "(" || tok.text === ")") continue; + if (tok.text === "OR" || tok.text === "|" || tok.text === "||") { + orPending = true; + q.hasDirectives = true; + continue; + } + if (tok.text === "AND" || tok.text === "&&") { + q.hasDirectives = true; + continue; + } + if (tok.text === "NOT" || tok.text === "!") { + negateNext = true; + q.hasDirectives = true; + continue; + } + if (tok.text === "-" || tok.text === "+") { + // The tokenizer splits `-"exact phrase"` into `-` + phrase; carry the negation over. + if (tok.text === "-" && tokens[idx + 1]?.quoted) negateNext = true; + continue; + } + + const match = DIRECTIVE_PATTERN.exec(tok.text); + const name = match?.[2].toLowerCase(); + const allMatch = name ? ALL_MODES[name] : undefined; + const field = name ? DIRECTIVE_FIELDS[name] : undefined; + + if (match && allMatch) { + allMode = allMatch; + q.hasDirectives = true; + // `allintitle:budget tips` — inline value plus every following term. + const inline = match[3].trim(); + if (inline) pushConstraint(allMatch, inline, match[1] === "-"); + orPending = false; + lastWasTerm = false; + continue; + } + + if (match && field) { + let value = match[3].trim(); + // `site: example.com` — lenient: adopt the next plain token as the value. + if (!value) { + const next = tokens[idx + 1]; + if (next && (next.quoted || !isReservedToken(next.text))) { + value = next.text.trim(); + idx++; + } + } + if (!value) { + q.hasDirectives = true; + continue; + } + const negated = match[1] === "-" || negateNext; + negateNext = false; + switch (field) { + case "before": + case "after": { + const iso = parseDateValue(value); + if (!iso) { + pushTerm(tok.text, false); + continue; + } + if (field === "before") q.before = iso; + else q.after = iso; + q.hasDirectives = true; + orPending = false; + lastWasTerm = false; + break; + } + case "lang": + q.lang = value.toLowerCase(); + q.hasDirectives = true; + orPending = false; + lastWasTerm = false; + break; + default: + pushConstraint(field, value, negated); + } + continue; + } + + // Plain term with optional +/- prefix. + let text = tok.text; + if (text.startsWith("-") && text.length > 1) { + negateNext = true; + q.hasDirectives = true; + text = text.replace(/^-+/, ""); + if (!text) continue; + // `-site:x` arrives pre-split only when written `- site:x`; re-check directive. + const negMatch = DIRECTIVE_PATTERN.exec(text); + const negName = negMatch?.[2].toLowerCase(); + const negField = negName ? DIRECTIVE_FIELDS[negName] : undefined; + if (negMatch && negField && negField !== "before" && negField !== "after" && negField !== "lang") { + negateNext = false; + pushConstraint(negField, negMatch[3].trim(), true); + continue; + } + pushTerm(text, false); + continue; + } + if (text.startsWith("+") && text.length > 1) { + // Legacy Google `+term`: verbatim/required — treat as an exact phrase. + pushTerm(text.slice(1), true); + q.hasDirectives = true; + continue; + } + pushTerm(text, false); + } + + q.text = renderTerms(q.terms, { phrases: true, negation: true, or: true }); + q.hasConstraints = + q.sites.length > 0 || + q.excludedSites.length > 0 || + q.inUrl.length > 0 || + q.excludedInUrl.length > 0 || + q.inTitle.length > 0 || + q.excludedInTitle.length > 0 || + q.filetypes.length > 0 || + q.excludedFiletypes.length > 0 || + q.before !== undefined || + q.after !== undefined; + return q; +} + +/** Quote a directive value when it contains whitespace. */ +function quoteValue(value: string): string { + return /\s/.test(value) ? `"${value}"` : value; +} + +/** Render the free-text terms per the target syntax. */ +function renderTerms(terms: readonly QueryTerm[], syntax: QuerySyntax): string { + const parts: string[] = []; + for (let i = 0; i < terms.length; i++) { + const term = terms[i]; + if (term.group !== undefined && syntax.or) { + const members: string[] = []; + let j = i; + for (; j < terms.length && terms[j].group === term.group; j++) { + const rendered = renderTerm(terms[j], syntax); + if (rendered) members.push(rendered); + } + i = j - 1; + if (members.length > 1) parts.push(`(${members.join(" OR ")})`); + else if (members.length === 1) parts.push(members[0]); + continue; + } + const rendered = renderTerm(term, syntax); + if (rendered) parts.push(rendered); + } + return parts.join(" "); +} + +function renderTerm(term: QueryTerm, syntax: QuerySyntax): string | undefined { + if (term.negated && !syntax.negation) return undefined; + const body = term.phrase && syntax.phrases ? `"${term.text}"` : term.text; + return term.negated ? `-${body}` : body; +} + +/** + * Rebuild a query string for an engine with the given {@link QuerySyntax}. + * + * Constraints whose syntax the engine lacks are omitted (the caller maps + * them onto API parameters or relies on {@link applyQueryConstraints}). + * Never returns an empty string for a non-empty input: a directives-only + * query falls back to the constraint values as keywords, then to `raw` — an + * engine searching *something* beats an empty-query error. + */ +export function formatQuery(q: StructuredQuery, syntax: QuerySyntax = {}): string { + const parts: string[] = []; + const text = renderTerms(q.terms, syntax); + if (text) parts.push(text); + + if (syntax.site) { + if (q.sites.length > 1 && syntax.or) parts.push(`(${q.sites.map(s => `site:${s}`).join(" OR ")})`); + else parts.push(...q.sites.map(s => `site:${s}`)); + parts.push(...q.excludedSites.map(s => `-site:${s}`)); + } + if (syntax.inUrl) { + parts.push(...q.inUrl.map(v => `inurl:${quoteValue(v)}`)); + parts.push(...q.excludedInUrl.map(v => `-inurl:${quoteValue(v)}`)); + } + if (syntax.inTitle) { + parts.push(...q.inTitle.map(v => `intitle:${quoteValue(v)}`)); + parts.push(...q.excludedInTitle.map(v => `-intitle:${quoteValue(v)}`)); + } + if (syntax.inText) { + parts.push(...q.inText.map(v => `intext:${quoteValue(v)}`)); + parts.push(...q.excludedInText.map(v => `-intext:${quoteValue(v)}`)); + } + if (syntax.filetype) { + if (q.filetypes.length > 1 && syntax.or) parts.push(`(${q.filetypes.map(f => `filetype:${f}`).join(" OR ")})`); + else parts.push(...q.filetypes.map(f => `filetype:${f}`)); + parts.push(...q.excludedFiletypes.map(f => `-filetype:${f}`)); + } + if (syntax.dateRange) { + if (q.after) parts.push(`after:${q.after}`); + if (q.before) parts.push(`before:${q.before}`); + } + + let result = parts.join(" ").trim(); + if (!result) { + // Directives-only query and no directive syntax: search the constraint + // values as plain keywords so the engine still gets a meaningful query. + const fallback = [...q.sites, ...q.inTitle, ...q.inUrl, ...q.inText, ...q.filetypes]; + result = fallback.join(" ").trim(); + } + return result || q.raw.trim(); +} + +/** + * Build the engine query for a credential-free HTML engine (Google, + * Startpage, DuckDuckGo, Ecosia, Mojeek, SearXNG, and the Public Web + * fan-out over them). + * + * Canonicalizes directives via {@link formatQuery} with the engine's + * {@link QuerySyntax} (default: full Google syntax), after demoting the + * operators that zero-match across the scraper set: engines only match + * `site:` against a bare domain (a path yields zero results everywhere), + * and DuckDuckGo ignores `inurl:` entirely — so either operator silently + * empties the result set. The raw URL as a plain term matches fine, so + * bare-domain `site:` filters are kept while path-carrying `site:` and all + * `inurl:` values become plain keywords; the demotion is structural (before + * formatting), so OR-grouped and quoted directives are covered. Negated + * forms (`-site:`, `-inurl:`) pass through untouched — demoting them would + * invert an exclusion into a search term; the pipeline post-filter + * ({@link applyQueryConstraints}) enforces every demoted or unsupported + * constraint on the returned sources. Directive-free queries pass through + * byte-identical. + */ +export function formatScraperQuery( + query: string, + parsedQuery?: StructuredQuery, + syntax: QuerySyntax = GOOGLE_QUERY_SYNTAX, +): string { + const parsed = parsedQuery ?? parseSearchQuery(query); + if (!parsed.hasDirectives) return query; + const demoted = [...parsed.sites.filter(site => site.includes("/")), ...parsed.inUrl]; + const downgraded: StructuredQuery = { + ...parsed, + sites: parsed.sites.filter(site => !site.includes("/")), + inUrl: [], + terms: [...parsed.terms, ...demoted.map(text => ({ text }))], + }; + return formatQuery(downgraded, syntax); +} + +/** Hostname (lowercased) and pathname of a URL, or undefined when unparsable. */ +function hostAndPath(url: string): { host: string; path: string } | undefined { + try { + const u = new URL(url); + return { host: u.hostname.toLowerCase(), path: u.pathname }; + } catch { + return undefined; + } +} + +/** + * `site:` matcher: exact host or subdomain of `site`; when `site` carries a + * path (`github.com/anthropics`), the URL path must start with it. + */ +export function matchesSite(url: string, site: string): boolean { + const parsed = hostAndPath(url); + if (!parsed) return false; + const slash = site.indexOf("/"); + const siteHost = slash === -1 ? site : site.slice(0, slash); + const sitePath = slash === -1 ? "" : site.slice(slash); + if (parsed.host !== siteHost && !parsed.host.endsWith(`.${siteHost}`)) return false; + if (sitePath && !parsed.path.toLowerCase().startsWith(sitePath.toLowerCase())) return false; + return true; +} + +/** `filetype:` matcher: URL pathname ends with `.ext`. */ +function matchesFiletype(url: string, ext: string): boolean { + const parsed = hostAndPath(url); + if (!parsed) return false; + return parsed.path.toLowerCase().endsWith(`.${ext}`); +} + +const RELATIVE_AGE_PATTERN = /^(\d+)\s*(minute|min|hour|hr|day|week|month|mo|year|yr|[mhdwy])s?\s+ago$/i; + +const RELATIVE_UNIT_SECONDS: Record = { + m: 60, + min: 60, + minute: 60, + h: 3600, + hr: 3600, + hour: 3600, + d: 86_400, + day: 86_400, + w: 604_800, + week: 604_800, + mo: 2_592_000, + month: 2_592_000, + y: 31_536_000, + yr: 31_536_000, + year: 31_536_000, +}; + +/** Best-effort publish time (ms epoch) of a source from `ageSeconds`, ISO, or relative dates. */ +function sourceTime(source: SearchSource): number | undefined { + if (typeof source.ageSeconds === "number" && Number.isFinite(source.ageSeconds)) { + return Date.now() - source.ageSeconds * 1000; + } + if (!source.publishedDate) return undefined; + const rel = RELATIVE_AGE_PATTERN.exec(source.publishedDate.trim()); + if (rel) { + const seconds = Number(rel[1]) * (RELATIVE_UNIT_SECONDS[rel[2].toLowerCase()] ?? 0); + return seconds > 0 ? Date.now() - seconds * 1000 : undefined; + } + const parsed = Date.parse(source.publishedDate); + return Number.isNaN(parsed) ? undefined : parsed; +} + +/** + * Strict per-source constraint check: every filterable dimension of `q` must + * pass. Sources without a resolvable date pass date bounds (a missing date + * is not proof of violation). For custom provider flows; the standard path + * is {@link applyQueryConstraints}. + */ +export function matchesQueryConstraints(source: SearchSource, q: StructuredQuery): boolean { + for (const dim of constraintDimensions(q)) { + if (!dim.pred(source)) return false; + } + return true; +} + +interface ConstraintDimension { + /** Directive rendering for relaxation notes (`site:arxiv.org`). */ + label: string; + pred: (source: SearchSource) => boolean; +} + +function constraintDimensions(q: StructuredQuery): ConstraintDimension[] { + const dims: ConstraintDimension[] = []; + const lower = (s: string | undefined): string => (s ?? "").toLowerCase(); + + if (q.sites.length > 0) { + dims.push({ + label: q.sites.map(s => `site:${s}`).join(" OR "), + pred: src => q.sites.some(site => matchesSite(src.url, site)), + }); + } + if (q.excludedSites.length > 0) { + dims.push({ + label: q.excludedSites.map(s => `-site:${s}`).join(" "), + pred: src => !q.excludedSites.some(site => matchesSite(src.url, site)), + }); + } + if (q.inUrl.length > 0) { + dims.push({ + label: q.inUrl.map(v => `inurl:${v}`).join(" "), + pred: src => q.inUrl.every(v => lower(src.url).includes(v.toLowerCase())), + }); + } + if (q.excludedInUrl.length > 0) { + dims.push({ + label: q.excludedInUrl.map(v => `-inurl:${v}`).join(" "), + pred: src => !q.excludedInUrl.some(v => lower(src.url).includes(v.toLowerCase())), + }); + } + if (q.inTitle.length > 0) { + dims.push({ + label: q.inTitle.map(v => `intitle:${v}`).join(" "), + pred: src => q.inTitle.every(v => lower(src.title).includes(v.toLowerCase())), + }); + } + if (q.excludedInTitle.length > 0) { + dims.push({ + label: q.excludedInTitle.map(v => `-intitle:${v}`).join(" "), + pred: src => !q.excludedInTitle.some(v => lower(src.title).includes(v.toLowerCase())), + }); + } + if (q.filetypes.length > 0) { + dims.push({ + label: q.filetypes.map(f => `filetype:${f}`).join(" OR "), + pred: src => q.filetypes.some(ext => matchesFiletype(src.url, ext)), + }); + } + if (q.excludedFiletypes.length > 0) { + dims.push({ + label: q.excludedFiletypes.map(f => `-filetype:${f}`).join(" "), + pred: src => !q.excludedFiletypes.some(ext => matchesFiletype(src.url, ext)), + }); + } + if (q.after !== undefined || q.before !== undefined) { + const afterMs = q.after !== undefined ? Date.parse(q.after) : undefined; + const beforeMs = q.before !== undefined ? Date.parse(q.before) : undefined; + const label = [q.after ? `after:${q.after}` : "", q.before ? `before:${q.before}` : ""].filter(Boolean).join(" "); + dims.push({ + label, + pred: src => { + const time = sourceTime(src); + if (time === undefined) return true; // undated → cannot prove violation + if (afterMs !== undefined && time < afterMs) return false; + if (beforeMs !== undefined && time >= beforeMs) return false; + return true; + }, + }); + } + return dims; +} + +/** + * Lenient post-filter: applies each constraint dimension of `q` in turn, + * skipping (and reporting) any dimension that would eliminate every + * remaining source. Guarantees a non-empty result for a non-empty input, so + * a mis-scoped directive degrades to unfiltered results plus a note instead + * of a dead search. + */ +export function applyQueryConstraints(sources: readonly SearchSource[], q: StructuredQuery): ConstraintFilterResult { + let current = [...sources]; + const dropped: string[] = []; + if (current.length === 0) return { sources: current, dropped }; + for (const dim of constraintDimensions(q)) { + const kept = current.filter(dim.pred); + if (kept.length > 0) current = kept; + else dropped.push(dim.label); + } + return { sources: current, dropped }; +} diff --git a/packages/coding-agent/test/tools/web-search-codex.test.ts b/packages/coding-agent/test/tools/web-search-codex.test.ts index f4832faf9..99e95b8f4 100644 --- a/packages/coding-agent/test/tools/web-search-codex.test.ts +++ b/packages/coding-agent/test/tools/web-search-codex.test.ts @@ -275,6 +275,39 @@ describe("searchCodex model selection", () => { expect(result.sources).toEqual([{ title: "Example Article", url: "https://example.com/article" }]); }); + function sentUserText(): string | undefined { + const input = capturedRequest?.body?.input as Array> | undefined; + const userItem = input?.find(item => item.role === "user"); + const content = userItem?.content as Array> | undefined; + return content?.[0]?.text as string | undefined; + } + + it("re-emits directive queries with normalized Google-style operators", async () => { + delete process.env.PI_CODEX_WEB_SEARCH_MODEL; + await searchCodex( + makeSearchParams( + 'bun runtime site:bun.sh -site:reddit.com after:2024-01-01 "exact phrase"', + mockCodexFetch("gpt-5.6-luna"), + ), + ); + + expect(capturedRequest).not.toBeNull(); + expect(sentUserText()).toBe('bun runtime "exact phrase" site:bun.sh -site:reddit.com after:2024-01-01'); + // Tool config stays untouched: the ChatGPT backend's filter support is + // unverified, so no `filters` field is added to the web_search tool. + const input = capturedRequest?.body?.input as Array>; + const additionalTools = input.find(item => item.type === "additional_tools"); + expect(additionalTools?.tools).toEqual([{ type: "web_search", search_context_size: "high" }]); + }); + + it("sends directive-free queries byte-identical", async () => { + delete process.env.PI_CODEX_WEB_SEARCH_MODEL; + const query = "how does the bun runtime schedule timers?"; + await searchCodex(makeSearchParams(query, mockCodexFetch("gpt-5.6-luna"))); + + expect(sentUserText()).toBe(query); + }); + it("uses configured Codex endpoint, API key, and headers without OAuth", async () => { process.env.PI_CODEX_WEB_SEARCH_MODEL = "gpt-5.4"; const result = await searchCodex({ diff --git a/packages/coding-agent/test/tools/web-search-exa.test.ts b/packages/coding-agent/test/tools/web-search-exa.test.ts index 6d491028d..9e424f15c 100644 --- a/packages/coding-agent/test/tools/web-search-exa.test.ts +++ b/packages/coding-agent/test/tools/web-search-exa.test.ts @@ -297,6 +297,38 @@ describe("searchExa", () => { contents: { summary: { query: "shape test" } }, }); }); + it("maps site:/before: directives to native Exa params with an operator-free query", async () => { + await withLocalAuthStorage(authStorage => + new ExaProvider().search({ + query: "vector db benchmarks site:qdrant.tech before:2025-01-01", + systemPrompt: "", + authStorage, + fetch: mockFetch(makeMockExaResponse()), + }), + ); + expect(capturedRequestBody!.query).toBe("vector db benchmarks"); + expect(capturedRequestBody!.includeDomains).toEqual(["qdrant.tech"]); + expect(capturedRequestBody!.endPublishedDate).toBe("2025-01-01"); + expect(capturedRequestBody!.startPublishedDate).toBeUndefined(); + expect(capturedRequestBody!.excludeDomains).toBeUndefined(); + }); + + it("sends directive-free queries byte-identical with no domain/date params", async () => { + await withLocalAuthStorage(authStorage => + new ExaProvider().search({ + query: "plain natural language question", + systemPrompt: "", + authStorage, + fetch: mockFetch(makeMockExaResponse()), + }), + ); + expect(capturedRequestBody).toEqual({ + query: "plain natural language question", + numResults: 10, + type: "auto", + contents: { summary: { query: "plain natural language question" } }, + }); + }); it("paces consecutive Exa API requests by the configured delay", async () => { resetSettingsForTest(); diff --git a/packages/coding-agent/test/tools/web-search-firecrawl.test.ts b/packages/coding-agent/test/tools/web-search-firecrawl.test.ts index 6fe9c17e5..3af554448 100644 --- a/packages/coding-agent/test/tools/web-search-firecrawl.test.ts +++ b/packages/coding-agent/test/tools/web-search-firecrawl.test.ts @@ -104,6 +104,54 @@ describe("Firecrawl web search provider", () => { authMode: "api_key", }); }); + it("maps before:/after: to a cdr tbs and strips dates from the operator query", async () => { + const captured: { body?: unknown } = {}; + const fetchMock: FetchImpl = async (_input, init) => { + captured.body = JSON.parse(String(init?.body ?? "null")) as unknown; + return new Response(JSON.stringify({ data: { web: [] } }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }); + }; + + await searchFirecrawl({ + ...makeParams("bun runtime site:github.com/oven-sh intitle:install after:2024-01-01 before:2024-06-30"), + recency: "month", + fetch: fetchMock, + }); + + expect(captured.body).toEqual({ + query: "bun runtime site:github.com/oven-sh intitle:install", + limit: 10, + sources: [{ type: "web" }], + // Explicit absolute bounds take precedence over the qdr:m recency window. + tbs: "cdr:1,cd_min:01/01/2024,cd_max:06/30/2024", + }); + }); + + it("re-emits non-date operators in the query while keeping recency tbs", async () => { + const captured: { body?: unknown } = {}; + const fetchMock: FetchImpl = async (_input, init) => { + captured.body = JSON.parse(String(init?.body ?? "null")) as unknown; + return new Response(JSON.stringify({ data: { web: [] } }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }); + }; + + await searchFirecrawl({ + ...makeParams('"exact phrase" -site:reddit.com filetype:pdf'), + recency: "week", + fetch: fetchMock, + }); + + expect(captured.body).toEqual({ + query: '"exact phrase" -site:reddit.com filetype:pdf', + limit: 10, + sources: [{ type: "web" }], + tbs: "qdr:w", + }); + }); it("uses the initially resolved credential for the first authenticated request", async () => { let resolutionCount = 0; diff --git a/packages/coding-agent/test/tools/web-search-gemini.test.ts b/packages/coding-agent/test/tools/web-search-gemini.test.ts index ec1bc628d..9da285387 100644 --- a/packages/coding-agent/test/tools/web-search-gemini.test.ts +++ b/packages/coding-agent/test/tools/web-search-gemini.test.ts @@ -111,6 +111,34 @@ describe("searchGemini tools serialization", () => { }); }); + it("normalizes query directive aliases to canonical Google forms in the grounding request", async () => { + const fetchMock = mockGeminiFetch(); + await searchGemini({ + ...makeParams("k8s domain:kubernetes.io since:2024"), + fetch: fetchMock, + }); + + expect(capturedRequest).not.toBeNull(); + const request = capturedRequest?.body?.request as Record; + expect(request).toMatchObject({ + contents: [{ role: "user", parts: [{ text: "k8s site:kubernetes.io after:2024-01-01" }] }], + }); + }); + + it("leaves directive-free queries untouched in the developer API request", async () => { + const fetchMock = mockGeminiFetch(DEVELOPER_SSE_RESPONSE); + await searchGemini({ + ...makeParams("plain query with no operators"), + authStorage: apiKeyAuthStorage, + fetch: fetchMock, + }); + + expect(capturedRequest).not.toBeNull(); + expect(capturedRequest?.body).toMatchObject({ + contents: [{ role: "user", parts: [{ text: "plain query with no operators" }] }], + }); + }); + it("uses configured developer API model and reports it when modelVersion is absent", async () => { const fetchMock = mockGeminiFetch(DEVELOPER_SSE_RESPONSE_WITHOUT_MODEL); const response = await searchGemini({ diff --git a/packages/coding-agent/test/tools/web-search-kagi.test.ts b/packages/coding-agent/test/tools/web-search-kagi.test.ts index d6d777a05..a38edc1aa 100644 --- a/packages/coding-agent/test/tools/web-search-kagi.test.ts +++ b/packages/coding-agent/test/tools/web-search-kagi.test.ts @@ -198,6 +198,32 @@ describe("Kagi search result parsing", () => { expect(requestBody?.filters?.after).toBe(expected); }); + it("canonicalizes Google-style directives into the upstream query", async () => { + let requestBody: KagiSearchRequest | undefined; + + const fetchMock: FetchImpl = async (input, init) => { + const urlStr = typeof input === "string" ? input : input instanceof URL ? input.toString() : input.url; + if (urlStr === "https://kagi.com/api/v1/search") { + requestBody = JSON.parse(init?.body as string) as KagiSearchRequest; + return new Response(JSON.stringify({ meta: { trace: "req-directives" }, data: { search: [] } }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }); + } + return new Response("not mocked", { status: 500 }); + }; + + await searchKagi({ + query: "x domain:kagi.com until:2025", + authStorage: fakeAuthStorage, + fetch: fetchMock, + }); + expect(requestBody?.query).toBe("x site:kagi.com before:2025-01-01"); + + await searchKagi({ query: "plain kagi query", authStorage: fakeAuthStorage, fetch: fetchMock }); + expect(requestBody?.query).toBe("plain kagi query"); + }); + it("uses a Bearer authorization header", async () => { let capturedAuth: string | null = null; const fetchMock: FetchImpl = async (input, init) => { diff --git a/packages/coding-agent/test/tools/web-search-kimi.test.ts b/packages/coding-agent/test/tools/web-search-kimi.test.ts index a99c82c3e..5135df75c 100644 --- a/packages/coding-agent/test/tools/web-search-kimi.test.ts +++ b/packages/coding-agent/test/tools/web-search-kimi.test.ts @@ -110,3 +110,50 @@ describe("searchKimi credential resolution", () => { }); }); }); + +describe("searchKimi query directives", () => { + afterEach(() => { + restoreSearchApiKeyEnv(); + vi.restoreAllMocks(); + }); + + function captureBodyFetch(capture: { body?: { text_query?: string } }): FetchImpl { + return (_url, init) => { + capture.body = JSON.parse(String(init?.body)) as { text_query?: string }; + return Promise.resolve( + new Response(JSON.stringify({ search_results: [] }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }), + ); + }; + } + + it("sends directive-free queries upstream unchanged", async () => { + process.env.KIMI_SEARCH_API_KEY = "test-key"; + const capture: { body?: { text_query?: string } } = {}; + await withLocalAuthStorage(async authStorage => { + await searchKimi({ query: "bun test runner docs", authStorage, fetch: captureBodyFetch(capture) }); + }); + expect(capture.body?.text_query).toBe("bun test runner docs"); + }); + + it("rebuilds directive queries with Bing-style operators and drops date directives", async () => { + process.env.KIMI_SEARCH_API_KEY = "test-key"; + const capture: { body?: { text_query?: string } } = {}; + await withLocalAuthStorage(async authStorage => { + await searchKimi({ + query: 'site:github.com intitle:changelog filetype:md after:2024-01-01 "bun runtime" -deprecated', + authStorage, + fetch: captureBodyFetch(capture), + }); + }); + const sent = capture.body?.text_query ?? ""; + expect(sent).toContain("site:github.com"); + expect(sent).toContain("intitle:changelog"); + expect(sent).toContain("filetype:md"); + expect(sent).toContain('"bun runtime"'); + expect(sent).toContain("-deprecated"); + expect(sent).not.toContain("after:"); + }); +}); diff --git a/packages/coding-agent/test/tools/web-search-mojeek.test.ts b/packages/coding-agent/test/tools/web-search-mojeek.test.ts index a13acdfc9..017be31b2 100644 --- a/packages/coding-agent/test/tools/web-search-mojeek.test.ts +++ b/packages/coding-agent/test/tools/web-search-mojeek.test.ts @@ -91,6 +91,19 @@ describe("Mojeek web search provider", () => { expect(url.searchParams.get("since")).toBeNull(); }); + it("re-emits supported operators (phrases, -, site:) and strips unsupported date bounds from the query", async () => { + let capturedUrl = ""; + const fetchMock: FetchImpl = input => { + capturedUrl = typeof input === "string" ? input : input.toString(); + return Promise.resolve(new Response(resultsPage(""), { status: 200 })); + }; + + await searchMojeek(makeParams("independent index site:mojeek.com after:2024", fetchMock)); + + const url = new URL(capturedUrl); + expect(url.searchParams.get("q")).toBe("independent index site:mojeek.com"); + }); + it("parses result rows, deduplicates targets, and skips junk and intra-Mojeek rows", async () => { const html = resultsPage( [ diff --git a/packages/coding-agent/test/tools/web-search-parallel.test.ts b/packages/coding-agent/test/tools/web-search-parallel.test.ts index 92369cfa1..1cda9781f 100644 --- a/packages/coding-agent/test/tools/web-search-parallel.test.ts +++ b/packages/coding-agent/test/tools/web-search-parallel.test.ts @@ -131,6 +131,45 @@ describe("Parallel web search", () => { ]); }); + it("maps site: directives onto source_policy.include_domains and strips them from the query", async () => { + const fetchMock = mockFetch({ + search_id: "search-parallel-3", + results: [], + warnings: null, + usage: null, + }); + + await searchParallel({ query: "web api site:parallel.ai", fetch: fetchMock }, fakeAuthStorage); + expect(capturedRequestBody).toEqual({ + objective: "web api", + search_queries: ["web api"], + mode: "fast", + excerpts: { max_chars_per_result: 10_000 }, + source_policy: { include_domains: ["parallel.ai"] }, + }); + }); + + it("maps -site: and after: onto exclude_domains/after_date, keeping phrases and negation", async () => { + const fetchMock = mockFetch({ + search_id: "search-parallel-4", + results: [], + warnings: null, + usage: null, + }); + + await searchParallel( + { query: '"web api" -legacy -site:reddit.com/r/node after:2025-06-01', fetch: fetchMock }, + fakeAuthStorage, + ); + expect(capturedRequestBody).toEqual({ + objective: '"web api" -legacy', + search_queries: ['"web api" -legacy'], + mode: "fast", + excerpts: { max_chars_per_result: 10_000 }, + source_policy: { exclude_domains: ["reddit.com"], after_date: "2025-06-01" }, + }); + }); + it("surfaces plain-text Parallel API errors", async () => { const fetchMock: FetchImpl = () => Promise.resolve(new Response("upstream unavailable", { status: 503 })); await expect(searchParallel({ query: "broken", fetch: fetchMock }, fakeAuthStorage)).rejects.toMatchObject({ diff --git a/packages/coding-agent/test/tools/web-search-searxng.test.ts b/packages/coding-agent/test/tools/web-search-searxng.test.ts index 8431a6f10..31991306a 100644 --- a/packages/coding-agent/test/tools/web-search-searxng.test.ts +++ b/packages/coding-agent/test/tools/web-search-searxng.test.ts @@ -59,6 +59,68 @@ describe("SearXNG web search provider", () => { }); }); + it("demotes engine-hostile operators while keeping bare-domain site: filters", async () => { + process.env.SEARXNG_ENDPOINT = "https://searx.example.org"; + + const captured: { q?: string | null } = {}; + const fetchMock: FetchImpl = input => { + captured.q = new URL(input.toString()).searchParams.get("q"); + return Promise.resolve( + new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com" }] }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }), + ); + }; + + await searchSearXNG({ + query: "site:github.com/can1357/oh-my-pi inurl:releases site:github.com 17.1.1 release", + fetch: fetchMock, + }); + + expect(captured.q).toBe("17.1.1 release github.com/can1357/oh-my-pi releases site:github.com"); + }); + + it("maps lang: to the language param and re-emits remaining directives in q", async () => { + process.env.SEARXNG_ENDPOINT = "https://searx.example.org"; + + const captured: { url?: URL } = {}; + const fetchMock: FetchImpl = input => { + captured.url = new URL(input.toString()); + return Promise.resolve( + new Response(JSON.stringify({ results: [{ title: "r", url: "https://searxng.org/docs" }] }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }), + ); + }; + + await searchSearXNG({ query: "docs lang:de site:searxng.org", fetch: fetchMock }); + + expect(captured.url?.searchParams.get("q")).toBe("docs site:searxng.org"); + expect(captured.url?.searchParams.get("language")).toBe("de"); + }); + + it("sends directive-free queries verbatim without a language param", async () => { + process.env.SEARXNG_ENDPOINT = "https://searx.example.org"; + + const captured: { url?: URL } = {}; + const fetchMock: FetchImpl = input => { + captured.url = new URL(input.toString()); + return Promise.resolve( + new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com" }] }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }), + ); + }; + + await searchSearXNG({ query: "plain metasearch query", fetch: fetchMock }); + + expect(captured.url?.searchParams.get("q")).toBe("plain metasearch query"); + expect(captured.url?.searchParams.get("language")).toBeNull(); + }); + it("reads Basic auth credentials from nested config.yml settings", async () => { const agentDir = await fs.mkdtemp(path.join(os.tmpdir(), "searxng-settings-")); try { @@ -231,6 +293,107 @@ describe("SearXNG web search provider", () => { expect(captured.headers?.get("Authorization")).toBe("Bearer bearer-token"); }); + it("resolves engine shortcuts via /config into canonical names for the engines parameter", async () => { + const agentDir = await fs.mkdtemp(path.join(os.tmpdir(), "searxng-engines-")); + try { + await Bun.write( + path.join(agentDir, "config.yml"), + [ + "searxng:", + " endpoint: https://searx-shortcuts.example.org", + ' engines: "ddg, Brave, unknown"', + "", + ].join("\n"), + ); + await Settings.init({ agentDir }); + + const requested: URL[] = []; + const fetchMock: FetchImpl = input => { + const url = new URL(input.toString()); + requested.push(url); + if (url.pathname === "/config") { + return Promise.resolve( + new Response( + JSON.stringify({ + engines: [ + { name: "duckduckgo", shortcut: "ddg" }, + { name: "brave", shortcut: "br" }, + ], + }), + { status: 200, headers: { "Content-Type": "application/json" } }, + ), + ); + } + return Promise.resolve( + new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com/r" }] }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }), + ); + }; + + await searchSearXNG({ query: "engine selection", fetch: fetchMock }); + + const searchUrl = requested.find(url => url.pathname === "/search"); + expect(searchUrl?.searchParams.get("engines")).toBe("duckduckgo,brave,unknown"); + } finally { + await removeWithRetries(agentDir); + } + }); + + it("passes configured engines verbatim when /config is unavailable", async () => { + const agentDir = await fs.mkdtemp(path.join(os.tmpdir(), "searxng-engines-fallback-")); + try { + await Bun.write( + path.join(agentDir, "config.yml"), + ["searxng:", " endpoint: https://searx-noconfig.example.org", ' engines: "ddg,brave"', ""].join("\n"), + ); + await Settings.init({ agentDir }); + + const requested: URL[] = []; + const fetchMock: FetchImpl = input => { + const url = new URL(input.toString()); + requested.push(url); + if (url.pathname === "/config") { + return Promise.resolve(new Response("forbidden", { status: 403 })); + } + return Promise.resolve( + new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com/r" }] }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }), + ); + }; + + await searchSearXNG({ query: "fallback engines", fetch: fetchMock }); + + const searchUrl = requested.find(url => url.pathname === "/search"); + expect(searchUrl?.searchParams.get("engines")).toBe("ddg,brave"); + } finally { + await removeWithRetries(agentDir); + } + }); + + it("strips external bang tokens but keeps engine bangs in the query", async () => { + process.env.SEARXNG_ENDPOINT = "https://searx-bangs.example.org"; + + const captured: { url?: URL } = {}; + const fetchMock: FetchImpl = input => { + captured.url = new URL(input.toString()); + return Promise.resolve( + new Response(JSON.stringify({ results: [{ title: "r", url: "https://example.com/r" }] }), { + status: 200, + headers: { "Content-Type": "application/json" }, + }), + ); + }; + + await searchSearXNG({ query: "!!g rust !ddg lifetimes", fetch: fetchMock }); + + expect(captured.url?.pathname).toBe("/search"); + expect(captured.url?.searchParams.get("q")).toBe("rust !ddg lifetimes"); + }); + it("treats empty SearXNG results with upstream failures as a provider error", async () => { process.env.SEARXNG_ENDPOINT = "https://searx.example.org"; diff --git a/packages/coding-agent/test/tools/web-search-tinyfish.test.ts b/packages/coding-agent/test/tools/web-search-tinyfish.test.ts index e09ee6a8b..efa8e2258 100644 --- a/packages/coding-agent/test/tools/web-search-tinyfish.test.ts +++ b/packages/coding-agent/test/tools/web-search-tinyfish.test.ts @@ -65,6 +65,48 @@ function expectTinyFishParams(url: URL, expectedParams: readonly string[]): void } describe("TinyFish web search provider", () => { + it("rewrites directive queries with supported operators only", async () => { + const captured: URL[] = []; + const fetchMock: FetchImpl = async input => { + const url = input instanceof URL ? input : new URL(typeof input === "string" ? input : input.url); + captured.push(url); + return new Response(JSON.stringify(tinyFishPage(tinyFishResults("tinyfish", 3))), { + status: 200, + headers: { "Content-Type": "application/json" }, + }); + }; + + await searchTinyFish({ + ...makeParams( + '"error handling" rust site:github.com -site:gitlab.com filetype:pdf intitle:tokio after:2024-01-01', + ), + fetch: fetchMock, + }); + + expect(captured).toHaveLength(1); + expect(captured[0].searchParams.get("query")).toBe( + '"error handling" rust site:github.com -site:gitlab.com filetype:pdf', + ); + expectTinyFishParams(captured[0], ["query", "num_results", "page"]); + }); + + it("sends directive-free queries verbatim", async () => { + const captured: URL[] = []; + const fetchMock: FetchImpl = async input => { + const url = input instanceof URL ? input : new URL(typeof input === "string" ? input : input.url); + captured.push(url); + return new Response(JSON.stringify(tinyFishPage(tinyFishResults("tinyfish", 3))), { + status: 200, + headers: { "Content-Type": "application/json" }, + }); + }; + + await searchTinyFish({ ...makeParams("plain query with ordinary words"), fetch: fetchMock }); + + expect(captured).toHaveLength(1); + expect(captured[0].searchParams.get("query")).toBe("plain query with ordinary words"); + }); + it("passes TinyFish num_results and applies numSearchResults across pages", async () => { const captured: { url: URL; init?: RequestInit }[] = []; const pages = new Map([ diff --git a/packages/coding-agent/test/tools/web-search-xai.test.ts b/packages/coding-agent/test/tools/web-search-xai.test.ts index 96deb62fd..d62b2e132 100644 --- a/packages/coding-agent/test/tools/web-search-xai.test.ts +++ b/packages/coding-agent/test/tools/web-search-xai.test.ts @@ -163,6 +163,48 @@ describe("xAI web search provider", () => { expect(capture.capturedRequest?.body).not.toHaveProperty("search_parameters"); }); + it("maps site: onto web_search allowed_domains and strips it from the query", async () => { + const capture = captureFetch({ id: "resp_directives", model: "grok-4.3", output_text: "directive answer" }); + + await searchXAI({ + ...makeParams(capture.fetchMock), + query: "grok api site:docs.x.ai after:2025-01-01", + }); + + const body = capture.capturedRequest?.body; + expect(body?.tools).toEqual([{ type: "web_search", filters: { allowed_domains: ["docs.x.ai"] } }]); + // The Responses web_search tool has no from_date/to_date, so the date + // bound stays in the query text for the agent while site: is stripped. + const input = body?.input as { role: string; content: string }[]; + expect(input[1]?.content).toBe("grok api after:2025-01-01"); + }); + + it("maps -site: onto excluded_domains as bare hosts only when no allow list is present", async () => { + const capture = captureFetch({ id: "resp_excludes", model: "grok-4.3", output_text: "exclude answer" }); + + await searchXAI({ + ...makeParams(capture.fetchMock), + query: "grok changelog -site:reddit.com/r/grok -site:news.ycombinator.com", + }); + + const body = capture.capturedRequest?.body; + expect(body?.tools).toEqual([ + { type: "web_search", filters: { excluded_domains: ["reddit.com", "news.ycombinator.com"] } }, + ]); + const input = body?.input as { role: string; content: string }[]; + expect(input[1]?.content).toBe("grok changelog"); + + await searchXAI({ + ...makeParams(capture.fetchMock), + query: "grok changelog site:docs.x.ai -site:reddit.com", + }); + // allowed_domains and excluded_domains are mutually exclusive per + // request: the allow list wins, exclusions fall to the central filter. + expect(capture.capturedRequest?.body?.tools).toEqual([ + { type: "web_search", filters: { allowed_domains: ["docs.x.ai"] } }, + ]); + }); + it("uses dedicated xAI OAuth credentials for Responses API bearer auth", async () => { const capture = captureFetch({ id: "resp_xai_oauth", model: "grok-4.3", output_text: "xAI OAuth answer" }); diff --git a/packages/coding-agent/test/web/search/anthropic.test.ts b/packages/coding-agent/test/web/search/anthropic.test.ts index 9bb4aaa5f..eeaa96476 100644 --- a/packages/coding-agent/test/web/search/anthropic.test.ts +++ b/packages/coding-agent/test/web/search/anthropic.test.ts @@ -75,4 +75,56 @@ describe("Anthropic search request body", () => { expect(userId.account_uuid).toBe(accountUuid); expect(userId.device_id).toMatch(/^[0-9a-f]{64}$/); }); + + it("maps site: to allowed_domains and strips the directive from the query", async () => { + using tempDir = TempDir.createSync("@pi-anthropic-search-sites-"); + const authStorage = await CodingAuthStorage.create(path.join(tempDir.path(), "auth.db")); + try { + authStorage.setRuntimeApiKey("anthropic", "test-key"); + + const cap = makeCaptureFetch(); + await searchAnthropic({ + query: "sdk docs site:docs.anthropic.com", + systemPrompt: "Use web search.", + sessionId: "session-2295", + authStorage, + fetch: cap.fetch, + }); + + const body = cap.body(); + const tool = (body?.tools as Record[] | undefined)?.[0]; + expect(tool?.allowed_domains).toEqual(["docs.anthropic.com"]); + expect(tool).not.toHaveProperty("blocked_domains"); + const messages = body?.messages as { content: string }[]; + expect(messages[0]?.content).toBe("sdk docs"); + } finally { + authStorage.close(); + } + }); + + it("maps -site: to blocked_domains when there are no site includes", async () => { + using tempDir = TempDir.createSync("@pi-anthropic-search-blocked-"); + const authStorage = await CodingAuthStorage.create(path.join(tempDir.path(), "auth.db")); + try { + authStorage.setRuntimeApiKey("anthropic", "test-key"); + + const cap = makeCaptureFetch(); + await searchAnthropic({ + query: "rust async runtime -site:reddit.com", + systemPrompt: "Use web search.", + sessionId: "session-2295", + authStorage, + fetch: cap.fetch, + }); + + const body = cap.body(); + const tool = (body?.tools as Record[] | undefined)?.[0]; + expect(tool?.blocked_domains).toEqual(["reddit.com"]); + expect(tool).not.toHaveProperty("allowed_domains"); + const messages = body?.messages as { content: string }[]; + expect(messages[0]?.content).toBe("rust async runtime"); + } finally { + authStorage.close(); + } + }); }); diff --git a/packages/coding-agent/test/web/search/perplexity.test.ts b/packages/coding-agent/test/web/search/perplexity.test.ts index 142939169..437817bf7 100644 --- a/packages/coding-agent/test/web/search/perplexity.test.ts +++ b/packages/coding-agent/test/web/search/perplexity.test.ts @@ -118,6 +118,22 @@ describe("Perplexity API-key request shape", () => { expect(body?.num_search_results).toBe(5); }); + it("maps site:/-site:/after: directives onto native filters and strips them from the query", async () => { + let body: Record | undefined; + const fetchMock = mockApi(b => (body = b), baseResponse()); + + await searchPerplexity({ + query: "rust site:docs.rs -site:reddit.com after:2024-06-01", + authStorage: apiKeyAuthStorage, + fetch: fetchMock, + }); + + expect(body?.search_domain_filter).toEqual(["docs.rs", "-reddit.com"]); + expect(body?.search_after_date_filter).toBe("6/1/2024"); + const messages = body?.messages as { role: string; content: string }[]; + expect(messages.at(-1)?.content).toBe("rust"); + }); + it("parses related_questions into relatedQuestions, preserving order and dropping blanks", async () => { const fetchMock = mockApi( () => {}, @@ -380,6 +396,26 @@ describe("Perplexity OAuth request shape", () => { expect(response.authMode).toBe("oauth"); expect(response.answer).toBe("OAuth answer"); }); + + it("maps directives onto ask-endpoint native filters and rewrites query_str", async () => { + let body: Record | undefined; + const fetchMock = mockOAuth(b => (body = b)); + + await searchPerplexity({ + query: "rust site:docs.rs -site:reddit.com after:2024-06-01", + search_recency_filter: "month", + authStorage: oauthAuthStorage, + fetch: fetchMock, + }); + + expect(body!.query_str).toBe("rust"); + const params = body!.params as Record; + expect(params.query_str).toBe("rust"); + expect(params.search_domain_filter).toEqual(["docs.rs", "-reddit.com"]); + expect(params.search_after_date_filter).toBe("6/1/2024"); + // Absolute date bounds take precedence over recency. + expect(params.search_recency_filter).toBeNull(); + }); }); describe("Perplexity OAuth transport failure (issue #5315)", () => { diff --git a/packages/coding-agent/test/web/search/query-pipeline.test.ts b/packages/coding-agent/test/web/search/query-pipeline.test.ts new file mode 100644 index 000000000..6de2b9e39 --- /dev/null +++ b/packages/coding-agent/test/web/search/query-pipeline.test.ts @@ -0,0 +1,70 @@ +/** + * Central directive pipeline: executeSearch parses the query once, hands the + * StructuredQuery to the provider, then lenient-filters the returned sources + * — enforcing constraints the provider ignored and relaxing (with a note) + * any dimension that would eliminate every result. + */ +import { afterEach, describe, expect, it, vi } from "bun:test"; +import type { AuthStorage } from "@oh-my-pi/pi-ai"; +import { runSearchQuery } from "@oh-my-pi/pi-coding-agent/web/search"; +import type { SearchParams } from "@oh-my-pi/pi-coding-agent/web/search/provider"; +import * as provider from "@oh-my-pi/pi-coding-agent/web/search/provider"; +import type { SearchProviderId, SearchResponse, SearchSource } from "@oh-my-pi/pi-coding-agent/web/search/types"; + +const SOURCES: SearchSource[] = [ + { title: "Docs page", url: "https://docs.example.com/guide" }, + { title: "Blog post", url: "https://blog.other.com/post" }, +]; + +function stubProvider(id: SearchProviderId, behaviour: (params: SearchParams) => Promise) { + const stub: provider.SearchProvider = { + id, + label: id, + isAvailable: () => true, + isExplicitlyAvailable: () => true, + search: behaviour, + }; + vi.spyOn(provider, "resolveProviderCandidates").mockReturnValue([{ id, explicit: true }]); + vi.spyOn(provider, "getSearchProvider").mockImplementation(async requested => { + if (requested !== id) throw new Error(`Unexpected provider: ${requested}`); + return stub; + }); +} + +describe("web search directive pipeline", () => { + afterEach(() => vi.restoreAllMocks()); + + it("passes the parsed query to the provider and post-filters sources it did not constrain", async () => { + let seen: SearchParams | undefined; + stubProvider("brave", async params => { + seen = params; + return { provider: "brave", sources: SOURCES }; + }); + + const result = await runSearchQuery( + { query: "guide site:docs.example.com", provider: "brave" }, + { authStorage: {} as AuthStorage }, + ); + + expect(seen?.parsedQuery?.sites).toEqual(["docs.example.com"]); + expect(seen?.parsedQuery?.text).toBe("guide"); + expect(result.details.response.sources.map(s => s.url)).toEqual(["https://docs.example.com/guide"]); + expect(result.content[0]?.text).not.toContain("Note:"); + }); + + it("relaxes a constraint that matches nothing and leads the LLM text with a note", async () => { + stubProvider("brave", async () => ({ provider: "brave", sources: SOURCES })); + + const result = await runSearchQuery( + { query: "guide site:nowhere.example", provider: "brave" }, + { authStorage: {} as AuthStorage }, + ); + + // Leniency: nothing matched site:nowhere.example, so all sources survive + // and the model is told the constraint was relaxed. + expect(result.details.response.sources).toHaveLength(SOURCES.length); + expect(result.content[0]?.text).toStartWith( + "Note: no results matched `site:nowhere.example`; the constraint was relaxed", + ); + }); +}); diff --git a/packages/coding-agent/test/web/search/query.test.ts b/packages/coding-agent/test/web/search/query.test.ts new file mode 100644 index 000000000..4ea0c0346 --- /dev/null +++ b/packages/coding-agent/test/web/search/query.test.ts @@ -0,0 +1,298 @@ +import { describe, expect, it } from "bun:test"; +import { + applyQueryConstraints, + formatQuery, + formatScraperQuery, + GOOGLE_QUERY_SYNTAX, + matchesQueryConstraints, + matchesSite, + parseDateValue, + parseSearchQuery, +} from "@oh-my-pi/pi-coding-agent/web/search/query"; +import type { SearchSource } from "@oh-my-pi/pi-coding-agent/web/search/types"; + +describe("parseSearchQuery", () => { + it("leaves plain queries untouched", () => { + const q = parseSearchQuery("rust async runtime comparison"); + expect(q.text).toBe("rust async runtime comparison"); + expect(q.hasDirectives).toBe(false); + expect(q.hasConstraints).toBe(false); + expect(q.terms.map(t => t.text)).toEqual(["rust", "async", "runtime", "comparison"]); + }); + + it("keeps unknown colon tokens (URLs, paths, error codes) as text", () => { + const q = parseSearchQuery("error TS2345: https://example.com/a?b=c C:\\Users\\me"); + expect(q.hasConstraints).toBe(false); + expect(q.terms.map(t => t.text)).toEqual(["error", "TS2345:", "https://example.com/a?b=c", "C:\\Users\\me"]); + }); + + it("parses site: aliases and exclusions with normalization", () => { + const q = parseSearchQuery( + "kubernetes site:HTTPS://Docs.K8s.IO/ domain:cncf.io -site:*.reddit.com host:github.com/kubernetes", + ); + expect(q.sites).toEqual(["docs.k8s.io", "cncf.io", "github.com/kubernetes"]); + expect(q.excludedSites).toEqual(["reddit.com"]); + expect(q.text).toBe("kubernetes"); + expect(q.hasConstraints).toBe(true); + }); + + it("adopts the next token when a directive value is space-separated", () => { + const q = parseSearchQuery("site: arxiv.org transformer scaling"); + expect(q.sites).toEqual(["arxiv.org"]); + expect(q.text).toBe("transformer scaling"); + }); + + it("does not adopt operators or directives as space-separated values", () => { + const q = parseSearchQuery("site: OR site:example.com"); + expect(q.sites).toEqual(["example.com"]); + }); + + it("parses inurl/intitle/intext variants including quoted values", () => { + const q = parseSearchQuery('inurl:docs intitle:"getting started" -intitle:deprecated inbody:websocket handshake'); + expect(q.inUrl).toEqual(["docs"]); + expect(q.inTitle).toEqual(["getting started"]); + expect(q.excludedInTitle).toEqual(["deprecated"]); + expect(q.inText).toEqual(["websocket"]); + expect(q.text).toBe("handshake"); + }); + + it("routes every following term for allintitle:", () => { + const q = parseSearchQuery("allintitle: budget planning tips"); + expect(q.inTitle).toEqual(["budget", "planning", "tips"]); + expect(q.text).toBe(""); + }); + + it("parses filetype/ext with normalization", () => { + const q = parseSearchQuery("quarterly report filetype:PDF ext:.xlsx -filetype:doc"); + expect(q.filetypes).toEqual(["pdf", "xlsx"]); + expect(q.excludedFiletypes).toEqual(["doc"]); + }); + + it("parses date bounds in common forms", () => { + const q = parseSearchQuery("llm evals after:2024 before:2025-06-15"); + expect(q.after).toBe("2024-01-01"); + expect(q.before).toBe("2025-06-15"); + expect(q.text).toBe("llm evals"); + + expect(parseSearchQuery("x since:2023/07/01").after).toBe("2023-07-01"); + expect(parseSearchQuery("x until:2024-02").before).toBe("2024-02-01"); + expect(parseSearchQuery("x after:6/15/2024").after).toBe("2024-06-15"); + expect(parseSearchQuery("x after:15/6/2024").after).toBe("2024-06-15"); + }); + + it("degrades unparseable date values to plain terms", () => { + const q = parseSearchQuery("release before:soon"); + expect(q.before).toBeUndefined(); + expect(q.terms.map(t => t.text)).toEqual(["release", "before:soon"]); + }); + + it("parses quoted phrases including smart quotes and negated phrases", () => { + const q = parseSearchQuery('"exact phrase" \u201csmart quoted\u201d -"not this"'); + expect(q.terms).toEqual([ + { text: "exact phrase", phrase: true }, + { text: "smart quoted", phrase: true }, + { text: "not this", phrase: true, negated: true }, + ]); + expect(q.text).toBe('"exact phrase" "smart quoted" -"not this"'); + }); + + it("groups OR alternatives and treats AND as default conjunction", () => { + const q = parseSearchQuery("(react OR vue OR svelte) AND hooks"); + const groups = q.terms.map(t => t.group); + expect(q.terms.map(t => t.text)).toEqual(["react", "vue", "svelte", "hooks"]); + expect(groups[0]).toBeDefined(); + expect(groups[0]).toBe(groups[1]); + expect(groups[1]).toBe(groups[2]); + expect(groups[3]).toBeUndefined(); + expect(q.text).toBe("(react OR vue OR svelte) hooks"); + }); + + it("supports pipe as OR and NOT as negation", () => { + const q = parseSearchQuery("deno | bun NOT node"); + expect(q.terms[0].group).toBe(q.terms[1].group); + expect(q.terms[2]).toMatchObject({ text: "node", negated: true }); + }); + + it("swallows OR between directives instead of grouping terms", () => { + const q = parseSearchQuery("caching site:redis.io OR site:memcached.org"); + expect(q.sites).toEqual(["redis.io", "memcached.org"]); + expect(q.terms).toEqual([{ text: "caching" }]); + }); + + it("keeps wikipedia-style parens inside directive values", () => { + const q = parseSearchQuery("site:en.wikipedia.org/wiki/Rust_(programming_language) borrow checker"); + expect(q.sites).toEqual(["en.wikipedia.org/wiki/rust_(programming_language)"]); + expect(q.text).toBe("borrow checker"); + }); + + it("treats legacy +term as an exact phrase", () => { + const q = parseSearchQuery("+immutable data"); + expect(q.terms[0]).toEqual({ text: "immutable", phrase: true }); + }); + + it("parses lang: into a language code", () => { + const q = parseSearchQuery("documentation lang:EN-us"); + expect(q.lang).toBe("en-us"); + expect(q.hasConstraints).toBe(false); + }); +}); + +describe("parseDateValue", () => { + it("rejects invalid components", () => { + expect(parseDateValue("2024-13-01")).toBeUndefined(); + expect(parseDateValue("2024-00-10")).toBeUndefined(); + expect(parseDateValue("notadate")).toBeUndefined(); + expect(parseDateValue("24-01-01")).toBeUndefined(); + }); +}); + +describe("formatQuery", () => { + it("re-emits full Google syntax", () => { + const q = parseSearchQuery( + 'release notes site:github.com -site:gist.github.com filetype:md after:2024-05-01 intitle:"v2"', + ); + expect(formatQuery(q, GOOGLE_QUERY_SYNTAX)).toBe( + "release notes site:github.com -site:gist.github.com intitle:v2 filetype:md after:2024-05-01", + ); + }); + + it("emits OR-grouped sites for multi-site queries", () => { + const q = parseSearchQuery("cve site:nvd.nist.gov site:mitre.org"); + expect(formatQuery(q, GOOGLE_QUERY_SYNTAX)).toBe("cve (site:nvd.nist.gov OR site:mitre.org)"); + }); + + it("produces plain keywords when the engine supports no syntax", () => { + const q = parseSearchQuery('(react OR vue) "state management" -redux site:dev.to filetype:pdf'); + expect(formatQuery(q, {})).toBe("react vue state management"); + }); + + it("falls back to constraint values for directive-only queries", () => { + const q = parseSearchQuery("site:kubernetes.io filetype:yaml"); + expect(formatQuery(q, {})).toBe("kubernetes.io yaml"); + }); +}); + +describe("formatScraperQuery", () => { + it("demotes path-carrying site: and inurl: to plain terms, keeping bare-domain site:", () => { + expect(formatScraperQuery("site:github.com/can1357/oh-my-pi inurl:releases site:github.com 17.1.1 release")).toBe( + "17.1.1 release github.com/can1357/oh-my-pi releases site:github.com", + ); + }); + + it("demotes every site in an OR-groupable multi-site query when all carry paths", () => { + expect(formatScraperQuery("cve site:nvd.nist.gov/vuln site:mitre.org/cgi-bin")).toBe( + "cve nvd.nist.gov/vuln mitre.org/cgi-bin", + ); + }); + + it("passes negated site:/inurl: through as operators", () => { + expect(formatScraperQuery("foo -site:github.com/x -inurl:bar")).toBe("foo -site:github.com/x -inurl:bar"); + }); + + it("passes directive-free queries through byte-identical", () => { + expect(formatScraperQuery("github.com/can1357/oh-my-pi 17.1.1 release")).toBe( + "github.com/can1357/oh-my-pi 17.1.1 release", + ); + }); + + it("respects a narrower engine syntax while still demoting hostile operators", () => { + expect( + formatScraperQuery("a site:x.com site:y.com/z inurl:w", undefined, { + phrases: true, + negation: true, + site: true, + }), + ).toBe("a y.com/z w site:x.com"); + }); +}); + +describe("matchesSite", () => { + it("matches exact hosts, subdomains, and path prefixes", () => { + expect(matchesSite("https://docs.k8s.io/setup", "k8s.io")).toBe(true); + expect(matchesSite("https://k8s.io/", "k8s.io")).toBe(true); + expect(matchesSite("https://notk8s.io/", "k8s.io")).toBe(false); + expect(matchesSite("https://github.com/anthropics/sdk", "github.com/anthropics")).toBe(true); + expect(matchesSite("https://github.com/other/sdk", "github.com/anthropics")).toBe(false); + expect(matchesSite("not a url", "k8s.io")).toBe(false); + }); +}); + +describe("applyQueryConstraints", () => { + const sources: SearchSource[] = [ + { title: "K8s docs — Install", url: "https://kubernetes.io/docs/setup/install.pdf" }, + { title: "Random blog", url: "https://blog.example.com/k8s", publishedDate: "2023-01-15" }, + { title: "Reddit thread", url: "https://www.reddit.com/r/kubernetes/post" }, + { title: "K8s blog", url: "https://kubernetes.io/blog/2024", ageSeconds: 3600 }, + ]; + + it("filters by included site", () => { + const q = parseSearchQuery("install site:kubernetes.io"); + const { sources: out, dropped } = applyQueryConstraints(sources, q); + expect(out.map(s => s.url)).toEqual([ + "https://kubernetes.io/docs/setup/install.pdf", + "https://kubernetes.io/blog/2024", + ]); + expect(dropped).toEqual([]); + }); + + it("drops a constraint that would eliminate every result and reports it", () => { + const q = parseSearchQuery("install site:nonexistent.example"); + const { sources: out, dropped } = applyQueryConstraints(sources, q); + expect(out).toHaveLength(sources.length); + expect(dropped).toEqual(["site:nonexistent.example"]); + }); + + it("relaxes dimensions independently", () => { + // site matches two sources, filetype:docx matches none of those → only filetype relaxed. + const q = parseSearchQuery("install site:kubernetes.io filetype:docx"); + const { sources: out, dropped } = applyQueryConstraints(sources, q); + expect(out.map(s => s.url)).toEqual([ + "https://kubernetes.io/docs/setup/install.pdf", + "https://kubernetes.io/blog/2024", + ]); + expect(dropped).toEqual(["filetype:docx"]); + }); + + it("applies exclusions and filetype filters", () => { + const q = parseSearchQuery("k8s -site:reddit.com filetype:pdf"); + const { sources: out, dropped } = applyQueryConstraints(sources, q); + expect(out.map(s => s.url)).toEqual(["https://kubernetes.io/docs/setup/install.pdf"]); + expect(dropped).toEqual([]); + }); + + it("filters by date bounds while letting undated sources pass", () => { + const q = parseSearchQuery("k8s after:2024-01-01"); + const { sources: out } = applyQueryConstraints(sources, q); + // 2023 blog post is provably too old; undated sources survive. + expect(out.map(s => s.url)).toEqual([ + "https://kubernetes.io/docs/setup/install.pdf", + "https://www.reddit.com/r/kubernetes/post", + "https://kubernetes.io/blog/2024", + ]); + }); + + it("parses relative published dates", () => { + const q = parseSearchQuery("news before:2020"); + const relative: SearchSource[] = [ + { title: "old", url: "https://a.example/1", publishedDate: "2 days ago" }, + { title: "undated", url: "https://a.example/2" }, + ]; + const { sources: out, dropped } = applyQueryConstraints(relative, q); + // "2 days ago" is provably after 2020 → violates before:2020 → only undated survives. + expect(out.map(s => s.url)).toEqual(["https://a.example/2"]); + expect(dropped).toEqual([]); + }); + + it("returns empty input untouched", () => { + const q = parseSearchQuery("x site:a.com"); + expect(applyQueryConstraints([], q)).toEqual({ sources: [], dropped: [] }); + }); +}); + +describe("matchesQueryConstraints", () => { + it("checks all dimensions strictly", () => { + const q = parseSearchQuery("site:kubernetes.io inurl:docs"); + expect(matchesQueryConstraints({ title: "t", url: "https://kubernetes.io/docs/setup" }, q)).toBe(true); + expect(matchesQueryConstraints({ title: "t", url: "https://kubernetes.io/blog/x" }, q)).toBe(false); + }); +}); diff --git a/packages/coding-agent/test/web/search/tavily.test.ts b/packages/coding-agent/test/web/search/tavily.test.ts index 2e5eeccfe..c08d5f8c0 100644 --- a/packages/coding-agent/test/web/search/tavily.test.ts +++ b/packages/coding-agent/test/web/search/tavily.test.ts @@ -45,6 +45,13 @@ describe("Tavily buildRequestBody", () => { expect(body.include_answer).toBe("advanced"); expect(body.include_raw_content).toBe(false); }); + + it("prefers explicit start_date/end_date over time_range", () => { + const body = buildRequestBody({ query: "q", recency: "week", start_date: "2026-01-01", end_date: "2026-02-01" }); + expect(body.start_date).toBe("2026-01-01"); + expect(body.end_date).toBe("2026-02-01"); + expect(body).not.toHaveProperty("time_range"); + }); }); describe("Tavily searchTavily request shape (integration)", () => { @@ -139,4 +146,72 @@ describe("Tavily searchTavily request shape (integration)", () => { expect(capturedBody).not.toHaveProperty("topic"); expect(capturedBody).not.toHaveProperty("time_range"); }); + + it("maps site: directives to include/exclude_domains and strips them from the query", async () => { + process.env.TAVILY_API_KEY = "test-key"; + + let capturedBody: Record | undefined; + const fetchMock: FetchImpl = async (input, init) => { + const url = + typeof input === "string" ? input : input instanceof URL ? input.toString() : (input as Request).url; + if (url === "https://api.tavily.com/search") { + capturedBody = JSON.parse(init?.body as string); + return new Response( + JSON.stringify({ + answer: "test answer", + results: [{ title: "Pricing", url: "https://tavily.com/pricing", content: "plans" }], + request_id: "req-1", + }), + { status: 200, headers: { "Content-Type": "application/json" } }, + ); + } + return new Response("not mocked", { status: 500 }); + }; + + await searchTavily({ + ...makeParams("pricing site:tavily.com -site:reddit.com"), + fetch: fetchMock, + }); + + expect(capturedBody).toBeDefined(); + expect(capturedBody?.include_domains).toEqual(["tavily.com"]); + expect(capturedBody?.exclude_domains).toEqual(["reddit.com"]); + expect(capturedBody?.query).toBe("pricing"); + }); + + it("maps after:/before: to start_date/end_date and retries without them on empty results", async () => { + process.env.TAVILY_API_KEY = "test-key"; + + const capturedBodies: Record[] = []; + const fetchMock: FetchImpl = async (input, init) => { + const url = + typeof input === "string" ? input : input instanceof URL ? input.toString() : (input as Request).url; + if (url === "https://api.tavily.com/search") { + capturedBodies.push(JSON.parse(init?.body as string)); + const empty = capturedBodies.length === 1; + return new Response( + JSON.stringify({ + answer: "", + results: empty ? [] : [{ title: "Post", url: "https://example.com/post", content: "text" }], + request_id: `req-${capturedBodies.length}`, + }), + { status: 200, headers: { "Content-Type": "application/json" } }, + ); + } + return new Response("not mocked", { status: 500 }); + }; + + const response = await searchTavily({ + ...makeParams('"llm agents" after:2026-01-01 before:2026-06-01'), + fetch: fetchMock, + }); + + expect(capturedBodies).toHaveLength(2); + expect(capturedBodies[0]?.start_date).toBe("2026-01-01"); + expect(capturedBodies[0]?.end_date).toBe("2026-06-01"); + expect(capturedBodies[0]?.query).toBe('"llm agents"'); + expect(capturedBodies[1]).not.toHaveProperty("start_date"); + expect(capturedBodies[1]).not.toHaveProperty("end_date"); + expect(response.sources).toHaveLength(1); + }); }); diff --git a/packages/coding-agent/test/web/search/zai.test.ts b/packages/coding-agent/test/web/search/zai.test.ts index a37a9ceda..7c98466a0 100644 --- a/packages/coding-agent/test/web/search/zai.test.ts +++ b/packages/coding-agent/test/web/search/zai.test.ts @@ -1,6 +1,6 @@ import { describe, expect, it } from "bun:test"; import type { AuthStorage, FetchImpl } from "@oh-my-pi/pi-ai"; -import { searchZai } from "@oh-my-pi/pi-coding-agent/web/search/providers/zai"; +import { searchZai, ZaiProvider } from "@oh-my-pi/pi-coding-agent/web/search/providers/zai"; interface CapturedRequest { method: string | undefined; @@ -108,4 +108,92 @@ describe("Z.AI web search provider", () => { }, ]); }); + + function createMcpFetch(): { fetchImpl: FetchImpl; capturedRequests: CapturedRequest[] } { + const capturedRequests: CapturedRequest[] = []; + const fetchImpl: FetchImpl = (_input, init) => { + const request = { + method: init?.method, + headers: new Headers(init?.headers), + body: JSON.parse(String(init?.body)) as Record, + }; + capturedRequests.push(request); + if (request.body.method === "initialize") { + return Promise.resolve( + new Response( + JSON.stringify({ + jsonrpc: "2.0", + id: request.body.id, + result: { + protocolVersion: "2025-03-26", + capabilities: { tools: {} }, + serverInfo: { name: "zai-web-search", version: "test" }, + }, + }), + { + status: 200, + headers: { "Content-Type": "application/json", "Mcp-Session-Id": "zai-session-1" }, + }, + ), + ); + } + if (request.body.method === "notifications/initialized") { + return Promise.resolve(new Response(null, { status: 202 })); + } + return Promise.resolve( + new Response( + JSON.stringify({ + jsonrpc: "2.0", + id: request.body.id, + result: { + content: [{ type: "text", text: JSON.stringify({ search_result: [] }) }], + }, + }), + { status: 200, headers: { "Content-Type": "application/json" } }, + ), + ); + }; + return { fetchImpl, capturedRequests }; + } + + const authStorage = { + resolver() { + return async () => "zai-test-key"; + }, + hasAuth(provider: string) { + return provider === "zai"; + }, + } as unknown as AuthStorage; + + function toolCallQuery(capturedRequests: CapturedRequest[]): unknown { + const toolCall = capturedRequests.find(request => request.body.method === "tools/call"); + const params = toolCall?.body.params as { arguments?: { query?: unknown } } | undefined; + return params?.arguments?.query; + } + + it("rewrites directive queries into Bing-flavored operator syntax, dropping date bounds", async () => { + const { fetchImpl, capturedRequests } = createMcpFetch(); + await new ZaiProvider().search({ + query: 'pytest "fixture scope" site:docs.pytest.org -inurl:changelog filetype:html after:2024-01-01', + systemPrompt: "", + authStorage, + fetch: fetchImpl, + }); + + expect(toolCallQuery(capturedRequests)).toBe( + 'pytest "fixture scope" site:docs.pytest.org -inurl:changelog filetype:html', + ); + }); + + it("sends directive-free queries upstream byte-identical", async () => { + const { fetchImpl, capturedRequests } = createMcpFetch(); + await new ZaiProvider().search({ + query: "latest bun release notes", + systemPrompt: "", + authStorage, + fetch: fetchImpl, + }); + + expect(toolCallQuery(capturedRequests)).toBe("latest bun release notes"); + }); });