26 Commits

Author SHA1 Message Date
can1357 12238f55ca feat: implemented native ctok tokenization engine with model scopes
- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
2026-08-19 23:27:29 +02:00
can1357 c25b008f4a Merge PR #8067: fix(compaction): manual /shake keeps a recent tail of tool results (@zhang17-24) 2026-08-13 02:00:49 +02:00
can1357 bbe0c6fafd test(plan-mode): aligned availability-gate assertions with compressed prompt wording 2026-08-13 01:19:22 +02:00
Slava Zavadsky 8b8e1931a4 fix(coding-agent): address review on plan-mode and orchestrate gating
- plan-mode-active: when ask is unavailable, no longer present a prose
  terminal action — plan mode cannot yield on text-only turns. Preferences
  are recorded as Assumptions with a recommended default instead.
- orchestrate-notice: name edit/write in the tool budget and inline-edit
  guidance only when each tool is individually present (no stray
  `edit`/`write` when only one is active).
- Tests updated to the new contracts (incl. write-only/edit-only render
  coverage).
2026-08-10 20:44:12 -04:00
Slava Zavadsky 0bbdfe2eaa fix(coding-agent): keep smoke-test fallback for uncovered UI surfaces
When a session has browser but not computer (or vice versa), the
behavioral/smoke-test fallback for UI changes was suppressed entirely,
leaving native-desktop (or web) verification guidance undefined. Render
the fallback whenever any visual runtime tool is missing.

Also strengthen the plan-mode tests per review: exercise the iterative
branch gating with iterative:true (previously passed trivially) and
restore a typed Overrides record for the render helper.
2026-08-10 20:37:28 -04:00
Slava Zavadsky b6a3862ebc fix(coding-agent): gate prompt tool mentions on session tool availability
System and mode prompts referenced tools that may be absent from the
session catalog, forcing the model to satisfy requirements it cannot
execute (generalizes #8139's browser-verification mismatch):

- system-prompt.md: browser verification now keys on the actual UI
  surface and available tools (browser/computer/TUI/CLI), with an
  explicit behavioral/smoke-test fallback when no runtime tool exists;
  todo workflow guidance and the AST-section grep hint are gated on
  tool presence; the auto-QA report_issue block additionally requires
  the write tool.
- project-prompt.md: workspace-tree drill-in and additional-roots tool
  hints only name tools present in the session.
- plan-mode-active.md: ask-tool directives get a prose fallback and
  scout-via-task dispatch is gated on the task tool (render site now
  passes askAvailable/taskAvailable).
- orchestrate-notice.md: the tool budget, verify gates, todo tracking,
  and inline-edit guidance are gated on tool presence; the notice is
  skipped entirely when the task tool is inactive (render fn takes the
  active tool list).

Adds regression tests for each gate; changelog entry.
2026-08-10 20:24:34 -04:00
left-to-right c3da093a1d fix(compaction): manual /shake keeps a recent tail of tool results
Manual /shake used protectTokens: 0, stripping every eligible tool result
including the ones the agent is still working from. Keep a small 4k-token
recent window (matching the automatic shake mechanism, at a quarter of its
16k budget) so the full escape hatch stays aggressive without destroying
the live tail. Two matcher-focused tests that implicitly relied on the
zero window now pin protectTokens: 0 explicitly.

Fixes #7776
2026-08-09 18:09:17 +08:00
can1357 a872d77068 chore: cleanup dumb tests 2026-08-02 20:39:23 +02:00
can1357 d16a251777 chore: reorg tests 2026-07-27 16:43:53 +02:00
roboomp 9bba7ac007 fix(plan-mode): compared canonical local urls for scan membership
The scan-membership test used a raw string includes, so a resumed single-slash local:/ state path failed to match the scanner's local:// entry and wrongly gained precedence over a newer draft. Normalize both sides via normalizeLocalScheme before comparing.

Fixes #6569
2026-07-25 01:58:12 +00:00
roboomp c7ac15e974 fix(plan-mode): kept out-of-scan state plan ahead of scanned artifacts
A state plan the artifact scan can't surface (cwd-relative, or a local file not ending in plan.md) has no mtime to compete on, so it now keeps precedence over scanned drafts; an in-scan state plan still competes on newest-first order.

Fixes #6569
2026-07-25 01:44:28 +00:00
roboomp 926f523a15 fix(plan-mode): preferred newest draft during review
Preferred the newest session-local plan artifact over a stale state path when the submitted title cannot reconstruct the draft filename.

Added regression coverage for completed-plan re-entry.

Fixes #6569
2026-07-25 01:38:29 +00:00
oldschoola a2854ba768 fix: migrate coding-agent tests from fs.rm to removeWithRetries
Migrate 203 test files (356 call sites) from fs.rm/fs.rmSync to
removeWithRetries/removeSyncWithRetries to reduce EBUSY test failures
on Windows. removeWithRetries is now exported from @oh-my-pi/pi-utils.

The migration uses a regex-based approach that:
- Replaces fs.rm(path, { recursive, force }) → removeWithRetries(path)
- Replaces fs.rmSync(path, { recursive, force }) → removeSyncWithRetries(path)
- Replaces fs.rm(path) → removeWithRetries(path) (no options)
- Skips fs.rm/fs.rmSync inside template literals (bun --eval scripts)
- Adds imports to existing @oh-my-pi/pi-utils import or creates new one
- Removes unused fs imports where fs.rm was the only fs usage (4 files)
2026-06-23 15:28:05 -07:00
can1357 f40e528073 fix(agent/compaction): skipped truncating tiny tool outputs during pruning
- Added a 50-token minimum floor constant for pruning decisions.
- Allowed non-superseded, non-useless tool results below that floor to avoid truncation.
2026-06-15 12:45:21 +02:00
can1357 0f043d4c29 fix(coding-agent): preserved approved plan paths during plan apply resolution
- Replaced approved-plan renaming with `resolveApprovedPlan` resolution and state/slug lookup.
- Updated ACP and interactive apply flows to propagate canonical `planFilePath` instead of renamed paths.
- Added local plan fallback lookup by mtime for unresolved slugs after plan approval.
- Restricted plan-mode writes to `local://` plan artifacts and simplified path handling.
2026-06-07 07:17:23 +02:00
can1357 3f22f939a2 feat(coding-agent/plan-mode): shared approved plan context with spawned subagents
- Added `loadOverallPlanReference` to resolve a session plan reference from local storage and skip empty or missing files.
- Updated task execution to read the active plan reference (except in plan mode) and pass it into each spawned subagent.
- Extended the subagent system prompt and session SDK/tools plumbing so subagents receive and render the approved plan path and contents.
2026-06-04 04:10:53 +02:00
can1357 e1b2809154 feat(coding-agent): added plan read matching and compaction protection behavior
- Added shared `getReadToolPath` API to extract paired read `path` values for protection matchers.
- Added `createPlanReadMatcher` and session wiring so compaction prune/shake keeps active plan reads intact.
- Updated `todo-write` instructions to initialize every user-supplied plan item as an individual task.
- Added compaction tests validating plan reads are protected from prune and shake while regular reads are still removable.
2026-06-03 16:14:15 +02:00
roboomp 14198e2ed3 fix(plan): derive plan title when resolve.extra.title is not a string
Grammar-constrained models (e.g. Qwen3.6-35B-MTP via llama.cpp) emit
`extra: { title: {} }` instead of `extra: { title: "<string>" }` because
the resolve schema declares `extra` as Record<string, unknown> with an
open value schema, leaving the model free to drop in an empty object.
The apply guard then threw 'Plan approval requires extra: { title: ... }'
on every retry, looping the model indefinitely (issue #1179).

Plan approval now uses a layered title resolution:
  1. `extra.title` if it is a non-empty string (and sanitizes to non-empty)
  2. First `# Heading` in the plan content
  3. Filename stem of `planFilePath` (`'/data/workspaces/can1357__oh-my-pi__1179/.omp-session/2026-05-19T03-59-21-254Z_019e3e63-62a6-7000-be63-371f2cd6d67d/local/PLAN.md'` → `PLAN`)
  4. Literal `plan` as a final safety net
Each candidate is run through `normalizePlanTitle`; rejected ones fall
through. Extracted as `resolvePlanTitle` in plan-mode/approved-plan.ts
so it's unit-testable.

Prompt language relaxed from MUST to SHOULD for `extra.title` in
plan-mode-active.md and plan-mode-tool-decision-reminder.md, noting the
fallback so models don't waste turns on a now-optional field.

Fixes #1179
2026-05-19 04:06:04 +00:00
can1357 226fe87345 fix(coding-agent/eval): hardened image display value coercion to strict base64
- Implemented strict base64 validation and normalization for image `displayValue` payloads in the JS runtime, supporting strict strings, `Uint8Array`, `Buffer`, `ArrayBuffer`, typed-array views, and JSON `Buffer` objects.
- Dropped unrecognized image payloads while emitting a warning and fallback text instead of forwarding malformed data.
- Added tests covering successful coercions and invalid image data rejection paths.
2026-05-19 04:36:30 +02:00
roboomp ee442cedea fix(ask,plan): fix renderer crash and plan-mode exit loop for Qwen3 and similar models
- Guard renderInlineMarkdown against non-string input: partial JSON during
  streaming can leave option label fields as undefined, causing marked.lexer
  to throw 'undefined is not an object (evaluating e.replace)'. The ask tool
  renderer now silently falls back to an empty string or baseColor output.

- Sanitize normalizePlanTitle instead of hard-rejecting: models that produce
  natural-language plan titles like 'My Improvement Plan' were getting a
  ToolError on every resolve call, causing an infinite retry loop. Spaces are
  now converted to hyphens, remaining invalid chars are dropped, and only
  truly unresolvable titles (empty after sanitization, path separators) throw.

- Fix ask.md prompt example: the example showed the legacy single-question
  format (question/options/recommended at the top level) while the schema
  requires questions: [{id, question, options}]. Models that follow examples
  closely (Qwen3) generated calls that always failed schema validation.

Fixes #1176
2026-05-19 02:09:47 +00:00
can1357 4f6e70f779 fix(coding-agent): auto-name approved plan sessions 2026-05-16 20:53:11 +02:00
can1357 975941aba4 chore: remove garbage tests 2026-05-12 04:09:33 +02:00
can1357 52719d1a7c refactor: restructured monorepo TypeScript config and build tasks for unified setup
- Migrated all package tsconfig files to extend tsconfig.workspace.json for unified TypeScript configuration across monorepo.
- Consolidated build and check scripts across 10+ packages to use biome for linting/formatting with separate type checking via tsgo.
- Renamed build scripts from build:native and build:binary to build for simplified command naming across packages/natives and packages/coding-agent.
- Refactored CI workflow to invoke bun tasks instead of inline shell scripts, reducing workflow complexity by 40+ lines.
- Removed sync-exports.ts and repro-stuck.ts scripts; deleted path aliases from tsconfig.base.json in favor of workspace-based configuration.
- Updated turbo.json with new task definitions (check:types, lint, fmt, fix) and removed build:native/embed:native tasks.
2026-04-08 17:05:20 +02:00
can1357 a21a542afd refactor(prompt-templates): migrated prompt utilities to pi-utils package
- Extracted prompt rendering and formatting utilities from coding-agent to centralized pi-utils package with new API surface (prompt.render, prompt.format, prompt.registerHelper).
- Migrated parseFrontmatter utility from coding-agent to pi-utils package; updated 8 files to import from @oh-my-pi/pi-utils.
- Removed 170-line prompt-format.ts module and consolidated 192 lines of Handlebars helper registrations into pi-utils prompt module.
- Updated 60+ files across coding-agent and typescript-edit-benchmark to use new prompt.render() and prompt.format() API from pi-utils.
- Simplified prompt-templates.ts by delegating core functionality to pi-utils while retaining custom helper registrations (jtdToTypeScript, jsonStringify, etc.).
2026-04-08 05:47:35 +02:00
can1357 cd5c9655aa refactor: renamed notes protocol to local
- Renamed the `notes://` protocol to `local://` for better clarity.
- Updated all internal references, prompts, and tool documentation.
- Migrated plan storage paths to use the new `local://` scheme.
2026-02-22 18:05:22 +01:00
can1357 563d6a0ab9 feat(coding-agent): introduced notes:// protocol for session-scoped artifact storage
- Replaced plan:// protocol with notes:// for session-scoped artifact storage and plan finalization.
- Added title parameter to exit_plan_mode tool to enable plan file renaming during approval workflow.
- Implemented NotesProtocolHandler for notes:// URL scheme with path traversal protection and session fallback.
- Added renameApprovedPlanFile function to handle plan artifact finalization with validation and error handling.
- Updated system prompt documentation to reference notes:// protocol and internal URL schemes for artifact access.
2026-02-22 17:41:01 +01:00