Commit Graph
14 Commits
Author SHA1 Message Date
can1357 f9f6ed9e8d feat(coding-agent): replaced legacy pi/ role alias prefix with
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
2026-07-13 23:26:33 +02:00
can1357 bdfc21df43 feat(coding-agent/tiny): added llama3.2:3b local tiny model option
- Added the `llama3.2:3b` model configuration pointing to the quantized `onnx-community/Llama-3.2-3B-Instruct-ONNX` repository.
- Registered the model in both the available local models registry and list of valid memory model values.
- Documented the new option as a shipped local model choice in the documentation and changelog.
2026-06-30 17:59:10 +02:00
can1357 f0f7a5ba89 feat(coding-agent): introduced tiny model role for background tasks
- Added `tiny` as a first-class model role to override online models for lightweight background tasks.
- Updated session title generation, auto-thinking difficulty classification, unexpected-stop detection, and mnemopi backend to resolve via the `tiny` role before falling back to `smol`.
- Updated configuration schema and documentation to reflect the new role precedence.
2026-06-27 07:56:27 +02:00
roboomp 9b9a6b38a7 fix(tiny): stopped respawning failed local model workers
Blocked the unsupported Qwen3 1.7B ONNX memory model before loading transformers and remembered local model execution failures so the client returns null instead of spawning another __omp_worker_tiny_inference process for the same failed model.

Fixes #3132
2026-06-20 13:24:53 +00:00
roboomp 648f41c710 fix(coding-agent): expanded local auto-thinking budget
- Gave reasoning-capable local auto-thinking classifiers the safe 1024-token answer budget used by online reasoning classifiers.
- Raised the non-reasoning local classifier floor to 16 tokens and allowed tiny completions to honor larger explicit budgets.
- Added regression coverage for qwen3-1.7b and qwen2.5-1.5b local classifier budgets.

Fixes #2808
2026-06-16 23:36:29 +00:00
can1357 14bd572f49 refactor(shake): removed shake-summary mode and local-model compressor
- Dropped `summarizeShakeRegions`, the shake-summary prompt, and related types.
- Removed `shake-summary` compaction strategy and `providers.shakeSummaryModel` setting.
- Migrated existing `shake-summary` configs to plain `shake` on load.
- Simplified `/shake` to `elide` and `images` modes only.
2026-05-31 14:14:37 +02:00
can1357 68430dee5c chore: renamed mnemosyne package to mnemopi
- Updated package name, directory, and binary from mnemosyne to mnemopi.
- Updated all lockfile references and workspace paths accordingly.
2026-05-31 08:45:12 +02:00
can1357 19be67921d feat(coding-agent): added shake configuration and type modeling
Introduce shake-related configuration and types: strategy options, action enums, and shake result types.
2026-05-31 07:39:51 +02:00
can1357 7f866a48a8 feat(coding-agent): added per-turn AUTO_THINKING in coding-agent session
- Added AUTO_THINKING as a configured thinking level in settings, schema, SDK, and session plumbing.
- Implemented per-turn auto reasoning classification with online/local prompts, effort clamping, and skip guards.
- Updated model selectors, ACP options, footer/status UI, and events to render auto and auto->resolved states.
- Added AUTO_THINKING parse/clamp tests and fixed local-module cycle and hashline preview regressions.
2026-05-31 03:32:21 +02:00
can1357 5d3adfab7f feat(tiny): added GPU-first device selection with CPU fallback for local models
- Added `PI_TINY_DEVICE` env var to control ONNX execution provider (`gpu` default, `cpu`, `metal`, `cuda`, `dml`, `coreml`).
- Local tiny-model inference now tries accelerated GPU provider first and retries on CPU if initialization fails.
- Added `device.ts` module with normalization, preference resolution, and load-order helpers.
- Updated docs and model descriptions to drop CPU-specific language.
2026-05-31 01:57:59 +02:00
can1357 f0b252449b feat(coding-agent): added tiny local model support for memory extraction and consolidation
- Added a new `providers.memoryModel` setting with tiny memory model options and `ONLINE_MEMORY_MODEL_KEY` default in settings.
- Updated Mnemosyne provider resolution so a configured local tiny model overrode remote completion and used new memory extraction and consolidation prompts.
- Expanded the tiny-model CLI registry to download and report all local tiny models (title plus memory) through a unified list.
2026-05-30 18:48:47 +02:00
can1357 70202360ff feat(coding-agent): added local model registry for title completion
- Added Mnemosyne runtime `extractionPrompt` and `consolidationPrompt` options and wired them into resolved LLM config.
- Added fact-extraction branch to call configured completion first (temp 0), then parse facts and safely fall back.
- Added tiny local model registry features for memory/title, including keys, specs, and validation helpers.
- Added `complete` protocol messages and abort-aware client/worker paths for local title completion generation.
- Added local-models documentation for tiny/memory transformer paths, defaults, and known parser caveats.
2026-05-30 18:41:48 +02:00
can1357 a53acf1431 test(coding-agent): updated hashline preview tests to use snapshot-tagged headers
- Updated hashline streaming preview tests to generate snapshot-tagged section headers and use an in-memory snapshot store.
- Replaced untagged file markers in multi-section preview inputs with `formatHashlineHeader` values derived from recorded file contents.
- Changed the bash command error test to expect a returned `isError` result with exit code 1 instead of a rejected promise.
2026-05-30 17:51:11 +02:00
can1357 826c3b932d refactor(packages/coding-agent): reorganized tiny-title runtime stack
- Added tiny-title protocol contracts, including progress-state unions, message payloads, and transport interfaces.
- Added title text utilities to truncate long inputs, wrap `<user-message>` blocks, and normalize generated titles.
- Added tiny-title model registry and helpers with type-safe keys and runtime optional loading via optionalDependencies.
- Added client-side worker orchestration with spawn fallback, request queueing, progress/error routing, and smoke-test APIs.
- Added worker runtime for model resolution, prompt-based inference, lock-based install retries, and close-time cache clear.
2026-05-30 17:40:18 +02:00