Commit Graph

348 Commits

Author SHA1 Message Date
can1357 eced7ab08a feat: implemented live status board utility and agent progress tracking
- Implement a live status board utility for transient multi-line CLI status displays with TTY fallback support.
- Add progress callback support and forward subagent progress events to runner hooks.
- Update cleanse execution flow to track checker runs, agent progress, and status rendering.
- Add comprehensive unit tests for live board repainting and cleanse progress assertions.
2026-08-20 03:22:54 +02:00
can1357 0cdd37fc15 feat(pi-natives/tools): implemented utok tokenizer for multiple models
- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
2026-08-20 01:45:04 +02:00
can1357 12238f55ca feat: implemented native ctok tokenization engine with model scopes
- Implemented the `ctok` Rust native tokenization engine with offline support for Claude V3, V47, V5, and V5Sonnet families.
- Replaced global token estimation with model-scoped `Tokenizer` instances and provider-anchored transcript accounting across packages.
- Added vocabulary generation scripts, test fixtures, and comprehensive unit tests for tokenizer routing and matching modes.
2026-08-19 23:27:29 +02:00
roboomp 4b07f409f6 fix(ai): hoist assistant message interleaved in responses tool batch
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.

buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.

Fixes #8789
2026-08-19 08:46:28 +00:00
can1357 75179e1dc0 fix(compaction): scale summary window floor to the model's context
The absolute 16,384-token floor plus the carried summary and output
reserves exceeds windows below ~58k outright, and overflow recovery then
bailed at the very floor that caused the rejection, leaving compaction
unusable on small-context models. Scale the floor to window/8 (min 1k)
and use the same floor in overflow recovery.
2026-08-19 01:39:17 +02:00
can1357 17e47b3eb7 Merge PR #8920: fix(compaction): bound summarization input and stop retrying overflow (@PaleRoses)
# Conflicts:
#	packages/agent/src/compaction/compaction.ts
2026-08-19 01:39:07 +02:00
can1357 996f562247 Merge PR #8727: fix(agent): harden compaction summaries against prompt injection (@koopmannleon19977-cmyk) 2026-08-19 01:36:04 +02:00
can1357 e88fb70afe Merge PR #8720: fix(compaction): honor /clear reset boundary in prepareCompaction (@roboomp) 2026-08-19 01:36:03 +02:00
PaleRoses 753c86ea72 fix(compaction): bound summarization input and stop retrying overflow
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.

Three independent defects:

1. `generateSummary` serialized the entire span into one prompt with no budget
   check. It now plans windows that fit the summarizer's context and folds them
   with the update prompt that iterative compaction already uses, so a stranded
   boundary is recovered instead of rejected. A provider that rejects a window
   the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
   beta-gated to 200k on OAuth credentials) halves what was actually sent and
   re-plans, because only the rejection knows the real cap.

2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
   the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
   appends to its own errors classified a deterministic 400 as a transient 503.
   Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.

3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
   identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
   an input that does not fit never fits; both layers now fail fast to the next
   candidate.

The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.

Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
2026-08-18 12:25:13 -07:00
can1357 8500092296 chore: bump version to 17.3.7
Retry: widened agent dequeue-hook deadline budgets from 25ms to 1s — the run loop checks the deadline before invoking dequeue hooks, so a cold or CPU-starved mock roundtrip expired the deadline first and the hooks never ran (deterministic failure in isolation, flaky under CI parallel load).
2026-08-18 11:34:10 +03:00
koopmannleon19977-cmyk 39c908bc90 fix(agent): harden compaction summarizer against prompt injection 2026-08-16 16:22:51 +02:00
roboomp 722c4aa0ae fix(compaction): honor /clear reset boundary in prepareCompaction
prepareCompaction walked the branch from the last compaction and ignored reset_boundary markers, so /compact (and auto-compaction) resurrected pre-/clear turns into the summary even though buildSessionContext already starts the model context after the boundary.

Model reset_boundary as a first-class agent-core session entry and start the summarization window after the latest boundary, dropping the superseded pre-reset compaction summary. A boundary before the last compaction stays superseded by it.

Fixes #8718
2026-08-16 11:37:59 +00:00
can1357 3e64a24714 test(agent): typed abort-signal listener without DOM lib 2026-08-16 02:19:32 +02:00
can1357 df6d1e1ac5 fix(ai): completed oneshot transient retry handling 2026-08-16 02:14:12 +02:00
can1357 b445134b4d Merge PR #8370: fix: retry transient Anthropic failures at oneshot LLM call sites (@wonjun3991)
# Conflicts:
#	packages/coding-agent/src/utils/title-generator.ts
2026-08-16 02:14:12 +02:00
roboomp 2996f16a61 fix(agent): added codex v2 compaction feature header
- Negotiated remote_compaction_v2 on Codex compatibility fetches.

- Covered explicit endpoints and non-Codex request scoping.

Fixes #8524
2026-08-14 06:56:57 +00:00
can1357 b279db1790 test: refactored test suites to eliminate time-based sleeps and polling loops
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
2026-08-13 19:32:22 +02:00
can1357 6b4823181b test: cleaned test suites and documented filtering guidelines
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
2026-08-13 08:28:42 +02:00
wonjun3991 720ac2168a fix(agent): retry handoff, branch summary and manual /compact on a blip
Each is a single, side-effect-free completion whose result is parsed after
it resolves, so one transient provider failure previously aborted the
whole operation - for /compact that left the user's context full.

Adds SummaryOptions.oneshotRetry, because both compaction paths call the
same generateSummary and the policy therefore cannot be a constant inside
it. Manual /compact has no outer loop and gets retry by default;
auto-compaction passes false because session-maintenance already retries
the whole attempt, and nesting would multiply the budget (10 outer x 3
inner) while stacking each outer wait on an inner backoff.
2026-08-13 10:54:42 +09:00
wonjun3991 0523c7112f feat(agent): add opt-in transient retry to instrumentedCompleteSimple
instrumentedCompleteSimple is the single funnel for every oneshot LLM
call in the agent, so retry belongs here rather than in a try/catch above
each caller - the failure arrives as a resolved AssistantMessage.

Opt-in rather than default-on: oneshotKind is free-form and callers may
pass arbitrary ctx.tools, so the funnel cannot itself prove a request is
replay-safe.

Response headers are captured per attempt and cleared between attempts,
so a stale retry-after can never be reused for a later failure.
2026-08-13 10:54:42 +09:00
left-to-right 835a7db15a fix(compaction): keep rescue shake able to elide the newest result
RESCUE_SHAKE_CONFIG spreads AGGRESSIVE_SHAKE_CONFIG, so the new 4k manual
tail leaked into dead-end recovery and could block eliding the very
result that caused the dead end. Override protectTokens back to 0 in the
rescue preset and add the coding-agent changelog entry for the manual
/shake behavior change.

Addresses review on #8067.
2026-08-11 22:33:40 +08:00
left-to-right c3da093a1d fix(compaction): manual /shake keeps a recent tail of tool results
Manual /shake used protectTokens: 0, stripping every eligible tool result
including the ones the agent is still working from. Keep a small 4k-token
recent window (matching the automatic shake mechanism, at a quarter of its
16k budget) so the full escape hatch stays aggressive without destroying
the live tail. Two matcher-focused tests that implicitly relied on the
zero window now pin protectTokens: 0 explicitly.

Fixes #7776
2026-08-09 18:09:17 +08:00
can1357 1b7b8bb0c7 Merge PR #7743: fix(agent): strip output statuses from remote compaction (@roboomp) 2026-08-05 22:15:46 +02:00
roboomp 86a856361a fix(agent): stripped output statuses from remote compaction
- Reused the Responses replay lifecycle policy for V1 and V2 compaction input.

- Covered persisted native history and converted assistant replay items.

Fixes #7742
2026-08-05 17:26:32 +00:00
can1357 e9888367d1 refactor: migrated packages to internal utility modules and removed external dependencies
- Implemented in-house, zero-dependency utility modules in `pi-utils` covering DOM manipulation, markdown parsing, templating, browser automation helpers, and terminal buffers.
- Migrated packages across the repository to consume the new internal utilities and `omptype` schema validators instead of external dependencies.
- Removed multiple external runtime and development dependencies including Zod, Marked, LRU cache, Turndown, and Puppeteer browser packages.
2026-08-05 13:39:09 +02:00
can1357 94c838faea Merge PR #7539: fix(coding-agent): complete usage-aware fallback integration (@eggpeat) 2026-08-05 01:12:02 +02:00
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
Brent 2da96d77a7 fix(coding-agent): close fallback lifecycle races 2026-08-03 17:51:32 +00:00
Brent 353bbc034b fix(coding-agent): harden usage fallback review races 2026-08-03 17:15:57 +00:00
Brent db97103c32 fix(coding-agent): complete usage-aware fallback integration 2026-08-03 16:34:35 +00:00
can1357 6a44844c57 Merge PR #7377: fix(hub): wake owners when supervised processes exit (@paralin) 2026-08-03 15:12:03 +02:00
can1357 d841569ee8 Merge PR #7460: fix(agent): distinguish parent steering from advisories (@fatihaziz) 2026-08-03 14:46:14 +02:00
roboomp 4997423f10 fix(agent): keep queued non-interruptible tools alive on peer-IRC interrupt
A peer-IRC interrupt (e.g. a subagent message) aborts only interruptible
waits and leaves already-running non-interruptible foreground work alone.
But runTool's `interruptState.triggered` early-return skipped every
not-yet-started tool regardless of source, so a non-interruptible tool
queued behind an interruptible wait in the same batch (a batched todo/write
after `hub wait`) was dropped with "Skipped due to pending peer interrupt".

Exclude non-interruptible tools from the early skip on the IRC path; user
and system steering still preempt all queued work.

Fixes #7493
2026-08-03 12:24:10 +00:00
Christian Stewart 302148523f fix(hub): wake owners when supervised processes exit
Publish terminal daemon completions to the session that started the
process so idle agents can resume without polling hub status.

Persist every unacknowledged generation with a stable completion ID and
immutable snapshot. Replay the collection after reconnect or broker
recovery, and clear each event only after the owning client acknowledges
it.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-08-03 01:31:39 -07:00
Fatih Al-Aziz e4de9bf176 fix(agent): classify attributed custom user steering
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 13:32:25 +07:00
Fatih Al-Aziz 2416134cf0 fix(agent): classify next queued steering source
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 13:18:45 +07:00
Fatih Al-Aziz 1e448304ac Merge remote-tracking branch 'origin/main' into fix/clarify-interrupt-skip-wording 2026-08-03 13:14:07 +07:00
Fatih Al-Aziz f8cc29e8db fix(agent): distinguish parent steering from advisories
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 13:05:20 +07:00
Fatih Al-Aziz c61ed66dd3 fix(agent): clarify internal steering skip message
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-03 12:54:08 +07:00
can1357 a58584ff3c feat: implemented shared serialization utilities and standardized inventories
- Added shared Python call and literal serialization utilities with multiline verbatim support.
- Standardized tool inventories to format as an OpenAI-Harmony functions namespace using TypeScript declarations.
- Updated tool normalization and rendering functions to accept options objects and default to Python-syntax examples.
- Refactored Gemini dialect rendering to leverage shared serialization functions directly.
2026-08-03 05:06:49 +02:00
can1357 a872d77068 chore: cleanup dumb tests 2026-08-02 20:39:23 +02:00
can1357 4e65a685ae fix(ai): addressed anthropic stream truncation errors for tool call recovery
- Added `isStreamEnvelopeErrorText` to packages/ai/src/error/flags.ts to recognize stream envelope truncation errors.
- Updated `streamAnthropicOnce` in packages/ai/src/providers/anthropic.ts to throw an envelope error when streams die mid-generation without a terminal frame.
- Updated `recoverTransientErrorToolTurn` in packages/agent/src/agent-loop.ts to recognize Anthropic stream envelope truncation errors for tool call salvage.
2026-08-02 06:14:01 +02:00
can1357 1088ee2127 Merge PR #7203: fix(tui): render mid-turn steering skips as info, not errors (@roboomp) 2026-08-01 20:14:40 +02:00
can1357 81cb1818fe fix(ai): roll back failed compaction metadata 2026-08-01 20:13:28 +02:00
roboomp 636f9eb0a5 fix(ai): preserved turn-state metadata from compaction websockets
- Applied response.metadata turn-state/models-etag refreshes while draining Codex V2 compaction WebSocket events.
- Deferred the update until the WebSocket attempt succeeds so a discarded attempt cannot leak turn state.
- Added coverage for a mid-turn compaction refreshing x-codex-turn-state.

Fixes #7198
2026-07-31 21:22:23 +00:00
roboomp af343f70d0 fix(agent): distinguished started steering aborts
Track entry into tool.execute separately from tool event emission. Never-started skips retain SyntheticToolResultDetails with executed:false; in-flight aborts now use distinct interrupted metadata with execution:started so consumers do not assume no partial work occurred.

Keep both interrupt states neutral in the TUI and cover the agent metadata boundary plus rendering behavior.

Fixes #7199
2026-07-31 21:16:28 +00:00
roboomp 6c205fab9f fix(ai): discarded partial websocket compaction replays
- Buffered WebSocket compaction events until terminal completion.
- Discarded buffered events when transport failure triggers an SSE replay.
- Added coverage for failure after a partial compaction output item.

Fixes #7198
2026-07-31 21:08:06 +00:00
roboomp 5283c4c99b fix(agent): routed codex v2 compaction through websockets
- Reused the live Codex provider session for WebSocket-first V2 compaction.
- Fell back to SSE V2 on WebSocket transport failure before the existing V1 fallback.
- Propagated the configured WebSocket preference through manual, automatic, and advisor compaction paths.
- Added transport reuse and fallback regression coverage.

Fixes #7198
2026-07-31 21:00:15 +00:00
can1357 5764021c81 fix(compaction): cap every summary transport 2026-07-31 19:42:43 +02:00
can1357 0d9c634665 Merge PR #7008: fix(compaction): cap the generated summary output budget (@terrxo) 2026-07-31 19:42:42 +02:00