Commit Graph

11338 Commits

Author SHA1 Message Date
can1357 2adf484ef1 Merge PR #7042: fix(lsp): handle quick exits before reader teardown (@roboomp) 2026-07-29 23:09:13 +02:00
roboomp dca8f44b73 fix(lsp): handled quick exits before reader teardown
Waited briefly for process exit publication after clean stdout EOF so the process handler preserves the real exit code and stderr, while genuine reader errors still tear down immediately.

Cleared only the matching initialization failure for explicit reloads and added regressions for quick exits, reader errors, ordinary backoff, and immediate reload retries.

Fixes #7041

(cherry picked from commit a76522b759f14421202d4cc437ec611b78be20d1)
2026-07-29 23:09:13 +02:00
can1357 672857d238 Merge PR #7025: fix(coding-agent): coalesce models config resource probe (@paralin) 2026-07-29 23:09:10 +02:00
Christian Stewart a43c7a4d36 fix(coding-agent): coalesce models config resource probe
The models config resource regression started two cold child processes inside one five-second test. Under parallel CI chunk load, the second child could still be waiting for pipe drain or process exit after the validator had completed.

Keep the process boundary as the owner of the lifecycle measurement and run the missing and custom phases in one child. The missing phase closes its storage and registry before the baseline snapshot, while the custom phase proves schema identity and retention without changing config loading or validator cleanup semantics.

Signed-off-by: Christian Stewart <christian@aperture.us>
(cherry picked from commit 81fa98491544c7fbcbf075922889e0a9e1a0a3b4)
2026-07-29 23:09:10 +02:00
can1357 1b3f01d197 Merge PR #7019: test(ai,catalog): give spawn-based lazy tests explicit timeouts (@roboomp) 2026-07-29 23:09:07 +02:00
roboomp cc04600a64 test(ai,catalog): give spawn-based lazy tests explicit timeouts
Spawn-based lazy-loading tests assert exitCode===0 on a child process but
set no per-test timeout, so bun's 5s default kills the child under CPU
contention and the assertion reports a dead child rather than a regression.

Give each spawn test an explicit 60s per-test timeout, matching the
existing precedent in auth-gateway-anthropic-caching.test.ts.

Fixes #7018

(cherry picked from commit 00e5ec8855eb8ba31c1bdf571bd7c16f2405d679)
2026-07-29 23:09:06 +02:00
can1357 1b4c9d7d1a Merge PR #7015: perf(prompts): streamline tool guidance (@usr-bin-roygbiv) 2026-07-29 23:09:05 +02:00
usr-bin-roygbiv 066a3239f8 fix(prompts): preserve supported tool routes
(cherry picked from commit 0b02fb9219f07e212d7a0f67dacd7747b622ee45)
2026-07-29 23:09:05 +02:00
usr-bin-roygbiv 1e0d352a72 perf(tools): streamline shell guidance
(cherry picked from commit 9039728d89f07852904962685581c752ffebcdc6)
2026-07-29 23:09:05 +02:00
can1357 def7c5fadd test(task): exercise bundled budget through subprocess
(cherry picked from commit 6e023cfcebebc429cc020f0f823604eb64eb9c18)
2026-07-29 23:09:03 +02:00
can1357 d2d9c81c84 Merge PR #7012: fix(task): let task.softRequestBudget lower bundled subagent budgets (@terrxo) 2026-07-29 23:09:03 +02:00
can1357 adad262ba9 fix(xdev): include truncation marker in summary byte cap
(cherry picked from commit aa2067bf7952191beab85b71002aafe812f544bc)
2026-07-29 23:09:01 +02:00
Nik Divjak c3011fff3c fix(task): let task.softRequestBudget lower bundled subagent budgets
The soft request budget resolved to `SOFT_REQUEST_BUDGET[agent.name] ??
configured`, so the bundled entries for scout and sonic replaced the
configured value outright. Lowering `task.softRequestBudget` to tighten
the guard therefore did nothing for exactly the two agents that spawn
most often: a scout kept its 100-request budget no matter how small the
user set the knob. Only 0 (disable) and raising the value for
non-bundled agents had any effect.

Treat both numbers as upper bounds and take the smaller one. The bundled
entries stay ceilings, so a runaway scout is still stopped at 100 by
default and existing behavior is unchanged for anyone who has not
lowered the setting; a configured 0 still disables the guard entirely.
Resolution moves into `resolveSoftRequestBudget`, which also normalizes
negative and fractional inputs, so the rule is testable without standing
up a subprocess run.

This composes with `task.maxEffort` on a separate axis: effort caps how
hard each request thinks, this caps how many requests a run may spend.

(cherry picked from commit f0db29f8f725f11390b64ca9342300c482ff5c5d)
2026-07-29 23:09:01 +02:00
can1357 4666b1ae41 Merge PR #7010: fix(xdev): bound device summaries in UTF-8 bytes and flag untrusted metadata (@terrxo) 2026-07-29 23:09:01 +02:00
Nik Divjak 630f9e5324 fix(xdev): bound device summaries in UTF-8 bytes and flag untrusted metadata
Catalog summaries of mounted xd:// devices are inlined verbatim into the
system prompt. External devices (MCP servers, plugins) supply that text, and
it was bounded only by character count: a summary of multi-byte script passed
roughly three times the intended budget, and control characters survived into
the prompt where they can forge structure.

Summaries now go through a single sanitize-and-bound step that strips C0/C1
control characters and bounds the result in UTF-8 bytes via the central
truncateHeadBytes helper, so a cut lands on a code point boundary and never
renders a partial code point. The built-in/external distinction is derived
once per entry, and that same boolean both selects the description cap and is
exposed as `dynamic`, so the cap and the flag cannot disagree. The prompt uses
the flag to state that dynamic summaries are untrusted metadata, and the mount
notice says the same for newly appeared devices.

(cherry picked from commit 5989da6235d820bc687779a791e655e6f1b2df0f)
2026-07-29 23:09:00 +02:00
can1357 f62526f111 fix(coding-agent): preserve VCS cache refresh semantics
(cherry picked from commit 3f3e475109deb7b9e7690720827876fc982b5b8d)
2026-07-29 23:08:59 +02:00
can1357 1d41b269f7 Merge PR #6997: perf(coding-agent): keep reftable branch resolution off the render path (@metaphorics) 2026-07-29 23:08:59 +02:00
robomp-bot 65707c7f4c fix(coding-agent): close VCS cache lifecycle gaps
(cherry picked from commit b88851334b07eb1da6743b51b542ea5d3222ccca)
2026-07-29 23:08:58 +02:00
robomp-bot c5b0348150 fix(coding-agent): harden async reftable resolution
(cherry picked from commit e451f1d6f3537d4b58dfe479a6daf8919d169397)
2026-07-29 23:08:58 +02:00
robomp-bot d966b3f8a0 perf(coding-agent): keep reftable branch resolution off the render path
(cherry picked from commit 5724c30ff66e2036ed03a2ce3514a9237a72a657)
2026-07-29 23:08:58 +02:00
can1357 650a8f0faf fix(collab): detach reconciled loader on idle
(cherry picked from commit d675de815644a5b02358fa90d83e5cd5c78141d2)
2026-07-29 23:08:57 +02:00
can1357 872a931795 Merge PR #6996: fix(collab): start the guest loader when the host reports streaming (@metaphorics) 2026-07-29 23:08:57 +02:00
robomp-bot 4557a4bb3d fix(collab): preserve maintenance loaders during state reconciliation
(cherry picked from commit ad611afa4885a71bbf8e8ad6c00328fb0be37d35)
2026-07-29 23:08:56 +02:00
robomp-bot d6dc8f5d14 fix(collab): start the guest loader when the host reports streaming
(cherry picked from commit 0bc0e38d865815b70c848bd00be3dc1fcd228ba4)
2026-07-29 23:08:56 +02:00
can1357 eb8f3e6c2c Merge PR #6993: fix(coding-agent): require web_search_call in codex search (@roboomp) 2026-07-29 23:08:55 +02:00
roboomp e5ea31b22c fix(coding-agent): require web_search_call in codex search
GPT-5.6 Responses-Lite models receive tool_choice "auto" (the forced
hosted choice is invalid under the lite shape, #5771/#5772), so the model
may answer without invoking the hosted web_search tool. The codex search
parser accepted any non-empty answer, returning a stale completion with
zero sources as a successful search.

callCodexSearch now tracks response.web_search_call.* events (and
web_search_call output items) and throws CodexNoWebSearchError when none
occurred. The candidate chain treats that error as retryable, advancing
default lite models to a non-lite model that forces web_search, and
surfaces a clear failure when the model was explicitly configured.

Fixes #6988

(cherry picked from commit a276cd0b3df1d0d041faf0a63fabcbb884e36a91)
2026-07-29 23:08:54 +02:00
can1357 9f80d24e20 fix(ai): preserve pre-stream provider error provenance
(cherry picked from commit d3c66195866de8886d05fc8002533160fa8f2767)
2026-07-29 23:08:52 +02:00
can1357 832d4f1506 Merge PR #6987: fix(coding-agent): treat streamed visible text as replay-unsafe in turn recovery (@metaphorics) 2026-07-29 23:08:51 +02:00
metaphorics 93819dbec4 fix(coding-agent): gate refusal retries on replay safety
(cherry picked from commit 531b25beffa2096c76adff6692f3697e649f0e95)
2026-07-29 23:08:51 +02:00
metaphorics 9733d13827 fix(coding-agent): make replay safety authoritative
(cherry picked from commit 118ecabb22f406afbd56864ff0fbafc8f4f1fa4c)
2026-07-29 23:08:50 +02:00
metaphorics c806a9b482 fix(coding-agent): guard Fireworks fallback after visible output
- reuse the centralized replay-unsafe predicate in Fast fallback
- exercise visible-text and replay-safe Fireworks branches

(cherry picked from commit f596392b891bb8ec82171a25c6651ed3cc8e0d61)
2026-07-29 23:08:50 +02:00
metaphorics b5602ddfc1 fix(coding-agent): treat streamed visible text as replay-unsafe in turn recovery
(cherry picked from commit 3aeac1a4afa756e9c7d579246d8f334cd0731cda)
2026-07-29 23:08:50 +02:00
can1357 4bad9b481d Merge PR #6973: fix(coding-agent): share parent local:// root with /tan clone (@roboomp) 2026-07-29 23:08:44 +02:00
roboomp bbf3d78c7e fix(coding-agent): keyed tan local root on session-manager id
Snapshot this.ctx.sessionManager.getSessionId() for the tan clone's local://
mapping instead of session.sessionId. The two diverge after /fresh or a
provider session override, and the Windows short-root fallback keys
%TEMP%/omp-local/<id> off the session-manager id used by the parent's
large-paste writes and '/data/workspaces/can1357__oh-my-pi__6971/.omp-session/2026-07-29T06-08-45-283Z_019fac7d-5ee3-7000-a7aa-16fe9394fdc9/local' reads, so the mismatched id left attachments
unreachable.

Diverge the mocked session id from the manager id in the regression test so
it pins the session-manager id.

Fixes #6971

(cherry picked from commit 1efcd22326d76fdb8b50c0977a836c892e80ab76)
2026-07-29 23:08:44 +02:00
roboomp dd4985f31e fix(coding-agent): scoped subagent local root overrides
Keep subagent localProtocolOptions on their ToolSession instead of installing
them as the process-global LocalProtocolHandler override. No-context URL
consumers therefore retain the active top-level session's mapping while tan
and task subagents continue to resolve through their caller context.

Add SDK regression coverage proving subagent creation preserves an existing
global mapping.

Fixes #6971

(cherry picked from commit a02eef174b03036dc960c842a0901d22333ad9cd)
2026-07-29 23:08:44 +02:00
roboomp fb4393ae16 fix(coding-agent): snapshotted tan parent local root
Capture the parent artifacts directory and session ID when /tan dispatches
instead of resolving them through the mutable interactive SessionManager.
This keeps background tan '/data/workspaces/can1357__oh-my-pi__6971/.omp-session/2026-07-29T06-08-45-283Z_019fac7d-5ee3-7000-a7aa-16fe9394fdc9/local' reads pinned to the dispatching transcript
after the user switches or resumes another session.

Extend the regression test to switch the mocked interactive session before
the background job starts and assert the original local mapping is retained.

Fixes #6971

(cherry picked from commit e05db428eabd5087e4b2b4a462f47927c0628e72)
2026-07-29 23:08:43 +02:00
roboomp d933cfbe05 fix(coding-agent): shared parent local root with tan clone
TanCommandController.start nests the tan clone at
<parent-artifacts>/Tan-<id>.jsonl, so the clone's session manager derived
its own artifacts dir and hence local root <parent-artifacts>/Tan-<id>/local.
Its sdk.createAgentSession call omitted localProtocolOptions, unlike the
task-subagent path which inherits the parent's mapping, so parent-session
'/data/workspaces/can1357__oh-my-pi__6971/.omp-session/2026-07-29T06-08-45-283Z_019fac7d-5ee3-7000-a7aa-16fe9394fdc9/local' attachments (pasted files, generated references) were unreadable.

Thread the parent session manager's localProtocolOptions into the tan clone
so local:// resolves against <parent-artifacts>/local.

Fixes #6971

(cherry picked from commit 1ded46e182fc24f9f57d8e9a907aaad58f790783)
2026-07-29 23:08:43 +02:00
can1357 756e872f64 Merge PR #6946: fix(tui): nest usage metrics in grouped reads (@joshrzemien) 2026-07-29 23:08:42 +02:00
joshrzemien daf1557b3d fix(tui): seal mixed read groups before usage
(cherry picked from commit 576b950b80583d8fd028edd08967b537f4eb18f0)
2026-07-29 23:08:42 +02:00
joshrzemien 4b753597fd docs(coding-agent): add grouped read changelog
(cherry picked from commit f23d576925906e6a39bfab44d421a52494a8f20b)
2026-07-29 23:08:42 +02:00
joshrzemien 57cb290fef fix(tui): preserve grouped read request order
(cherry picked from commit eb6767e3248ad0fe3ea371e5226a4af6e0b8f97b)
2026-07-29 23:08:42 +02:00
joshrzemien a3e90b3d95 fix(tui): nest usage metrics in read groups
(cherry picked from commit 101268adc632e80bb3f8304b95497ac7bd884131)
2026-07-29 23:08:41 +02:00
can1357 c6db470db5 test(coding-agent): cover drifted ACP write snapshot
(cherry picked from commit a8ef3465929e0ebdc026e0b6f94ea65d491946c8)
2026-07-29 23:08:39 +02:00
can1357 4cc259f6cf Merge PR #6934: fix(hashline,coding-agent): key snapshot tag on actually-persisted content after ACP bridge writes (@marton78) 2026-07-29 23:08:39 +02:00
Márton Danóczy 5c0b52a141 test(coding-agent): use declared AgentToolResult type in ACP bridge test
Replaces `Awaited<ReturnType<typeof executeHashlineSingle>>` with the
function's declared `AgentToolResult<EditToolDetails, typeof
hashlineEditParamsSchema>` return type. AGENTS.md bans ReturnType<>.

Review: https://github.com/can1357/oh-my-pi/pull/6934
(cherry picked from commit 59dbc5b79f8857bab67559656eae96cf7dfb6ac0)
2026-07-29 23:08:39 +02:00
can1357 d2283dd301 fix(model-registry): isolate modifier record mutations
(cherry picked from commit 3941103443576a6c8ab91a52e7da4ba1d8521d69)
2026-07-29 23:08:38 +02:00
Márton Danóczy a5d01deb85 test(coding-agent): hoist hashline import to module scope in ACP bridge test
Replaces three inline `await import("@oh-my-pi/hashline")` calls with a
single top-level `computeFileHash` import. AGENTS.md forbids dynamic
imports.

Review: https://github.com/can1357/oh-my-pi/pull/6934
(cherry picked from commit 9ffda46c94de0146a8e77ec816640918a565a0ea)
2026-07-29 23:08:38 +02:00
can1357 75b77c08b5 Merge PR #6930: fix(model-registry): preserve oauth.modifyModels projection across reloads (@abhishekbiyala) 2026-07-29 23:08:38 +02:00
Márton Danóczy 7708f372b5 fix(hashline,coding-agent): key snapshot tag on actually-persisted content after ACP bridge writes
Root cause of the reported "edit tool silently reformats the whole
file" corruption: fs/write_text_file has no verbatim guarantee. When
an ACP client (e.g. Zed with format_on_save: on) reformats a buffer
on save, routeWriteThroughBridge reported the pre-write content as
successfully written, and Patcher.commit keyed the returned snapshot
tag on that same pre-write text instead of what actually landed on
disk. The next edit anchored on that tag then resolved hunks against
a baseline the file had already drifted away from, which is what
produced whole-file "corruption" from single-line hunks -- reproduced
live in this session against real Swift/JSON/TypeScript files with
Zed as the ACP client.

- routeWriteThroughBridge reads the file back after the bridge write
  and returns the verified content plus a drift flag (best-effort:
  ACP defines no ordering between the client acking the write and its
  own async format-on-save settling, so this degrades gracefully to
  the old stale-tag-on-next-read failure mode, never to corruption).
- HashlineFilesystem.writeText propagates that verified content in
  view-space (the same space readText returns -- e.g. a notebook's
  editable cell text, not its raw JSON), not storage-space, so tag
  validation on the next edit compares like with like.
- Patcher.commit keys fileHash/header/snapshot on the verified
  post-write content (normalized, so BOM/line-ending restoration never
  produces a false "drift") when it diverges from what was sent, and
  appends a warning naming the drift -- but deliberately leaves the
  returned `after` (and therefore the model-visible diff) scoped to
  the intended hunk. Diffing against the full drifted file would
  balloon the tool response to span every reformatted line (measured
  ~6.8x inflation on a 245-line file with one touched line); the
  warning is the correct O(1) channel for "your editor reformatted
  this," not an O(file-size) diff.
- write.ts keys its own snapshot header on the verified bridge content
  too (no diff-size concern there since write always replaces the
  whole file).

Caught via code review (dispatched against the first pass of this
fix): a naive "just use the verified content everywhere" fix broke
.ipynb editing outright (write-space vs read-space content mismatch,
tag invalid on every notebook edit) and would have inflated every
drifted edit response by ~6.8x. Both are now covered by regression
tests that fail against the pre-fix code and pass against this one.

(cherry picked from commit 35ab80e43be5800b2f48728e4400eb9fd7f7f7d2)
2026-07-29 23:08:38 +02:00
Abhishek Sharma c7a113c9cb refactor(model-registry): derive projected catalog from an unprojected snapshot
The previous commits patched each rebuild path individually to avoid feeding a
modifyModels hook its own output. That left the invariant implicit and the
provider-scoped path applying only a subset of hooks, which is wrong for a hook
that inspects or suppresses another provider's models.

Keep #unprojectedModels as the canonical pre-projection catalog and derive
#models from it at every mutation point, so projections are always a pure
function of the unprojected base:

- #composeUnprojectedStaticModels builds the catalog; #composeStaticModels
  projects it. A scoped lookup with modifiers registered composes and projects
  the whole catalog before narrowing, matching getAll() followed by a filter.
  Providers without modifiers keep the cheap filtered path.
- Discovery completion, registerProvider, and runtime transport overrides
  update the unprojected snapshot and reproject, instead of mutating an
  already-projected array.
- Runtime metadata patches apply to the unprojected model, then reproject, so
  a later registration cannot discard them.
- Provider lookup snapshots are invalidated wherever the projection changes.

Hooks no longer take a providerFilter: a modifier is a whole-catalog transform
and every rebuild now runs the full ordered set exactly once.

(cherry picked from commit e5d2e9eac7c371cc196e9b362f77d3a5d7bdf507)
2026-07-29 23:08:37 +02:00