- Added throttling and debouncing to HUD data rendering and observer UI synchronization to coalesce update bursts.
- Constrained the subagent HUD display to a maximum of 8 rows with a truncation notice for hidden sessions.
- Enhanced the session observer registry to categorize update types, enabling more granular UI reconciliation.
- Verified render coalescing and display truncation behavior with comprehensive integration tests using fake timers.
Drained pending IRC asides before parking irc wait so replies that arrive between wait calls are returned instead of being treated only as queued interrupts.
Added regression coverage for the already-aborted queued-IRC signal path and documented the fix in the coding-agent changelog.
Fixes#4657
Reset per-turn maintenance counters before IRC wake prompts so yielded subagents do not carry stale yield termination into later wake turns.
Add regression coverage for empty-stop retry after an IRC wake following a yielded run.
Fixes#4658
- Implemented streaming synthesis in the `say` command to allow processing of arbitrarily long text without hitting model phoneme limits.
- Added file input support via the `--file` flag and updated the CLI to prevent conflicting arguments.
- Refined `SpeakableStream` segmentation logic to prioritize valid sentence and clause breaks within the maximum segment length when processing large text buffers.
- Added comprehensive tests for stream segmentation behavior under long-form input.
- Propagated llama.cpp /props input modalities through selected-model runtime refresh.
- Added a regression test for cached text-only local vision models becoming image-capable after refresh.
- Updated the coding-agent changelog for the local vision detection fix.
Fixes#4654
Skipped the home directory during Claude project skill walk-up so disabling Claude user skills cannot reload the same files as project skills.
Added regression coverage for the home-skill duplicate path with an enabled agents fallback.
Fixes#4648
Applied source toggles before skill-name dedup so disabled higher-priority providers no longer hide enabled lower-priority authored skills.
Added regression coverage for disabled claude versus enabled agents duplicate names and managed dead-last behavior.
Fixes#4648
Codex P2 finding on commit 9632ed5: #friendlyNameCollidesWithSecret tested
a regex-entry friendlyName's RAW spelling directly against the pattern,
which can never match a label that is already normalized (uppercased,
separators stripped) even when that label IS the normalized rendering of
a value the regex actually discovers -- e.g. friendlyName: "TOKABC123"
for content: "tok_[a-z0-9]+" discovering literal tok_abc123. Nothing
compared the label against the actual matched value either. The check
now also compares the sanitized label against the sanitized value of the
secret currently being minted (reusing #prefixIsSecretShaped), catching
this on the secret's first mint before it's recorded as previously
discovered.
Codex P2 finding on commit 9b2e14d: #friendlyNameCollidesWithSecret
compared a secret's full sanitized value against the already 32-char-capped
(sanitizeSecretFriendlyName) friendly name, so a secret whose sanitized
form exceeds MAX_FRIENDLY_NAME_LEN could never be fully contained in the
truncated label -- the collision went undetected and the secret's first
32 sanitized characters leaked as an accepted placeholder prefix. The
collision check now runs against the full, untruncated sanitized label
(sanitizeForCollisionCheck(friendlyName)); the 32-char cap is applied
only afterward, to the label actually used for display.
Codex P2 findings on commit 029e838:
- secrets/index.ts:189: loadFriendlyName pre-sanitized the friendlyName
before storing it on the SecretEntry, silently defeating the raw-label
regex collision check for every secrets.yml-loaded entry. The loader now
preserves the original, unsanitized string (still validating it sanitizes
to something non-empty).
- obfuscator.ts:1232: the forged-alias guard (isGeneratedPlaceholder)
compared the dropped prefix against RAW plain-secret values, so a
lowercase/punctuated secret's normalized rendering slipped through. Both
the plain-secret-value and obfuscateMappings loops now normalize the
compared value the same way the prefix is already constrained to.
Self-discovered while verifying the above: deobfuscate()'s bare-alias
fallback had NO prefix validation at all (unlike obfuscate()'s guard),
so a forged token wrapping any real placeholder's hash suffix in a
secret-shaped prefix would restore to that secret's raw value on the
live provider-output/tool-call-argument path -- strictly worse than the
obfuscate-direction leak. Extracted the shared check into
#prefixIsSecretShaped and reused it in a new #lookupLiveAlias gate for
deobfuscate(), verified a genuine friendly-name rename still round-trips.
Codex P2 finding on commit dff2a8d: #isGeneratedPlaceholder's forged-alias
guard only checked a dropped friendly-name prefix against exact previously-
discovered secret strings, recorded in whatever casing they first turned
up in. A case-insensitive (or other flag-variant) regex only ever records
the one casing it actually discovered, so a forged token wrapping a
differently-cased occurrence of that secret-shaped text around a real
bare-alias suffix matched neither exact-string check and sailed through
as an already-redacted placeholder, leaking the secret-shaped text
verbatim. The guard now also tests the dropped prefix directly against
every configured regex pattern.
Codex P2 finding on commit 7d3a3a2: #friendlyNameCollidesWithSecret ran a
configured regex entry's pattern against the already-sanitized (uppercased,
separator-stripped) friendly name, so a case-sensitive/punctuated pattern
like tok_[a-z0-9]+ never matched the sanitized label even when the raw
friendlyName was itself a live match for that regex — letting a
secret-shaped label slip through and stamp into every placeholder minted
for it. The regex check now runs against the raw, pre-sanitization label,
matching how the regex would encounter that text verbatim.
Two Codex P2 findings on commit 732f725:
- A default (no custom replacement) mode: "replace" regex that cannot
escape a 1-2 char match (e.g. ".", "[\\s\\S]", "[\\s\\S]{2}") had its
key-derived same-length fallback marker returned without checking it
against the matched value. Since that marker is drawn from an alphabet
the regex has already proven to match exhaustively, a real 1-2 byte
secret coinciding with it would ship unredacted. Such entries are now
rejected: dropped with a warning when loaded from secrets.yml, dropped
silently as a construction-time backstop otherwise.
- #friendlyNameCollidesWithSecret compared the sanitized (uppercased,
alnum-only) friendly name against each secret's raw value, so a
friendlyName that was a lowercase or punctuated variant of its own
secret slipped through and stamped most of the secret into the
placeholder. The secret value is now sanitized the same way before
comparing.
Stop advertising eval in the default prompt and workflow notice when no eval
backend is enabled. Gate bash guidance on live eval backend availability and
cover the disabled-backend rendering contract.
Agent-Milestone: tooling: hide eval prompt guidance when eval backends are disabled
Signed-off-by: Christian Stewart <christian@aperture.us>
Treat timeout 0 as an explicit no-deadline contract across the bash tool, executor, async job, and PTY paths.
Signed-off-by: Christian Stewart <christian@aperture.us>
sendErrorNotification now reads the settled turn from event.messages,
but sendCompletionNotification still read viewSession.getLastAssistantMessage().
For a classifier-refusal turn that stale/undefined lookup no longer
matched 'aborted'/'error', so with completion.notify=on the same
failed turn fired both the error toast and a misleading 'Complete'
toast.
Thread the same agent_end event into sendCompletionNotification so
both gates read one consistent source of truth.
- Fix #generateRegexReplacement's pathological (match-everything) fallback
emitting a 1-2 byte matched value unchanged when it was exactly `Z`/`ZZ`,
the shared sentinel #generateReplacement uses for such short values.
Falls back to a same-length, key-derived run instead, which stays a fixed
point under re-obfuscation without being a public, guessable constant.
- Default getSecretPlaceholderKey()/getExistingSecretPlaceholderKey() to
getAgentDir() instead of getConfigRootDir(), matching the directory
createAgentSession() actually passes.
- Isolate the getSecretPlaceholderKey test suite under a fresh $HOME/temp
agent dir instead of the real homedir, fixing an EACCES failure in
sandboxed review environments.
- Revert an unrelated gc session-ordering tie-breaker bundled into this
branch's history; out of scope for the secrets/friendly-name feature.
- Relocate this PR's CHANGELOG.md entries out of already-released sections
(16.3.0, 16.3.5) into [Unreleased], where a stale merge had left them,
and drop a duplicate blank line and duplicate serverSideFallback/
softRequestBudgetNotice entries the same merge introduced.
This branch tracked main forward through many merge commits over its
long life; four files carried stale fixups for intermediate states of
main that current main never needed (a test-title rename, retimed
pi-native stream fixtures, a mermaid-cache type refactor, and a
multi-path test rewrite). None relate to error.notify, and current
upstream/main's own versions of these files already pass. Restore them
to keep this PR scoped to the error-notification feature.
Classifier-refusal failures end a turn with stopReason === "error" but
get pruned from the active context (agent-session.ts's
#removeAssistantMessageFromActiveContext) before agent_end fires.
sendErrorNotification() read viewSession.getLastAssistantMessage(),
which reflects that mutated context and silently missed the
notification for exactly the turns it should fire on.
Thread the agent_end event through #handleAgentEnd -> #finishAgentEnd
so sendErrorNotification reads the turn's own outcome from
agent_end.messages instead.
Resolves the two Codex P2s raised on #4420 that merged unaddressed:
- wrapUrlRows indented every continuation chunk. A multi-row terminal
selection includes the newline plus that indent; address bars strip
newlines but preserve or percent-encode embedded spaces, so the
reassembled URL was corrupted at every chunk boundary - silently,
when the damage landed inside a query value. Chunk rows now carry
zero leading bytes (label rows keep their indent), and the test
reassembly helper concatenates chunks raw instead of stripping the
indent that previously masked exactly this defect.
- #launchUrlIfSafe advertised a localhost /launch copy target for
flows whose redirectUri never returns to the loopback server. Its
catch-comment assumed custom-scheme URIs are non-parseable, but
new URL('vscode://gitlab.gitlab-workflow/authentication') parses
fine and sailed through the pathname check. The guard now requires
an http(s) loopback redirectUri (localhost / 127.0.0.1 / [::1]);
custom schemes, non-loopback hosts, and unparseable URIs all
suppress the launch URL. Regression tests cover the GitLab Duo
vscode:// shape and a fixed non-loopback HTTPS redirect.
Refs #4418
Follow-up to the #4420 opener hardening: absolute-path rundll32 fixes the
stripped-PATH spawn throw, but rundll32 exits 0 unconditionally, so the
delayed-failure telemetry added there can never observe a Windows launch
failure. Replace it with %SystemRoot%-resolved PowerShell Start-Process
via -EncodedCommand:
- failures ShellExecute itself reports (missing target, no handler
executable, access denied) surface as exit code 1 and reach the
existing non-zero-exit logging (verified live on Windows 11: missing
file exits 1; unregistered schemes exit 0 on any opener because the
OS hands them to the app-picker — documented limitation);
- the UTF-16LE/base64 payload keeps OAuth query strings opaque to
cmd/PowerShell metacharacter parsing; embedded single quotes are
doubled into a PS literal;
- %SystemRoot% anchoring with a bare-name PATH fallback preserves the
stripped-PATH resilience from #4420.
Also pins the WSL-mount test's path.resolve against Windows dev hosts
so the mocked linux platform stays deterministic.
Refs #4418
Without --binary, `git diff --cached` writes `Binary files ... differ` stubs
for staged binary files. runSplitCommit reset the index and then fed those
stubs to `git apply --cached --binary`, which cannot reconstruct the content,
so a split plan that included `bun.lockb` (or any other staged binary) would
crash the apply step after the reset had already cleared the index.
Pass `binary: true` when capturing `stagedDiff` so the patch text carries the
real binary payload and the executor can re-stage it per commit group.
Refs #4632, #4634 review
git_overview hides EXCLUDED_LOCK_FILES from the model so lock files never drive
split decisions, but runSplitCommit then re-fetched the raw staged set and
rejected any plan that failed to enumerate them, aborting `omp commit` with
"Split commit plan missing staged files: <lockfile>". Skipping the validator
would have masked a real drop — the executor resets the index and only
re-stages files listed in each commit group.
Introduce packages/coding-agent/src/commit/agentic/lock-files.ts with a
LOCK_FILE_MANIFESTS map and an assignLockFilesToPlan helper that attaches each
orphaned lock file to (1) the commit group touching a sibling manifest in the
same directory, (2) any commit group touching a matching manifest, or (3) the
last commit group. git-overview.ts imports EXCLUDED_LOCK_FILES from the shared
module so the filter and the pairing table stay in sync.
Fixes#4632
Two Codex bot findings from earlier PR reviews were still open.
1. local:// URL selector shadow (read.ts): the local:// branch resolved
`local://foo:1-2` and rewrote readPath to `${localFile.path}:${sel}`,
then let splitPathAndSelPreferringLiteral run on the synthesized
string. A sibling literal `${localFile.path}:${sel}` file would win
over the intended URL selector semantics. The branch now promotes the
URL selector into the explicit-selector state and sets
readPath = localFile.path, so downstream literal-preferring routing
never re-splits the concatenation.
2. Delimited expansion before literal probe (path-utils.ts): grep called
expandDelimitedPathEntries before parsePathSpecs, and
splitDelimitedPathEntry only checked whether the peeled base of the
entry resolved. A real POSIX file whose name contained a delimiter
plus a selector-shaped tail (a;b:1-2) got split into ["a", "b:1-2"]
and never reached the literal-preferring probe. splitDelimitedPathEntry
now short-circuits on probeLiteralPathExists — "missing" is the only
outcome that lets delimiter expansion run.
Added regressions: `read local://notes.md:1-2` still slices the base file
when a sibling `notes.md:1-2` literal exists, and grep searches a real
`a;b:1-2` file without semicolon-splitting.
- Allowed implicit default fallback resolution when other role fallback chains are configured.
- Covered the mixed-role first-run fallback case.
Fixes#4533
The generate_image tool gains an optional 'provider' param
(auto|openai|openai-codex|antigravity|xai|gemini|openrouter): say the
provider in chat and the tool uses it for that call; absent, it falls back
to the providers.image setting, then auto-detect. openai-codex now works
INDEPENDENT of the active chat model: a connected Codex (ChatGPT OAuth)
subscription drives OpenAI's hosted image_generation tool (model priority
gpt-5.5 -> gpt-5.4 -> gpt-5.1 -> gpt-5 -> gpt-5-codex), so images ride the
subscription instead of the metered API key. providers.image accepts
openai-codex. (Recovered from parked lane dbf87cd04; oauth.html rebrand
left parked.)
resolveExistingReadPath treated any stat failure other than ENOENT/ENOTDIR as
"exists" and any other resolved path was considered a hit. That silently
reinterpreted a real literal path such as test:1-2 as test plus selector 1-2
whenever the raw path was a dangling symlink, sat under an unreadable parent,
or hit a transient I/O error.
The new probeLiteralPathExists returns "exists" / "missing" / "unknown" from
an lstat probe. splitPathAndSelPreferringLiteral now falls back to the strict
selector split only on "missing"; both "exists" and "unknown" keep the raw
path, so an unreachable literal is never reinterpreted. Grep and read use
the same probe: the explicit selector branch keeps the literal path when
existence is uncertain, and only a definitive ENOENT/ENOTDIR lets structured
archive/sqlite/pdf dispatch take over.
Added regressions covering probeLiteralPathExists exists/missing/dangling-
symlink cases and splitPathAndSelPreferringLiteral over a dangling symlink.
The literal-path stat fallback made selector-shaped filenames accessible, but it did not give callers a deterministic way to read or grep a range from a literal filename such as test:1-2. Encoding that as test:1-2:1-2 remained recursively ambiguous if a longer literal file later appeared.
Read now accepts an optional selector field that is parsed independently from path. When selector is present, path is treated as the exact path first, so { path: "test:1-2", selector: "1-2" } always means lines 1-2 from the literal file test:1-2. Inline :<sel> remains supported for compatibility.
Grep now accepts an optional line-range selector field with the same literal-path behavior. Explicit selectors bypass path suffix peeling, while archive/internal/URL routing still handles non-literal structured paths.
Updated read/grep tool prompts and added deterministic regressions proving that a longer literal file like test:1-2:5-6 or test:1-2:2-2 does not change the meaning of { path: "test:1-2", selector: ... }.
splitPathAndSelPreferringLiteral only statted resolveToCwd(rawPath), so shell-escaped paths such as dir/a\ b:1-2 missed the existing dir/a b:1-2 file and fell back to the strict selector peel. That let read target dir/a\ b with a range instead of the literal filename.
The helper now probes resolveReadPath(rawPath, cwd), reusing the read path resolver's existing escaped-space and filesystem variant normalization before deciding whether the literal file exists.
Grep also stores the resolved filesystem path for literal matches so the later search-scope parser does not reinterpret backslashes as path separators. Regression coverage now includes helper, read, and grep cases for dir/a\ b:1-2.
parsePathSpecs preserved an existing literal path like data.zip:1-2, but resolveArchiveSearchPaths only received the cleaned path strings and reparsed the same literal as archive data.zip plus member 1-2. If data.zip existed, grep materialized or errored on the archive member before searching the literal file.
GrepPathSpec now carries whether a local entry was kept because the raw filesystem path exists. Archive materialization consumes the specs instead of bare strings and skips those literal matches, while ordinary archive selectors still materialize as before.
Regression coverage adds grep over data.zip:1-2 with a real data.zip alongside, proving the literal file is searched instead of the archive member.
The prior hunk placed the literal-preferring split after resolveArchiveReadPath,
resolveSqliteReadPath, and splitPdfImageMemberReadPath, so a real POSIX file
such as data.zip:1-2 or notes.db:1-2 still got hijacked: the archive/sqlite
resolvers matched the base extension, opened data.zip / notes.db, and errored
on the phantom :1-2 member before the literal file was ever considered.
Now the async splitter runs first. When the strict grammar would have peeled
a suffix but the literal path stats successfully, all three structured
dispatchers decline. Otherwise the ordering is unchanged, so archive/sqlite/
pdf-image reads keep working when the literal file does not exist.
Regression coverage adds `data.zip:1-2` and `notes.db:1-2` cases where the
base archive/sqlite file also exists on disk, exercising the exact ordering
bug the reviewer flagged.
splitPathAndSel unconditionally peels a trailing :<sel> chunk whenever it
matches the read-tool selector grammar (raw, conflicts, N-M, N+K, ...). On
POSIX, filenames may legitimately contain colons, so a real file named
test:1-2 or log:raw was shredded to test/log before either read.ts or
grep.parsePathSpecs stated anything and both surfaced "Path not found".
Added splitPathAndSelPreferringLiteral(rawPath, cwd) alongside the strict
splitter: it only overrides the peel when fs.stat succeeds against the raw
path. Read (execute) and grep (parsePathSpecs) call the async variant for
non-URL paths; internal-URL splitting stays unchanged. Regression covers
splitter fallbacks, read/grep behavior on literal-colon files, and that
:1-2 selectors still work when the base file is the only real match.
Fixes#4618