CI evidence overturns the chunking half of this branch. Splitting the
singleton bucket into 8 chunks made the job run 11m27s and go red, where
the unchunked bucket passes in ~1m31s on a concurrent PR. The duration
matches the 600s chunk watchdog, i.e. a chunk wedged rather than merely
running slowly -- the "these suites co-locate in one process" comment was
describing a real coupling, so restore it and the original plan entry.
The failure-attribution change is kept and is independently justified:
this run is precisely the case where a bare "exit code 137" cannot
distinguish the watchdog from the OOM killer.
The singleton/global-state bucket was the only bucket left unchunked, and
it is the one that got OOM-killed: 79 files in a single `bun test`
process exited 137. Chunking exists precisely to hold peak RSS under the
runner's ceiling, so the bucket that opts out is the bucket that hits it.
The stated reason for leaving it whole does not hold. That bucket is
selected because its suites mutate process-wide state, and what keeps
them from colliding is sequencing, which `parallel: 1` already
guarantees. Splitting them across processes strictly increases isolation
rather than reducing it. Use 10 files per chunk, the width the 650-file
native bucket already sustains on the same runner.
Separately, the sequential path could not say why a chunk was SIGKILLed.
It arms a watchdog but never records when the watchdog fires, so an OOM
kill and a watchdog kill both surfaced as a bare "failed with exit code
137" - the exact message this failure produced, which points at neither
remedy. The parallel path already tracks this; apply the same idea here
and name the cause.
The runner stacked two independent parallelism knobs: OMP_TEST_CONCURRENCY
spawned N `bun test` processes while each process ran its own --parallel=M
test files, so the workspace bucket put 4 x 8 = 32 files in flight on a
4-core runner. bun's per-test timeout is wall-clock, so CPU-starved suites
crossed the 5s default and failed at random - mnemopi's sqlite/CLI files
tripped a different pair every run.
- TestCommand now carries a `parallel` request instead of baking the flag
into argv; the dispatcher resolves it against one shared budget
(availableParallelism x 2, split by the live pool width) so total
in-flight files track the machine. A chunk that runs alone still gets
its full requested width, leaving the sequential CI path unchanged.
- Raised the per-test timeout to 30s (OMP_TEST_TIMEOUT to override).
Suites here build real SQLite schemas and spawn CLIs, already running
1-4s per case on a quiet runner; the 10-minute chunk watchdog stays the
backstop for an actual hang.
- --dry-run now resolves the same budget, so it prints the argv a real run
would use, and the pool header reports cores plus the granted width.
- Dropped the dead `{ smol: true }` argument: workspaceTestCommand never
accepted it, and this file's own findings say a smaller heap makes bun
1.3.14's GC crash more often, so it should not be wired up.
The stated reason for the exclusion no longer holds: no test in
packages/mnemopi loads the fastembed model. The suites that touch embeddings
inject a deterministic in-process provider via setEmbeddingProviderForTests,
and the model-cache suite mocks globalThis.fetch, so nothing downloads or
reads a ~270MB model. On a clean checkout with no model cached, all 448 tests
across 73 files pass on linux-x64 in 11.7s.
packages/mnemopi moves into fastWorkspacePackages, which the workspace bucket
runs and which local-ts already covered through localOnlyWorkspacePackages, so
the local full run is unchanged.
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
Scoped test-runtime detection to explicit runner markers and Bun test entrypoints, so application NODE_ENV/BUN_ENV values no longer make ProcessTerminal headless.
Added subprocess regression coverage and propagated the private marker to test children.
Fixes#7261
- Replaced the napi-cli/cargo-zigbuild/cargo-xwin/sccache build path with
Bazel: rules_rust + crate_universe over Cargo.lock, hermetic zig cc
toolchains (linux-gnu pinned to glibc 2.17, linux-musl), host Xcode for
darwin, and a repo-local hermetic clang-cl + llvm-ml + xwin toolchain for
windows-msvc (bazel/toolchains/msvc).
- All eight shipped addons build as //:natives-<target> via the release
transition in bazel/defs.bzl (opt, thin LTO, cgu=16, stripped, canonical
.node naming); scripts/bazel-natives.ts is the single driver for local
dev and CI.
- Rust validation moved to bazel test + clippy aspects (strict workspace
policy for opted-in crates, default lints elsewhere, mirroring cargo
semantics) and the rustfmt aspect; cargo stays as the dev-iteration
surface, with brush-core/brush-builtins promoted to workspace members
and excluded from cargo dev tasks to keep their historical scope.
- CI caches through an in-cluster bazel-remote action cache (TLS + basic
auth, cluster-internal only); GitHub-hosted runners never touch the
infrastructure and use an actions/cache-backed disk cache instead.
- Deleted the hand-rolled caching machinery: ci-target-cache,
ci-native-artifact-cache, ci-build-native, native-source-hash,
find-native-artifacts, restore-linux-native, native-prewarm workflow,
ensure-* toolchain actions, and all sccache/Swatinem wiring.
- Warm native rebuilds drop from ~20 minutes to seconds; a cold client
with a warm remote cache rebuilds the linux x64 pair in ~2.5 minutes.
- Enhanced CI workflows and GitHub actions to support native artifact caching and parallel builds.
- Added composite actions and scripts for computing sources, finding artifacts, and managing caches.
- Updated infrastructure documentation and runner deployment scripts with revised resource limits.
- Detect bun crash exits (signals 132-139) and retry up to MAX_CHUNK_ATTEMPTS in a fresh heap.
- Add `retries` field to ChunkOutcome and surface crash retries in progress output.
- Fix test assertion expecting `stopReason: "aborted"` to use `toolUse` (aborted turns are dropped from context).
- Initialized in-memory settings in the repro issue test by resetting settings before theme setup.
- Reduced the UI/TUI CI test bucket chunk size from 10 to 5 to avoid cumulative Bun GC heap aborts.
- Built the generic Windows release binary with Bun's baseline x64 runtime so older Windows 10 CPUs do not hit the AVX2-only modern executable.
- Forced pi-natives release builds to link PCRE2 statically so macOS installs do not depend on Homebrew's libpcre2 dylib.
- Added release dry-run and native-build regression coverage for the portable artifact contracts.
Fixes#5172
- Added scripts/fix-dts-extensions.test.ts to the test:scripts npm script.
- Included the .d.ts extension rewrite test in the CI workspace test suite.
- Updated repoScriptTests list in scripts/ci-test-ts.ts to ensure consistent tracking.
- Extracted failure formatting logic into a reusable `formatChunkFailure` function.
- Updated quiet mode to display detailed failure output immediately when a test chunk fails.
- Adjusted final failure summary to improve output readability and conditional flow.
- Added a per-chunk watchdog that terminates wedged test processes with SIGKILL after a configurable timeout.
- Implemented incremental stream draining to capture test output even when a process is forcibly killed.
- Stabilized test execution by forcing single-threaded GC markers (`BUN_JSC_numberOfGCMarkers=1`) to prevent segmentation faults during heap-heavy operations.
- Replaced the non-existent `ci-test-ts.test.ts` with the active `link-omp.test.ts` in `package.json` and `ci-test-ts.ts`.
- Documented that the invalid test path was previously ignored silently by the bun test runner.
- Added `scripts/**/*.ts` to the linter and formatter inclusion list in `biome.json`.
- Replaced manual array index searches with standard `indexOf` operations.
- Cleaned up unused variables, obsolete helpers, and command-line flags.
- Applied consistent formatting and code style across tooling scripts.
- Refactored `formatProgressLine` to display passes and failures in the visual style of `bun test` including elapsed time.
- Implemented `extractFailingTests` to parse individual test failures and their raw output details out of captured stdout/stderr.
- Added a `formatSummaryFooter` to output a clean concluding tally of passed/failed chunks and the total wall time.
- Added `local` and `local-ts` modes to `ci-test-ts.ts` to coordinate full local runs of TypeScript and Rust packages in a single progress stream.
- Implemented a quiet output formatter that displays a one-line progress entry and defers verbose output until failure.
- Included previously skipped workspace packages `packages/mnemopi` and `python/robomp/web` in the local testing scope.
- Configured colorized ANSI formatting for progress logs when running in interactive TTY terminals.
- Bound repo-level scripts and rust task orchestration command under the consolidated runner.
- Added a worker pool mechanism to run independent test chunks concurrently in local environments.
- Introduced `OMP_TEST_CONCURRENCY` to allow manual control over parallel worker counts.
- Retained sequential, fail-fast execution for CI environments to ensure stability within memory-constrained runner jobs.
- Consolidated environment variable scrubbing into a shared helper function.
- Updated the `ai` package E2E tests to retrieve environment variables via `e2eApiKey` to ensure they only execute when E2E testing is enabled.
- Refactored the `scripts/ci-test-ts` runner to conditionally apply the `--only-failures` flag based on provided arguments.
- Removed configurable tab width support and the `display.tabWidth` setting across all packages.
- Deleted obsolete utility functions `getIndentation`, `getIndentationNoescape`, and `setDefaultTabWidth`.
- Standardized tab expansion logic to use a fixed `DEFAULT_TAB_WIDTH` globally.
- Cleaned up related configuration schemas, test suites, and internal API signatures to remove path-dependency.
- Added `chunkSize` configuration to coding-agent test buckets to prevent OOM errors in CI.
- Updated `codingAgentTestCommands` to generate multiple test processes when a bucket exceeds its chunk limit.
- Refactored `commandsForMode` to support flattened arrays of chunked test commands.
- Removed the `--only-failures` flag from workspace test commands so CI runs complete test passes.
- Raised native test mode parallelism from 1 to 4 for native and integration package execution.
- Removed the mnemopi Bun preload file and added explicit `./setup` imports in the changed test suites.
- Deleted the `RUN_EMBEDDINGS` flag path and shifted `MNEMOPI_NO_EMBEDDINGS` handling into per-suite before/after hooks.
- Excluded `packages/mnemopi` from the fast parallel CI test bucket with a note about missing fastembed models.
- Updated workspace-mode CI to run `workspaceTestCommand` with 8 workers instead of 4.
- Raised the documented ARC runner CPU limit from 8 to 16 in caching docs.
- Added explicit environment variable blacklists for credential and cache-related prefixes and names used by CI test runs.
- Added a helper that also strips provider API, OAuth, and bearer token variables by pattern.
- Updated test command spawning to pass a filtered environment, making suites run without inherited secret credentials.
- Added a setup-system-deps action with preloaded-runner guards and apt fallbacks.
- Updated CI workflows to download Linux x64 native artifacts and gate on native job success.
- Renamed coding-agent fast mode to singleton in scripts and test partitioning logic.
- Added settings test-state begin/restore helpers with recursive cleanup in affected tests.
- Added a mode-based `ci-test-ts.ts` runner with `--dry-run` support.
- Partitioned coding-agent tests into fast/ui/runtime/native/heavy buckets and separated workspace/native runs.
- Added coding-agent bucket modes that fail CI when a target bucket has no matching tests.
- Updated CI scripts/workflow to run the new TS buckets, use `omp-kata`, and gate releases on them.