Commit Graph

34 Commits

Author SHA1 Message Date
Muhammad Mustaqeem 21a4d0b2fa revert(ci): keep the singleton bucket unchunked
CI evidence overturns the chunking half of this branch. Splitting the
singleton bucket into 8 chunks made the job run 11m27s and go red, where
the unchunked bucket passes in ~1m31s on a concurrent PR. The duration
matches the 600s chunk watchdog, i.e. a chunk wedged rather than merely
running slowly -- the "these suites co-locate in one process" comment was
describing a real coupling, so restore it and the original plan entry.

The failure-attribution change is kept and is independently justified:
this run is precisely the case where a bare "exit code 137" cannot
distinguish the watchdog from the OOM killer.
2026-08-13 23:42:37 +05:00
Muhammad Mustaqeem 6b78c4e112 fix(ci): chunk the singleton bucket and name the cause of a SIGKILLed chunk
The singleton/global-state bucket was the only bucket left unchunked, and
it is the one that got OOM-killed: 79 files in a single `bun test`
process exited 137. Chunking exists precisely to hold peak RSS under the
runner's ceiling, so the bucket that opts out is the bucket that hits it.

The stated reason for leaving it whole does not hold. That bucket is
selected because its suites mutate process-wide state, and what keeps
them from colliding is sequencing, which `parallel: 1` already
guarantees. Splitting them across processes strictly increases isolation
rather than reducing it. Use 10 files per chunk, the width the 650-file
native bucket already sustains on the same runner.

Separately, the sequential path could not say why a chunk was SIGKILLed.
It arms a watchdog but never records when the watchdog fires, so an OOM
kill and a watchdog kill both surfaced as a bare "failed with exit code
137" - the exact message this failure produced, which points at neither
remedy. The parallel path already tracks this; apply the same idea here
and name the cause.
2026-08-13 23:22:53 +05:00
can1357 08819b279c ci(test): budgeted bun test parallelism across the chunk pool
The runner stacked two independent parallelism knobs: OMP_TEST_CONCURRENCY
spawned N `bun test` processes while each process ran its own --parallel=M
test files, so the workspace bucket put 4 x 8 = 32 files in flight on a
4-core runner. bun's per-test timeout is wall-clock, so CPU-starved suites
crossed the 5s default and failed at random - mnemopi's sqlite/CLI files
tripped a different pair every run.

- TestCommand now carries a `parallel` request instead of baking the flag
  into argv; the dispatcher resolves it against one shared budget
  (availableParallelism x 2, split by the live pool width) so total
  in-flight files track the machine. A chunk that runs alone still gets
  its full requested width, leaving the sequential CI path unchanged.
- Raised the per-test timeout to 30s (OMP_TEST_TIMEOUT to override).
  Suites here build real SQLite schemas and spawn CLIs, already running
  1-4s per case on a quiet runner; the 10-minute chunk watchdog stays the
  backstop for an actual hang.
- --dry-run now resolves the same budget, so it prints the argv a real run
  would use, and the pool header reports cores plus the granted width.
- Dropped the dead `{ smol: true }` argument: workspaceTestCommand never
  accepted it, and this file's own findings say a smaller heap makes bun
  1.3.14's GC crash more often, so it should not be wired up.
2026-08-08 07:06:07 +02:00
can1357 afe926329d docs(ci): update local TypeScript runner comment 2026-08-05 22:16:29 +02:00
Cyrus 8bbe1ed01d ci: ran the mnemopi suite in the workspace bucket
The stated reason for the exclusion no longer holds: no test in
packages/mnemopi loads the fastembed model. The suites that touch embeddings
inject a deterministic in-process provider via setEmbeddingProviderForTests,
and the model-cache suite mocks globalThis.fetch, so nothing downloads or
reads a ~270MB model. On a clean checkout with no model cached, all 448 tests
across 73 files pass on linux-x64 in 11.7s.

packages/mnemopi moves into fastWorkspacePackages, which the workspace bucket
runs and which local-ts already covered through localOnlyWorkspacePackages, so
the local full run is unchanged.
2026-08-05 21:44:00 +08:00
can1357 bc39ffa265 feat: introduced omptype validation package and migrated workspace dependencies
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
2026-08-03 21:56:48 +02:00
can1357 fc04aa6fa7 refactor(coding-agent): deferred startup parsing and schema compilation
- Wrapped security contract schemas in a lazy initializer with jitless scopes to eliminate startup JIT compilation tax.
- Enabled jitless scope configuration in auth-broker wire schemas to skip definition-time codegen.
- Deferred startup changelog parsing to overlap with interactive session creation.
2026-08-03 16:31:27 +02:00
can1357 8cd6e3e964 chore: delete dumb tests 2026-08-03 15:45:44 +02:00
roboomp 9f9c9758a1 fix(tui): prevented test env from suppressing interactive launch
Scoped test-runtime detection to explicit runner markers and Bun test entrypoints, so application NODE_ENV/BUN_ENV values no longer make ProcessTerminal headless.

Added subprocess regression coverage and propagated the private marker to test children.

Fixes #7261
2026-08-01 11:39:57 +00:00
can1357 d16a251777 chore: reorg tests 2026-07-27 16:43:53 +02:00
can1357 8facd237d5 feat(build): migrated native pipeline to bazel with remote caching
- Replaced the napi-cli/cargo-zigbuild/cargo-xwin/sccache build path with
  Bazel: rules_rust + crate_universe over Cargo.lock, hermetic zig cc
  toolchains (linux-gnu pinned to glibc 2.17, linux-musl), host Xcode for
  darwin, and a repo-local hermetic clang-cl + llvm-ml + xwin toolchain for
  windows-msvc (bazel/toolchains/msvc).
- All eight shipped addons build as //:natives-<target> via the release
  transition in bazel/defs.bzl (opt, thin LTO, cgu=16, stripped, canonical
  .node naming); scripts/bazel-natives.ts is the single driver for local
  dev and CI.
- Rust validation moved to bazel test + clippy aspects (strict workspace
  policy for opted-in crates, default lints elsewhere, mirroring cargo
  semantics) and the rustfmt aspect; cargo stays as the dev-iteration
  surface, with brush-core/brush-builtins promoted to workspace members
  and excluded from cargo dev tasks to keep their historical scope.
- CI caches through an in-cluster bazel-remote action cache (TLS + basic
  auth, cluster-internal only); GitHub-hosted runners never touch the
  infrastructure and use an actions/cache-backed disk cache instead.
- Deleted the hand-rolled caching machinery: ci-target-cache,
  ci-native-artifact-cache, ci-build-native, native-source-hash,
  find-native-artifacts, restore-linux-native, native-prewarm workflow,
  ensure-* toolchain actions, and all sccache/Swatinem wiring.
- Warm native rebuilds drop from ~20 minutes to seconds; a cold client
  with a warm remote cache rebuilds the linux x64 pair in ~2.5 minutes.
2026-07-27 12:22:19 +02:00
can1357 5f988a8270 ci: configured native artifact caching and parallel execution in ci workflows
- Enhanced CI workflows and GitHub actions to support native artifact caching and parallel builds.
- Added composite actions and scripts for computing sources, finding artifacts, and managing caches.
- Updated infrastructure documentation and runner deployment scripts with revised resource limits.
2026-07-27 07:53:28 +02:00
can1357 0babb55fbd ci: retried test chunks when bun crashes due to GC bugs
- Detect bun crash exits (signals 132-139) and retry up to MAX_CHUNK_ATTEMPTS in a fresh heap.
- Add `retries` field to ChunkOutcome and surface crash retries in progress output.
- Fix test assertion expecting `stopReason: "aborted"` to use `toolUse` (aborted turns are dropped from context).
2026-07-23 14:40:28 +02:00
Victor Araújo ad36bbeae3 test(ci): run release publish coverage 2026-07-22 00:58:29 -03:00
can1357 c69c048361 ci: stabilized test runs by resetting settings and shrinking UI/TUI chunks
- Initialized in-memory settings in the repro issue test by resetting settings before theme setup.
- Reduced the UI/TUI CI test bucket chunk size from 10 to 5 to avoid cumulative Bun GC heap aborts.
2026-07-13 18:51:34 +02:00
roboomp bc7a143c1e fix(setup): used portable native install artifacts
- Built the generic Windows release binary with Bun's baseline x64 runtime so older Windows 10 CPUs do not hit the AVX2-only modern executable.

- Forced pi-natives release builds to link PCRE2 statically so macOS installs do not depend on Homebrew's libpcre2 dylib.

- Added release dry-run and native-build regression coverage for the portable artifact contracts.

Fixes #5172
2026-07-11 09:47:37 +00:00
can1357 336e1edf7a chore: organized test scripts in CI workspace and package configuration
- Added scripts/fix-dts-extensions.test.ts to the test:scripts npm script.
- Included the .d.ts extension rewrite test in the CI workspace test suite.
- Updated repoScriptTests list in scripts/ci-test-ts.ts to ensure consistent tracking.
2026-07-05 16:01:51 +02:00
can1357 5c171b2860 chore: refactored failure reporting and update quiet mode output
- Extracted failure formatting logic into a reusable `formatChunkFailure` function.
- Updated quiet mode to display detailed failure output immediately when a test chunk fails.
- Adjusted final failure summary to improve output readability and conditional flow.
2026-07-03 00:47:56 +02:00
can1357 83176e159f chore: implemented watchdog and stabilize GC for test execution
- Added a per-chunk watchdog that terminates wedged test processes with SIGKILL after a configurable timeout.
- Implemented incremental stream draining to capture test output even when a process is forcibly killed.
- Stabilized test execution by forcing single-threaded GC markers (`BUN_JSC_numberOfGCMarkers=1`) to prevent segmentation faults during heap-heavy operations.
2026-07-03 00:46:46 +02:00
can1357 388534802c chore(repo): corrected invalid test script paths in ci configuration
- Replaced the non-existent `ci-test-ts.test.ts` with the active `link-omp.test.ts` in `package.json` and `ci-test-ts.ts`.
- Documented that the invalid test path was previously ignored silently by the bun test runner.
2026-07-02 02:40:07 +02:00
can1357 8687899728 ci: disabled concurrent garbage collection in test environment
- Set the `BUN_JSC_useConcurrentGC` environment variable to `"0"` in the child process environment built for test suites.
2026-07-01 00:07:03 +02:00
can1357 cbb471b666 chore: standardized code quality and formatting for tooling scripts
- Added `scripts/**/*.ts` to the linter and formatter inclusion list in `biome.json`.
- Replaced manual array index searches with standard `indexOf` operations.
- Cleaned up unused variables, obsolete helpers, and command-line flags.
- Applied consistent formatting and code style across tooling scripts.
2026-06-30 23:51:16 +02:00
can1357 6d0569788c ci: formatted test runner progress, summary, and failure parsing to match bun style
- Refactored `formatProgressLine` to display passes and failures in the visual style of `bun test` including elapsed time.
- Implemented `extractFailingTests` to parse individual test failures and their raw output details out of captured stdout/stderr.
- Added a `formatSummaryFooter` to output a clean concluding tally of passed/failed chunks and the total wall time.
2026-06-30 21:27:41 +02:00
can1357 bd77bd507c feat: introduced a unified local runner for rust and typescript tests
- Added `local` and `local-ts` modes to `ci-test-ts.ts` to coordinate full local runs of TypeScript and Rust packages in a single progress stream.
- Implemented a quiet output formatter that displays a one-line progress entry and defers verbose output until failure.
- Included previously skipped workspace packages `packages/mnemopi` and `python/robomp/web` in the local testing scope.
- Configured colorized ANSI formatting for progress logs when running in interactive TTY terminals.
- Bound repo-level scripts and rust task orchestration command under the consolidated runner.
2026-06-30 20:39:32 +02:00
can1357 469046fbcb chore: implemented parallel execution for test scripts
- Added a worker pool mechanism to run independent test chunks concurrently in local environments.
- Introduced `OMP_TEST_CONCURRENCY` to allow manual control over parallel worker counts.
- Retained sequential, fail-fast execution for CI environments to ensure stability within memory-constrained runner jobs.
- Consolidated environment variable scrubbing into a shared helper function.
2026-06-28 17:23:31 +02:00
can1357 af01c63e3b test(ai): updated credential access in E2E tests and CI test runner
- Updated the `ai` package E2E tests to retrieve environment variables via `e2eApiKey` to ensure they only execute when E2E testing is enabled.
- Refactored the `scripts/ci-test-ts` runner to conditionally apply the `--only-failures` flag based on provided arguments.
2026-06-25 21:50:55 +02:00
can1357 3998d95088 feat: standardized tab expansion to fixed width
- Removed configurable tab width support and the `display.tabWidth` setting across all packages.
- Deleted obsolete utility functions `getIndentation`, `getIndentationNoescape`, and `setDefaultTabWidth`.
- Standardized tab expansion logic to use a fixed `DEFAULT_TAB_WIDTH` globally.
- Cleaned up related configuration schemas, test suites, and internal API signatures to remove path-dependency.
2026-06-19 04:48:41 +02:00
can1357 5875b20866 chore(scripts): implemented chunked test execution for coding-agent
- Added `chunkSize` configuration to coding-agent test buckets to prevent OOM errors in CI.
- Updated `codingAgentTestCommands` to generate multiple test processes when a bucket exceeds its chunk limit.
- Refactored `commandsForMode` to support flattened arrays of chunked test commands.
2026-06-18 03:49:20 +02:00
can1357 ee6e036073 chore(scripts): raised native test parallelism and removed failure-only test flag
- Removed the `--only-failures` flag from workspace test commands so CI runs complete test passes.
- Raised native test mode parallelism from 1 to 4 for native and integration package execution.
2026-06-15 08:27:59 +02:00
can1357 4952b28a8e test(mnemopi): reworked mnemopi test embedding setup and CI bucketing
- Removed the mnemopi Bun preload file and added explicit `./setup` imports in the changed test suites.
- Deleted the `RUN_EMBEDDINGS` flag path and shifted `MNEMOPI_NO_EMBEDDINGS` handling into per-suite before/after hooks.
- Excluded `packages/mnemopi` from the fast parallel CI test bucket with a note about missing fastembed models.
2026-06-15 02:45:03 +02:00
can1357 12950af096 perf(scripts): raised workspace test workers and ARC CPU limits
- Updated workspace-mode CI to run `workspaceTestCommand` with 8 workers instead of 4.
- Raised the documented ARC runner CPU limit from 8 to 16 in caching docs.
2026-06-15 02:44:16 +02:00
can1357 173e9829ab chore(scripts): scrubbed cloud credentials from CI test command environment
- Added explicit environment variable blacklists for credential and cache-related prefixes and names used by CI test runs.
- Added a helper that also strips provider API, OAuth, and bearer token variables by pattern.
- Updated test command spawning to pass a filtered environment, making suites run without inherited secret credentials.
2026-06-15 00:14:45 +02:00
can1357 b8e4da23d0 ci(ci): refactored CI setup and test-state isolation for coding-agent workflows
- Added a setup-system-deps action with preloaded-runner guards and apt fallbacks.
- Updated CI workflows to download Linux x64 native artifacts and gate on native job success.
- Renamed coding-agent fast mode to singleton in scripts and test partitioning logic.
- Added settings test-state begin/restore helpers with recursive cleanup in affected tests.
2026-06-14 22:37:05 +02:00
can1357 a734c29234 ci(scripts): reworked CI test execution with mode-based TypeScript buckets
- Added a mode-based `ci-test-ts.ts` runner with `--dry-run` support.
- Partitioned coding-agent tests into fast/ui/runtime/native/heavy buckets and separated workspace/native runs.
- Added coding-agent bucket modes that fail CI when a target bucket has no matching tests.
- Updated CI scripts/workflow to run the new TS buckets, use `omp-kata`, and gate releases on them.
2026-06-14 21:55:55 +02:00