fix(robomp): harden sandbox cleanup, git-probe, and worktree-add paths

A diff-scoped review of the event-loop-hang fixes surfaced gaps in the new
timeout/error-handling code and its tests. All at/above the medium floor,
each mutation-verified.

- remove_workspace: prune on any nonzero `git worktree remove` (not just a
  present checkout) and RAISE on a failed prune, so a killed remove that
  leaves a dangling pool registration is cleared or retried instead of
  recording success over stale metadata. Gate git ops on the pool being a
  real clone (ensure_clone mkdir's the dir before cloning, so a failed first
  clone leaves a non-git dir where `git worktree prune` would error), and
  only speculatively prune a missing checkout when ws_root still exists.
- _worktree_add: new helper wrapping the three worktree-add sites; on a
  failed add (incl. the new 124 timeout) it removes the partial checkout and
  prunes the pool before re-raising, so the event retry starts clean. Raises
  a failed prune chained from the add error.
- _reset_origin_url: a timed-out (124) `git remote get-url origin` probe is
  indeterminate; raise before fetch instead of silently skipping the rewrite,
  so a legacy credentialed origin cannot persist and be reused.
- tests: assert the subprocess timeout is passed in the _safe_run/_run
  timeout fakes; add a real-`git worktree prune` integration test; make the
  cancel-drain test deterministic (loop-turn pump, no wall-clock sleep) and
  cover the repeated-cancel branch; add regressions for the prune-failure,
  checkout-gone-on-entry, non-git-pool, and repeat-close cleanup paths.

Op: correct
Restores: spec:pool-cleanup-clears-or-retries-dangling-registration
Restores: spec:indeterminate-git-probes-raise-not-silently-proceed
This commit is contained in:
metaphorics
2026-07-02 13:04:19 +09:00
parent dfaa41f6b3
commit 16600e89ab
3 changed files with 394 additions and 37 deletions
+20 -9
View File
@@ -102,19 +102,30 @@ async def test_run_workspace_op_drains_thread_before_propagating_cancel():
await asyncio.to_thread(started.wait, 1.0)
assert started.is_set()
# Cancel the AWAITING coroutine while the thread is mid-flight.
task.cancel()
# Let the loop deliver the cancellation into the helper's drain loop.
await asyncio.sleep(0.05)
async def pump(turns: int = 20) -> None:
# Deterministically advance the loop without a wall-clock sleep: each
# sleep(0) drains the ready queue, so a DETACHING (pre-fix) helper would
# resolve `task` within these turns. A draining helper keeps it pending
# while the worker thread is still blocked on `proceed`.
for _ in range(turns):
await asyncio.sleep(0)
# Cancel the AWAITING coroutine while the thread is mid-flight, then a SECOND
# time while it is still blocked. The repeated cancel must land on the drain
# loop's re-`await` and be swallowed by its `continue` branch, NOT abandon
# the thread. The whole sequence runs under try/finally so any failed assert
# still releases the worker and cannot leak a blocked thread into later tests.
try:
# The thread must NOT have been abandoned: it is still blocked on
# `proceed`, so `finished` is not set and the task has not resolved yet.
task.cancel()
await pump()
assert not task.done(), "helper propagated the first cancel before the thread completed (thread abandoned)"
task.cancel()
await pump()
# The thread is still blocked on `proceed`, so it has not finished and
# the task has not resolved despite two cancels.
assert not finished.is_set(), "thread finished before we released it — impossible unless abandoned"
assert not task.done(), "helper propagated cancel before the thread completed (thread abandoned)"
assert not task.done(), "helper abandoned the thread after a repeated cancel"
finally:
# Always release the worker, even if an assert above fails, so a failed
# run cannot leave a blocked thread leaking into later tests.
proceed.set()
# The helper must now let the thread finish, THEN raise CancelledError.