Per #3740 review: many short strings (e.g. a tool result whose content array holds thousands of small text blocks) could sum past MAX_REPLICATED_PAYLOAD_BYTES without any individual field crossing the per-string floor, so the helper exited the truncation loop and shipped an oversized frame — the relay close/reconnect loop the helper was meant to prevent.
Replace the string-only truncation pass with a single walker that head-truncates strings AND head-clips arrays in one descent, driven by a concrete SHRINK_PASSES schedule that tightens both axes together. The final pass clamps every string to 64 B and every array to one element, so any payload converges. Add direct unit tests for shrinkForReplication covering: identity for small values, single-giant-string clamp, many-short-strings array clamp (no field above floor), and discriminator preservation on a fully-shrunk payload.
CollabHost shipped the first entry of every snapshot-chunk batch unconditionally, and broadcast live entry/event frames verbatim, so a single multi-megabyte tool result (read/bash/search) overflowed the relay's per-frame maxPayloadLength. The relay closed the host's WebSocket with 1006 ("Received too big message"), CollabSocket treated 1006 as non-fatal and reconnected, the next guest hello triggered the same oversized send, and the host status line cycled "Collab relay connection lost, reconnecting…" indefinitely.
Add shrinkForReplication: any host->guest payload whose JSON exceeds MAX_REPLICATED_PAYLOAD_BYTES (1 MB) is deep-cloned with long strings head-truncated and an "[…N chars elided for collab session]" marker; otherwise the original reference passes through. Apply it to snapshot chunk entries, live entry broadcasts, and live event broadcasts (including large tool_execution_end results). Regression test stands up a Bun.serve relay with 8 MB maxPayloadLength and a snapshot containing a 5 MB entry; asserts the host stays connected, the snapshot train finalizes, and the guest sees the entry with the elision marker.
Fixes#3739