Two failures surfaced scaling the DE sweep up. The replay sink prints a
per-second progress line with an explicit flush; under
subprocess.run(capture_output=True) those thousands of writes fill a fixed
OS pipe buffer that nothing drains until exit, so the long films blocked on
write to stderr and looked like hangs. Discard the child's stdout/stderr
(DEVNULL) — it was captured and thrown away anyway; the long films then
finish in seconds.
Separately, the per-film timeout is only a backstop for the rare,
intermittent ROCm GEMM wedge (a wedged replay hangs forever and must be
killed so the sweep continues), not a performance bound. It had been set
huge, which let a single flake stall the whole sweep; set it to a sane 180s
(overridable via REPLAY_TIMEOUT) — well above a healthy replay, short
enough to reap a wedge quickly.