[Parallel] Terminate timed out and unconnected workers - #8522
Conversation
37251f5 to
54ce386
Compare
|
Thank you 👍 This is pretty complex, how did you figure it out? |
Thanks for merging both! It started at work: we updated Rector on a large project and I wanted to understand the parallel run better. I read ParallelFileProcessor and ParallelProcess, compared them with PHPStan's Process.php, and found #7711, which describes exactly this hang. To reproduce it, I used a worker that never answers (just sleep) and aborted the run. Rector waited forever, on macOS and on Linux. On Linux there was one more catch: the worker runs in |
|
Cool 👍 Btw, how many loc your project has and how long does it take to perform one Rector run? |
|
Around 4,400 PHP files and 470k lines. A full run with a cold cache takes about 80 seconds on an M1 Max (10 cores, 6 parallel processes) with 700-850 MB per worker. With a warm cache and only a few changed files it takes 1-2 seconds. |
|
I see. Not bad |
When a worker gets stuck, e.g. in an endless loop, the timeout is reported, but Rector never ends:
quit()only closes the connection, and a stuck worker never reads it. A worker that is still starting when the run is aborted has no connection at all. The main process waits for both forever (see rectorphp/rector#7711).