feat: let a worker declare the process unable to serve - #36
Merged
Merged
Conversation
A sticky CUDA error kept /health passing while every later job failed, so FatalWorkerError now fails the request, reports 503, and exits for a restart.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A worker can now declare that its process can no longer serve by raising
FatalWorkerError, alongside the existingValidationError. The server answers that request with 500 and the full error chain. After that,/healthand/inferenceanswer 503, and the process sends itself SIGTERM so the orchestrator restarts it.Why: some failures leave a process unusable without making it look unhealthy. A typical case is a sticky CUDA error such as an illegal memory access: once it happens, every later CUDA call in that process fails the same way.
/healthkeeps passing, so jobs keep being routed to a process that can only fail them. The existing consecutive-error breaker eventually catches this, but only after several more failed jobs. On a GPU shared between processes, the stuck process can also block other clients until it exits.Ownership: this package owns the mechanism: health, refusing requests, exiting. Deciding which errors are fatal is left to the worker.
Behaviour
FatalWorkerErrorhas its own exception handler, so it doesn't feed the consecutive-500 counter. Ordinary failures andValidationErrorbehave as before.maestro-server: the 500 was delivered, the listener closed, and the process exited with code 143.Testing
tests/test_serve.py:/healthto 503, refuses later requests without calling the worker, and terminates the process exactly once.ValidationErrorleave the process serving.uv run pytest,ruffandtypass.Bumps the version to 5.2.0 for the new public API.