Skip to content

Handle typed poll capacity backpressure without warning or duplicate transport retries #75

Description

@rmcdaniel

Problem

Server can return a typed HTTP 429 long_poll_capacity_exhausted when its bounded wait pool is full. The Python SDK's generic transport policy currently retries that same poll as a generic 429, then the worker counts/logs it as a poll error and sleeps a fixed one second. This creates extra admission traffic and an alarming warning for an expected, retryable capacity response. The response already supplies retry_after_seconds and Retry-After.

Accepted fix

  • Recognize only the complete typed poll-capacity response on the workflow, activity, and query poll routes. Return it to the worker without consuming generic transport retries.
  • Release the worker's poll reservation, record a distinct backpressure outcome, and wait for the bounded advertised delay before a fresh poll. Do not log an error warning for this expected response.
  • Keep malformed, unrelated, storage-admission, and true HTTP/transport failures on their existing error paths. Avoid printing large server bodies.
  • Test all three poll loops, advertised delay, malformed 429 behavior, and generic non-poll 429 retry behavior.

This changes SDK handling only; it does not raise Server capacity or alter poll/task ownership semantics. Recheck the behavior against a published Server artifact before describing it as Cloud qualification.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions