Skip to content

Reuse async Kubernetes API client across trigger polls - #71349

Open
xvega wants to merge 6 commits into
apache:mainfrom
xvega:fix-async-k8s-client-reuse
Open

xvega wants to merge 6 commits into
apache:mainfrom
xvega:fix-async-k8s-client-reuse

Conversation

@xvega

@xvega xvega commented Aug 9, 2026 •

Copy link
Copy Markdown
Contributor

I've been chasing triggerer issues on our production deployment: sustained triggers.blocked_main_thread bursts whenever a wave of deferrable KubernetesPodOperator/GKEStartPodOperator tasks started or finished, with the async thread blocked 2–10s at a time. #69661 helped the log-parsing path, but the metric barely moved. After reading the source and profiling the TriggerRunner with py-spy during blocking windows, the loop thread showed up inside ssl.create_default_context() reached from AsyncKubernetesHook.get_conn(): every API method builds a fresh ApiClient per call, which parses the CA bundle synchronously on the event loop and opens a new aiohttp session, a full TCP+TLS handshake per status poll (default 2s per trigger) with zero reuse. The GKE hook additionally fetched a new OAuth token and wrote the cluster CA to a temp file each call.

The change itself is simple: instead of building a new API client for every call, the hook now builds one and keeps it for its whole lifetime. For in-cluster and static-auth kubeconfigs Airflow loads the configuration once per hook, and anything that does rotate (the in-cluster service-account token) is read from the configuration per request rather than baked into the client, so one client can serve the hook's lifetime. Exec-based auth (EKS/GKE kubeconfigs) is the exception and keeps the old per-call behavior on purpose, reloading the config on every call is exactly what refreshes its short-lived credentials, so those clients can't be reused. To pair every cached client with an owner, hooks get an idempotent close() and the triggers call it from cleanup(), so the client lives exactly as long as its trigger. The GKE hook gets the same treatment plus token caching: it re-sets the bearer header every time the connection is used, so tokens refresh on schedule instead of being fetched from Google on every call (and since all that was left of its _load_config() was building the client, it's now called _build_client()).

A few things follow from the caching that are worth knowing. If you instantiate these hooks yourself or subclass the triggers and override cleanup(), call hook.close() when you're done, otherwise the client is only cleaned up at garbage collection and aiohttp logs an "unclosed session" warning; calling the hook again after close() just builds a fresh client. mTLS client certs from static kubeconfigs are now read once per hook rather than on every call, which I think is fine for hooks that live only as long as a trigger (if cert rotation mid-trigger ever becomes a real problem, rebuilding the client on auth errors would be the fix). The pre-existing gap left by #65212 the default-kubeconfig path sets _config_loaded without exec-auth detection, is untouched. And one expectation to set: each new trigger still builds its one client on the event loop, so blocked_main_thread should drop dramatically but won't be exactly zero when a batch of triggers starts.

To make sure this was really the problem, I reproduced it locally on a minikube cluster: real pod triggers polling real pods, with Airflow's own block_watchdog running, counting how many SSL contexts and API clients got created. I ran the exact same load twice, once with the old per-call get_conn patched back in as a baseline, once with this change. The baseline created a new client (and SSL context) for every single API call; with this change each trigger creates exactly one and releases it on cleanup. Unit tests in both providers assert the same behavior.

related: #69661

Was generative AI tooling used to co-author this PR?

No

@boring-cyborg boring-cyborg Bot added area:providers provider:cncf-kubernetes Kubernetes (k8s) provider related issues provider:google Google (including GCP) related issues labels Aug 9, 2026
@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch 5 times, most recently from a3d8b31 to e6838f1 Compare August 12, 2026 11:51

@Miretpl Miretpl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you adjust your PR description to match our project guidelines - https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst?

Comment thread providers/google/src/airflow/providers/google/cloud/hooks/kubernetes_engine.py Outdated
@Miretpl

Miretpl commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

I would appreciate a review from a Kubernetes Google maintainer on that part of the code.

@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch from e6838f1 to 0bf66bf Compare August 17, 2026 07:02
@xvega

xvega commented Aug 17, 2026

Copy link
Copy Markdown
Contributor Author

Could you adjust your PR description to match our project guidelines - https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst?

Besides the AI disclosure block what's missing or not part of the guideline?

@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch from 0bf66bf to 4ec850e Compare August 17, 2026 08:38
@xvega
xvega requested a review from Miretpl August 17, 2026 08:41
@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch from 4ec850e to a708b4b Compare August 17, 2026 17:51
@Miretpl

Miretpl commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Besides the AI disclosure block what's missing or not part of the guideline?

Only this was missing here. Thanks for adding it. I've rerun the whole CI as there were a couple of unrelated failures. Let's see how it will be now.

@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch from a708b4b to 2f15574 Compare August 28, 2026 17:52
@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch from 2f15574 to 73adafa Compare September 8, 2026 12:37
@soupam05

soupam05 commented Sep 8, 2026

Copy link
Copy Markdown

Since the ApiClient is now reused for the lifetime of the trigger, can we explicitly document how credential refresh is handled independently of the client lifecycle? In particular, I'm wondering about long-running triggers where credentials may expire while the client is still alive.

@soupam05

soupam05 commented Sep 8, 2026

Copy link
Copy Markdown

What happens if client initialization succeeds but a later poll fails due to an authentication/configuration error? Should that error cause the cached client to be invalidated so that the next poll can rebuild it, or is the expectation that the trigger will terminate/retry? It would be good to define this now that the client is no longer recreated on every poll.

@xvega

xvega commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

Since the ApiClient is now reused for the lifetime of the trigger, can we explicitly document how credential refresh is handled independently of the client lifecycle? In particular, I'm wondering about long-running triggers where credentials may expire while the client is still alive.

Credential refresh is handled separately from the cached client:

  • In-cluster service-account tokens refresh through the Kubernetes configuration’s token-refresh callback.
  • When exec auth is detected, the hook still reloads the config and creates a client on each call.
  • For GKE, each get_conn() checks the cached token and updates the Authorization header when it refreshes.
    The caveat is static mTLS certificates: those are now loaded once per client, so certificate rotation during a running trigger won’t be picked up automatically. That tradeoff is noted in the PR description.

@xvega

xvega commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

What happens if client initialization succeeds but a later poll fails due to an authentication/configuration error? Should that error cause the cached client to be invalidated so that the next poll can rebuild it, or is the expectation that the trigger will terminate/retry? It would be good to define this now that the client is no longer recreated on every poll.

The cached client isn’t invalidated when a poll fails. Existing retries still handle transient errors, but authentication errors such as 401/403 cause the pod trigger to report an error, and cleanup closes the client. Whether the task retries afterward depends on its retry settings.
I’d keep that behavior for this PR unless we identify a case where client reuse requires additional recovery. Did you have a particular authentication setup or failure scenario in mind? That would help determine whether rebuilding the client would actually resolve it.

@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch from 73adafa to f64718a Compare September 9, 2026 08:03
@xvega
xvega requested a review from Fokko as a code owner September 9, 2026 10:52
@xvega

xvega commented Sep 13, 2026

Copy link
Copy Markdown
Contributor Author

Thanks, that clarifies the 401/403 handling. One follow-up: for a long-running trigger, if a refreshable credential expires and the Kubernetes API returns a 401, can we guarantee that the credential-refresh path is attempted before the 401 is treated as terminal? Ideally, an expected credential refresh should be transparent to the task rather than causing the trigger to fail and relying on a task-level retry.

We couldn’t guarantee that before. The regression test I added showed that retries could reuse the rejected token, so I also added a fix. The GKE client now refreshes credentials and retries once after a 401, using the same client. If the refresh or retry fails, the error still propagates.

@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch 2 times, most recently from d304e6f to 13c3bd3 Compare September 14, 2026 09:14

@aaron-y-chen aaron-y-chen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall, the Google provider parts LGTM. However, it now relies on _cached_kube_client and close(), which are added to AsyncKubernetesHook in this PR. Without marking this cross-provider dependency, the Google provider could still be installed with an older Kubernetes provider and fail with an AttributeError on the first GKE trigger poll.

Should we add # use next version at here?

"apache-airflow-providers-cncf-kubernetes>=10.1.0",

@xvega

xvega commented Sep 14, 2026

Copy link
Copy Markdown
Contributor Author

Overall, the Google provider parts LGTM. However, it now relies on _cached_kube_client and close(), which are added to AsyncKubernetesHook in this PR. Without marking this cross-provider dependency, the Google provider could still be installed with an older Kubernetes provider and fail with an AttributeError on the first GKE trigger poll.

Should we add # use next version at here?

"apache-airflow-providers-cncf-kubernetes>=10.1.0",

I added # use next version 👍

@aaron-y-chen aaron-y-chen left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Kubernetes Google part LGTM.

@xvega
xvega force-pushed the fix-async-k8s-client-reuse branch from bc15bfd to ca71fc5 Compare September 16, 2026 08:22

@Miretpl Miretpl left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking at the changes in the code, I'm reverting my approve to prevent accidental merge of it. I would recommend separate this PR to cncf and google provider - it should be easier to merge at least one of them.

try:
return await super().call_api(*args, **kwargs)
except async_client.ApiException as error:
if error.status != 401:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can receive a 401 code for more than just an expired token, so I don't think that this handling is correct.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How about following google-auth’s refresh-on-401 approach, capped at one attempt? I’d retry only if the token changes and propagate any refresh or retry failure. This preserves the refresh attempt requested by @soupam05

@xvega

xvega commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor Author

Looking at the changes in the code, I'm reverting my approve to prevent accidental merge of it. I would recommend separate this PR to cncf and google provider - it should be easier to merge at least one of them.

@Miretpl I opened #73399 for the cncf changes. The Google changes depend on them, so I’ll keep them here until that PR merges, then rebase this PR to leave only the Google changes.

@Miretpl

Miretpl commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

@Miretpl I opened #73399 for the cncf changes. The Google changes depend on them, so I’ll keep them here until that PR merges, then rebase this PR to leave only the Google changes.

Cool, thanks! So let's focus on the CNCF Kubernetes PR first.

@VladaZakharova VladaZakharova left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

get_conn() previously closed its ApiClient when the async context exited. For cached configurations it now requires a separate await hook.close(), so existing direct hook users and trigger subclasses can silently leak aiohttp sessions. Can client reuse be opt-in for the built-in triggers, preserving the existing context-manager lifecycle by default? If the lifecycle change is intentional, it should at least be documented as a provider behavioral change.

@xvega

xvega commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

get_conn() previously closed its ApiClient when the async context exited. For cached configurations it now requires a separate await hook.close(), so existing direct hook users and trigger subclasses can silently leak aiohttp sessions. Can client reuse be opt-in for the built-in triggers, preserving the existing context-manager lifecycle by default? If the lifecycle change is intentional, it should at least be documented as a provider behavioral change.

@VladaZakharova I'll address this as soon as #73399 is merged

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:providers provider:cncf-kubernetes Kubernetes (k8s) provider related issues provider:google Google (including GCP) related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants