Skip to content

Retry transient Kubernetes watch failures without relisting #3273

Description

@grolu

What would you like to be added:

Extend the reflector in the packages/kube-client workspace so that transient Kubernetes watch failures are retried without restarting the complete LIST/WATCH cycle.

The reflector should:

  • Retry watch establishment after an HTTP 429 Too Many Requests response.
  • Retry an established watch after receiving a retryable 429 status event.
  • Support bounded retries for Kubernetes InternalError status responses.
  • Reopen the watch using the last successfully observed resource version.
  • Honor Retry-After when it is available and otherwise use exponential backoff.
  • Make retry and backoff waits abortable through the reflector's AbortSignal.
  • Continue retrying 429 responses with backoff until the reflector is stopped.
  • Fall back to a full relist when the resource version has expired, the error is not retryable, or the bounded InternalError retry window has been exceeded.

Retry handling must be limited to safe, idempotent LIST and WATCH requests.

This behavior is implemented by the upstream Kubernetes reflector and REST client:

Acceptance criteria:

  • A 429 while establishing a watch retries the watch without performing another LIST.
  • A 429 status event from an established watch retries the watch without performing another LIST.
  • Kubernetes InternalError responses can be retried for a bounded duration.
  • The watch is reopened using the last successfully observed resource version.
  • Retry-After is honored when present.
  • Retry waits use exponential backoff when no retry delay is provided.
  • Retry and backoff waits terminate immediately when the reflector is stopped.
  • Repeated 429 responses continue to retry the watch with backoff until cancellation.
  • Expired resource versions and non-retryable errors trigger the normal relist flow.
  • Exceeding the bounded InternalError retry window triggers the normal relist flow.- [ ] Tests cover watch establishment errors, watch status events, retry exhaustion, and cancellation.

Why is this needed:

The reflector currently restarts the complete listAndWatch cycle when watch establishment or processing fails with a transient Kubernetes API error such as HTTP 429 Too Many Requests. This triggers another full LIST request even though the watch could normally be reopened from the last observed resource version.

Full relists can be expensive for resources with many or large objects. Repeating them after transient watch failures increases Kubernetes API server load, network traffic, recovery time, and memory pressure in the dashboard backend. It also increases the likelihood of encountering the oversized LIST responses addressed by the related pagination work.

Retrying only the watch preserves the existing cache and follows the behavior of the official Kubernetes Go client.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions