What would you like to be added:
Extend the reflector in the packages/kube-client workspace so that transient Kubernetes watch failures are retried without restarting the complete LIST/WATCH cycle.
The reflector should:
- Retry watch establishment after an HTTP
429 Too Many Requests response.
- Retry an established watch after receiving a retryable
429 status event.
- Support bounded retries for Kubernetes
InternalError status responses.
- Reopen the watch using the last successfully observed resource version.
- Honor
Retry-After when it is available and otherwise use exponential backoff.
- Make retry and backoff waits abortable through the reflector's
AbortSignal.
- Continue retrying
429 responses with backoff until the reflector is stopped.
- Fall back to a full relist when the resource version has expired, the error is not retryable, or the bounded
InternalError retry window has been exceeded.
Retry handling must be limited to safe, idempotent LIST and WATCH requests.
This behavior is implemented by the upstream Kubernetes reflector and REST client:
Acceptance criteria:
Why is this needed:
The reflector currently restarts the complete listAndWatch cycle when watch establishment or processing fails with a transient Kubernetes API error such as HTTP 429 Too Many Requests. This triggers another full LIST request even though the watch could normally be reopened from the last observed resource version.
Full relists can be expensive for resources with many or large objects. Repeating them after transient watch failures increases Kubernetes API server load, network traffic, recovery time, and memory pressure in the dashboard backend. It also increases the likelihood of encountering the oversized LIST responses addressed by the related pagination work.
Retrying only the watch preserves the existing cache and follows the behavior of the official Kubernetes Go client.
What would you like to be added:
Extend the reflector in the
packages/kube-clientworkspace so that transient Kubernetes watch failures are retried without restarting the complete LIST/WATCH cycle.The reflector should:
429 Too Many Requestsresponse.429status event.InternalErrorstatus responses.Retry-Afterwhen it is available and otherwise use exponential backoff.AbortSignal.429responses with backoff until the reflector is stopped.InternalErrorretry window has been exceeded.Retry handling must be limited to safe, idempotent LIST and WATCH requests.
This behavior is implemented by the upstream Kubernetes reflector and REST client:
Acceptance criteria:
429while establishing a watch retries the watch without performing another LIST.429status event from an established watch retries the watch without performing another LIST.InternalErrorresponses can be retried for a bounded duration.Retry-Afteris honored when present.429responses continue to retry the watch with backoff until cancellation.InternalErrorretry window triggers the normal relist flow.- [ ] Tests cover watch establishment errors, watch status events, retry exhaustion, and cancellation.Why is this needed:
The reflector currently restarts the complete
listAndWatchcycle when watch establishment or processing fails with a transient Kubernetes API error such as HTTP429 Too Many Requests. This triggers another full LIST request even though the watch could normally be reopened from the last observed resource version.Full relists can be expensive for resources with many or large objects. Repeating them after transient watch failures increases Kubernetes API server load, network traffic, recovery time, and memory pressure in the dashboard backend. It also increases the likelihood of encountering the oversized LIST responses addressed by the related pagination work.
Retrying only the watch preserves the existing cache and follows the behavior of the official Kubernetes Go client.