Skip to content

fix(core): feed discovery failures into the critical-failure detector #629

Description

@jamezrin

Problem

Discovery failures never reach the critical-failure detector in production.

The scheduler tick records a failing critical-health observation (apiErrorObservation) only in its legacy in-tick refreshCampaigns branch. Since the committed discovery snapshot lane (#394), runTick always passes discovery to the tick, so that branch never ran in production. An incomplete discovery (complete: false) gives the tick a neutral observation. A platform whose discovery fails on every tick therefore never counts failing ticks and never raises the critical-failure prompt.

#599 removed the dead branch, together with the scheduler tests that were the only ones asserting "discovery throws → failing tick" (scheduler.test.ts, "scheduler critical health observations"). Those tests only exercised the legacy path, so they gave false confidence.

Expected

A failed discovery refresh for a platform feeds a failing observation (with the SafeFetchError breadcrumb: kind, status, detail) into observeCriticalHealth for that platform. A discarded refresh (settings changed mid-flight) does not.

Notes

  • The discovery lane (background/discovery.ts) knows the failure. The tick only sees complete: false, plus discarded for a thrown-away refresh. Either pass the failure through as a decide input or record the observation from the lane.
  • Keep the existing rule: an accrual precondition break in the same tick wins over the failure.

Found while doing #599.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    • Status
      Backlog

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions