Skip to content

[FEATURE]: Liveness policy when a run's last claim detaches #784

Description

@justintime4tea

Summary

Losing the session's last claimant while a run is live cancels the run. With both seams on the runtime (#782, #783) that is the only behavior the runtime has, and it is the right one for a chat endpoint, where a run nobody reads is a run spending provider turns for no one. It is the wrong answer for a run meant to outlive its request: a task started headless, or a chat whose client is reconnecting.

Make the reaction a policy the run is started with. The vocabulary — LivenessPolicy { Cancel, Continue, Park } and the ClaimsExhausted / LivenessDecided lifecycle events — is already in aura_events::run (#778); this gives it behavior.

Goals

Data structures

LivenessPolicy is as implemented in PR #730. Liveness is proposed beside it in aura_events::run, since it is serialized on the wire (#785) and in RunFacts (#780) and reuses that module's duration_ms:

// aura_events::run — as implemented (PR #730)
#[non_exhaustive] pub enum LivenessPolicy { Cancel, Continue, Park }

// aura_events::run — proposed
#[derive(Clone, Copy, Debug, Serialize, Deserialize)]
pub struct Liveness {
    pub policy: LivenessPolicy,
    /// How long after the last claim detaches before `policy` acts. A claim
    /// arriving within it cancels the timer. Zero applies the policy at once.
    /// On the wire as `grace_ms`, like every other duration in the schema.
    #[serde(rename = "grace_ms", with = "duration_ms")]
    pub grace: Duration,
}

impl Default for Liveness {
    fn default() -> Self { Self { policy: LivenessPolicy::Cancel, grace: Duration::ZERO } }
}

Config, so a deployment can change a seam's default without the seam knowing. Seconds here, milliseconds on the wire — the codebase's existing split (timeout_secs in TOML, *_ms in events):

[runtime.liveness]
chat = { policy = "cancel", grace_secs = 0 }       # /v1/chat/completions
headless = { policy = "continue", grace_secs = 0 } # POST /v1/runs (#785)

[runtime.park]
resume_on_start = false   # resume this instance's parked sessions on boot

There is no a2a entry. An A2A task holds its run's claim for the task's whole life (#783), so ClaimsExhausted fires only after the run is already terminal, and a policy would never act.

Additional Context

Park is the only policy that leaves durable state behind, and it reuses the checkpoint and continuation surfaces in crates/aura/src/orchestration/park rather than growing a second set. A parked run holds nothing in memory: the runtime's record of it is AgentState::Parked on the session (#780), and resume-on-start reads only the checkpoint and the session stream.

Continue keeps the run alive, but what a later observer can see of it is bounded by the retention window and the journal store's bound (#779, #780). Until a durable journal exists (#210 / #325), Continue means "finish the work and keep it attachable for a while", not "keep it forever".

A disconnect under the default policy writes three events — ObserverDetached, ClaimsExhausted, LivenessDecided — and with grace zero the last two are always adjacent. They stay separate because ClaimsExhausted marks the start of a non-zero window and LivenessDecided its end; a consumer that only cares about the outcome reads the terminal event.

Searched Issues

  • No similar issues found

Code of Conduct

  • I agree to follow this project's Code of Conduct

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions