You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[FEATURE]: Liveness policy when a run's last claim detaches #784
Losing the session's last claimant while a run is live cancels the run. With both seams on the runtime (#782, #783) that is the only behavior the runtime has, and it is the right one for a chat endpoint, where a run nobody reads is a run spending provider turns for no one. It is the wrong answer for a run meant to outlive its request: a task started headless, or a chat whose client is reconnecting.
Make the reaction a policy the run is started with. The vocabulary — LivenessPolicy { Cancel, Continue, Park } and the ClaimsExhausted / LivenessDecided lifecycle events — is already in aura_events::run (#778); this gives it behavior.
Goals
Three policies. Cancel: what the runtime does today, and the default. Continue: run to completion unclaimed, journal kept for a later attach. Park: checkpoint through the park machinery that exists for orchestration under [hitl.park] and end the run resumable.
Park is refused at start for an agent with no checkpoint support — single-agent runs today — rather than silently downgraded to Continue. An operator who asked for Park gets StartError::PolicyUnsupported, not a run that behaves otherwise.
A grace window before the policy acts, so a client reconnecting within it re-claims a run that never noticed it was gone. Zero by default, so [FEATURE]: Chat completions run on the runtime #782's behavior holds until a deployment opts in. The window's outcome is a lifecycle event either way.
The decision itself is a lifecycle event ([FEATURE]: A session's event envelope carries run lifecycle alongside agent activity #778): ClaimsExhausted, then LivenessDecided { policy }, so an observer arriving later reads why the run ended or kept going. A Cancel decision ends the run with Cancelled { reason: Unclaimed }, distinguishable from an operator's /cancel (External) and from shutdown (Shutdown).
Cancel-on-drop stays on the handle, for a caller that holds an AgentRun directly and never goes through the runtime.
A parked run comes back by policy, not only by request. With [runtime.park] resume_on_start = true, a runtime that boots and finds a session it holds — or, with [FEATURE]: Atomic Fence for VFS Claims #581, claims — whose last run is Parked resumes it through [FEATURE]: Endpoints to list, inspect, attach to, and control sessions and runs #785's /resume path, seeded from the checkpoint and the session stream, with no caller involved. That is the "reified by policy on runtime start" case. Off by default, because an instance restarting under load should not fan out every parked run at once.
Data structures
LivenessPolicy is as implemented in PR #730. Liveness is proposed beside it in aura_events::run, since it is serialized on the wire (#785) and in RunFacts (#780) and reuses that module's duration_ms:
// aura_events::run — as implemented (PR #730)#[non_exhaustive]pubenumLivenessPolicy{Cancel,Continue,Park}// aura_events::run — proposed#[derive(Clone,Copy,Debug,Serialize,Deserialize)]pubstructLiveness{pubpolicy:LivenessPolicy,/// How long after the last claim detaches before `policy` acts. A claim/// arriving within it cancels the timer. Zero applies the policy at once./// On the wire as `grace_ms`, like every other duration in the schema.#[serde(rename = "grace_ms", with = "duration_ms")]pubgrace:Duration,}implDefaultforLiveness{fndefault() -> Self{Self{policy:LivenessPolicy::Cancel,grace:Duration::ZERO}}}
Config, so a deployment can change a seam's default without the seam knowing. Seconds here, milliseconds on the wire — the codebase's existing split (timeout_secs in TOML, *_ms in events):
There is no a2a entry. An A2A task holds its run's claim for the task's whole life (#783), so ClaimsExhausted fires only after the run is already terminal, and a policy would never act.
Additional Context
Park is the only policy that leaves durable state behind, and it reuses the checkpoint and continuation surfaces in crates/aura/src/orchestration/park rather than growing a second set. A parked run holds nothing in memory: the runtime's record of it is AgentState::Parked on the session (#780), and resume-on-start reads only the checkpoint and the session stream.
Continue keeps the run alive, but what a later observer can see of it is bounded by the retention window and the journal store's bound (#779, #780). Until a durable journal exists (#210 / #325), Continue means "finish the work and keep it attachable for a while", not "keep it forever".
A disconnect under the default policy writes three events — ObserverDetached, ClaimsExhausted, LivenessDecided — and with grace zero the last two are always adjacent. They stay separate because ClaimsExhausted marks the start of a non-zero window and LivenessDecided its end; a consumer that only cares about the outcome reads the terminal event.
Summary
Losing the session's last claimant while a run is live cancels the run. With both seams on the runtime (#782, #783) that is the only behavior the runtime has, and it is the right one for a chat endpoint, where a run nobody reads is a run spending provider turns for no one. It is the wrong answer for a run meant to outlive its request: a task started headless, or a chat whose client is reconnecting.
Make the reaction a policy the run is started with. The vocabulary —
LivenessPolicy { Cancel, Continue, Park }and theClaimsExhausted/LivenessDecidedlifecycle events — is already inaura_events::run(#778); this gives it behavior.Goals
[hitl.park]and end the run resumable.Parkis refused atstartfor an agent with no checkpoint support — single-agent runs today — rather than silently downgraded toContinue. An operator who asked for Park getsStartError::PolicyUnsupported, not a run that behaves otherwise.ClaimsExhausted, thenLivenessDecided { policy }, so an observer arriving later reads why the run ended or kept going. A Cancel decision ends the run withCancelled { reason: Unclaimed }, distinguishable from an operator's/cancel(External) and from shutdown (Shutdown).RunContext(on nightly since PR Deliver a run's events through its handle and retire the request registries #738) and reads the run's token from it; [FEATURE]: HITL webhook route has no cancellation contract, stranding abandoned approvals #613 is the gap on the webhook arm and is a prerequisite for Cancel to mean what it says.stream_shutdown_tokenstill ends every run through the two-phase drain ([FEATURE]: A runtime that owns sessions and their runs #780); a Park policy parks instead of cancelling where it can.AgentRundirectly and never goes through the runtime.[runtime.park] resume_on_start = true, a runtime that boots and finds a session it holds — or, with [FEATURE]: Atomic Fence for VFS Claims #581, claims — whose last run isParkedresumes it through [FEATURE]: Endpoints to list, inspect, attach to, and control sessions and runs #785's/resumepath, seeded from the checkpoint and the session stream, with no caller involved. That is the "reified by policy on runtime start" case. Off by default, because an instance restarting under load should not fan out every parked run at once.Data structures
LivenessPolicyis as implemented in PR #730.Livenessis proposed beside it inaura_events::run, since it is serialized on the wire (#785) and inRunFacts(#780) and reuses that module'sduration_ms:Config, so a deployment can change a seam's default without the seam knowing. Seconds here, milliseconds on the wire — the codebase's existing split (
timeout_secsin TOML,*_msin events):There is no
a2aentry. An A2A task holds its run's claim for the task's whole life (#783), soClaimsExhaustedfires only after the run is already terminal, and a policy would never act.Additional Context
Park is the only policy that leaves durable state behind, and it reuses the checkpoint and continuation surfaces in
crates/aura/src/orchestration/parkrather than growing a second set. A parked run holds nothing in memory: the runtime's record of it isAgentState::Parkedon the session (#780), and resume-on-start reads only the checkpoint and the session stream.Continuekeeps the run alive, but what a later observer can see of it is bounded by the retention window and the journal store's bound (#779, #780). Until a durable journal exists (#210 / #325),Continuemeans "finish the work and keep it attachable for a while", not "keep it forever".A disconnect under the default policy writes three events —
ObserverDetached,ClaimsExhausted,LivenessDecided— and with grace zero the last two are always adjacent. They stay separate becauseClaimsExhaustedmarks the start of a non-zero window andLivenessDecidedits end; a consumer that only cares about the outcome reads the terminal event.Searched Issues
Code of Conduct