Outcome
AURA can continue or recover long-running work across request timeouts, process restarts, and supported multi-instance deployments without losing the state required to inspect, resume, or complete that work.
A running task should not disappear simply because the process or node that originally owned it goes away.
Current state
Today, important AURA state still lives in process memory and is tied to the instance currently executing the work.
- Long-running work can outlive the request that started it.
- Restarting the process can destroy conversation, task, approval, or execution state that has not been persisted elsewhere.
- State owned by one node may not be visible to another node.
- OTEL provides a durable record of what happened, but does not provide resumable execution state.
- HITL parking, A2A tasks, and other headless workflows already expose places where process-local state is insufficient.
Use cases
Durable execution should support scenarios such as:
- A headless task continues after the web request that started it has timed out.
- If an AURA process restarts while work is running, the new process can recover the state required to resume or safely reconcile that work.
- A task waiting for human approval can remain parked for an extended period and resume after the decision arrives, even if the AURA process restarts.
- A different AURA instance can inspect or continue work that was started elsewhere.
- Streaming, cancellation, and task status continue to work correctly in a multi-node deployment.
- AURA can perform maintenance on the infrastructure running AURA itself, including restarting or upgrading that infrastructure, without losing responsibility for the task.
Scope
Durable runtime state
Persist the state required for supported headless tasks to survive beyond the lifetime of a single request or process.
The durability contract should clearly distinguish:
- state required to resume execution;
- state required to inspect or control work;
- durable conversation/session state;
- transient telemetry or derived state that can be reconstructed.
Recovery and reification
AURA should be able to recover supported parked or interrupted work from durable state.
Recovery must define what can safely resume, what must be retried or reconciled, and what should fail rather than risk duplicate or unsafe execution.
Multi-instance execution
Supported runtime state must be accessible across AURA instances where required.
Multiple nodes should be able to:
- discover existing work;
- inspect its state;
- route control actions to the instance currently responsible for it;
- recover ownership when that instance is no longer available.
Coordination
Durable storage alone is not sufficient for active work.
The runtime needs a coordination mechanism for cases such as:
- waking parked work;
- propagating task events;
- routing cancellation or steering;
- preventing multiple instances from executing work that should have a single owner.
Related work
#325 Aura session storage and persistence provides much of the existing storage and cross-node foundation for this epic.
Existing work under #325 already addresses durable A2A task state, cross-node event delivery, and related storage abstractions.
#271 HITL Park/Reify depends on durable execution state and is an important proving use case for recovery after long waits.
#562 Agents Runtime owns the protocol-independent runtime, lifecycle, and control model. This epic makes that runtime durable across process and infrastructure boundaries.
Parent initiative: #747
Outcome
AURA can continue or recover long-running work across request timeouts, process restarts, and supported multi-instance deployments without losing the state required to inspect, resume, or complete that work.
A running task should not disappear simply because the process or node that originally owned it goes away.
Current state
Today, important AURA state still lives in process memory and is tied to the instance currently executing the work.
Use cases
Durable execution should support scenarios such as:
Scope
Durable runtime state
Persist the state required for supported headless tasks to survive beyond the lifetime of a single request or process.
The durability contract should clearly distinguish:
Recovery and reification
AURA should be able to recover supported parked or interrupted work from durable state.
Recovery must define what can safely resume, what must be retried or reconciled, and what should fail rather than risk duplicate or unsafe execution.
Multi-instance execution
Supported runtime state must be accessible across AURA instances where required.
Multiple nodes should be able to:
Coordination
Durable storage alone is not sufficient for active work.
The runtime needs a coordination mechanism for cases such as:
Related work
#325 Aura session storage and persistence provides much of the existing storage and cross-node foundation for this epic.
Existing work under #325 already addresses durable A2A task state, cross-node event delivery, and related storage abstractions.
#271 HITL Park/Reify depends on durable execution state and is an important proving use case for recovery after long waits.
#562 Agents Runtime owns the protocol-independent runtime, lifecycle, and control model. This epic makes that runtime durable across process and infrastructure boundaries.
Parent initiative: #747