You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[FEATURE]: A prepared agent stays warm across a session's runs #787
Every request builds its agent from scratch: provider client, tool discovery, MCP connections. PR #719 separates the config-derived half of that into PreparedAgent, built once by prepare and serving runs through begin_run, and stops there — RigBuilder::prepare_agent exposes it "for session-scoped reuse to build on" and nothing builds on it. The runtime is where a session's runs meet, so it is where reuse happens.
Goals
The session record's agent slot ([FEATURE]: A runtime that owns sessions and their runs #780) keeps the PreparedAgent across the session's runs, built on first use and reused by every later run that matches the forwarded-header set it was prepared with. Preparation already lives in the runtime, so the cache is the slot plus eviction bookkeeping, and no seam changes.
Preparing becomes a once-per-session state. With the slot warm, AgentState::Preparing ([FEATURE]: A runtime that owns sessions and their runs #780) is observable on the first run only; later runs go to Running as fast as begin_run. It is also where the session claim ([FEATURE]: Atomic Fence for VFS Claims #581) and, later, the write claim on the session's memory directory are taken, so "booting the agent" and "owning the session" are one step.
The cache key includes the forwarded headers. PR refactor(agent): separate a prepared agent from its runs #719's begin_run refuses a request whose forwarded headers differ from the ones the agent was prepared with (BeginRunError::ForwardedHeaderDiffers), because the MCP connections were opened with the preparing request's credentials. A cache keyed on session alone would hit that error on every credential change; keyed on both, it prepares a fresh agent instead.
One run at a time per prepared agent is PR refactor(agent): separate a prepared agent from its runs #719's rule (begin_run returns RunInProgress while a RunLease is alive) and one live run per session is [FEATURE]: A runtime that owns sessions and their runs #780's (StartError::RunInProgress, a 409 naming the run); this issue relies on both and tightens the first, so a run's lease outlasts the tool calls it started (below). What it changes is that the refused second start would have hit the same prepared agent, so the cache never has to spin up a second agent for one session. Whether a second start waits instead of being refused is discussion question 10.
A tool call keeps the run it started in. Tools and wrappers find the current run through the prepared agent's BoundRun, which each begin_run rebinds. Rig's tool server finishes a call that is in flight when its stream is dropped, and the stream's drop frees the RunLease, so with a warm agent the next run can begin while that call is still executing — and the call then charges the new run's scratchpad budget and turn nudge, and once the new stream binds the MCP manager, reports its progress to the new run's observer. A call captures its run once, when it starts, and carries it through pre_call, transform_output, and post_call; and it holds the RunLease until it returns, so begin_run refuses with RunInProgress until the previous run's calls have finished. Every slot reader changes: the scratchpad wrapper and read tools, the turn nudge, read_artifact, the skill tools, the HITL gate and request_approval, and McpManager::bind_call.
A warm agent carries no request's tools.prepare takes additional_tools and client_tools, and StartRun ([FEATURE]: A runtime that owns sessions and their runs #780) supplies both per run, but the agent is prepared once per session. The server's run tools arrive through additional_tools, and slack_search holds the asking user's Slack action_token and a per-message search budget, so a warm agent would run the next request's searches as the first asker. Per-run tools reach the run the way run state does: a stable tool registered at prepare that dispatches to the instance begin_run hands the run. Client tools are definitions only, so they join CacheKey as a hash of their definitions. session_id, which the HITL gate stamps into its scope at prepare, is already in the key.
Idle eviction with a configured TTL; eviction closes the MCP connections it holds. Config reload invalidates.
Warm discovery demonstrable: a second turn in a session emits aura.mcp_status without reconnecting, and a test asserts discovery ran once.
Cache size bounded, oldest-idle evicted first.
Data structures
Proposed, on the runtime (#780). The session record's AgentSlot is the cache entry; this is the bookkeeping around it:
pubstructPreparedAgentCache{entries:RwLock<HashMap<CacheKey,CachedAgent>>,config:CacheConfig,}#[derive(Hash,Eq,PartialEq)]pubstructCacheKey{session:SessionId,headers:HeadersFingerprint,// keyed hash of the forwarded headers PR #719 records on the agentclient_tools:ToolsFingerprint,// hash of the request's client-tool definitions; empty when noneagent:String,// AgentInfo::id — a session may switch agents}/// Derived from credentials, so: a keyed hash with a per-process random key,/// no `Serialize`, no `Debug` that prints it, never logged. Comparable only/// within one process lifetime, which is all a cache of live connections needs.pubstructHeadersFingerprint([u8;32]);/// Hash of the client-tool definitions (name, description, parameters) a run supplies.pubstructToolsFingerprint([u8;32]);pubstructCachedAgent{agent:Arc<PreparedAgent>,last_used:Instant,config_version:u64,// bumped on reload; a stale entry is evicted on next use}pubstructCacheConfig{pubidle_ttl:Duration,pubmax_entries:usize,}implPreparedAgentCache{pubasyncfnget_or_prepare(&self,key:CacheKey,prepare:implFuture<Output = Result<Arc<PreparedAgent>,BuilderError>>) -> Result<Arc<PreparedAgent>,BuilderError>;pubasyncfnevict_idle(&self);// on a timerpubasyncfninvalidate_all(&self);// on config reload}
The rig-side constraint PR #719 and #732 describe — a tool call cannot see which run, or which stream within a run, it belongs to — is what caps this at one run per prepared agent. The first of the tool-call goals above is what makes "one at a time" hold across a run boundary rather than only between leases. Lifting the cap needs per-run tool instances or a rig change that hands each call its run; this issue lives within the constraint.
Summary
Every request builds its agent from scratch: provider client, tool discovery, MCP connections. PR #719 separates the config-derived half of that into
PreparedAgent, built once byprepareand serving runs throughbegin_run, and stops there —RigBuilder::prepare_agentexposes it "for session-scoped reuse to build on" and nothing builds on it. The runtime is where a session's runs meet, so it is where reuse happens.Goals
PreparedAgentacross the session's runs, built on first use and reused by every later run that matches the forwarded-header set it was prepared with. Preparation already lives in the runtime, so the cache is the slot plus eviction bookkeeping, and no seam changes.Preparingbecomes a once-per-session state. With the slot warm,AgentState::Preparing([FEATURE]: A runtime that owns sessions and their runs #780) is observable on the first run only; later runs go toRunningas fast asbegin_run. It is also where the session claim ([FEATURE]: Atomic Fence for VFS Claims #581) and, later, the write claim on the session's memory directory are taken, so "booting the agent" and "owning the session" are one step.begin_runrefuses a request whose forwarded headers differ from the ones the agent was prepared with (BeginRunError::ForwardedHeaderDiffers), because the MCP connections were opened with the preparing request's credentials. A cache keyed on session alone would hit that error on every credential change; keyed on both, it prepares a fresh agent instead.begin_runreturnsRunInProgresswhile aRunLeaseis alive) and one live run per session is [FEATURE]: A runtime that owns sessions and their runs #780's (StartError::RunInProgress, a409naming the run); this issue relies on both and tightens the first, so a run's lease outlasts the tool calls it started (below). What it changes is that the refused second start would have hit the same prepared agent, so the cache never has to spin up a second agent for one session. Whether a second start waits instead of being refused is discussion question 10.BoundRun, which eachbegin_runrebinds. Rig's tool server finishes a call that is in flight when its stream is dropped, and the stream's drop frees theRunLease, so with a warm agent the next run can begin while that call is still executing — and the call then charges the new run's scratchpad budget and turn nudge, and once the new stream binds the MCP manager, reports its progress to the new run's observer. A call captures its run once, when it starts, and carries it throughpre_call,transform_output, andpost_call; and it holds theRunLeaseuntil it returns, sobegin_runrefuses withRunInProgressuntil the previous run's calls have finished. Every slot reader changes: the scratchpad wrapper and read tools, the turn nudge,read_artifact, the skill tools, the HITL gate andrequest_approval, andMcpManager::bind_call.preparetakesadditional_toolsandclient_tools, andStartRun([FEATURE]: A runtime that owns sessions and their runs #780) supplies both per run, but the agent is prepared once per session. The server's run tools arrive throughadditional_tools, andslack_searchholds the asking user's Slackaction_tokenand a per-message search budget, so a warm agent would run the next request's searches as the first asker. Per-run tools reach the run the way run state does: a stable tool registered at prepare that dispatches to the instancebegin_runhands the run. Client tools are definitions only, so they joinCacheKeyas a hash of their definitions.session_id, which the HITL gate stamps into its scope at prepare, is already in the key.aura.mcp_statuswithout reconnecting, and a test asserts discovery ran once.Data structures
Proposed, on the runtime (#780). The session record's
AgentSlotis the cache entry; this is the bookkeeping around it:Additional Context
The rig-side constraint PR #719 and #732 describe — a tool call cannot see which run, or which stream within a run, it belongs to — is what caps this at one run per prepared agent. The first of the tool-call goals above is what makes "one at a time" hold across a run boundary rather than only between leases. Lifting the cap needs per-run tool instances or a rig change that hands each call its run; this issue lives within the constraint.
Searched Issues
Code of Conduct