Skip to content

refactor(agent): separate a prepared agent from its runs - #719

Merged
justintime4tea merged 1 commit into
nightlyfrom
justingross/GH-628-separate-config-from-state
Oct 7, 2026
Merged

justintime4tea merged 1 commit into
nightlyfrom
justingross/GH-628-separate-config-from-state

Conversation

@justintime4tea

@justintime4tea justintime4tea commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Agent held config-derived values (model, turn depth, system prompt, context window) alongside state that belongs to one request (scratchpad budget, turn-nudge counters, the skill-invocation recorder) and was rebuilt for every chat request.

Split it in two. PreparedAgent is the reusable half, built once by PreparedAgent::prepare: the provider client, the tools discovered at build time, and the open MCP connections. Agent is one run of it, begun with PreparedAgent::begin_run, and owns the run's RunContext: its request id, a fresh scratchpad budget, fresh turn-nudge counters, and the recorder for the turn's skill invocations, alongside the event channel, tool-call queue and cancel token the context already carries. The recorder is handed to begin_run because what it records under, the turn's position in the history and its sequence within the turn, belongs to the turn. An orchestration worker or coordinator begins its run within the orchestration's, through begin_run_within: the same id, observer and cancellation, with tool state of its own.

Tools and wrappers are built with the prepared agent, so they can no longer hold run state directly. They reach the current run through BoundRun, the slot the HITL gate and request_approval tool already read: now also the turn-nudge wrapper and NudgedTool, the scratchpad wrapper and read tools, read_artifact, and the skill tools. The gate and tool take the request id from the run they are bound to. AgentRuntimeConfig loses its turn_nudge field accordingly.

A prepared agent serves one run at a time. Rig spawns its tool server once per agent, so two runs sharing one could not tell their tool calls apart; begin_run refuses a second run with RunInProgress while the first is alive, which RunLease counts: the Agent holds one, and so does every stream it produced.

A prepared agent is also bound to the request credentials it forwarded. headers_from_request values are applied to the MCP servers and the HITL webhook route when the agent is prepared, and the MCP connections open with them, so ForwardedHeaders records which inbound headers were forwarded and with what values, and begin_run refuses a request that forwards different ones. Without that, preparing once and running per request would run a later request's tool calls under the first request's credentials.

Behavior is unchanged: every request still prepares a fresh agent, and orchestration workers and the coordinator go through the same prepare-then-run path, with the coordinator's run recording skill invocations and workers' runs recording none. RigBuilder::prepare_agent exposes the reusable half for session-scoped reuse to build on.

Fixes: GH-628

@greptile-apps

greptile-apps Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High risk] Refactors agent lifecycle into prepare-once and run-many phases.

The PR appears safe to merge; no outstanding findings remain.

Summary

The PR separates reusable agent preparation from individual runs, moving request state into RunContext and making tools access it through BoundRun.

  • It prevents simultaneous runs on one prepared agent and rejects runs whose forwarded request headers differ from those used at preparation.
  • Direct agents, orchestration workers, and coordinators use the prepare-then-run lifecycle.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  C[Agent configuration and forwarded headers] --> P[PreparedAgent]
  P --> B[begin_run]
  B --> R[Agent and RunContext]
  R --> L[RunLease]
  R --> S[BoundRun]
  S --> T[Prepared tools and wrappers]
Loading

Reviews (9) · Last reviewed commit: "refactor(agent): separate a prepared age..." · Reviewed by Greptile

Comment thread crates/aura/src/rig_builder.rs
Comment thread crates/aura/src/builder.rs Outdated
@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch from 760f421 to 7205594 Compare September 24, 2026 14:55
Comment thread crates/aura/src/forwarded_headers.rs Outdated
@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch 2 times, most recently from e04038f to aa84e50 Compare September 24, 2026 15:03
@justintime4tea

Copy link
Copy Markdown
Collaborator Author

Waiting for #710 to merge so I can rebase afterwards since there will be conflicts and it'll be this branches code that needs updating to satisfy.

@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch from aa84e50 to a44af75 Compare September 29, 2026 17:57
@justintime4tea
justintime4tea changed the base branch from nightly to jakedipity/event-schema-default-path September 29, 2026 17:57
Comment thread crates/aura/src/config.rs Outdated
Comment thread crates/aura/src/run_context.rs
@jakedipity
jakedipity force-pushed the jakedipity/event-schema-default-path branch 6 times, most recently from 84cbbdd to 6221e16 Compare September 29, 2026 22:39
Base automatically changed from jakedipity/event-schema-default-path to nightly September 29, 2026 23:14
@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch from a44af75 to a9a6776 Compare September 30, 2026 17:23
@greptile-apps

This comment has been minimized.

@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch from a9a6776 to 99e1f78 Compare September 30, 2026 17:40
Comment thread crates/aura/src/scratchpad/context_budget.rs Outdated
@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch from 99e1f78 to 04074b9 Compare September 30, 2026 18:40
Comment thread crates/aura/src/builder.rs Outdated
@justintime4tea
justintime4tea marked this pull request as ready for review October 1, 2026 18:22
@justintime4tea
justintime4tea requested a review from a team October 1, 2026 18:22
@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch from 04074b9 to fe87379 Compare October 5, 2026 19:39
Shearerbeard
Shearerbeard previously approved these changes Oct 6, 2026
`Agent` held config-derived values (model, turn depth, system prompt,
context window) alongside state that belongs to one request (scratchpad
budget, turn-nudge counters, the skill-invocation recorder) and was
rebuilt for every chat request.

Split it in two. `PreparedAgent` is the reusable half, built once by
`PreparedAgent::prepare`: the provider client, the tools discovered at
build time, and the open MCP connections. `Agent` is one run of it,
begun with `PreparedAgent::begin_run`, and owns the run's `RunContext`:
its request id, a fresh scratchpad budget, fresh turn-nudge counters,
and the recorder for the turn's skill invocations, alongside the event
channel, tool-call queue and cancel token the context already carries.
The recorder is handed to `begin_run` because what it records under,
the turn's position in the history and its sequence within the turn,
belongs to the turn. An orchestration worker or coordinator begins its
run within the orchestration's, through `begin_run_within`: the same
id, observer and cancellation, with tool state of its own.

Tools and wrappers are built with the prepared agent, so they can no
longer hold run state directly. They reach the current run through
`BoundRun`, the slot the HITL gate and `request_approval` tool already
read: now also the turn-nudge wrapper and `NudgedTool`, the scratchpad
wrapper and read tools, `read_artifact`, and the skill tools. The gate
and tool take the request id from the run they are bound to.
`AgentRuntimeConfig` loses its `turn_nudge` field accordingly.

A prepared agent serves one run at a time. Rig spawns its tool server
once per agent, so two runs sharing one could not tell their tool calls
apart; `begin_run` refuses a second run with `RunInProgress` while the
first is alive, which `RunLease` counts: the `Agent` holds one, and so
does every stream it produced.

A prepared agent is also bound to the request credentials it forwarded.
`headers_from_request` values are applied to the MCP servers and the
HITL webhook route when the agent is prepared, and the MCP connections
open with them, so `ForwardedHeaders` records which inbound headers
were forwarded and with what values, and `begin_run` refuses a request
that forwards different ones. Without that, preparing once and running
per request would run a later request's tool calls under the first
request's credentials.

Behavior is unchanged: every request still prepares a fresh agent, and
orchestration workers and the coordinator go through the same
prepare-then-run path, with the coordinator's run recording skill
invocations and workers' runs recording none.
`RigBuilder::prepare_agent` exposes the reusable half for
session-scoped reuse to build on.

Fixes: GH-628
@justintime4tea
justintime4tea force-pushed the justingross/GH-628-separate-config-from-state branch from fe87379 to f9bb190 Compare October 7, 2026 18:38
@justintime4tea
justintime4tea enabled auto-merge (rebase) October 7, 2026 18:42

@Shearerbeard Shearerbeard left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

approved with questions for later

request_id = %self.request_id,
"no run bound to this gate; its approval reaches no observer"
),
None => tracing::warn!("no run bound to this gate; its approval reaches no observer"),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If this logs does it need more context or is that implicit somewhere?

@justintime4tea justintime4tea Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not implicit, no. It only fires on the park path (park_pre_call is the only caller of emit), and only if the worker's gate has no run. The gate picks up whatever run is in scope when create_worker builds it, and create_worker always runs inside the factory's with_run, so it shouldn't happen. But if it did, the warn would be close to useless. It used to log the request id from the gate; that id now comes from the run, and this branch is exactly the case where there's no run.

We could do a fast follow in two parts?

  • Bind the worker's gate the way a single agent's is bound: set hitl_gate / hitl_approval_tool on the worker's PreparedAgent (they're None today) so begin_run_within binds them.
    • Then the park path always has a run.
    • Approvals behave the same, since the worker's run has the orchestration's id and observer and is cancelled with it.
  • For whatever's left (a gate used with no run at all), add the scope (run id, task, session) and agent name to the warn?

This also lines up with #786, where the gate reads presence from the run it's bound to - binding worker gates to their own run is what makes that work for workers, as long as child runs share the session's presence and claim (adding that to #781).

Comment on lines +60 to +72
// No run bound means no budget to check the artifact against;
// withhold it rather than inline an unbounded read, as the
// scratchpad read tools do with `ScratchpadToolError::NoRun`.
let Some(budget) = sp.run.scratchpad_budget() else {
tracing::warn!(
"read_artifact: artifact {} withheld — no run is bound",
filename
);
return format!(
"[artifact '{filename}' withheld: no run is bound to count it against. \
Retry the read_artifact call.]"
);
};

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What situation are we in where we don't have data to run scratchpad budget? Just omitted model data in the config?

@justintime4tea justintime4tea Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not config - if the context window isn't configured, scratchpad never gets wired up, self.scratchpad is None, and read_artifact takes the first branch and inlines.

This branch is scratchpad on, but no run bound to the tool's slot. Tools are built once when the agent is prepared, before any run exists, so they reach the run's budget through BoundRun instead of holding it. begin_run / begin_run_within bind the run before rig can call any tool, so we can't get here through an Agent. If we ever did, it withholds the artifact rather than inline an unbounded read, same as the scratchpad read tools do with ScratchpadToolError::NoRun. Before the split the tool held the budget directly, so this state couldn't exist.

It isn't impossible by construction, though: the slot starts empty, and rig calls tools on its own task with no context, so a slot is the only way the tool can find its run. Making it impossible means handing each call its run, which needs a change in our rig fork, #787 already says it lives within that rig constraint, so if we want it, it could be its own issue.

The fast follow I'm thinking here is the tool gets its slot from the same place begin_run binds instead of matching it by convention (related to Jake's comment item 11), and the "retry" goes away from the withheld message, since a retry would be withheld too (related to Jake's comment item 7).

@justintime4tea
justintime4tea merged commit 513f1ec into nightly Oct 7, 2026
11 checks passed
@justintime4tea
justintime4tea deleted the justingross/GH-628-separate-config-from-state branch October 7, 2026 19:13
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 7, 2026
@jakedipity

Copy link
Copy Markdown
Collaborator

Went through this after it merged. A couple of these change behavior for callers today. Most of the rest can't happen yet because nothing reuses a prepared agent outside tests, but they will as soon as something does.

Changes behavior now

  1. stream ignores the request_id it's given and uses the one from begin_run, which for build_streaming_agent is config.request_id and defaults to "" (builder.rs#L2001). If you build without setting it and then call stream(..., "req_123") like the streaming.rs doc example, the hook key, MCP in-flight tracking and approvals are all keyed by "", and cancel_and_close_mcp("req_123") matches nothing. Two agents built that way share "" and can release each other's approvals. Before this the run used the id passed in. I think this should either be an error or stream shouldn't take an id at all. A debug log is easy to miss.

  2. The RunContext is built once in begin_run, so a second stream call on the same Agent reuses it (builder.rs#L2020). The first AgentRun's drop guard has already cancelled the token, so the second stream stops at the first hook. It also gets no observer since events was taken, and it inherits leftover ids in the tool-call FIFO. Two concurrent calls on one Arc<dyn StreamingAgent> would mix up tool_call_ids. It'd be good if a second call failed loudly.

Shows up once a prepared agent is reused

  1. The lease doc says "nothing calls a tool of a run that has ended", but rig's tool server awaits toolset.call inline (builder.rs#L1031). If the client disconnects mid-call, the stream drops and frees the lease while the call keeps going. If the next request has hit begin_run by the time it finishes, the old call counts its scratchpad tokens against the new run's budget and picks up the new run's nudge, and once the new stream does bind_call its MCP progress goes to the new observer. Capturing the run when the call starts instead of reading the slot afterwards would fix it.

  2. prepare says nothing in it depends on a request, but additional_tools, client_tools and session_id all go in there (builder.rs#L355). The web handler passes run_tools() through additional_tools, and that includes slack_search with the asker's action_token and the per-message search budget. begin_run only checks forwarded headers, so prepare once and begin per request means request B searches with A's Slack token and B's approvals carry A's session_id. That's the crossover ForwardedHeaders is meant to stop. I'd move those into begin_run or fingerprint them with the headers.

  3. RunContext::child hands the child parent.cancel.clone() rather than child_token() (run_context.rs#L103). Cancelling a child cancels the whole orchestration. Nothing calls .stream() on a worker or coordinator Agent yet, but if it did, that stream's drop guard would kill the sibling workers and the client stream when it ends.

  4. The HITL gate's pre_call reads the slot twice, once for run and again for request_id (gate.rs#L320). If the slot is rebound in between, the approval carries B's id but waits on A's token. RequestApprovalTool::call uses one run for both, and this should too.

Other bugs

  1. With no run bound, the scratchpad wrapper withholds the output and tells the model to retry (wrapper.rs#L116). The MCP call has already run by then, so a side-effecting tool runs twice, and the retry is withheld too because the slot is still unbound. At minimum I'd drop the retry instruction and say the call ran.

  2. worker_config.skill_recorder = None was removed, so worker configs now carry the coordinator's session recorder (orchestrator.rs#L780). Workers only avoid recording because begin_run_within gets None. Anything that builds a worker through Agent::new or build_streaming_agent_with_tools would write task-scoped skill loads into the session again. I'd put that line back.

Cleanup

  1. stream spawns a task per stream just to forward the caller's cancellation, because the run's token exists before the caller's (builder.rs#L2013). If begin_run took the parent token and used child_token(), the task goes away and an already-cancelled parent is seen right away.

  2. The Deref to PreparedAgent puts the seed fields scratchpad_budget and turn_nudge right next to the scratchpad_budget() and turn_nudge() methods that return live run state (builder.rs#L192). Forget the parens and it compiles but reads a budget that never moves. Renaming them seed_* or dropping the Deref would rule that out.

  3. add_all_tools gets the run slot from both its run param and config.scratchpad_tools_config.run (builder.rs#L1271). They only match by convention, and the tests already pass different ones. Dropping ScratchpadToolsConfig.run would leave one source.

  4. ForwardedHeaders::of/resolve redo the case-insensitive lookup from apply_request_header_mappings and walk the same headers_from_request maps as resolve_mcp_headers_in (forwarded_headers.rs#L60). The credential check is only right while they agree, so I'd derive both from one helper.

  5. The commit footer says Fixes: GH-628. It needs to be Fixes: #628 for GitHub to link and close the issue.

@justintime4tea

justintime4tea commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

Went through this after it merged. A couple of these change behavior for callers today. Most of the rest can't happen yet because nothing reuses a prepared agent outside tests, but they will as soon as something does.

Thanks for going through this Jacob, sorry it got merged before I could get to your feedback. Agree with almost all of it. None of it is reachable through the server today (every caller prepares per request, begins one run under the id it streams with, and streams once). But 1 and 2 are real regressions, and 3/4 have to be fixed before anything reuses a prepared agent.

This is what I'm thinking for fast follows:

  1. stream errors when the id doesn't match the run's, and fix the streaming.rs example. Dropping the id from stream entirely would be great, but that's trait-wide, so I'd do that in [FEATURE]: A runtime that owns sessions and their runs #780 - the runtime mints the run's id once and hands it to begin_run, so after that nothing needs to pass it to stream.
  2. A second StreamingAgent::stream fails loudly. Only the trait method though? The coordinator retry loop re-streams the same Agent through stream_chat_with_depth and so it needs to keep working. [FEATURE]: A runtime that owns sessions and their runs #780 makes this moot for the server since the runtime streams each run once, but library callers can hit it today.
  3. Right, that doc line is wrong once a call outlives its stream. Maybe we can have an in-flight call capture the run and hold the lease, so begin_run can't rebind under it? It touches every slot reader plus bind_call. It's really a prerequisite for [FEATURE]: A prepared agent stays warm across a session's runs #787 (warm prepared agent), which counts on one run at a time per prepared agent and this is the hole in that, so I'll add it to [FEATURE]: A prepared agent stays warm across a session's runs #787's goals.
  4. Agreed, and the prepare doc is wrong about it. We should fix the doc in a fast follow. The actual fix belongs to [FEATURE]: A prepared agent stays warm across a session's runs #787: session is already in its cache key, but the tools aren't, so I'll add that there (run tools reached through the run rather than baked in at prepare, client tools in the key).
  5. You're right it should be child_token(). Not a regression as of yet though because workers shared the orchestration token before, and nothing needs a worker to cancel its parent. We can fast follow this?
  6. That makes sense, we should read it once like the tool does.
  7. Good catch - the call already ran, and a retry gets withheld too. Fast follow: say it ran and was withheld, no retry. Same wording fix in read_artifact?
  8. Good call.
  9. Agreed. begin_run needs the caller's token for that, and the handler only has it at stream time. [FEATURE]: A runtime that owns sessions and their runs #780's StartRun carries the caller's RunOptions, so the runtime has the token when it calls begin_run - this goes with 1 in [FEATURE]: A runtime that owns sessions and their runs #780.
  10. Agreed, we should rename to seed_scratchpad_budget / seed_turn_nudge.
  11. Agreed, we should drop ScratchpadToolsConfig.run and just have the one.
  12. Agreed, we should have one helper for the lookup and one walk of the mappings.
  13. GH- closes too - it just closes when the commit reaches main (our default branch), not when the PR merges to nightly. #396 and #724 both closed from their Fixes: GH-… footers seconds after #749 synced nightly into main. #628 will close the same way on the next sync. I got push back from some of the other devs for using "#" related to something in semantic release I believe? I can't recall exactly where/why the push back but was instructed to use either "gh-" or "GH-"?

Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants