Skip to content

[FEATURE]: Endpoints to list, inspect, attach to, and control sessions and runs #785

Description

@justintime4tea

Summary

Once a run can outlive its request, something has to be able to find it. Today the only management surface is A2A's, shaped for agents: a task id you already have, a subscribe that is the execution stream, a cancel. The epic's "native API intent based layer for kicking off a headless task", and its endpoints for "understanding what is running at any given time", have no home. Neither does the resume endpoint the park work leaves for later.

Give the runtime an HTTP surface of its own. It is the runtime's public form: every route maps onto one runtime call, and the handlers are the only code outside the runtime that hold it.

Goals

Inspect, attach, cancel — needs only the runtime and observer kinds:

  • GET /v1/sessions — sessions this instance holds, with each one's AgentState ([FEATURE]: A runtime that owns sessions and their runs #780): Preparing, Running, Blocked and on what, Parked, Done. GET /v1/sessions/{id} — one session's summary: its state, agent, live and last run, latest sequence, and this instance's observer counts. This is "what is running right now", answered per agent rather than per request.
  • GET /v1/sessions/{id}/events — an SSE stream of SessionEvents from a cursor; a collecting attach by default, with claim and presence as query parameters for the other kinds ([FEATURE]: Observers declare whether they collect, claim, or attend #781). This is "zoom in": the projection is the envelope itself, not the chat wire. Attaching to a session with no live run is allowed; the stream delivers the next run from its Started. A second claim is refused with the current claimant named; take-over is not offered over HTTP.
  • GET /v1/runs — live and recently ended runs, filterable by session, status, and pending approval. GET /v1/runs/{id} — one run's summary: its RunFacts ([FEATURE]: A runtime that owns sessions and their runs #780) plus the session's observer counts. GET /v1/runs/{id}/events — the session stream from the run's Started; the same attach, with the cursor filled in.
  • POST /v1/runs/{id}/cancel — with a reason; the lifecycle event carries External and the message.

Start and continue — needs the liveness policy:

  • POST /v1/runs — start a run: agent, the conversation's messages in the chat endpoint's shape, session, liveness policy ([FEATURE]: Liveness policy when a run's last claim detaches #784), and optionally the run it continues. One request serves every case a seam has: a headless start on a new session; the next turn after a run finished with a ClarificationNeeded or an answer, where continues records the lineage and liveness defaults to the prior run's; a start into an existing session. It returns 202 with the run id while the run is Preparing; Started arrives on the events stream. A start into a session with a live run is 409 naming that run ([FEATURE]: A runtime that owns sessions and their runs #780). (background: true belongs to the Responses API, which AURA does not serve; if it ever does, it can front this same start.)
  • The messages travel with the request because the server keeps no session history today. The session stream ([FEATURE]: A session's event journal and late attach #779) now makes history derivable — Started { prompt } and each run's Completed are the turns in order — but whether the server reconstructs it from the stream or [EPIC] State Management #210 stores messages beside it is [EPIC] State Management #210's decision, and until it is made messages is required. Nothing here injects a prompt into a running turn; HITL decisions and the turn nudge are the only input that enters a run, between its turns.
  • POST /v1/runs/{id}/resume — rehydrate a parked run through the continuation surfaces in orchestration::park and start it under the runtime with its recorded decisions. Also a new run that continues the parked one; it is a separate route because its seed is the checkpoint, not messages. (It could become POST /v1/runs { continues } with no messages on a Parked run, the runtime dispatching on the prior run's status — left as an option. [FEATURE]: Liveness policy when a run's last claim detaches #784's resume-on-start does the same thing without a request.) PR Park and resume for human-in-the-loop approvals #760 (open) adds a resume endpoint at POST /v1/sessions/{session_id}/runs/{run_id}, guarded by an in-process ResumeClaimTable; this route is that endpoint under the runtime, and the two paths need to agree.

Cross-instance — needs the session claim:

  • A session this instance does not hold is reachable: the summary comes from the holder, and an events attach relays over EventBus on session:{session_id} with the sequence numbers and fell-behind contract bus_bridge has, generalized from A2A tasks to every session. The relay is a lease in the holder's LeaseStore ([FEATURE]: Atomic Fence for VFS Claims #581), so a relaying instance that dies is detached with Expired. A cancel routes the way a2a:cancel:{task_id} does, and the holder records the outcome. Until [FEATURE]: Atomic Fence for VFS Claims #581 lands, a session elsewhere is reported as not found.

Throughout:

  • Authorization the same as the existing routes for now; these follow whatever strategy the existing endpoints adopt.

Data structures

Proposed wire types, in aura-web-server:

// POST /v1/runs
pub struct StartRunRequest {
    pub agent: Option<String>,             // AgentInfo::id; the default agent when absent
    pub messages: Vec<ChatMessage>,        // the chat endpoint's shape; the last user message is the prompt
    pub session_id: Option<SessionId>,     // a new session when absent
    pub continues: Option<RunId>,          // the run this one follows; liveness defaults to its
    pub liveness: Option<Liveness>,        // #784; the `headless` config default when absent
}
pub struct StartRunResponse {              // 202 — the run is Preparing; 409 — RunInProgress { run_id }
    pub run_id: RunId,
    pub session_id: SessionId,
    pub continues: Option<RunId>,
}

// GET /v1/sessions, GET /v1/sessions/{id}
pub struct ListSessionsQuery { pub state: Option<AgentStateKind>, pub agent: Option<String>, pub limit: Option<u32>, pub cursor: Option<String> }
pub struct SessionSummary {
    pub session_id: SessionId,
    pub agent: String,
    #[serde(flatten)]
    pub state: AgentState,                 // #780 — { "state": "blocked", "on": { "kind": "approval", ... } }
    pub current_run: Option<RunId>,
    pub last_run: Option<RunId>,
    pub latest_seq: Option<SequenceNumber>,
    pub observers: ObserverCounts,         // #781
}

// GET /v1/runs
pub struct ListRunsQuery {
    pub session_id: Option<SessionId>,
    pub status: Option<RunStatus>,
    pub pending_approval: Option<bool>,
    pub limit: Option<u32>,
    pub cursor: Option<String>,
}
pub struct ListRunsResponse { pub runs: Vec<RunSummary>, pub next_cursor: Option<String> }

// GET /v1/runs/{id} — RunFacts (#780) plus what only this instance knows
pub struct RunSummary {
    #[serde(flatten)]
    pub facts: RunFacts,                   // usage snapshotted live while running
    pub observers: ObserverCounts,         // the session's, #781
}

// GET /v1/sessions/{id}/events, GET /v1/runs/{id}/events
pub struct EventsQuery {
    pub after: Option<u64>,                // replay events with seq > after; absent = from the start
                                           // (/runs/{id}/events: absent = from the run's Started)
    pub claim: Option<bool>,               // attach as the claimant (#781); refused if one exists
    pub presence: Option<bool>,            // a human is at this observer
}
// body: `text/event-stream` of SessionEvent JSON, one per `data:` line, plus a
// `gap` event carrying `resumed_at` when the subscriber fell behind (#779)

// POST /v1/runs/{id}/cancel
pub struct CancelRunRequest { pub reason: Option<String> }

Additional Context

There is no user or tenant model, so "the same as the existing routes" means any caller with the API key can list and attach to every session on the instance. That is the current exposure of /v1/chat/completions extended to sessions the caller did not start. It is accepted for V1 because the auth work in flight is where the answer belongs, not here.

?after= is the session sequence number the client has already seen, so a reconnecting client asks for after=<last seq> and gets no duplicate and no hole. It matches Cursor::After (#779); a from that was ambiguous about inclusivity was the reconnect off-by-one waiting to happen. On the run-scoped route the default cursor is the run's first_seq; an explicit after is still a session sequence.

The session summary's state is constructed, not stored — the derivation table is in #780 — so an instance can answer it for any session it holds without a store read, and a different implementation could construct it differently without this API noticing.

Presence on ?presence=true is self-declared, as it is for every observer (#781).

Searched Issues

  • No similar issues found

Code of Conduct

  • I agree to follow this project's Code of Conduct

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions