Repository navigation
[BUG]: Otel does not drain before server shutdown #305
Description
Activity
- added a parent issue
on Jul 1, 2026 An audit of the mainline binary at
7a0f0651for the TerminalBench work
turned up a more specific mechanism than this issue describes, and it
shifts where the fix belongs.Aura does wait for OTel.
shutdown_tracerwrapsprovider.shutdown()in
a 5s bounded timeout (crates/aura/src/logging.rs:638-656), and the
graceful-shutdown path waits for in-flight requests to drain before it
runs. The traces still go missing because of what happens before the
flush rather than the flush itself.Spans that are still OPEN at shutdown never end, so the flush has nothing
to export for them.execute_completionruns in a detached
tokio::spawnholding the rootagent.streamspan
(crates/aura-web-server/src/handlers.rs:599-602), and nothing aborts
it. Three paths lose spans:- Grace expires while the client is already gone.
axum::servehas no
connection left to wait on, returns immediately, and the tracer drains
while the task still holds its span open. - A handler parked in an MCP connect never returns. The connect,
initialize, and discover chain has no deadline (crates/aura/src/mcp.rs:501-520;
no.connect_timeout()atmcp_sse.rs:118or
mcp_streamable_http.rs:148), so the HTTP connection stays open,
axum::servenever returns, andshutdown_tracernever runs at all.
An external SIGKILL then loses queued spans too, which is the case
this issue reports. - The orchestration inner task is spawned before its cancellable select
(crates/aura/src/orchestration/factory.rs:62,64) and holds a clone of
the root span. Ending the outer task dropscancel_tx, and the watcher
reads a dropped sender as normal completion
(orchestrator.rs:227-228), so the inner task leaks with spans open.
aura-clihas no SIGTERM handler at all (only SIGINT, as a REPL flag at
crates/aura-cli/src/ui/state.rs:443-450), so standalone runs lose
everything on a signal.Two repairs have no counterpart here yet: aborting live request tasks
after the drain and before the flush so their spans end, and a drain
deadline onaxum::serveso a stuck task cannot hold shutdown past its
budget. Both were proven in a prototype against a kill-with-open-span
probe, which failed pre-fix with nothing exported.One adjacent finding for whoever picks this up: the benchmark adapter
sends SIGTERM and SIGKILLs 5s later while mainline's grace default is
30s, so any in-flight request at stop guarantees the tracer is never
reached. The budgets need to agree.- Grace expires while the client is already gone.
- added 11 commits that reference this issue
on Sep 16, 2026 - added 4 commits that reference this issue
on Oct 5, 2026
Summary
We wait on SSE drain when we send a shutdown signal to AURA however we do not wait for otel to drain. Under normal loads this isn't the end of the world but this really bites us when benchmarks timeout and docker kills aura - under rapid load a lot of otel traces can build up in the buffer waiting for a flush. This leads to timed out benchmark runs that leave us no way to diagnose.
Reproduction
Relevant log output
Additional Context
No response
Upload screenshots
No response
Searched Issues
Code of Conduct