Skip to content

[BUG]: Otel does not drain before server shutdown #305

Description

@Shearerbeard

Summary

We wait on SSE drain when we send a shutdown signal to AURA however we do not wait for otel to drain. Under normal loads this isn't the end of the world but this really bites us when benchmarks timeout and docker kills aura - under rapid load a lot of otel traces can build up in the buffer waiting for a flush. This leads to timed out benchmark runs that leave us no way to diagnose.

Reproduction

  • run aura over a long series of benchmarks with the otel collector pointing at Phoenix - notice the latency can get pretty high over large output runs like Terminal Bench
  • Kill server part way through noting the most recent trace
  • Notice that the bulk of the trace never makes it out of AURA

Relevant log output

Additional Context

No response

Upload screenshots

No response

Searched Issues

  • No similar issues found

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

  1. self-assigned this
    on Jul 1, 2026
  2. Shearerbeard commented on Aug 12, 2026

    @Shearerbeard
    CollaboratorAuthor

    An audit of the mainline binary at 7a0f0651 for the TerminalBench work
    turned up a more specific mechanism than this issue describes, and it
    shifts where the fix belongs.

    Aura does wait for OTel. shutdown_tracer wraps provider.shutdown() in
    a 5s bounded timeout (crates/aura/src/logging.rs:638-656), and the
    graceful-shutdown path waits for in-flight requests to drain before it
    runs. The traces still go missing because of what happens before the
    flush rather than the flush itself.

    Spans that are still OPEN at shutdown never end, so the flush has nothing
    to export for them. execute_completion runs in a detached
    tokio::spawn holding the root agent.stream span
    (crates/aura-web-server/src/handlers.rs:599-602), and nothing aborts
    it. Three paths lose spans:

    • Grace expires while the client is already gone. axum::serve has no
      connection left to wait on, returns immediately, and the tracer drains
      while the task still holds its span open.
    • A handler parked in an MCP connect never returns. The connect,
      initialize, and discover chain has no deadline (crates/aura/src/mcp.rs:501-520;
      no .connect_timeout() at mcp_sse.rs:118 or
      mcp_streamable_http.rs:148), so the HTTP connection stays open,
      axum::serve never returns, and shutdown_tracer never runs at all.
      An external SIGKILL then loses queued spans too, which is the case
      this issue reports.
    • The orchestration inner task is spawned before its cancellable select
      (crates/aura/src/orchestration/factory.rs:62,64) and holds a clone of
      the root span. Ending the outer task drops cancel_tx, and the watcher
      reads a dropped sender as normal completion
      (orchestrator.rs:227-228), so the inner task leaks with spans open.

    aura-cli has no SIGTERM handler at all (only SIGINT, as a REPL flag at
    crates/aura-cli/src/ui/state.rs:443-450), so standalone runs lose
    everything on a signal.

    Two repairs have no counterpart here yet: aborting live request tasks
    after the drain and before the flush so their spans end, and a drain
    deadline on axum::serve so a stuck task cannot hold shutdown past its
    budget. Both were proven in a prototype against a kill-with-open-span
    probe, which failed pre-fix with nothing exported.

    One adjacent finding for whoever picks this up: the benchmark adapter
    sends SIGTERM and SIGKILLs 5s later while mainline's grace default is
    30s, so any in-flight request at stop guarantees the tracer is never
    reached. The budgets need to agree.

  3. self-assigned this
    on Aug 18, 2026
  4. added 11 commits that reference this issue on Sep 16, 2026
    ffe1d8f
    413c43a
    57a4da6
    825e0ce
    46482c7
    643427f
    8427ce3
    0088768
    1222763
    c74a5f0
    3883d36
  5. added 4 commits that reference this issue on Oct 5, 2026
    3a2fd56
    6c278e8
    9b95c00
    6adf23b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions