Skip to content

HITL approval gates for MCP tool calls #191

Description

@Shearerbeard

Problem

Agents execute tool calls unconditionally. There is no way to require a human to sign off before a destructive call runs, which blocks SRE dogfooding of the orchestrated agent: incident remediation needs an operator's approval before anything irreversible happens.

Two deployment shapes need approval and differ structurally:

  • Unattended (A2A, background): no human on the stream, so an external service answers.
  • Attended (CLI, dev, dogfooding): the operator on the SSE stream is the approver. SSE is one-way, so the decision needs a path back up.

Decision

Recorded in the ADR (docs/adr/2026-06-16-hitl-approval-architecture.md) with the type-level design in docs/design/hitl.md: dual-channel routing chosen by config, fail-closed, per-process registry.

Routes, fixed per deployment in a [hitl] table:

  • Webhook (unattended): one synchronous round-trip; decision in the response body.
  • Conversational (attended): the gated call parks in-process while the SSE stream stays open; the decision returns via POST /v1/approvals/{decision_id}. The park lasts as long as the connection, and a dropped stream fails closed.

Where a single-agent client advertises its own request_approval tool, the attended surface reuses the existing turn-ending client-tool cycle.

[hitl]
require_approval = ["kubectl_*", "restart_*"]

[hitl.route]
mode = "conversational"   # or: mode = "webhook", url = "https://..."
timeout_secs = 300

Plan

Each phase has a workspace build/test/clippy gate and an e2e smoke against a local mock webhook:

  1. Domain plus module split, webhook arm only.
  2. Conversational route: registry, decision ingress, disconnect/shutdown cancellation.
  3. Attended client plus orchestration verification.

Landed: the id newtypes (RunId/SessionId/TaskIdentity) and the aura::hitl typed skeleton (signatures complete, todo!() bodies).

Out of scope (tracked elsewhere)

Durable park-and-resume and exactly-once binding (#209), ingress auth, A2A InputRequired wiring, confidence-based auto-approve, multi-webhook fan-out, an Auto route, coordinator-mediated surfaces.

Notes

Activity

  1. added 3 commits that reference this issue on Jun 3, 2026
    bbe222c
    8053483
    80fb6f5
  2. changed the title [-]Add human in the loop (HITL) gating both config based and agent discretionary[/-] [+]HITL approval gates for MCP tool calls[/+] on Jun 4, 2026
  3. added a commit that references this issue on Jun 4, 2026
    2773d08
  4. rpelevin commented on Jun 5, 2026

    @rpelevin

    This looks like the right boundary for destructive MCP tools. One invariant I would make explicit in the webhook contract: the approval response should bind to the exact intercepted tool call, not just to the matching glob or to the fact that a human approved this tool once.

    For the approval request payload, I would include:

    • stable decision/request id;
    • mode: config gate or agent-requested approval;
    • agent/run identity;
    • MCP server and tool;
    • target/resource if known;
    • canonical args digest;
    • matched rule or policy version;
    • expiry/timeout;
    • decision state: allow, deny, or needs_more_context;
    • reason code once resolved.

    Then, immediately before execution, Aura can revalidate that server, tool, target, args digest, policy version, and expiry still match. If any of those changed, the previous approval is stale and the call should fail closed or raise a new approval request.

    That also gives single-agent and orchestration mode one shared contract: single-agent can trace the same decision id, and orchestration can emit aura.approval_requested / aura.approval_completed around it.

  5. self-assigned this
    on Jun 5, 2026
  6. Keesan12 commented on Jun 5, 2026

    @Keesan12

    The webhook contract looks solid. The extra invariant I’d add is exact-once execution after approval.

    Approval should not just say approved=true; it should mint a decision/consumption id that is marked spent when the guarded tool call actually runs. That closes the replay hole where the same approved payload gets executed twice after retries, reconnects, or orchestration restarts.

  7. 155 remaining items

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions