Skip to content

[FEATURE]: HITL agent called semantics need guardrails #306

Description

@Shearerbeard

Summary

We have seen evidence of agents pre-calling tools they don't know are hard gated with HITL. There is evidence of them shoving tool call data into these fields thinking it will be helpful for the user. We need to look into providing specific guidlines on when to call HITL manually + maybe some context around what tools are hard gated to avoid the approval -> approval flowl

Additional Context

Some diagnosis of a real agent called HITL that ended up being confusing UX - this is lifted from slack:

Mike Shearer  [10:42 AM]
So this part of the breakdown should be pretty deterministic other than the "reason" field.

{
"version":1,
"decision_id":"019f1416-71bb-7c01-b5b3-ad3b3176b37a",
"request_id":"req_8608a5bed4874ba2b79d6d87aa8587de",
"scope":{
"kind":"worker",
"run_id":"019f1416-1df4-71b1-9134-5b0fc379e332",
"task_id":2,
"worker":"researcher",
"session_id":"eaa390db-18aa-444d-810d-aa3d8dd707f7"
},
"origin":{
"kind":"agent_requested",
"reason":"Requires access to potentially sensitive production logs in Mezmo; needs human approval before querying."
}, The confusing part is where it maps to Items in the webhook structure. Its using theses LLM fields from the tool schema:

"properties": {
"action_description": {
"type": "string",
"description": "What you want to do (the action awaiting approval)."
},
"risk_rationale": {
"type": "string",
"description": "Why this action requires human approval."
},
"context": {
"type": "object",
"description": "Optional additional structured metadata for the reviewer."
}
},Which gets oddly mapped back to items[0] and reads strange because the LLM is essentially meta describing the tool it wants to run in an effort to explain why its deciding to call HITL manually.
"items":[
{
"tool_name":"request_approval",
"arguments":{
"action_description":"Run get_correlated_timeline_relative_time for the last 5 minutes (window ending at 2026-06-29T15:53:41Z) to build a correlated cross-service timeline.",
"context":{
"dedup_mode":"template",
"grouping_field":"_app",
"max_logs_per_source":20,
"max_timeline_events":50,
"query":null,
"since":"last 5 minutes",
"tool":"functions.get_correlated_timeline_relative_time"
},
"risk_rationale":"Requires access to potentially sensitive production logs in Mezmo; needs human approval before querying."
}
}[10:43 AM]Its asking for approval on a situation it deems high risk but might be off the mark because we don't put anything into explaining when it should ask for approval or what tools are already hard gated.
[10:44 AM]I think we should make the output schema for Agent called HITL more explicit and add some guidance as to when it should be calling that manually

Upload screenshots

No response

Searched Issues

  • No similar issues found

Code of Conduct

  • I agree to follow this project's Code of Conduct

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions