Problem
The duplicate-call guard has two coupled issues:
-
Surrender-inducing phrasing + output hiding. When a tool call repeats identically, the guard returns a ToolError ending with "Write your final answer now. Do not call this tool again." The phrasing tips the worker into surrender — models respond with "I was unable to complete the task" despite having turn budget remaining. Returning Err at validate_args means the tool is never called and the real output is hidden.
-
Over-firing on legitimate retries. The guard counts ANY repeated (tool, args, result) triple, including transient MCP errors (503, timeouts, rate limits). An LLM retrying a temporarily-unavailable tool is penalized the same as one stuck in a schema-error loop.
The surrender output is then mislabeled success=true by the downstream grading path — see "Orchestrator marks soft-failed worker tasks as success=true" (filed separately).
Solution
Two-stage escalation via annotation on the real tool output (never hidden):
nudge_threshold — appends [DUPLICATE_CALL_GUIDANCE] with observation framing and alternatives
block_threshold — appends [DUPLICATE_CALL_ABORT] and sets an Arc<AtomicBool> escalation flag
Error-kind-aware counting using CallOutcome from #68: GeneralToolError does not increment the counter (legitimate retries). Success and SchemaError with identical output do.
Per-fingerprint tracking via HashMap<CallFingerprint, CallState> prevents ping-pong evasion where alternating tools reset each other's counters.
Related code
crates/aura/src/orchestration/duplicate_call_guard.rs — guard implementation
crates/aura/src/prompts/duplicate_call_guidance.md — nudge template
crates/aura/src/prompts/duplicate_call_abort.md — abort template
crates/aura/src/orchestration/orchestrator.rs — AgentWithPreamble escalation flag plumbing
Problem
The duplicate-call guard has two coupled issues:
Surrender-inducing phrasing + output hiding. When a tool call repeats identically, the guard returns a
ToolErrorending with "Write your final answer now. Do not call this tool again." The phrasing tips the worker into surrender — models respond with "I was unable to complete the task" despite having turn budget remaining. ReturningErratvalidate_argsmeans the tool is never called and the real output is hidden.Over-firing on legitimate retries. The guard counts ANY repeated
(tool, args, result)triple, including transient MCP errors (503, timeouts, rate limits). An LLM retrying a temporarily-unavailable tool is penalized the same as one stuck in a schema-error loop.The surrender output is then mislabeled
success=trueby the downstream grading path — see "Orchestrator marks soft-failed worker tasks as success=true" (filed separately).Solution
Two-stage escalation via annotation on the real tool output (never hidden):
nudge_threshold— appends[DUPLICATE_CALL_GUIDANCE]with observation framing and alternativesblock_threshold— appends[DUPLICATE_CALL_ABORT]and sets anArc<AtomicBool>escalation flagError-kind-aware counting using
CallOutcomefrom #68:GeneralToolErrordoes not increment the counter (legitimate retries).SuccessandSchemaErrorwith identical output do.Per-fingerprint tracking via
HashMap<CallFingerprint, CallState>prevents ping-pong evasion where alternating tools reset each other's counters.Related code
crates/aura/src/orchestration/duplicate_call_guard.rs— guard implementationcrates/aura/src/prompts/duplicate_call_guidance.md— nudge templatecrates/aura/src/prompts/duplicate_call_abort.md— abort templatecrates/aura/src/orchestration/orchestrator.rs—AgentWithPreambleescalation flag plumbing