Problem
The aura.usage SSE event currently has dual semantics:
- In single-agent mode, it reports the usage mathed out in a way that its attempting to only capture turn by turn conversation context fils
- In orchestration mode the results seem mostly accurate (verify ground truth with provider - for example Bedrock runs can be verified after in Cloudwatch metrics)
The semantics are inconsistent and we need to follow normal OpenAI style SSE events for usage that ONLY report the turn usage in full. We should add another aura event to help client accurately calculate only the context window additive tokens per turn (basically total - inner react loop tokens) or simply a calculation of the outboud output tokens. This is partially provided by rig already and accurate on rigs end (at least in bedrock - more ground truth needs verified with other providers).
Expected Behavior
- Standard
usage field just reports 1:1 with what we get from our provider for both single and orchestrated
- some sort of custom SSE field that accurately calculates just the context window payloads (basically returned agent turns). Preferable if we don't have to calculate our own with tiktoken.
Problem
The
aura.usageSSE event currently has dual semantics:The semantics are inconsistent and we need to follow normal OpenAI style SSE events for usage that ONLY report the turn usage in full. We should add another aura event to help client accurately calculate only the context window additive tokens per turn (basically total - inner react loop tokens) or simply a calculation of the outboud output tokens. This is partially provided by rig already and accurate on rigs end (at least in bedrock - more ground truth needs verified with other providers).
Expected Behavior
usagefield just reports 1:1 with what we get from our provider for both single and orchestrated