Skip to content

Proposal: attribute recall and adoption to the agent session in every agent #884

Description

@SaulMoro

Today. Recall feedback only half works, and only on Claude. #883 part 1 (recall quality never reached the agent's session) is fixed by #887; this issue now carries part 2.

teamai recall        → recalled_count++ per doc; recall quality under the agent session (#887)   every agent
adoption (upvotes)   → Claude transcript parser at Stop                                            ≈ 0 on the subagent path, none elsewhere

On Claude's recommended path the recall runs inside the teamai-recall subagent, whose turns go to <session>/subagents/agent-*.jsonl. The main transcript keeps only the summary and the recalled-doc-ids comment, with ids and no paths, so a later Read of the doc can't be matched:

main agent → teamai-recall subagent → teamai recall
main transcript   <!-- teamai:recalled-doc-ids: [redis-timeout] -->
main agent        Read /…/learnings/redis-timeout.md
Stop → parseTranscriptForVotes   recalled: [redis-timeout]   adopted: []

Where the parser does run, it counts too much and too little. A Glob, find or ls that only lists the doc counts, and so does a Grep whose path is the doc even with no match. Cursor's Shell, Copilot's view and sed -n don't count. Other agents never had adoption: the parser reads only Claude's JSONL. Upvotes feed KB Health, search ranking and promotion, which needs at least 5 upvotes (docs/usage-guide.md:1287).

Proposal. Every recall run is recorded with its session, and adoption comes from the PostToolUse hook that almost every agent already has.

teamai recall   → run record {run, session, actor, docs[key, scope, printed path]}   active scope, every agent
PostToolUse     → claim: the shell call that ran `teamai recall` confirms the run's session
                  evidence: a read of a file under the knowledge roots
reducer         → evidence × settled runs, same session, within 24 h, not the recall subagent itself
                  → upvote through the existing per-session ledger        at Stop, SubagentStop and pull
teamai stats    → recall for the last 10 sessions: runs, recalled, adopted
  • Upvotes count again on Claude's subagent path, and start counting on Codex, CodeBuddy, Qoder, Copilot, Cursor, Pi, OpenCode and OMP.
  • The query is never stored, and nothing new goes to git.

This serves the roadmap in #647 (Team Context: "Improve Recall's recall and precision, and strengthen feedback mechanisms"). It doesn't depend on OpenTelemetry; #877 can later export these records as-is.

Goal. For any supported agent, TeamAI knows which session ran each recall and whether that session then opened what it got.

Terms.

  • A recall run is one teamai recall process that searched. --check is not a run.
  • Adoption means "opened after recall": the same session reads a recalled doc within 24 h of the run. A main agent that uses the subagent's summary without opening the doc doesn't adopt. Only the opt-in judge (TEAMAI_UPVOTE_JUDGE) sees that kind of use.
  • An actor is (session, agent_id): the main agent, or one subagent inside the session.
  • A run is settled when a hook claim confirms it, or when agentSessionIdFromEnv() saw a single candidate at run time. Only settled runs vote.
  • A bridge is the plugin or extension TeamAI generates for agents without settings-file hooks (OpenCode, Pi, OMP).
flowchart LR
  R["teamai recall<br/>stdout: … run=&lt;uuid&gt;"] -->|"env session, provisional"| L["recall log<br/>active scope"]
  P["PostToolUse<br/>shell call running teamai recall"] -->|"claim: session + actor"| L
  P2["PostToolUse<br/>read under the knowledge roots"] -->|"evidence"| L
  L --> X["reducer<br/>Stop · SubagentStop · pull"]
  X --> V["votes ledger → upvote"]
  L --> S["teamai stats"]
  L -.->|"later"| O["#877 export"]
Loading

Which session ran the recall. Two sources, so either can be missing:

at run time    agentSessionIdFromEnv() from #887; a run id printed on the region's start line
at hook time   the PostToolUse of the shell call that ran it carries session_id, agent_id and the output
--- [teamai:recall:start] --- (3 results) run=5f0c2e1a-9b7d-4c1e-8a54-2f6d0b3e9c71

A claim is valid only when all three hold:

  • the hook's command invokes teamai recall itself (parsed, not a substring);
  • the run id exists in the local log;
  • it is the first valid claim.

A later claim that disagrees is recorded but not applied. So an outer codex exec "run teamai recall … and return its stdout" can't take the run from the inner session that really ran it. The run id is a crypto.randomUUID(), and every run in one shell call is extracted.

The hook wins over the env value because nesting fools the env. #887 picks the session whose current run started last, but that fails when the outer agent compacts during a background inner run. The env value covers the cases with no hook call, such as a Codex command still running when its tool call returns. It votes only when it was unambiguous.

The recall subagent's own reads don't count; other subagents' reads do. A run is marked as the recall subagent's when either:

  • the command carries --caller teamai-recall, a hidden flag that agents/teamai-recall.md passes; or
  • the hook says agent_type == teamai-recall.

Reads by the actor that ran a marked run don't count. Reads by the main agent or by any other subagent in the same session do. The flag covers agents whose hooks name no subagent type (Cursor, Copilot). agent_type covers the case where the model drops the flag.

The recall log is local, in the active scope, one line per event under the usage file's lock:

{"kind":"run","ts":"2026-09-28T10:02:11Z","run":"5f0c2e1a-…","session":"5624be7c-…","agent":"claude","actor":{"agent_id":"a-19"},"caller":"teamai-recall","via":"env",
 "docs":[{"key":"redis-timeout","type":"learning","scope":"project","path":"/…/learnings/redis-timeout.md","score":7.2,"eligible":true}]}
{"kind":"claim","ts":"…","run":"5f0c2e1a-…","session":"5624be7c-…","actor":{"agent_id":"a-19","agent_type":"teamai-recall"}}
{"kind":"evidence","ts":"…","session":"5624be7c-…","actor":{},"path":"/…/learnings/redis-timeout.md"}
{"kind":"link","ts":"…","child":"ses_child","parent":"ses_parent"}
  • Each doc keeps its vote key, its source scope, and the path recall printed. Inherited user-scope hits keep eligible: false, so they stay read-only as they are today.
  • A full recall with no hits is recorded with docs: []. TEAMAI_RECALL_DISABLED turns off both recall recording and hook recording.
  • No query, no prompt, no file content. votes keeps its per-doc counters. The team sees nothing new unless Proposal: export TeamAI's team-level telemetry over OpenTelemetry #877 exports it.
  • Retention: 30 days, with 5000 lines as a backstop. Pruning happens at pull under the lock, never on the hook path, and never drops evidence younger than 24 h that is still pending.

Adoption is decided by tool category, not by tool name. The normalizer src/utils/tool-names.ts maps every agent's names, and a name it doesn't know never counts:

read     Read, view, read_file, read, ReadFile, NotebookRead…     the input path is a recalled doc
shell    Bash, Shell, PowerShell, execute_command, run_in_terminal, bash…
         one simple command, or a pipeline starting with a reader; no &&, ||, ;
         read verbs    cat bat batcat less more head tail nl, sed print form without -i,
                       Get-Content gc type                         → the doc is a file operand
         search verbs  rg grep egrep fgrep ag ack, git grep         → as the search row
         list verbs    ls find fd tree, rg --files, git ls-files    → never
search   Grep, grep_code, search_content, grep…                   no -c, --count or count mode, and either
         the input path is the doc and the output isn't empty, or an output line starts with the doc's path
         followed by ':' (content shown). A bare path line is a listing. Structured {filenames[]} don't count;
         their `content` string goes through the same line rule. OMP's markdown-tree grep and Cursor's Grep
         don't count until an adapter or a live payload exists.
list     Glob, glob, list_dir, search_file, list_files              never
  • Path. Compared against the path recall printed. It must be equal once resolved against the call's cwd or search root. The suffix match (basename plus at least 2 trailing segments) applies only when no base is known, and is rejected when ambiguous. Windows paths are normalized: drive-letter case, separators, and Git Bash /c/.
  • Status. A failed call doesn't count. When the status is unknown, as with Codex's shell output, the call counts only for a simple read of a just-recalled path.
  • Window. Same session, from the run to 24 h after it. That matches the votes ledger's TTL (src/votes.ts:77), so a resumed session can't credit the same evidence twice.

The reducer does the join and calls incrementUpvoted, and evidence is marked consumed when credited. votes-sync stops parsing transcripts for adoption and no longer needs transcript_path, which Copilot, OpenCode, OMP and Pi don't send. The judge keeps the parser for its candidates and skips docs already in the ledger.

Bridges forward what their hosts already give them. Today they send only cwd, tool_name and tool_input, so their session falls back to pid-<ppid>-<cwd>:

 payload to teamai hook-dispatch, on every event (SessionStart, prompt, PostToolUse, Stop)
   cwd, tool_name, tool_input
+  session_id            OpenCode event sessionID · Pi ctx.sessionManager.getSessionId() · OMP ctx.sessionManager.getSessionId()
+  tool output, status   so PostToolUse sees the run id and failed reads
 OMP
+  agent_id, agent_type  from ctx.agent when kind == "sub" (read only if present: OMP ≥ 18.3.2)
 OpenCode
+  shell.env → TEAMAI_AGENT_SESSION_ID                          so the recall process knows its session
+  `task` tool.execute.after → link {child → parent}            from output.metadata.{sessionId, parentSessionId}
 src/utils/session-id.ts
+  TEAMAI_AGENT_SESSION_ID and PI_SESSION_ID in AGENT_SESSION_ENV; OPENCODE and PI_SESSION_ID out of BRIDGE_AGENT_ENV
+  deriveSessionId reads conversation_id (Cursor) after sessionId

Evidence per agent.

Agent Session at run time Subagent tool calls in PostToolUse Direct recall Subagent path
Claude Code CLAUDE_CODE_SESSION_ID root session_id + agent_id + agent_type yes yes
Codex CODEX_SESSION_ID (≥ 0.148) root session_id + agent_id + agent_type (≥ 0.134) yes, simple shell reads yes
CodeBuddy, WorkBuddy CODEBUDDY_SESSION_ID parent session_id + agent_id + agent_type (≥ 2.103.1; WorkBuddy to check) yes yes; nested subagents to check
Qoder none (hook claim only) session_id + agent_id + agent_type; root session to check yes yes if the session is the root
Copilot CLI COPILOT_AGENT_SESSION_ID the subagent's own sessionId, no agent field yes no: no link to the parent
Cursor CURSOR_CONVERSATION_ID the subagent's own conversation_id, no parent field (a confirmed Cursor gap) yes no: no link to the parent
OpenCode TEAMAI_AGENT_SESSION_ID via the bridge the child session; parent from the task link yes yes
OMP none (hook claim only) own session + ctx.agent; parent not linked yes no in v1
Pi PI_SESSION_ID once the bridge sends it no subagents yes n/a
ZCode none (hook claim only) hooks don't fire in subagents yes no

What I'd put in v1, as one PR:

  • The run id on the region's start line and the recall log in the active scope, with claims, evidence, links and retention.
  • Session from env at run time, settled by the claim of the same call. The --caller flag in agents/teamai-recall.md, and dropping its stale referenced-doc-ids instruction.
  • Adoption from PostToolUse with the rule above, through one reducer at Stop, SubagentStop and pull, replacing the transcript parser for adoption.
  • Bridges forwarding the session id on every event, plus the tool output, status and OMP's ctx.agent; OpenCode's shell.env and task link.
  • teamai stats: a "Recall (last 10 sessions)" section in the default output (session, agent, runs, recalled, adopted), shown only when the active scope's log has runs.
  • Docs: the recall section of docs/usage-guide.md and docs/usage-guide.zh-CN.md, with the per-agent table above, replacing the notes that say adoption is Claude-only (e.g. OpenCode, docs/usage-guide.md:1839); stats in both guides; the core skill; commands.md regenerated for --caller. The README support table has no recall column, so it doesn't change.

What could wait.

  • A dashboard view of the same records.
  • OMP's parent link, through getHeader().parentSession, once verified on a real OMP.
  • Search adapters for OMP's grep tree and Cursor's Grep.
  • Attributing each recall to a turn inside the session.
  • Exporting the recall log through Proposal: export TeamAI's team-level telemetry over OpenTelemetry #877.
  • Adoption for tools without PostToolUse (OpenClaw, Hermes, Kiro, JoyCode).

Acceptance check.

  • Alice asks Claude Code a question. The teamai-recall subagent finds redis-timeout and reads it, then the main agent reads it too. Her session upvotes it once; the subagent's own read adds nothing. Today the same session gets adoptedDocIds: [].
  • Bob does the same in Codex. sed -n '1,80p' learnings/redis-timeout.md counts, and test -e x && cat learnings/redis-timeout.md || true doesn't.
  • Carol's OpenCode session runs a recall and never opens the doc. The log has her session and the doc, and nothing is upvoted.
  • Dan runs codex exec from inside Claude. The recall is attributed to the Codex session, even though the Claude shell's output shows the same run id.
  • Erin resumes Alice's session the next day and reads redis-timeout again without a new recall. No second vote.
  • Frank's Cursor recall subagent reads the doc it found. No vote.

In every case the log contains no query, and teamai stats lists each session with its runs, recalled and adopted docs.

Design notes
  • Why not process ancestry. Walking up the parent processes to the agent's recorded pid breaks too often to be the key. A resumed Claude session kept running under a new pid while its recorded monitorPid was dead. Two Codex sessions share one app-server pid. No process start time is stored, so reused pids can't be told apart. On Windows each lookup costs about 1.3 s of PowerShell. The run id makes the join exact without it.

  • Why not a transcript parser per agent. Codex rollouts, CodeBuddy's index.json, Copilot's events.jsonl and OpenCode's database each hold tool calls in a different shape. Copilot's log is deliberately not read, for privacy (src/dashboard-collector.ts:1234-1240). PostToolUse gives the same facts through one path that TeamAI already receives.

  • Phantom recalls. The parser trusts assistant text blocks (src/transcript-parser.ts:228-241), so an agent that quotes the recalled-doc-ids marker creates recalls that never ran. A run now exists only when a teamai recall process records it, and a claim only when the hook sees that process's command. The judge still takes candidates from the transcript. That is a known limit of the opt-in path.

  • What the metric misses. The recall subagent inlines what the main agent needs (agents/teamai-recall.md:365-366), so the main agent may never open the doc. That use is real, but "opened after recall" doesn't see it. The log itself will show how often the main agent opens the doc after a subagent summary. Forcing extra reads just to create votes is not an option.

  • Reading Claude's subagent transcript: considered, not pursued. <session>/subagents/agent-*.jsonl holds the full recall region, so reading it at Stop would restore adoption on Claude's subagent path alone. It covers one agent and adds parser code this proposal removes.

  • Run id placement. Older CLIs find the region by the --- [teamai:recall:start] --- prefix (src/transcript-parser.ts:11), so text after (N results) doesn't break them. The recalled-doc-ids comment has a strict pattern (src/transcript-parser.ts:429) and stays as it is.

  • Hook cost. PostToolUse already runs for every tool call, in the foreground: 4.5 s per handler in TeamAI, and 10 s at most in Cursor, Copilot and CodeBuddy. The new work is limited to:

    • a category check on the tool name;
    • a path check against the knowledge roots, or a string match for the run marker in shell output;
    • one append, under a short lock wait.

    The hook never reads the log. The join runs in the reducer.

  • Privacy. A recall query can contain user text, so it isn't logged. Tool output is read in the hook to find the run id, and isn't stored. The log files are owner-only.

  • Older CLIs. A member on an older CLI keeps the transcript path and upvotes as today. Votes from both paths go through the same per-session ledger (incrementUpvoted), so a doc isn't upvoted twice in one session within the ledger's window. The agent file with --caller is deployed by the same binary that parses it.

  • Copilot and Cursor subagents. The subagent runs under its own session, and no hook field ties it to the parent. Copilot 1.0.81's "re-emitted on its parent" concerns hook.start/hook.end lifecycle records, not the tool payload. So the subagent path gets no adoption there, and the direct path works. --caller keeps the subagent's own reads from voting.

To verify during the spec, each with one live payload:

  • Cursor's PostToolUse session field and its Grep input and output.
  • Copilot's shell-hook tool names, and the session a subagent's shell hook carries.
  • Qoder's root session inside subagents, CodeBuddy's nested subagents, and WorkBuddy's tool names.
  • Codex's agent_type for a custom agent.
  • Claude's Grep content-mode format.
  • OMP's ctx.agent below 18.3.2.
  • Whether Qoder exports a session variable in its shell, for AGENT_SESSION_ENV.

Feedback from anyone who uses recall in more than one agent would help:

  • Do the upvotes you see today match how often your agents actually use recalled docs?
  • Is "opened within 24 h in the same session" the right bar for adoption?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions