Skip to content

CR-AR-001: Risk-Tiered Deferred Human Review - #153

Merged
coreytshaffer merged 2 commits into
mainfrom
claude/cr-ar-001-deferred-review
Aug 8, 2026
Merged

coreytshaffer merged 2 commits into
mainfrom
claude/cr-ar-001-deferred-review

Conversation

@coreytshaffer

Copy link
Copy Markdown
Owner

Scope

Exactly one new documentation file. 506 insertions, zero other changes.

  • Added: docs/change/requests/CR-AR-001-risk-tiered-deferred-human-review.md
  • Not touched: docs/current_backlog.md, docs/backlog.md, docs/futures/futures_register.md, docs/change/change_log.md, and everything under triage_core/, schemas/, tests/.
  • No relationship to docs: nominate governed knowledge substrate research venue #152 (the GBrain venue-candidate nomination). This branch was cut directly from main at c86a117 and does not carry that commit. Independent docs-only slice, independent review.

Status of the CR itself

  • Status: Proposed
  • Implementation Authority: Not authorized

Merging this PR approves the governance/design proposal recorded in the document — nothing more. It does not authorize the Initial Implementation Slice described in the CR; that slice requires its own separate approval and bounded file allowlist, the same sequencing already used for CR-DD-012B and CR-OC-001C through CR-OC-001E. No source code, schema, CLI, ledger, or runtime change is authorized by this document or by merging it.

What the CR proposes

A deferred-terminal-review authorization mode for TriageCore's agentic execution control plane, addressing approval fatigue from requiring human sign-off before every low-risk tool call.

  • R0–R5 risk tiers (Observe / Ephemeral / Reversible / Staged External / Consequential / Privileged-Critical), each with a default review mode.
  • Monotonic risk escalation — a workflow may escalate tier, never silently de-escalate, and cannot exceed its claim's ceiling without an explicit CAPABILITY_ESCALATION_REQUIRED result the agent cannot reason its way past.
  • Cumulative/compositional risk budgets — individually low-risk actions (e.g. many reads, many small sends) can collectively cross a threshold; budgets are claimed atomically so concurrent workflows can't each believe they hold the same remaining budget.
  • Risk floors — conditions (credential access, security-policy mutation, untrusted-content-driven mutation, missing audit evidence, etc.) that impose a minimum tier regardless of an otherwise lower calculated score, and are never averaged away.
  • Untrusted-input escalation — a mutation causally influenced by untrusted retrieved content gets at least R3 treatment, based on provenance/trust classification rather than the model judging whether the input "looks malicious."
  • A capability-claim schema extension (review_mode, effect_ceiling, budgets, reversibility, terminal_review, evidence) building on the existing atomic capability-claim work.
  • Eleven runtime requirements (AR-001 through AR-011) — e.g. the mediated executor enforces the claim on every call regardless of review mode; deferral changes review timing only, never validation; an agent cannot modify its own risk tier, budget, or authorization policy.
  • A mechanically-derived terminal evidence bundle — the reviewer sees intent digest, capability-claim ID, policy hash, tools invoked, exact deltas, denials, attempted escalations, and what was explicitly not executed — not an agent-authored summary.
  • Terminal dispositions (accepted / revised / rejected) with required compensation on rejection, and a state model with a staged path for R3.
  • 13 required tests and acceptance criteria scoped to a future, separately-authorized Initial Implementation Slice (R0–R3 only; no autonomous R4/R5 in this CR at all).

Terminology invariant preserved throughout

Deferred human review ≠ deferred authorization. Runtime authorization and enforcement remain contemporaneous — on every tool call, regardless of review mode — even when the human review step itself moves to the workflow boundary. This distinction is stated explicitly in the Status, Design Principle, Runtime Requirements, and Security Invariants sections so it can't be read as "review is deferred, therefore enforcement is too."

Structural notes

  • Adapted from a longer draft into the repository's ##-heading CR format used by larger existing CRs (e.g. CR-YK-002) rather than a numbered-list style.
  • No cross-reference added to current_backlog.md or futures_register.md: an earlier proposed-only CR (CR-004) was never listed in either backlog document while in Proposed status, so a standalone CR doc with no backlog entry is an established pattern, not an omission.
  • Acceptance criteria are explicitly gated on the Initial Implementation Slice being separately authorized — accepting this CR document cannot be read as having silently granted that authorization.

Verification

Check Result
git diff --cached --stat 1 file changed, 506 insertions, 0 deletions
git diff --cached --check no whitespace errors
git status --short single A entry, no residue

Tests not run: no executable, schema, or test surface changed. No effect on the open daily-use evidence window.

🤖 Generated with Claude Code

Add CR-AR-001 as a proposed, unauthorized CR document defining a
deferred-terminal-review authorization mode: an R0-R5 risk-tier scheme,
monotonic escalation, cumulative/compositional risk budgets, risk floors,
untrusted-input escalation, a capability-claim extension, eleven runtime
requirements (AR-001..AR-011), a mechanically-derived terminal evidence
bundle, and required tests/acceptance criteria for a future Initial
Implementation Slice.

Status: Proposed. Implementation Authority: Not authorized. Accepting this
document approves the governance/design proposal only; the Initial
Implementation Slice requires its own separate approval and bounded file
allowlist, matching the CR-DD-012B and CR-OC-001C-E sequencing pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@netlify

netlify Bot commented Aug 8, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for poetic-quokka-0fd859 ready!

Name Link
🔨 Latest commit b8c1810
🔍 Latest deploy log https://app.netlify.com/projects/poetic-quokka-0fd859/deploys/6a76b657c18e5f000824e7fa
😎 Deploy Preview https://deploy-preview-153--poetic-quokka-0fd859.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@coreytshaffer

Copy link
Copy Markdown
Owner Author

Pushed `b8c1810`: a documentation-only revision of CR-AR-001 responding to a read-only adversarial design review (chat-only memo, not committed anywhere in the repo). No implementation, schema, test, or runtime change — Status: Proposed and Implementation Authority: Not authorized are unchanged.

The review found the design's problems were specification gaps at the CR-YK-002 integration boundary, not evidence the deferred-review concept itself is unsound. Four items were treated as genuine blockers (one reframed) and resolved in this revision:

  • B1 — the two state machines weren't related. State Model now states explicitly that the workflow/review lifecycle is a separate state machine layered above the existing closed issued/claimed/completed/failed capability-claim enum, not an extension of it, with a full state-mapping table.
  • B2 — Revised's effect on the original claim was undefined. The original claim must now reach failed (never completed, never left open); already-executed effects become provisional evidence attributed to that failed claim; a new/amended claim must explicitly reference which provisional effects it retains, by claim ID and effect digest. Added a matching Required Test and Acceptance Criterion.
  • B3 (reframed) — compensation was assumed, not specified. New Compensation Contract section defines what a policy-recognized, testable compensation contract must specify, states an agent's own reversibility claim never establishes eligibility, and explicitly defers any concrete mechanism to a future CR — TriageCore has none today (mediated_executor.py explicitly performs no automatic rollback).
  • B4 — terminal-review assurance was unstated. New Terminal-Review Assurance section distinguishes pre-effect authorization / deferred review / retrospective disposition / acknowledgement, and states a disposition is evidence-minimum and not equivalent to the existing WebAuthn HumanAuthorizationReceipt unless a future CR says so.

Also addressed, not labeled blockers: risk-classification evidence now records which rule/floor produced the tier (new Classification Basis section + Terminal Evidence Bundle fields); REVIEW_PENDING is required to stay durable/visible/queryable and distinct from ACCEPTED — deliberately without an invented timeout, since an indefinite pending state is fine as long as it never silently reads as accepted; and the disposition vocabulary (accepted/revised/rejected) is explicitly reconciled against evidence_schema.md's review_decision, supervisor_decision, and the informal daily-use-evidence-window "operator disposition" — boundary drawn, nothing renamed.

Full diff and per-finding mapping in the session record. Suggested next step: a second, narrower review of just this revision — whether B1–B4 are genuinely closed and whether the new language introduces any contradictions — before any implementation-authorization discussion.

@coreytshaffer
coreytshaffer merged commit aea7ec9 into main Aug 8, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant