Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
59 changes: 54 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,9 +65,9 @@ Sometimes a customer, auditor, or partner needs evidence that a human approved a

This is an **experimental cryptographic preview**, not a production compliance claim. It uses an embedded BN254/Groth16 circuit and a development single-party setup; the circuit has not received an independent audit. Use it to evaluate the disclosure model, then replace the setup through a ceremony before relying on it in production. See [zero-knowledge approval proofs](docs/zk-approval-proofs.md) for the trust model, exact statement, and limitations.

> **New in v0.6.0:** the gate holds under load. The [tool-call gate](docs/tool-gate.md) answers asynchronously (`mode: async`, `wait:`) so a harness with a short HTTP timeout never loses a decision, and a tool call waiting on a human is durable across a restart. Rules constrain arguments (`args:` — glob, regex, `one_of`, `min`/`max`) and never widen on a mismatch. A repeat guard stops an agent that loops on one call from paging you, the operator hears about denials the gate made on its own, `/pending` and `draftcat pending` list every open gate, `/status` shows spend against caps, cost caps enforce the provider's real charge, rate limits back off instead of failing the run — and one Telegram update pump fixes taps that were silently lost while two gates were open at once.
> **New in v0.7.0:** execution decisions now carry their proof. Every tool-gate route is authenticated, each request has a stable action identity and exact policy binding, and an allowed decision becomes an atomic consume-once permit before the side effect runs. Webhook acceptance is durable before HTTP 202 and can be polled after handoff. Versioned receipts bind action, payload, policy, and expiry, with `draftcat receipts list|show|export` for verification-ready JSONL. Ordered `model_policy` rules can deny or send matching model input/output to a human, while `/healthz` and `/readyz` give orchestrators a safe listener contract.
>
> **In v0.5.0:** approvals reach any operator surface via the [`hitl/v0` protocol](docs/hitl-protocol.md) — Microsoft Teams through a Power Automate flow in your own tenant, with no bot and no Azure app registration. Plus a tool-call gate for an agent's MCP/SDK calls (`POST /gate/tool-call`), risk tiers with pre-declared `approval_policy` exemptions, run-correlated audit rows, spend shown at the moment of decision, and `escalate_after` reminders before a gate times out.
> **In v0.6.0:** the gate holds under load. The [tool-call gate](docs/tool-gate.md) answers asynchronously (`mode: async`, `wait:`) so a harness with a short HTTP timeout never loses a decision, and a tool call waiting on a human is durable across a restart. Rules constrain arguments (`args:` - glob, regex, `one_of`, `min`/`max`) and never widen on a mismatch. A repeat guard stops an agent that loops on one call from paging you, the operator hears about denials the gate made on its own, `/pending` and `draftcat pending` list every open gate, `/status` shows spend against caps, cost caps enforce the provider's real charge, rate limits back off instead of failing the run - and one Telegram update pump fixes taps that were silently lost while two gates were open at once.

![Demo](demo.gif)

Expand Down Expand Up @@ -106,17 +106,20 @@ However your agent runs, draftcat sits between it and your customer systems as a
- **Cost budgets** — `per_day_cost` / `per_pipeline_cost` cap spend in money. On OpenRouter the caps are enforced on the charge the provider reports for each call (cached and reasoning tokens included); elsewhere on your configured per-1k rates. The approval prompt shows what the run has spent, and `/status` shows the day against every cap.
- **Human-in-the-loop** — every outbound action requires an explicit operator decision, made live or declared in advance.
- **Any operator channel** — the [`hitl/v0` protocol](docs/hitl-protocol.md) keeps draftcat as the gate and lets an untrusted relay own presentation. Teams runs through a Power Automate flow in your own tenant: no bot, no Azure app registration, no admin consent. Check yours with `draftcat hitl verify <relay-url>`.
- **Tool-call gate** — `POST /gate/tool-call` puts an agent's MCP or SDK calls through the same gate as a pipeline step. Denies by default; the approval binds to a hash of the exact arguments. Rules can constrain the arguments themselves (`args:`) and a mismatch only ever tightens — ask a human, or refuse. A decision that needs a human can be collected asynchronously (`mode: async`, `wait:`, `GET /gate/tool-call/<id>`), and the open gate is durable across a restart. See [`docs/tool-gate.md`](docs/tool-gate.md).
- **Consume-once tool permits** - `POST /gate/tool-call` puts an agent's MCP or SDK call through the same gate as a pipeline step. Bearer authentication covers ask, poll, and consume. A stable `action_id` makes retries idempotent; the binding covers the exact arguments, policy, and expiry; only the first successful `POST /gate/tool-call/<id>/consume` carries `permit: execute`. See [`docs/tool-gate.md`](docs/tool-gate.md).
- **Repeat guard** — inside `repeat_window` an identical tool call (same agent, tool, arguments) gets the gate's remembered answer instead of a new prompt: a denied call stays denied, an in-flight call joins the open prompt, and `max_repeats` stops a looping agent from paging you.
- **Denial notices** — a refusal the gate makes on its own (unlisted tool, argument outside a rule, repeat guard) is reported to the operator channel, one notice per agent, tool and reason per window, so nothing is refused silently.
- **Open gates** — `/pending` on the channel and `draftcat pending` on the host list every approval waiting on a human, pipeline steps and tool calls alike, with how long each has waited and how long it has left.
- **Risk tiers** — steps declare `risk: low | normal | high`, and `approval_policy` can pre-approve a declared class. High risk never qualifies, and each exemption is audited as `policy_approve` with the rule that fired.
- **Escalation** — `escalate_after` re-notifies before a gate times out; `escalate_to` widens who is told, never who may decide.
- **Durable, run-correlated gates** — every gate is written to SQLite before the draft goes out, so an approval in flight survives a restart, and each decision records the run it released.
- **Approver scoping** — `approvers:` on a step narrows who may decide it to a subset of `allowed_users`. Quorum says *how many*; this says *which ones*. It can only narrow, never widen.
- **Model I/O policy** - ordered `model_policy` regex rules check exact input before it reaches the provider and output before it leaves Draftcat. A match can deny or enter the existing human approval gate, and the decision is written as a versioned receipt.
- **Input sanitization** — operator input is scrubbed for prompt-injection patterns before the LLM.
- **Output validation** — AI output is checked against the skill's `output_schema` (field types, numeric `min`/`max`, `enum` membership) and rejected if it doesn't conform.
- **Checked action receipts** — approval decisions can be tied to a payload hash and verified later; see [`docs/action-receipts.md`](docs/action-receipts.md).
- **Checked action receipts** - v2 receipts bind immutable action ID, payload hash, policy digest, validity window, run, and decision. List, inspect, or stream JSONL from SQLite with `draftcat receipts`; see [`docs/action-receipts.md`](docs/action-receipts.md).
- **Durable webhook admission** - Draftcat writes a body-hash-only admission row before returning HTTP 202. The response includes `admission_id` and an authenticated poll URL; unfinished admissions become `interrupted` after restart.
- **Health contract** - `GET /healthz` reports process liveness and `GET /readyz` succeeds only while the SQLite decision store is available.
- **Private approval proofs** — share proof that a direct human approval met quorum without sharing the action, approver, or counts; see [`docs/zk-approval-proofs.md`](docs/zk-approval-proofs.md).
- **Encrypted approval tally** — combine three or more encrypted votes without letting the collector read any individual vote; see [`docs/fhe-vote-tally.md`](docs/fhe-vote-tally.md).
- **Rate limiting** — per-user, per-minute caps on operator interactions.
Expand Down Expand Up @@ -192,6 +195,14 @@ curl -X POST https://draftcat.yourco.eu/hooks/<pipeline> \
```

The POST only **starts** a gated pipeline — the approval step still runs, so inbound can never make the LLM fire a customer-facing action.
Draftcat writes the admission to SQLite before returning `202`:

```json
{"admission_id":"wh_...","status":"accepted","poll":"/hooks/status/wh_..."}
```

Poll that path with the same bearer token. `GET /healthz` is a liveness check;
`GET /readyz` verifies that the decision store is reachable.

## How it works

Expand Down Expand Up @@ -283,6 +294,41 @@ tool_gate:
on_mismatch: deny # outside it: refuse without asking
```

Every gate request uses the webhook bearer token. Send a stable `action_id`,
then consume an allowed binding exactly once before running the side effect:

```bash
curl -X POST http://127.0.0.1:8088/gate/tool-call \
-H "Authorization: Bearer $DRAFTCAT_WEBHOOK_SECRET" \
-H 'Content-Type: application/json' \
-d '{"action_id":"send-invoice-4821","tool":"send_email","args":{"to":"billing@example.com"}}'

curl -X POST http://127.0.0.1:8088/gate/tool-call/send-invoice-4821/consume \
-H "Authorization: Bearer $DRAFTCAT_WEBHOOK_SECRET" \
-H 'Content-Type: application/json' \
-d '{"binding_hash":"sha256:..."}'
```

Model input and output policy is ordered and deterministic. `deny` fails the
LLM call closed; `review` pauses at the configured operator channel:

```yaml
model_policy:
max_preview_chars: 800
rules:
- id: credentials-in-input
phase: input
pattern: '(?i)(api[_ -]?key|password)'
action: review
reason: Credentials require an explicit operator decision.
- id: unsupported-claim
phase: output
roles: [drafter]
pattern: '(?i)guaranteed results'
action: deny
reason: Do not send unsupported guarantees.
```

Skills are YAML prompt templates in `skills/` with an `output_schema` the engine enforces. With `-tags voice`, a `voice:` block configures the webhook receivers, Dograh endpoints, and pre-call lookup — see [docs/voice.md](docs/voice.md).

## Commands
Expand All @@ -293,6 +339,9 @@ draftcat validate [--strict] # lint config + skills
draftcat test <pipeline> # dry-run against fixtures/<pipeline>/ (never touches real APIs)
draftcat runs [pipeline] # recent runs + the approval decisions in each (--json to archive)
draftcat pending # approval gates waiting on a human right now (--json)
draftcat receipts list # approval receipts and verification status (--json)
draftcat receipts show <id> # one versioned receipt
draftcat receipts export # JSONL to stdout (--out path writes mode 0600)
draftcat audit-verify # verify signed approval receipts
draftcat hitl verify <url> # run the hitl/v0 conformance suite against a relay
```
Expand Down Expand Up @@ -328,7 +377,7 @@ curl -X POST http://127.0.0.1:8088/hooks/invoice-due-diligence \
-H "Authorization: Bearer $DRAFTCAT_WEBHOOK_SECRET" -d '{"path": "/inbox/invoice.pdf"}'
```

The body reaches the pipeline as `{{webhook_body}}` / `{{input}}`; bearer auth is constant-time, and a second trigger while the pipeline is running gets `409`. A webhook only *starts* a pipeline — the approval gate still runs, so an inbound request can never make the LLM fire an outbound action.
The body reaches the pipeline as `{{webhook_body}}` / `{{input}}`; bearer auth is constant-time, and a second trigger while the pipeline is running gets `409`. Before `202`, Draftcat stores an admission ID, pipeline, body hash, and status in SQLite. `GET /hooks/status/<admission_id>` returns the authenticated status without retaining the request body. A webhook only *starts* a pipeline - the approval gate still runs, so an inbound request can never make the LLM fire an outbound action.

**Signed requests.** Bind each trigger to its exact body and a timestamp with an HMAC receipt, on top of the bearer token:

Expand Down
27 changes: 22 additions & 5 deletions config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,18 @@ roles:
classifier: gpt-4o-mini
drafter: gpt-4o-mini

# Ordered checks applied before model input is sent and after output returns.
# A matching rule can fail closed or enter the normal operator approval gate.
# model_policy:
# max_preview_chars: 800
# rules:
# - id: credentials-in-input
# phase: input # input | output | both
# roles: [drafter] # optional; empty means every role
# pattern: '(?i)(api[_ -]?key|password)'
# action: review # deny | review
# reason: Credentials require an explicit operator decision.

budgets:
per_step_tokens: 2048
per_pipeline_tokens: 10000
Expand All @@ -67,9 +79,14 @@ observability:
# addr: 127.0.0.1:8088
# secret_env: DRAFTCAT_WEBHOOK_SECRET # required when enabled (bearer token)
# max_body_bytes: 65536

# Tool-call gate for an agent's MCP/SDK calls — POST /gate/tool-call on the
# webhook listener (requires webhook.enabled). Anything not listed is denied.
# require_signature: true # also signs gate POST bodies when enabled
# GET /healthz reports liveness. GET /readyz checks SQLite. Accepted triggers
# return an admission_id for authenticated GET /hooks/status/<admission_id>.

# Tool-call gate for an agent's MCP/SDK calls - POST /gate/tool-call on the
# webhook listener (requires webhook.enabled). Every route requires the webhook
# bearer token. Send a stable action_id and consume an allowed binding once at
# POST /gate/tool-call/<action_id>/consume before executing the side effect.
# See docs/tool-gate.md.
# tool_gate:
# enabled: true
Expand All @@ -82,8 +99,8 @@ observability:
# risk: high
# require_approval: true
# args:
# to: {glob: "*@example.com"} # arguments inside the rule → ask as usual
# on_mismatch: deny # outside it → refuse without asking
# to: {glob: "*@example.com"} # arguments inside the rule: ask as usual
# on_mismatch: deny # outside it: refuse without asking

pipelines:
# Uses skill reference instead of inline prompt
Expand Down
75 changes: 50 additions & 25 deletions docs/action-receipts.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,10 +16,12 @@ a persisted fact that can be inspected later.
5. The approved action executes.
6. A receipt can be exported for audit or incident review.

## Current integrity layer
## Integrity layer

The approval package signs immutable decision fields with HMAC-SHA256:
New decisions use receipt schema v2. The approval package signs immutable
decision fields with HMAC-SHA256:

- receipt, run, and action IDs
- pipeline
- step
- decision time
Expand All @@ -28,10 +30,14 @@ The approval package signs immutable decision fields with HMAC-SHA256:
- payload hash
- quorum requirement
- quorum result
- policy and policy digest
- binding digest
- permit expiry and lifecycle
- nonce

See [`internal/approval/receipt.go`](../internal/approval/receipt.go). If any
covered field changes after signing, verification fails.
covered field changes after signing, verification fails. Existing v1 receipts
continue to verify with their original canonical field set.

## Receipt shape

Expand All @@ -48,40 +54,59 @@ Example:

```json
{
"receipt_id": "act_20260704_001",
"version": 2,
"receipt_id": "rcpt_9d4d...",
"run_id": "run_abc123",
"action_id": "run_abc123:lead_reply:send_email",
"pipeline": "lead_reply",
"step": "send_email",
"action_type": "outbound_message",
"status": "executed",
"proposed_by": "agent",
"approved_by": "operator:12345",
"approved_at": "2026-07-04T09:30:00Z",
"executed_at": "2026-07-04T09:31:00Z",
"schema_version": 1,
"decision": "approve",
"operator_id": 12345,
"decided_at": "2026-09-17T09:30:00Z",
"payload_hash": "sha256:...",
"signature_status": "valid",
"policy_checks": [
{"id": "recipient_allowlist", "status": "pass"},
{"id": "budget_limit", "status": "pass"}
],
"human_decision": {
"decision": "edit_then_approve",
"notes": "Tightened the CTA and removed an unsupported claim."
}
"policy": "human-approval",
"policy_hash": "sha256:...",
"binding_hash": "sha256:...",
"expires_at": "2026-09-17T13:30:00Z",
"lifecycle": "decided",
"quorum_n": 1,
"quorum_got": 1,
"nonce": "...",
"signature": "...",
"verification": "ok"
}
```

## CLI direction
The SQLite row stores hashes and identifiers, not the customer payload.

A small receipt surface should be enough for operators and auditors:
## CLI

List recent receipts across all pipelines, or narrow to one pipeline:

```bash
draftcat receipts list --limit 100
draftcat receipts list --pipeline lead_reply --json
```

Inspect one receipt by its stable ID (legacy rows also accept their numeric row
ID):

```bash
draftcat receipts list --run run_abc123
draftcat receipts show act_20260704_001
draftcat receipts export --format jsonl --out receipts.jsonl
draftcat receipts show rcpt_9d4d...
```

Export newline-delimited JSON in chronological order. Stdout makes it easy to
pipe into an auditor or log shipper; `--out` creates a mode `0600` file:

```bash
draftcat receipts export --pipeline lead_reply > receipts.jsonl
draftcat receipts export --out receipts.jsonl
```

Set `DRAFTCAT_APPROVAL_SECRET` while reading to receive `verification: ok` or
`tampered`. Signed rows without the key report `unverified`; unsigned rows
report `unsigned` explicitly.

## Design rule

Do not let the model decide whether the approval boundary was satisfied. The
Expand Down
Loading
Loading