You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Show what the demos and the quickstart actually print (#25)
Both pages describe outcomes in prose and never show the output. For
pages whose whole pitch is "don't take the spec on trust, run it", the
output is the argument.
I ran the published quickstart end to end against cmcp-runtime 0.4.0 and
demo 6 against weight-custody-manifest 0.25.0. Both do exactly what the
pages say, so nothing here corrects a claim about behaviour. What
changes is that a reader can now see it before they install anything.
Quickstart:
- Step 4 shows the response body the runtime really returns, replacing a
one-line prose summary of it.
- That body carries a call_id, and the same id lands in the audit chain,
which is the thread from the block in step 4 to the signed record in
step 5. The page never mentioned it, and it is the point.
- Step 5 shows the CRYPTO-001 advisory the CLI prints before the checks.
It is not a failure, but it is the first thing on screen and the page
did not prepare anyone for it.
- Says which version the page was last verified against, and that a step
not behaving as written is a bug worth reporting.
Demos:
- The governance section said "these five govern the tool boundary" above
six cards. Demo 10 is in that section and governs model calls, not the
tool boundary.
- The card times sum to thirteen minutes, not twelve. Fixed in the hero
and in the meta, OG and Twitter descriptions.
- Dropped export CMCP_BEARER_TOKEN from the quick start: demo.py sets it,
as the repo README says. Added --no-pause, without which the runner
waits for a keypress before every demo.
- Added demo 6's verbatim output.
Claude-Session: https://claude.ai/code/session_013EQx4N5BzTQbY8kvXUsdkY
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
<title>Model Weight Protection and Agent Governance Demos | AgenTrust</title>
7
-
<metaname="description" content="Ten runnable demos for securing AI model weights and governing agent tool and model calls. Run them on your laptop in about twelve minutes, no confidential-computing hardware required.">
7
+
<metaname="description" content="Ten runnable demos for securing AI model weights and governing agent tool and model calls. Run them on your laptop in about thirteen minutes, no confidential-computing hardware required.">
<metaproperty="og:title" content="Model Weight Protection and Agent Governance Demos">
15
-
<metaproperty="og:description" content="Ten runnable demos for model-weight custody and governed agent tool and model calls. About twelve minutes on your laptop, no special hardware.">
15
+
<metaproperty="og:description" content="Ten runnable demos for model-weight custody and governed agent tool and model calls. About thirteen minutes on your laptop, no special hardware.">
<metaname="twitter:title" content="Model Weight Protection and Agent Governance Demos">
26
-
<metaname="twitter:description" content="Model-weight custody plus governed agent tool and model calls. Ten runnable demos, about twelve minutes, no special hardware.">
26
+
<metaname="twitter:description" content="Model-weight custody plus governed agent tool and model calls. Ten runnable demos, about thirteen minutes, no special hardware.">
<h1>Don't take the spec<br>on trust. <em>Run it.</em></h1>
52
-
<pclass="sub">Ten demos, about twelve minutes end to end. Four cover custody of model weights. Five govern what an agent does at the tool boundary. One governs model calls through an OpenAI-compatible endpoint.</p>
52
+
<pclass="sub">Ten demos, about thirteen minutes end to end. Four cover custody of model weights. Five govern what an agent does at the tool boundary. One governs model calls through an OpenAI-compatible endpoint.</p>
53
53
<div><spanclass="status">Software mode · <code>CMCP_DEV_MODE=1</code> · no special hardware</span></div>
54
54
</div></div>
55
55
@@ -63,11 +63,32 @@ <h2>Clone it and run all ten</h2>
<pre>git clone https://github.com/agentrust-io/demos && cd demos
65
65
pip install -r requirements.txt
66
-
export CMCP_BEARER_TOKEN=demo-token
67
66
python demo.py <spanclass="c"># all ten, pausing before each</span>
67
+
python demo.py --no-pause <spanclass="c"># straight through, no prompts</span>
68
68
python demo.py 6 <spanclass="c"># just demo 6</span></pre>
69
69
</div>
70
-
<p>The requirements install cMCP for demos 1 to 5, Weight Custody Manifest for demos 6 to 9, and the OpenAI client for demo 10. Source: <ahref="https://github.com/agentrust-io/demos">github.com/agentrust-io/demos</a>.</p>
70
+
<p>The requirements install cMCP for demos 1 to 5, Weight Custody Manifest for demos 6 to 9, and the OpenAI client for demo 10. <code>demo.py</code> sets dev mode and the bearer token for you, so there is nothing to export. Source: <ahref="https://github.com/agentrust-io/demos">github.com/agentrust-io/demos</a>.</p>
71
+
72
+
<pclass="meta" style="margin-top:2.25rem;">What demo 6 actually prints, verbatim from a run on <code>weight-custody-manifest 0.25.0</code></p>
<divclass="dim">Weight Custody Manifest: possession is not provenance.<br>Real WCM code with a software (mock) attestation provider, no hardware.</div>
77
+
<divstyle="margin-top:0.9rem;">1. The builder signs a manifest binding the exact weight hash</div>
<div>matches manifest : <spanclass="alert">False -> REFUSE to load</span></div>
85
+
<divclass="dim">no human reads 2.8T parameters; the hash does the reading.</div>
86
+
<divstyle="margin-top:0.9rem;">4. The fine-tune is the real IP: lineage back to the signed base</div>
87
+
<div>lineage verified : <spanclass="ok">True</span> depth 1 root is a base: <spanclass="ok">True</span></div>
88
+
<divclass="dim" style="margin-top:0.9rem;">honest scope : accountability-grade against an operator who physically<br> owns the silicon (see TEE.fail), not silicon-proof custody.</div>
89
+
</div>
90
+
</div>
91
+
<pclass="meta">Hashes are truncated here for width; the run prints them in full. Every demo ends with a scope statement like that last one.</p>
71
92
</section>
72
93
73
94
<sectionid="weights">
@@ -133,7 +154,7 @@ <h2>Securing model weights</h2>
133
154
<sectionid="governance">
134
155
<spanclass="label">Agent governance</span>
135
156
<h2>Governing what an agent does</h2>
136
-
<pclass="lead-serif">Demos 6 to 9 protect the weights. These five govern the tool boundary: what the agent is allowed to call, under which workflow, with what compliance attributes, and what evidence survives afterwards. Cedar policy is enforced on every call and each session closes with a signed TRACE claim.</p>
157
+
<pclass="lead-serif">Demos 6 to 9 protect the weights. These six put the policy at the boundary the agent has to cross: five at the tool call, and demo 10 at the model call. What is it allowed to invoke, under which workflow, with what compliance attributes, and what evidence survives afterwards. Cedar is enforced on every call and each session closes with a signed TRACE claim.</p>
Copy file name to clipboardExpand all lines: quickstart/index.html
+16-4Lines changed: 16 additions & 4 deletions
Original file line number
Diff line number
Diff line change
@@ -135,6 +135,7 @@ <h2>The quickstart</h2>
135
135
136
136
<divclass="callout">
137
137
<p><strong>Before you start.</strong> Python 3.11+ · pip · macOS or Linux · two terminal windows · about ten minutes · no special hardware.</p>
138
+
<pstyle="margin-top:0.6rem;">Every command and every output on this page was last run end to end against <code>cmcp-runtime 0.4.0</code> on 20 August 2026. If a step does not do what it says here, that is a bug and worth <ahref="https://github.com/agentrust-io/cmcp/issues" target="_blank" rel="noopener">reporting</a>.</p>
138
139
</div>
139
140
140
141
<divclass="steps">
@@ -271,10 +272,17 @@ <h3>Fire a bad action, watch it get blocked</h3>
271
272
}'</pre>
272
273
</div>
273
274
<pclass="hint"><code>workflow_id</code> is the only field the runtime reads out of <code>_cmcp</code>. The session id is a label for your own logs: the runtime mints its own session id, which is why step 5 looks it up instead of assuming it.</p>
275
+
<divclass="term" style="margin-top:1.15rem;">
276
+
<divclass="term-bar"><spanclass="dot"></span><spanclass="dot"></span><spanclass="dot"></span><spanclass="term-title">what the runtime returns</span></div>
<divstyle="margin-top:0.7rem;">{"jsonrpc":"2.0","error":{"code":-32000,<br> "message":"Request denied by policy",<br> "data":{"error_code":<spanclass="alert">"POLICY_DENY"</span>,<br> "call_id":"51da9a46-149f-40c4-b83f-82d48fd654bd"}},"id":2}</div>
280
+
</div>
281
+
</div>
274
282
<divclass="verdict">
275
-
<divclass="flag"><spanclass="check">✓</span> What you'll see — 403 Forbidden</div>
276
-
<p>Your policy stops a PII record from leaving on a tool call, <strong>before it reaches Salesforce</strong>, decided by the rule you wrote, enforced where the agent can't tamper with it. That's the barrier most teams can't cross today: shipping an agent you can actually <strong>prove</strong> is governed.</p>
<divclass="flag"><spanclass="check">✓</span> What just happened</div>
284
+
<p>Your policy stopped a PII record from leaving on a tool call, <strong>before it reached Salesforce</strong>, decided by the rule you wrote, enforced where the agent can't tamper with it. That's the barrier most teams can't cross today: shipping an agent you can actually <strong>prove</strong> is governed.</p>
285
+
<p>Keep an eye on that <code>call_id</code>. The same id lands in the audit chain, so the deny you just watched is the deny you can hand to someone else in step 5. A refusal nobody can check afterwards is just a log line.</p>
278
286
</div>
279
287
</div>
280
288
</div>
@@ -291,7 +299,11 @@ <h3>Walk away with proof</h3>
<pclass="hint">Expected output in dev mode. The <code>CRYPTO-001</code> line comes first and is an advisory, not a failure: it is the CLI saying up front that a software-mode key binding proves nothing about hardware.</p>
0 commit comments