Conversation
Add security and privacy wg terms for discussion
MatthewKhouzam
left a comment
There was a problem hiding this comment.
These are great, could you add some scope, i feel defining examples with sabotage vs derailment would be good to have, also they are related terms. They probably will help accuracy too.
| category: '', | ||
| aliases: [], | ||
| broaderTerm: null, | ||
| definition: 'An AI agent that operates outside its authorized boundaries, whether due to compromise, misconfiguration, or emergent behavior.', |
There was a problem hiding this comment.
Does intention factor in here? Thinking out loud, it could be operating outside of authorized or intended boundaries.
|
Thanks for pulling the Security & Privacy set together, and for flagging the two entries yourself. One observation on why they need work, and three changes that would close it. The set classifies by attributed cause, and attribution is not observable. Read the five entries together and they separate on whether an external party caused the behaviour. The two flagged entries are not separable at the point of evidence. As written, the cause clause in both definitions is identical:
Nothing in the set states who establishes the attribution. Both definitions are compatible with the affected party acting and reporting on itself, and a reader has no way to tell. A self-reported attribution and an independently witnessed one are not interchangeable for any downstream purpose: a control, a disclosure, or a risk score. Concretely, take two incident records identical in every producer-authored field:
Both satisfy the definition of Suggested changes
The distinction in (2) is already carried in the published evidence-grade ladder E0 to E4 (witnessos, Empire Labs), where each rung names what a claim requires and the establishing party is recorded separately from the claim. Offered as a reference for wording, not as something for this workstream to adopt. Happy to draft the exact entry text as a follow-up PR in the shape of #37, or to supply Records A and B as fixture files if the workstream wants a test case. |
|
As offered on this thread: the two scope notes are drafted and open as a PR against this branch. Each note records, on its own entry, that the attributed cause is a claim rather than a field in a record, and that the establishing party is recorded separately from the claim, so a self-reported attribution is not interchangeable with one established by material the affected operator did not produce. Each note points at the vocabulary already filed in this workstream: #52 (self-reported), #53 (witness-scope) and #54 (three-state reconciliation). One note per entry rather than one shared note, per CONTRIBUTING section 5.C.2. Two lines, scopeNote only. Definitions, category, aliases, broaderTerm, relatedTerms, contrastsWith and workgroups are unchanged. PR: kdruckman#1 If it is easier to apply the two lines directly, take the text and close the PR. If you would rather they land after this merges, say so and I will re-target main. |
This adds the initial Security & Privacy terms agreed upon in the taxonomy and landscape workstream.
Definitions to be reviewed by the Security and Privacy WG before approving.
In particular,
Agent sabotageandAgent mishandling/misuseneed refinement.