Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
57 changes: 57 additions & 0 deletions .agents/council/QUORUM-RESULT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
# 12-Factor AgentOps — re-derivation quorum result

> **3-0 quorum** reached 2026-06-07 by a cross-model NTM council: **Opus 4.8** (claude),
> **Codex GPT-5.5** (codex, xhigh), **Gemini 3.5 Flash High** (antigravity/agy). Verified against
> the artifacts (`proposals/{claude,codex,gemini}.md`), not self-report. Round 1 independent →
> Round 2 converge → Round 3 resolve the single residual (measurement placement; Codex conceded
> to last). Supersedes the v3.2 patch branch — this is a re-derivation (v4.0 candidate).

## Organizing principle
A **closed operational control loop** / dependency chain: each factor earns its position because
later factors can't be trusted without it. Four phases — **Prepare → Bound → Select → Govern** —
are one pass through the loop; Govern (compounding + measurement) feeds back into Prepare.

## The 12 (agreed order + grouping)

### Prepare (I–III) — set up the environment
| # | Factor | Rule |
|---|--------|------|
| I | Context Is Everything | Manage what enters the window like what enters production. |
| II | Track Everything in Git | Durable record (or a committed reference) lives in git. |
| III | One Agent, One Job | Scoped task, fresh context per phase. |

### Bound (IV–VI) — constrain what may act
| # | Factor | Rule |
|---|--------|------|
| IV | **Enforce Least Privilege** *(new)* | Least-privilege envelope; sandbox; untrusted input can't widen it; bound the blast radius. |
| V | Research Before You Build | Understand the integration surface before changing it. |
| VI | Isolate Workers | Concurrent workers share only gated coordination state, never mutable working state. |

### Select (VII–IX) — decide what survives
| # | Factor | Rule |
|---|--------|------|
| VII | Validate Externally | Worker emits claims+evidence; an independent checker writes the binding verdict. |
| VIII | Lock Progress Forward | Validated work ratchets; regression needs an explicit recorded reversal. |
| IX | Extract Learnings | Every non-trivial session yields the work product + provenance-backed lessons (incl. failures). |

### Govern (X–XII) — steer and feed back
| # | Factor | Rule |
|---|--------|------|
| X | Compound Knowledge *(absorbs Harvest Failures)* | Gate, inject, cite, decay learnings — positive and negative — so future sessions start smarter. |
| XI | Supervise Hierarchically | One escalation path per worker; failures move up with evidence, authority flows down. |
| XII | Measure Outcomes | Track fitness toward goals, not activity; the feedback that closes the loop. |

## Changes vs the current set
- **Added IV Enforce Least Privilege** — the security/permissions gap (unanimous round 1).
- **Merged old XII Harvest Failures → X Compound Knowledge** (unanimous; negative knowledge is the same flywheel).
- **Reordered by dependency control loop**, replacing the Foundation/Flow/Knowledge/Scale adoption tiers (the "feels random" complaint).
- **Measure Outcomes → XII** (governance capstone), out of the old Knowledge tier.
- **Groups renamed** to verbs/phases: Prepare / Bound / Select / Govern.
- Accuracy fixes from the v3.2 audit carried in (no pseudo-math, invented numbers, or absolutism).

## Residual (non-blocking)
- Naming: Gemini preferred "Track Durable State"; 2-1 kept "Track Everything in Git."

## Status
Quorum is on the SET + ORDER + GROUPING. Adoption (writing v4.0 across canonical repo + showcase
+ redirects) is a separate operator decision — not yet executed.
57 changes: 57 additions & 0 deletions .agents/council/accuracy-audit-input.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
# Adversarial accuracy audit of the current 12 factors (council input)

> Four independent adversarial auditors (mandate: refute, not affirm) reviewed the current
> factors. Condensed findings below. Use as evidence for the re-derivation — these are the
> known defects of the incumbent set.

## Per-factor problems
- **I Context Is Everything** — "40% utilization rule" is an invented number dressed as if it
follows from Liu et al. (it doesn't); "lost in the middle is how attention works" stated as a
fixed law (it's a training-dependent, shrinking empirical tendency).
- **II Track Everything in Git** — "if it's not in git it didn't happen" self-contradicted by its
own LFS/S3 carve-outs; a `git bisect` misuse; overstates that raw git-merge handles concurrent
structured-data edits.
- **III One Agent, One Job** — invented "50-exchange / 70%" thresholds; conflates exchange count
with token fill; "research-warm agent is the worst to implement" overstated.
- **IV Research Before You Build** — "no exceptions / always" absolutism; unfalsifiable "every
agent eventually does research"; unsupported "simple tasks benefit more than complex."
- **V Validate Externally** — strongest factor. "single-writer / sole writer" was overstated
(worker can author its own gate via TDD); "the moat" claimed here AND in VIII (can't both be).
- **VI Lock Progress Forward** — OBJECTIVE BUG: cited "Factor III (Validation First)" — III is
One-Agent-One-Job, validation is V. "agent quality doesn't matter / filter is perfect" overstated.
- **VII Extract Learnings** — clean. Genuine producer (write) half of the knowledge loop.
- **VIII Compound Knowledge** — signature inequality `retrieval × citation > decay` is
dimensionally incoherent pseudo-math; "decays to zero" false (artifacts persist). Hero factor.
- **IX Measure What Matters** — "dormancy is success" false for continuous-ops/SRE agents;
"harder to game" overstated (goal-redefinition games it). **Likely mis-tiered**: it's a
governance/feedback factor, not a Knowledge factor — wedged into the Knowledge tier to fill 4×3.
- **X Isolate Workers** — "zero shared mutable state" overstated (the tracker + main ARE shared
mutable state by design); true claim is "no shared mutable *working* state."
- **XI Supervise Hierarchically** — OBJECTIVE BUGS: "Further Reading" linked factors from a
different framework (Dispose Gracefully / Orchestrate Declaratively); "root supervisor never
crashes" is false about Erlang/OTP; OTP analogy misapplied (OTP restarts deterministic processes
to a known state; agents are stochastic).
- **XII Harvest Failures as Wisdom** — likely **collapses into VIII** (it's the flywheel applied to
negative knowledge; "prune the search space" is a metaphor, no literal tree). The genuinely
distinct ideas (negative-knowledge value, fresh-agent-on-failure) survive but may not need a slot.

## Set-level findings
1. **Security / permissions / sandboxing / untrusted-input is absent from all 12** — the biggest
gap. A doctrine for operating write-capable agent fleets with no permission/blast-radius/
prompt-injection primitive. MUST be added.
2. **"12" looks padded** to Heroku's number; honest count is ~9-10. XII→VIII; IX is governance.
(Operator decision: keep 12, but the extra slots must be real primitives — security is one.)
3. **Tiers may be post-hoc.** Foundation/Flow/Knowledge/Scale maps onto the old product partition.
The order and grouping feel backfilled, not derived. THIS is the core thing to fix.
4. **Genuine distinctions that DO hold:** III (temporal: one agent over time) vs X (concurrent:
many agents at once); X (independence between peers) vs XI (authority up a chain); VII (write/
capture) vs VIII (read/inject loop). Preserve these axes if you keep these factors.

## Current grouping (the incumbent to beat)
- Foundation (I–III): Context, Track, Scope
- Flow (IV–VI): Research, Validate, Lock
- Knowledge (VII–IX): Extract, Compound, Measure
- Scale (X–XII): Isolate, Supervise, Harvest

The operator finds this grouping unprincipled. Propose a better organizing principle or defend
this one with a real argument.
68 changes: 68 additions & 0 deletions .agents/council/factor-recut-charter.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
# Council Charter — re-derive the 12 Factors (quorum required)

> 3-model council: **Opus 4.8** (claude), **Codex GPT-5.5** (codex), **Gemini 3.5 Flash** (antigravity/agy).
> Goal: agree on what the 12 factors SHOULD be — and crucially their **order** and **grouping** —
> from something closer to first principles. The operator's verdict on the current set: it "feels
> random," the factors aren't ordered properly, and the four tiers feel backfilled to hit twelve.
> Your job is to fix that, with a real organizing principle, and reach **quorum** (all three agree).

## The problem you are solving
The current 12-Factor AgentOps doctrine (read `factors/*.md`) is a *list*, not a *system*. The
adversarial accuracy audit (read `.agents/council/accuracy-audit-input.md`) found: overstatement,
a security/permissions gap, a probable XII→VIII redundancy, IX wedged into the wrong tier, and a
"12" that looks padded to match Heroku's number. Don't just patch it. Re-derive it so the **order
is meaningful** (each factor earns its position) and the **grouping reflects a genuine organizing
principle** (a lifecycle? a control loop? a dependency order? a maturity ladder? argue for one).

## Hard constraints (operator decisions — not up for debate)
1. **Exactly 12 factors.** The number is brand-load-bearing. If you believe the honest count is ~9,
say so in your rationale, but deliver 12 (no padding-for-padding's-sake — if you must reach 12,
the extra slots must be genuinely distinct primitives, e.g. security, cost, observability, HITL).
2. **Security/permissions MUST be represented** — least-privilege, sandboxing, untrusted-input /
prompt-injection, blast-radius. The current set has none; that's the single biggest gap.
3. **Vendor-neutral, runtime-agnostic.** Applies to Claude/Codex/Gemini/Cursor alike.
4. **Each factor genuinely distinct** — no two are the same underlying primitive in different words.
5. **Voice:** punchy aphoristic headline rule + honest body (no invented numbers stated as fact,
no pseudo-math, no absolutism the body contradicts).

## What "good order and grouping" means
- The **order** should tell a story — e.g. you cannot do factor N well without N-1, or the factors
trace a work lifecycle, or they ascend a maturity ladder. Make the through-line explicit.
- The **grouping** must have ONE stated organizing principle. Name it. "Foundation/Flow/Knowledge/
Scale" is the incumbent — improve on it or defend it, but justify the principle, don't assume it.

## Your output (each model, every round)
Write/overwrite your proposal to `.agents/council/proposals/<you>.md` where `<you>` is
`claude`, `codex`, or `gemini`. Structure:

```
# <model> proposal — round <N>
## Organizing principle
<the ONE principle behind the order + grouping, in 2-3 sentences>
## The 12 factors (in order)
| # | Factor name | One-line rule | Group |
(12 rows)
## Grouping
<the groups, the principle, why this order>
## What changed vs the current set & why
<additions (incl. security), merges (e.g. XII→VIII?), reorders, renames — with reasons>
## Open disagreements with the other two proposals
<after round 1: where you differ and your argument>
```

## Process / how to reach quorum
- **Round 1:** propose independently. Do NOT read the others first — derive your own best answer
from the current factors + the audit, then write your file.
- **Round 2+:** read all three files in `.agents/council/proposals/`. Adopt what's better, argue
what's worse, converge. Update your own file each round with a new "round N" version.
- **Quorum = all three proposals agree** on: the set of 12, their order, and the grouping/principle.
Minor wording differences are fine; the structure must match. When you believe quorum is reached,
state `QUORUM: yes` at the top of your file and list the agreed 12. Otherwise `QUORUM: no` + the
remaining disagreement.
- Be intellectually honest, not agreeable-for-the-sake-of-it. A forced false consensus is worse than
a logged disagreement. But genuinely try to converge — find the best answer, not a compromise.

## Inputs to read first
- `factors/*.md` — the current 12 (the thing you're improving).
- `.agents/council/accuracy-audit-input.md` — the adversarial audit's findings.
- `README.md` — current framing (tiers, heritage, the operator model).
Loading
Loading