Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,38 @@ All notable changes to 12-Factor AgentOps will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [3.1.0] - 2026-06-06

### Changed - Whole-System Constitution Alignment

The 12 factors are reframed as the **constitution of the whole agent system**,
lived at altitudes (one agent → a fleet), correcting the earlier implicit
"AgentOps owns I–IX / scale tier is optional" partition. **No factor was
renamed, renumbered, or deleted — names and numbers remain frozen.** This is an
expression + framing change, not a redefinition.

- **Factor V (Validate Externally)** now leads with the claims-vs-verdicts /
single-writer rule: the worker emits claims plus evidence; an independent
checker is the sole writer of the binding verdict. Promoted from a buried
sentence to the spine; expressed at two altitudes (worker honesty / factory
authority). This is the moat.
- **Factor XII (Harvest Failures)** rewritten from "failed attempts are data"
(a restatement of VII) to its distinct mechanism: failures as routing hints
that prune the next agent's search space, plus fresh-agent-on-failure.
- **Factors VII ↔ VIII** boundary sharpened (capture/write vs inject/read); the
flywheel diagram no longer owns EXTRACT (shown as an inbound handoff from VII).
- **Factors III ↔ X** disambiguated (temporal vs concurrent "fresh context");
fungible/disposable workers named (the bead is the durable unit).
- **Factors X ↔ XI** framed as independence (peers) vs authority (chain).
- **Scale tier (X–XII)** reframed from "Advanced, Optional / skip if solo" to
the **factory altitude** — the same factors at fleet scale, lived in miniature
solo. The progressive adoption path is preserved; the "optional / you lose
nothing" claim is retired across README, docs, and GOALS.
- Added the **Core + Skin** rule and a **Doctrine Stability** (no-renumber /
no-delete) guard to the README.
- Aligned all downstream docs (00-SUMMARY, principles/*, explanation/*,
reference/*) and the showcase site to the corrected framing.

## [3.0.0] - 2026-02-15

### Changed - Operational Discipline Pivot
Expand Down
8 changes: 4 additions & 4 deletions GOALS.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -59,13 +59,13 @@ goals:
grep -qi 'flywheel\|compounding equation\|retrieval rate\|decay rate' factors/08-*.md
weight: 9

# Scale factors (X-XII) marked as optional/advanced
- id: scale-factors-optional
description: "Factors 10-12 are explicitly marked as optional or advanced tier"
# Scale factors (X-XII) framed as the factory altitude (same factors at fleet scale)
- id: scale-factors-factory-altitude
description: "Factors 10-12 are framed as the factory altitude (same factors at fleet scale), not an optional/skippable tier"
check: |
count=0
for f in factors/10-*.md factors/11-*.md factors/12-*.md; do
[ -f "$f" ] && grep -qi 'optional\|advanced\|not prerequisite' "$f" && count=$((count+1))
[ -f "$f" ] && grep -qi 'factory altitude\|fleet scale\|grow into\|in miniature' "$f" && count=$((count+1))
done
[ "$count" -ge 3 ]
weight: 8
Expand Down
25 changes: 17 additions & 8 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ primitives, and flows that compound.

<!-- Build & Status -->
[![CI](https://img.shields.io/github/actions/workflow/status/boshu2/12-factor-agentops/validate-factors.yml?label=CI)](https://github.com/boshu2/12-factor-agentops/actions)
[![Version](https://img.shields.io/badge/Version-3.0.0-blue.svg)](https://github.com/boshu2/12-factor-agentops/releases)
[![Version](https://img.shields.io/badge/Version-3.1.0-blue.svg)](https://github.com/boshu2/12-factor-agentops/releases)

<!-- Technology -->
[![12 Factors](https://img.shields.io/badge/Factors-12-00CED1.svg)](factors/)
Expand Down Expand Up @@ -99,7 +99,7 @@ In Cursor, add to `.cursorrules`. In Codex, add to `AGENTS.md`. The mechanism va

## The 12 Factors

Twelve vendor-neutral principles organized in four tiers. Start at the top. Each tier builds on the previous one. You can stop at any tier and keep the value.
Twelve vendor-neutral principles. No factor belongs to one tool or one tier — each is the same rule lived at whatever altitude you're working, from a single agent on one task to a fleet running many. The four tiers below are an **on-ramp, not a partition**: start at the top, stop at any tier and keep the value. But "stopping" means you haven't *automated* the later factors yet — it doesn't mean they no longer apply. A solo developer still lives isolation and supervision; they just live them with a worktree and their own judgment instead of a control plane.

### Foundation (I–III) — Start Here

Expand All @@ -120,7 +120,7 @@ How work flows through agents. The discipline that separates "prompting and hopi
| # | Factor | The Rule |
|---|--------|----------|
| **[IV](./factors/04-research-before-you-build.md)** | **Research Before You Build** | Understand the problem space before generating a single line of code. |
| **[V](./factors/05-validate-externally.md)** | **Validate Externally** | No agent grades its own work. Ever. |
| **[V](./factors/05-validate-externally.md)** | **Validate Externally** | The worker emits claims plus evidence; an independent checker is the sole writer of the binding verdict. No agent grades its own work. Ever. |
| **[VI](./factors/06-lock-progress-forward.md)** | **Lock Progress Forward** | Once work passes validation, it ratchets — it cannot regress. |

**Without tooling:** Research before implementing. Have a different session (or human) review the work. Commit validated work to protected branches.
Expand All @@ -142,20 +142,28 @@ better models don't replace durable bookkeeping.

**Without tooling:** Manually update `learnings.md` after each session. Review it weekly and prune stale entries. It's tedious but it works. The AgentOps plugin automates this — but the principle is portable.

### Scale (X–XII) — Advanced, Optional
### Scale (X–XII) — The Factory Altitude

Multi-agent orchestration patterns. **Skip this entire tier if you work solo.** You lose nothing. These patterns apply when you're running parallel agents on complex projects.
The same factors at fleet scale. Working solo, you live these at a small altitude — a git worktree is isolation, your own judgment is supervision, your `learnings.md` is failure harvesting. Running parallel agents on complex projects, the same three rules need real machinery. You don't *adopt* this tier so much as *grow into* its altitude; the rules were always there.

| # | Factor | The Rule |
|---|--------|----------|
| **[X](./factors/10-isolate-workers.md)** | **Isolate Workers** | Each worker gets its own workspace, its own context, and zero shared mutable state. |
| **[XI](./factors/11-supervise-hierarchically.md)** | **Supervise Hierarchically** | Escalation flows up, never sideways. |
| **[XII](./factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Failed attempts are data. Extract and index them with the same rigor as successes. |
| **[XII](./factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Turn failed attempts into routing hints that prune the next agent's search; on repeat failure, hand the context to a fresh agent. |

**Without tooling:** Use git worktrees for parallel work. Designate one person (or agent) as coordinator. Document what doesn't work alongside what does.

---

## Core and Skin

Every layer of an agent system is a universal **core** plus a removable **skin**, and the two are never conflated. The core is the invariant — these twelve factors, the operator model below them. The skin is house style: your naming, your personas, your rituals, the story you tell yourself about the work. The skin is never imposed. You adopt the constitution without anyone's mythology, and you dress it in your own. That separation is what makes the doctrine portable across teams, tools, and vendors: take the rules, leave the costume.

## Doctrine Stability

The twelve names and numbers are frozen at v3.0.0. Corrections are **expression-only** — sharpen the prose, fix a diagram, re-cut a boundary that reads ambiguously. A factor may be rewritten, and a soft one may be reduced to an emphasis that points at its neighbor, but **no factor is ever deleted and the set is never renumbered.** The number is load-bearing (URLs, the badge, every inbound link, twenty years of "twelve-factor" recognition). When two factors feel redundant, the fix is to make the distinction legible, not to merge the slots.

## The Operator Model Underneath the Factors

The twelve factors are the public operating rules behind bookkeeping,
Expand Down Expand Up @@ -203,10 +211,10 @@ Quickstart (5 min) → learnings.md file, zero tooling
Foundation (I-III) → Context discipline, git tracking, fresh sessions
Flow (IV-VI) → Research, validation, ratcheting
Knowledge (VII-IX) → Extraction, compounding, measurement
Scale (X-XII) → Multi-agent isolation, supervision, failure harvesting (OPTIONAL)
Scale (X-XII) → Multi-agent isolation, supervision, failure harvesting (grow into it)
```

**Key principle:** You can stop at any level and keep the value. Each level justifies the next, but none requires it.
**Key principle:** You can stop adopting at any level and keep the value. Each level justifies the next, but none requires it. Stopping means you haven't automated the higher factors yet — not that they stopped applying.

**When to level up:**
- **Quickstart → Foundation:** When your `learnings.md` gets unwieldy or you notice repeated context problems
Expand Down Expand Up @@ -268,3 +276,4 @@ The factors evolve through production validation and community feedback.
- **v1.0** (2025-01-27): Initial twelve factors — coding agent validation focus
- **v2.0** (2025-12-27): Production implementation patterns added
- **v3.0** (2026-02-15): Pivot to full operational discipline. Factors rewritten. Adoption model inverted (results-first, not manifesto-first). Knowledge compounding as hero differentiator. Scale factors marked optional.
- **v3.1** (2026-06-06): Whole-system constitution alignment. The 12 are reframed as one constitution lived at altitudes (one agent → a fleet), not a product partition. Factor V leads with the claims-vs-verdicts / single-writer moat; Factor XII rewritten to routing-hints + fresh-agent-on-failure; VII↔VIII, III↔X, X↔XI boundaries sharpened; Scale tier reframed from "optional" to the factory altitude. No factor renamed, renumbered, or deleted.
2 changes: 1 addition & 1 deletion VERSION
Original file line number Diff line number Diff line change
@@ -1 +1 @@
v2.0.0
v3.1.0
12 changes: 7 additions & 5 deletions docs/00-SUMMARY.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@ from a reliable operating model.
| # | Factor | One-Line Rule |
|---|--------|---------------|
| **[IV](../factors/04-research-before-you-build.md)** | **Research Before You Build** | Understand the problem space before generating a single line of code. |
| **[V](../factors/05-validate-externally.md)** | **Validate Externally** | No agent grades its own work -- tests, linters, separate sessions, or humans validate. |
| **[V](../factors/05-validate-externally.md)** | **Validate Externally** | The worker reports evidence; an independent checker writes the binding verdict. No agent grades its own work. |
| **[VI](../factors/06-lock-progress-forward.md)** | **Lock Progress Forward** | Once work passes validation, it ratchets forward and cannot regress. |

### Knowledge (VII--IX) -- Where Compounding Kicks In
Expand All @@ -72,16 +72,18 @@ getting measurably smarter over time.
> repeat your mistakes. Knowledge compounding is the one capability no amount of
> model improvement replaces.

### Scale (X--XII) -- Advanced, Optional
### Scale (X--XII) -- The Factory Altitude

Multi-agent orchestration. Skip this tier entirely if you work solo. These
patterns apply only when running parallel agents on complex projects.
The same factors at fleet scale. Working solo, you live them in miniature -- a
git worktree is isolation, your own judgment is supervision, your `learnings.md`
is failure harvesting -- so you grow into the altitude rather than skipping the
factors.

| # | Factor | One-Line Rule |
|---|--------|---------------|
| **[X](../factors/10-isolate-workers.md)** | **Isolate Workers** | Each worker gets its own workspace, context, and zero shared mutable state. |
| **[XI](../factors/11-supervise-hierarchically.md)** | **Supervise Hierarchically** | Escalation flows up, never sideways -- one coordinator dispatches, workers execute. |
| **[XII](../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Failed attempts are data -- extract and index them with the same rigor as successes. |
| **[XII](../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Turn dead ends into routing hints that prune the next agent's search. |

---

Expand Down
4 changes: 3 additions & 1 deletion docs/explanation/from-theory-to-production.md
Original file line number Diff line number Diff line change
Expand Up @@ -338,7 +338,9 @@ Each factor maps to concrete implementation patterns from Houston, Fractal, and

---

### Scale Tier (X-XII, optional): Multi-Agent Operations
### Scale Tier (X-XII): The Factory Altitude

The same three factors at fleet scale. Working solo you live them in miniature — a git worktree is isolation, your own judgment is supervision, your `learnings.md` is failure harvesting. Running parallel agents on complex projects, the same rules need real machinery. You grow into this altitude; you don't skip the factors.

**Factor X: Isolate Workers**
- **Philosophy:** Independent, focused execution
Expand Down
4 changes: 2 additions & 2 deletions docs/explanation/vibe-coding-integration.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,11 +68,11 @@ This is where vibe coding transforms from a series of isolated sessions into a c

---

### Tier 4: Scale (Factors X-XII, Optional)
### Tier 4: Scale (Factors X-XII) — The Factory Altitude

**Making vibe coding work across teams and complex systems.**

These factors are optional for solo developers but become essential when vibe coding scales beyond one person and one agent.
These are the same three factors at fleet scale. A solo developer already lives them in miniature — a git worktree is isolation, your own judgment is supervision, the note you write after a failed session is failure harvesting. When vibe coding scales beyond one person and one agent, the same rules need real machinery. You grow into this altitude; you don't skip the factors.

| Factor | What It Solves |
|--------|---------------|
Expand Down
11 changes: 7 additions & 4 deletions docs/principles/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ Twelve vendor-neutral principles organized in four tiers. Each tier builds on th
| # | Factor | The Rule |
|---|--------|----------|
| **[IV](../../factors/04-research-before-you-build.md)** | **Research Before You Build** | Understand the problem space before generating a single line of code. |
| **[V](../../factors/05-validate-externally.md)** | **Validate Externally** | No agent grades its own work. Ever. |
| **[V](../../factors/05-validate-externally.md)** | **Validate Externally** | The worker reports evidence; an independent checker writes the binding verdict. No agent grades its own work. |
| **[VI](../../factors/06-lock-progress-forward.md)** | **Lock Progress Forward** | Once work passes validation, it ratchets -- it cannot regress. |

### Knowledge (VII-IX) -- Where compounding kicks in
Expand All @@ -36,15 +36,18 @@ Twelve vendor-neutral principles organized in four tiers. Each tier builds on th

**Factor VIII is the hero.** It is the knowledge flywheel: extract learnings, gate for quality, inject into future sessions, measure retrieval, let stale knowledge decay. This is the differentiator that no amount of model improvement replaces -- better models with amnesia still repeat your mistakes.

### Scale (X-XII) -- Advanced, optional
### Scale (X-XII) -- The Factory Altitude

Multi-agent orchestration patterns. Skip this tier if you work solo.
The same factors at fleet scale. Working solo, you live them in miniature -- a
git worktree is isolation, your own judgment is supervision, your `learnings.md`
is failure harvesting -- so you grow into the altitude rather than skipping the
factors.

| # | Factor | The Rule |
|---|--------|----------|
| **[X](../../factors/10-isolate-workers.md)** | **Isolate Workers** | Each worker gets its own workspace, its own context, and zero shared mutable state. |
| **[XI](../../factors/11-supervise-hierarchically.md)** | **Supervise Hierarchically** | Escalation flows up, never sideways. |
| **[XII](../../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Failed attempts are data. Extract and index them with the same rigor as successes. |
| **[XII](../../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Turn dead ends into routing hints that prune the next agent's search. |

---

Expand Down
16 changes: 8 additions & 8 deletions docs/principles/comparison-table.md
Original file line number Diff line number Diff line change
Expand Up @@ -140,10 +140,10 @@ This compression does not replace the factors. It explains the mechanism beneath
- *Key Practice*: Simple APIs for agent lifecycle management

**12-Factor AgentOps (V: Validate Externally)**
- *Evolution*: No agent grades its own work
- *Evolution*: The worker reports evidence; an independent checker writes the binding verdict. No agent grades its own work.
- *Why Different*: Agents are confident but not reliable -- they cannot objectively evaluate their own output
- *Key Practice*: External validation via tests, linters, a different agent, or human review
- *Unique Aspect*: Zero-trust principle applied to cognition -- validate the output, not the source
- *Key Practice*: The worker emits claims plus evidence; an independent checker -- tests, linters, a different agent, or a human reviewer -- is the sole writer of the binding verdict
- *Unique Aspect*: Zero-trust principle applied to cognition -- claims and verdicts are written by different parties, so a worker cannot launder its own confidence into a trusted result (the single-writer moat)

---

Expand Down Expand Up @@ -242,7 +242,7 @@ This compression does not replace the factors. It explains the mechanism beneath
- *Evolution*: Each worker gets its own workspace, its own context, and zero shared mutable state
- *Why Different*: Parallel agents sharing state create cascading conflicts
- *Key Practice*: Git worktrees, separate context windows, independent validation
- *Unique Aspect*: Scale tier -- skip if working solo; essential for multi-agent orchestration
- *Unique Aspect*: Scale tier (factory altitude) -- solo you live it every time you spin up a second worktree; structural for multi-agent orchestration

---

Expand All @@ -262,7 +262,7 @@ This compression does not replace the factors. It explains the mechanism beneath
- *Evolution*: Escalation flows up, never sideways
- *Why Different*: Multi-agent systems without hierarchy devolve into circular coordination
- *Key Practice*: Coordinator agents delegate and escalate; worker agents execute and report
- *Unique Aspect*: Scale tier -- skip if working solo; prevents coordination chaos in multi-agent setups
- *Unique Aspect*: Scale tier (factory altitude) -- solo you are the supervisor and break ties yourself; structural when the escalation path no longer fits in one head

---

Expand All @@ -279,10 +279,10 @@ This compression does not replace the factors. It explains the mechanism beneath
- *Key Practice*: Human contact is a first-class operation, not exception

**12-Factor AgentOps (XII: Harvest Failures as Wisdom)**
- *Evolution*: Failed attempts are data -- extract and index them with the same rigor as successes
- *Evolution*: Turn dead ends into routing hints that prune the next agent's search space
- *Why Different*: Failures contain the highest-value learnings but are typically discarded
- *Key Practice*: Document what did not work and why; index failure patterns for future avoidance
- *Unique Aspect*: Scale tier -- feeds directly into Factor VII (Extract Learnings) and Factor VIII (Compound Knowledge)
- *Key Practice*: Index negative knowledge for retrieval at decision time, and hand a stuck worker's failure trace to a fresh agent rather than looping the saturated one
- *Unique Aspect*: Scale tier (factory altitude) -- negative knowledge that prunes the search, distinct from Factor VII's generic capture; feeds Factor VIII (Compound Knowledge)

---

Expand Down
Loading
Loading