diff --git a/CHANGELOG.md b/CHANGELOG.md index 399d3c3..71e746f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,38 @@ All notable changes to 12-Factor AgentOps will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [3.1.0] - 2026-06-06 + +### Changed - Whole-System Constitution Alignment + +The 12 factors are reframed as the **constitution of the whole agent system**, +lived at altitudes (one agent → a fleet), correcting the earlier implicit +"AgentOps owns I–IX / scale tier is optional" partition. **No factor was +renamed, renumbered, or deleted — names and numbers remain frozen.** This is an +expression + framing change, not a redefinition. + +- **Factor V (Validate Externally)** now leads with the claims-vs-verdicts / + single-writer rule: the worker emits claims plus evidence; an independent + checker is the sole writer of the binding verdict. Promoted from a buried + sentence to the spine; expressed at two altitudes (worker honesty / factory + authority). This is the moat. +- **Factor XII (Harvest Failures)** rewritten from "failed attempts are data" + (a restatement of VII) to its distinct mechanism: failures as routing hints + that prune the next agent's search space, plus fresh-agent-on-failure. +- **Factors VII ↔ VIII** boundary sharpened (capture/write vs inject/read); the + flywheel diagram no longer owns EXTRACT (shown as an inbound handoff from VII). +- **Factors III ↔ X** disambiguated (temporal vs concurrent "fresh context"); + fungible/disposable workers named (the bead is the durable unit). +- **Factors X ↔ XI** framed as independence (peers) vs authority (chain). +- **Scale tier (X–XII)** reframed from "Advanced, Optional / skip if solo" to + the **factory altitude** — the same factors at fleet scale, lived in miniature + solo. The progressive adoption path is preserved; the "optional / you lose + nothing" claim is retired across README, docs, and GOALS. +- Added the **Core + Skin** rule and a **Doctrine Stability** (no-renumber / + no-delete) guard to the README. +- Aligned all downstream docs (00-SUMMARY, principles/*, explanation/*, + reference/*) and the showcase site to the corrected framing. + ## [3.0.0] - 2026-02-15 ### Changed - Operational Discipline Pivot diff --git a/GOALS.yaml b/GOALS.yaml index 73c673a..6c84a0d 100644 --- a/GOALS.yaml +++ b/GOALS.yaml @@ -59,13 +59,13 @@ goals: grep -qi 'flywheel\|compounding equation\|retrieval rate\|decay rate' factors/08-*.md weight: 9 - # Scale factors (X-XII) marked as optional/advanced - - id: scale-factors-optional - description: "Factors 10-12 are explicitly marked as optional or advanced tier" + # Scale factors (X-XII) framed as the factory altitude (same factors at fleet scale) + - id: scale-factors-factory-altitude + description: "Factors 10-12 are framed as the factory altitude (same factors at fleet scale), not an optional/skippable tier" check: | count=0 for f in factors/10-*.md factors/11-*.md factors/12-*.md; do - [ -f "$f" ] && grep -qi 'optional\|advanced\|not prerequisite' "$f" && count=$((count+1)) + [ -f "$f" ] && grep -qi 'factory altitude\|fleet scale\|grow into\|in miniature' "$f" && count=$((count+1)) done [ "$count" -ge 3 ] weight: 8 diff --git a/README.md b/README.md index 203b30c..0824299 100644 --- a/README.md +++ b/README.md @@ -9,7 +9,7 @@ primitives, and flows that compound. [![CI](https://img.shields.io/github/actions/workflow/status/boshu2/12-factor-agentops/validate-factors.yml?label=CI)](https://github.com/boshu2/12-factor-agentops/actions) -[![Version](https://img.shields.io/badge/Version-3.0.0-blue.svg)](https://github.com/boshu2/12-factor-agentops/releases) +[![Version](https://img.shields.io/badge/Version-3.1.0-blue.svg)](https://github.com/boshu2/12-factor-agentops/releases) [![12 Factors](https://img.shields.io/badge/Factors-12-00CED1.svg)](factors/) @@ -99,7 +99,7 @@ In Cursor, add to `.cursorrules`. In Codex, add to `AGENTS.md`. The mechanism va ## The 12 Factors -Twelve vendor-neutral principles organized in four tiers. Start at the top. Each tier builds on the previous one. You can stop at any tier and keep the value. +Twelve vendor-neutral principles. No factor belongs to one tool or one tier — each is the same rule lived at whatever altitude you're working, from a single agent on one task to a fleet running many. The four tiers below are an **on-ramp, not a partition**: start at the top, stop at any tier and keep the value. But "stopping" means you haven't *automated* the later factors yet — it doesn't mean they no longer apply. A solo developer still lives isolation and supervision; they just live them with a worktree and their own judgment instead of a control plane. ### Foundation (I–III) — Start Here @@ -120,7 +120,7 @@ How work flows through agents. The discipline that separates "prompting and hopi | # | Factor | The Rule | |---|--------|----------| | **[IV](./factors/04-research-before-you-build.md)** | **Research Before You Build** | Understand the problem space before generating a single line of code. | -| **[V](./factors/05-validate-externally.md)** | **Validate Externally** | No agent grades its own work. Ever. | +| **[V](./factors/05-validate-externally.md)** | **Validate Externally** | The worker emits claims plus evidence; an independent checker is the sole writer of the binding verdict. No agent grades its own work. Ever. | | **[VI](./factors/06-lock-progress-forward.md)** | **Lock Progress Forward** | Once work passes validation, it ratchets — it cannot regress. | **Without tooling:** Research before implementing. Have a different session (or human) review the work. Commit validated work to protected branches. @@ -142,20 +142,28 @@ better models don't replace durable bookkeeping. **Without tooling:** Manually update `learnings.md` after each session. Review it weekly and prune stale entries. It's tedious but it works. The AgentOps plugin automates this — but the principle is portable. -### Scale (X–XII) — Advanced, Optional +### Scale (X–XII) — The Factory Altitude -Multi-agent orchestration patterns. **Skip this entire tier if you work solo.** You lose nothing. These patterns apply when you're running parallel agents on complex projects. +The same factors at fleet scale. Working solo, you live these at a small altitude — a git worktree is isolation, your own judgment is supervision, your `learnings.md` is failure harvesting. Running parallel agents on complex projects, the same three rules need real machinery. You don't *adopt* this tier so much as *grow into* its altitude; the rules were always there. | # | Factor | The Rule | |---|--------|----------| | **[X](./factors/10-isolate-workers.md)** | **Isolate Workers** | Each worker gets its own workspace, its own context, and zero shared mutable state. | | **[XI](./factors/11-supervise-hierarchically.md)** | **Supervise Hierarchically** | Escalation flows up, never sideways. | -| **[XII](./factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Failed attempts are data. Extract and index them with the same rigor as successes. | +| **[XII](./factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Turn failed attempts into routing hints that prune the next agent's search; on repeat failure, hand the context to a fresh agent. | **Without tooling:** Use git worktrees for parallel work. Designate one person (or agent) as coordinator. Document what doesn't work alongside what does. --- +## Core and Skin + +Every layer of an agent system is a universal **core** plus a removable **skin**, and the two are never conflated. The core is the invariant — these twelve factors, the operator model below them. The skin is house style: your naming, your personas, your rituals, the story you tell yourself about the work. The skin is never imposed. You adopt the constitution without anyone's mythology, and you dress it in your own. That separation is what makes the doctrine portable across teams, tools, and vendors: take the rules, leave the costume. + +## Doctrine Stability + +The twelve names and numbers are frozen at v3.0.0. Corrections are **expression-only** — sharpen the prose, fix a diagram, re-cut a boundary that reads ambiguously. A factor may be rewritten, and a soft one may be reduced to an emphasis that points at its neighbor, but **no factor is ever deleted and the set is never renumbered.** The number is load-bearing (URLs, the badge, every inbound link, twenty years of "twelve-factor" recognition). When two factors feel redundant, the fix is to make the distinction legible, not to merge the slots. + ## The Operator Model Underneath the Factors The twelve factors are the public operating rules behind bookkeeping, @@ -203,10 +211,10 @@ Quickstart (5 min) → learnings.md file, zero tooling Foundation (I-III) → Context discipline, git tracking, fresh sessions Flow (IV-VI) → Research, validation, ratcheting Knowledge (VII-IX) → Extraction, compounding, measurement -Scale (X-XII) → Multi-agent isolation, supervision, failure harvesting (OPTIONAL) +Scale (X-XII) → Multi-agent isolation, supervision, failure harvesting (grow into it) ``` -**Key principle:** You can stop at any level and keep the value. Each level justifies the next, but none requires it. +**Key principle:** You can stop adopting at any level and keep the value. Each level justifies the next, but none requires it. Stopping means you haven't automated the higher factors yet — not that they stopped applying. **When to level up:** - **Quickstart → Foundation:** When your `learnings.md` gets unwieldy or you notice repeated context problems @@ -268,3 +276,4 @@ The factors evolve through production validation and community feedback. - **v1.0** (2025-01-27): Initial twelve factors — coding agent validation focus - **v2.0** (2025-12-27): Production implementation patterns added - **v3.0** (2026-02-15): Pivot to full operational discipline. Factors rewritten. Adoption model inverted (results-first, not manifesto-first). Knowledge compounding as hero differentiator. Scale factors marked optional. +- **v3.1** (2026-06-06): Whole-system constitution alignment. The 12 are reframed as one constitution lived at altitudes (one agent → a fleet), not a product partition. Factor V leads with the claims-vs-verdicts / single-writer moat; Factor XII rewritten to routing-hints + fresh-agent-on-failure; VII↔VIII, III↔X, X↔XI boundaries sharpened; Scale tier reframed from "optional" to the factory altitude. No factor renamed, renumbered, or deleted. diff --git a/VERSION b/VERSION index 46b105a..6c8dc7e 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -v2.0.0 +v3.1.0 diff --git a/docs/00-SUMMARY.md b/docs/00-SUMMARY.md index 1df5a25..b45ecb8 100644 --- a/docs/00-SUMMARY.md +++ b/docs/00-SUMMARY.md @@ -53,7 +53,7 @@ from a reliable operating model. | # | Factor | One-Line Rule | |---|--------|---------------| | **[IV](../factors/04-research-before-you-build.md)** | **Research Before You Build** | Understand the problem space before generating a single line of code. | -| **[V](../factors/05-validate-externally.md)** | **Validate Externally** | No agent grades its own work -- tests, linters, separate sessions, or humans validate. | +| **[V](../factors/05-validate-externally.md)** | **Validate Externally** | The worker reports evidence; an independent checker writes the binding verdict. No agent grades its own work. | | **[VI](../factors/06-lock-progress-forward.md)** | **Lock Progress Forward** | Once work passes validation, it ratchets forward and cannot regress. | ### Knowledge (VII--IX) -- Where Compounding Kicks In @@ -72,16 +72,18 @@ getting measurably smarter over time. > repeat your mistakes. Knowledge compounding is the one capability no amount of > model improvement replaces. -### Scale (X--XII) -- Advanced, Optional +### Scale (X--XII) -- The Factory Altitude -Multi-agent orchestration. Skip this tier entirely if you work solo. These -patterns apply only when running parallel agents on complex projects. +The same factors at fleet scale. Working solo, you live them in miniature -- a +git worktree is isolation, your own judgment is supervision, your `learnings.md` +is failure harvesting -- so you grow into the altitude rather than skipping the +factors. | # | Factor | One-Line Rule | |---|--------|---------------| | **[X](../factors/10-isolate-workers.md)** | **Isolate Workers** | Each worker gets its own workspace, context, and zero shared mutable state. | | **[XI](../factors/11-supervise-hierarchically.md)** | **Supervise Hierarchically** | Escalation flows up, never sideways -- one coordinator dispatches, workers execute. | -| **[XII](../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Failed attempts are data -- extract and index them with the same rigor as successes. | +| **[XII](../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Turn dead ends into routing hints that prune the next agent's search. | --- diff --git a/docs/explanation/from-theory-to-production.md b/docs/explanation/from-theory-to-production.md index f5ab675..d953829 100644 --- a/docs/explanation/from-theory-to-production.md +++ b/docs/explanation/from-theory-to-production.md @@ -338,7 +338,9 @@ Each factor maps to concrete implementation patterns from Houston, Fractal, and --- -### Scale Tier (X-XII, optional): Multi-Agent Operations +### Scale Tier (X-XII): The Factory Altitude + +The same three factors at fleet scale. Working solo you live them in miniature — a git worktree is isolation, your own judgment is supervision, your `learnings.md` is failure harvesting. Running parallel agents on complex projects, the same rules need real machinery. You grow into this altitude; you don't skip the factors. **Factor X: Isolate Workers** - **Philosophy:** Independent, focused execution diff --git a/docs/explanation/vibe-coding-integration.md b/docs/explanation/vibe-coding-integration.md index edf7056..a6b2b0e 100644 --- a/docs/explanation/vibe-coding-integration.md +++ b/docs/explanation/vibe-coding-integration.md @@ -68,11 +68,11 @@ This is where vibe coding transforms from a series of isolated sessions into a c --- -### Tier 4: Scale (Factors X-XII, Optional) +### Tier 4: Scale (Factors X-XII) — The Factory Altitude **Making vibe coding work across teams and complex systems.** -These factors are optional for solo developers but become essential when vibe coding scales beyond one person and one agent. +These are the same three factors at fleet scale. A solo developer already lives them in miniature — a git worktree is isolation, your own judgment is supervision, the note you write after a failed session is failure harvesting. When vibe coding scales beyond one person and one agent, the same rules need real machinery. You grow into this altitude; you don't skip the factors. | Factor | What It Solves | |--------|---------------| diff --git a/docs/principles/README.md b/docs/principles/README.md index 6342980..054dd7b 100644 --- a/docs/principles/README.md +++ b/docs/principles/README.md @@ -23,7 +23,7 @@ Twelve vendor-neutral principles organized in four tiers. Each tier builds on th | # | Factor | The Rule | |---|--------|----------| | **[IV](../../factors/04-research-before-you-build.md)** | **Research Before You Build** | Understand the problem space before generating a single line of code. | -| **[V](../../factors/05-validate-externally.md)** | **Validate Externally** | No agent grades its own work. Ever. | +| **[V](../../factors/05-validate-externally.md)** | **Validate Externally** | The worker reports evidence; an independent checker writes the binding verdict. No agent grades its own work. | | **[VI](../../factors/06-lock-progress-forward.md)** | **Lock Progress Forward** | Once work passes validation, it ratchets -- it cannot regress. | ### Knowledge (VII-IX) -- Where compounding kicks in @@ -36,15 +36,18 @@ Twelve vendor-neutral principles organized in four tiers. Each tier builds on th **Factor VIII is the hero.** It is the knowledge flywheel: extract learnings, gate for quality, inject into future sessions, measure retrieval, let stale knowledge decay. This is the differentiator that no amount of model improvement replaces -- better models with amnesia still repeat your mistakes. -### Scale (X-XII) -- Advanced, optional +### Scale (X-XII) -- The Factory Altitude -Multi-agent orchestration patterns. Skip this tier if you work solo. +The same factors at fleet scale. Working solo, you live them in miniature -- a +git worktree is isolation, your own judgment is supervision, your `learnings.md` +is failure harvesting -- so you grow into the altitude rather than skipping the +factors. | # | Factor | The Rule | |---|--------|----------| | **[X](../../factors/10-isolate-workers.md)** | **Isolate Workers** | Each worker gets its own workspace, its own context, and zero shared mutable state. | | **[XI](../../factors/11-supervise-hierarchically.md)** | **Supervise Hierarchically** | Escalation flows up, never sideways. | -| **[XII](../../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Failed attempts are data. Extract and index them with the same rigor as successes. | +| **[XII](../../factors/12-harvest-failures-as-wisdom.md)** | **Harvest Failures as Wisdom** | Turn dead ends into routing hints that prune the next agent's search. | --- diff --git a/docs/principles/comparison-table.md b/docs/principles/comparison-table.md index 02dd158..f609582 100644 --- a/docs/principles/comparison-table.md +++ b/docs/principles/comparison-table.md @@ -140,10 +140,10 @@ This compression does not replace the factors. It explains the mechanism beneath - *Key Practice*: Simple APIs for agent lifecycle management **12-Factor AgentOps (V: Validate Externally)** -- *Evolution*: No agent grades its own work +- *Evolution*: The worker reports evidence; an independent checker writes the binding verdict. No agent grades its own work. - *Why Different*: Agents are confident but not reliable -- they cannot objectively evaluate their own output -- *Key Practice*: External validation via tests, linters, a different agent, or human review -- *Unique Aspect*: Zero-trust principle applied to cognition -- validate the output, not the source +- *Key Practice*: The worker emits claims plus evidence; an independent checker -- tests, linters, a different agent, or a human reviewer -- is the sole writer of the binding verdict +- *Unique Aspect*: Zero-trust principle applied to cognition -- claims and verdicts are written by different parties, so a worker cannot launder its own confidence into a trusted result (the single-writer moat) --- @@ -242,7 +242,7 @@ This compression does not replace the factors. It explains the mechanism beneath - *Evolution*: Each worker gets its own workspace, its own context, and zero shared mutable state - *Why Different*: Parallel agents sharing state create cascading conflicts - *Key Practice*: Git worktrees, separate context windows, independent validation -- *Unique Aspect*: Scale tier -- skip if working solo; essential for multi-agent orchestration +- *Unique Aspect*: Scale tier (factory altitude) -- solo you live it every time you spin up a second worktree; structural for multi-agent orchestration --- @@ -262,7 +262,7 @@ This compression does not replace the factors. It explains the mechanism beneath - *Evolution*: Escalation flows up, never sideways - *Why Different*: Multi-agent systems without hierarchy devolve into circular coordination - *Key Practice*: Coordinator agents delegate and escalate; worker agents execute and report -- *Unique Aspect*: Scale tier -- skip if working solo; prevents coordination chaos in multi-agent setups +- *Unique Aspect*: Scale tier (factory altitude) -- solo you are the supervisor and break ties yourself; structural when the escalation path no longer fits in one head --- @@ -279,10 +279,10 @@ This compression does not replace the factors. It explains the mechanism beneath - *Key Practice*: Human contact is a first-class operation, not exception **12-Factor AgentOps (XII: Harvest Failures as Wisdom)** -- *Evolution*: Failed attempts are data -- extract and index them with the same rigor as successes +- *Evolution*: Turn dead ends into routing hints that prune the next agent's search space - *Why Different*: Failures contain the highest-value learnings but are typically discarded -- *Key Practice*: Document what did not work and why; index failure patterns for future avoidance -- *Unique Aspect*: Scale tier -- feeds directly into Factor VII (Extract Learnings) and Factor VIII (Compound Knowledge) +- *Key Practice*: Index negative knowledge for retrieval at decision time, and hand a stuck worker's failure trace to a fresh agent rather than looping the saturated one +- *Unique Aspect*: Scale tier (factory altitude) -- negative knowledge that prunes the search, distinct from Factor VII's generic capture; feeds Factor VIII (Compound Knowledge) --- diff --git a/docs/principles/evolution-of-12-factor.md b/docs/principles/evolution-of-12-factor.md index fcdc691..0928fdd 100644 --- a/docs/principles/evolution-of-12-factor.md +++ b/docs/principles/evolution-of-12-factor.md @@ -77,7 +77,7 @@ v3 restructured the factors around operational reality rather than theoretical t | **Organization** | Flat list of 12 | Four tiers: Foundation, Workflow, Knowledge, Scale | | **Adoption model** | All-or-nothing manifesto | Progressive -- stop at any tier, keep the value | | **Hero concept** | Distributed | Factor VIII (Compound Knowledge) is the differentiator | -| **Scale factors** | Required | Optional (X-XII) -- skip if working solo | +| **Scale factors** | Required | Factory altitude -- lived small solo, structural at fleet scale (never skipped) | | **Framing** | Framework for AI infrastructure | Operational discipline for working with agents | | **Entry point** | Read the theory first | Start with a `learnings.md` file and zero tooling | @@ -139,13 +139,13 @@ AI operations broke every assumption: **12-Factor AgentOps v3** added: - **Knowledge compounding** -- the flywheel that makes each session smarter (Factors VII, VIII) -- **External validation** -- no agent grades its own work (Factor V) +- **External validation** -- the worker reports evidence; an independent checker writes the binding verdict (Factor V) - **Progress ratcheting** -- validated work cannot regress (Factor VI) - **Research-first workflow** -- understand before generating (Factor IV) - **Outcome measurement** -- track what matters, not activity (Factor IX) - **Fitness gradient** -- define better versus worse states through goals, metrics, and gates (Factor IX) - **Provenance-backed learning** -- know where a learning came from before trusting or promoting it (Factors II, VII) -- **Failure harvesting** -- failed attempts are high-value data (Factor XII) +- **Failure harvesting** -- dead ends become routing hints that prune the next agent's search (Factor XII) - **Tiered adoption** -- start with zero tooling, scale when needed ### The Compression Beneath v3 @@ -200,9 +200,9 @@ Research before building. Validate externally. Lock progress forward. The discip Extract learnings. Compound knowledge. Measure outcomes. This is where sessions start getting measurably smarter over time. -### Scale (X-XII): Advanced, optional +### Scale (X-XII): The factory altitude -Isolate workers. Supervise hierarchically. Harvest failures. Multi-agent orchestration patterns. Skip this tier entirely if you work solo -- you lose nothing. +Isolate workers. Supervise hierarchically. Harvest failures. These are the same factors at fleet scale -- you grow into the altitude, you don't skip the factors. Working solo you live them in miniature: a worktree is isolation, your own judgment is the supervisor, your `learnings.md` is failure-harvesting. The machinery becomes structural when one head can no longer hold the whole thing. --- @@ -214,13 +214,13 @@ v3 reframes the core insight. The original framing was "zero-trust cognitive inf The v3 framing: **operational discipline for working with AI agents.** The same way DevOps transformed ad-hoc deployment into a reliable practice, 12-Factor AgentOps transforms ad-hoc agent usage into a reliable, compounding practice. -The zero-trust principle survives as Factor V (Validate Externally): no agent grades its own work. But the framework is bigger than validation. It is about: +The zero-trust principle survives as Factor V (Validate Externally): the worker reports evidence, an independent checker writes the binding verdict, and no agent grades its own work. But the framework is bigger than validation. It is about: 1. **Managing context** so agents get good input (Factor I) 2. **Persisting knowledge** so nothing is lost between sessions (Factor II) 3. **Scoping work** so agents operate in their effective range (Factor III) 4. **Understanding before building** so agents solve the right problem (Factor IV) -5. **Validating externally** so quality is objective (Factor V) +5. **Validating externally** so quality is objective -- claims from the worker, the binding verdict from an independent checker (Factor V) 6. **Ratcheting progress** so validated work is protected (Factor VI) 7. **Extracting learnings** so every session produces knowledge (Factor VII) 8. **Compounding knowledge** so each session is smarter than the last (Factor VIII) diff --git a/docs/reference/README.md b/docs/reference/README.md index fc9048b..fd9aaff 100644 --- a/docs/reference/README.md +++ b/docs/reference/README.md @@ -30,7 +30,7 @@ | **VIII** | [Compound Knowledge](../../factors/08-compound-knowledge.md) | HERO pattern: knowledge grows across sessions | | **IX** | [Measure What Matters](../../factors/09-measure-what-matters.md) | Track the metrics that drive improvement | -### Scale (X-XII, optional) +### Scale (X-XII) — the factory altitude | # | Factor | Purpose | |---|--------|---------| @@ -50,7 +50,7 @@ The 12 factors are organized into four tiers of increasing sophistication: **Knowledge (VII-IX)** -- The compounding engine. Extract learnings, compound them across sessions, and measure the metrics that actually matter. -**Scale (X-XII)** -- Optional. For teams running multiple agents. Worker isolation, hierarchical supervision, and systematic failure harvesting. +**Scale (X-XII)** -- The factory altitude: the same factors at fleet scale. Solo, you live them in miniature (a worktree is isolation, your judgment is supervision); running multiple agents, they become structural — worker isolation, hierarchical supervision, and systematic failure harvesting. You grow into the altitude, you don't skip the factors. --- diff --git a/factors/03-one-agent-one-job.md b/factors/03-one-agent-one-job.md index b7721e4..662c2c5 100644 --- a/factors/03-one-agent-one-job.md +++ b/factors/03-one-agent-one-job.md @@ -9,6 +9,8 @@ An agent that just finished researching your auth system is the worst agent to i When an agent completes a phase of work, end its session. The next phase gets a new agent with a clean window, loaded with only what it needs to execute. +This factor is about *one* agent across time — fresh context phase to phase. Keeping *concurrent* workers from contaminating each other is a different axis, and it belongs to [Factor X](./10-isolate-workers.md). Same instinct ("fresh context"), two altitudes: III is temporal, X is parallel. + ## The Rationale #### The Context Saturation Problem diff --git a/factors/05-validate-externally.md b/factors/05-validate-externally.md index d5466e6..a7bd7ee 100644 --- a/factors/05-validate-externally.md +++ b/factors/05-validate-externally.md @@ -2,15 +2,26 @@ ## Rule -**No agent grades its own work. Ever.** +**The worker emits claims plus evidence; an independent checker is the sole writer of the binding verdict. No agent grades its own work. Ever.** -Self-validation is confirmation bias with extra steps. The same context that generated the work is the worst possible context to validate it. Validation must come from outside the agent that did the work: a different agent, a different model, a test suite, or a human reviewer. +Separate the two roles cleanly. The worker that did the work produces a **claim** — "this is done, here is the proof." Something outside that worker — a different agent, a different model, a test suite, a human reviewer — turns the claim into a **verdict**, and only that verdict is binding. The worker can assert; it cannot ratify. Self-validation is confirmation bias with extra steps: the same context that generated the work is the worst possible context to validate it. If you can't validate externally, you haven't validated at all. +**This single-writer rule is the moat.** Anyone can have an agent that does work and reports success. The durable advantage is structural: claims and verdicts are written by different parties, and the authority to write the verdict is held outside the worker. When that separation is enforced in the write path — not just requested in a prompt — a worker *cannot* launder its own confidence into a result the rest of the system trusts. + In operator-model terms, validation is a **selection gate**. The environment, not the author, decides what survives: tests, review, ratchets, deploy checks, and other external gates accept or reject work before it becomes shared state. -This is also where **governance** becomes concrete. Governance sets the objective, the risk tolerance, and the boundaries; selection gates enforce them. A solo developer can do that with tests and a deliberate review pass. A team can add independent reviewers, deployment approvals, and escalation paths. The principle is the same. +This is also where **governance** becomes concrete. Governance sets the objective, the risk tolerance, and the boundaries; selection gates enforce them. + +### Lived at two altitudes + +This factor is the same rule whether you're one developer or a fleet, but it shows up at two altitudes: + +- **Worker altitude (honesty).** The worker is responsible for fresh-eyes review and for reporting claims *with* their evidence — command output, diffs, repro steps — never bare confidence. A solo developer does this with tests and a deliberate second pass; the discipline is "report what's true, attach the proof." +- **Factory altitude (authority).** A separate layer is the *sole writer* of the binding verdict and of anything promoted to shared, fleet-trusted state. The worker proposes; the gate disposes. At a team or fleet, this is independent reviewers, deployment approvals, and an assurance layer that holds write authority the workers don't have. + +The seam between them is the whole point: **honesty lives with the worker, authority lives with the gate.** Keep them in the same party and you're back to self-validation theater. --- diff --git a/factors/07-extract-learnings.md b/factors/07-extract-learnings.md index 6bc2971..0d83e19 100644 --- a/factors/07-extract-learnings.md +++ b/factors/07-extract-learnings.md @@ -12,6 +12,8 @@ This knowledge exists for exactly as long as the session stays open. The moment **Extraction is the difference between organizational learning and organizational amnesia.** +This factor is the **write** half of the knowledge loop: capture the lesson and give it provenance so it outlives the session. Getting it *back* into a future session — retrieval, injection, decay — is a different mechanism, and it lives in [Factor VIII](./08-compound-knowledge.md). Extraction without VIII is a write-only archive; VIII without extraction has nothing to serve. The two are a producer/consumer pair, not one habit. + Most people think they'll remember. They won't. Most people think the commit message is enough. It isn't. Most people think "the code is the documentation." The code shows what you built, not why you built it that way, not what you tried first, not what traps to avoid. Seen through the operator model, extraction is the act of turning transient session experience into durable traces. But durable traces only become reusable intelligence when they carry provenance: where the learning came from, what evidence supports it, and how confident you should be in reusing it. diff --git a/factors/08-compound-knowledge.md b/factors/08-compound-knowledge.md index ea9d0b4..4f85117 100644 --- a/factors/08-compound-knowledge.md +++ b/factors/08-compound-knowledge.md @@ -6,7 +6,7 @@ ## Rule -Every session must extract validated learnings and inject them at startup. Knowledge must compound over time through a closed-loop flywheel: extract what worked, gate for quality, decay what's stale, inject at session start. If learnings don't flow back automatically, you're running a write-only database that your agents will never read. +Learnings must flow back into future sessions automatically. This factor is the **read** half of the knowledge loop: take what [Factor VII](./07-extract-learnings.md) captured and gate it, store it, inject it at session start, cite it, and decay what stops earning citations. Factor VII writes the lesson down; Factor VIII is what makes writing it down pay off. If learnings don't flow back automatically, you're running a write-only database that your agents will never read. Seen through the operator model, this is where a **stateful environment** becomes smarter than any single session. The actors remain replaceable. The environment carries continuity through learnings, citations, checkpoints, and reusable rules. Intelligence compounds when those traces move through **promotion loops** instead of sitting in storage. @@ -27,6 +27,11 @@ If this inequality fails, your knowledge decays to zero. If it holds, session 50 This is institutional memory that actually works—not a wiki nobody reads. +### Lived at two altitudes + +- **Worker altitude (the flywheel).** Inside a project, the loop compounds: retrieve relevant prior lessons at task boundaries, cite them, let uncited knowledge decay. Session 50 beats session 1. +- **Factory altitude (promotion authority).** Deciding what is assured-enough to promote into shared, fleet-trusted knowledge is its own gate — and, like the verdict in [Factor V](./05-validate-externally.md), that promotion is written by the assurance layer, not by the worker that produced the lesson. Honesty proposes a learning; authority promotes it. + --- ## Rationale @@ -41,29 +46,32 @@ So they make the same mistake. Again. And the team extracts the same learning. A ### The Flywheel Pattern -Compound knowledge requires a closed loop: +Compound knowledge requires a closed loop. Extraction (Factor VII) feeds the loop from outside; this factor owns everything from the quality gate onward: ``` + Factor VII + (EXTRACT) + │ captured lesson + provenance + ▼ ┌─────────────────────────────────────────────────┐ -│ │ -│ ┌─────────┐ ┌──────┐ ┌───────┐ │ -│ │ EXTRACT │ ──> │ GATE │ ──> │ STORE │ │ -│ └─────────┘ └──────┘ └───────┘ │ -│ ▲ │ │ -│ │ ▼ │ -│ ┌─────────┐ ┌────────┐ │ -│ │ CITE │ <────────────── │ INJECT │ │ -│ └─────────┘ └────────┘ │ -│ │ │ -│ ▼ │ -│ ┌─────────┐ │ -│ │ DECAY │ (prune uncited knowledge) │ -│ └─────────┘ │ -│ │ +│ ── Factor VIII: the compounding loop ── │ +│ ┌──────┐ ┌───────┐ │ +│ │ GATE │ ──> │ STORE │ │ +│ └──────┘ └───────┘ │ +│ ▲ │ │ +│ │ ▼ │ +│ ┌─────────┐ ┌────────┐ │ +│ │ CITE │ <─│ INJECT │ │ +│ └─────────┘ └────────┘ │ +│ │ │ +│ ▼ │ +│ ┌─────────┐ │ +│ │ DECAY │ (prune uncited knowledge) │ +│ └─────────┘ │ └─────────────────────────────────────────────────┘ ``` -**EXTRACT**: At session end, capture what worked, what failed, what you learned. This is post-mortem, retrospective, or simply structured reflection. The key: make it specific, actionable, and tagged for retrieval. +**EXTRACT** *(Factor VII, feeding in)*: At session end, capture what worked, what failed, what you learned, with provenance. This is the *write* half and it belongs to Factor VII — the loop below consumes its output. The key: make it specific, actionable, and tagged for retrieval. **GATE**: Not all learnings are equal. Bad learnings pollute the knowledge base. Gate entries through quality filters: - Is it actionable? ("Don't use `rm -rf`" without safer alternatives is noise) diff --git a/factors/10-isolate-workers.md b/factors/10-isolate-workers.md index 62b69c1..b08c7c5 100644 --- a/factors/10-isolate-workers.md +++ b/factors/10-isolate-workers.md @@ -1,13 +1,17 @@ # X. Isolate Workers -**This factor is part of the Scale tier (X-XII) — advanced patterns for multi-agent workflows. Not a prerequisite for getting value from Factors I-IX.** +**Scale tier (X–XII) — lived at the factory altitude. Solo, you live it every time you spin up a second worktree; at fleet scale it becomes structural. The axis here is *independence between peers* — distinct from [Factor XI](./11-supervise-hierarchically.md), which is *authority up a chain*. Isolation keeps workers from corrupting each other; supervision decides who resolves it when they conflict. Different problems, often deployed together.** ## Rule **Each worker gets its own workspace, its own context, and zero shared mutable state.** +Where [Factor III](./03-one-agent-one-job.md) scopes *one* agent over time — a fresh window between phases so research context doesn't bleed into implementation — this factor keeps *many* agents from contaminating each other at once. Same word, "fresh context," two different axes: III is temporal (one worker, phase to phase), X is concurrent (many workers, side by side). + When you run multiple agents in parallel, isolation is everything. Two agents sharing a context window corrupt each other. Two agents sharing a working directory create race conditions. Two agents sharing mutable state create cascading failures that are impossible to debug. +Isolation is also what makes workers **fungible**. When each worker is sealed off behind its own worktree and context, any one of them is disposable: if it fails, delete the worktree, spawn a fresh worker, reassign the job. The durable unit is the *tracked task*, not the worker that happens to be holding it — workers are interchangeable and ephemeral by design. + True isolation means: - **Separate git worktrees** for each worker (not just branches) - **Fresh context per worker** (no shared conversation history) diff --git a/factors/11-supervise-hierarchically.md b/factors/11-supervise-hierarchically.md index 783b36e..58afded 100644 --- a/factors/11-supervise-hierarchically.md +++ b/factors/11-supervise-hierarchically.md @@ -6,7 +6,9 @@ ## Tier -This factor is part of the **Scale tier (X-XII)** — advanced patterns for multi-agent workflows. Not a prerequisite for getting value from Factors I-IX. You can run effective single-agent operations without supervision hierarchies. This becomes critical when you scale to multiple agents working concurrently, where failures compound and coordination overhead explodes without clear escalation paths. +This factor is part of the **Scale tier (X–XII)**, lived at the factory altitude. Working solo you *are* the supervisor — you hold the escalation path in your head and break the ties yourself. The rule doesn't appear when you scale; it just stops fitting in one head, and the tree has to become structural. It becomes critical when multiple agents work concurrently and coordination overhead explodes without clear escalation paths. + +The axis here is **authority up a chain** — distinct from [Factor X](./10-isolate-workers.md), which is **independence between peers**. Isolation stops workers from corrupting each other; supervision decides who has the authority to resolve things when they conflict or get stuck. You can have one without the other: a flat swarm of isolated workers with no supervisor, or a supervisor over workers that share state. You usually want both, but they are answering different questions. --- diff --git a/factors/12-harvest-failures-as-wisdom.md b/factors/12-harvest-failures-as-wisdom.md index 6c09049..a893f0a 100644 --- a/factors/12-harvest-failures-as-wisdom.md +++ b/factors/12-harvest-failures-as-wisdom.md @@ -1,18 +1,19 @@ # XII. Harvest Failures as Wisdom -> **This factor is part of the Scale tier (X-XII) — advanced patterns for multi-agent workflows. Not a prerequisite for getting value from Factors I-IX.** +> **Scale tier (X–XII) — lived at the factory altitude. Working solo, your `learnings.md` and your own memory of dead ends already do this in miniature; at fleet scale it needs real machinery. Not something you bolt on — something you grow into.** ## The Rule -**Failed attempts are data. Extract and index them with the same rigor as successes.** +**Turn failed attempts into routing hints that prune the next agent's search space.** -Every failed approach is a negative result. Negative results are knowledge. Knowledge compounds. Most systems treat failures as noise to be suppressed. 12-Factor AgentOps treats them as signal to be harvested. +[Factor VII](./07-extract-learnings.md) captures what a session learned. This factor is about the distinct power of *negative* knowledge and the machinery that exploits it. A recorded dead end doesn't merely stop the next agent from repeating a mistake — it removes whole branches from the search *before the agent starts*. "Don't try X when Y holds" is often worth more than a positive pattern, because positive knowledge tells you one thing to do while negative knowledge eliminates many things not to. The first prunes the tree; the second only adds a leaf. -When an agent tries three approaches before the fourth works, you don't just have one success — you have three documented learnings about what doesn't work under specific conditions. That's the wisdom that prevents the next agent from burning cycles on the same dead ends. +Two mechanisms make this factory-grade, and neither is just "extract learnings, but sad": -In operator-model terms, failures are durable traces. They coordinate future work across sessions by narrowing the search space for the next actor. That matters because actors are replaceable. The environment has to remember the dead ends so the next worker does not pay the same tuition again. +- **Failure as a routing hint.** A failure record is not a diary entry — it is an input the next worker reads *before choosing an approach*, so its search starts already narrowed. Index negative knowledge for retrieval at decision time, not for reading after the fact. +- **Fresh-agent-on-failure.** When a worker is stuck, don't loop it through the hole it already mapped. Hand the failure context — what was tried, the conditions, the error signature — to a *fresh* agent. The stuck worker's saturated context is the worst place to recover from; a clean worker carrying the failure trace is the best. -Harvesting failures also creates promotion loops for negative knowledge: failed attempt → validated pattern → preventative rule. Successes are not the only things worth promoting. +When an agent tries three approaches before the fourth works, you don't just have one success — you have three documented dead ends that narrow the next agent's search to the branch that pays. In operator-model terms, failures are durable traces that coordinate future work by shrinking the search space, and they ride **promotion loops for negative knowledge**: failed attempt → validated dead end → preventative rule, gate, or test. ## The Rationale