Status: Architecture of a platform in production use (September 2026). The reusable layers are published as standalone reference implementations, linked below; client data, credentials, provider adapters and company-specific policy stay private. · Portfolio map ›
This is the architecture of the internal platform I run for a team of six GTM engineers, who build and operate outbound, research and signal systems for 34 client companies (September 2026). The engineers are the platform's users. The platform gives them a catalog that resolves every client's systems from one slug, a read-only CLI scoped from telemetry, published skills for the common jobs, unattended jobs with drift detection, a record of every agent tool call, guardrails derived from that record, an eval-gated model allowlist, and propose-then-apply approval for anything that writes, sends or spends.
The document is organized as ten capability layers. For each one it states the invariant the layer enforces, the failure it prevents, what it measured, and the public reference implementation where one exists. The last sections list the failures the platform caught in itself and the gaps a platform engineer would find, because the platform has them.
The team side: how the engineers who use this platform are organized and led is in gtm-engineering-operating-model.
platform (one person) → GTM engineers (build client systems) → client revenue programs
this document the platform's users 34 clients
In Team Topologies terms the GTM engineers are stream-aligned: each owns a book of clients end to end. The platform exists to lower their cognitive load: they should not need to know which database, repository, sending workspace or model a client's work touches. RevOps is not the platform team here; the platform sits under the GTM engineers, not beside sales.
Platform and workloads are kept separate. List building, research, deliverability and campaign copy are workloads the engineers run on the platform. They are listed at the end, not as layers, because they are what the platform serves, not part of it.
| # | Layer | What it gives GTM engineers | Nearest standard term |
|---|---|---|---|
| 1 | Catalog and tenancy | One client slug resolves every system that client has; the model never chooses a tenant | Service catalog; developer control plane |
| 2 | Self-service interfaces and golden paths | A read-only CLI, APIs and published skills for the common jobs | APIs and CLIs; golden path templates |
| 3 | Delivery and fleet | Unattended jobs as separate apps, with drift detection between repo and deployed image | Integration and delivery plane |
| 4 | Observability | Every agent session and tool call in one warehouse | Monitoring and logging plane |
| 5 | Policy and guardrails | Rules at the tool boundary, derived from recorded failures | Policy enforcement; security plane |
| 6 | Harness feedback loop | Telemetry → proposed change → human approval → measured against a control | Feedback on the outcome of tasks (DORA) |
| 7 | Model gateway | Which model may take which task, earned by an eval and re-checked for drift | AI gateway / LLM gateway |
| 8 | Context and knowledge | Versioned, governed context: client repositories, memory, knowledge graph, graded copy rules | Data services; context as code |
| 9 | Human approval and authority | Propose-then-apply loops; a ledger of what each unattended app may touch | Identity, scoped permissions, audit trail |
| 10 | Team interfaces | Where engineers and colleagues reach the platform: Slack, a daily digest, per-client cards | Portal |
- Invariant: the model never constructs a tenant id, path, namespace or credential. A generated registry resolves the client first; the agent then gets content-level operations only.
- What it holds: one slug per client → its repository, database, Slack channel, sending workspace, task-board filter, sector and the engineer who owns it. The registry and the engineer-assignment map are generated from their sources and never edited by hand.
- Failure it prevents: cross-tenant reads and hallucinated identifiers.
- What it measured: a fuzzy client-name match once read another client's data at 0.53 similarity; client resolution now refuses matches below 0.80.
- Reference implementation: agent-tenancy.
- Invariant: the common read paths are one command with JSON out and categorical exit codes, so an agent branches on the exit code instead of parsing prose. Writes are not on the CLI; they go through dedicated paths with an approval step.
- Why it exists: telemetry showed agents writing a fresh inline program for routine lookups, about 6 inline snippets for every call to a committed script, and thousands of snippets that existed only to import the shared library. The CLI's command set was chosen from that telemetry: 14 commands cover about 95% of real use.
- Golden paths: published skills for campaign copy, dial research, government-sector launch, list upload and queue grooming, each mirrored to a read-only team repository with its exact dependency list and a secret check before every push. A dial-research API gives each engineer a worker pool: 100 organisations researched in 14.2 minutes, with a free title preview so an engineer can correct the target titles before research is spent.
- Reference implementation: the list-building path in gtm-pipeline. The CLI and skills are private.
- Invariant: every unattended job is its own app, and the repo can say which deployed apps are running code older than the repo.
- How: each app bakes its dependencies into its image at build time and does no pull at boot. The drift detector derives each app's file set from that app's own build file, hashes it, and compares it with what was last deployed. It also flags code in an image that refers to files the image does not contain.
- Failure it prevents: an edit that "shipped" in the repo while every machine keeps running the old text with no error anywhere.
- What it measured: drift is detected, not yet kept at zero. It also cannot see a secret that was frozen at deploy time; see "Failures it caught in itself" below.
- Invariant: every agent session, sub-agent and tool call is recorded (Claude Code, Codex, Cursor) alongside model-router spend, in one warehouse.
- Failure it prevents: a workflow that drifts silently. Two runs can produce the same output while one recovered from three failures and got lucky.
- What it measured: about 238,000 tool calls and model requests over 150 days fed the feedback loop in layer 6. A verification rule came out of this layer's own failure: a job is checked by reading the written row back, never by its exit code.
- Reference implementations: cc-logger · codex-logger · cursor-logger.
- Invariant: recoverable mistakes get a nudge; irreversible actions are blocked; a guard bug never stops a session. Guards fail open; policy fails closed.
- Where rules come from: failures that actually recurred in recorded runs, not a list of what might go wrong. Rules start as monitors and are promoted only on evidence.
- What it measured: in a controlled comparison, guards changed token use by +0.2%, which is below the noise between control runs. Guards here limit damage; they do not make agents faster.
- Reference implementations: callusguard (on PyPI), with its components agent-guard, codex-guard and wroteonly.
- Invariant: nothing changes the agents' runtime without a human approving it, and every approved change is measured against a control.
- The loop: recorded runs → proposed guard rules, skills, sub-agent types and route checks → approval by reply to a digest → the change lands as a reviewable diff → measured, then kept, reviewed or retired.
- What it measured: two rules cut their failure class by 19.1 and 15.4 points against a control. The first version of a third made failures 26.4 points worse and was rewritten. Five of the loop's own recommendations changed once checked against production, and one live experiment was inconclusive (Fisher's exact p = 0.43). All of it is published.
- Reference implementation: runtune and its evidence doc.
- Invariant: a task may run on a cheaper model only after an evaluation on real task data has cleared that exact task. Modes are named for the task, not the model. Everything without a clearance stays on the frontier model or is refused.
- What it holds: one allowlist, read live by every caller, with a verification date per mode, an optional provider pin so production runs on the endpoint the eval scored, and a frozen regression spec that is re-run to catch drift after a provider changes a model underneath its id. Nine task types hold a clearance; one more is on trial.
- What it measured: a model that was never on the allowlist ran 1,464 requests and classified 34,162 production rows before it was found, and scored worst of five candidates when it was evaluated afterwards. The failure was a missing checkpoint, not the model. The feedback loop in layer 6 later found other model traffic bypassing the router, traced it to a helper, and it was fixed.
- Reference implementation: model-eval-gate.
- Invariant: shared context is versioned and carries its provenance. Agents propose writes to durable memory; only validated writes land.
- What it holds: each client's repository, refreshed from new call transcripts with recent calls weighted higher; a memory store where every fact records the session that wrote it and the date it was last verified, swept periodically for staleness; a knowledge graph behind a write gate; and cold-email rules graded against outcomes.
- What it measured: the cold-email rules are graded against 171,151 sends and 3,002 human replies. About 74% of raw "replies" were autoresponders, so dashboard reply rates run roughly 4x high. The rules are enforced at draft time by a linter that shares its feature extraction with the analysis, and a quarterly regression checks that each enforced rule still predicts replies.
- Reference implementation: knowledge-graph-governance.
- Invariant: anything that writes to a shared system, sends, or spends is proposed first and applied only after a person approves the specific change.
- Authority ledger: a daily record of what each unattended app may touch, which models are cleared for which kind of contribution, and who may decide what. A row is written only when that ceiling changes, so the table reads as the list of moments it moved.
- What it measured: see "Failures it caught in itself" below; most of them are this layer's.
- Reference implementation: private.
- Invariant: the platform reaches people where they already work, and silence is a designed outcome, not a failure.
- What it holds: a Slack agent that answers in one person's name without lending that person's authority to whoever asked; one shared attention inbox that every unattended job writes to instead of notifying directly, drained into one daily digest; a weekly prep card per client.
- Reference implementation: private.
| Workload | Reference implementation |
|---|---|
| List building: plain-English brief → deduped, qualified, sequencer-ready list | gtm-pipeline |
| Source-verified research, free sources first | gtm-research |
| Deliverability: mail-gateway cohorts and per-account throttling before launch | gtm-deliverability |
| Dial research: 0 fabricated contacts in a 1,472-record evaluation | private |
| Cold-email campaigns, linted against the graded rules in layer 8 | private |
I expected the platform to double how many clients each engineer could carry. It did not.
- From Q1 to Q3 2026, clients with a launched campaign per engineer rose from 5.4 to 7, while launches per engineer rose to 20 a month.
- Emails sent rose 2.81x over the same period.
- Current load is 7.5 active clients per engineer, against a published agency norm of 3–7 per operator (GTME Pulse, n=228).
Headcount and process changed in the same months, so these numbers do not fully show that the platform caused the rise, though theres strong correlation. A baseline is being taken before per-engineer load changes again, to be re-read at 30, 60 and 90 days.
Each one returned a plausible answer, not an error. The fix in each case was a test over the whole class of mistake, not the one instance.
| What happened | What catches it now |
|---|---|
| A new column on the task board renumbered its status options; 793 of 1,063 cards read as the wrong status, and one card was written to the wrong column | Statuses are addressed by name and read back after every write; a sweep test fails on any pinned index |
A secret frozen at deploy time left 14 clients without database credentials on one sync app, and the run reported errors: 0 |
Jobs are verified by reading rows back; the gap is listed below |
| A reassigned sending workspace mapped one client onto another's | Workspaces are verified by asking each key which workspace it owns, not by name |
| A fuzzy client-name match read another client's data at 0.53 similarity | A 0.80 similarity floor |
| A partial approval parsed as "approve all 25" (nothing was written) | A 22-case parsing test on that loop's reply parser |
| A roster email change silently unassigned one engineer's clients | An alias bridges the join, and the map builder warns on any engineer it cannot resolve |
| The telemetry database filled up and dropped 32,811 writes while every job exited 0 | Moved to a larger warehouse; writers are checked by row count |
These are the questions a platform engineer should ask, answered.
- Adoption is measured for about half the team. The usage record covers some engineers well and others not at all, so "the platform powers the team" is not yet a measured claim.
- No per-agent identity yet. Unattended apps share credentials more broadly than their declared reach; the authority ledger records this rather than hiding it.
- Drift is detected, not yet kept at zero, and the detector cannot see a secret frozen at deploy time.
- Missing platform-team basics: a catalog page engineers can browse, a public roadmap, an engineer survey, stated reliability targets, a deprecation policy, and a contribution path.
- A platform team of one. The bus-factor question is fair.
- One team's evidence. Every number above comes from one platform and one team.
- Deterministic code owns consequences; the model owns content. Tenants, resource ids and credentials are resolved before the agent runs.
- Refuse by default, prove to allow. Cheaper models, new graph writes and new rules earn their place with an evaluation on real data, and lose it when the evidence goes stale.
- Grow controls from telemetry. Rules written from what might go wrong over-block and under-cover.
- Choose the failure direction per layer. Guards fail open so a guard bug never stops a session; policy fails closed so a malformed allowlist refuses everything.
- Human gates on spend and irreversibility. Everything recoverable stays fast and unattended.
- Provenance and audit by default. Every binding, fact and durable write records who wrote it, in which run, and when.
- Verify the written state, not the exit code. Most failures above exited cleanly.
Client data, credentials, provider adapters, company-specific policy (eval verdicts, allowlists, controlled vocabularies) and the orchestration code that wires the layers together. The public repositories are reference implementations extracted from the running platform.
Maintained by Kai Karlstrom. Figures are as of September 2026 and are updated from their sources, not by hand.