Skip to content
162 changes: 120 additions & 42 deletions docs/azure-requirements.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,11 +36,19 @@ Added the same day:
- A cost guard, so that a day's spend cannot quietly reach 100 dollars.
- "For cross-check, just use the same pi fleet (or copy it in, whatever works)."

Amended by the owner on 2026-08-19: the crosscheck requirement is model-family bidirectionality,
not the literal codex/claude pairing quoted above. The second reviewer family is Kimi-K2.7-Code
on Azure AI Foundry, reviewing codex-authored work; claude-authored work keeps the codex
crosscheck. pi-anthropic was rejected the same day (API pricing), as was an interim same-day
Claude-Code-CLI-subscription direction. Details and work in R6.
Amended by the owner on 2026-08-19: the crosscheck requirement is that no author's work is
reviewed only by its own model family, not the literal codex/claude pairing quoted above.
The second reviewer family is GLM-5.2 on Azure AI Foundry.
pi-anthropic was rejected the same day (API pricing), as was an interim same-day
Claude-Code-CLI-subscription direction, and an interim Kimi-K2.7-Code pick was revised to
GLM-5.2 the same day on model quality. Details and work in R6.

Amended again by the owner later on 2026-08-19: crosscheck routes through GLM only. A single
reviewer family outside both author families satisfies the paradigm for every author, unbinds
review capacity from the pi-codex subscription profiles the fleet's own work depends on, and
makes the lane scalable for firstmate and for engineers' on-demand use. pi-codex stays as an
explicit fallback whose activation must be easy to see. no-mistakes stays on pi-codex.
The owner also directed a Slack v1 team exposure of crosscheck, recorded as R10.

## R1. Crewmates run in Azure

Expand Down Expand Up @@ -87,7 +95,9 @@ A stranding strand was fixed: a passed run carrying a short receipt set is now d
so the cell collects and retains legibly.
That buys legibility, not closing, and the root cause (an in-cell bridge producing no receipts) is
untouched, so a demoted cell still cannot `close`.
One stranded work disk is on the subscription today.
The one stranded work disk that sat on the subscription was deleted by the owner's direction on
2026-08-19 after no cell record could be found anywhere for the sanctioned close, so the lane's
inability to close now has no live residue, only the unfixed root cause.
The runner-offload lane (`FM_AZURE_RUNNER_REMOTE_CLASSES`) has no caller: nothing in the
repository sets `FM_AZURE_RUNNER_TASK`, `FM_AZURE_RUNNER_GENERATION`, or
`FM_AZURE_RUNNER_CONFIRM_SUBSCRIPTION`, and the placeholders in `docs/azure-runner.md` are an
Expand All @@ -108,9 +118,11 @@ separately, one validation cell reaches `close` with its worktree disk released.

Status: PARTIAL.

Crosscheck is done: eight pi profiles across eight distinct upstream accounts, projected into
single-profile account homes by `bin/fm-pi-account-home.py`, with the roster repointed and read
back through the real `bin/fm-crosscheck.py` reader and policy screen.
Crosscheck on the pi fleet is done at the roster level: eight pi profiles across eight distinct
upstream accounts, projected into single-profile account homes by `bin/fm-pi-account-home.py`,
with the roster repointed and read back through the real `bin/fm-crosscheck.py` reader and
policy screen. Under the second 2026-08-19 amendment this roster is now the dormant crosscheck
fallback; the pi fleet's primary duties are authors and no-mistakes.
Crewmate placement is not wired: the cell image carries pi, but placement does not select across
the eight profiles.

Expand All @@ -119,9 +131,9 @@ tool and the account-lease identity already present in the worker request path.

Acceptance: concurrent crewmates run on distinct pi profiles with no account collision.

## R6. Crosscheck is bidirectional across model families
## R6. Crosscheck reviews outside the author's model family

Status: NOT DONE. Direction decided by the owner 2026-08-19 (see the amendment above).
Status: NOT DONE. Direction decided by the owner 2026-08-19 (see the second amendment above).

The requirement is that no author's work is reviewed only by its own model family.
The roster was instead made eight reviewers all on `openai-codex`, by reading "just use the same
Expand All @@ -133,47 +145,79 @@ pi-anthropic is dead twice over: pi's anthropic OAuth authenticates as its own c
own token endpoint, so a Claude CLI refresh token cannot be spent by pi, and a fresh pi-anthropic
login would bill at API pricing, which the owner rejected.
An interim directive the same day routed claude reviews through the Claude Code CLI on the
owner's subscription profile; the standing decision superseded it: the second reviewer family is
Kimi-K2.7-Code on Azure AI Foundry (Direct-from-Azure lane; at decision time deployable Global
Standard from eastus with tool calling and published pay-per-token pricing - re-verify at deploy
time), chosen for tool calling, the Microsoft-hosted custody lane, and lineage independent of
both OpenAI and Anthropic.
Codex-authored work is reviewed by a Kimi-backed reviewer; claude-authored work keeps the codex
crosscheck, which R5 records as proven at the roster level while R9 still owes the live proof.
The model pick is explicitly provisional: reevaluate after live review data (GLM-5.2 was the
runner-up; the comparison is in the owner's evidence folder,
R6-FOUNDRY-RESEARCH-2026-08-19.md).

No owner login is needed anymore: the Kimi lane authenticates with a Foundry deployment and an
owner's subscription profile; a later interim pick was Kimi-K2.7-Code; the standing decision is
GLM-5.2 through the Fireworks AI lane on Foundry (`FW-GLM-5.2`, pay-per-token Data Zone Standard
US; at decision time it carries tool calling, a 1M-token context, and `reasoning_effort` with
`max` as the default - re-verify at deploy time), chosen on model quality with lineage
independent of both OpenAI and Anthropic.
The custody trade was accepted by the owner on 2026-08-19 knowing its exact shape: Microsoft
disclaims data handling for the Fireworks lane (data is shared between Microsoft and Fireworks
and processed on Fireworks infrastructure inside the US data zone), and the zero-data-retention
promise (volatile memory only, no logging by default, on the chat-completions surface) is
Fireworks' own policy, not Microsoft's. The lane must therefore use plain chat completions
only; the Responses API retains data for 30 days under its default store flag and is forbidden
here.
Billing was verified before acceptance: Fireworks pay-per-token on Foundry bills as Azure
consumption (a feature registration, not a Marketplace SaaS purchase), is MACC-eligible, and
Microsoft's own Startups material states startup credits apply to exactly this SKU; a small live
spend must confirm the credit decrement before volume. Fireworks pay-per-token models can retire
on 15 days notice, which is an accepted operational risk and one more reason the fallback below
stays armed.
Later the same day the owner simplified the routing: ALL crosscheck reviews route through the
GLM lane, for codex-authored and claude-authored work alike. GLM belongs to neither author
family, so single-family review is avoided for every author with one reviewer lane, and review
capacity stops competing with the pi-codex subscription profiles that no-mistakes and the
author fleet consume. The pi-codex roster (which R5 records as proven at the roster level while
R9 still owes the live proof) is retained as a dormant fallback behind a config flip, never
deleted; every review must name the lane that produced it, and a status read must show whether
GLM is serving or the fallback is active, so a silent fallback is impossible.
Fallback operation is a recorded degradation, not free service restoration: with the fallback
active, codex-authored work is reviewed by its own family again (the flip therefore includes
`config/crosscheck-same-model` on for the duration, which the policy screen otherwise refuses),
which is exactly the defect this requirement removes - accepted only while GLM is unavailable,
and one more reason fallback activation must be loud.
The model pick is explicitly provisional: reevaluate after live review data (Kimi-K2.7-Code was
the prior same-day pick; the comparison and the custody/billing verification are in the owner's
evidence folder, R6-FOUNDRY-RESEARCH-2026-08-19.md).

No owner login is needed anymore: the GLM lane authenticates with a Foundry deployment and an
api-key, not a subscription session.
The Kimi credential is an api-key, not a pi OAuth slot, so the literal `openai-codex` slot key in
The GLM credential is an api-key, not a pi OAuth slot, so the literal `openai-codex` slot key in
`bin/fm-pi-account-home.py`, `bin/fm-crosscheck.py` (`inspect_pi_credential`,
`account_identity`), and the Azure credential archive in `bin/fm-crosscheck-azure.py` stays as it
is - those three still refuse any non-codex pi OAuth slot, which no longer blocks R6 and remains
the recorded constraint if a second pi OAuth provider is ever added.
The identity question those tools answered with `accountId` still needs an answer for an api-key
credential: reviewer identity for the Kimi lane must bind the Foundry resource and deployment,
credential: reviewer identity for the GLM lane must bind the Foundry resource and deployment,
since an api-key carries no account identity of its own.

Work: deploy `Kimi-K2.7-Code` in a Foundry resource and store the key in the fleet's secret
custody, never in the repo; a pi custom provider entry (`models.json` `baseUrl` +
Work: register the `Fireworks.EnableDeploy` subscription feature, deploy `FW-GLM-5.2`
(pay-per-token Data Zone Standard) and store the key in the fleet's secret custody, never in
the repo; confirm by a small live spend that the charge decrements startup credit; verify the
endpoint passes `reasoning_effort` through and pin the lane to chat completions only; a pi
custom provider entry (`models.json` `baseUrl` +
`openai-completions` api) pointing at the resource's OpenAI-compatible `/openai/v1` endpoint
with the deployment name as the model id; verify pi tolerates `reasoning_content` in streamed
deltas before rollout; extend `bin/fm-crosscheck.py` `allowed_profiles` and roster validation to
carry the Kimi lane (today they pin pi to `("pi", "gpt-5.6-sol", "xhigh")` and would refuse it)
carry the GLM lane (today they pin pi to `("pi", "gpt-5.6-sol", "xhigh")` and would refuse it)
with `config/crosscheck-same-model` off; extend the `bin/fm-crosscheck-azure.py` endpoint
allowlist and credential archive for the Foundry host and api-key shape; define reviewer
identity for api-key credentials as above; retire or re-point the interim claude reviewer
artifacts (the `("claude", "claude-opus-5", "xhigh")` `allowed_profiles` entry, the
`api.anthropic.com` allowlist entry, and the claude-profile boot copy described in
`docs/azure-crosscheck.md`); review guards sized to the model's context window at deploy time: a
strict findings schema, path-existence validation before filing, and a per-review context cap;
and routing so codex-authored changes draw the Kimi reviewer while claude-authored changes draw
codex reviewers.

Acceptance: a codex-authored change is reviewed by a Kimi-backed reviewer and a claude-authored
change by a codex-backed reviewer, on distinct credentials with bound reviewer identities, with
the same evidence discipline as the codex lane.
routing so every crosscheck draws the GLM reviewer, with the pi-codex roster behind a config
flip as fallback (the flip sets `config/crosscheck-same-model` on, accepting same-family review
of codex-authored work as the recorded degraded mode while it is active); and lane visibility:
the review evidence and report name the reviewing lane, and a status command answers whether
GLM is serving or the fallback is active.

Acceptance: a codex-authored change and a claude-authored change are each reviewed by a
GLM-backed reviewer with bound reviewer identity and the same evidence discipline as the codex
lane; the fallback flip to pi-codex is demonstrated once, its activation is visible in the review
evidence and the status read, and the demonstration records the degraded same-family mode it
accepts for codex-authored work.

## R7. Everything is logged in

Expand All @@ -182,9 +226,8 @@ Status: HOLDS, through R8.
The eight pi profiles renew on their own now, which is R8.
Two of the three profiles in `~/.local/share/agent-fleet/accounts/claude/` hold blanked,
length-zero tokens; the third is `refreshable` with material declared valid to 2026-09-10.
None of the three is needed for R6 anymore: the Kimi lane authenticates with a Foundry
deployment key, and claude-authored work is reviewed by the codex fleet, so R7 holds with no
owner login outstanding.
None of the three is needed for R6 anymore: the GLM lane authenticates with a Foundry
deployment key and reviews all authors, so R7 holds with no owner login outstanding.

## R8. Auth refreshes on its own

Expand Down Expand Up @@ -235,6 +278,33 @@ The existing smoke assignments cannot be reused because their `repository_genera
commit that exists.
Proving it requires a real spawned crewmate task that commits.

## R10. Crosscheck is exposed to team engineers through Slack

Status: NOT DONE. Directed by the owner 2026-08-19; builds after R6.

The owner's v1 shape: an engineer tags the crosscheck bot in a Slack channel with a pull request
link; the GLM lane reviews it; the bot posts the findings as a thread reply on the engineer's
own message. No engineer wires up a harness or touches an endpoint, and the same path works for
deliberate on-demand use. Cursor Bugbot continues to run for engineers' pull requests (it stays
disabled on the owner's), so this lane complements rather than replaces it.

Constraints the build must honor:

- The listener uses Slack Socket Mode, so no public inbound endpoint is added to the private
lane posture. It is a resident process; where it runs is decided at build time and its
standing cost is recorded under C3.
- v1 accepts pull request links only, and only for repositories in the organization allowlist.
The bot's repository read credential must never be pointed at a repository outside that
allowlist, because a review pulls untrusted content into a credentialed context.
- Every thread reply names the lane that produced it (GLM, or the pi-codex fallback), the same
visibility R6 requires, so engineers and the owner can always see what is serving.
- Team usage is metered per submitter under a daily cost bound (C3); when the bound is reached
the bot says so in the thread instead of silently dropping the request.

Acceptance: an engineer other than the owner tags the bot with a pull request link and receives
threaded findings produced by the GLM lane, with the lane named in the reply, the request
metered, and an out-of-allowlist link refused with a clear message.

## C1. Crosscheck completes in 20 to 30 minutes

Status: NOT ADDRESSED.
Expand Down Expand Up @@ -290,11 +360,16 @@ rather than daily, so nothing today refuses a single expensive day.
Workers also do not deallocate on idle.
Compute is released only when an exact release receipt is followed by a controller `reconcile`,
and the sole self-acting bound is a per-VM shutdown schedule at a wall-clock deadline.
Four worker slots are assigned with no release proof, one of their VMs is running with no live
task, and one validation work disk is unattached and stranded.
The four stranded worker slots were cleaned to zero on 2026-08-19 through the new surrender
lane, and the unattached validation disk by the owner's directed delete, but only by hand:
wkr-04 had
idled about four hours with no release proof until its TTL fired, which is the live example the
idle-release work exists to remove.

Work: a daily spend bound that refuses a mutation once the day's spend crosses it, and an idle
release path so an assigned worker whose task ended returns its compute without a human.
Work: a daily spend bound that refuses a mutation once the day's spend crosses it; an idle
release path so an assigned worker whose task ended returns its compute without a human; and,
once R10 exists, the per-submitter daily metering and the listener's standing cost that R10
books here.

Acceptance: a day cannot cross the bound without an explicit operator override, and a worker whose
task ended releases and deallocates unattended.
Expand All @@ -304,11 +379,14 @@ task ended releases and deallocates unattended.
1. R8, done 2026-08-19.
2. C2, because contention blocks demonstrating anything at scale. One of its three changes landed.
3. R2/R3, the largest architectural gap and the requirement most misread by the current build.
4. R6, whose direction is decided (Kimi-K2.7-Code on Foundry) and which no longer needs an owner login.
4. R6, whose direction is decided (GLM-5.2 on the Fireworks Foundry lane) and which no longer
needs an owner login.
5. R4, which needs the runner caller built and one validation cell closed.
6. R5.
7. C1, measured before it is changed.
8. R9, which is the proof of the rest.
9. R10, the Slack team exposure, which needs R6's lane and can be pulled forward right after R6
if the owner wants engineers on it sooner.

## Standing constraints

Expand Down
Loading