diff --git a/docs/azure-requirements.md b/docs/azure-requirements.md index 69b4ee8effd..110fbd6fd4d 100644 --- a/docs/azure-requirements.md +++ b/docs/azure-requirements.md @@ -36,11 +36,19 @@ Added the same day: - A cost guard, so that a day's spend cannot quietly reach 100 dollars. - "For cross-check, just use the same pi fleet (or copy it in, whatever works)." -Amended by the owner on 2026-08-19: the crosscheck requirement is model-family bidirectionality, -not the literal codex/claude pairing quoted above. The second reviewer family is Kimi-K2.7-Code -on Azure AI Foundry, reviewing codex-authored work; claude-authored work keeps the codex -crosscheck. pi-anthropic was rejected the same day (API pricing), as was an interim same-day -Claude-Code-CLI-subscription direction. Details and work in R6. +Amended by the owner on 2026-08-19: the crosscheck requirement is that no author's work is +reviewed only by its own model family, not the literal codex/claude pairing quoted above. +The second reviewer family is GLM-5.2 on Azure AI Foundry. +pi-anthropic was rejected the same day (API pricing), as was an interim same-day +Claude-Code-CLI-subscription direction, and an interim Kimi-K2.7-Code pick was revised to +GLM-5.2 the same day on model quality. Details and work in R6. + +Amended again by the owner later on 2026-08-19: crosscheck routes through GLM only. A single +reviewer family outside both author families satisfies the paradigm for every author, unbinds +review capacity from the pi-codex subscription profiles the fleet's own work depends on, and +makes the lane scalable for firstmate and for engineers' on-demand use. pi-codex stays as an +explicit fallback whose activation must be easy to see. no-mistakes stays on pi-codex. +The owner also directed a Slack v1 team exposure of crosscheck, recorded as R10. ## R1. Crewmates run in Azure @@ -87,7 +95,9 @@ A stranding strand was fixed: a passed run carrying a short receipt set is now d so the cell collects and retains legibly. That buys legibility, not closing, and the root cause (an in-cell bridge producing no receipts) is untouched, so a demoted cell still cannot `close`. -One stranded work disk is on the subscription today. +The one stranded work disk that sat on the subscription was deleted by the owner's direction on +2026-08-19 after no cell record could be found anywhere for the sanctioned close, so the lane's +inability to close now has no live residue, only the unfixed root cause. The runner-offload lane (`FM_AZURE_RUNNER_REMOTE_CLASSES`) has no caller: nothing in the repository sets `FM_AZURE_RUNNER_TASK`, `FM_AZURE_RUNNER_GENERATION`, or `FM_AZURE_RUNNER_CONFIRM_SUBSCRIPTION`, and the placeholders in `docs/azure-runner.md` are an @@ -108,9 +118,11 @@ separately, one validation cell reaches `close` with its worktree disk released. Status: PARTIAL. -Crosscheck is done: eight pi profiles across eight distinct upstream accounts, projected into -single-profile account homes by `bin/fm-pi-account-home.py`, with the roster repointed and read -back through the real `bin/fm-crosscheck.py` reader and policy screen. +Crosscheck on the pi fleet is done at the roster level: eight pi profiles across eight distinct +upstream accounts, projected into single-profile account homes by `bin/fm-pi-account-home.py`, +with the roster repointed and read back through the real `bin/fm-crosscheck.py` reader and +policy screen. Under the second 2026-08-19 amendment this roster is now the dormant crosscheck +fallback; the pi fleet's primary duties are authors and no-mistakes. Crewmate placement is not wired: the cell image carries pi, but placement does not select across the eight profiles. @@ -119,9 +131,9 @@ tool and the account-lease identity already present in the worker request path. Acceptance: concurrent crewmates run on distinct pi profiles with no account collision. -## R6. Crosscheck is bidirectional across model families +## R6. Crosscheck reviews outside the author's model family -Status: NOT DONE. Direction decided by the owner 2026-08-19 (see the amendment above). +Status: NOT DONE. Direction decided by the owner 2026-08-19 (see the second amendment above). The requirement is that no author's work is reviewed only by its own model family. The roster was instead made eight reviewers all on `openai-codex`, by reading "just use the same @@ -133,34 +145,61 @@ pi-anthropic is dead twice over: pi's anthropic OAuth authenticates as its own c own token endpoint, so a Claude CLI refresh token cannot be spent by pi, and a fresh pi-anthropic login would bill at API pricing, which the owner rejected. An interim directive the same day routed claude reviews through the Claude Code CLI on the -owner's subscription profile; the standing decision superseded it: the second reviewer family is -Kimi-K2.7-Code on Azure AI Foundry (Direct-from-Azure lane; at decision time deployable Global -Standard from eastus with tool calling and published pay-per-token pricing - re-verify at deploy -time), chosen for tool calling, the Microsoft-hosted custody lane, and lineage independent of -both OpenAI and Anthropic. -Codex-authored work is reviewed by a Kimi-backed reviewer; claude-authored work keeps the codex -crosscheck, which R5 records as proven at the roster level while R9 still owes the live proof. -The model pick is explicitly provisional: reevaluate after live review data (GLM-5.2 was the -runner-up; the comparison is in the owner's evidence folder, -R6-FOUNDRY-RESEARCH-2026-08-19.md). - -No owner login is needed anymore: the Kimi lane authenticates with a Foundry deployment and an +owner's subscription profile; a later interim pick was Kimi-K2.7-Code; the standing decision is +GLM-5.2 through the Fireworks AI lane on Foundry (`FW-GLM-5.2`, pay-per-token Data Zone Standard +US; at decision time it carries tool calling, a 1M-token context, and `reasoning_effort` with +`max` as the default - re-verify at deploy time), chosen on model quality with lineage +independent of both OpenAI and Anthropic. +The custody trade was accepted by the owner on 2026-08-19 knowing its exact shape: Microsoft +disclaims data handling for the Fireworks lane (data is shared between Microsoft and Fireworks +and processed on Fireworks infrastructure inside the US data zone), and the zero-data-retention +promise (volatile memory only, no logging by default, on the chat-completions surface) is +Fireworks' own policy, not Microsoft's. The lane must therefore use plain chat completions +only; the Responses API retains data for 30 days under its default store flag and is forbidden +here. +Billing was verified before acceptance: Fireworks pay-per-token on Foundry bills as Azure +consumption (a feature registration, not a Marketplace SaaS purchase), is MACC-eligible, and +Microsoft's own Startups material states startup credits apply to exactly this SKU; a small live +spend must confirm the credit decrement before volume. Fireworks pay-per-token models can retire +on 15 days notice, which is an accepted operational risk and one more reason the fallback below +stays armed. +Later the same day the owner simplified the routing: ALL crosscheck reviews route through the +GLM lane, for codex-authored and claude-authored work alike. GLM belongs to neither author +family, so single-family review is avoided for every author with one reviewer lane, and review +capacity stops competing with the pi-codex subscription profiles that no-mistakes and the +author fleet consume. The pi-codex roster (which R5 records as proven at the roster level while +R9 still owes the live proof) is retained as a dormant fallback behind a config flip, never +deleted; every review must name the lane that produced it, and a status read must show whether +GLM is serving or the fallback is active, so a silent fallback is impossible. +Fallback operation is a recorded degradation, not free service restoration: with the fallback +active, codex-authored work is reviewed by its own family again (the flip therefore includes +`config/crosscheck-same-model` on for the duration, which the policy screen otherwise refuses), +which is exactly the defect this requirement removes - accepted only while GLM is unavailable, +and one more reason fallback activation must be loud. +The model pick is explicitly provisional: reevaluate after live review data (Kimi-K2.7-Code was +the prior same-day pick; the comparison and the custody/billing verification are in the owner's +evidence folder, R6-FOUNDRY-RESEARCH-2026-08-19.md). + +No owner login is needed anymore: the GLM lane authenticates with a Foundry deployment and an api-key, not a subscription session. -The Kimi credential is an api-key, not a pi OAuth slot, so the literal `openai-codex` slot key in +The GLM credential is an api-key, not a pi OAuth slot, so the literal `openai-codex` slot key in `bin/fm-pi-account-home.py`, `bin/fm-crosscheck.py` (`inspect_pi_credential`, `account_identity`), and the Azure credential archive in `bin/fm-crosscheck-azure.py` stays as it is - those three still refuse any non-codex pi OAuth slot, which no longer blocks R6 and remains the recorded constraint if a second pi OAuth provider is ever added. The identity question those tools answered with `accountId` still needs an answer for an api-key -credential: reviewer identity for the Kimi lane must bind the Foundry resource and deployment, +credential: reviewer identity for the GLM lane must bind the Foundry resource and deployment, since an api-key carries no account identity of its own. -Work: deploy `Kimi-K2.7-Code` in a Foundry resource and store the key in the fleet's secret -custody, never in the repo; a pi custom provider entry (`models.json` `baseUrl` + +Work: register the `Fireworks.EnableDeploy` subscription feature, deploy `FW-GLM-5.2` +(pay-per-token Data Zone Standard) and store the key in the fleet's secret custody, never in +the repo; confirm by a small live spend that the charge decrements startup credit; verify the +endpoint passes `reasoning_effort` through and pin the lane to chat completions only; a pi +custom provider entry (`models.json` `baseUrl` + `openai-completions` api) pointing at the resource's OpenAI-compatible `/openai/v1` endpoint with the deployment name as the model id; verify pi tolerates `reasoning_content` in streamed deltas before rollout; extend `bin/fm-crosscheck.py` `allowed_profiles` and roster validation to -carry the Kimi lane (today they pin pi to `("pi", "gpt-5.6-sol", "xhigh")` and would refuse it) +carry the GLM lane (today they pin pi to `("pi", "gpt-5.6-sol", "xhigh")` and would refuse it) with `config/crosscheck-same-model` off; extend the `bin/fm-crosscheck-azure.py` endpoint allowlist and credential archive for the Foundry host and api-key shape; define reviewer identity for api-key credentials as above; retire or re-point the interim claude reviewer @@ -168,12 +207,17 @@ artifacts (the `("claude", "claude-opus-5", "xhigh")` `allowed_profiles` entry, `api.anthropic.com` allowlist entry, and the claude-profile boot copy described in `docs/azure-crosscheck.md`); review guards sized to the model's context window at deploy time: a strict findings schema, path-existence validation before filing, and a per-review context cap; -and routing so codex-authored changes draw the Kimi reviewer while claude-authored changes draw -codex reviewers. - -Acceptance: a codex-authored change is reviewed by a Kimi-backed reviewer and a claude-authored -change by a codex-backed reviewer, on distinct credentials with bound reviewer identities, with -the same evidence discipline as the codex lane. +routing so every crosscheck draws the GLM reviewer, with the pi-codex roster behind a config +flip as fallback (the flip sets `config/crosscheck-same-model` on, accepting same-family review +of codex-authored work as the recorded degraded mode while it is active); and lane visibility: +the review evidence and report name the reviewing lane, and a status command answers whether +GLM is serving or the fallback is active. + +Acceptance: a codex-authored change and a claude-authored change are each reviewed by a +GLM-backed reviewer with bound reviewer identity and the same evidence discipline as the codex +lane; the fallback flip to pi-codex is demonstrated once, its activation is visible in the review +evidence and the status read, and the demonstration records the degraded same-family mode it +accepts for codex-authored work. ## R7. Everything is logged in @@ -182,9 +226,8 @@ Status: HOLDS, through R8. The eight pi profiles renew on their own now, which is R8. Two of the three profiles in `~/.local/share/agent-fleet/accounts/claude/` hold blanked, length-zero tokens; the third is `refreshable` with material declared valid to 2026-09-10. -None of the three is needed for R6 anymore: the Kimi lane authenticates with a Foundry -deployment key, and claude-authored work is reviewed by the codex fleet, so R7 holds with no -owner login outstanding. +None of the three is needed for R6 anymore: the GLM lane authenticates with a Foundry +deployment key and reviews all authors, so R7 holds with no owner login outstanding. ## R8. Auth refreshes on its own @@ -235,6 +278,33 @@ The existing smoke assignments cannot be reused because their `repository_genera commit that exists. Proving it requires a real spawned crewmate task that commits. +## R10. Crosscheck is exposed to team engineers through Slack + +Status: NOT DONE. Directed by the owner 2026-08-19; builds after R6. + +The owner's v1 shape: an engineer tags the crosscheck bot in a Slack channel with a pull request +link; the GLM lane reviews it; the bot posts the findings as a thread reply on the engineer's +own message. No engineer wires up a harness or touches an endpoint, and the same path works for +deliberate on-demand use. Cursor Bugbot continues to run for engineers' pull requests (it stays +disabled on the owner's), so this lane complements rather than replaces it. + +Constraints the build must honor: + +- The listener uses Slack Socket Mode, so no public inbound endpoint is added to the private + lane posture. It is a resident process; where it runs is decided at build time and its + standing cost is recorded under C3. +- v1 accepts pull request links only, and only for repositories in the organization allowlist. + The bot's repository read credential must never be pointed at a repository outside that + allowlist, because a review pulls untrusted content into a credentialed context. +- Every thread reply names the lane that produced it (GLM, or the pi-codex fallback), the same + visibility R6 requires, so engineers and the owner can always see what is serving. +- Team usage is metered per submitter under a daily cost bound (C3); when the bound is reached + the bot says so in the thread instead of silently dropping the request. + +Acceptance: an engineer other than the owner tags the bot with a pull request link and receives +threaded findings produced by the GLM lane, with the lane named in the reply, the request +metered, and an out-of-allowlist link refused with a clear message. + ## C1. Crosscheck completes in 20 to 30 minutes Status: NOT ADDRESSED. @@ -290,11 +360,16 @@ rather than daily, so nothing today refuses a single expensive day. Workers also do not deallocate on idle. Compute is released only when an exact release receipt is followed by a controller `reconcile`, and the sole self-acting bound is a per-VM shutdown schedule at a wall-clock deadline. -Four worker slots are assigned with no release proof, one of their VMs is running with no live -task, and one validation work disk is unattached and stranded. +The four stranded worker slots were cleaned to zero on 2026-08-19 through the new surrender +lane, and the unattached validation disk by the owner's directed delete, but only by hand: +wkr-04 had +idled about four hours with no release proof until its TTL fired, which is the live example the +idle-release work exists to remove. -Work: a daily spend bound that refuses a mutation once the day's spend crosses it, and an idle -release path so an assigned worker whose task ended returns its compute without a human. +Work: a daily spend bound that refuses a mutation once the day's spend crosses it; an idle +release path so an assigned worker whose task ended returns its compute without a human; and, +once R10 exists, the per-submitter daily metering and the listener's standing cost that R10 +books here. Acceptance: a day cannot cross the bound without an explicit operator override, and a worker whose task ended releases and deallocates unattended. @@ -304,11 +379,14 @@ task ended releases and deallocates unattended. 1. R8, done 2026-08-19. 2. C2, because contention blocks demonstrating anything at scale. One of its three changes landed. 3. R2/R3, the largest architectural gap and the requirement most misread by the current build. -4. R6, whose direction is decided (Kimi-K2.7-Code on Foundry) and which no longer needs an owner login. +4. R6, whose direction is decided (GLM-5.2 on the Fireworks Foundry lane) and which no longer + needs an owner login. 5. R4, which needs the runner caller built and one validation cell closed. 6. R5. 7. C1, measured before it is changed. 8. R9, which is the proof of the rest. +9. R10, the Slack team exposure, which needs R6's lane and can be pulled forward right after R6 + if the owner wants engineers on it sooner. ## Standing constraints