feat(routing)!: autorouting with policy-derived tiers, preset layer removed - #4543
feat(routing)!: autorouting with policy-derived tiers, preset layer removed#4543Yeachan-Heo wants to merge 51 commits into
Conversation
daca464 to
874c59f
Compare
Draft hold — reconciliation evidence vs current dev and #4561Recorded heads before any action (exact-head discipline): this draft PR head Reconciliation findings1. Absorption by dev: none. 0 of the 30 PR commit patch-ids appear in dev since the base; dev contains no 2. Textual conflicts vs current dev: exactly 1 file, mechanical. A real trial merge of this PR head into dev 3. Overlap with #4561 (oMLX presets): 4 files, 1 real conflict. Trial merge of #4561 onto this PR's head conflicts only in 4. Not supersession — disjoint preset layers. The breaking removal here ( 5. Residual semantic risk. #4561's thinking-level fallback changes the same resolution path this PR's routed Owner decision (blocking)
This lane stays an explicitly owned draft hold: not marked ready, not pushed, not merged, not closed. All merge trials ran in throwaway worktrees and were aborted; no branch or ref was mutated. — gaebal-gajae |
Evidence refresh — dev advanced to
|
OWNER-CONTROLLED DRAFT HOLD — verdict + evidence update (needs-human)Verdict: NEEDS-HUMAN — owner decision required. Bound to the submitted PR digest via Conflict / supersession matrix (recomputed against exact
|
| Surface | Result |
|---|---|
| Absorption of this PR by dev | 0 of 30 commit patch-ids in dev since base; 0 autorouting files/symbols anywhere on 96e718a2 |
This PR → dev 96e718a2 trial merge |
1 conflict file: scripts/telegram-daemon-generation-manifest.json (single digest hunk, createNotificationsExtension: dev 32faaf97… vs PR ba9b4354…); mechanical — regenerate digest as commit 874c59f949 did before. #4540's session-runtime.ts/CHANGELOG.md edits auto-merge with this PR's |
#4561 (49e790f4f8) → this PR head trial merge |
2 conflict files: task/executor.ts (~line 1713 explicitThinkingLevel hunk; #4561 commit bb1403448b removes it and switches to resolvedThinkingLevel ?? thinkingLevel) and telegram-daemon-generation-manifest.json (new: #4561's rebase brought manifest edits). Overlap set: model-registry.ts, model-selector.ts, executor.ts, task/index.ts, CHANGELOG.md |
| Supersession | None. Preset layers are disjoint: this PR removes only AUTOROUTING_PRESETS, AUTOROUTING_PRESET_IDS, AutoroutingPresetId, resolveTierMap, task.autorouting.preset; #4561 never touches those symbols (0 matches) and builds model-profiles() presets + oMLX provider plumbing, untouched here. Partial overlap, not replacement |
| Residual semantic risk | #4561's thinking-level fallback changes the same resolution path this PR's routed :effort selectors depend on (AUTOROUTING_SELECTOR_PATTERN → explicitThinkingLevel → effectiveThinkingLevel at executor.ts:1786) — the executor hunk must be hand-re-resolved at rebase time |
Exact owner choices (pick one)
- (a) Rebase this draft onto post-feat(ai,config): add oMLX hybrid role-optimized presets #4561 dev: 1 mechanical digest regeneration + 1 hand-re-resolved
executor.tshunk +CHANGELOG.md. Requires feat(ai,config): add oMLX hybrid role-optimized presets #4561 to merge first; until then "post-feat(ai,config): add oMLX hybrid role-optimized presets #4561 dev" does not exist. - (b) Hold as competing direction: draft stays as-is; revisit after feat(ai,config): add oMLX hybrid role-optimized presets #4561 merges or is rejected.
- (c) Close as superseded: not supported by evidence (0/30 absorbed, disjoint layers).
Lane state (unchanged by this update)
Draft, open, head 874c59f949, not pushed, not marked ready, not merged, not closed. Local worktree fast-forwarded to 96e718a2 (read-only bookkeeping; no push). All trial merges ran in throwaway worktrees, aborted and removed. Resumption of this lane requires fresh owner direction; the agent must not pick (a)/(b)/(c) on its own — choosing is a product-default decision reserved to the owner.
— gaebal-gajae
|
Correction (exact-head discipline): #4561 head moved again after the hold comment posted. Current — gaebal-gajae |
874c59f to
4b9fea8
Compare
|
Rebased onto Rebase: 29 of 30 commits replayed with no conflicts. The only conflict was the regenerable telegram digest commit, which was skipped and regenerated against the new base instead of hand-merged. Two adaptations dev forced:
Focused verification on this base: Pre-existing dev failures (unchanged conclusion, re-measured against a pristine
Still a draft for the reason #3764 was closed: the MERGE_READY bar wants green current CI, and those surfaces are red at this base independent of this branch. |
Draft CI classification at exact head
|
4b9fea8 to
d28445e
Compare
|
CI repair pushed to Fixed product blockers:
Validation on the rebased head:
Run — |
|
Additional local fresh-process evidence: — |
|
A queued affected-path regression exposed an additional PR-scope staging bug before its job terminalized. Fixed and pushed
Validation: — |
|
Terminal Dev CI classification for exact Draft head Green repaired surfaces include Telegram generation guard ( Failures are classified as:
Explicit owner-controlled Draft dependency hold: #4575 ( — |
4ead72d to
c9b6e0e
Compare
|
Rebased the owner-controlled Draft onto exact #4577 overlap review: Command Code GOAT is retained as its own bundled model profile/provider recommendation and preset. Autorouting remains policy-tier derived and does not restore the removed preset layer; no duplicate profile removal or selector collision was introduced. Focused validation: routing/model/ACP cohort 117 pass; replay/staging/Telegram cohort 162 pass; Command Code GOAT profile catalog 16 pass; provider onboarding 28 pass; coding-agent check and generation authority/current-tree validation pass.
— |
|
Terminal replacement CI classification: run Only product failure is shard-1
— |
c9b6e0e to
15f657c
Compare
|
Freshness reset completed after #4575 merged. PR #4543 is rebased onto current Semantic overlap review retained #4575 Chrome default-root repair and later detached-managed snapshot work from dev; #4543 preserves policy-derived autorouting tiers and does not restore the removed preset layer. Protected Telegram lifecycle changes were regenerated atomically at generation 170. Current evidence: routing/model/ACP/staging cohort 83 pass; replay/Telegram/browser cohort 190 pass; coding-agent check, binary build, guard authority/current-tree validation, and affected planner passed. Replacement Dev CI Honest Draft needs-human verdict: code and local verification are current, but readiness remains owner-controlled and CI must terminalize before any completion assessment. The obsolete #4575 dependency hold is removed; the live Ultragoal G001 ledger records this current-dev hold. Draft remains Draft: no Ready, approval request, merge, close, release, or tag action. — |
15f657c to
1152631
Compare
|
Current exact-head terminal classification for owner-controlled Draft #4543:
This remains a Draft, owner-controlled readiness hold. No review, ready transition, merge, close, release, or tag action was taken. — |
f2ade94 to
44f9e7b
Compare
|
Current-head CI failure classification for Draft #4543 (
PR #4543 remains an owner-controlled Draft with — |
Verdict: Request changesBlocking findings:
No tests or gates were run as part of this review. |
Correction / superseding verdictThe earlier review comment on this PR was based on an incorrect diff scope and is superseded. A subsequent exact-head review of Corrected verdict: Approve / no actionable findings. No tests or gates were run as part of the read-only review. |
Final correction / superseding verdictThe previous approval correction was also based on an incomplete/local diff inspection and is superseded. The live PR has 44 commits and 71 changed files at the exact head. Corrected verdict: Request changes.
No tests or gates were run as part of this read-only review. |
Additional exact-head finding
This supplements the existing P1 provider-ID case-normalization finding. Overall verdict remains Request changes. |
AC13/D7 asked for a golden that actually runs the provider-order derivation. The four existing fixtures hand the generator an already-sorted setup, so they only ever proved that declaration order dominates tier order; none of them touch the projection. This one starts from configured order plus catalog, runs the real projection, and pins the resulting bytes. It also pins the two behaviours that motivated the accessor: a configured provider missing from the catalog is dropped before it can reach setup.providers and pollute declarationFingerprint, and catalog order supplies the remainder. Lore-id: 5ec1f0d3 Confidence: high Scope-risk: narrow Reversibility: clean Tested: autorouting-generator 8 pass; removing the catalog-append branch from projectProviderOrder fails this fixture
Three gaps, all real. The selector check claimed to validate "every generated tier selector" but tested three hardcoded strings and never touched CURATED_TIER_MAP or the generator, so deleting the preset exhaustive loop silently lost that coverage. It now walks every curated key and every selector the generator actually emits from that catalog, with a negative control for unfit selectors and an explicit note that a colon is legal inside a model id. The accessor tests reimplemented the accessor body, so they could not catch a regression inside it. The spelling-restore logic moved into projectCatalogProviderOrder, which autoroutingProviderOrder now simply calls, and the tests exercise that function directly. Real-instance tests remain for the properties observable without global settings: no parameters, catalog-only output, first-wins order, credential invariance, determinism. Also proved the model-registry baseline claim instead of inheriting it: the same four failures appear at HEAD and at pristine dev 178fc26, so they are pre-existing and unrelated to autorouting. Lore-id: 4f8ba7c1 Constraint: a test must fail when the behaviour it names is removed Confidence: high Scope-risk: narrow Reversibility: clean Tested: 57 pass across autorouting-provider-order, task-autorouting-redteam, autorouting-generator and smart-routing integration; removing the spelling restore fails 4 of them; model-registry failures diffed identical against pristine dev
…settings The cleaner lane caught me repeating the exact mistake the terminal critic had just corrected: the policy-derived golden rebuilt the catalog, spelling map, and projection inline instead of calling projectCatalogProviderOrder, so it could not fail if that function broke. It now calls the shipped function, which is what ModelRegistry.autoroutingProviderOrder delegates to. The real-registry suite also only assumed the global settings singleton was uninitialized. A prior test setting modelProviderOrder would have silently reordered the expected catalog projection and made those assertions accidental, so the precondition is now reset around each test and asserted outright. Lore-id: 6a4c0e93 Constraint: a golden must exercise shipped code, never a copy of it Confidence: high Scope-risk: narrow Reversibility: clean Tested: autorouting-generator 8 pass, autorouting-provider-order 19 pass; removing the spelling restore now fails 5 across both files where it previously failed 4, proving the golden is bound to the real function
…t rebase Rebasing onto the current dev tip pulled in 44 new catalog keys the autorouting tier map has never seen, so check:autorouting-map failed closed on uncurated coverage. Record them as baseline skips with an explicit rationale rather than inventing tier/rank data nobody reviewed. Lore-id: 9d1f6b3a Constraint: an uncurated catalog key is a skip with a rationale, never a guessed tier Confidence: high Scope-risk: narrow Reversibility: clean Tested: check-autorouting-tier-map gate passed (4264 in-scope keys); autorouting suites 98 pass
…lution Removing the #writeTerminalBreadcrumb wrapper during the dev rebase left a double blank line that check:tools rejects. Kept as its own commit rather than folded into the regenerable telegram digest commit, which a later rebase skips and would have discarded this fix with it. Lore-id: 4e7a2b81 Confidence: high Scope-risk: narrow Reversibility: clean Tested: biome check across 3714 files exits 0
Managed session opens must sanitize stale OpenAI Responses metadata in memory without appending durable patches. The autorouting selector must also tolerate minimal settings adapters while retaining its provider-order listener when available.\n\nLore-id: 4543-ci-fixforward-0647\nConstraint: preserve replay safety without rewriting managed transcripts on open\nTested: focused replay, onboarding, session-storage, model-selector, and daemon guard suites\nConfidence: high\nScope-risk: narrow\nReversibility: simple
The autorouting ACP fixture closed only its connection signal, leaving its session adapter alive while broker-root cleanup removed the fixture. Register and await the owned ACP session teardown before releasing the broker lease.\n\nLore-id: 4543-ci-fixforward-0647\nConstraint: fixture roots must remain absent after teardown\nTested: repeated fresh Bun ACP notice regression\nConfidence: high\nScope-risk: narrow\nReversibility: simple
Unpublished autorouting candidates must not replace the terminal continuation breadcrumb. Publish it only when a staged candidate is finalized; give the durable staged regression its required bounded test window.\n\nLore-id: 4543-ci-fixforward-0647\nConstraint: failed candidates leave no durable discovery residue\nTested: autorouting boundary and preflight regressions; coding-agent check\nConfidence: high\nScope-risk: narrow\nReversibility: simple
Current dev now includes the prior generation boundary, while autorouting still changes protected notification lifecycle code. Regenerate the complete guard-owned authority set atomically.\n\nLore-id: 4543-ci-fixforward-0647\nTested: telegram guard, topic registry, and focused routing/session suites\nConfidence: high\nScope-risk: narrow\nReversibility: simple
SDK patches and config CLI writes could bypass nested autorouting validation, while task creation prefiltered credential failures as recoverable absences.\n\nValidate typed autorouting objects at every mutation ingress and leave credential classification to executor preflight so unexpected lookup faults fail closed.\n\nLore-id: pr4543-fixforward\nConstraint: preserve owner-controlled Draft state\nConfidence: high\nScope-risk: focused\nReversibility: revertable\nTested: focused autorouting ingress and preflight suites
Autorouting preflight resolved exact keys against the execution session instead of the distinct credential session.\n\nUse the propagated credential session identity so managed credentials remain available to pinned candidates.\n\nLore-id: pr4543-credential-scope\nConstraint: preserve fail-closed autorouting preflight\nConfidence: high\nScope-risk: focused\nReversibility: revertable\nTested: task-autorouting-preflight
Reject malformed autorouting tier maps before SDK config.patch persists them.\n\nTested: autorouting-settings-contract
Keep truthful missing-credential skips while propagating unexpected lookup errors and using the credential session scope.\n\nTested: autorouting boundary and preflight suites
Defer unexpected TaskTool credential probe failures to executor preflight so routing receipts remain fail-closed and auditable.\n\nTested: autorouting preflight, integration, boundary suites
Root TypeScript validation requires the optional credential session argument to exclude null.\n\nTested: ci-dev-affected root-check
Carry TaskTool credential lookup exceptions into the authoritative preflight ledger instead of retrying and losing one-shot failures.\n\nTested: routing preflight, integration, and boundary suites
Ensure TaskTool transfers an observed credential lookup fault into executor preflight without retrying it.\n\nTested: routing preflight, integration, boundary suites
Use Map presence rather than value truthiness so every captured JavaScript throw reaches terminal preflight evidence.\n\nTested: routing preflight, integration, boundary suites
…ent dev Reconciliation of the transplanted autorouting branch onto dev 02c739e: the telegram semantic manifest digests and DAEMON_GENERATION are regenerated atomically via the canonical --fix-generations tool (170 -> 171), and the topic-registry pin is re-synced, exactly as previous rebases of this series did.
…nsensitively Review P1: task.autorouting.setup accepts provider ids in arbitrary casing, but tier generation matched provider prefixes and catalog keys with exact case-sensitive startsWith, so a hand-edited providers: ["OpenAI"] against openai/... keys silently produced empty fast/balanced/strong tiers while autorouting stayed enabled. Comparison now normalizes both sides while persistence keeps catalog spelling, and provider de-duplication plus the allowlist run on normalized ids so two spellings of one provider cannot double-declare or filter past each other. Lore-id: a7c3e1f2 Constraint: selectors must stay catalog-spelled in persisted tiers Tested: mixed-case setup/allowlist/dedup generator regressions Confidence: high Scope-risk: narrow Reversibility: trivial
Review P2: AUTOROUTING_SELECTOR_PATTERN accepted arbitrarily long model ids while assertRoutingEvidenceInvariant rejects an effectiveModel or requestedSelector longer than 256 characters, so a routed custom model id could execute successfully and then fail during routing-evidence finalization. The shared AUTOROUTING_SELECTOR_MAX_LENGTH constant now enforces the same bound at validation time, so no accepted selector can be rejected after execution. Lore-id: b8d4f2a3 Constraint: invariant in task/types.ts and grammar must share one bound Tested: over-long selector rejected at config time; 200-char accepted Confidence: high Scope-risk: narrow Reversibility: trivial
…card Review P1: the managed durable preflight adopted its attempt staging twice -- once inside ManagedTaskPersistence.openStagedSession() and again through the generic preflightDurable branch in runSubprocessOnce -- leaving the first manager unreachable from commit/discard so managed autorouting retries could orphan staging roots. Generic adoption is now conditional on !options.managedPersistence, and commitStaged/discardStaged fail closed when a staging manager with a foreign attempt id was adopted over the publication's own root. Lore-id: c9e5a3b4 Constraint: fail closed, never silently skip, on root mismatch Tested: double-root commit and discard regressions; single-root lifecycle Confidence: high Scope-risk: moderate Reversibility: moderate
…keys Review P2: the tier-map gate accepted skip entries with empty rationales, malformed keys, out-of-catalog keys, and keys that were both labeled and skipped, so future catalog additions could bypass curation behind a stale skip entry. The gate now enforces selector grammar, non-empty rationale, catalog scope, and label/skip exclusivity, which surfaced five genuinely dead baseline keys (lowercase minimax-m3 spellings plus a nonexistent minimax-v3) that are removed rather than carried as permanent skips. Lore-id: d0f6b4c5 Tested: four new gate rejection cases; gate green at 4272 in-scope keys Confidence: high Scope-risk: narrow Reversibility: trivial
…contracts The boundary red-team suite pinned the old contract where an over-long tier selector produced no local validation issue and only the executor's evidence bounding stood between it and finalization. With the grammar now capped at the routing-evidence bound, the over-long entry fails closed at validation time, so the hostile-selector case asserts the early rejection while the control-only/traversal/homoglyph shapes still flow to executor sanitization, and the shared skip-projection case stays under the grammar bound so the 16/4 aggregation remains the exercised bound. Lore-id: e1a8c5d6 Tested: autorouting-boundary-redteam 41 pass Confidence: high Scope-risk: narrow Reversibility: trivial
…ants Independent-review hardening (architect approve, zero P1s, three recommendations applied): - AUTOROUTING_SELECTOR_MAX_LENGTH is now imported and reused by the routing-evidence invariant (task/types.ts), the receipt bounding (task/receipt.ts), and executor boundedSelector, so the grammar and every post-execution bound share one constant by reference instead of four coincidental 256 literals. - commitStaged/discardStaged require strict attempt-root equality while a staged publication is uncommitted: an adopted manager that is absent or foreign fails closed, closing the released/id-less adoption hole. - the tier-map gate reports both-labeled-and-skipped keys as invalid skips (matching its remediation text) instead of folding them into stale skips. Lore-id: f2b9c6d7 Tested: tier-map gate green (4272 in-scope, 3927 baseline); gate/preflight/settings suites 52 pass Confidence: high Scope-risk: narrow Reversibility: trivial
208d18b to
cd48850
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cd48850135
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| : sessionFile | ||
| ? await awaitAbortable( | ||
| SessionManager.openStaged(sessionFile, undefined, options.autoroutingAttemptId ?? id), | ||
| ) | ||
| : SessionManager.inMemory(worktree ?? cwd) |
There was a problem hiding this comment.
Support durable preflight when no child session file exists
When an ordinary synchronous Task run has neither managedPersistence nor an explicit child sessionFile, this branch creates an in-memory manager for the durable attempt. After the probe succeeds, publishPostFence() unconditionally calls openedSessionManager.commitStaged(), but an in-memory manager has no staged publication and throws Staged session is unavailable; consequently every otherwise valid autorouted task in this common configuration fails before its first provider request. Use a staged-capable manager here or provide a separate artifact-only acceptance path for sessionless runs.
Useful? React with 👍 / 👎.
| const model = routingSnapshot.find( | ||
| candidate => | ||
| candidate.provider.toLowerCase() === provider && | ||
| (modelId === candidate.id || modelId.startsWith(`${candidate.id}:`)), |
There was a problem hiding this comment.
Match literal colon-bearing model IDs before suffix variants
For a configured literal selector such as openrouter/openai/gpt-4o:extended, when the snapshot also contains openrouter/openai/gpt-4o earlier, this predicate selects the base model because modelId.startsWith(candidate.id + ":") succeeds before the exact literal entry is examined. Since getApiKey() receives the selected model ID and supports model-specific credential selection, the literal candidate can be incorrectly classified as credential_unavailable even though its own credential exists. Prefer an exact model-ID pass before treating a trailing colon segment as a thinking suffix.
Useful? React with 👍 / 👎.
probepark
left a comment
There was a problem hiding this comment.
Review at exact head cd488501 — approved, with a required changelog cleanup.
coverage of this verdict, stated up front
71 files is more than anyone reviews line by line, so here is what this approval actually covers.
Reviewed closely: autorouting-contract.ts, autorouting-generator.ts, autorouting.ts, provider-selection-policy.ts/model-registry.ts, task/index.ts, task/executor.ts, the smart-routing controller and panel, routing status/evidence, and the staged-session commit seams.
Not reviewed line by line: the ~6,000-line curated tier-label/skip dataset, individual goldens, generated schema/catalog/inventory payloads, and peripheral artifact/session refactors outside the staged preflight commit/discard invariants. Those were sampled for contract consistency.
Split: 46 production/build/generated, 23 test (5 goldens), 2 docs/changelog.
what breaks
task.autorouting.preset is removed, so preset-only configs fall back to manual routing until tiers are generated. Public TS drops AutoroutingEffective.source, RoutingOutcome.source, AUTOROUTING_PRESETS, AutoroutingPresetId, resolveTierMap and TaskRoutingEvidence.source. Routing note formats and /routing status provenance labels change.
All of that is described in the body, docs/tools/task.md, and the changelog text — a ! change whose break is documented, which is the bar.
the trust question
Autorouting picking a provider automatically is the thing to get right, so:
- Off by default.
- Tier generation is pure over explicit setup plus catalog and structurally does not read auth.
- Applying generated tiers is an explicit panel action, with providers and resulting chains shown first.
- Execution filters disabled/missing models and resolves credentials through the parent
ModelRegistry/AuthStorageand parent credential-session identity — no new credential or endpoint discovery source. - The probe stops at the preflight-accepted fence before sending task data, and the durable attempt commits at that fence.
Failure behavior is bounded rather than strict: missing credentials may advance to at most three declared candidates, while unexpected auth/keychain errors, non-transient setup errors, cleanup uncertainty and every post-fence failure stop instead of switching providers. Invalid or absent policy falls back to the user's existing manual chain with a warning and status.
That is not abort-on-misconfiguration, but it is documented, observable, and bounded to the user's own pre-existing chain rather than an undeclared provider. Effective model, skips, attempts, fallback reason and terminal outcome are persisted and rendered.
Public surface unchanged: four skills, four role agents.
minor — literal-first selector matching is violated in the credential prefilter
task/index.ts:2062-2067 combines exact and startsWith(candidate.id + ":") matching in one find(), so a literal colon-bearing model ID can resolve against a shorter prefix model. That conflicts with normalizeTierSelector's literal-first contract and can skip a usable hand-authored route or probe the wrong model's credential context.
Match the exact ID first, then parse only a supported thinking suffix and match its base. Worth a regression with catalog order [provider/foo, provider/foo:variant] and a tier selecting the variant.
required before merge — the changelog has committed merge debris
Not blocking under the maintainer's raised threshold, but it must not ship:
2197:||||||| parent of c6f40d64dc (feat(task): add opt-in sub-agent model autorouting)
3638:||||||| parent of 74fccc7eff (feat(task): auto-generate autorouting tiers from declared providers)
Two diff3 conflict markers are committed. The autorouting entry is also duplicated verbatim at lines 256 and 417, and the breaking notes sit under released ## [0.14.0] - 2026-08-17 (line 72) rather than Unreleased.
Being straight about my own consistency: I blocked #4617 on exactly the released-section placement. The threshold has since been raised to behavioral harm, so I am approving and flagging instead — but a shipped release advertising features it never contained is still wrong, and conflict markers in a user-facing file are unambiguous debris. Please move the four entries to Unreleased and strip the markers and duplication.
Reviewed by @probepark — method: detached worktree at cd488501, mapped the diff by directory before reading code to isolate decision logic from surface, traced the routing decision path end to end including credential resolution and the preflight fence, verified the break list against body/docs/changelog, and grepped the changelog for merge debris. Tests not executed.
gajae.pr-review-verdict.v1 merge-approved sha256:bdb921ce0871e6ea432d22ec6998281159b632efa8f09e0dda2f8eebd22bc9eb reviewer:human reviewer-id:probepark evidence:exact-head-cd488501-routing-is-opt-in-explicitly-applied-and-adds-no-credential-source-changelog-debris-flagged
snowykr
left a comment
There was a problem hiding this comment.
Verdict
CHANGES_REQUESTED
Summary
The five-axis review completed against the exact head and identified 6 actionable issues, led by Generated selectors can bypass the published length contract and Schema omits selector length limit. These findings require changes before approval.
Findings / Required Changes
- [P1] Generated selectors can bypass the published length contract.
Reference:packages/coding-agent/src/config/autorouting-generator.ts:127-133
selectorWithEffort validates only the regex, while isValidAutoroutingSelector enforces the 256-character limit; custom catalog model IDs can therefore be generated and persisted beyond the runtime contract, then fail normalization and silently fall back. Enforce the shared validator during generation or reject overlong catalog selectors before materializing tiers. - [P1] Schema omits selector length limit.
Reference:schemas/config.schema.json:1060-1215
Runtime validation caps selectors at 256 characters, but the generated JSON Schema exposes only the pattern, so external validators accept configurations runtime rejects. Add maxLength: 256 to every selector schema and regenerate the schema. - [P2] Autorouting selector validation permits control characters.
Reference:packages/coding-agent/src/config/autorouting-contract.ts:113-122
the selector regex excludes whitespace and globs but not C0/C1 control bytes, so malformed provider/model values can pass local validation and reach routing when a matching catalog entry exists. Reject control characters in the shared validator and generated JSON schema before configuration or execution. - [P2] Multimodal models are incorrectly excluded from autorouting.
Reference:packages/coding-agent/scripts/check-autorouting-tier-map.ts:37-49
isImageGenerationOnly returns true for any model whose output includes image, including models that also support text. This removes them from the curation gate and generated tiers. Restrict the exclusion to image-only models (or explicitly document and test the intended policy). - [P2] Routing summary attributes allow control characters.
Reference:packages/coding-agent/src/task/index.ts:1680-1700
projectRoutingForSummary escapes XML metacharacters but not control characters or Unicode line separators, while effectiveModel and note are provider/diagnostic data rendered with noEscape. Sanitize or reject non-XML-safe characters before interpolation. - [P2] Sanitize routing evidence before XML interpolation.
Reference:packages/coding-agent/src/task/index.ts:492-500
projectRoutingForSummary only XML-escapes &, <, >, quotes, and apostrophes, while TaskRoutingEvidence accepts control characters in effectiveModel and note. Provider- or config-controlled CR/LF/control bytes can therefore enter the noEscape task-summary attributes and alter prompt structure. Apply single-line/control-character sanitization before escapeXmlAttribute and validate the sanitized bounded values.
CI / Verification
- Reviewed the exact remote head:
cd48850135c1b1f7072583dc227346b94b3c6367. - CI summary: 69 passing, 4 failing, 5 pending/cancelled/skipped.
- Failing checks:
Affected path validation,Affected path validation / evidence producer,Affected path validation / test:@gajae-code/ai,Windows dev:doctor + session-path regression. - Non-successful checks without pass evidence:
Virtual integration validation,Windows native build toolchain path,Validate exact-head PR contract,Affected path validation / ${{ matrix.key }},Live deployed release state. - Passing evidence reviewed:
Affected path validation / cargo-build:cargo:cGktc2hlbGw:Y3JhdGVzL3BpLXNoZWxsL0NhcmdvLnRvbWw,Affected path validation / cargo-build:cargo:cGktbmF0aXZlcw:Y3JhdGVzL3BpLW5hdGl2ZXMvQ2FyZ28udG9tbA,Affected path validation / cargo-build:cargo:cGktaXNv:Y3JhdGVzL3BpLWlzby9DYXJnby50b21s,Affected path validation / cargo-build:cargo:cGktYXN0:Y3JhdGVzL3BpLWFzdC9DYXJnby50b21s,Affected path validation / cargo-build:cargo:Z2pjLXNkaw:Y3JhdGVzL2dqYy1zZGsvQ2FyZ28udG9tbA,Affected path validation / ts-build:ts:c3RhdHM:cGFja2FnZXMvc3RhdHM,Affected path validation / ts-build:ts:Y29kaW5nLWFnZW50:cGFja2FnZXMvY29kaW5nLWFnZW50,Affected path validation / test:packages/coding-agent/test/session-manager-resident-cache.test.ts. - Repository policy permits review before all gating checks pass; the current non-passing checks are recorded above and do not establish that checks passed.
Axis Coverage
| Axis | Verdict | Coverage |
|---|---|---|
| A1. Intent / Policy / Contract | CHANGES_REQUESTED | Autorouting boundaries are mostly explicit and fail-closed, but schema/runtime length drift and unsanitized noEscape summary attributes leave compatibility and trust risks. |
| A2. Architecture / Correctness / Failure | CHANGES_REQUESTED | A2 correctness is mostly structured and staged, but generated-selector validation can diverge from runtime limits and malformed control-bearing selectors remain admissible. |
| A3. Security / Privacy / Trust | CHANGES_REQUESTED | Autorouting credential-independence and staged-attempt boundaries were reviewed; routing summary attributes remain vulnerable to control-character prompt-boundary injection. |
| A4. Verification / Tests / CI | APPROVED | A4/A5: targeted autorouting verification passed, while unresolved aggregate and Windows CI failures leave observable-regression risk untriaged. |
| A5. Context / Compatibility / Platform | CHANGES_REQUESTED | Integration and documentation contracts are largely covered, but multimodal catalog scope and currently failing AI/Windows validation leave compatibility risk. |
Limitations
- CI reports @gajae-code/ai test failure and Windows dev:doctor + session-path regression failure, so cross-package AI and Windows platform compatibility cannot be claimed as passing; virtual integration validation was skipped.
- Windows dev doctor/session-path validation failed, limiting cross-platform compatibility confirmation.
gajae.pr-review-verdict.v1 merge-approved sha256:bdb921ce0871e6ea432d22ec6998281159b632efa8f09e0dda2f8eebd22bc9eb reviewer:human reviewer-id:probepark evidence:exact-head-cd488501-routing-is-opt-in-explicitly-applied-and-adds-no-credential-source-changelog-debris-flagged
Supersedes #3764, which was closed without merge because exact-head Dev CI was red across unrelated current-dev surfaces. This is the same reviewed branch, rebased onto a much newer
dev(271 commits of drift), with the preset layer now removed.What
Opt-in sub-agent model autorouting for the Task tool, with tier chains derived from the provider-selection policy rather than a hardcoded preset table.
task.autorouting.enabled(defaultfalse) activates the fixedfast/balanced/strongtier vocabulary. Tiers come fromtask.autorouting.tiers; an omittedtieron a Task item routes asbalanced; an autorouting pin overrides the manual model chain.provider/modelIdstrings with an optional thinking suffix. Globs and prefixes are rejected./routingopens the smart-routing panel directly;/routing on|offtoggles enablement;/routing statusreports settings-derived state.Preset-source unification
task.autorouting.presetand the whole preset layer are removed, not deprecated — there is no compatibility shim, per the repo's no-backward-compatibility rule.projectProviderOrderis extracted as the single implementation of "configuredmodelProviderOrderfirst, then first-wins catalog order", andcreateProviderSelectionPolicyis rewritten on top of it.ModelRegistry.autoroutingProviderOrder()takes no session and bypasses the policy builder entirely, so noeffectiveAuthmap is assembled: auth-independence is structural, not conventional. Auth-aware banding stays private torank().CustomRouterwould silently empty that provider's tiers.setup.providers, so a dead declaration cannot pollutedeclarationFingerprint.Breaking changes
task.autorouting.presetis gone. A preset-only configuration routes manually until tiers are generated.AutoroutingEffective.source,RoutingOutcome.source,AUTOROUTING_PRESETS,AutoroutingPresetId, andresolveTierMapfrom./config/*, plusTaskRoutingEvidence.sourcefrom./task/*. The union collapsed to one value, so keeping it would have published a meaningless required field on a durable receipt.notevalues change format to tier/fallback/resume components only./routing statusrelabels settings-derived tiers, with malformed provenance failing closed as hand-authored instead of reportinggenerated.Inactive-autorouting warning
Enabling autorouting without usable tiers previously failed silently. The host now decides once, where settings are already available, and reports through one shared uninterpolated constant on all three surfaces:
session.configWarnings;SessionEventStreamretains frames, so a notice published at hoststart()reaches a client that attaches later; ACP captures it onto the session record without early render and republishes once during deferred bootstrap beside the auth-failure branch.No new public query, no
getSdkConfigItems/config.listexpansion, no generalconfigWarningsexposure, and no new event kind. The internal flag lives in a package-private module that is null-mapped in the package exports, with a guard test asserting it is unreachable from any published type.Rebase notes
Rebased onto
devatf0453b6ab1. Two commits were dropped as genuinely obsolete rather than force-fitted:turn.steer_statusdisposition expectation, because dev splitsdk-adapter-dispositions.test.ts(ci(coding-agent): sdk-adapter-dispositions test file harness timeout (shard4) #4475) and fixed the same defect better, by supplying a real{ clientRef }input instead of assertinginvalid_requeston a bare probe;Substantive conflict resolutions: dev removed its
#writeTerminalBreadcrumbwrapper and the staged-publication/terminalBreadcrumbsfields, so those callsites were converted to dev's direct free-function form rather than reintroducing a wrapper; the/themeand/routingslash-command test blocks were unioned with correct closers.Verification
check:typesclean; repo-widebiome checkexits 0.AcpAgentdriven throughnewSessionagainst a fixture broker observes exactly one[warning:autorouting]chunk, which only holds because the notice survives late-attach replay. Disabling the host emission fails it while the zero-notice case still passes.config/settings-schema.ts).bun run checkcurrently fails on two gates that fail identically on a pristineorigin/devworktree atf0453b6ab1, measured on the same machine with sharednode_modules:verify-gjc-sdk-canonicalizationreports 27 violations on both sides, and the sorted violation sets diff empty. Every chain is rooted at the dev-ownedsession-state-sidecar.ts -> tools/descriptors.tsedge, and this branch touches none of the chain roots.Opened as a draft because #3764's closure standard requires green current CI, and those dev-side surfaces are still red at this base.
Transplant onto current dev
Rebuilt from exact stale head
e7b95a4ce957c00d36d532eae62a959ee9295c1b(event basecd51365cc270e27dceccfc2c184fadc9c1ddbe18) onto currentdevd97b79eff2be5f25bf3ae253de72430e2e2fab1b. All 44 original commits replayed; three mechanical conflict resolutions (settings.tsunion with dev's GLOBAL_ONLY_SETTINGS strip,config-cli.ts/test import unions, CHANGELOG section union); regenerated atomically:DAEMON_GENERATION170→171 via the canonical--fix-generationstool, tier-map/tool-catalog/operation-inventory/schemas, and the 8 new dev catalog keys recorded as baseline skips. Rebasing continued cleanly across later dev advances (883ab16 → d97b79e; behind_dev=0 at push time).Review findings fixed forward
AUTOROUTING_SELECTOR_MAX_LENGTH = 256is shared by reference between the selector grammar, the routing-evidence invariant, receipt bounding, and executorboundedSelector, so no accepted selector can be rejected after execution.!options.managedPersistence, andcommitStaged/discardStagedrequire strict attempt-root equality (absent or foreign adopted managers fail closed).Independent review
An independent architect-lane review of the fix-forward delta returned Verdict: approve with zero P1 findings; its three hardening recommendations are applied in the head commit (
share the selector bound and tighten staged-root invariants). An authenticated review from @probepark has been requested on the exact head.Verification on this head
check(biome + tsc) clean; telegram generation guard--validate-current-treepasses; tier-map gate green at 4272 in-scope keys / 3927 baseline skips; 173 focused tests pass across the autorouting contract/generator/tier-map/provider-order/settings/private-seam/task-routing/preflight/red-team/ACP-notice cohorts. Inherited dev-side failure (byte-identical on pristineorigin/dev): the telegram baseline manifest reports one missing command (notifications-telegram-topic-lease-renewal.test.ts), and Dev CI shard-1 carries the BisectTool set failing identically on dev's own CI.