Repository navigation
fix(permission): budget permission review output and classify its failures - #57
Merged
Merged
Conversation
…lures
Permission review asked the reviewer for at most 512 output tokens. On a
reasoning reviewer those tokens are billed with hidden reasoning, so the
response ended as `response.incomplete` long before the review JSON, and the
runtime reported the truncation as invalid reviewer output. Sessions using a
reasoning model as the approval reviewer therefore escalated nearly every
permission request to the human fallback.
Raise the reviewer budget to one order of magnitude above a review JSON,
request the lowest standard reasoning effort so thinking cannot consume the
whole allowance, and keep the two knobs documented where a reader looks for
them. The reviewer request shape stays owned by runtime and independent of the
primary provider's reasoning_effort, which examples/config.toml now states.
Split the non-stop finish classification so a response the reviewer never
controlled reads as a failed or truncated review instead of a contract
violation:
- Length -> new ReviewOutputTruncated { finish_reason }
- Blocked/Error -> ReviewFailed with the provider cause
- anything else -> InvalidReviewOutput
InvalidReviewOutput now means only that the reviewer returned output outside
the review contract, which lets runtime policy retry or escalate a truncated
review without parsing error text.
Verification: cargo fmt --all --check; cargo clippy --all-targets
--all-features -- -D warnings; cargo test --all (all targets ok).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the "AI review unavailable: permission review output is invalid: permission review completed without stop finish reason" escalation loop reported for session
a4daeba5-7108-47b9-80ba-7eb9b9b592d0.Root cause
Permission review asked the reviewer for at most 512 output tokens. On a reasoning reviewer those tokens are billed together with hidden reasoning, so the response ended as
response.incompletelong before the review JSON arrived.complete_single_textonly accepts astopfinish, so the truncation surfaced asInvalidReviewOutput, andmodel_then_humanescalated to the human dialog.Evidence from that session: the reviewer stream (19:36:03.184 -> .646) contains only
response.reasoning_text.deltaevents, ends inresponse.incomplete, and never reaches the review schema; across the log, all 35 truncated reviews are the reasoning reviewer while all 50 responses from the non-reasoning reviewer completed.Changes
crates/merry-runtime/src/permission/review.rs: reviewer budget 512 -> 2048, and the reviewer request now asks for a low reasoning effort so hidden thinking cannot consume the whole allowance.crates/merry-runtime/src/permission.rs: new typedPermissionAdmissionError::ReviewOutputTruncated { finish_reason };InvalidReviewOutputnow means only that the reviewer returned output outside the review contract.crates/merry-runtime/src/permission/review.rs:classify_non_stop_review_finishsplitsLength(truncated),Blocked/Error(failed review with the provider cause), and everything else (contract violation).examples/config.toml: documents that runtime owns the reviewer request shape, that review effort is independent of the primaryreasoning_effort, and that the role should point at a provider accepting a reasoning effort.The reviewer request shape stays owned by runtime and provider-neutral; no provider wire type or policy crosses a layer boundary, and the primary step request is unchanged (its
reasoning_effortremains optional and configuration-driven).Verification
cargo fmt --all --checkcargo clippy --all-targets --all-features -- -D warningscargo test --all(all targets ok)Lengthmaps toReviewOutputTruncated;Blocked/Errorstay failed reviews.Not verified: no live model call was made, so the endpoint behavior of
loweffort is untested; the change is documented inexamples/config.tomlas the mitigation for that.Follow-ups not in this change
[models.approval_review]if a provider needs it omitted.