End a task on the claim's success verdict, not the attacker's - #49
Merged
Merged
Conversation
The run loop broke only on the optimizer's RunEndResponse(done=True), so the security claim's verdict never terminated anything. evaluation.success was computed, logged, and discarded. Under a blind threat model (include_feedback=False) the optimizer is never told it won, so it cannot stop itself and spends its whole run budget attacking a target it has already broken. Measured on a DTAP sweep: 52 of 52 won tasks in that scope kept going, 785 further victim episodes, and none of them recorded stop_reason="success" because no such reason existed -- 44 were filed as "done" and 8 as "timeout", indistinguishable from a task that exhausted its budget. Add Controller(stop_on_success=True) and a "success" stop reason. The check runs before reset_ephemeral_state(), since there is no next run to reset for and that reset is the most expensive step in the loop for a container-backed target. Ending the task leaks nothing back to the attacker: the optimizer is torn down and no evaluation is ever sent, so a blind threat model stays blind. Pass stop_on_success=False to keep running after a win, which is what an attack-reliability study wants. "success" is added to all three completed-reason tuples (controller aggregate, persistence summary, live reporter). Omitting it from any of them would drop every win out of both the ASR numerator and denominator and report a perfect sweep as 0/0. Behaviour change: tests that ran on after a win now stop at it. Those exercising loop mechanics use a non-succeeding task; those whose premise is what happens after a win pass stop_on_success=False explicitly.
RoldSI
marked this pull request as ready for review
August 27, 2026 22:36
This was referenced Aug 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bug
The run loop breaks only on the optimizer's
RunEndResponse(done=True):evaluation.successis right there. It is computed, logged, and thrown away. Theframework will cheerfully log
success=True done=Falseand launch another run.This contradicts the project's own stated principle — the SecurityClaim decides
success, and attacker feedback is "for steering the next attempt, not for
self-certifying a win" — while termination is delegated entirely to the
attacker's self-report.
Why it has gone unnoticed
Almost every attacker/scope stops at the first win anyway, because the optimizer
is told it won (
include_feedback=True) and setsdoneitself. The gap onlyopens for a blind threat model with a multi-run budget, where the optimizer
cannot see the verdict and so cannot act on it.
Measured on a DecodingTrust-Agent sweep (16,390 tasks), every post-success run in
the whole campaign came from the one cell with that combination:
stop_reasondone44,timeout8The record damage outlasts the wasted compute: a task that stopped because it
won is indistinguishable from one that exhausted its budget, and eight tasks that
had already succeeded are filed as timeouts. It also skews cross-scope
comparison — mean runs/task was 18.96 in the blind scope against 16.00 in the
sighted one, with attacker spend tracking it, so the blind scope reads as a more
persistent attacker as an artifact of not being told it had won.
The change
Controller(stop_on_success: bool = True)and a new"success"stop reason.reset_ephemeral_state(). There is no next run to resetfor, and targets already document that reset is not called after the final run.
For a container-backed target that reset is the most expensive step in the loop.
"success"wins over"done"."done"is ambiguous —an attacker that gave up returns it too.
"success"says the framework ended iton a win.
down and no evaluation is ever sent to it.
stop_on_success=Falserestores the old loop exactly, which is what anattack-reliability study wants (how many of N attempts succeed, not whether
any did).
The sharp edge
"success"had to be added to three separate completed-reason tuples —_threat_model_end_event, the persistence summary, and the live reporter. Missingany one of them drops every win out of both the ASR numerator and denominator and
reports a perfect sweep as
0/0. There is a regression test for exactly this(
test_a_win_still_counts_in_the_asr).Behaviour change
Tasks that previously ran on after a win now stop at it, so run counts and stop
reasons are not comparable across this change. ASR is unaffected:
successlatched before and latches now. Resume is unaffected:
is_keptkeys on thederived status, which was already
"success"wheneversuccess=True.Existing tests split cleanly, which is the change being made auditable:
max_runs, budget) used asucceeding stub only incidentally → now use
StubTask(success=False)read on run 2, truncated-but-successful, attacker spend surviving a cancel) →
pass
stop_on_success=Falseexplicitly, with a comment saying whyNew tests
donestops at the winstop_on_success=Falsekeeps running after a win"success"beats the optimizer's owndonereset_ephemeral_state()Checks
594 passed(589 + 5 new) ·mypyclean on 25 files ·ruff checkandruff formatclean · docs updated inreference/controller.mdandreference/results.md, including an explicit behaviour-change note.