Skip to content

End a task on the claim's success verdict, not the attacker's - #49

Merged
RoldSI merged 1 commit into
mainfrom
feat/stop-on-success
Aug 27, 2026
Merged

RoldSI merged 1 commit into
mainfrom
feat/stop-on-success

Conversation

@RoldSI

@RoldSI RoldSI commented Aug 27, 2026

Copy link
Copy Markdown
Collaborator

The bug

The run loop breaks only on the optimizer's RunEndResponse(done=True):

done = isinstance(end_response, RunEndResponse) and end_response.done   # :1677
logger.info("... score=%.4f success=%s done=%s", ..., evaluation.success, done)
...
if done:            # :1507 — the only way the loop breaks
    stop_reason = "done"

evaluation.success is right there. It is computed, logged, and thrown away. The
framework will cheerfully log success=True done=False and launch another run.

This contradicts the project's own stated principle — the SecurityClaim decides
success, and attacker feedback is "for steering the next attempt, not for
self-certifying a win" — while termination is delegated entirely to the
attacker's self-report.

Why it has gone unnoticed

Almost every attacker/scope stops at the first win anyway, because the optimizer
is told it won (include_feedback=True) and sets done itself. The gap only
opens for a blind threat model with a multi-run budget, where the optimizer
cannot see the verdict and so cannot act on it.

Measured on a DecodingTrust-Agent sweep (16,390 tasks), every post-success run in
the whole campaign came from the one cell with that combination:

won tasks that kept running 52 of 52
further victim episodes 785
won on run 1, then ran ~19 more 31
one task scored 1.0 on all 20 runs
recorded stop_reason done 44, timeout 8

The record damage outlasts the wasted compute: a task that stopped because it
won is indistinguishable from one that exhausted its budget, and eight tasks that
had already succeeded are filed as timeouts. It also skews cross-scope
comparison — mean runs/task was 18.96 in the blind scope against 16.00 in the
sighted one, with attacker spend tracking it, so the blind scope reads as a more
persistent attacker as an artifact of not being told it had won.

The change

Controller(stop_on_success: bool = True) and a new "success" stop reason.

  • Checked before reset_ephemeral_state(). There is no next run to reset
    for, and targets already document that reset is not called after the final run.
    For a container-backed target that reset is the most expensive step in the loop.
  • When both would apply, "success" wins over "done". "done" is ambiguous —
    an attacker that gave up returns it too. "success" says the framework ended it
    on a win.
  • Blindness is preserved. Ending the task leaks nothing: the optimizer is torn
    down and no evaluation is ever sent to it.
  • stop_on_success=False restores the old loop exactly, which is what an
    attack-reliability study wants (how many of N attempts succeed, not whether
    any did).

The sharp edge

"success" had to be added to three separate completed-reason tuples —
_threat_model_end_event, the persistence summary, and the live reporter. Missing
any one of them drops every win out of both the ASR numerator and denominator and
reports a perfect sweep as 0/0. There is a regression test for exactly this
(test_a_win_still_counts_in_the_asr).

Behaviour change

Tasks that previously ran on after a win now stop at it, so run counts and stop
reasons are not comparable across this change
. ASR is unaffected: success
latched before and latches now. Resume is unaffected: is_kept keys on the
derived status, which was already "success" whenever success=True.

Existing tests split cleanly, which is the change being made auditable:

  • tests exercising loop mechanics (run counts, max_runs, budget) used a
    succeeding stub only incidentally → now use StubTask(success=False)
  • tests whose premise is what happens after a win (success latching, feedback
    read on run 2, truncated-but-successful, attacker spend surviving a cancel) →
    pass stop_on_success=False explicitly, with a comment saying why

New tests

  • a blind attacker that never says done stops at the win
  • stop_on_success=False keeps running after a win
  • "success" beats the optimizer's own done
  • a won task skips the final reset_ephemeral_state()
  • a win still counts in the ASR

Checks

594 passed (589 + 5 new) · mypy clean on 25 files · ruff check and
ruff format clean · docs updated in reference/controller.md and
reference/results.md, including an explicit behaviour-change note.

The run loop broke only on the optimizer's RunEndResponse(done=True), so
the security claim's verdict never terminated anything. evaluation.success
was computed, logged, and discarded.

Under a blind threat model (include_feedback=False) the optimizer is never
told it won, so it cannot stop itself and spends its whole run budget
attacking a target it has already broken. Measured on a DTAP sweep: 52 of
52 won tasks in that scope kept going, 785 further victim episodes, and
none of them recorded stop_reason="success" because no such reason existed
-- 44 were filed as "done" and 8 as "timeout", indistinguishable from a
task that exhausted its budget.

Add Controller(stop_on_success=True) and a "success" stop reason. The
check runs before reset_ephemeral_state(), since there is no next run to
reset for and that reset is the most expensive step in the loop for a
container-backed target.

Ending the task leaks nothing back to the attacker: the optimizer is torn
down and no evaluation is ever sent, so a blind threat model stays blind.
Pass stop_on_success=False to keep running after a win, which is what an
attack-reliability study wants.

"success" is added to all three completed-reason tuples (controller
aggregate, persistence summary, live reporter). Omitting it from any of
them would drop every win out of both the ASR numerator and denominator
and report a perfect sweep as 0/0.

Behaviour change: tests that ran on after a win now stop at it. Those
exercising loop mechanics use a non-succeeding task; those whose premise
is what happens after a win pass stop_on_success=False explicitly.
@RoldSI
RoldSI marked this pull request as ready for review August 27, 2026 22:36
@RoldSI
RoldSI merged commit 510720b into main Aug 27, 2026
3 checks passed
@RoldSI
RoldSI deleted the feat/stop-on-success branch August 27, 2026 22:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant