Add MT-JailBench modules (CoA, X-Teaming, Mix + 3-judge claim) - #191
Conversation
Port SafetyArena/mt-jailbench (MIT, pinned cb8184e0) into SuperRed: the multi-turn attacks it has that SuperRed lacks, plus its evaluation. - optimizers/mtjb_coa, mtjb_xteaming, mtjb_mix (multi-turn attackers) - security_claims/mt_jailbench: the 3-judge unanimous vote (X-Teaming + CoA + FITD, 3-of-3) over HarmBench's standard subset (159) - upstream vendored byte-identical at the pinned commit, sync-checked; model calls routed through the framework LLM client; no temperature set - actor/crescendo/fitd omitted (SuperRed already ships those techniques); interactive omitted (upstream console harness, not an attack) - refinement moved pre-send (no mid-conversation rewind in the runtime); textgrad refine path is an optional extra and is env-gated in tests
55ab656 to
1e8a192
Compare
|
@claude review |
| tg.set_backward_engine(_BridgeEngine(), override=True) | ||
| attack = tg.Variable( | ||
| attacker_prompt, | ||
| role_description="attacker input to refine for a higher jailbreak score", | ||
| ) | ||
| target = _FixedResponseLLM()(attack) | ||
| target.set_role_description("victim model response to the attacker input") | ||
| optimizer = tg.TGD(parameters=[attack]) | ||
| optimizer.zero_grad() | ||
| loss = tg.TextLoss(loss_text)(target) | ||
| if _REFUSAL_MARKER in str(loss.value): | ||
| raise _RefineError("loss evaluation refused") | ||
| loss.backward() | ||
| optimizer.step() |
There was a problem hiding this comment.
Bug: global TextGrad backward engine can be clobbered by concurrent tasks
tg.set_backward_engine(_BridgeEngine(), override=True) stores the engine on TextGrad's process-wide SingletonBackwardEngine (textgrad/config.py), not something thread-local or contextvar-scoped. _run executes inside asyncio.to_thread(...) (line 121), so if two tasks call refine_prompt concurrently, task B's set_backward_engine call can overwrite task A's engine between task A's set_backward_engine (line 105) and A's later reads at tg.TGD(...) (112), tg.TextLoss(...) (114), and loss.backward() (117) — all of which fall back to the last-set global engine since no engine= is passed explicitly.
The result: task A's TextGrad LLM calls (loss evaluation, gradient, step) can silently route through task B's LLMClient, breaking the per-task isolation this codebase otherwise relies on (see the sibling _client_shim.py's "Concurrency-safe: each task sets its own ContextVar value" comment, which this path doesn't honor).
Suggested fix: pass engine= explicitly to tg.TGD(...), tg.TextLoss(...), and loss.backward(...) instead of relying on the global singleton, or hold a lock around the whole _run body if the API doesn't support explicit engine passing everywhere.
| for loader in (ast.literal_eval, json.loads): | ||
| try: | ||
| parsed = loader(json_str) | ||
| except (SyntaxError, ValueError, json.JSONDecodeError): |
There was a problem hiding this comment.
Bug: _loads_dict doesn't catch all exceptions ast.literal_eval can raise, breaking the "returns None on failure" contract
The module docstring (lines 6-7) promises "All parsers return None on failure ... rather than raising", but this except clause only covers SyntaxError, ValueError, and json.JSONDecodeError. ast.literal_eval also raises TypeError for syntactically-valid-but-type-invalid literals — e.g. a set containing a dict, like {{"improvement": "...", "prompt": "..."}} (double braces, which is exactly what the escaped {{ }} prompt templates elsewhere in this PR would produce if an LLM echoes them literally) parses as a set literal containing a dict, which raises TypeError: unhashable type: 'dict'.
Since extract_chain, extract_update_prompt, and extract_similarity (which call _loads_dict) are invoked from optimizer.py without a surrounding try/except for TypeError (e.g. _build_chain → initialize, _update_attack → the post-call handler), a single malformed attacker-LLM reply can crash the optimizer run instead of being discarded/retried as intended.
| except (SyntaxError, ValueError, json.JSONDecodeError): | |
| except (SyntaxError, ValueError, TypeError, RecursionError, json.JSONDecodeError): |
|
@claude review — pushed fixes for the review findings (commit f49d968), including regression tests. Please re-review the current head. |
Code reviewNo issues found. Checked for bugs and CLAUDE.md compliance. |
Ports SafetyArena/mt-jailbench into SuperRed. Benjamin Erichson added an MIT licence at our request, so the code is vendored unmodified at a pinned commit (
cb8184e0) rather than reimplemented.What's here
optimizers/mtjb_coaoptimizers/mtjb_xteamingoptimizers/mtjb_mixsecurity_claims/mt_jailbenchThe claim runs MT-JailBench's
majority_voteat its shipped threshold, which is unanimous: X-Teaming + CoA + FITD judges must all vote success (3-of-3). Its task set is HarmBench's standard subset (159 behaviours), so it evaluates the same behaviours as the existingharmbenchclaim but with the three-judge vote instead of the HarmBench classifier. HarmBench is attributed separately inLICENSES/.What I left out
actor,crescendo,fitd: SuperRed already ships these techniques. Re-porting them would double-count, so they're omitted and noted in each README.interactive: it is a consoleinput()/print()harness (upstream comments it "not a real attack"), so it isn't an optimizer.Deviations (in each
ASSUMPTIONS.md)base_response_sem, Mix/FITD refusal re-test) are disabled.refineextra, lazy-imported; without it the optimizers import and degrade cleanly.Vetting
SHA256SUMS+ an offline integrity test +scripts/sync_upstream.py --check), and ship in the built wheels.pytest: coa 24, xteaming 24 (+1 TextGrad-gated), mix 19 (+1 TextGrad-gated), claim 37. ruff clean. Entry points resolve. Tests cover per-run state re-arming, the pre-send refinement, unanimous-vs-majority voting, abstain-raises, and thatevaluateemits neither the behaviour nor the response.refineextra installed, so they are env-gated here; worth exercising once withpip install '.[refine]'.