Skip to content

Add MT-JailBench modules (CoA, X-Teaming, Mix + 3-judge claim) - #191

Merged
RoldSI merged 1 commit into
mainfrom
feat/mt-jailbench
Sep 24, 2026
Merged

RoldSI merged 1 commit into
mainfrom
feat/mt-jailbench

Conversation

@rishabhsinha17

@rishabhsinha17 rishabhsinha17 commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Ports SafetyArena/mt-jailbench into SuperRed. Benjamin Erichson added an MIT licence at our request, so the code is vendored unmodified at a pinned commit (cb8184e0) rather than reimplemented.

What's here

Module What it is
optimizers/mtjb_coa Chain-of-Attack multi-turn optimizer
optimizers/mtjb_xteaming X-Teaming multi-turn optimizer (optional TextGrad refine)
optimizers/mtjb_mix Mix optimizer, drives the vendored multi-family engine
security_claims/mt_jailbench MT-JailBench's evaluation: the three judges

The claim runs MT-JailBench's majority_vote at its shipped threshold, which is unanimous: X-Teaming + CoA + FITD judges must all vote success (3-of-3). Its task set is HarmBench's standard subset (159 behaviours), so it evaluates the same behaviours as the existing harmbench claim but with the three-judge vote instead of the HarmBench classifier. HarmBench is attributed separately in LICENSES/.

What I left out

  • actor, crescendo, fitd: SuperRed already ships these techniques. Re-porting them would double-count, so they're omitted and noted in each README.
  • interactive: it is a console input()/print() harness (upstream comments it "not a real attack"), so it isn't an optimizer.

Deviations (in each ASSUMPTIONS.md)

  • The runtime resets only between runs, with no mid-conversation rewind, so each attack's refinement moves to pre-send and the target-requery paths (CoA base_response_sem, Mix/FITD refusal re-test) are disabled.
  • Every model call routes through the framework's constrained LLM client; the vendored SDK client hub is replaced and no temperature is forwarded.
  • CoA similarity uses the LLM-approx path, so it needs no SimCSE weights or external endpoint.
  • TextGrad (X-Teaming/Mix refine) is an optional refine extra, lazy-imported; without it the optimizers import and degrade cleanly.

Vetting

  • Vendored files are byte-identical to the pinned commit (SHA256SUMS + an offline integrity test + scripts/sync_upstream.py --check), and ship in the built wheels.
  • pytest: coa 24, xteaming 24 (+1 TextGrad-gated), mix 19 (+1 TextGrad-gated), claim 37. ruff clean. Entry points resolve. Tests cover per-run state re-arming, the pre-send refinement, unanimous-vs-majority voting, abstain-raises, and that evaluate emits neither the behaviour nor the response.
  • The TextGrad refine paths (X-Teaming/Mix) run only with the refine extra installed, so they are env-gated here; worth exercising once with pip install '.[refine]'.

Port SafetyArena/mt-jailbench (MIT, pinned cb8184e0) into SuperRed: the
multi-turn attacks it has that SuperRed lacks, plus its evaluation.

- optimizers/mtjb_coa, mtjb_xteaming, mtjb_mix (multi-turn attackers)
- security_claims/mt_jailbench: the 3-judge unanimous vote
  (X-Teaming + CoA + FITD, 3-of-3) over HarmBench's standard subset (159)
- upstream vendored byte-identical at the pinned commit, sync-checked;
  model calls routed through the framework LLM client; no temperature set
- actor/crescendo/fitd omitted (SuperRed already ships those techniques);
  interactive omitted (upstream console harness, not an attack)
- refinement moved pre-send (no mid-conversation rewind in the runtime);
  textgrad refine path is an optional extra and is env-gated in tests
@RoldSI

RoldSI commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

@claude review

Comment on lines +105 to +118
tg.set_backward_engine(_BridgeEngine(), override=True)
attack = tg.Variable(
attacker_prompt,
role_description="attacker input to refine for a higher jailbreak score",
)
target = _FixedResponseLLM()(attack)
target.set_role_description("victim model response to the attacker input")
optimizer = tg.TGD(parameters=[attack])
optimizer.zero_grad()
loss = tg.TextLoss(loss_text)(target)
if _REFUSAL_MARKER in str(loss.value):
raise _RefineError("loss evaluation refused")
loss.backward()
optimizer.step()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: global TextGrad backward engine can be clobbered by concurrent tasks

tg.set_backward_engine(_BridgeEngine(), override=True) stores the engine on TextGrad's process-wide SingletonBackwardEngine (textgrad/config.py), not something thread-local or contextvar-scoped. _run executes inside asyncio.to_thread(...) (line 121), so if two tasks call refine_prompt concurrently, task B's set_backward_engine call can overwrite task A's engine between task A's set_backward_engine (line 105) and A's later reads at tg.TGD(...) (112), tg.TextLoss(...) (114), and loss.backward() (117) — all of which fall back to the last-set global engine since no engine= is passed explicitly.

The result: task A's TextGrad LLM calls (loss evaluation, gradient, step) can silently route through task B's LLMClient, breaking the per-task isolation this codebase otherwise relies on (see the sibling _client_shim.py's "Concurrency-safe: each task sets its own ContextVar value" comment, which this path doesn't honor).

Suggested fix: pass engine= explicitly to tg.TGD(...), tg.TextLoss(...), and loss.backward(...) instead of relying on the global singleton, or hold a lock around the whole _run body if the API doesn't support explicit engine passing everywhere.

tg.set_backward_engine(_BridgeEngine(), override=True)
attack = tg.Variable(
attacker_prompt,
role_description="attacker input to refine for a higher jailbreak score",
)
target = _FixedResponseLLM()(attack)
target.set_role_description("victim model response to the attacker input")
optimizer = tg.TGD(parameters=[attack])
optimizer.zero_grad()
loss = tg.TextLoss(loss_text)(target)
if _REFUSAL_MARKER in str(loss.value):
raise _RefineError("loss evaluation refused")
loss.backward()
optimizer.step()

for loader in (ast.literal_eval, json.loads):
try:
parsed = loader(json_str)
except (SyntaxError, ValueError, json.JSONDecodeError):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bug: _loads_dict doesn't catch all exceptions ast.literal_eval can raise, breaking the "returns None on failure" contract

The module docstring (lines 6-7) promises "All parsers return None on failure ... rather than raising", but this except clause only covers SyntaxError, ValueError, and json.JSONDecodeError. ast.literal_eval also raises TypeError for syntactically-valid-but-type-invalid literals — e.g. a set containing a dict, like {{"improvement": "...", "prompt": "..."}} (double braces, which is exactly what the escaped {{ }} prompt templates elsewhere in this PR would produce if an LLM echoes them literally) parses as a set literal containing a dict, which raises TypeError: unhashable type: 'dict'.

Since extract_chain, extract_update_prompt, and extract_similarity (which call _loads_dict) are invoked from optimizer.py without a surrounding try/except for TypeError (e.g. _build_chain → initialize, _update_attack → the post-call handler), a single malformed attacker-LLM reply can crash the optimizer run instead of being discarded/retried as intended.

Suggested change
except (SyntaxError, ValueError, json.JSONDecodeError):
except (SyntaxError, ValueError, TypeError, RecursionError, json.JSONDecodeError):

def _loads_dict(json_str: str) -> dict | None:
"""Parse a dict from a string, trying ``ast.literal_eval`` then ``json``."""
for loader in (ast.literal_eval, json.loads):
try:
parsed = loader(json_str)
except (SyntaxError, ValueError, json.JSONDecodeError):
continue
if isinstance(parsed, dict):
return parsed
return None

@rishabhsinha17

Copy link
Copy Markdown
Collaborator Author

@claude review — pushed fixes for the review findings (commit f49d968), including regression tests. Please re-review the current head.

@claude

claude Bot commented Sep 23, 2026

Copy link
Copy Markdown

Code review

No issues found. Checked for bugs and CLAUDE.md compliance.

@RoldSI
RoldSI merged commit feae01d into main Sep 24, 2026
4 checks passed
@RoldSI
RoldSI deleted the feat/mt-jailbench branch September 24, 2026 05:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants