Add ReNeLLM optimizer (nested prompt jailbreak) - #195
Conversation
Port ReNeLLM (NJUNLP, MIT) into SuperRed as an optimizer. The upstream torch-free helpers (prompt rewrite ops + scenario nesting + harmful classification) are vendored byte-identical and executed live; only the openai/anthropic SDK helper is replaced by a stdlib shim that forwards to the framework LLM client. - optimizers/renellm; no temperature set; RunStart re-arms per-run state - rewritten+nested prompt produced pre-send (no mid-conversation rewind); the harmful-classification judge decides stop/continue up to iter_max - both rewrite and judge roles collapse to the single constrained self.llm; llama/, defenses, torch/transformers not ported. Deviations in ASSUMPTIONS.md Paper: Ding, Kuang, Ma, Cao, Xian, Chen, Huang, "A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool LLMs Easily", NAACL 2024, arXiv:2311.08268. Vendored from NJUNLP/ReNeLLM @ a61c39e. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
| if success: | ||
| self._succeeded = True | ||
| return RunEndResponse(event=event, done=True) | ||
| if not self._injected: | ||
| # No eligible surface fired (or no payload) -> nothing can advance. | ||
| return RunEndResponse(event=event, done=True) | ||
| if self._run_count >= self._iter_max: |
There was a problem hiding this comment.
A transient failure in _produce_nested (any non-budget exception from the aux-LLM rewrite/judge call) is caught and turned into _pending = None, which means self._injected never becomes True for that run. Here, the if not self._injected: check then returns done=True unconditionally — ending the entire sweep after a single flaky aux-LLM call, regardless of iter_max (checked only afterward on line 349).
This contradicts both the comment on line 225 ("a transient aux-LLM failure ends this run, not the sweep") and ASSUMPTIONS.md ("Any other transient auxiliary-LLM failure ends the current run ... not the sweep"): the current code lumps "production failed this run" together with "no eligible surface exists" (which legitimately can't make progress and should stop), while a transient failure should just move on to the next run up to iter_max.
Suggested fix: track "production failed" separately from "no eligible surface fired", and only short-circuit with done=True for the latter case (or when self._run_count >= self._iter_max).
| def _handle_post_call(self, event: ControllablePostCallEvent) -> ControllableNoInjection: | ||
| # PostCall requires an injection decision; this attack reads the reply | ||
| # but never rewrites it, so it always declines. | ||
| if self._injected and event.controllable.name == self._channel: | ||
| self._pending_answer = event.answer | ||
| return ControllableNoInjection(event=event, controllable=event.controllable) |
There was a problem hiding this comment.
self._pending_answer is overwritten on every post-call on self._channel while self._injected is True, not just the first one after injection. Since the nested prompt is injected only once (subsequent pre-calls on the same channel decline via self._injected at line 311), any later call on that channel within the same run sends the original (non-attack) input through, and its reply silently overwrites the attack reply here. _handle_run_end then judges self._pending_answer (line 332), i.e. the reply to the wrong turn — causing false negatives (a real jailbreak scored as failure) or false positives, in any run with more than one call on the injected surface.
Suggested fix: capture only the first post-call reply after injection, e.g. gate the assignment on self._pending_answer is None (it's reset to None each run in _reset_run_state):
| def _handle_post_call(self, event: ControllablePostCallEvent) -> ControllableNoInjection: | |
| # PostCall requires an injection decision; this attack reads the reply | |
| # but never rewrites it, so it always declines. | |
| if self._injected and event.controllable.name == self._channel: | |
| self._pending_answer = event.answer | |
| return ControllableNoInjection(event=event, controllable=event.controllable) | |
| def _handle_post_call(self, event: ControllablePostCallEvent) -> ControllableNoInjection: | |
| # PostCall requires an injection decision; this attack reads the reply | |
| # but never rewrites it, so it always declines. Only the first reply | |
| # after injection is captured, so a later non-attack turn on the same | |
| # channel can't overwrite it. | |
| if self._injected and self._pending_answer is None and event.controllable.name == self._channel: | |
| self._pending_answer = event.answer | |
| return ControllableNoInjection(event=event, controllable=event.controllable) |
|
@claude review — pushed fixes for the review findings (commit efe71c8), including regression tests. Please re-review the current head. |
Code reviewNo issues found. Checked for bugs and CLAUDE.md compliance. |
Ports ReNeLLM (NJUNLP, MIT) into SuperRed as an optimizer — a distinct jailbreak (generalized nested prompts) the repo doesn't yet ship.
What's here
optimizers/renellm— the four torch-free upstream helpers (prompt-rewrite ops, scenario nesting, harmful classification, data utils) are vendored byte-identical fromNJUNLP/ReNeLLMand executed live; only the openai/anthropic SDK helper is replaced by a stdlib shim forwarding to the framework LLM client. Runtime dep:superredonly.iter_max.Faithfulness / deviations (
ASSUMPTIONS.md)self.llm(replacing upstream'sclaudeCompletion).while Trueis capped (default 20).llama/, defenses, torch/transformers, and result-file persistence are not ported.Path(__file__)+spec_from_file_location, so the module needs an unzipped install (which SuperRed uses) — it is not zip-safe, unlike a pureimportlib.resourcestext read.Vetting
pytest25 passed, ruff clean, vendored files byte-identical to the pin and shipped in the wheel; no vendored template text in any authored file (tests assert the injected value equals exactly one loaded nested scenario, using synthetic strings). Matches repo convention (class via__init__, no entry-points block).Paper: Ding, Kuang, Ma, Cao, Xian, Chen, Huang, "A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily", NAACL 2024, arXiv:2311.08268. Vendored from NJUNLP/ReNeLLM @ a61c39e.
🤖 Generated with Claude Code