Skip to content

Add ReNeLLM optimizer (nested prompt jailbreak) - #195

Merged
RoldSI merged 1 commit into
mainfrom
feat/renellm
Sep 24, 2026
Merged

RoldSI merged 1 commit into
mainfrom
feat/renellm

Conversation

@rishabhsinha17

Copy link
Copy Markdown
Collaborator

Ports ReNeLLM (NJUNLP, MIT) into SuperRed as an optimizer — a distinct jailbreak (generalized nested prompts) the repo doesn't yet ship.

What's here

  • optimizers/renellm — the four torch-free upstream helpers (prompt-rewrite ops, scenario nesting, harmful classification, data utils) are vendored byte-identical from NJUNLP/ReNeLLM and executed live; only the openai/anthropic SDK helper is replaced by a stdlib shim forwarding to the framework LLM client. Runtime dep: superred only.
  • One run = one upstream outer iteration: RunStart re-arms state and produces the rewritten+nested prompt pre-send; inject at PreCall; read the reply at PostCall; the harmful-classification judge decides stop/continue up to iter_max.

Faithfulness / deviations (ASSUMPTIONS.md)

  • Attacked model = the SuperRed target; rewrite and judge roles both collapse to the single constrained self.llm (replacing upstream's claudeCompletion).
  • No sampling temperature is ever sent (house rule); upstream's unbounded rewrite-retry while True is capped (default 20). llama/, defenses, torch/transformers, and result-file persistence are not ported.
  • Reviewer note: the vendored helpers are executed via Path(__file__) + spec_from_file_location, so the module needs an unzipped install (which SuperRed uses) — it is not zip-safe, unlike a pure importlib.resources text read.

Vetting

  • pytest 25 passed, ruff clean, vendored files byte-identical to the pin and shipped in the wheel; no vendored template text in any authored file (tests assert the injected value equals exactly one loaded nested scenario, using synthetic strings). Matches repo convention (class via __init__, no entry-points block).

Paper: Ding, Kuang, Ma, Cao, Xian, Chen, Huang, "A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily", NAACL 2024, arXiv:2311.08268. Vendored from NJUNLP/ReNeLLM @ a61c39e.

🤖 Generated with Claude Code

Port ReNeLLM (NJUNLP, MIT) into SuperRed as an optimizer. The upstream
torch-free helpers (prompt rewrite ops + scenario nesting + harmful
classification) are vendored byte-identical and executed live; only the
openai/anthropic SDK helper is replaced by a stdlib shim that forwards to
the framework LLM client.

- optimizers/renellm; no temperature set; RunStart re-arms per-run state
- rewritten+nested prompt produced pre-send (no mid-conversation rewind);
  the harmful-classification judge decides stop/continue up to iter_max
- both rewrite and judge roles collapse to the single constrained self.llm;
  llama/, defenses, torch/transformers not ported. Deviations in ASSUMPTIONS.md

Paper: Ding, Kuang, Ma, Cao, Xian, Chen, Huang, "A Wolf in Sheep's Clothing:
Generalized Nested Jailbreak Prompts can Fool LLMs Easily", NAACL 2024,
arXiv:2311.08268. Vendored from NJUNLP/ReNeLLM @ a61c39e.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Comment on lines +343 to +349
if success:
self._succeeded = True
return RunEndResponse(event=event, done=True)
if not self._injected:
# No eligible surface fired (or no payload) -> nothing can advance.
return RunEndResponse(event=event, done=True)
if self._run_count >= self._iter_max:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A transient failure in _produce_nested (any non-budget exception from the aux-LLM rewrite/judge call) is caught and turned into _pending = None, which means self._injected never becomes True for that run. Here, the if not self._injected: check then returns done=True unconditionally — ending the entire sweep after a single flaky aux-LLM call, regardless of iter_max (checked only afterward on line 349).

This contradicts both the comment on line 225 ("a transient aux-LLM failure ends this run, not the sweep") and ASSUMPTIONS.md ("Any other transient auxiliary-LLM failure ends the current run ... not the sweep"): the current code lumps "production failed this run" together with "no eligible surface exists" (which legitimately can't make progress and should stop), while a transient failure should just move on to the next run up to iter_max.

Suggested fix: track "production failed" separately from "no eligible surface fired", and only short-circuit with done=True for the latter case (or when self._run_count >= self._iter_max).

Comment on lines +323 to +328
def _handle_post_call(self, event: ControllablePostCallEvent) -> ControllableNoInjection:
# PostCall requires an injection decision; this attack reads the reply
# but never rewrites it, so it always declines.
if self._injected and event.controllable.name == self._channel:
self._pending_answer = event.answer
return ControllableNoInjection(event=event, controllable=event.controllable)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

self._pending_answer is overwritten on every post-call on self._channel while self._injected is True, not just the first one after injection. Since the nested prompt is injected only once (subsequent pre-calls on the same channel decline via self._injected at line 311), any later call on that channel within the same run sends the original (non-attack) input through, and its reply silently overwrites the attack reply here. _handle_run_end then judges self._pending_answer (line 332), i.e. the reply to the wrong turn — causing false negatives (a real jailbreak scored as failure) or false positives, in any run with more than one call on the injected surface.

Suggested fix: capture only the first post-call reply after injection, e.g. gate the assignment on self._pending_answer is None (it's reset to None each run in _reset_run_state):

Suggested change
def _handle_post_call(self, event: ControllablePostCallEvent) -> ControllableNoInjection:
# PostCall requires an injection decision; this attack reads the reply
# but never rewrites it, so it always declines.
if self._injected and event.controllable.name == self._channel:
self._pending_answer = event.answer
return ControllableNoInjection(event=event, controllable=event.controllable)
def _handle_post_call(self, event: ControllablePostCallEvent) -> ControllableNoInjection:
# PostCall requires an injection decision; this attack reads the reply
# but never rewrites it, so it always declines. Only the first reply
# after injection is captured, so a later non-attack turn on the same
# channel can't overwrite it.
if self._injected and self._pending_answer is None and event.controllable.name == self._channel:
self._pending_answer = event.answer
return ControllableNoInjection(event=event, controllable=event.controllable)

@rishabhsinha17

Copy link
Copy Markdown
Collaborator Author

@claude review — pushed fixes for the review findings (commit efe71c8), including regression tests. Please re-review the current head.

@claude

claude Bot commented Sep 23, 2026

Copy link
Copy Markdown

Code review

No issues found. Checked for bugs and CLAUDE.md compliance.

@RoldSI
RoldSI merged commit 8d1ebc5 into main Sep 24, 2026
5 of 6 checks passed
@RoldSI
RoldSI deleted the feat/renellm branch September 24, 2026 05:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants