Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
130 changes: 130 additions & 0 deletions optimizers/renellm/ASSUMPTIONS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,130 @@
# ASSUMPTIONS & DEVIATIONS

Deliberate deviations from upstream ReNeLLM (`NJUNLP/ReNeLLM` @ `a61c39e`) made
while porting it to superred's event-driven `Optimizer`. Upstream is the
authoritative source for the technique; this document records where the port
differs and why.

## Event-model mapping (pre-send, no rewind)

- **One superred *run* == one ReNeLLM outer iteration.** Upstream `renellm.py`
loops per behaviour up to `iter_max`, and every iteration re-rewrites from the
*original* goal (`harm_behavior` is reset to `temp_harm_behavior` after each
failed jailbreak). This port produces a fresh rewrite+nest on each `RunStart`
from `goal.description`, so runs are independent iterations. `iter_max` maps to
the number of runs before the optimizer reports `done` (upstream default 20).
- **The rewritten+nested prompt is produced pre-send.** Upstream rewrites, nests,
*then* queries the attacked model within one iteration. superred has no
mid-conversation rewind, so the rewrite-until-harmful loop and scenario nesting
run entirely inside the `RunStart` handler; the finished nested prompt is
injected at the first eligible `ControllablePreCall`. Nothing is rewritten
after a send.
- **The model under attack is the target, not an optimizer-built client.**
Upstream's `claudeCompletion(attack_model, nested_prompt, ...)` is replaced by
injecting `nested_prompt` on the target's controllable surface and reading the
reply from `ControllablePostCall.answer` (with a trajectory-observable
fallback). Upstream's `--attack_model`/`--claude_*` parameters and the entire
anthropic client path are dropped.
- **`ControllablePostCall` always returns `ControllableNoInjection`.** This
attack reads the reply but never rewrites the target's answer; the channel type
requires an explicit injection decision, so it declines every post-call.

## Auxiliary LLM routing (rewrite + judge via self.llm)

- Upstream uses `--rewrite_model` (default `gpt-3.5-turbo`) for the six rewrite
operations and `--judge_model` for the harmful-classification judge. In this
port **both roles are the single constrained `self.llm`** (superred's
`LLMClient`). The judge is used, as upstream, for two things: (a) confirming a
rewrite is still harmful, and (b) scoring the target reply.
- **Upstream `utils/llm_completion_utils.py` is the only util not vendored.** It
builds `openai`/`anthropic` SDK clients; it is replaced by
`src/renellm_optimizer/_shim.py`, a stdlib `chatCompletion` with the identical
signature that forwards only `messages` to `self.llm`. The vendored helpers
import it via `from utils.llm_completion_utils import chatCompletion`
unchanged.
- **No sampling temperature is ever sent** (house rule; see
`tests/test_no_temperature.py`). Upstream's rewrite operations pass
`random.uniform(0, 1)` and the judge passes `args.temperature` (default 0) to
the SDK helper; the shim drops both, so `self.llm` uses each provider's
default. This is the documented reason the rewrite may occasionally fail the
still-harmful check and be retried.

## Synchronous vendored code bridged to async self.llm

- The vendored rewrite/judge helpers are **synchronous**. They are executed in a
worker thread via `asyncio.to_thread`; each model call hops back to the event
loop with `asyncio.run_coroutine_threadsafe(self.llm.complete(...), loop)`. The
bridge is passed through a `contextvars.ContextVar` (copied into the worker by
`to_thread`), so the module-level vendored `chatCompletion` reaches the right
optimizer instance's client. The shim mirrors upstream's `.strip()` on the
returned content.
- Upstream's `round_sleep` / `fail_sleep` / `retry_times` and their
`time.sleep()` calls live only in the replaced SDK helper and are dropped;
retry/backoff is `self.llm`'s responsibility.

## Rewrite-retry loop is bounded

- Upstream's inner rewrite loop is `while True`: it retries indefinitely until
the judge labels the rewrite harmful (relying on the random temperature to
eventually produce one). This port caps it at `max_rewrite_attempts` (default
20); on exhaustion it uses the last rewritten candidate so a nested prompt is
still produced. This prevents an unbounded loop when the judge never returns
`"1"` (e.g. a mocked or unusually conservative judge).

## Preserved upstream behaviours

- **Refusal short-circuit, verbatim.** Before judging a reply, upstream skips the
judge when the reply looks like a refusal:
`"apologize" not in output and "sorry" not in output.lower()`. This asymmetric
casing (`apologize` matched case-sensitively, `sorry` case-insensitively) is
reproduced exactly.
- **Rewrite operation selection.** A random count `randint(1, 6)` and a random
order `sample(range(6), n)` of the six operations, applied in sequence,
matching upstream. Selection uses a seedable RNG for reproducibility; the
vendored `shortenSentence` candidate pick still uses the global `random`
module, as upstream.
- **Scenario nesting.** A random choice among the three vendored scenarios, with
`<>` replaced by the rewritten goal (`scenario.replace("<>", ...)`).

## Injection surface

- ReNeLLM is a **single-turn user-input** attack. It injects only into a
free-text user surface and never into the reserved `system_prompt` (which would
clobber a Task-configured system prompt and does not match the technique). If
`target_controllable_name` is set, only that surface is used; otherwise the
first eligible free-text surface is pinned for the run.

## Stop / continue and error handling

- **Success** ends the attack (`done=True`): either the vendored judge labels the
reply harmful, or the framework `RunEndEvent.evaluation.success` is `True` (a
framework success verdict is authoritative). Otherwise the run continues until
`iter_max`.
- **No-progress guard.** If no eligible surface fires in a run (e.g. a
system-prompt-only target) or no payload could be produced, the optimizer
reports `done=True` rather than spinning the full budget injecting nothing.
- **Budget / transient errors.** A genuine `BudgetExhaustedError` with non-zero
spent cost propagates (a spent run must not be reported as a clean finish); the
zero-cost noop client handed to non-LLM optimizers degrades quietly (no
injection). Any other transient auxiliary-LLM failure ends the current run
(no injection / non-success), not the sweep.

## Not ported (out of scope for the optimizer)

- The AdvBench dataset reader and `data/advbench/harmful_behaviors.csv` are not
vendored: the goal is supplied by superred as `goal.description`, so the
rewrite/nest paths load no data files. (`data_utils.py` is still vendored
because `prompt_rewrite_utils` imports `remove_number_prefix` from it.)
- Upstream's `llama/`, `defense/`, `get_responses.py`, `check_gpt_asr.py`,
`check_kw_asr.py`, `renellm_tcps.py`, and all `torch`/`transformers` code paths
are not bundled. Result-file JSON persistence and temp-file checkpointing are
handled by superred's trajectory/persistence layer instead.

## Vendored-code execution hygiene

- Because the vendored files use absolute `from utils.X import Y` imports that
cannot be edited, `vendored.load()` installs a private `utils` package (plus
the shim as `utils.llm_completion_utils`) into `sys.modules` only for the
duration of one hermetic import, then restores `sys.modules`. The imported
module objects keep working afterwards because their imported names are bound
at import time. The load is cached and guarded by a lock.
21 changes: 21 additions & 0 deletions optimizers/renellm/LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) 2026 Rishabh Sinha

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
33 changes: 33 additions & 0 deletions optimizers/renellm/LICENSES/NOTICE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# NOTICE

This module ports ReNeLLM (generalized nested jailbreak: prompt rewriting +
scenario nesting) from NJUNLP/ReNeLLM, pinned at commit `a61c39e`. Upstream is
redistributed under the MIT License; the verbatim text is in `ReNeLLM-MIT.txt`.

## Code (this module)

MIT, see the module `LICENSE`. The ReNeLLM control flow (the rewrite-until-
harmful loop, scenario nesting, refusal short-circuit, and reply scoring) is
reimplemented against superred's async event model in
`src/renellm_optimizer/`. Upstream's SDK completion helper is replaced by a
stdlib shim that routes model calls through superred's constrained LLMClient.

## Vendored payloads

Vendored verbatim (sha256-pinned in `_vendor/SHA256SUMS`), executed but never
reproduced in authored code (torch-free helpers only; `llama/` and
`torch`/`transformers` paths are not bundled):

- `_vendor/renellm/utils/prompt_rewrite_utils.py` (six rewrite operations)
- `_vendor/renellm/utils/scenario_nest_utils.py` (three nesting scenarios)
- `_vendor/renellm/utils/harmful_classification_utils.py` (LLM judge)
- `_vendor/renellm/utils/data_utils.py` (rewrite post-processing helper)

License: MIT (Copyright (c) 2024 NJUNLP).

## Citation

Cite ReNeLLM (Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun
Chen, Shujian Huang, "A Wolf in Sheep's Clothing: Generalized Nested Jailbreak
Prompts can Fool Large Language Models Easily", NAACL 2024, arXiv:2311.08268)
when reporting numbers produced with this module.
21 changes: 21 additions & 0 deletions optimizers/renellm/LICENSES/ReNeLLM-MIT.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) 2024 NJUNLP

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
48 changes: 48 additions & 0 deletions optimizers/renellm/NOTICE
Original file line number Diff line number Diff line change
@@ -0,0 +1,48 @@
superred-optimizer-renellm
Copyright (c) 2026 the superred module authors

This product includes original code licensed under the MIT License (see LICENSE).

It bundles and builds upon third-party material from ReNeLLM, the official
implementation of:

Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and
Shujian Huang. "A Wolf in Sheep's Clothing: Generalized Nested Jailbreak
Prompts can Fool Large Language Models Easily." NAACL 2024.
arXiv:2311.08268.

The upstream source is NJUNLP/ReNeLLM (https://github.com/NJUNLP/ReNeLLM),
pinned at commit a61c39ea4e5311a5fbfaee9869e5325783d1522f. ReNeLLM is
redistributed under the MIT License (Copyright (c) 2024 NJUNLP); the upstream
license text is preserved verbatim in LICENSES/ReNeLLM-MIT.txt.

--------------------------------------------------------------------------------
Bundled material (vendored byte-for-byte; sha256-pinned in _vendor/SHA256SUMS)
--------------------------------------------------------------------------------
Only ReNeLLM's torch-free rewrite/nest/judge helpers are vendored (the upstream
llama/ and torch/transformers code paths are NOT bundled):

- src/renellm_optimizer/_vendor/renellm/utils/prompt_rewrite_utils.py
The six prompt-rewriting operations.
- src/renellm_optimizer/_vendor/renellm/utils/scenario_nest_utils.py
The three scenario-nesting templates (code completion / table filling /
text continuation).
- src/renellm_optimizer/_vendor/renellm/utils/harmful_classification_utils.py
The LLM harmful-classification judge.
- src/renellm_optimizer/_vendor/renellm/utils/data_utils.py
Dataset reader and the rewrite candidate post-processing helper.

Upstream's utils/llm_completion_utils.py (an openai/anthropic SDK client) is
NOT vendored: it is replaced by a stdlib shim (src/renellm_optimizer/_shim.py)
that routes every vendored model call through superred's constrained LLMClient.

The ReNeLLM control flow (rewrite-until-harmful loop, scenario nesting, and
reply scoring) is reimplemented against superred's async event model in
src/renellm_optimizer/ (original to this module, MIT). No prompt/scenario/judge
body is reproduced in authored code.

--------------------------------------------------------------------------------
Citation
--------------------------------------------------------------------------------
When reporting numbers produced with this module, cite ReNeLLM
(Ding et al., NAACL 2024, arXiv:2311.08268).
81 changes: 81 additions & 0 deletions optimizers/renellm/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
# superred-optimizer-renellm

ReNeLLM generalized nested-jailbreak optimizer for
[superred](https://superred.simonsure.com).

Ports **ReNeLLM** from
[NJUNLP/ReNeLLM](https://github.com/NJUNLP/ReNeLLM) (pinned at commit
`a61c39e`), the official implementation of:

> Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and
> Shujian Huang. "A Wolf in Sheep's Clothing: Generalized Nested Jailbreak
> Prompts can Fool Large Language Models Easily." NAACL 2024.
> [arXiv:2311.08268](https://arxiv.org/abs/2311.08268).

## Technique

ReNeLLM generalizes prompt jailbreaks into two auxiliary-LLM-driven stages:

1. **Prompt rewriting** — a random count and order of six semantics-preserving
rewrite operations (paraphrase-shorten, misspell sensitive words, reorder
words, insert meaningless characters, partial translation, and style/slang
change) are applied to the goal. A binary LLM judge confirms the rewrite is
still harmful; if not, it is retried from the original goal.
2. **Scenario nesting** — the rewritten goal is embedded into one of three
benign carriers (code completion, table filling, or text continuation).

The nested prompt is sent to the model under attack, and its reply is scored by
the same judge. The loop repeats up to a budget until the reply is judged
harmful.

## Mapping to superred

- **One superred run == one ReNeLLM outer iteration.** On `RunStart` the per-run
state is re-armed and the rewritten+nested prompt is produced **pre-send**
(superred has no mid-conversation rewind, so all refinement happens before the
send).
- The nested prompt is injected on the first eligible free-text **user** surface
at `ControllablePreCall`. The model under attack is the **target itself**, not
an LLM this optimizer constructs.
- The reply is read from `ControllablePostCall` (with a trajectory-observable
fallback) and scored at `RunEnd` to decide stop/continue, up to `iter_max`
runs.
- Every **auxiliary** LLM call (the six rewrite operations and the judge) runs
the byte-identical vendored upstream code, routed to the constrained
`self.llm` (superred's `LLMClient`). No sampling temperature is ever set.

The upstream rewrite/nest/judge helpers are vendored byte-for-byte under
`src/renellm_optimizer/_vendor/renellm/utils/` and executed; only upstream's
openai/anthropic SDK completion helper is replaced (by
`src/renellm_optimizer/_shim.py`). See `ASSUMPTIONS.md` for every deviation and
`NOTICE` for attribution.

## Usage

```python
from renellm_optimizer import ReNeLLMOptimizer

optimizer = ReNeLLMOptimizer(
iter_max=20, # max runs (outer iterations); upstream default
max_rewrite_attempts=20, # cap on the rewrite-retry loop
seed=None, # reproducible operation/scenario selection
)
```

The optimizer is instantiated by the superred controller, which supplies the
goal, controllables, observables, and the constrained `LLMClient`.

## Development

```shell
# offline test battery
PYTHONPATH=src python -m pytest -q tests

# verify vendored files are byte-identical to the pinned upstream commit
python scripts/sync_upstream.py --check
```

## License

MIT (see `LICENSE`). Vendored ReNeLLM code is MIT (Copyright (c) 2024 NJUNLP);
its text is preserved in `LICENSES/ReNeLLM-MIT.txt`.
Loading
Loading