Refusal is the functional behavior that enables safety-aligned language models to reject harmful or unethical requests. Recent work in mechanistic interpretability has shown that refusal can be represented through directions or decision boundaries in the model's latent space. This repository explores a complementary perspective: refusal suppression can be formulated as a latent-space evasion problem against linear refusal probes.
We introduce Controlled Latent-space Evasion (CLE), a family of attacks that intervene on internal activations in order to move harmful-prompt representations across per-layer refusal boundaries. Instead of removing a fixed refusal direction, CLE uses trained linear probes to compute controlled perturbations at selected transformer layers. The attack is governed by a small set of interpretable hyperparameters: the intervened layers, the projection margin, the probe type, and the perturbation strength.
This codebase contains the main scripts for training refusal probes, searching CLE hyperparameters, generating completions, and evaluating attack success with LLM-based judges.
We start from contrastive harmful and harmless prompt splits and extract post-instruction latent representations from the target language model. For each transformer layer, we train a linear probe that separates harmful from harmless representations. These probes define layer-wise refusal boundaries in activation space.
Given a harmful prompt, CLE modifies the model's internal states so that selected activations evade these learned boundaries. The intervention uses the probe normal vector and intercept to compute a controlled movement toward the non-refusal side of the boundary. A margin parameter controls how far the activation is pushed beyond the boundary, while a beta parameter scales the strength of the intervention.
The repository implements two variants:
cle-a.py: CLE-A, the additive variant. It computes a fixed per-layer delta from the prompt forward pass and applies that delta during generation.cle-p.py: CLE-P, the projective variant. It applies the controlled projection directly during both prefill and autoregressive generation.
Both variants share the same probe format and expose a similar command-line interface for models, layers, margins, datasets, batching, and evaluation.
classifier/train_latent.py: trains per-layer harmful/harmless refusal probes.cle-a.py: runs the additive CLE-A attack.cle-p.py: runs the projective CLE-P attack.optuna_search.py: performs Bayesian optimization over layer windows and margins.utils/eval_jailbreaks.py: evaluates saved completions with HarmBench judges.utils/probes.py: loads SVM and single-direction probe artifacts.utils/hooks.py: contains the activation hooks used by CLE-A and CLE-P.utils/runtime.py: handles model loading, dataset loading, evaluation dispatch, and common runtime setup.models/: contains model wrappers and the Hugging Face model registry.dataset/: contains prompt splits and processed validation/test prompt files.
Generated probes, completions, evaluations, plots, and Optuna studies are written under dataset/representations/, completions/, and runs/.
We show here the general pipeline used to reproduce the experiments.
Create a fresh Python environment and install the required dependencies:
python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install -r requirements.txtIf you use gated Hugging Face models, set your token before running the scripts:
export HF_TOKEN=<your_huggingface_token>Train per-layer linear probes on the model-specific harmful and harmless training prompts:
python -m classifier.train_latent \
--model_name <model_name> \
--device cuda \
--artifact_dir ./dataset/representations \
--n_samples 128The script writes probe checkpoints and train representations under:
dataset/representations/<model_name>/train_svm/
The training splits used by the probe-training script are stored in dataset/splits/.
Run Bayesian optimization on the validation set to search for the best contiguous layer window and shared margin:
python optuna_search.py \
--target pipeline \
--model_name llama2-7b \
--device cuda:0 \
--dataset harmbench_val \
--margin_low 0.1 \
--margin_high 5.0 \
--margin_step 0.1 \
--layer_start_low 3 \
--layer_start_high 24 \
--layer_end_low 6 \
--layer_end_high 32 \
--trials 500 \
--sampler tpeUse --target projection to optimize CLE-P instead of CLE-A.
During validation, optuna_search.py evaluates completions with the HarmBench validation judge. Each trial writes generated completions, evaluations, and Optuna artifacts under:
completions/<model_name>/<target>/
runs/<model_name>/<target>_optuna/
The best configuration is saved in the Optuna result summary:
{
"best_params": {
"layer_start": 5,
"layer_end": 25,
"margin": 3.1
},
"best_user_attrs": {
"layers_arg": "5-25",
"margin": 3.1,
"evaluation_path": "..."
}
}Use best_user_attrs.layers_arg and best_user_attrs.margin for final test-set generation.
Run the additive variant with the selected layers and margin:
python cle-a.py \
--model_name llama2-7b \
--device cuda:0 \
--svm_dir ./dataset/representations/llama2-7b/train_svm \
--layers 5-25 \
--margin 3.1 \
--batch_size 32 \
--dataset harmbench_test \
--out_dir ./completions/llama2-7b/cle-aCLE-A first computes a per-layer additive delta at the prompt position and then keeps that delta active while the model generates.
Run the projective variant with the selected layers and margin:
python cle-p.py \
--model_name llama2-7b \
--device cuda:0 \
--svm_dir ./dataset/representations/llama2-7b/train_svm \
--layers 5-25 \
--margin 1.6 \
--batch_size 32 \
--dataset harmbench_test \
--out_dir ./completions/llama2-7b/cle-pCLE-P keeps projection hooks active throughout prefill and generation.
Add --evaluate to either CLE command if you want to run evaluation immediately after generation.
You can also evaluate an existing completions file:
python utils/eval_jailbreaks.py \
--completions_path ./completions/llama2-7b/cle-a/completions_<run_tag>.json \
--methodologies harmbench \
--evaluation_path ./completions/llama2-7b/cle-a/evaluation/evaluation_<run_tag>.jsonFinal test-set evaluation uses the Llama-based HarmBench judge through --methodologies harmbench. This is intentionally different from validation-time optimization, which uses the Mistral-based validation judge.
To intervene on all available probe layers, pass:
--layers allTo intervene on a contiguous layer window, pass an end-exclusive interval:
--layers 5-25To intervene on specific layers, pass a comma-separated list:
--layers 10,14,20CLE also supports per-layer margins aligned with the selected layers:
python cle-a.py \
--model_name llama2-7b \
--device cuda:0 \
--svm_dir ./dataset/representations/llama2-7b/train_svm \
--layers 5-8 \
--margin 3.1 \
--batch_size 32 \
--layer_margins 2.9,3.1,3.3 \
--dataset harmbench_testThe scripts can load two probe types:
--probe_type svm: loads learned SVM probes fromsvm_layerXX.pt.--probe_type single_direction: loadssd_layerXX.ptdirections and derives the midpoint bias fromHFx_train.ptandHLx_train.pt.
This work was partly supported by the EU-funded Horizon Europe projects Sec4AI4Sec (GA No. 101120393) and CoEvolution (GA No. 101168560), by project FISA-2023-00128, funded under the MUR program “Fondo Italiano per le Scienze Applicate,” and by Fondazione di Sardegna under the project “LatentShield: Protecting Large Language Models from Prompt Injection in Latent Space” (CUP: F83C26000350007).



