Localising and editing LLM traits and fine-tuned habits via component-level causal patching. Demonstrates repairing misaligned LoRAs by swapping 17 attention heads with clean base model activations.
alignment mechanistic-interpretability qwen llm-safety llama3 causal-patching function-vectors model-diff
-
Updated
Aug 2, 2026 - Jupyter Notebook