This personal learning and research project grew out of the ARENA Intro to Mechanistic Interpretability exercises. I drew inspiration from the course and adapted parts of its induction-head analysis workflow, including repeated-token experiments, attention-pattern inspection, and QK/OV circuit analysis. Credit for the educational material and these foundations belongs to the ARENA authors and contributors. Here I explore the workflow on a different pretrained checkpoint and extend it with additional controls, interventions, and uncertainty reporting.
The initial full analysis is complete, with committed code, figures, tables, and experiment manifests. Further validation and methodological improvements remain ongoing. Results are exploratory and should not be treated as a definitive identification of the canonical induction circuit.
Remaining work includes evaluating causal interventions on held-out seeds, reporting paired contrasts of how ablations change the repetition advantage, and extending runtime Q/K decomposition and logit attribution. This project uses a different checkpoint from the ARENA induction-head exercises and is not a completed replication of their demonstration.
This repository asks whether a canonical induction circuit can be identified
and causally validated in TransformerLens’ pretrained attn-only-2l model. The
investigation uses repeated random-token sequences, automatic head discovery,
weight-level OV/QK analysis, activation ablation, activation patching, direct
path controls, and stress tests.
The model shows a modest relationship between the first copy and clean reference prediction at the second copy: permuting the first copy while holding the second-copy query and clean target fixed increases NLL by 0.050 on average across 64 examples. But the proposed circuit is not established. The strongest raw induction-stripe score is only 0.0185, below the position-matched uniform causal reference of 0.0413; the global winner changes from L0H6 on discovery seeds to L1H2 on held-out seeds. The strongest previous-token head is L0H3, but the L0H3 → L1H2 composition statistic is −0.026, while another pair reaches 2.052. Candidate ablations are not selective relative to controls, and direct path effects are small.
The defensible conclusion is therefore deliberately narrow: this checkpoint has repeat-related behavior and a clear predecessor-head signature, but these experiments do not identify a uniquely necessary canonical induction circuit. The repository preserves the nulls and causal ambiguity rather than tuning the analysis to force a positive result.
- Behavior: clean second-copy expected-token NLL is 12.039, versus 12.090 when only the first copy is permuted and the second-copy query and clean target are preserved. The clean second-copy NLL is 0.019 higher than the first-copy reference NLL, so this is not a positive first-to-second “advantage.”
- Discovery: L0H3 has previous-token score 0.673. The train induction winner is L0H6 (0.0185); the held-out winner is L1H2 (0.0177), so the global winner is not stable even though all five top candidates overlap.
- OV: L1H2 has copying margin 0.452, but L1H6 is a much stronger copier (5.265). Runtime projection checks have maximum error below 6e−6; this validates the hook/weight orientation, not the full copying mechanism.
- Ablation: zero ablation changes clean NLL by −0.303 for L0H6 and −0.114 for L1H2, while controls reach +1.004 and +1.116. Negative means that this loss endpoint improved after ablation, so these are not necessity results.
- Activation patching: raw paired NLL improvement is 0.0095 (95% bootstrap CI [−0.0016, 0.0199]) for the candidate head output and 0.0120 [0.0024, 0.0218] for its query. These raw effects are primary; normalized restoration is secondary and has unstable per-example behavior.
- Path patching: the source-head → candidate-key approximation has raw NLL effect 0.00089 [0.00035, 0.00145], versus 0.00036 [−0.00001, 0.00074] for the source-head control path. This is not exact runtime path isolation because the hook bypasses intervening layer normalization.
All report numbers above come from the committed full-run tables under
results/full. Standard deviations and bootstrap intervals
are labeled in those tables; intervals are not significance tests.
The central question is:
Can we identify an induction circuit in a pretrained Transformer, decompose the mechanism that implements it, and causally verify that the identified components are responsible for induction behavior?
The workflow separates observations, correlations, mechanistic hypotheses, and causal evidence. The data are synthetic random-token sequences so that the analysis does not depend on English semantics.
For a repeated structure
[A] [B] ... [A] -> [B]
an induction head can use the earlier occurrence of [A] to find the earlier
[B] and copy it to the new occurrence. The code evaluates the second-copy
[A] positions and points the induction stripe at the first-copy continuation
[B] positions.
A previous-token head attends from [B] to the immediately preceding [A].
Its output can write information about [A] into the residual stream at the
[B] position. A later head could use that signal when forming keys.
The QK circuit determines where a head attends; the OV circuit determines what is moved from the attended source. The installed TransformerLens convention is checked in code with explicit shapes:
W_E: [vocab, d_model]
W_V: [d_model, d_head]
W_O: [d_head, d_model]
W_U: [d_model, vocab]
The sampled token-to-token OV map is W_E @ W_V @ W_O @ W_U. The composition
analysis uses analogous products for a previous-head output entering a later
head’s key projection. These are weight-level hypotheses, not complete runtime
causal proofs: positional terms, layer normalization, and other residual paths
remain possible.
- Model:
attn-only-2l, loaded on CPU through the verifiedHookedTransformer.from_pretrained("attn-only-2l", device="cpu", dtype=torch.float32)API in TransformerLens 3.7.2. The loader is isolated insrc/induction_circuits/model.py. - Data: unique random token IDs from a fixed vocabulary subset, with an explicit BOS and two copies of the sampled sequence.
- Controls: actual next-token prediction; clean-reference prediction; first-copy versus second-copy scoring; source-corrupted inputs that permute only the first copy; second-copy-permuted inputs; alternative stripes; uniform causal attention; held-out seeds; and unrelated heads.
- Uncertainty: full defaults use four batches of 16 examples. Intervention tables preserve per-example paired effects and report percentile bootstrap CIs. Example SD is descriptive variability, not a significance test.
- Hardware: no training or gradients. Quick and full end-to-end runs are practical on CPU after the checkpoint is cached.
Every head is scored without hardcoded identities. The raw induction score is
attention mass on the indexed stripe; it is compared against the uniform
causal baseline, an alternative stripe, corrupted inputs, and held-out seeds.
The full clean discovery table is
head_scores_summary.csv.
The uniform reference is 0.0413, while every train raw induction score is below it. This means the ranking identifies relative candidates, but does not establish that any head has a meaningful induction signature above this basic null. L0H3 is still a clear previous-token candidate, with score 0.673.
The heatmaps aggregate configured batches rather than displaying a selected prompt. The stripe statistic and predecessor diagonal are computed from the same explicit indices used by the saved metrics.
For sampled vocabulary IDs, the copying score is the source-token output-logit margin over sampled alternatives. L1H2 is positive but L1H6 is much stronger, so a copying-like OV map does not identify the induction head by itself.
The implementation also independently captures hook_z and the model’s
attention-output hook, checking z @ W_O + b_O against runtime output. The
algebraic composed/sequential equality is retained as an orientation check,
not described as independent validation of the complete runtime mechanism.
L0H3 has a strong predecessor diagonal. This is a behavioral/attention observation and supports a mechanistic hypothesis, but does not show that the head supplies information used by a later induction head.
The hypothesized L0H3 → L1H2 key-composition statistic is −0.026. The strongest listed pair is L0H3 → L1H6 at 2.052. This weakens the narrow candidate-pair story and is reported as an alternative interpretation rather than replaced with a more favorable pair.
The ablation script zeros selected hook_z head activations and also performs
candidate mean ablation. It evaluates clean and source-corrupted inputs under
the same intervention. Candidate effects are not selective: L0H6 and the
previous-token candidate improve the clean NLL after zero ablation, while some
controls strongly worsen it. This could reflect redundancy, a general
prediction role, or an out-of-distribution zero-ablation state.
This experiment uses clean and second-copy-permuted inputs with the same length and first-copy tokens, then scores the clean-reference targets. The primary estimand is the raw paired effect:
NLL effect = corrupted NLL − patched NLL
logit effect = patched logit − corrupted logit
The tables additionally contain the clean–corrupted denominator, its distribution, ratio-of-means, valid per-example ratios, and the number of poorly conditioned cases. No undefined ratio is silently replaced with zero. The source-key and source-previous-head patches are explicit zero-effect controls under this corruption because those source activations are unchanged.
Here the first copy is permuted while the second-copy query remains fixed. The
hook adds the clean-minus-corrupted source-head contribution through
z -> W_O -> W_K, with source-to-V, destination-to-Q, source-head, and
source-to-Q controls. The source-to-Q control is labeled structural: in this
two-layer model, changing final-layer queries at first-copy positions cannot
affect the evaluated second-copy outputs. Its zero effect therefore says
nothing about Q being less important. The destination-Q control is the
meaningful query-position comparison.
The direct hook bypasses intervening layer normalization, so the result is a linear direct-path approximation and not an exact runtime path decomposition.
The compact sweeps preserve per-example behavior and plot bootstrap CIs. The clean-versus-permuted NLL contrast is largest at sequence length 4 (0.367) and falls to 0.065 at length 24. Repeat-distance means range from 0.116 to 0.137 for gaps 0–8. Second-copy corruption effects are noisy and near zero across corruption levels. These are robustness probes, not a claim of universal induction behavior.
- The main behavioral effect is modest and the raw discovery score is below the uniform causal reference.
- Correlational attention and weight products do not prove mechanism use.
- Candidate ablations are not selectively associated with repeated behavior.
- Zero ablation can create an out-of-distribution residual state; mean ablation is a check, not a complete replacement for resampling ablation.
- Direct path patching bypasses layer normalization and must not be read as exact path isolation.
- The random-token vocabulary subset and sequence design may affect signal.
- GPT-2 Small generalization, a fully normalization-aware exact path decomposition, and a full factorial stress grid were deferred to preserve a reproducible scope.
- The legacy-compatible model loader is pinned and isolated, but upstream API deprecation may require a future migration.
The project demonstrates a full mechanistic workflow: controlled behavior, automatic head discovery, weight-level OV/QK analysis, composition hypotheses, ablation, activation patching, direct-path controls, and stress testing. It does not establish the proposed canonical induction circuit in this checkpoint. The positive behavioral and predecessor-head observations are consistent with the story, but the below-uniform discovery score, unstable winner, nonselective ablations, weak composition alignment, and small path effects prevent a stronger causal conclusion. The experiments also do not constitute strong evidence against induction in general: the intervention design and direct-path approximation leave important alternatives open.
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytest
# Fast end-to-end smoke run; writes only results/quick and figures/quick.
python scripts/run_all.py --quick
# Full CPU run used for the committed report artifacts.
python scripts/run_all.pyThe first model run downloads attn-only-2l and its tokenizer through
TransformerLens/Hugging Face. Checkpoints are not committed. Every experiment
writes a manifest under its output directory. If a downstream script sees a
different configuration fingerprint, it stops instead of silently mixing
artifacts.
src/induction_circuits/ reusable data, metrics, model, circuits, hooks
scripts/ ordered reproducible experiments and run_all.py
tests/ indexing, shape, hook, and integration checks
results/full/ committed full-run metrics, tables, manifests
results/quick/ ignored smoke-test outputs
figures/full/ committed report figures
figures/quick/ ignored smoke-test figures
notebooks/ optional walkthrough, never the source of truth
- Elhage et al., A Mathematical Framework for Transformer Circuits, 2021.
- Olsson et al., In-context Learning and Induction Heads, 2022.
- TransformerLens, the activation-level analysis library used here.
- TransformerLens patching utilities, implementation reference for activation interventions.
- Fred Zhang and Neel Nanda, Towards Best Practices of Activation Patching in Language Models, 2024.
- Goldowsky-Dill et al., Towards Automated Circuit Discovery for Mechanistic Interpretability, 2023.
This investigation builds on my previous implementation of the GPT-2 architecture from scratch in PyTorch, where I implemented embeddings, layer normalization, causal multi-head attention, MLPs, residual Transformer blocks, and the final language-model forward pass: https://github.com/pegruk/gpt2-from-scratch.
That repository does not contain pretrained weights or induction heads, and this project does not depend on it. The conceptual connection is:
building the architecture -> reverse-engineering learned computations inside a pretrained Transformer








