Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ jobs:
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install -e ".[dev]"
- run: pip install -e ".[dev,sklearn,parallel,progress]"
- run: ruff check .
- run: ruff format --check .
- run: mypy
Expand All @@ -33,5 +33,5 @@ jobs:
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- run: pip install -e ".[dev]"
- run: pip install -e ".[dev,sklearn,parallel,progress]"
- run: pytest
3 changes: 2 additions & 1 deletion .pre-commit-config.yaml
Original file line number Diff line number Diff line change
@@ -1,6 +1,7 @@
repos:
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.14.2
# Keep in step with the ruff constraint in pyproject.toml's dev extra.
rev: v0.16.8
hooks:
- id: ruff-check
args: [--fix]
Expand Down
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,17 @@ is `0.x` the public API may change with a minor bump, always with an entry here.
latter needs the research repository present and is removed once the
extraction is complete.

- Public API: `score`, `score_edges`, `block_measures_frame`, and the
`GargAmlScorer` scikit-learn wrapper (optional, loaded on first use so that
scikit-learn stays out of the import path). `__all__` now lists exactly what
is supported.
- `smurfing_graph`, a networkx generator of graphs with known patterns, so the
documentation and tests run with no data download.
- `n_jobs` on the scoring entry points, and `progress` for a progress bar.
- A documentation site: quickstart, how it works, five guide pages, scaling,
limitations, API reference and the decision log. Every `>>>` example in the
docs is executed by the test suite.

### Changed

- PEP 8 names throughout, with the two per-node measure functions merged into
Expand Down
58 changes: 58 additions & 0 deletions docs/guide/alerts.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# Turning scores into alerts

A score per account is not an alert list. Two decisions turn one into the other,
and both are yours.

## Threshold or top-K?

**Top-K** — take the K highest-scoring accounts — matches how alerting actually
works: a team can review a fixed number of cases per day, and K is that number.
It is also what the evaluation metrics in the paper are built around.

**A fixed threshold** — everything above 0.7, say — is tempting but brittle. The
score distribution shifts with the graph, the preprocessing and the period, so
a threshold tuned in March is a different alert volume in June.

Start with top-K sized to your review capacity. Read
[the note on ties](interpreting.md#ties) first — at the top of the distribution
they are common enough to make a naive top-K partly arbitrary.

## The imbalance

Laundering labels are extremely rare: **under 0.1%** of accounts in the paper's
datasets. Three consequences worth internalising before reading any metric:

- **Accuracy is meaningless.** Flagging nothing scores over 99.9%.
- **A "low" precision may be excellent.** At a 0.1% base rate, precision of 5%
is a fifty-fold lift over random review.
- **AUC-ROC flatters everything.** With this much imbalance it is dominated by
the easy negatives. Prefer precision@K, recall@K, and average precision.

## Measuring

If you have labels, measure at the K you will actually use:

```python
scores = ga.score(graph)["GARGAML"]
ranked = scores.sort_values(ascending=False, kind="stable")

k = 100
flagged = ranked.head(k).index
hits = int(labels.reindex(flagged).sum())

print(f"precision@{k}: {hits / k:.1%}")
print(f"recall@{k}: {hits / labels.sum():.1%}")
print(f"lift@{k}: {(hits / k) / labels.mean():.1f}x")
```

Lift is the honest headline: how much better than reviewing accounts at random.

## Using it alongside what you have

GARG-AML is a structural signal, not a complete system. It is at its most useful
as one feature among several, not as a standalone alert generator — a high score
plus an unusual amount profile is far more actionable than either alone. See
[using scores as features](features.md).

The paper positions it the same way: the score is strong on its own, and
stronger when a model combines it with neighbourhood statistics.
60 changes: 60 additions & 0 deletions docs/guide/data.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# Preparing your data

## What GARG-AML reads

Two columns: who paid, and who was paid. That is all.

Amounts, timestamps, currencies and transaction counts are **not** used. The
method reads the shape of the network. Repeated transactions between the same
pair of accounts collapse to a single edge, and self-transfers are dropped.

This is a real limitation as well as a strength — see
[Limitations](../limitations.md).

## From a table

```python
scores = ga.score_edges(transactions, "payer_account", "payee_account")
```

## From a graph

If you already build a networkx graph, pass it directly:

```python
scores = ga.score(graph)
```

Node ids come back exactly as you supplied them — strings, integers, tuples —
as the index of the returned frame, so you can join the scores straight back
onto your own tables.

## What counts as an account

Whatever you make a node. That choice matters more than any parameter in this
package:

- **One node per account** is the usual choice and what the paper evaluates.
- **One node per customer**, merging their accounts, will find schemes that
spread across accounts of the same person — and will hide schemes that use
several accounts of one customer as the mules.
- **One node per bank** is too coarse; the pattern disappears into aggregate.

## Directed or undirected?

Both are supported. `score()` follows the graph you give it: a `DiGraph` gets
the directed analysis, a `Graph` the undirected one.

Start with **undirected**. In the paper's experiments the undirected score
outperforms the directed one, which is counter-intuitive given that
one-directional flow is definitional for smurfing. The directed variant is
stricter about which accounts count as senders and receivers, and that
strictness appears to cost more than the extra information gains.

If you have direction, it is still worth trying both on your own data.

## Scale

Nothing here needs the whole graph in one piece conceptually, but networkx does
hold it in memory. For rough numbers and what to do about a graph that does not
fit, see [Scaling](../scaling.md).
68 changes: 68 additions & 0 deletions docs/guide/features.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
# Using scores as features

The score is one number per account. The paper's stronger results come from
feeding it, plus a summary of the neighbourhood around it, into an ordinary
classifier.

## The feature table

```python
scores = ga.score(graph)
features = ga.build_features(graph, scores)
```

Ten columns:

| Column | Meaning |
|---|---|
| `GARGAML` | the account's own score |
| `GARGAML_min/mean/max/std` | the same over its direct counterparties |
| `degree` | its number of counterparties |
| `degree_min/mean/max/std` | the same over its counterparties |

The neighbourhood statistics are what let a model distinguish an account that is
*in* a pattern from one that merely *touches* one. A mule sits among other
high-scoring accounts; a legitimate business that happens to pay one mule does
not.

!!! warning "Use the same graph"
Pass `build_features` the graph the scores were computed on. If you scored a
reduced graph, the degrees must come from the reduced graph too, or the
features describe two different networks.

## With scikit-learn

```python
from garg_aml import GargAmlScorer

features = GargAmlScorer(reduce=True, resolution=10).fit_transform(graph)
```

Needs the `sklearn` extra. Nothing is learned — GARG-AML is closed-form, so
`fit` has no parameters to estimate. The estimator exists so the step composes
in a pipeline, not because there is a model inside it.

## A caution on evaluation

Training a classifier on these features is **transductive**: every account was
scored as part of one graph, so a train/test split on the feature table does not
separate the test accounts from the training ones structurally — they were
neighbours when the scores were computed.

That is not wrong, but it is not the same as a held-out evaluation, and a
split-on-the-feature-table number will be optimistic relative to scoring a
genuinely unseen period. Say which one you are reporting.

## Building your own

The pieces are public if the ten columns are not what you want:

```python
from garg_aml import neighbour_score_stats, neighbour_degree_stats

stats = neighbour_score_stats(graph, "account", scores["GARGAML"].to_dict())
```

And the raw block densities are available through
`ga.score(graph, return_measures=True)` if you would rather let a model see the
components than the aggregated score.
75 changes: 75 additions & 0 deletions docs/guide/interpreting.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
# Interpreting a score

## The range

| Score | Structure |
|---|---|
| **1** | a textbook pattern: on-diagonal blocks empty, off-diagonal full |
| **~0.5** | clearly smurfing-shaped, with some background activity mixed in |
| **~0** | ordinary: a group that trades among itself scores here |
| **-1** | no neighbourhood at all — nothing to measure |

The score is a **structural resemblance, not a probability**. 0.8 does not mean
80% likely to be laundering. It means this account's neighbourhood looks 80% of
the way towards the idealised pattern, on a scale where a clique sits near 0.

## Scores are comparable within a run, not across runs

Two things change the scale underneath you: the preprocessing
([`reduce_graph`](preprocessing.md) changes which neighbourhoods exist) and the
`score_type`. Comparing a score computed on a reduced graph against one computed
on the full graph is meaningless. Fix both before comparing anything.

## Everyone in the pattern scores high

A mule's neighbourhood has the same block structure as the source's: its
counterparties (source and target) do not deal with each other. So the mules
score as high as the account that organised the scheme, sometimes higher.

Read a high score as **"this account sits in a smurfing-shaped subgraph"**, and
expect to investigate the subgraph rather than the single account. Pull the
account's second-order neighbourhood and look at it.

## Ties

**The score ties heavily**, especially at the top. Many accounts land on exactly
1.0 — in the paper's bank-level analysis, 522 of one institution's 2,639
customers shared exactly 1.0.

This matters the moment you rank. `sort_values().head(50)` will silently break
those ties by whatever order the frame happens to be in, which is an artefact of
your data loading, not a signal. If you take a top-K, check how many accounts
sit at the boundary score first:

```python
scores = ga.score(graph)["GARGAML"]
cutoff = scores.nlargest(50).min()
print((scores == cutoff).sum(), "accounts tied at the cut-off")
```

If that number is large, the top-50 is not 50 accounts — it is an arbitrary
50 drawn from a bigger tied set. Break the tie on something you trust
(transaction volume, exposure, customer risk rating) rather than on sort order.

## Why did this account score high?

Ask for the block measures:

```python
detailed = ga.score(graph, return_measures=True)
detailed.loc["account_of_interest"]
```

`measure_2` is the density of the off-diagonal block — how completely the
counterparties connect through. `measure_1` and `measure_3` are the on-diagonal
blocks that should be empty. A high score with a tiny `size_2` means the
pattern is real but tiny, and probably not interesting.

That decomposition is the point of the method: the score is always reducible to
a handful of interpretable densities.

## What a low score does not mean

A low score is not evidence of legitimacy. It means *this particular structure*
is absent. Laundering that does not route through mules — cash, trade
mis-invoicing, a single large transfer — has no reason to score high.
74 changes: 74 additions & 0 deletions docs/guide/preprocessing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# Preprocessing

Two optional steps. Neither is applied unless you ask.

## Community reduction

```python
reduced = ga.reduce_graph(graph, resolution=10)
scores = ga.score(reduced)

# or, equivalently
scores = ga.score(graph, reduce=True, resolution=10)
```

`reduce_graph` partitions the graph with Louvain and **keeps only the edges
whose endpoints are in the same community**. This is Algorithm 1 of the paper,
and it is how the published results were produced.

!!! warning "This is lossy, and the loss is large"
On the paper's IBM dataset this removes the large majority of edges. It does
not merely thin the graph — it changes which second-order neighbourhoods
exist at all, and therefore which patterns *can* be found.

Why do it anyway? Cost. Scoring is driven by neighbourhood size, and in a dense
transaction network a handful of hub-adjacent accounts have enormous
neighbourhoods. Cutting between-community edges makes those tractable.

### Choosing a resolution

Higher resolution means smaller communities and so more edges cut.

```pycon
>>> import networkx as nx
>>> import garg_aml as ga
>>> graph = nx.disjoint_union(nx.complete_graph(4), nx.complete_graph(4))
>>> graph.add_edge(0, 4)
>>> ga.reduce_graph(graph, resolution=1).number_of_edges()
12
>>> ga.reduce_graph(graph, resolution=10).number_of_edges()
0

```

Two cliques joined by one edge. At resolution 1 the bridge is cut and both
cliques survive. At the published resolution of 10 the communities are smaller
than a clique, so **every** edge goes and all eight accounts are left isolated.

Nothing is dropped — isolated accounts remain in the output and score -1 — but
they can no longer be scored meaningfully. If a large share of your accounts
come back at -1, the resolution is too high for your data.

The default is 10 because that is what the paper used. It is not a
recommendation for your network. Sweep it and look at how many accounts survive
with a neighbourhood.

## Hub removal

```python
pruned = ga.drop_hubs(graph, quantile=0.01)
```

Removes the highest-degree accounts — payment processors, exchanges, salary
accounts — which have huge neighbourhoods and are rarely what you are looking
for. Not part of the published pipeline, but useful when a few accounts
dominate the runtime.

Note that accounts at the threshold degree are all removed, so ties can push the
number removed above the nominal fraction.

## Reproducibility

Both steps are deterministic given a seed. `reduce_graph` takes `seed=1997` by
default — the value used throughout the paper — because Louvain is otherwise
not reproducible run to run. Keep it fixed when comparing runs.
Loading
Loading