Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
74c0bd0
Document and automate GPU predictor smoke tests
jsalsman Sep 19, 2026
4212de8
Fix reusable GPU setup version handling
jsalsman Sep 19, 2026
efefa58
Merge pull request #1 from jsalsman/codex/apply-patch-to-readme-and-a…
jsalsman Sep 19, 2026
a70fcaa
Fix GPU setup reruns and repository-relative smoke tests
jsalsman Sep 19, 2026
5ac47d5
Resolve GPU virtual environment paths
jsalsman Sep 19, 2026
2b1af1d
Merge pull request #2 from jsalsman/codex/fix-gpu-setup-reruns-and-sm…
jsalsman Sep 19, 2026
5bf24f4
Use credential-free phonemizer dependency
jsalsman Sep 19, 2026
fea749c
Update dependency installation documentation
jsalsman Sep 19, 2026
e0543ae
Merge pull request #3 from jsalsman/codex/update-phonemizer-dependenc…
jsalsman Sep 19, 2026
bcee634
Add opt-in cached KenLM model download
jsalsman Sep 19, 2026
cbfb775
Honor forced language model refreshes
jsalsman Sep 19, 2026
af50dfc
Accept manually installed KenLM binaries
jsalsman Sep 19, 2026
f575e90
Block ARPA fallback for rejected KenLM binary
jsalsman Sep 20, 2026
055d50e
Remove stale binary when installing ARPA model
jsalsman Sep 20, 2026
7912a86
Merge pull request #4 from jsalsman/codex/update-test_gpu_predictors-…
jsalsman Sep 20, 2026
c1861d4
Range-extract Zenodo language model
jsalsman Sep 20, 2026
dabbc7d
Merge pull request #5 from jsalsman/codex/revise-downloader-in-tools/…
jsalsman Sep 20, 2026
71ec9e3
Correct Zenodo language model checksum
jsalsman Sep 20, 2026
6197d0c
Merge pull request #6 from jsalsman/codex/update-zenodo-model-checksu…
jsalsman Sep 20, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,34 @@ on:
branches: ["**"]

jobs:
dependency-resolution:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"

- name: Resolve dependencies without GitHub credentials
env:
GIT_TERMINAL_PROMPT: "0"
run: |
python -m venv /tmp/pathbench-clean
/tmp/pathbench-clean/bin/python -m pip install --upgrade pip
/tmp/pathbench-clean/bin/python -m pip install \
--dry-run --ignore-installed --report /tmp/install-report.json .
/tmp/pathbench-clean/bin/python - <<'PY'
import json
from pathlib import Path

report = json.loads(Path("/tmp/install-report.json").read_text())
urls = [item["download_info"]["url"] for item in report["install"]]
vcs_urls = [url for url in urls if url.startswith("git+") or "github.com" in url]
assert not vcs_urls, f"VCS/GitHub dependencies remain: {vcs_urls}"
PY

test:
runs-on: ubuntu-latest
strategy:
Expand Down
239 changes: 230 additions & 9 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,25 @@ score = evaluator.score("utt1", "/path/to/audio.wav", transcription="the cat sat
print(f"ArtP score: {score}")
```

ArtP is reference-based: it force-aligns the phonemes in the supplied transcription
with the audio, so the transcription and its language are required. DArtP is
reference-free: a language-specific ASR model and n-gram language model first
produce a transcription, which is then scored by the same phonetic model:

```python
from pathbench import ArtPDoubleASREvaluator

evaluator = ArtPDoubleASREvaluator(language="en-us")
score = evaluator.score("utt1", "/path/to/audio.wav")
print(f"DArtP score: {score}")
```

DArtP currently supports `en`/`en-us`, `es`, `nl`, `it`, and `cmn`. It also
requires the corresponding file from the [n-gram model download](#n-gram-models)
to be placed in `lms/`; without it, scoring returns `None`. Run commands from
the repository root because DArtP resolves `lms/` relative to the working
directory.

### I want to contribute a new predictor to this repository, how do I do that?

See [CONTRIBUTING.md](CONTRIBUTING.md) for a step-by-step guide.
Expand Down Expand Up @@ -120,18 +139,85 @@ Results are written to the `results_11/` directory as timestamped text files con

## Installation

We are continously trying to make the installation easier for your use case.
We are continuously trying to make the installation easier for your use case.

### Complete GPU installation (install missing components only)

The following Ubuntu procedure is safe to re-run: it installs only absent apt
packages, builds the pinned `espeak-ng` unless its commit marker matches,
clones PathBench only when the checkout is absent, and lets the GPU helper reuse
an existing virtual environment and matching Python packages.

```bash
# 1. Install missing build prerequisites.
packages=(git python3 python3-venv build-essential cmake libfftw3-dev liblapack-dev)
missing=()
for package in "${packages[@]}"; do
dpkg-query -W -f='${Status}' "$package" 2>/dev/null | grep -q "ok installed" \
|| missing+=("$package")
done
if ((${#missing[@]})); then
sudo apt-get update -qq
sudo apt-get install -y "${missing[@]}"
fi

# 2. Build the reproducible phonemizer backend unless the pinned commit is installed.
espeak_ng_commit=2ea41210
espeak_ng_marker=/usr/local/share/pathbench/espeak-ng-commit
if ! command -v espeak-ng >/dev/null \
|| [[ ! -r "$espeak_ng_marker" ]] \
|| [[ "$(cat "$espeak_ng_marker")" != "$espeak_ng_commit" ]]; then
if { test -d /tmp/espeak-ng/.git \
|| git clone https://github.com/espeak-ng/espeak-ng.git /tmp/espeak-ng; } \
&& git -C /tmp/espeak-ng fetch origin "$espeak_ng_commit" \
&& git -C /tmp/espeak-ng checkout --detach "$espeak_ng_commit" \
&& cmake -S /tmp/espeak-ng -B /tmp/espeak-ng/build \
-DUSE_ASYNC=OFF -DBUILD_SHARED_LIBS=ON \
&& cmake --build /tmp/espeak-ng/build -j"$(nproc)" \
&& sudo cmake --install /tmp/espeak-ng/build \
&& sudo ldconfig \
&& sudo install -d "$(dirname "$espeak_ng_marker")"; then
printf '%s\n' "$espeak_ng_commit" \
| sudo tee "$espeak_ng_marker" >/dev/null
else
echo "Failed to install pinned espeak-ng; commit marker was not written." >&2
exit 1
fi
fi

# 3. Reuse the current checkout, or clone to a stable absolute destination.
if pathbench_root=$(git rev-parse --show-toplevel 2>/dev/null) \
&& test -f "$pathbench_root/tools/test_gpu_predictors.py"; then
: # Already anywhere inside a PathBench checkout.
else
pathbench_root=${PATHBENCH_ROOT:-"$PWD/pathbench"}
test -d "$pathbench_root/.git" \
|| git clone https://github.com/karkirowle/pathbench.git "$pathbench_root"
fi
cd "$pathbench_root"
python3 tools/test_gpu_predictors.py --download-language-model --cuda-version 12.4
```

This procedure assumes that a working NVIDIA driver is already installed;
`nvidia-smi` must list the assigned GPU. Driver installation is host- and
cloud-specific and is deliberately not attempted by the script.

If you have the opportunity to start from a clean AWS/GCE instance, please do so and follow the make installation.

If you are working on a highly restricted HPC cluster, I would recommend starting from the singularity container [provided](https://github.com/karkirowle/pathbench/releases/download/v0.1.0/pathbench.sif).

Package installation is the recommended pathway when you are trying to incorporate into your existing stuff. In this case, you are kind of your own figuring out
Package installation is the recommended pathway when incorporating PathBench
into an existing environment. In that case, you are responsible for resolving
dependency conflicts.

### Package installation
All Python runtime dependencies are available from package indexes rather than
VCS URLs. In particular, `phonemizer-fork==3.3.2` (which installs the
`phonemizer` import package) and `pyctcdecode==0.5.0` use versioned PyPI
releases, so installing PathBench does not require GitHub credentials. The
system-level espeak-ng revision below remains separately pinned because its
language-specific IPA output is part of the metric definition.

PathBench cannot be published to PyPI because it depends on Git-hosted forks of `phonemizer` and `pyctcdecode`.
### Package installation

**System dependencies** (not installable via pip — must be installed separately):
- `espeak-ng` at commit [`2ea41210`](https://github.com/espeak-ng/espeak-ng/commit/2ea41210) (post-1.52.0) — required by the phonemizer for grapheme-to-phoneme conversion. The exact commit matters: different espeak-ng versions produce different IPA symbols for some languages (e.g. Italian `ɾ` vs `r`), which affects phoneme-based metrics (PER, dPER, ArtP). Build from source:
Expand All @@ -141,7 +227,8 @@ PathBench cannot be published to PyPI because it depends on Git-hosted forks of
cmake -B build -DUSE_ASYNC=OFF -DBUILD_SHARED_LIBS=ON
cmake --build build -j$(nproc) && sudo cmake --install build
```
- PyTorch with CUDA support — install following [pytorch.org](https://pytorch.org/get-started/locally/) *before* installing pathbench
- PyTorch — install the CPU or CUDA build appropriate for your system by following
[pytorch.org](https://pytorch.org/get-started/locally/) *before* installing PathBench.

**Option A — Install from a GitHub Release:**
```bash
Expand All @@ -160,7 +247,10 @@ pip install "pathbench[scripts] @ git+https://github.com/karkirowle/pathbench.gi

### Make installation

The `make` installation route assumes the default setup of a standard Ubuntu 22.04 image (`ubuntu-2204-jammy`).
The `make` installation route assumes the default setup of a standard Ubuntu
22.04 image (`ubuntu-2204-jammy`) and Python 3.10–3.12. It creates
`tools/venv`. The default is a CPU-only PyTorch installation, which works for
inference and tests but is slower than a supported GPU.

```bash
sudo apt-get update -qq
Expand All @@ -171,12 +261,120 @@ cd /tmp/espeak-ng && git checkout 2ea41210
cmake -B build -DUSE_ASYNC=OFF -DBUILD_SHARED_LIBS=ON
cmake --build build -j$(nproc) && sudo cmake --install build && sudo ldconfig
cd -
git clone git@github.com:karkirowle/pathbench.git
git clone https://github.com/karkirowle/pathbench.git
cd pathbench/tools && make
cd ..
source tools/venv/bin/activate
```

For a CUDA build, select a wheel index supported by the pinned PyTorch version.
For example, PyTorch 2.6.0 provides CUDA 12.4 wheels:

```bash
cd pathbench/tools
make CUDA_VERSION=12.4
```

You can select a particular interpreter with, for example,
`make PYTHON=python3.12`. Re-running `make` resumes after completed stages;
run `make clean` first to rebuild the environment with a different Python,
PyTorch, or CUDA selection.

### GPU installation and predictor smoke test

After installing the pinned `espeak-ng` build above, systems with an NVIDIA GPU
and driver can use the helper script to create a separate CUDA environment and
run the focused ArtP and DArtP tests:

```bash
python tools/test_gpu_predictors.py --download-language-model --cuda-version 12.4
```

The script checks for `nvidia-smi` and `espeak-ng`, creates
`tools/gpu_venv`, installs the CUDA 12.4 builds of PyTorch and torchaudio 2.6.0,
installs PathBench and its test dependencies, verifies that PyTorch can access
the GPU, and runs both predictor tests. It can be invoked from any directory.
The first run downloads the Python packages and model checkpoints and therefore
requires network access and several gigabytes of free disk space.

Override its defaults with command-line options (or the corresponding
`PYTHON`, `PATHBENCH_CUDA_VERSION`, `PYTORCH_VERSION`, and `VENV` environment variables)
when needed. The selected CUDA wheel must exist for the selected PyTorch release:

```bash
python tools/test_gpu_predictors.py --python python3.11 --cuda-version 12.6 \
--pytorch-version 2.6.0 --venv /path/to/pathbench-gpu-venv
```

The script requires an NVIDIA driver compatible with the chosen CUDA wheel;
installing the wheel does not install a host GPU driver or the CUDA toolkit.
Language-model download is deliberately opt-in. With
`--download-language-model`, the helper uses HTTP range requests against the
immutable [35 GB Zenodo archive](https://zenodo.org/api/records/18738598/files/lms.zip/content)
to retrieve only the compressed English member: **8,582,666,912 bytes**
(approximately 8.0 GiB). It streams the raw DEFLATE data into a temporary file,
producing a **14,600,342,241-byte** model (approximately 13.6 GiB), and checks
the member metadata, CRC-32, expanded-model SHA-256
(`d786eec55174c696c0bf3327928ff496684f482194ba3c6ebdf4311acb823d00`),
and KenLM readability before atomically installing it. The server or any proxy
must support standards-compliant byte ranges (HTTP 206 and `Content-Range`);
the helper refuses an HTTP 200 response rather than accidentally downloading
the complete archive. Allow space for the installed model, its temporary
expanded copy, and at least a 1 GiB safety margin. The record is openly
accessible and licensed **CC BY 4.0**, which permits automatic download and
redistribution with attribution. Without the option, a missing model still
causes DArtP to be reported as skipped. ArtP does not need the model. Custom
standalone file/archive mirrors remain supported but must be supplied together with
`--language-model-sha256`; the project-specific environment equivalents are
`PATHBENCH_LANGUAGE_MODEL_URL`, `PATHBENCH_LANGUAGE_MODEL_SHA256`, and
`PATHBENCH_LANGUAGE_MODEL_CACHE`. The built-in range mode retains no compressed
archive and treats the verified model under `lms/` as its cache; the cache
directory applies to custom downloads.

The built-in Zenodo artifact identity used by the live integration is:

| Field | Published value |
| --- | --- |
| Archive (`lms.zip`) size | `35017940434` bytes |
| Member | `lms/wiki_en_token.arpa.bin` |
| Decompressed member size | `14600342241` bytes |
| Compressed member size | `8582666912` bytes |
| Member CRC-32 | `5afb90ef` |
| Decompressed member SHA-256 | `d786eec55174c696c0bf3327928ff496684f482194ba3c6ebdf4311acb823d00` |

The SHA-256 value is specifically the digest of the fully transferred and
decompressed `lms/wiki_en_token.arpa.bin` member. It is **not** a digest of
`lms.zip` or of the member's compressed DEFLATE stream.

For both tests together, allow **at least 12 GB of system RAM and 8 GB of GPU
VRAM**; **16 GB system RAM and 12–16 GB VRAM are recommended** to leave room
for both wav2vec2 models, the decoder, and transient activations. Any NVIDIA
CUDA GPU supported by the selected PyTorch wheel is acceptable; a T4 (16 GB),
L4 (24 GB), A10/A10G (24 GB), V100 (16/32 GB), or A100 works. Smaller 8 GB
cards may require closing other GPU processes and can run out of memory on
long audio. AMD ROCm GPUs, Apple GPUs, and CPU-only runtimes do not satisfy
this CUDA smoke test.

Google Colab GPUs can be used. Select a GPU runtime and confirm that
`nvidia-smi` works; the commonly assigned T4 and higher-memory L4/A100 options
meet the recommendation. Colab does not guarantee a particular GPU, RAM
amount, availability, or uninterrupted runtime, and its temporary filesystem
means the environment and downloaded checkpoints may need to be recreated in
a later session. If Colab assigns a smaller GPU or low-RAM runtime, inspect
`nvidia-smi` and available system memory before running the helper.
On a Python 3.12 T4 runtime, use PyTorch 2.6.0's CUDA 12.4 wheels:

```bash
python tools/test_gpu_predictors.py --download-language-model --cuda-version 12.4
```

A successful run ends with `2 passed`. A second invocation reuses both
`tools/gpu_venv` and the verified model cache rather than downloading them
again. Colab's local disk is ephemeral, however, so the installed 13.6 GiB
model and environment are lost when its runtime is recycled. Persist the
checkout itself if reuse across sessions is important (`PATHBENCH_LANGUAGE_MODEL_CACHE`
only controls custom standalone downloads).

**Without sudo access:** A containerised environment such as Docker is recommended.

## Downloads
Expand Down Expand Up @@ -213,7 +411,18 @@ find /path/to/your/datasets/easycall/EasyCall/m13 -name "m13 _*" -exec bash -c '

### N-gram models

The n-gram models required for DArtP and ArtP are included in the [Oral Cancer - YouTube](https://zenodo.org/records/18738598) download.
The n-gram models required by DArtP are included in the
[Oral Cancer - YouTube](https://zenodo.org/records/18738598) download. ArtP does
not require an n-gram model. Create `lms/` at the repository root and copy the
models there with these exact names:

| Language | Filename |
| --- | --- |
| English | `wiki_en_token.arpa` or `wiki_en_token.arpa.bin` |
| Dutch | `wiki_nl_token.arpa` or `wiki_nl_token.arpa.bin` |
| Spanish | `wiki_es_token.arpa.bin` |
| Italian | `wiki_it_token.arpa.bin` |
| Mandarin Chinese | `wiki_zh_token.arpa` or `wiki_zh_token.arpa.bin` |

## Testing

Expand All @@ -228,6 +437,19 @@ python -m pytest tests/test_evaluators.py::TestEvaluatorMethods -v

All tests should pass. If all evaluator tests fail simultaneously, the reference audio file in `tests/data/test_audio.wav` may be corrupted — the `test_audio_integrity` test will confirm this.

To test only ArtP and DArtP after downloading the English n-gram model, run:

```bash
python -m pytest \
tests/test_evaluators.py::TestEvaluatorMethods::test_articulatory_precision \
tests/test_evaluators.py::TestEvaluatorMethods::test_artp_double_asr -v
```

The first run downloads the Hugging Face checkpoints used by the phonetic and
English ASR models and therefore requires network access and several gigabytes
of free disk space. The DArtP test is reported as skipped, rather than failed,
when neither English n-gram filename listed above exists.

> **Note:** During the NAD evaluator tests you will see a `Wav2Vec2Model LOAD REPORT` table listing several keys (e.g. `project_q`, `quantizer`) as **UNEXPECTED**. These warnings are harmless — the keys belong to pre-training heads that are not needed for feature extraction and can be safely ignored.

### Dataset integrity
Expand Down Expand Up @@ -280,4 +502,3 @@ This work is partly financed by the Dutch Research Council (NWO) under project n
## Author

Bence Mark Halpern, Nagoya University

7 changes: 5 additions & 2 deletions docs/installation.rst
Original file line number Diff line number Diff line change
Expand Up @@ -39,5 +39,8 @@ Without sudo access, a containerised environment such as Docker is recommended.

.. note::

PathBench cannot be published to PyPI because it depends on Git-hosted forks
of ``phonemizer`` and ``pyctcdecode``.
PathBench's Python dependencies are available as versioned package-index
releases, including ``phonemizer-fork==3.3.2`` and
``pyctcdecode==0.5.0``. Installing them does not require GitHub credentials.
The ``espeak-ng`` shared library remains a separate system dependency; use
the revision documented in the project README for reproducible IPA output.
6 changes: 4 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -50,8 +50,10 @@ dependencies = [
"transformers",
"dtw-python",
"jiwer",
"phonemizer-fork @ git+https://github.com/thewh1teagle/phonemizer-fork.git",
"pyctcdecode @ git+https://github.com/kensho-technologies/pyctcdecode.git",
# PyPI release of the fork; it continues to expose the ``phonemizer`` API.
"phonemizer-fork==3.3.2",
# Keep every runtime dependency installable without VCS/GitHub access.
"pyctcdecode==0.5.0",
"praat-parselmouth>=0.4.4",
"scikit-learn",
]
Expand Down
Loading