Official implementation of the research thesis: "Scribe Verification in Chinese Manuscripts: From Siamese Baselines to Co-Tuplet Vision Transformers."
This repository provides a modular and reproducible framework for authorial authenticity verification in historical Chinese manuscripts. By transitioning from isolated pairwise comparisons to a Co-Tuplet metric learning framework, this project successfully unlocks the global feature extraction capabilities of Vision Transformers (ViT) for paleographic analysis.
This research builds upon and extends the following foundational works. If you utilize this codebase, please cite the primary thesis and these core references:
- Co-Tuplet Loss: The multi-target repulsion logic is adapted from the official MS-SigNet repository by Huang & Lu (2023).
- Siamese/Triplet Baseline: This framework extends the original scribe verification pipeline developed by Liakopoulos et al. (2026).
- Tsinghua Dataset: Based on the manuscript analysis methodologies established by Wang et al. (2026).
Full BibTeX entries are available in the CITATIONS.bib file.
Beyond the baseline Siamese/Triplet models, this repository introduces:
- Co-Tuplet Implementation: Adapted for historical document feature extraction.
- Multi-Target Loaders: Development of specialized Tuplet dataset loaders and batch collators.
- ViT Optimization: Custom model implementations compatible with dense relational losses.
- Uncertainty Quantization: Extensions to the conformal prediction module to handle the geometric gap regions produced by Tuplet distributions.
ScribeVerification_Tuplet/
├── configs/ # YAML configuration files
├── scripts/ # Data prep: check_images, make_test_pairs, make_calib
├── src/ # Core modular logic
│ ├── dataset_tuplets.py <-- Added for Co-Tuplet
│ ├── losses.py <-- Extended with Co-Tuplet Loss
│ ├── models/ <-- ViT and CNN implementations
│ └── metrics.py <-- ROC, AUC, FAR, FRR computation
├── train_tuplet.py # Entry point for Co-Tuplet training
├── evaluate_tuplet_perclass_full.py # Full diagnostic evaluation
└── requirements.txt
This framework requires Python 3.10+ and PyTorch 2.5.1.
# Create and activate virtual environment
python -m venv .venv
source .venv/bin/activate # Or .venv\Scripts\activate on Windows
# Install core dependencies
pip install -r requirements.txt
# For GPU execution with CUDA 12.1
pip install torch==2.5.1+cu121 torchvision==0.20.1+cu121 --index-url [https://download.pytorch.org/whl/cu121](https://download.pytorch.org/whl/cu121)
Datasets must be organized in a class-based folder structure.
data/
└── Tsinghua/
├── train/
│ ├── Scribe_A/
│ └── Scribe_B/
└── test/
├── Scribe_A/
└── Scribe_B/
Generate fixed test pairs for deterministic evaluation:
python scripts/make_test_pairs.py
Train the proposed Co-Tuplet Vision Transformer:
python train_tuplet.py --config configs/default.yaml
Run the full-class diagnostic evaluation to generate ROC curves and per-class statistics:
python evaluate_tuplet_perclass_full.py \
--model_module src.models.vit_tuplet \
--embedding_dim 10 \
--data_root ./data/Tsinghua/test \
--pairs_csv ./data/Tsinghua/test_pairs_tsinghua.csv \
--ckpt ./checkpoints/vit_tsinghua_tuplet/vit_tsinghua_tuplet_e30.pt \
--out_dir ./vit_tsinghua_tuplet_eval_outputs
First, generate a disjoint calibration set:
python scripts/make_calib_pairs_disjoint.py --train_root ./data/Tsinghua/train --out ./data/Tsinghua/calib_pairs.csv
Run the evaluation with the --epsilon confidence flag:
python evaluate_tuplet_perclass_full.py \
--ckpt ./checkpoints/vit_tsinghua_tuplet_e30.pt \
--calib_pairs_csv ./data/Tsinghua/calib_pairs.csv \
--epsilon 0.05
The pipeline dynamically generates:
- Per-Class ROC Curves and one-vs-rest confusion matrices.
- Distance Histograms mapping the geometric separation of scribes.
- Conformal Prediction Sets (Confident, Ambiguous, or Empty).
This project is released for academic research purposes. Please ensure appropriate attribution to the authors of the foundational papers cited herein.