Official implementation of DispViT, published at ICLR 2026.
DispViT ViT-G is currently the only released checkpoint.
| Model card | #Params |
|---|---|
| aeolusguan/dispvit-vitg | 1.26B |
The checkpoint includes the model weights and architecture configuration,
which DispViT.from_pretrained() loads automatically.
- Release a ViT-G model self-distilled on real-world data.
- Release the full model with NMRF refinement.
- Release checkpoints for smaller model variants.
Use Python 3.10 or newer. Clone the repository and enter its root directory:
git clone https://github.com/aeolusguan/DispViT.git
cd DispViTInstall PyTorch 2.0 or newer and a matching torchvision build for your CUDA environment using the PyTorch installation instructions. Then install the dependencies needed for evaluation:
python -m pip install 'hydra-core>=1.3,<1.4' huggingface-hub numpy Pillow opencv-python imageio iopath wcmatchThe default evaluation uses BF16 automatic mixed precision on a CUDA GPU.
Load the released checkpoint directly from Hugging Face:
from dispvit.models.dispvit import DispViT
model = DispViT.from_pretrained("aeolusguan/dispvit-vitg").cuda().eval()A downloaded checkpoint can be loaded with the same API:
model = DispViT.from_pretrained("/path/to/model_final.pt").cuda().eval()Place the evaluation datasets under datasets/ in the repository root.
The evaluation configuration covers KITTI 2012,
KITTI 2015, ETH3D, and Middlebury. Run the following command from the repository root:
torchrun --standalone --nproc_per_node=1 launch.py --config-name zero_shot_eval \
pretrained=aeolusguan/dispvit-vitgTo use a local checkpoint:
torchrun --standalone --nproc_per_node=1 launch.py --config-name zero_shot_eval \
pretrained=/path/to/model_final.ptEvaluation metrics are printed to the terminal.
All error rates are reported as percentages; lower is better.
| Model | KITTI 2012 (D1 ↓) | KITTI 2015 (D1 ↓) | ETH3D (BP-1 ↓) | Middlebury (BP-2 ↓) |
|---|---|---|---|---|
| dispvit-vitg | 3.03 | 3.30 | 1.00 | 3.82 |
This repository includes code adapted from VGGT. We thank the authors for making their code publicly available.
If you find DispViT useful in your research, please cite our ICLR 2026 paper:
@inproceedings{guan2026dispvit,
title = {{DispViT}: Direct Stereo Disparity Regression with a Single-Stream Vision Transformer},
author = {Guan, Tongfan and Guo, Jiaxin and Huang, Tianyu and Dong, Jinhu and Wang, Chen and Liu, Yun-Hui},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://openreview.net/forum?id=c21yqwf02V}
}