Skip to content

Latest commit

 

History

History
176 lines (129 loc) · 7.9 KB

File metadata and controls

176 lines (129 loc) · 7.9 KB

Datamint logo

Datamint Python API

Build Status Python 3.10+

Datamint turns medical imaging ML work. Dataset management, annotation, training, and deployment into a few lines of Python, with built-in support for DICOM/NIfTI/PNG, PyTorch Lightning trainers, and MLflow tracking.

Common use cases: 🩻 Segmentation · 🏷️ Classification · 📦 Detection

Datamint handles the full journey from raw files to a deployed model:

flowchart LR
    Files(["📁 Your Files"])

    subgraph s1 [1 · Ingest]
        Resource(["📦 Resource"])
    end

    subgraph s2 [2 · Organize & Annotate]
        direction TB
        Project(["🗂️ Project"])
        Annotations(["🏷️ Annotations"])
        Project -.->|annotate| Annotations
    end

    subgraph s3 [3 · Train]
        direction LR
        Dataset(["🧮 Dataset"])
        Trainer(["🧠 Trainer"])
        Model(["📈 Model"])
        Dataset -->|train| Trainer -->|register| Model
    end

    subgraph s4 [4 · Deploy & Predict]
        direction LR
        DeployJob(["🚀 Deploy Job"])
        Inference(["🔮 Inference"])
        DeployJob -->|predict| Inference
    end

    Files -->|upload| Resource
    Resource -->|organize| Project
    Project -->|load| Dataset
    Model -->|deploy| DeployJob

    classDef ingestNode fill:#ffffff,stroke:#1f6feb,stroke-width:2px,color:#0b2b4c
    classDef organizeNode fill:#ffffff,stroke:#1a7f37,stroke-width:2px,color:#0b3a1c
    classDef mlNode fill:#ffffff,stroke:#8250df,stroke-width:2px,color:#2c1a4d
    classDef deployNode fill:#ffffff,stroke:#d1720f,stroke-width:2px,color:#4d2b00
    classDef fileNode fill:#f6f8fa,stroke:#57606a,stroke-width:2px,color:#24292f

    class Files fileNode
    class Resource ingestNode
    class Project,Annotations organizeNode
    class Dataset,Trainer,Model mlNode
    class DeployJob,Inference deployNode

    style s1 fill:#dceeff,stroke:#1f6feb,stroke-width:2px,color:#0b2b4c
    style s2 fill:#dbf5df,stroke:#1a7f37,stroke-width:2px,color:#0b3a1c
    style s3 fill:#ecdcff,stroke:#8250df,stroke-width:2px,color:#2c1a4d
    style s4 fill:#ffe8c7,stroke:#d1720f,stroke-width:2px,color:#4d2b00
Loading

📋 Table of Contents

🎬 See it in action

Create a project, split the data, train, and deploy, all through the API:

Datamint pipeline demo

🚀 Features

  • Dataset Management: Download, upload, and manage medical imaging datasets using intuitive object-based APIs or CLI tools
  • Annotation Tools: Create, upload, and manage annotations (segmentations, labels, measurements) with ease
  • Experiment Tracking: Seamless support for experiment management via MLflow integration
  • One-line Trainers: Train segmentation, classification, and detection models with built-in PyTorch Lightning trainers, skipping the dataset class, training loop, and logging setup
  • Model Benchmarking: Compare several trainers against the same dataset and split, and get a ranked leaderboard of their performance
  • DICOM Support: Native handling of DICOM files, including powerful anonymization capabilities during upload to protect patient privacy
  • Multi-format Support: Robust support for a wide range of medical imaging formats: PNG, JPEG, NIfTI (NIfTI/NRRD), DICOMs and more

⚡ Quick Start

1. Install

pip install -U datamint

Using a virtual environment (recommended)

We recommend that you install Datamint in a dedicated virtual environment, to avoid conflicting with your system packages. For instance, create the enviroment once with python3 -m venv datamint-env and then activate it whenever you need it with:

  1. Create the environment (one-time setup):

    python3 -m venv datamint-env
  2. Activate the environment (run whenever you need it):

    Platform Command
    Linux/macOS source datamint-env/bin/activate
    Windows CMD datamint-env\Scripts\activate.bat
    Windows PowerShell datamint-env\Scripts\Activate.ps1
  3. Install the package:

    pip install datamint

2. Configure your API key

datamint config

Follow the prompts (ask your administrator if you don't have a key yet). Environment variable and programmatic options are in the Setup API Key guide.

3. Scaffold a project — the fastest way to start

datamint init

This is the recommended on-ramp: it asks for a project name and task type (segmentation, classification, or detection), then generates a ready-to-run, numbered set of scripts (01_upload_data.py06_deploy.py) — upload data, train, and deploy by running them in order.

4. ...or write it yourself

from datamint import Api
from datamint.lightning import UNetPPTrainer

api = Api()
api.projects.create(name="my-project", exists_ok=True)

trainer = UNetPPTrainer(project="my-project")
results = trainer.fit()

📚 Documentation

Difficulty levels:

  • Beginner no ML knowledge needed
  • Intermediateassumes SDK familiarity, introduces ML/dataset concepts
  • Advanced full training pipelines, custom models, 3D data, multi-step workflows.
Resource Level Description
🚀 Getting Started Beginner Step-by-step setup and basic usage
📖 API Reference Intermediate Complete API documentation
🔥 PyTorch Integration Intermediate ML workflow integration
🧠 Trainer Guide Intermediate Built-in trainers, trainer lifecycle, and custom model integration
🔍 Bringing an External Model into Datamint Intermediate Integrate, log, and deploy a model trained outside Datamint for inference through the UI
🛠️ Command Line Tools Beginner Full reference for datamint upload, datamint init, and datamint config
🔒 SSL Troubleshooting Fixing SSLCertVerificationError
📓 Notebooks Beginner Intermediate Advanced Numbered, runnable tutorials. Start at 01_getting_started and work through annotations, datasets, experiment tracking, deployment, and a full end-to-end example

🆘 Support

Full Documentation
GitHub Issues