Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

vision-label-skill

License: MIT Claude Code Skill Python 3 Formats

Claude Code skill for batch object detection / instance segmentation labeling via any multimodal vision API. Export YOLO · YOLO-seg · Labelme · VOC · COCO · CSV under <images_dir>/Labels/.

用外部多模态 API 给图片文件夹批量自动标注,导出 YOLO / Labelme / VOC / COCO / CSV。类别完全自定义(螺丝、标签、商标、纸卷、缺陷……)。


Features

  • Any classes you define (name + optional description + example images / URLs)
  • bbox or polygon (SAM-style multipoint outlines)
  • API styles: OpenAI Chat Completions · OpenAI Responses · Anthropic Messages (and compatible gateways)
  • Canonical model output: JSON array in 0–1000 coordinate space → exporters convert to training formats
  • Re-export from .raw.json without calling the API again (if you keep raw files)
  • Stdlib + Pillow only (no official OpenAI/Anthropic SDKs required)

Install (Claude Code)

# User skills directory (example)
git clone https://github.com/Autsunset/vision-label-skill.git \
  ~/.claude/skills/vision-label-skill

cd ~/.claude/skills/vision-label-skill
pip install Pillow
cp env/.env.example env/.env
# edit env/.env with your API key / base / model / format

Windows (PowerShell example):

git clone https://github.com/Autsunset/vision-label-skill.git "$env:USERPROFILE\.claude\skills\vision-label-skill"
cd "$env:USERPROFILE\.claude\skills\vision-label-skill"
pip install Pillow
copy env\.env.example env\.env

Layout:

vision-label-skill/
├── SKILL.md                 # Claude Code skill instructions
├── LICENSE
├── env/
│   └── .env.example
├── scripts/
│   ├── vision_client.py
│   ├── label_batch.py
│   ├── class_spec.py
│   └── export_labels.py
├── references/
│   ├── formats.md
│   ├── prompts.md
│   └── class_spec.md
└── evals/
    └── evals.json

Configure API

Edit env/.env (never commit this file):

VISION_API_KEY=sk-...
VISION_API_BASE=https://api.openai.com/v1
VISION_MODEL=gpt-4o
# openai_chat | openai_responses | anthropic
VISION_API_FORMAT=openai_chat
VISION_MAX_TOKENS=8192
VISION_API_FORMAT Request path
openai_chat {base}/chat/completions
openai_responses {base}/responses
anthropic {base}/v1/messages or {base}/messages

Also supported: process env VISION_*, or ~/.claude/vision-config.json.

python scripts/vision_client.py --config-check

Use with Claude Code

Say for example:

Label ./images as YOLO boxes; classes screw, nut, washer

The skill will ask (if needed) for format, bbox vs polygon, class names / descriptions, and the image folder. Default folder preset is ./images under the current working directory (not datasets).

After a batch finishes, it asks whether to keep or delete Labels/*.raw.json.

CLI

# Recommended: class_spec with descriptions / examples
python scripts/label_batch.py \
  --images-dir "./images" \
  --format yolo \
  --mode bbox \
  --class-spec "./images/Labels/class_spec.json" \
  --skip-existing

# Names only
python scripts/label_batch.py \
  --images-dir "./images" \
  --format yolo \
  --mode bbox \
  --classes "screw,nut,washer" \
  --skip-existing

# Dry-run (count + prompt only)
python scripts/label_batch.py \
  --images-dir "./images" \
  --format yolo \
  --mode bbox \
  --classes "screw,nut,washer" \
  --dry-run

Re-export from existing .raw.json without another API call:

python scripts/export_labels.py \
  --image "./images/001.jpg" \
  --ann "./images/Labels/001.raw.json" \
  --format labelme \
  --mode bbox \
  --classes "screw,nut,washer"

Output

Path Content
Labels/<stem>.raw.json Canonical 0–1000 JSON (optional keep)
Labels/<stem>.txt YOLO / YOLO-seg lines
Labels/classes.txt + data.yaml Class id map
Labels/<stem>.json Labelme (if format=labelme)
Labels/session_meta.json Run summary

Details: references/formats.md.

Example class_spec.json

{
  "classes": [
    {
      "name": "Label",
      "description": "White logistics shipping label on cartons, with barcode",
      "examples": ["./refs/label1.jpg", "https://example.com/label2.png"]
    },
    {
      "name": "Logo",
      "description": "Printed brand trademark on packaging",
      "examples": []
    }
  ]
}

Security & open-source notes

  • Secrets only in env/.env (gitignored). Never put keys in SKILL.md or commits.
  • Docs/examples use relative paths like ./images/....
  • Local images/ and Labels/ run outputs are gitignored — do not publish private photos or credentials.
  • Spot-check labels; vision models can miss or invent boxes.

Acknowledgements

本项目分享于 Linux.do 社区。

This project is shared with the Linux.do community.

License

MIT — use freely with your own datasets and API keys.

About

batch auto-label images with any vision API and export YOLO, Labelme, VOC, COCO, or CSV.基于多模态视觉模型的目标检测 / 实例分割自动标注skill。自定义类别、矩形框或多边形、支持示例图;导出 YOLO、YOLO-seg、Labelme、VOC、COCO、CSV。

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages