RLHF preference data curation pipeline: HH-RLHF + UltraFeedback + OASST1 → quality filter → MinHash dedup → DPO-ready JSONL
-
Updated
Jun 7, 2026 - Python
RLHF preference data curation pipeline: HH-RLHF + UltraFeedback + OASST1 → quality filter → MinHash dedup → DPO-ready JSONL
Measuring verbosity bias in UltraFeedback (61K GPT-4-annotated preference pairs): statistically unambiguous, but a small effect. Python + Streamlit.
To associate your repository with the ultrafeedback topic, visit your repo's landing page and select "manage topics."