Text in the Dark dataset. The first extremely low light text dataset based on SID and LOL datasets with a total of 3,075 images and 59,975 texts annotated.
-
Updated
Feb 24, 2026 - Python
Text in the Dark dataset. The first extremely low light text dataset based on SID and LOL datasets with a total of 3,075 images and 59,975 texts annotated.
Open-source Arabic NLP datasets by GDG KSU: collecting, cleaning, and publishing high-quality Arabic and dialectal text data, powered by a modular spaCy pipeline that outputs ML-ready Parquet files for Hugging Face and GitHub.
Persian News Dataset
nepali websites data extractor
To associate your repository with the text-dataset topic, visit your repo's landing page and select "manage topics."