The current profanity dataset is sourced from multiple language files and contains ~55K+ entries.
We need to improve the quality of this dataset by:
Tasks:
- Identify and remove duplicate entries across languages
- Detect and eliminate false positives (valid words incorrectly flagged)
- Normalize formatting (case sensitivity, spacing, special characters)
- Ensure consistency across all language files
Optional:
- Add a script/tool to automate dataset validation
This will significantly improve filtering accuracy and reduce unintended masking.
The current profanity dataset is sourced from multiple language files and contains ~55K+ entries.
We need to improve the quality of this dataset by:
Tasks:
Optional:
This will significantly improve filtering accuracy and reduce unintended masking.