A foundational Open-Source Dataset for Kurdish NLP, Machine Learning, and Large Language Models (LLMs).
โ ๏ธ This dataset has been superseded. The current, actively maintained version is KurdishCorpus-clean on Hugging Face โ 2.7M documents / 2.97B tokens across Kurmancรฎ, Soranรฎ, and Zazakรฎ, deduplicated and quality-filtered, under clear open licenses. Use that one for new work; this repo is kept for historical reference.
The Kurd-LLM-Dataset is a massive, meticulously compiled linguistic repository created by the [Kurdish-Tech] organization. Designed to bridge the gap in low-resource language modeling, this dataset serves as the core infrastructure for developing advanced AI applications, text processing tools, and translation engines for the Kurdish language.
- Total Records: 500,000+ words, terms, and definitions.
- Size: ~166 MB (Raw Text/Data).
- Dialects Covered: Kurmancรฎ & Soranรฎ.
- Origin: Fully compiled and engineered from the massive wq ferheng database, formatted for machine-readability.
This dataset is published with a mission to empower the Kurdish language in the digital and AI era. It is highly optimized for:
- Training Large Language Models (LLMs): Fine-tuning foundation models (like LLaMA, GPT, etc.) to understand and generate Kurdish text natively.
- Natural Language Processing (NLP): Sentiment analysis, tokenization, stemming, and syntax parsing.
- Machine Translation: Building high-accuracy translation matrices between Kurdish dialects and global languages.
- Speech Recognition (STT/TTS): Providing the textual baseline for phonetic mapping and voice AI training.
This project is maintained by Kurdish-Tech. We believe in the power of open data to preserve culture and accelerate technological adoption. Researchers, data scientists, and developers are encouraged to fork, analyze, and build upon this dataset.
Licensed under CC BY-SA 4.0 โ matching the license used by its successor, KurdishCorpus-clean. You're free to share and adapt this data, including commercially, as long as you give attribution and license any derivative under the same terms.
Built with โค๏ธ for the Kurdish Language & Tech Community.