Skip to content

Latest commit

ย 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿง  Kurd-LLM-Dataset

A foundational Open-Source Dataset for Kurdish NLP, Machine Learning, and Large Language Models (LLMs).

โš ๏ธ This dataset has been superseded. The current, actively maintained version is KurdishCorpus-clean on Hugging Face โ€” 2.7M documents / 2.97B tokens across Kurmancรฎ, Soranรฎ, and Zazakรฎ, deduplicated and quality-filtered, under clear open licenses. Use that one for new work; this repo is kept for historical reference.

๐Ÿ“Œ Overview

The Kurd-LLM-Dataset is a massive, meticulously compiled linguistic repository created by the [Kurdish-Tech] organization. Designed to bridge the gap in low-resource language modeling, this dataset serves as the core infrastructure for developing advanced AI applications, text processing tools, and translation engines for the Kurdish language.

๐Ÿ“Š Dataset Specifications

  • Total Records: 500,000+ words, terms, and definitions.
  • Size: ~166 MB (Raw Text/Data).
  • Dialects Covered: Kurmancรฎ & Soranรฎ.
  • Origin: Fully compiled and engineered from the massive wq ferheng database, formatted for machine-readability.

๐Ÿš€ Purpose & Use Cases

This dataset is published with a mission to empower the Kurdish language in the digital and AI era. It is highly optimized for:

  1. Training Large Language Models (LLMs): Fine-tuning foundation models (like LLaMA, GPT, etc.) to understand and generate Kurdish text natively.
  2. Natural Language Processing (NLP): Sentiment analysis, tokenization, stemming, and syntax parsing.
  3. Machine Translation: Building high-accuracy translation matrices between Kurdish dialects and global languages.
  4. Speech Recognition (STT/TTS): Providing the textual baseline for phonetic mapping and voice AI training.

๐Ÿค Contribution & Organization

This project is maintained by Kurdish-Tech. We believe in the power of open data to preserve culture and accelerate technological adoption. Researchers, data scientists, and developers are encouraged to fork, analyze, and build upon this dataset.

๐Ÿ“„ License

Licensed under CC BY-SA 4.0 โ€” matching the license used by its successor, KurdishCorpus-clean. You're free to share and adapt this data, including commercially, as long as you give attribution and license any derivative under the same terms.


Built with โค๏ธ for the Kurdish Language & Tech Community.

About

A foundational Open-Source Dataset for Kurdish NLP, Machine Learning, and Large Language Models (LLMs).

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors