Tool for converting error corpora to parallel datasets
-
Updated
Jun 7, 2022 - Python
Tool for converting error corpora to parallel datasets
Dataset of Estonian L2 writings and source code used to train and test machine learning models for CEFR-based classification.
Statistics on some error categories from the REALEC corpus.
Writing assistant
Text Normalization on Learner Texts (South Tyrolean German as a L2)
Add a description, image, and links to the learner-corpus topic page so that developers can more easily learn about it.
To associate your repository with the learner-corpus topic, visit your repo's landing page and select "manage topics."