Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
-
Updated
May 13, 2026 - Python
Multi-modal OCR pipeline optimized for ML training (text, figure, math, tables, diagrams)
Dafonts Free Dataset and python scripts used to make it
simLIBS provides Python class to simulate LIBS spectra with NIST LIBS Database interface.
Library to programmatically build labeled datasets for Named-Entity Recognition (NER) and Relation Extraction (RE) Machine Learning tasks
Systematic quality evaluation suite for AI/ML datasets. 103 ego datasets audited. ISO 5259-2 aligned.
XFetcher is Python Library for downloading data from different sources for machine learning purposes
Python JSON ORM (simple module all in one file)
API for showing the crawled wiki movies content and list of movies
Image classification subset of Google's Open Images V7 dataset with 100 classes. Inspired by ImageNet100.
Privacy Act 2020 compliance auditor for ML datasets — scans CSV / Parquet / HuggingFace datasets and flags New Zealand-specific PII (IRD, NHI, driver licence, phone, address, te reo names) using a hybrid regex + NER + LLM verification pipeline.
Генерирует изображения (100x100) с случайными символами (от 1 до 5), создав изображения со всеми видами шрифтов (15)
🫙 Event datasets used for training machine learning models.
Add a description, image, and links to the ml-datasets topic page so that developers can more easily learn about it.
To associate your repository with the ml-datasets topic, visit your repo's landing page and select "manage topics."