Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
-
Updated
Jan 13, 2026 - Python
Cleanlab's open-source library is the standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
[SIGMOD'27] Easy Data Preparation with latest LLMs-based Operators and Pipelines.
A light-weight, flexible, and expressive statistical data testing library
Machine learning with dataframes
🚚 Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
Continuously updated paper list on advancements in Data Agents. Companion repo to our paper "A Survey of Data Agents: Emerging Paradigm or Overstated Hype?"
Easy to use Python library of customized functions for cleaning and analyzing data.
The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.
🗣️ A book and repo to get you started programming voice computing applications in Python (10 chapters and 200+ scripts).
Deal with bad samples in your dataset dynamically, use Transforms as Filters, and more!
LLM-based text extraction from unstructured data like PDFs, Words and HTMLs. Transform and cluster the text into your desired format. Less information loss, more interpretation, and faster R&D!
Pydantic extension for annotating autocorrecting fields.
🗺️ Data Cleaning and Textual Data Visualization 🗺️
🤖 An automated machine learning framework for audio, text, image, video, or .CSV files (50+ featurizers and 15+ model trainers). Python 3.6 required.
pyDVL is a library of stable implementations of algorithms for data valuation and influence function computation
A full data warehouse infrastructure with ETL pipelines running inside docker on Apache Airflow for data orchestration, AWS Redshift for cloud data warehouse and Metabase to serve the needs of data visualizations such as analytical dashboards.
Zero-config entity resolution feeding a durable identity layer: messy records from any source become stable golden entities, a Customer 360 with provenance, merge/split and audit. Fellegi-Sunter beats hand-tuned Splink. Arrow-native/Rust, 250M rows in 11.2 min. Python + edge TypeScript (WASM), SQL-native in Postgres & DuckDB, 97 MCP tools + REST.
OpenDataVal: a Unified Benchmark for Data Valuation in Python (NeurIPS 2023)
Data trust for Python. Validate, clean, and profile DataFrames before everything else.
🚢 Data Toolkit for Sailor Language Models
To associate your repository with the data-cleaning topic, visit your repo's landing page and select "manage topics."