Data deduplication engine, supporting optional compression and public key encryption.
-
Updated
Aug 25, 2022 - Rust
Data deduplication engine, supporting optional compression and public key encryption.
Official Repository of "LLM × DATA" Survey Paper
🚢 Data Toolkit for Sailor Language Models
Self-contained C# library for data deduplication using Sqlite
Fast and efficient content-defined chunking for data deduplication. Java implementation of FastCDC as library.
Optimal distributed data deduplication and supervised learning pipeline using Apache Spark
A JAVA project that splits data using hashing techniques and removes duplicate blocks to save cloud storage. This project also uses the CloudSim framework for cloud storage simulation.
Sievio turns GitHub, local repos, and web PDFs into clean JSONL for LLM pretraining, fine-tuning, and RAG. It offers structure-aware chunking, reliable Unicode decoding, pluggable QC and safety checks, plus optional dataset cards and deduplication.
A pure-JS, content-addressed, copy-on-write virtual filesystem for the browser, featuring: deduplication, filesystem universes (snapshots), events, and optional asynchronous sync.
PolyDeDupe: Multi-Lingual Data Deduplication
Enterprise-grade SaaS platform for importing, cleaning, and managing large-scale mailing lists with advanced deduplication and enrichment.
JSON + CSV to PostgreSQL pipeline with three-level deduplication, validation and idempotent re-runs, three analytical queries, and a /companies page on Next.js App Router
Fellow is a package for creating people that can be unified by their shared values via a singleton list on the class
Этот проект представляет собой мощный инструмент для поиска и анализа дублирующихся файлов в указанной директории. Программа позволяет эффективно выявлять одинаковые файлы на основе их содержимого, используя алгоритм хеширования SHA-256. Она поддерживает настройку параметров, таких как минимальный размер файла для проверки и игнорирование определен
Probabilistic record linkage across sites like Shopify & Stripe — 100% precision, zero false positives.
A calculator for storage and transmission of deduplicated data. Output: charts and tables
A web tool that compares two URL lists, identifies unique and matching domains, and creates a detailed report. It streamlines URL data analysis efficiently.
面向 Excel/CSV 的本地数据研判工具箱,支持多表碰撞、跨表检索、智能提取、模糊去重、差异对比和取最新行,并可打包为 Docker 离线部署包。
Add a description, image, and links to the data-deduplication topic page so that developers can more easily learn about it.
To associate your repository with the data-deduplication topic, visit your repo's landing page and select "manage topics."