Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
-
Updated
Aug 21, 2026 - Python
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
文本挖掘和预处理工具(文本清洗、新词发现、情感分析、实体识别链接、关键词抽取、知识抽取、句法分析等),无监督或弱监督方法
🧹 Python package for text cleaning
A Python toolkit for file processing, text cleaning and data splitting. 文件处理,文本清洗和数据划分的python工具包。
NLP预/后处理工具。
Text preprocessing tools in python.
Dataiku DSS plugin to detect languages, correct misspellings, and clean text data 🧼
Korean text data preprocess toolkit for NLP
A Python package to get useful information from documents using TopicRank Algorithm.
Text preprocessing package for use in NLP tasks https://pypi.org/project/textcl/
Portfolio-grade audit of a student mental health & academic pressure survey. Measures coverage and sample imbalance, runs validity checks, highlights measurement and selection bias risks, and converts messy open-text “stress causes” into a transparent taxonomy. Ships a Markdown report, figures, and a Streamlit dashboard.
Text preprocessing and PII anonymisation for NLP/ML. ONNX NER ensemble, language detection, stopword removal. Built for statistical ML and language models.
Remove AI text artifacts, hidden unicode characters and file metadata before you publish. Deterministic, offline, zero dependencies. Python and Node CLIs with one shared rule set.
Corpora and scripts for cleaning political science texts. Scripts are translated into transformations that support SAGE Texti.
Common Text Pre-Processing for Portuguese
ValX is an open-source Python package for text cleaning tasks, including profanity detection and removal. Now also includes sensitive information detection, and removal.
A Simple Easy To Use Text Cleaning Package For NLP Built In Python. It Can Clean and Analyze Your Text Data In One Line of Code.
TextLasso is a Simple Python library for extracting structured data from raw text, with special focus on processing LLM (Large Language Model) responses.
Structured EVTAL pipeline for extracting, cleaning, and analyzing HTML web data with spaCy-based NLP features.
A lightweight Python library for cleaning raw text and documents before NLP and LLM workflows.
Add a description, image, and links to the text-cleaning topic page so that developers can more easily learn about it.
To associate your repository with the text-cleaning topic, visit your repo's landing page and select "manage topics."