Dust library for html processing
-
Updated
Jul 18, 2026 - Java
Dust library for html processing
⚡ 2.3x faster file scraping for Java — powered by native Win32 JNI (FastGLOB). Directory trees, content extraction & chunking for LLMs, agents and code analysis. Parallel multi-threaded. Part of the FastJava ecosystem.
Configurable and schedulable web scrapping tool. Used to extract raw article content and metadata for aggregated news feeds.
Mirror Discord server data into SQLite for fast, local searching and analysis of guild messages, members, and history without relying on Discord's search.
Add a description, image, and links to the content-extraction topic page so that developers can more easily learn about it.
To associate your repository with the content-extraction topic, visit your repo's landing page and select "manage topics."