Skip to content
#

html-extraction

Here are 16 public repositories matching this topic...

A Java-based server leveraging Apache Tika to extract content and metadata from files (PDF, DOCX, TXT, etc.) in a local files-to-extract directory. Supports HTML (with CSS styling) and text extraction, file listing, and metadata retrieval via MCP-compliant tools and REST APIs. Built with Spring Boot, Jetty, and MCP SDK.

  • Updated Aug 30, 2025
  • Java

Improve this page

Add a description, image, and links to the html-extraction topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the html-extraction topic, visit your repo's landing page and select "manage topics."

Learn more