Skip to content
#

multi-modal-ai

Here are 14 public repositories matching this topic...

RAPTOR (Robust AI-Powered Toolkit for Operational Robots) is an AI-native Content Insight Engine that transforms passive media storage into an intelligent knowledge platform through automated analysis, semantic search, and actionable insights. RAPTOR reducing manual tagging by 85% and making content discovery 10x faster.

  • Updated Jul 18, 2026
  • Python

Sales Forge is a high performance, real time voice interaction platform designed to train sales representatives through adaptive AI personas. It provides a low latency, immersive roleplay experience that simulates real world sales challenges.

  • Updated May 6, 2026
  • Python

为PDF格式的电子书生成目录。通过多模态AI实现,自动生成有层次的目录书签。专为扫描版/纯图片PDF设计。 Generate a table of contents for PDF ebooks via multimodal AI. Auto-produces hierarchical bookmarks. Designed for scanned / image-only PDFs.

  • Updated Jul 18, 2026
  • Python

🔍 Enhance reasoning accuracy with the Reflective Reasoning Transformer, leveraging causal reasoning graphs for better dynamic reasoning performance.

  • Updated Jul 27, 2026
  • Python

PolySensor is an AI-powered agentic multi-modal content analysis tool that analyzes textual information from different 108 file formats or more. Using Google Gemini and advanced RAG techniques, it transforms documents, images, audio, and video into actionable insights through intelligent summarization and contextual understanding.

  • Updated Nov 3, 2025
  • Python

A Multi-Modal RAG Knowledge Engine An intelligent knowledge graph system that ingests video, audio, and PDF documents to create a connected semantic web. Features a graph-based retrieval engine (GraphRAG), multi-modal search, and an interactive React Flow visualization dashboard. Built with FastAPI, Next.js, Neo4j, and LlamaIndex.

  • Updated Jul 23, 2026
  • Python

Color-based semantic routing for Apache Kafka - Tag events with RGB hex codes for flexible consumer-side filtering. Eliminates topic proliferation and enables dynamic routing without payload deserialization. Python reference implementation with validated 5x speedup over content-based routing.

  • Updated Jun 19, 2026
  • Python

An intelligent travel planning platform powered by GPT-4 and DALL-E 3 that generates personalized, optimized itineraries with route optimization, budget allocation, and AI-generated visual content through advanced prompt engineering and multi-modal AI integration.

  • Updated Feb 3, 2026
  • Python

Text-Vision-Agent is an AI-powered assistant that generates images from text descriptions and provides detailed image descriptions. It combines image generation using FluxPipeline with vision-based language models like ChatOllama, enabling seamless text-to-image and image interpretation interactions.

  • Updated Feb 16, 2025
  • Python

Improve this page

Add a description, image, and links to the multi-modal-ai topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the multi-modal-ai topic, visit your repo's landing page and select "manage topics."

Learn more