indicTranslate v1 - Machine Translation for 11 Indic languages. For latest v2, check: https://github.com/AI4Bharat/IndicTrans2
-
Updated
Jan 2, 2024 - Jupyter Notebook
indicTranslate v1 - Machine Translation for 11 Indic languages. For latest v2, check: https://github.com/AI4Bharat/IndicTrans2
A pipeline for transliteration, spell correction, POS tagging and word sense disambiguation of Hinglish code mixed data to Hindi Devanagari script.
Vyākarana: A Colorless Green Benchmark for Syntactic Evaluation in Indic Languages
Classical Telugu prosody (chandassu) identifier — detects and classifies traditional poetic meters using NLP
A community-maintained catalog of Bangla (Bengali) NLP resources — papers, datasets, models, tools, and benchmarks.
Non-contextual : Word2Vec, FastText Contextual : BERT, RoBERTa, ELECTRA, CamemBERT, Distil-BERT, XLM-RoBERTa Analyzed embedding models, used the best one to build a Flask web app for Hindi NER and data collection from user feedback, deployed on AWS.
Lightweight on-device Hindi TTS for Android & iOS — fine-tuned on AI4Bharat IndicVoices, ONNX export, runs offline on CPU in real-time.
Resource-aware Telugu news summarization using morphology-aware TF-IDF, mT5, adaptive routing, and neural speech synthesis.
Bengali-first, dialect-aware LLM research. Preserving Bengali and its dialects.
Bharat Multimodal EO AI - ISRO Satellite Vision + Indic NLP for Disaster Management, Agriculture & Climate Monitoring
13 executed Jupyter labs, 10 decks and 2 Agent Skills for the Sarvam AI stack. Every API call priced in rupees.
Point your existing ElevenLabs, OpenAI Audio or Deepgram app at Sarvam AI by changing one line. Drop-in compatibility gateway for Indic TTS & STT on Bulbul and Saaras: script-aware language detection, grapheme-safe Hindi/Tamil chunking, CER-based voice mapping, caching + request coalescing. TypeScript, Docker, 168 tests, MIT.
Prove what your model was trained on. Open-source data factory for LLM training corpora: provenance gating, dedup, PII redaction, contamination checks — sealed with a reproducible build manifest. Emits the EU AI Act Article 53(1)(d) training-content summary.
🇮🇳 Google Gemini 2.0 Flash Multimodal Document Intelligence & Indic Vernacular Q&A (Tamil, Hindi, English) with security guardrails & OCR.
Python toolkit to decode legacy Hindi font-encoded PDFs (KrutiDev, Chanakya, DevLys) into Unicode Devanagari. Built for Hindi PDF & govt document ingestion pipelines.
Akshara-aware tokenizer for six Brahmic scripts, with a trained SentencePiece model
A high-performance, community-driven tokenizer for Indian languages. Built for NLP, LLMs, and multilingual AI.
A full-stack ML system for Hindi news classification powered by fine-tuned IndicBERTv2. Features real-time multi-modal inputs: manual text entry, live scraping from major news portals, and OCR-based heading extraction. Deployed via Docker on Hugging Face Spaces.
Multi-agent RAG system for high-quality Telugu story generation — Planner → Drafter → Critic loop
To associate your repository with the indic-nlp topic, visit your repo's landing page and select "manage topics."