Projects
Research projects, datasets, and AI systems for Law and Healthcare
KAMAL Health
2 papersProblem: Healthcare systems in multilingual regions lack interpretable AI tools for clinical decision support and patient-centric care.
Method: Developing responsible AI systems that support clinical decision-making across Arabic and English languages with explainable outputs.
Impact: Improves patient care accessibility and clinical decision quality in multilingual healthcare environments.
BharatGen / BharatGPT
Problem: India needs indigenous generative AI technologies tailored to its diverse linguistic and cultural landscape.
Method: Contributed to BharatGPT, a suite of generative AI technologies for India funded by DST.
Impact: Advances indigenous AI capabilities for Indian languages and domains.
NyayaMind
1 paperProblem: Legal AI systems often lack transparency in their reasoning processes, reducing trust among legal professionals.
Method: Framework for transparent legal reasoning and judgment prediction using structured thinking and explainable AI.
Impact: Enhances trust and interpretability in AI-assisted legal decision-making.
NyayaRAG
1 paperProblem: Existing LJP systems ignore statutory provisions and judicial precedents, core elements of common law reasoning.
Method: RAG framework integrating case facts, statutes, and semantically retrieved precedents for realistic legal judgment prediction.
Impact: Significantly improves predictive accuracy and explanation quality by grounding predictions in external legal knowledge.
TathyaNyaya / FactLegalLlama
1 paperProblem: Prior LJP datasets use complete judgments including reasoning, unlike real-world early-stage decision-making based only on facts.
Method: Created TathyaNyaya dataset focusing on factual statements, and FactLegalLlama, an instruction-tuned LLaMa-3-8B for fact-based prediction and explanation.
Impact: Enables more realistic legal prediction scenarios and improves transparency in AI-assisted legal analysis.
NyayaAnumana / InLegalLLaMA
1 paperProblem: Existing Indian legal datasets lack scale, diversity across court levels, and comprehensive coverage.
Method: Compiled NyayaAnumana (702,945 cases) and developed INLegalLlama through continual pretraining and supervised fine-tuning on Indian legal documents.
Impact: Achieves ~90% F1-score, setting a new benchmark for Indian legal judgment prediction.
LegalSeg
1 paperProblem: Legal documents are unstructured and lengthy, making automatic processing difficult without semantic segmentation.
Method: Created the largest annotated dataset (7,000+ docs, 1.4M sentences) with 7 rhetorical roles and benchmarked multiple SOTA models including RhetoricLLaMA.
Impact: Enables structured legal document understanding and improves downstream tasks like summarization and retrieval.
VidhikDastaavej
1 paperProblem: Generating structured legal documents requires domain-specific formatting and content organization.
Method: Model-agnostic wrapper approach for structured legal document generation in India.
Impact: Automates legal document drafting while remaining adaptable to various underlying generative models.
AILQA
1 paperProblem: There is no comprehensive evaluation framework for AI-driven legal question answering in the Indian context.
Method: Developed AILQA, a multi-faceted evaluation framework for legal QA systems tailored to Indian law.
Impact: Sets benchmarks and identifies challenges for legal QA system development in India.
IBPS
1 paperProblem: Bail prediction is a critical judicial decision that lacks AI-assisted tools in the Indian context.
Method: Developed the Indian Bail Prediction System (IBPS) using machine learning on historical bail data.
Impact: Supports judicial efficiency by providing data-driven insights for bail decision-making.
IndicMedDialog
1 paperProblem: Indic languages lack medical dialogue datasets for building accessible healthcare AI.
Method: Created a parallel multi-turn medical dialogue dataset covering multiple Indic languages.
Impact: Enables development of multilingual medical dialogue systems for underserved language communities.
MedAidDialog
1 paperProblem: Global healthcare AI is limited by lack of diverse multilingual medical conversation data.
Method: Developed a multilingual multi-turn medical dialogue dataset for accessible healthcare.
Impact: Supports creation of inclusive healthcare dialogue agents across languages.