Projects

Research projects, datasets, and AI systems for Law and Healthcare

2 papers

Problem: Healthcare systems in multilingual regions lack interpretable AI tools for clinical decision support and patient-centric care.

Method: Developing responsible AI systems that support clinical decision-making across Arabic and English languages with explainable outputs.

Impact: Improves patient care accessibility and clinical decision quality in multilingual healthcare environments.

Healthcare AIMultilingual AIExplainable AICurrent

Problem: India needs indigenous generative AI technologies tailored to its diverse linguistic and cultural landscape.

Method: Contributed to BharatGPT, a suite of generative AI technologies for India funded by DST.

Impact: Advances indigenous AI capabilities for Indian languages and domains.

Generative AIMultilingual AIIndia
1 paper

Problem: Legal AI systems often lack transparency in their reasoning processes, reducing trust among legal professionals.

Method: Framework for transparent legal reasoning and judgment prediction using structured thinking and explainable AI.

Impact: Enhances trust and interpretability in AI-assisted legal decision-making.

Legal AIExplainable AIJudgment Prediction
1 paper

Problem: Existing LJP systems ignore statutory provisions and judicial precedents, core elements of common law reasoning.

Method: RAG framework integrating case facts, statutes, and semantically retrieved precedents for realistic legal judgment prediction.

Impact: Significantly improves predictive accuracy and explanation quality by grounding predictions in external legal knowledge.

Legal AIRAGJudgment PredictionRetrieval
1 paper

Problem: Prior LJP datasets use complete judgments including reasoning, unlike real-world early-stage decision-making based only on facts.

Method: Created TathyaNyaya dataset focusing on factual statements, and FactLegalLlama, an instruction-tuned LLaMa-3-8B for fact-based prediction and explanation.

Impact: Enables more realistic legal prediction scenarios and improves transparency in AI-assisted legal analysis.

Legal AIFact-based PredictionLLMs
1 paper

Problem: Existing Indian legal datasets lack scale, diversity across court levels, and comprehensive coverage.

Method: Compiled NyayaAnumana (702,945 cases) and developed INLegalLlama through continual pretraining and supervised fine-tuning on Indian legal documents.

Impact: Achieves ~90% F1-score, setting a new benchmark for Indian legal judgment prediction.

Legal AIDatasetLLMsJudgment Prediction
1 paper

Problem: Legal documents are unstructured and lengthy, making automatic processing difficult without semantic segmentation.

Method: Created the largest annotated dataset (7,000+ docs, 1.4M sentences) with 7 rhetorical roles and benchmarked multiple SOTA models including RhetoricLLaMA.

Impact: Enables structured legal document understanding and improves downstream tasks like summarization and retrieval.

Legal AISegmentationRhetorical RolesDataset
1 paper

Problem: Generating structured legal documents requires domain-specific formatting and content organization.

Method: Model-agnostic wrapper approach for structured legal document generation in India.

Impact: Automates legal document drafting while remaining adaptable to various underlying generative models.

Legal AIDocument GenerationNLP
1 paper

Problem: There is no comprehensive evaluation framework for AI-driven legal question answering in the Indian context.

Method: Developed AILQA, a multi-faceted evaluation framework for legal QA systems tailored to Indian law.

Impact: Sets benchmarks and identifies challenges for legal QA system development in India.

Legal AIQuestion AnsweringEvaluation
1 paper

Problem: Bail prediction is a critical judicial decision that lacks AI-assisted tools in the Indian context.

Method: Developed the Indian Bail Prediction System (IBPS) using machine learning on historical bail data.

Impact: Supports judicial efficiency by providing data-driven insights for bail decision-making.

Legal AIBail PredictionJudicial Decision Support
1 paper

Problem: Indic languages lack medical dialogue datasets for building accessible healthcare AI.

Method: Created a parallel multi-turn medical dialogue dataset covering multiple Indic languages.

Impact: Enables development of multilingual medical dialogue systems for underserved language communities.

Healthcare AIMultilingual AIDatasetIndic Languages
1 paper

Problem: Global healthcare AI is limited by lack of diverse multilingual medical conversation data.

Method: Developed a multilingual multi-turn medical dialogue dataset for accessible healthcare.

Impact: Supports creation of inclusive healthcare dialogue agents across languages.

Healthcare AIMultilingual AIDataset