A curated list of papers and resources for language models (LLMs) for Low-Resourced Languages in South Asia
Despite the rise of large language models (LLMs), South Asian languages still struggle with limited resources, code-mixing, and diversity, making it necessary to develop inclusive resources for these languages. This repository brings together models, adaptation strategies, and resources on bias and evaluation to help bridge the gap.
We aim to record and provide a structured overview of existing resources to promote future research work for improving NLP technologies in South Asia.
This repository is organized into three main sections:
- LLMs and Model Resources
- Adaptation and Fine-Tuning Techniques
- Bias and Evaluation in South Asian NLP
- (IndicBART) - IndicBART: A Pre-trained Model for Indic Natural Language Generation [paper]
- (AxomiyaBERTa) - AxomiyaBERTa: A Phonologically-aware Transformer Model for Assamese [paper]
- (IndicBERT) - IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages. [paper]
- (IndicBERTv2) - Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages [paper]
- (IndicTrans) - Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages [paper]
- (IndicTrans2) - IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages [paper]
- (Indic-ColBERT) - IndicIRSuite: Multilingual dataset and neural information models for Indian languages. [paper]
- (RelateLM) - Exploiting language relatedness for low web-resource language model adaptation: An indic languages study [paper]
- (HindiLLM-Small, HindiLLM-Medium) - HindiLLM: Large Language Model for Hindi [paper]
- (Tri-Distil-BERT, Mixed-Distil-BERT) - Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi [paper]
- (TigerLLM) - TigerLLM - A Family of Bangla Large Language Models [paper]
- (MahaBERT) - L3Cube-MahaCorpus and MahaBERT: Marathi Monolingual Corpus, Marathi BERT Language Models, and Resources [paper]
- (CPT-R, IFT-R) - RomanSetu: Efficiently un-locking multilingual capabilities of large language models via Romanization [paper]
- (BUQRNN, PN-BUQRNN) - Application of Quantum Recurrent Neural Network in Low-Resource Language Text Classification [paper]
- (DC-LM) - The best of both worlds: Dual Channel LM for Hope Speech Detection in low-resourced Kannada [paper]
- (Nepali DistilBERT, Nepali DeBERTa) - Nepali Encoder Transformers: An Analysis of Auto Encoding Transformer Language Models for Nepali Text Classification [paper]
- (EmojiPredictor) - Predicting multi-label emojis, emotions, and sentiments in code-mixed texts using an emojifying sentiments framework [paper]
- (Hindi-WMT) - Multimodal Machine Translation for Low-Resource Indic Languages: A Chain-of-Thought Approach Using Large Language Models [paper]
- (MAPE model) - Together We Can: Multilingual Automatic Post-Editing for Low-Resource Languages [paper]
- (IndIE) - IndIE: A Multilingual Open Information Extraction Tool For Indic Languages [paper]
- (LambaniNMT) - Machine translation for a very low-resource language – layer freezing approach on transfer learning [paper]
- (BASE, MED, RETRAIN) - Abstractive Hindi text summarization: A challenge in a low-resource setting [paper]
- (MedSumm) - MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries [paper]
- (AI-Tutor) - AI-Tutor: Interactive Learning of Ancient Knowledge from Low-Resource Languages [paper]
- (TPPoet) - TPPoet: Transformer-Based Persian Poem Generation using Minimal Data and Advanced Decoding Techniques [paper]
- Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi [paper]
- MedSumm: A Multimodal Approach to Summarizing Code-Mixed Hindi-English Clinical Queries [paper]
- Predicting multi-label emojis, emotions, and sentiments in code-mixed texts using an emojifying sentiments framework [paper]
- IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages [paper]
- Does Transliteration Help Multilingual Language Modeling? [paper]
- Context-aware Transliteration of Romanized South Asian Languages [paper]
- Multimodal Machine Translation for Low-Resource Indic Languages: A Chain-of-Thought Approach Using Large Language Models [paper]
- Neural Methods for Aligning Large-Scale Parallel Corpora from the Web for South and East Asian Languages [paper]
- Generalists vs. Specialists: Evaluating Large Language Models for Urdu [paper]
- Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages [paper]
- IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages [paper]
- AxomiyaBERTa: A Phonologically-aware Transformer Model for Assamese [paper]
- Overlap-based Vocabulary Generation Improves Cross-lingual Transfer Among Related Languages [paper]
- IndicBART: A Pre-trained Model for Indic Natural Language Generation [paper]
- RomanSetu: Efficiently un-locking multilingual capabilities of large language models via Romanization [paper]
- Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings [paper]
- Enhancing Low Resource NER Using Assisting Language And Transfer Learning [paper]
- Improving NER Tagging Performance in Low-Resource Languages via Multilingual Learning [paper]
- Improving Pretraining Techniques for Code-Switched NLP [paper]
- Cultural Compass: Predicting Transfer Learning Success in Offensive Language Detection with Cultural Features [paper]
- A Tulu Resource for Machine Translation [paper]
- LLMs Are Few-Shot In-Context Low-Resource Language Learners [paper]
- Bidirectional English-Nepali Machine Translation(MT) System for Legal Domain [paper]
- Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi [paper]
- Nepali Encoder Transformers: An Analysis of Auto Encoding Transformer Language Models for Nepali Text Classification [paper]
- Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters [paper]
- Cross-Lingual Named Entity Recognition for Low-Resource Languages: A Hindi-Nepali Case Study Using Multilingual BERT Models [paper]
- BenLLM-Eval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP [paper]
- Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation [paper]
- Large Language Models as a Normalizer for Transliteration and Dialectal Translation [paper]
- Abstractive Summarization of Low resourced Nepali language using Multilingual Transformers [paper]
- Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs [paper]
- IndiText Boost: Text Augmentation for Low Resource India Languages [paper]
- Table Question Answering for Low-resourced Indic Languages [paper]
- Language Model Tokenizers Introduce Unfairness Between Languages [paper]
- Re-contextualizing Fairness in NLP: The Case of India [paper]
- On Evaluating and Mitigating Gender Biases in Multilingual Settings [paper]
- Cultural Compass: Predicting Transfer Learning Success in Offensive Language Detection with Cultural Features [paper]
- DweshVaani: An LLM for Detecting Religious Hate Speech in Code-Mixed Hindi-English [paper]
- Natural Language Understanding of Devanagari Script Languages: Language Identification, Hate Speech and its Target Detection [paper]
- TRO(F)LL or ROFL ? : Exploring Troll Detection in Tamil Memes [paper]
- Investigating the Effect of Backtranslation for Indic Languages [paper]
- Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision–Language Models [paper]
- Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore [paper]
- BenLLM-Eval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP [paper]
- IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages [paper]
- IndicXNLI: Evaluating Multilingual Inference for Indian Languages [paper]
- INDICGENBENCH: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages [paper]
- Information Parity: Measuring and Predicting the Multilingual Capabilities of Language Models [paper]
If you find this repository useful, please kindly consider citing the following:
@inproceedings{poria2025bhaasha,
title = {Bhaasha, Bhāṣā, Zabān: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challenges},
author = {Sampoorna Poria and Xiaolei Huang},
booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
year = {2025},
address = {Suzhou, China},
publisher = {Association for Computational Linguistics}
}