Language intelligence for Odia and low-resource Indian languages.
We are a language infrastructure company building the AI layer for the Odia language (~50M speakers). Based in Dhenkanal, Odisha, India.
- Tokenizer research — Efficient Unicode-aware tokenization for Brahmic scripts
- Lekhani model family — Odia-optimized LLMs (Pada, Chhanda, Kavya, Mahakavya)
- Shruti — Odia speech recognition and synthesis
- Khoja — Odia semantic search
- Anuvada — Odia translation
- Patra — Odia OCR
- Deep specialization — One language, done well. Not 22 languages done adequately.
- Open source by default — Apache 2.0 license. The ecosystem grows when the foundation is open.
- Built for Odisha — Government, media, education, and enterprise-grade reliability.
- Efficiency first — Every API call should cost less. Our tokenizer is 3x more efficient for Odia.
| Repository | Description |
|---|---|
| odia-tokenizer-benchmarks | Benchmarking framework for Odia tokenizer evaluation |
| lekhani-pada | Lekhani Pada — lightweight Odia language model |
| odia-pretrain-dataset | CC-BY-4.0 corpus for Odia language model pretraining |
- Twitter/X: @maelisresearch
- Hugging Face: maelis-research