Antariksha.ai is a domain-specific foundation model family developed by Tanumanasa, focused on India's languages — including Hindi, Odia, Bengali, Tamil, Telugu, Marathi and others. Rather than treating Indian languages as an afterthought, Antariksha.ai is built around them: continued pre-training on Indian-language corpora, instruction tuning, and alignment for the contexts where it will actually be used.
Global models under-serve Indian languages, especially low-resource and tribal ones.
Critical infrastructure should not depend entirely on models built and governed elsewhere.
Indian knowledge, idioms and use-cases need models trained on them.
Frugal model development makes capable AI affordable at Indian scale.
Built in the practical 7B–32B parameter range: a capable open base, continued pre-training on Indian-language corpora, supervised fine-tuning, and preference alignment. Parameter-efficient methods and quantisation keep both training and inference frugal.
[Final base, sizes and context length to be confirmed before publishing.]
Phase 1 focuses on five major Indian languages, with a roadmap toward all 22 scheduled languages and selected tribal and low-resource languages.
Our dataset initiative catalogues existing corpora and creates new, high-quality data across India's 22 scheduled languages plus tribal and low-resource languages, in partnership with universities. The goal is not just more data, but representative, ethically-sourced, well-documented data — the foundation of a model India can trust.
Results publish only once verified, with methodology and reproducible details.
Core Indian languages, base + instruct models, early access.
Expansion toward 22 scheduled languages, multimodal exploration.
Tribal & low-resource coverage, domain variants, broader availability.
Join the waitlist for early model and API access. We'll confirm by email and respond as cohorts open.
Building Bharat's next foundation models — from Odisha, for India, for the world.