Antariksha.ai | Foundation Models for India's Languages
TANUMANASA
RESEARCH
Book a consultation
FLAGSHIP FOUNDATION MODEL FAMILY

Antariksha.ai

A foundation model family built for Indian languages — with the reasoning depth, efficiency, and sovereignty India's future demands.

Request early access Read the technical brief
WHAT IS ANTARIKSHA.AI?

Indian languages as first-class citizens, not an afterthought.

Antariksha.ai is a domain-specific foundation model family developed by Tanumanasa, focused on India's languages — including Hindi, Odia, Bengali, Tamil, Telugu, Marathi and others. Rather than treating Indian languages as an afterthought, Antariksha.ai is built around them: continued pre-training on Indian-language corpora, instruction tuning, and alignment for the contexts where it will actually be used.

Why India needs its own foundation model

Linguistic coverage

Global models under-serve Indian languages, especially low-resource and tribal ones.

Sovereignty

Critical infrastructure should not depend entirely on models built and governed elsewhere.

Contextual fidelity

Indian knowledge, idioms and use-cases need models trained on them.

Cost & efficiency

Frugal model development makes capable AI affordable at Indian scale.

ARCHITECTURE

Large enough to reason. Efficient enough to deploy.

Built in the practical 7B–32B parameter range: a capable open base, continued pre-training on Indian-language corpora, supervised fine-tuning, and preference alignment. Parameter-efficient methods and quantisation keep both training and inference frugal.

[Final base, sizes and context length to be confirmed before publishing.]

MODEL CARD · ANTARIKSHA.AI v-preview
PARAMETERS
7B / 14B / 32B — tiered for edge → server deployment
BASE APPROACH
Open base + continued pre-training (Qwen 2.5 / Sarvam-1 class)
ALIGNMENT
SFT + DPO — instruction-following and preference tuning
EFFICIENCY
LoRA / QLoRA, quantisation — GGUF / AWQ inference formats
CONTEXT
[TBD — confirm before publishing]
LANGUAGES SUPPORTED

Phase 1 languages, honestly labelled.

Phase 1 focuses on five major Indian languages, with a roadmap toward all 22 scheduled languages and selected tribal and low-resource languages.

Hindi · हिन्दी Odia · ଓଡ଼ିଆ Bengali · বাংলা Tamil · தமிழ் Telugu · తెలుగు ← PHASE 1
Punjabi Gujarati Kannada Punjabi Santali + 22 scheduled & tribal languages ← TARGET
TRAINING VISION

Data India can trust.

Our dataset initiative catalogues existing corpora and creates new, high-quality data across India's 22 scheduled languages plus tribal and low-resource languages, in partnership with universities. The goal is not just more data, but representative, ethically-sourced, well-documented data — the foundation of a model India can trust.

BENCHMARKS
BENCHMARK
ANTARIKSHA.AI
OPEN BASELINE
Indic language understanding
pending
pending
Reasoning (Indian context)
pending
pending
Translation quality
pending
pending

Results publish only once verified, with methodology and reproducible details.

ROADMAP

Three phases to sovereignty.

I
PHASE 1 · NOW

Core Indian languages, base + instruct models, early access.

II
PHASE 2

Expansion toward 22 scheduled languages, multimodal exploration.

III
PHASE 3

Tribal & low-resource coverage, domain variants, broader availability.

EARLY ACCESS

Build with Antariksha.ai first.

Join the waitlist for early model and API access. We'll confirm by email and respond as cohorts open.

TANUMANASA
RESEARCH

Building Bharat's next foundation models — from Odisha, for India, for the world.

COMPANY
PRODUCTS
CAPABILITIES
RESOURCES
© 2026 Tanumanasa Research Pvt. Ltd. · EmTeck CoE, STPI Bhubaneswar, Odisha, India Privacy Policy Terms of Use Responsible AI