Release·Sep 5, 2026

Bodhan AI AI4Bharat

Indic-Speak

Write naturally. Hear it come alive.

3.36Bparameters
2+ minsingle-shot generation
45voices
22 / 12languages / scripts

Write the sentence exactly as you would send it: Hindi and English, Tamil and English, native and Latin scripts side by side. Pick a voice. Indic-Speak reads it aloud across 22 languages and 45 voices, without asking you to label where one language ends and another begins.

Think of the last message, lesson or support reply you read in an Indian language. Chances are, it did not stay in one script. A Hindi science lesson keeps mammal and Blue Whale in Latin script inside a Devanagari sentence; a Tamil explainer keeps alveolar capillaries in English. That is not messy text. It is how millions of us write.

Most speech systems treat that ordinary sentence as an edge case: they want a language flag for every span and a phoneme frontend for every language. Indic-Speak needs neither. Give it the sentence as written, choose a voice, and it reads the whole thing. Its 3.36B-parameter stack extends a language model with audio codes, so the same network that models Hindi and English in one token stream can also model how they sound.

At AI4Bharat we have worked on Indian language AI for years, and have open-sourced translation, speech recognition and speech generation across all 22 languages. Speech itself was never the missing piece. Reading that sentence was: the one that changes script mid-clause, the way people actually write. That is the gap we set out to close.

Today, as a joint effort between Bodhan AI and AI4Bharat, we are introducing Indic-Speak: a text-to-speech model built to read text exactly as it is written, across 22 Indian languages and 12 scripts.

The 45 voices retain where they come from: a Kashmiri voice reading Tamil sounds like a Kashmiri speaker reading Tamil. For some products, that travelling accent is the point; for others, it is the reason to choose one of the recommended native voices. Either way, the identity of the speaker does not disappear when the language changes.

Hear it

Do not take our word for it. Listen first; everything after this section is the evidence behind what you hear.

01 Press play before you read another word

Before any claim about the model, the longest thing it made. If this holds up for five and a half minutes, the rest of the page will hold up too.

Chapter one of Ponniyin Selvan, read by Arun in the single person narration audiobook style. Six paragraphs generated one at a time and joined: five minutes and thirty-six seconds in a single voice.

The hard part

Indian writing does not arrive clean. It changes script inside a sentence, borrows English mid-clause, and carries notation nobody says out loud. That is the problem this model was built for, and the rest of the page is what it takes to solve it.

02 It reads the way India writes

No language tag, no pronunciation dictionary, no cleanup pass. The sentence goes in the way a person typed it (two scripts, an English noun in the middle, a number, a formula) and comes back spoken. Every clip here shows the text exactly as it was sent, and each one carries the judge’s top content score.

code-mixing

One sentence. No language labels.

Native script and embedded English in one sentence, read straight through. The language is inferred from the text; there is nothing to tag and no phoneme dictionary to maintain.

delivery

Thirteen ways to say it

One tag moves the delivery, from an All India Radio bulletin to a children's story. With no tag the model uses its conversational register, which is the right choice for most plain reading.

length

A chapter without losing the thread

A chapter is not a sentence. The voice has to stay the same person for minutes at a time, hold its register through description and dialogue, and not drift. The chapter below runs five and a half minutes in one voice.

prosody

Your punctuation directs the performance

A comma creates a breath. A full stop or danda gives the voice room to settle. A question rises; an exclamation adds lift. The pauses you write become the pacing you hear.

Ten languages. English in the middle.

One female and one male native voice per language, reading a benchmark sentence with English terms embedded. Switch language with the tabs.

Text has a sound before it has a voice

A model that reads text as written has to be handed text a person would actually say out loud. Nobody says "one two three comma four five six point seven eight rupees", and nobody says "backslash int". The normaliser in front of the model turns written notation into spoken words first, and it is the part of the stack that does the least glamorous and most necessary work.

What you write is not always what we say

Written symbols have to become spoken words before the model can read them. Here is one real transformation.

03 Forty-five voices. Find yours.

Reading the text correctly is one thing; sounding like a person is another. Forty-five voices cover the 22 languages, every one of them reads every language, and the accent travels with the voice rather than with the text. Ten of the 22 are in the public benchmark and carry measured scores; the other twelve have voices and audio but not yet the same evidence.

Every voice, mapped

One choice drives the rest. Pick a language and you get the voices we would start with in it, where those voices sit on the library’s pitch axis, and the same sentence read by one of them against a voice from somewhere else.

1Pick a language22 languages, 12 scripts, grouped by the script they arrive in

2See where its voices sitall 45 on one axis of measured median pitch; click any of them to hear it

3Hear the accent travelthe same sentence, once in a voice native to the language and once in a voice from elsewhere

All 22 languages as a table(22)
LanguageFemaleMaleJudge, native text
The same 45 as a plot, pitch against pace(45)
  • female
  • male
  • click to play

Every voice speaks every language. Conditioning is fully cross-lingual and there is no separate model per language to switch between: across the benchmark, native casting scored 4.94 and cross-lingual casting 4.90 out of 5: the accent stays, the words survive. The five most portable voices by measured cross-lingual score are . The pair in step 3 is chosen by rule: for each language, the sentence whose transcripts matched most closely among those a native and a non-native voice both placed in the judge’s top fidelity band. Pitch is the median fundamental over each voice’s own generations at the default sampling settings, and it is the only measure here on a shared axis; pace is quoted beside each voice’s own language, because 73% of the variation in characters per second across this library is explained by which script the voice reads (Meitei Mayek averages 9.9 characters per second and Tamil 13.9), against 15% for pitch.

Fourteen deliveries

A voice can be pointed at a register. Eight name a context the speech is going into and six name an emotion. Matching is literal (the capitals and the apostrophe are part of the value), so send each string exactly as it appears.

04 Under the hood

There is no speech pipeline here. There is a language model that learned to spell sound, and a vocoder that turns what it spells back into air.

A 24 kHz waveform is cut into frames of 2,048 samples (85.3 milliseconds each), and every frame is written as seven codes in a fixed 1:2:4 interleave. Those codes were added to the model’s vocabulary alongside its words, so predicting the next instant of audio is the same operation as predicting the next word. Speech comes out at 82.03 audio tokens per second.

Nothing in front of the model converts letters to phonemes, and nothing tells it which language it is reading. It reads the script. What does sit in front is a normaliser, which turns written notation into the words a person would say. You can inspect it in section 02.

One sentence in. One waveform out.

Six steps, one backbone and one vocoder. Everything between the prompt and the audio is the same next-token loop a text model runs.

Only the first two steps are the model. Everything after the token stream is arithmetic: split the frame, look up the codes, run the vocoder.

The stack, counted

Two parts run when you make a request: the backbone that predicts the next token, and the vocoder that turns the codes it emits back into a waveform. The SNAC quantizer is counted with the vocoder: at 139,824 parameters it is a rounding error beside it, and they are one step.

The evidence

Everything above is a demonstration, and a demonstration is chosen. This is the part you can check instead: thirty thousand scored readings, and what they say once you look past the average.

05 Thirty thousand readings. One clear result.

Fifteen thousand code-mixed sentences across ten languages, each read twice by different voices: thirty thousand readings spread over all 45 voices, mostly cross-lingual casting. Every reading was transcribed by an Indic speech recogniser, then scored by an independent judge model on a 0 to 5 content-fidelity scale that treats script, spelling and number-format differences as correct.

Ten languages, tightly grouped

Share of readings in the judge’s top fidelity band, on a scale that starts at 85%, with each language’s 95% confidence interval at the tip of its bar. Ten languages, 3,000 readings each, spread across all 45 voices. Ordered alphabetically, because the ordering by rate would not mean anything.

The axis starts at 85%, marked by the break at each bar’s base. All ten fall between 92.0% and 93.6%, a spread of 1.6 percentage points against a measurement uncertainty of about ±0.9 points on each. Nine of the ten intervals contain the all-language mean, and the gap between the highest and lowest does not survive correcting for the number of comparisons ten languages allow. There is no strong or weak language here; there is one number, measured ten times.

View data table
Judge score, top-band rate, low-score rate and reading count per text language
Text languageJudge (0–5)Scored 5Scored ≤2Readings

the three numbers that matter

What happens beyond the average

Thirty thousand readings, looked at past the mean: how much the casting matters, how far the voices spread, and how often a reading needs a second take.

native vs cross-lingual
4.94 vs 4.90
readings scored ≤2
0.70%
weakest to strongest voice
4.53 to 4.96
Why not word error rate? The recogniser writes English words in native script, so mammal in a Bengali sentence comes back as ম্যামাল and counts as an error against the Latin-script source even when it was spoken cleanly. Across the readings in the judge's top band, string word error rate still averages 48.8 and character error rate 52.9: on code-mixed text those numbers measure the script mismatch, not the speech. We keep both in the data and let the judge give the verdict.

coming soonPreference ranking tests are ongoing and will land soon. Content fidelity is one axis; human preference against other Indian-language systems, on naturalness and voice quality, is in progress and this post will be updated with the results.

Judge: Gemma-4-31B-IT with a content-fidelity rubric. Recogniser: IndicCanary. Reference text is the normalised input, so number expansion is scored as spoken rather than as written.

Before you choose

What it cannot do yet, where the evidence stops, and what we are building next.

06 Know the edges before you ship

Three things the results above do not show on their own, each measured on the same benchmark run.

  • The voices are not interchangeable. The judge mean runs from 4.53 to 4.96 across the library, and the weakest voice needs a second take about nine times as often as the strongest. The recommended voice for each language stays out of the thin end of that range, so start there rather than picking at random.
  • Benchmark coverage currently spans ten languages. The other twelve have voices and playable audio, but not yet the same scored evidence. Treat them as supported and still being measured, rather than assuming they match the benchmarked set.
  • Most readings work first time; 0.7% need a second take. That is 211 of 30,000. Sampling is not deterministic, so a second take draws a fresh reading rather than reproducing the first one.

07 The next voice should feel more human

This release is a starting point, not a victory lap. These are the gaps we are working on next. We are not attaching dates; each improvement lands here when it holds up to the same evidence as this release.

Six things we are working on next

Make the model think before it speaks

thinking TTSToday the prosody of a reading comes from what you send: the punctuation carries the breaths and a style tag sets the register. Next the model reasons about the context before it speaks and decides those for itself (where the emphasis belongs, which clause to slow down, what register the passage is actually in), so one sentence reads differently as a news bulletin, as a bedtime story, and as an answer to someone who sounds worried.
preference studyPreference ranking against other Indian-language speech systems, on the code-mixed text these voices are built for. Content fidelity is what this post measures; naturalness is not, and we do not claim it yet. Those tests are running now and will be published here whichever way they come out.

Put better speech in more hands

smaller modelA smaller, cheaper Indic-Speak: the same voices and the same languages at a fraction of the cost per second of audio, so speech can sit inside an interactive product rather than only inside a batch job.
expressive voicesVoices built for range rather than neutrality: laughter, hesitation, warmth, urgency, and the shifts inside a single passage that separate a person talking from a narrator reading. More voices in each language arrive with them.
long formSingle-pass generation of multi-minute passages, so a chapter holds together end to end instead of being assembled from pieces.
timestampsWord-level timestamps alongside the audio, for subtitling, read-along and anything that has to line text up against audio.
India does not speak one language at a time. Its speech technology should not ask people to flatten a sentence, hide an accent or rewrite themselves for the machine. Indic-Speak is our step towards the opposite: meet people in the words they already use, then let those words sound like someone they know.

Cite this work

Released under the Indic Open Model License v1.0.

bibtex
@misc{indic-speak-2026,
  title  = {Indic-Speak: Text-to-Speech for 22 Indian Languages in 45 Voices},
  author = {Bodhan AI and AI4Bharat},
  year   = {2026},
  url    = {https://bodhan.ai/research/blogs/indic-speak}
}