Think of the last message, lesson or support reply you read in an Indian language. Chances are, it did not stay in one script. A Hindi science lesson keeps mammal and Blue Whale in Latin script inside a Devanagari sentence; a Tamil explainer keeps alveolar capillaries in English. That is not messy text. It is how millions of us write.
Most speech systems treat that ordinary sentence as an edge case: they want a language flag for every span and a phoneme frontend for every language. Indic-Speak needs neither. Give it the sentence as written, choose a voice, and it reads the whole thing. Its 3.36B-parameter stack extends a language model with audio codes, so the same network that models Hindi and English in one token stream can also model how they sound.
At AI4Bharat we have worked on Indian language AI for years, and have open-sourced translation, speech recognition and speech generation across all 22 languages. Speech itself was never the missing piece. Reading that sentence was: the one that changes script mid-clause, the way people actually write. That is the gap we set out to close.
Today, as a joint effort between Bodhan AI and AI4Bharat, we are introducing Indic-Speak: a text-to-speech model built to read text exactly as it is written, across 22 Indian languages and 12 scripts.
The 45 voices retain where they come from: a Kashmiri voice reading Tamil sounds like a Kashmiri speaker reading Tamil. For some products, that travelling accent is the point; for others, it is the reason to choose one of the recommended native voices. Either way, the identity of the speaker does not disappear when the language changes.
Hear it
Do not take our word for it. Listen first; everything after this section is the evidence behind what you hear.
01 Press play before you read another word
Before any claim about the model, the longest thing it made. If this holds up for five and a half minutes, the rest of the page will hold up too.
Chapter one of Ponniyin Selvan, read by Arun in the single person narration audiobook style. Six paragraphs generated one at a time and joined: five minutes and thirty-six seconds in a single voice.
The hard part
Indian writing does not arrive clean. It changes script inside a sentence, borrows English mid-clause, and carries notation nobody says out loud. That is the problem this model was built for, and the rest of the page is what it takes to solve it.
02 It reads the way India writes
No language tag, no pronunciation dictionary, no cleanup pass. The sentence goes in the way a person typed it (two scripts, an English noun in the middle, a number, a formula) and comes back spoken. Every clip here shows the text exactly as it was sent, and each one carries the judge’s top content score.
code-mixing
One sentence. No language labels.
Native script and embedded English in one sentence, read straight through. The language is inferred from the text; there is nothing to tag and no phoneme dictionary to maintain.
delivery
Thirteen ways to say it
One tag moves the delivery, from an All India Radio bulletin to a children's story. With no tag the model uses its conversational register, which is the right choice for most plain reading.
length
A chapter without losing the thread
A chapter is not a sentence. The voice has to stay the same person for minutes at a time, hold its register through description and dialogue, and not drift. The chapter below runs five and a half minutes in one voice.
prosody
Your punctuation directs the performance
A comma creates a breath. A full stop or danda gives the voice room to settle. A question rises; an exclamation adds lift. The pauses you write become the pacing you hear.
Ten languages. English in the middle.
One female and one male native voice per language, reading a benchmark sentence with English terms embedded. Switch language with the tabs.
Text has a sound before it has a voice
A model that reads text as written has to be handed text a person would actually say out loud. Nobody says "one two three comma four five six point seven eight rupees", and nobody says "backslash int". The normaliser in front of the model turns written notation into spoken words first, and it is the part of the stack that does the least glamorous and most necessary work.
What you write is not always what we say
Written symbols have to become spoken words before the model can read them. Here is one real transformation.
03 Forty-five voices. Find yours.
Reading the text correctly is one thing; sounding like a person is another. Forty-five voices cover the 22 languages, every one of them reads every language, and the accent travels with the voice rather than with the text. Ten of the 22 are in the public benchmark and carry measured scores; the other twelve have voices and audio but not yet the same evidence.
Every voice, mapped
One choice drives the rest. Pick a language and you get the voices we would start with in it, where those voices sit on the library’s pitch axis, and the same sentence read by one of them against a voice from somewhere else.
1Pick a language22 languages, 12 scripts, grouped by the script they arrive in
2See where its voices sitall 45 on one axis of measured median pitch; click any of them to hear it
3Hear the accent travelthe same sentence, once in a voice native to the language and once in a voice from elsewhere
All 22 languages as a table(22)
| Language | Female | Male | Judge, native text |
|---|
The same 45 as a plot, pitch against pace(45)
- female
- male
- click to play
Every voice speaks every language. Conditioning is fully cross-lingual and there is no separate model per language to switch between: across the benchmark, native casting scored 4.94 and cross-lingual casting 4.90 out of 5: the accent stays, the words survive. The five most portable voices by measured cross-lingual score are . The pair in step 3 is chosen by rule: for each language, the sentence whose transcripts matched most closely among those a native and a non-native voice both placed in the judge’s top fidelity band. Pitch is the median fundamental over each voice’s own generations at the default sampling settings, and it is the only measure here on a shared axis; pace is quoted beside each voice’s own language, because 73% of the variation in characters per second across this library is explained by which script the voice reads (Meitei Mayek averages 9.9 characters per second and Tamil 13.9), against 15% for pitch.
Fourteen deliveries
A voice can be pointed at a register. Eight name a context the speech is going into and six name an emotion. Matching is literal (the capitals and the apostrophe are part of the value), so send each string exactly as it appears.
04 Under the hood
There is no speech pipeline here. There is a language model that learned to spell sound, and a vocoder that turns what it spells back into air.
A 24 kHz waveform is cut into frames of 2,048 samples (85.3 milliseconds each), and every frame is written as seven codes in a fixed 1:2:4 interleave. Those codes were added to the model’s vocabulary alongside its words, so predicting the next instant of audio is the same operation as predicting the next word. Speech comes out at 82.03 audio tokens per second.
Nothing in front of the model converts letters to phonemes, and nothing tells it which language it is reading. It reads the script. What does sit in front is a normaliser, which turns written notation into the words a person would say. You can inspect it in section 02.
One sentence in. One waveform out.
Six steps, one backbone and one vocoder. Everything between the prompt and the audio is the same next-token loop a text model runs.
Only the first two steps are the model. Everything after the token stream is arithmetic: split the frame, look up the codes, run the vocoder.
The stack, counted
Two parts run when you make a request: the backbone that predicts the next token, and the vocoder that turns the codes it emits back into a waveform. The SNAC quantizer is counted with the vocoder: at 139,824 parameters it is a rounding error beside it, and they are one step.
The evidence
Everything above is a demonstration, and a demonstration is chosen. This is the part you can check instead: thirty thousand scored readings, and what they say once you look past the average.
05 Thirty thousand readings. One clear result.
Fifteen thousand code-mixed sentences across ten languages, each read twice by different voices: thirty thousand readings spread over all 45 voices, mostly cross-lingual casting. Every reading was transcribed by an Indic speech recogniser, then scored by an independent judge model on a 0 to 5 content-fidelity scale that treats script, spelling and number-format differences as correct.
Ten languages, tightly grouped
Share of readings in the judge’s top fidelity band, on a scale that starts at 85%, with each language’s 95% confidence interval at the tip of its bar. Ten languages, 3,000 readings each, spread across all 45 voices. Ordered alphabetically, because the ordering by rate would not mean anything.
The axis starts at 85%, marked by the break at each bar’s base. All ten fall between 92.0% and 93.6%, a spread of 1.6 percentage points against a measurement uncertainty of about ±0.9 points on each. Nine of the ten intervals contain the all-language mean, and the gap between the highest and lowest does not survive correcting for the number of comparisons ten languages allow. There is no strong or weak language here; there is one number, measured ten times.
View data table
| Text language | Judge (0–5) | Scored 5 | Scored ≤2 | Readings |
|---|
the three numbers that matter
What happens beyond the average
Thirty thousand readings, looked at past the mean: how much the casting matters, how far the voices spread, and how often a reading needs a second take.
- native vs cross-lingual
- 4.94 vs 4.90
- readings scored ≤2
- 0.70%
- weakest to strongest voice
- 4.53 to 4.96
coming soonPreference ranking tests are ongoing and will land soon. Content fidelity is one axis; human preference against other Indian-language systems, on naturalness and voice quality, is in progress and this post will be updated with the results.
Judge: Gemma-4-31B-IT with a content-fidelity rubric. Recogniser: IndicCanary. Reference text is the normalised input, so number expansion is scored as spoken rather than as written.
Before you choose
What it cannot do yet, where the evidence stops, and what we are building next.
06 Know the edges before you ship
Three things the results above do not show on their own, each measured on the same benchmark run.
- The voices are not interchangeable. The judge mean runs from 4.53 to 4.96 across the library, and the weakest voice needs a second take about nine times as often as the strongest. The recommended voice for each language stays out of the thin end of that range, so start there rather than picking at random.
- Benchmark coverage currently spans ten languages. The other twelve have voices and playable audio, but not yet the same scored evidence. Treat them as supported and still being measured, rather than assuming they match the benchmarked set.
- Most readings work first time; 0.7% need a second take. That is 211 of 30,000. Sampling is not deterministic, so a second take draws a fresh reading rather than reproducing the first one.
07 The next voice should feel more human
This release is a starting point, not a victory lap. These are the gaps we are working on next. We are not attaching dates; each improvement lands here when it holds up to the same evidence as this release.