Transliteration – AI4BHĀRAT

Machine Transliteration

Indic languages are written in a variety of scripts (Brahmi family of abugida scripts, Arabic-derived abjad scripts, and even alphabetic Roman script). This diversity makes it challenging to support mechanisms which are convenient for typing or creating content in these diverse languages and scripts. Most Indian users are comfortable with the Roman keyboard and thus an optimal solution that users find beneficial is automatic transliteration of the romanized input into the native script. To enable this, at AI4Bharat, we have undertaken the task of creating large-scale transliteration corpora for Indic languages along with models for transliteration of romanized inputs into native scripts.