Digitising Kannada: A work in progress

From electronic typesetting to ASCII-based proprietary fonts to Unicode to AI-enabled translation, Kannada has come a long way in its digital journey. But Indian languages face unique challenges in today’s fast-evolving tech landscape, writes Shilpa Elizabeth.

C.S. Yogananda, founder of Sriranga Digital Software Technologies, remembers a funny instance of coming across a post on X around two years ago. Originally written in Kannada, the post, after machine translation, took on an entirely different meaning.

In Kannada, he wrote ‘ಸಹಾಯ ಮಾಡಿದವರನ್ನ ಮರೀಬಾರ್ದು (Sahaya madidavaranna mareebardu), which translates to ‘One should not forget those who helped’. X’s default translation in 2024-25 read, “Kill those who helped you.”

In yet another instance of hilariously wrong machine translations, a Bengaluru-based journalist recollects how when she used the “translate page” option on a Kannada website to read the content in English, an MLA’s name was translated as “ready to eat.”

The quirks of automated translation often make for amusing anecdotes. But it doesn’t take much time for things to take a more serious turn, especially when mistranslations appear as errors in documents like land records or voter ID, leading to documentation inconsistencies and larger consequences.  

From electronic typesetting to ASCII-based proprietary fonts to Unicode to AI-enabled translation, Kannada has come a long way in its digital journey. However, a deeper look shows that lacunae still persist: some of it translating into real life problems for many, others resulting in the homogenisation of a language and culture that is richly diverse.

A pioneering effort

Kannada’s digital journey began fairly early. In fact, it was the second Indian language to be electronically typeset, after Hindi, in the 1970s, notes K.P. Rao, a pioneer in Indian-language computing. He was instrumental in creating a digital font or set of glyphs for the Kannada characters.

The Hindu archives from 1976 show how scholar I. Mahadevan wanted to publish a digital concordance of over 9,000 Indus Valley seals. But computers and printing systems were designed for English, and digital fonts didn’t exist either for Indus Valley script or for any Indian languages. The archives show all the basic symbols in the Indus script were drawn by an artist and the same were microfilmed. The films were prepared by Tata Press, Bombay, where Rao used to work then.

The same system was soon adapted to print Kannada on an experimental basis at Tata Press, Bombay, in the early 1970s, says Rao. Both these efforts were experimental and worked by substituting language glyphs for English-language characters, he adds.

But these early systems were only a workaround, where English letters were simply replaced with Kannada characters. This allowed Kannada text to be printed, but computers could not still process Kannada as a language.

Typing Indian scripts, which have far more character combinations than English, was a huge challenge. While English uses 26 letters that are typed one after another, Indian scripts such as Kannada have hundreds of character combinations created by joining consonants, vowels and other symbols. Fitting all of these onto a keyboard was far from straightforward, unlike in English.

“Fortunately, a good number of computer enthusiasts started working on the problem of putting Indian-language text into computers around this time,” says Rao, recollecting names like R. Kasturi Rangan and Vishweshwar Dixit. However, the non-standard layout of keyboards continued to pose a challenge.

From phonetic input to Unicode

In the early 1980s, Monotype India’s Bengaluru office developed a phonetic input system, where, instead of creating a separate key for every character, users typed words as they sounded on a standard English keyboard, and the software converted them into the correct Indian script.

This became the basis for one of the earliest Indian-language typing systems. “This led many Bangalore computer hackers to take up the wider range of input and output problems across languages,” says Rao.

In the 1980s, with manual typesetting giving way to desktop publishing or DTP, more Indian language fonts, page layout software and publishing tools started coming into the picture. Software such as Sediyapu by K.P. Rao, Baraha by Sheshadrivasu Chandrasekharan, and later Nudi by Kannada Ganaka Parishad brought Kannada typing to personal computers and government offices.

U.B. Pavanaja created the first Kannada programme called Kannada Kali in 1993 and the first Kannada website named ‘Vishva Kannada’ in 1996. “After Hindi and Marathi, it was the third Indian language to go online,” reminisces Pavanaja.

However, every software developer used their own font systems. In other words, Kannada text was essentially treated as pictures. If the recipient computer or operating system did not have the same font installed, the text appeared as gibberish.

From printable to digital

It was the entry of Unicode that elevated Kannada from ‘printable’ to ‘digital’. Unicode assigned every Kannada character a universal code understood by all modern computers. This meant that the language would appear correctly across computers, could be searched, copied and pasted, indexed by search engines and be exchanged across operating systems, among other things.

“It helped with standardisation and platform independence,” says Pavanaja.

Several technologies such as speech-to-text, Optical Character Recognition, and AI followed taking ahead the digitisation of the language.

Absence of data

Today, the problems of Kannada digitisation have moved away from keyboard layouts and fonts to more nuanced, yet complex, challenges. According to linguists and technologists, one of the most persistent problems today is the lack of a large, high-quality, annotated Kannada corpus.

“AI systems can only learn from the data they are trained on. Without a rich corpus, technologies such as machine translation, speech recognition and text-to-speech will always lag behind English, which has had a 30- to 40-year headstart in building such resources. Kannada has a substantial amount of digital content, but much of it has not been systematically cleaned, annotated or organised for AI,” says Yogananda, a prominent mathematician who has played a major role in Kannada computing, typesetting, and digital archiving.

Sriranga Digital Software Technologies, the organisation founded by him, has been digitising many Kannada books, magazines and literary works over the last two decades.

The team also developed one of the earliest Kannada OCR systems around 2008-09 to make scanned texts searchable. Yogananda, however, notes that scanning books is only the first step. The next stage is to convert this material into clean, structured datasets that AI systems can learn from.

Linguistic challenges

The complex characteristics of agglutinative languages like Kannada (where thousands of variations are created by adding prefixes, suffixes and markers onto a root word) pose unique computational challenges.

The grammar rule of “sandhi” is an example. When mane + alli becomes maneyalli, a new ‘y’ is thrown into the mix to bridge the vowels, which confuses AI and uses up more computational power. Further, if a slang or typo makes it to the sentence, it could lead to misinterpretations by AI and crank up the computational costs.

Yogananda points out that such complexities of the language make it harder to build language models for Kannada than for languages with simpler morphology, and increases the need for carefully annotated datasets.

Apart from linguistic differences, Pavanaja notes that the current AI systems are still largely shaped by Western ways of thinking, both linguistically and culturally. While India has begun developing its own large language models through initiatives such as Sarvam AI, he feels these efforts are still far from the likes of Gemini, ChatGPT, Claude or Perplexity.

Interconnected issues

Community-driven sister organisations Sanchi and Sanchaya have played a critical role in preserving Karnataka’s cultural and linguistic heritage by developing digital tools, digitising literary work, archives and manuscripts, reviving Kannada typefaces and improving the accessibility of all of those making them open-access.

Om Shivaprakash H.L., who co-founded the initiatives, paints a picture of how problems related to absence of data, OCR, dictionaries and typefaces are all interconnected.

“Digitising books creates the historical data needed to build better OCR. Better OCR helps extract text from older books, manuscripts and inscriptions, which in turn contributes to richer dictionaries, corpora and language resources. These resources then become the foundation for grammar tools, spell-checkers, translation systems and AI applications.”

He further adds, “Many people think the problem is solved because tools like Google Lens can recognise printed Kannada. But that is far from enough. OCR still struggles with old books, manuscripts and inscriptions, where much of our linguistic and cultural heritage lies. AI can help, but only if it is trained on high-quality data. That data has to be created first.”

Given that many government institutions still use legacy ANSI fonts instead of Unicode, a large amounts of government data remain invisible to search engines. “If official documents cannot be searched or indexed, valuable knowledge and good governance practices remain inaccessible,” he notes.

While typing has become easier, Shivaprakash highlights that spell-checkers, grammar checkers and other language technologies are still underdeveloped for Indian languages including Kannada.

Real-life issues, linguistical divide

Stakes are much higher when transliteration errors make their way into government documents. In records such as voter ID cards and Aadhaar, details entered in English are often automatically rendered in regional languages. Many a time, this results in errors.

“My name is Pavanaja Ubaradka Bellippady. But they have entered it as Bellippaddy in Kannada in the voter card. Now SIR has begun. My name on the Aadhaar and the voter ID does not match,” worries Pavanaja, who has already faced difficulties accessing government services owing to name mismatch.

Garbage typing and incorrect rendering have contributed to mistakes in land records during digitisation efforts.

“Correcting mistakes in Karnataka’s land records requires lengthy legal and administrative procedures, making these errors especially serious. These as examples of the linguistic digital divide, where seemingly small technical issues ultimately affect citizens’ identities, property records and access to public services,” says an expert who didn’t wish to be named.

Government initiatives

Project Vaani, an initiative by IISc and ARTPARK, aims to capture India’s diverse languages, dialects, and speech patterns and build an open source corpus of 1,50,000 hours of natural speech data from around 1 million people across nearly 800 Indian districts.

The team has completed two phases of the project, covering 165 Indian districts, 156,534 speakers, and 105 languages.

Project Vaani is being carried out under MeitY’s Bhashini. The other prominent government and publicly funded initiatives in Indian language technologies include the Technology Development for Indian Languages (TDIL) programme, C-DAC’s GIST technologies, AI4Bharat at IIT Madras, and BharatGen. These span machine translation, speech recognition, text-to-speech, multilingual AI models, digital public infrastructure and open datasets for Indian languages.

In Karnataka, Kannada Development Authority Chairman Purushotham Bilimale says that KDA has submitted a proposal to the State government to create an archive of folklore and oral traditions and has requested ₹2 crore for the same. The government is yet to respond. “The State government has given almost ₹1 crore to Mysore University to digitise some of the manuscripts. There are also entities, such as e-Kannada, which is now looking into digitisation,” Bilimale notes.

The e-Kannada project, formulated by the Department of Personnel and Administrative Reforms and the Kannada Development Authority (KDA), was formed in 2006 under the Centre for e-Governance to bridge the gap between Kannada and digital advancement. Initiatives including the e-Kannada Learning Portal, Padakanaja, transliteration tools, spell checker and e-Kalika Kannada Academy, among others, were set up as part of the project.

Community- and market-driven

Experts, however, feel that most of the progress in digitising Kannada has been driven by volunteer community and market.

“It is worth noting that, for much of this period, it was not government or government institutions that helped develop Kannada linguistic technology — it was the effort of private individuals, contributing and refining these tools and making them freely available to users,” Rao notes.

“Universities, government departments and research institutions should have taken the lead in creating open datasets, language tools and digital resources. Instead, much of this work has been carried out by volunteers,” argues Shivaprakash, who adds that many of the datasets also remain inaccessible for researchers.

“If publicly funded institutions generate language data, it should not remain behind paywalls or in isolated repositories,” he demands.

Sustained effort missing

Around two years ago, KDA announced plans to launch Kannada Kasturi, a proposed alternative to Google Translate. While the work seems to have been completed, it has not been launched yet for reasons unknown. This underscores a criticism that people like Pavanaja have often been raising – that governments repeatedly launch new projects instead of strengthening and sustaining existing ones.

Kanaja, the Karnataka government’s digital knowledge repository, developed the Karnataka Knowledge Commission. Although the website still exists, Pavanaja says very little new content has been added over the years. According to him, the portal also crashed multiple times, resulting in the loss of several years of data. While old back-ups were restored, a significant amount of content could not be recovered, he says.

According to him, a large part of corpus-building, in India, has been driven by private companies such as Google and Microsoft rather than the government.

Foundational datasets

Shivaprakash, however, cautions that while AI-powered translation and chatbots are useful, they should not be mistaken for complete language development.

“They are built on data created by others. If we fail to build the foundational datasets, dictionaries, corpora and language tools ourselves, we risk depending entirely on proprietary technologies.”

Karnataka is a treasure trove of an enormous amount of knowledge in old books, manuscripts, inscriptions and family collections that have never been digitised. But as these materials deteriorate, we risk losing them permanently.

“Minority languages and dialects are especially vulnerable because much of their knowledge survives only through oral traditions or unpublished manuscripts,” Shivaprakash warns.

Technology alone cannot preserve a language, he points out. “Linguists, technologists, volunteers, universities, government agencies and language communities all have to work together. Digitisation is not just about technology; it is about preserving cultural memory.”

source/content: thehindu.com (headline edited)

Leave a Reply

Your email address will not be published. Required fields are marked *