The problem

Kodava Takk is spoken by roughly 114,000 people and is classified by UNESCO as definitely endangered: children are increasingly not learning it at home. It has no script in daily use — the language lives in conversation, in the Palame folk songs, and in ritual speech, which means every unrecorded elder is an archive lost.

No large speech corpus or language AI exists for Kodava Takk. That gap is also the opportunity: modern speech models can learn a language from a few hundred hours of well-described recordings.

What Kodava Thakk does

The project preserves the language the way it actually exists — as sound. Level 0 is live now: anyone can record a story, song, proverb, or blessing at kodavathakk.kodagu.ai, straight from the browser. Every clip carries speaker, village, okka, dialect, and tiered consent, and lands in a community-governed tracker. A public dial shows voices contributed and hours recorded, in real time.

On that corpus the project fine-tunes open models — speech recognition first, then text-to-speech, then translation — following the playbook proven by the Māori (Te Hiku Media), Mozilla Common Voice, AI4Bharat, and Project Vaani. Everything ships back to the community: a listening archive, a talking dictionary, classroom packs, and Ainmane, a speech-to-speech companion that gives learners someone patient to talk to.

Owned by the community

The corpus is held under a Kodava data guardianship licence modelled on Te Hiku Media's Kaitiakitanga licence: contributors keep moral ownership, a community trust is custodian, commercial use requires council approval and benefit-sharing, and the data can never be sold. Elders' voices are never cloned without family and council consent.

Kodava Thakk is a Nada Kodagu initiative run through the Kodagu.ai community, with the Karnataka Kodava Sahitya Academy, Mangalore University's Kodava MA program, and India's open speech-AI ecosystem as natural partners.

How you can help

  • Every speaker: record 30 seconds of Thakk at kodavathakk.kodagu.ai — then get three relatives to do it
  • Facilitators: sit with the elders of your okka and record the long versions — kits and training provided
  • Transcribers & validators: paid, work-from-home listening and Kannada-script transcription
  • Developers & ML engineers: the whole stack is open source, from this site to the coming speech models
Get Involved View Repository