The First Interface of Indigenous Language AI Is Not the Microphone, It Is Consent
Original Chinese title: 族語 AI 不是把錄音倒進模型:資料主權才是第一個介面
The core of Indigenous language AI is not collecting more recordings, but establishing mechanisms for consent, classification, withdrawal, refusal, and feedback so that technology truly serves language revitalization.
Two-Eyed Seeing Lab
Research group on Indigenous language technology, cultural data sovereignty, and Two-Eyed Seeing.

The First Interface of Indigenous Language AI Is Not the Microphone, It Is Consent
Working on Indigenous language AI can easily make people excited: recording, transcribing, translating, chatbots, teaching apps — it seems as if simply putting elders' voices into a model would automatically revitalize the language. This is the most common illusion in technical presentations. Indigenous languages are not data mines, and elders are not free voice suppliers. The true first interface is not the microphone, but consent: who consents, to what extent, whether they can withdraw, whether future models must be re-asked, and whether commercial use incurs additional charges.
If these questions are not addressed first, Indigenous language AI will quickly turn from a revitalization tool into a new extraction machine. It will flatten a sentence carrying family, ritual, place names, and context into "audio plus transcription." It will push knowledge that should remain private into open data, turning learning relationships requiring long-term companionship into download buttons. Worst of all, when platforms fail, communities are blamed for having insufficient data.
Open Source Is Not a Get-Out-of-Jail Card
Global low-resource language technology is accelerating. Initiatives like Common Voice have made many languages available with usable voice datasets; research on Quechua, Pashto, and others shows that open corpora can bring long-neglected languages into speech recognition and natural language processing. But "can be opened" does not equal "should be opened," nor does "clear licensing" mean "culturally appropriate." For certain oral traditions, CC0-style full openness may not be democratic; it may instead hand forward control to future generations.
The CARE principles remind us that data governance must look beyond FAIR — findable, accessible, interoperable, reusable — and also consider collective benefit, control, responsibility, and ethics. Applied to Indigenous language AI, this means datasets are not better simply because they are larger; they are better when clearer: clearly marking sources, purposes, visibility scopes, prohibited contexts, feedback mechanisms, and stop conditions.
Models Must Know How to Speak — And When to Stay Silent
The most important capability of an Indigenous language chatbot may not be answering, but refusing. When encountering taboo knowledge, family-exclusive stories, unpublicized rituals, sensitive locations, or plant knowledge easily commercialized, the system should say: "This is not suitable for reply here; please return to the community authorization context." This is not a functional shortfall; it is a cultural safety feature. General product managers may find this troublesome because refusals lower interaction rates; but for traditional knowledge, only models that know how to stay silent can be trusted.
Technically, what can be done? First, data classification: separate public teaching corpora, community-internal corpora, restricted-access corpora, and non-machine-readable corpora. Second, versioned authorization: bind each batch's consent forms and purposes so that "last year agreed to teaching" does not automatically become "this year may train commercial models." Third, output auditing: record when the model uses which data types to generate answers, enabling communities to trace usage. Fourth, withdrawal mechanisms: when community members or families request removal, the system should not say "the model has been trained and cannot be deleted," but at least provide a freeze, replacement, retraining, and risk notification process.
Revitalization Is Not Handing People Over to Models
The best role of Indigenous language AI is not replacing teachers, elders, or families, but reducing the time they spend on administrative forms and low-quality teaching materials. It can assist in organizing curricula, generating exercises, comparing pronunciations, searching corpora, and performing initial transcription; yet where languages truly live remains human relationships. A child learning a greeting does more than get phonemes right — they learn when to speak, to whom, and with what posture.
Therefore, the success of Indigenous language AI should not be judged solely by recognition rates, but also by governance quality: can community members understand what the system is doing? Do data return to the community? Are teachers less burdened? Are elders respected? Have young people shifted from users to co-designers? If the answers are all no, then even high model scores amount only to colonial technology wearing a trendier skin.
Indigenous language technology is worth pursuing — and must be pursued. But it cannot begin with "give me the data"; it must start with "let us together decide how the data will live on."
---
Sources retained from the Chinese original
AI use and content-safety disclosure
This article was assisted by AI in organizing information on low-resource language technology and Indigenous Data Sovereignty. Human editors were responsible for cultural context, taboo knowledge boundaries, and governance perspectives.