Low-resource languages are not cheap data: Indigenous language AI must learn manners first
Original Chinese title: 低資源語言不是低價資料:族語 AI 要先學會禮貌
Starting from research on digitizing minority languages in Bangladesh, this article discusses low-resource NLP, Indigenous language data governance, community consent and cultural data sovereignty.
Two-Eyed Seeing Lab
Consultant: 高德生
The Two-Eyed Seeing Lab focuses on Indigenous Peoples' knowledge, data sovereignty and AI applications; this piece features 高德生 as cultural context consultant.

Low-resource languages are not cheap data
The AI industry loves one term: low-resource languages. It sounds neutral, even carrying a touch of technical pity, as if these languages simply lack enough data, models are too small and annotation is too expensive. But for many communities, language is not about being "resource-poor"; it has been systematically excluded from resources by state education systems, market platforms and colonial classifications. What is truly low is not the value of the language but the patience mainstream technology is willing to invest.
A recent study on digitizing minority and Indigenous languages in Bangladesh proposes large-scale parallel, multimodal corpus construction, including text, translation, IPA phonetic transcription and audio data. Such work is important for low-resource NLP because without corpora it is very difficult to build speech recognition, machine translation, teaching tools and digital archives. Yet it immediately raises an old question: who decides how the data are collected? Who can download them? Who may train models? Who will profit in the future? If the answers remain only researchers and platforms, then "preservation" may quickly become another form of extraction.
Low-resource languages are not cheap data. This sentence must be written on the first page of every Indigenous language AI project, preferably also posted on the login screen of each annotation platform. Language is not fuel waiting to be consumed by models nor material for papers when research projects end. Language is the sum of people and place, ancestors, taboos, humour, shame, intimacy, power and ways of classifying the world. If AI is to speak for Indigenous languages, it should first learn manners.
Corpora are not warehouses; they are governance systems
Building a corpus is like constructing a house. Outsiders only notice whether the roof looks pretty; what truly matters are the foundation, pipes, keys and who has the right to enter. Language data are especially so. Recordings may contain elders' voices, ritual vocabulary, place-name memories, kinship terms, jokes, ways of scolding, and also narratives unsuitable for public release. Throwing them all under the banner of "open data" seems progressive but can be extremely crude.
For Taiwan's Indigenous language technology this point is especially critical. Taiwan has rich experiences in revitalizing Indigenous languages, Romanized writing systems, teaching platforms and demand for voice data, yet also suffers from scattered data, unclear licensing, insufficient community feedback and risks of model leakage. If AI wants to enter Indigenous languages it should not only ask "can we train on this?" but rather "what problems do the people want solved?" Is it to assist elders in oral transcription? To support teachers' lesson planning? To let young people query vocabulary? Or to build a knowledge assistant controllable within the community? Different purposes require different data boundaries.
Data governance is not about making consent forms longer. True governance must address the lifecycle: how to explain before collection, who should be present during collection, who can review after collection, how data are classified, who may withdraw, how models are constrained, how errors are corrected, how results return to the community, and whether commercial use is prohibited or renegotiated. Without these arrangements a corpus is merely a warehouse that looks noble; the larger the warehouse, the greater the risk.
Stepping back from technical problems to power questions
Low-resource language engineering often lists challenges: little data, inconsistent standards, spelling variation, poor audio quality, lack of annotation manpower. These are real but insufficient. The deeper issue is how power is distributed. If a corpus enables external researchers to publish papers, platforms to train models and companies to build products while the community receives only a thank-you certificate, that is not Two-Eyed Seeing; it is a polished one-way transfer.
A better approach designs governance from the start of the data lifecycle: clear consent forms, revocable mechanisms, tiered authorization, marking sensitive content, community review, feedback tools, benefit or outcome sharing, and model usage limits. Some data can be public, some only usable within the community, some should not be digitized at all. The AI world often treats "more data is better" as common sense; cultural data sovereignty reminds us that sometimes less data with clearer permissions is more responsible.
高德生 has long reminded from Tsou ceremonial ecology, hunter schools and cultural heritage work that knowledge never floats alone; it always has a place, a season, an audience and responsibilities. This reminder matters for AI. Models excel at breaking context apart: turning voice into text, text into vectors, vectors into answers. But cultural knowledge often cannot be judged by content alone; one must also consider when to speak, who may speak, to whom, and where to stop. If technology does not know how to stop, manners do not exist.
Do not turn taboos into search functions
What Indigenous language technology needs most to beware of is not weak models but overly capable ones. Many systems strive for "answer everything" to showcase ability. Yet in cultural knowledge answering everything may be rude rather than clever. Certain ritual details, hunting-ground knowledge, family stories, place-name contexts and taboo words are not always suitable for public query. Even if data already exist in some texts it does not mean they can be repackaged as instant chatbot answers.
This is not anti-technology; it demands technology acknowledge boundaries. A good Indigenous language AI can say "I cannot provide such content; please return to community authorization and elders' guidance." It can mark "this answer is for general language learning purposes only, not representing ritual or family context." It can exclude sensitive data from training sets or use them only in closed, community-governed environments. More importantly it must admit it is not an elder, priest, hunter or tribal council. AI may assist but should never impersonate cultural authority.
If Taiwan's Indigenous language revitalization enters the AI stage it should first build this humility rather than rush to show models can sing, translate and tell stories. Being able to speak does not mean one should speak; sounding like a speaker does not make one right. The core of cultural data sovereignty is not locking data away but empowering communities to decide the form, speed and purpose of opening.
Let language technology learn manners first
A truly useful Indigenous language AI need not start by writing poetry, debating or imitating elders. It can first master simple things: listen clearly, transcribe accurately, mark sources, admit ignorance, avoid random taboo translations, do not mix different languages into one pot. These seem basic yet are harder than showmanship because they require the people behind the model to respect language community order.
The minority language corpus project in Bangladesh offers an international reference case; Taiwan need not copy it exactly. What Taiwan truly needs is to place Indigenous language technology within tribal, school, family and public service contexts. Low-resource languages are not weaklings waiting for AI salvation; they are teachers reminding the AI industry to learn manners again. Language is not a data mine; voice is not free fuel. If models want to speak, they should first know which piece of land they stand on.
Conclusion: governance must arrive before models open their mouths
The technical challenges of low-resource NLP deserve investment because language tools can indeed help learning, preservation and public services. The question is not whether to do it but how. Corpora without data sovereignty may become new colonial archives; models without community consent may be external powers that speak Indigenous languages; search functions without taboo boundaries may turn cultural dignity into product experience.
The first lesson of Indigenous language AI is not transformer, embedding or RAG but manners. Manners are not mere politeness; they are knowing one cannot speak for others, recognizing data are not picked up freely and understanding some knowledge must return to relationships rather than be broken into parts. When the AI industry accepts this lesson low-resource languages will no longer be processed cheaply.
Further reading and sources
AI use and content-safety disclosure
This article was assisted by AI for data organization, structural drafting and sentence polishing; human editors set the viewpoint and fact-checking direction