Don't Turn Ancestors into Datasets: Who Has the Right to Say 'No' Before Traditional Knowledge Enters AI?
Original Chinese title: 不要把祖先變成資料集:傳統知識進入 AI 之前,誰有權說「不」?
One of the most dangerous misunderstandings in the AI era is equating 'can be digitized' with 'can be used for training.' Images, songs, plant knowledge, ritual vocabulary, Indigenous language stories, and elder interviews, once entered into databases, can easily be repackaged by platforms, models, and research projects as 'open data.' But traditional knowledge is not raw material lying on the internet; it has context, relationships, taboos, authorizations, and the right to refuse use. This article argues that Indigenous Peoples AI should not merely make culture visible but establish a data governance system capable of saying 'yes, pause, no'.
全明正
Long-term focus on ecological documentation, cultural transmission, and knowledge governance; emphasizes the subjectivity of traditional knowledge in modern technological environments.

In the AI era, there is a phrase that is both fascinating and dangerous: 'We can digitize traditional knowledge.'
The first half of this sentence is beautiful. Digitization can preserve language, organize images, allow younger generations to rediscover ancestral wisdom, enable Indigenous people living in cities to access tribal memories through platforms. It can also assist teaching, research, exhibition, and intergenerational communication. The problem lies not with digitization itself but with the second half that is often quietly appended: 'Since it has already been digitized, shouldn't it then be usable for training AI?'
This is where the danger begins.
Traditional knowledge is not a pile of data waiting to be tagged. It is not free ore lying in cloud storage nor premium feedstock for technology models. It has sources, storytellers, seasons, contexts, families, taboos, and conditions of use. Some knowledge can be taught publicly; some must only be spoken at specific occasions; some require particular identities to access; some are best preserved by not being platformed at all.
But AI systems do not understand these nuances. Models excel at flattening differences into calculable patterns. When they encounter songs, they want to convert them to audio features; when they see plants, they want image classification; when they read elder stories, they want text corpora; when they observe rituals, they want 'cultural content.' Thus a living knowledge system may be dismantled within data processing pipelines into filenames, tags, fields, vectors, and APIs. Ancestors do not become stars; they become embeddings. This is harsh, yet today's technical workflows can indeed proceed this way.
So we must reframe the question: instead of asking 'Can traditional knowledge enter AI?', ask 'Who has the right to decide whether it enters AI?'
This is at the heart of Indigenous Peoples data sovereignty. Data sovereignty is not merely copyright, nor does signing a consent form settle everything. It involves collective rights, cultural responsibilities, governance procedures, and control over future use. Tribal data do not belong permanently to researchers once obtained, nor become public goods simply by uploading to platforms. Public does not mean ownerless; open does not mean rights abandoned.
The CARE principles often cited internationally offer an important reminder. Data governance cannot pursue only FAIR—findable, accessible, interoperable, reusable—but for Indigenous Peoples data must also adhere to CARE: Collective benefit, Authority to control, Responsibility, and Ethics. Simply put, data should be useful but used correctly; shareable but with awareness of who benefits; trainable models but with authority to halt.
For Yuan Media AI this has direct implications. If future chatbots are built incorporating images, voice, Indigenous language, plant knowledge, traditional crafts, and tribal narratives, they must involve not just database engineering but governance engineering. Technical architecture should record authorization conditions; data fields must mark sensitivity levels; model training should exclude specific datasets; outputs must avoid reproducing taboo knowledge erroneously. More importantly, tribes and knowledge holders should be able to know how their data is used, where, what responses it generates, and even withdraw or restrict access.
Does this sound complicated? Of course. Yet culture does not exist merely for the convenience of data engineers. If an AI system can only operate by ignoring cultural rules, the problem lies not in culture's complexity but in the system's rudeness.
Taiwan has a chance to forge its own path here. The integration of Indigenous traditional knowledge and digital technology need not replicate large platforms' script of 'collect first, apologize later, then form an ethics committee.' A better approach is to establish data classification from the outset, tribal authorization, feedback loops for knowledge holders, usage records, and protection for sensitive content.
AI should not become a new colonial mining tool but rather assist communities in managing memory, educating younger generations, and connecting cross-disciplinary knowledge.
This also requires Two-Eyed Seeing. Scientific data and traditional knowledge are not substitutes; they must see each other's limits. AI can help organize vast images, voice recordings, and texts, yet cannot replace elders' judgment nor tribal contextual mastery. Models may identify plant leaf shapes but might not know which story segments should never be spoken; they can compile Indigenous language vocabularies but may lack awareness of certain terms' weight within ritual contexts.
A truly mature Indigenous Peoples AI is not the one that answers the most questions, but knows which questions must remain unanswered, which content needs to return to the community, and which knowledge must stay silent.
In a data economy, 'silence' is often seen as waste—unused means no commercial value. Yet in traditional knowledge, silence can be protection, courtesy, right, and governance wisdom. If AI enters this domain, it must learn: not everything should become data; not all data should train models; not every answer should be generated.
Ancestors may be remembered but should not be mined.
Traditional knowledge may be transmitted but should not be swallowed by platforms.
AI can be clever, yet it must first learn to respect others' right to say 'no'.
AI use and content-safety disclosure
This article was compiled and edited through Yuan Media AI editorial workflow.