Who Turned Ancestors into Datasets? — Cultural Data Sovereignty in the AI Era, Not Cloud Storage Permissions
Original Chinese title: 誰把祖靈變成資料集?——AI時代的文化資料主權,不是雲端硬碟的權限設定
The AI industry is accustomed to breaking down the world into trainable data, as if files that can be downloaded, text that can be extracted, and images that can be labeled naturally become food for models. But for Indigenous communities, knowledge is not merely information, and data is not merely a resource; it connects to land, kinship, care, taboos, responsibilities, and collective futures. Drawing on Indigenous Data Sovereignty and the CARE Principles, this article argues that AI governance cannot treat ancestral knowledge as free cloud fertilizer.
Yuan Media AI Editorial Desk

# Who Turned Ancestors into Datasets? — Cultural Data Sovereignty in the AI Era, Not Cloud Storage Permissions
From CARE Principles to community consent, re-examining whether AI can touch ancestral knowledge
Author: Yuan Media AI Editorial Desk
Editorial Preface
The AI industry is accustomed to breaking down the world into trainable data, as if files that can be downloaded, text that can be extracted, and images that can be labeled naturally become food for models. But for Indigenous communities, knowledge is not merely information, and data is not merely a resource; it connects to land, kinship, care, taboos, responsibilities, and collective futures. Drawing on Indigenous Data Sovereignty and the CARE Principles, this article argues that AI governance cannot treat ancestral knowledge as free cloud fertilizer.
Not Just “Data”: How a People Remembers Itself
AI companies love the word “data.” These two words are clean, convenient, odorless—like a whiteboard freshly wiped in a conference room. But before cultural knowledge enters a database, it is often not files; it is responsibilities between people: who may speak, when to speak, to whom, and what obligations follow after speaking.
Flattening these relationships into datasets raises issues beyond technical accuracy—it concerns governance positioning. Once data is stripped of context, the first thing that disappears is judgment. Models see fragments but cannot perceive why communities keep certain matters silent. The black humor is that the tech industry often claims models need “more context,” yet excels at dismantling human context first.
Ancestral Data Is Not a Public Stockpile
What appears online does not mean it can be trained; past research citations do not mean platforms may commercially use them; museum photos of exhibits do not equal ancestral consent to become generative AI style packs. If these boundaries are not clarified, AI will misread “public” as “ownerless.”
Indigenous Data Sovereignty reminds us that data rights belong not only to individuals but also to families, communities, nations, and future generations. This is not anti-technology; it demands technology recognize: collective memory requires collective governance. If models ask only about license terms without inquiring about relationships and consequences, they mistake a legal notice for an ethical framework.
CARE Principles: From Consent to Collective Benefit
The CARE Principles focus on Collective Benefit, Authority to Control, Responsibility, and Ethics. They do not seek to lock data away nor banish AI from the gates; they require that data use serves community interests, acknowledges community control, holds users accountable, and places ethics before convenience.
Applied to AI training, this raises concrete questions: Is model purpose clear? Can communities refuse or opt out? How are errors corrected? How is benefit returned? Which knowledge should not be collected? If the answer is “we’ll decide later,” it usually means “we want to take first.”
The AI Black Box Cannot Open Backdoors for Cultural Context
AI systems often evade scrutiny under the guise of black boxes: data too large, models too complex, supply chains too long, responsibilities too dispersed. Yet once cultural knowledge is trained into a model, it may be recombined, misattributed, and commodified across countless outputs, leaving communities to laboriously explain afterward that this was not their intent.
For sensitive knowledge, the best protection may not be more precise watermarks but simply not putting it into models. Not all knowledge should be disseminated; not every memory needs searching; not every trace left by ancestors must become “creative material.” Some boundaries exist not out of backwardness, but because they embody practices of care.
Indigenous Data Governance Is Not Anti-Technology
Mature AI governance does not treat community consent as an obstacle but as part of quality. Data without consent is not only ethically risky—it is unreliable knowledge. Innovation without feedback mechanisms is merely a colonial economy in cloud clothing.
Communities can collaborate with researchers, schools, and platforms to establish tiered licensing, usage restrictions, data review processes, shared benefit arrangements, and error-correction systems. The point is not to keep AI forever away from cultural data but to ensure every encounter is accountable. If AI wishes to approach ancestral knowledge, it must first learn how to knock.
Conclusion: Return Data to Relationships; Do Not Let Models Swallow It
What we need is neither locking all cultural knowledge into safes nor dumping all data into models. The more demanding path is to recognize that knowledge has stewards and relationships, that some parts should remain private, and that others may be shared when doing so serves community interests.
In the AI era, talking about cultural data sovereignty ultimately asks an old question: who has the right to explain a group’s memory? If the answer always remains platforms, research institutions, and capital, ancestors will not become more visible—they will simply turn into another batch of training data with invisible sources. That is not preservation; that is swallowing.
AI use and content-safety disclosure
This article was compiled and edited through Yuan Media AI editorial workflow.