Ancestral Spirits Are Not Cloud Storage: Who Gets Backed Up First Before Traditional Knowledge Is On-Chain?
Original Chinese title: 祖靈不是雲端硬碟:傳統知識上鏈前,誰先被備份?
When Indigenous community knowledge is scanned, tagged, uploaded, and trained, is it being preserved or re-colonized? The AI industry loves to call data the new oil, but Indigenous experience reminds us that someone else's data are actually called ancestral spirits, kinship, and responsibility.
山海資料庫筆記

The most common opening line for digital curation is: 'If we don't preserve it now, it will disappear.' Then the scanners open, recorders boot up, field tables are laid out, and knowledge gets sliced into names, places, species, rituals, stories, and keywords. Preservation certainly matters, but more urgently we should ask: who decides what to preserve, where to store it, who can search for it, and after preservation, who still has the right to say no?
If these questions go unanswered, backup becomes the quietest form of colonization. In the past they took physical objects; now they copy files. Ships used to dock at ports; today an API connection is enough. Efficiency has gone up, but courtesy hasn't necessarily followed.
The Quietest Colonization Is Usually Called 'Backup'
For platforms, backup is risk management; for communities, what gets backed up may be ancestors' voices, family responsibilities, stories that can only be told at certain seasons, and knowledge that shouldn't leave a particular relational network. Having an extra copy of the file doesn't mean the community has more control. In fact, the easier replication becomes, the harder it is to track where things go.
'We're just preserving' cannot automatically absolve responsibility either. If data later gets handed off to research institutions, fed into models, or turned into products, does the original consent still hold? If preservers can keep changing purposes while providers only signed once a decade ago, that's not collaboration—it's more like an ever-valid cultural ATM card.
Data Is Not Neutral, Especially When It Comes From Colonized Peoples
Data looks like objective record-keeping, but every field hides the classifier's worldview. Who gets written as 'interviewee' and who is called 'expert'; which content is tagged as folklore and which is recognized as knowledge; which place names follow official versions and which Indigenous-language terms are flagged by systems as spelling errors—all reflect power choices.
Indigenous data sovereignty therefore asks not only about individual privacy but also how Indigenous Peoples, communities, families, and future generations can govern data collectively. The CARE principles—collective benefit, authority to control, responsibility, and ethics—are precisely reminding data workers that being able to take doesn't mean you should; being able to analyze doesn't give you the right to draw conclusions for a community.
Blind Spots in Open Data: Who Is Asked to Open? Who Profits?
Open data is often described as a shortcut to democracy and innovation, but not all data starts on an equal footing. Resource-rich companies can download vast amounts, cross-reference them, train models, then sell back curated services to public agencies; the communities providing knowledge may never know where their data ends up.
More absurdly, communities are often asked to prove 'why it shouldn't be open,' while corporations rarely have to justify 'why they can profit.' When openness is only transparent for the weak and profits concentrate in capital hands, so-called public goods become a one-way door.
Licensing Is Not Signing—It's Relational Governance
True licensing shouldn't stop at consent forms. It must include purpose, duration, who may access it, conditions for reuse, withdrawal mechanisms, error correction, benefit sharing, and dispute resolution. More importantly, communities need the ability to renegotiate when contexts change, rather than being bound forever by an old spreadsheet uploaded to the cloud.
Some knowledge can be public; some fit only within a community; some require family or specific role consent; some shouldn't be digitized at all. These layers aren't database design headaches—they're realities databases must respect.
AI Model Black Boxes Make Traditional Knowledge Harder to Return Home
Once traditional knowledge enters large models, sources, contexts, and limits can get scattered. Models may generate a fluent cultural explanation but might not say where the knowledge came from, whether it's open, or who should receive benefits. Erroneous outputs can then be cited by search engines, textbooks, and media—false data gets dressed up in suits while true knowledge is asked to show ID.
Therefore, the most reliable protection for sensitive knowledge sometimes isn't finer-grained tagging but keeping it out of general-purpose models altogether. Technical feasibility should never be swapped for ethical permission.
Indigenous Community Data Trusts and Tiered Knowledge Repositories
Viable directions include community-led data trusts, tiered repositories, and purpose-review systems. Data can be handled at levels such as public, restricted, community-internal, family-managed, or prohibited from digitization; external researchers must state their aims, storage duration, and how results will benefit the community, while AI training requires separate explicit consent.
Systems also need to let communities see access logs, request deletion of copies, correct erroneous descriptions, and decide how revenues return to language work, education, care, and cultural practice. Sovereignty isn't locking data away—it's ensuring keys aren't held only in someone else's server room.
Conclusion: Not Anti-Technology, but Against Treating Knowledge as Ownerless
Opposing unrestricted extraction doesn't mean opposing digital curation; demanding community control isn't rejecting research or collaboration. What is truly rejected is treating knowledge that has owners, relationships, and responsibilities as ownerless raw material.
Ancestral spirits are not cloud storage, and Indigenous communities are not free data centers. For preservation to have meaning, it must protect not just files but the community's right to say yes, no, or 'not now,' and to decide how knowledge returns home.
AI use and content-safety disclosure
This article was generated as a draft by Yuan Media AI's daily article factory, then reviewed by human editors; content involving Indigenous traditional knowledge, data sovereignty, and cultural governance still requires ongoing correction within community contexts.