AI Learned What Indigenous People Look Like, But Who Authorized the Lesson? The UN Pushes Data Sovereignty Into the Generative AI Era
Original Chinese title: AI學會了「原住民長什麼樣」,但誰准它學?聯合國把資料主權問題推進生成式AI時代
A UN article, the GIDA 2026 AI communiqué, and a Human Rights Council side event move the question beyond whether AI can generate outputs toward access, consent, governance, reciprocity, and the right to say no.
雙向知識實驗室
A knowledge exchange lab focused on Indigenous knowledge, cultural data sovereignty, artificial intelligence applications, and two-way knowledge translation.

# AI Learned What Indigenous People Look Like, But Who Authorized the Lesson? The UN Pushes Data Sovereignty Into the Generative AI Era
Begin with a Stereotyped Image
A United Nations article begins with generated images of Indigenous Peoples and notes that models may repeat feathered headdresses, beadwork, and travel-brochure figures. This is not merely an aesthetic failure; it is about how data compresses diverse communities into a few symbols. After asking what AI learned, ask who supplied data, who defines accuracy, and who can demand a stop.
Artificial Intelligence Depends on Existing Data
Generative artificial intelligence does not learn from a blank slate. It depends on text, images, languages, maps, archives, and user feedback. The UN article notes that Indigenous lands, knowledge, languages, and bodies have long been studied and governed by outsiders. If new systems repeat extraction, technology may only turn old power relations into a faster interface.
Data Sovereignty Asks Who Has Authority
Data sovereignty is not only an individual checkbox. It concerns community authority over how data is collected, stored, accessed, interpreted, and used. The UN article links this to self-determination, while GIDA’s communiqué brings Indigenous Data Sovereignty and Indigenous Data Governance into artificial intelligence. Both require institutions to treat communities as rights-holders, not only respondents.
Consent Must Travel Through the Life Cycle
The Human Rights Council side-event description places risk at the intersection of data extraction, artificial intelligence, and neurotechnology, and emphasizes that consent and governance cannot appear only at collection. Training, labeling, fine-tuning, deployment, evaluation, updating, reuse, and public output can all change risk. One signature does not authorize every future use.
CARE and Local Contexts Are Not Universal Fixes
CARE principles and Local Contexts labels can articulate collective benefit, authority, responsibility, and local context, but they do not automatically provide consent, representation, reciprocity, security, or legal accountability. A label applied unilaterally by an outside institution can become decoration. Governance needs community decisions, refusal, correction, traceability, and resources to operate.
RAG Systems Must Also Know How Not to Answer
Retrieval-augmented generation is often described as safer because it retrieves trusted material. But if an index contains culturally restricted content or mixes distinct communities, a citation does not make use legitimate. Permissions must operate before retrieval; refusal and partial answers must be testable; versions, purposes, retention, and withdrawal state must reach the output layer.
Break the Data Flow into at Least Eight Questions
In practice, ask at least eight questions: how was data acquired, who consented, who labeled it, who trained on it, who can query it, how is output presented, how do benefits return, and how can use be withdrawn? Each needs an owner and verifiable record. Compressing them into an ethics review or open-data claim hides the transfer of power.
The Output Is Not Finished on the Screen
An image, voice, or summary may enter classrooms, exhibitions, advertising, databases, and another model. Outputs should identify generation, provenance, and uncertainty, without treating model guesses as community self-description. For people represented, the ability to see, challenge, and request correction is governance, not an optional post-publication service.
Archives Still Carry Historical Responsibility
A museum or archive may legally hold material without every reuse being automatically legitimate. Digitization increases visibility while amplifying misnaming, colonial classification, and unlimited downloading. AI projects should check acquisition history, community protocols, and possibilities for return, allowing description and access rules to be rewritten together rather than feeding old metadata to a model.
The Right to Refuse Must Be Real
Refusal is not a policy slogan; it is a state in the data workflow. Communities need to know which use stops, what happens to embeddings, caches, backups, and partner copies, who responds, and when. If complete deletion is technically impossible, state the limitation honestly rather than creating false control with the phrase withdrawable.
Reciprocity Is More Than Acknowledgment
When community data creates value for a model, study, or product, reciprocity may include funding, infrastructure, language tools, training, data copies, joint authorship, or long-term governance roles. A one-time acknowledgment or a name on a website does not answer who benefits and who bears risk. Communities should also be able to refuse a form of benefit they do not want.
Community Governance Must Reach the Decision Table
UN sources and the side event place Indigenous leadership at the center of governance. Community representatives should participate in standards, risk classification, access, audits, procurement, and stop decisions, not only review a finished model. Representation cannot rest on one adviser; disclose authorization, terms, conflicts, and how disagreements among communities are handled.
Some Knowledge Should Not Be Digitized
A genuinely inclusive artificial intelligence future does not require every cultural material to become searchable. Some knowledge is shared only through particular relationships, seasons, places, or ceremonies; some is restricted to certain identities. Recognizing what should not be digitized is not anti-technology. It treats refusal and protection as knowledge governance.
Taiwan Should Begin with Processes, Not Slogans
Taiwan’s language, image, community-research, and museum data have distinct institutional contexts; international principles are not completion certificates. Start with one data set, map consent, access, correction, reciprocity, and exit with the relevant community, then align engineering, legal, and procurement documents. This reveals accountability better than claiming compliance with CARE.
Users Also Carry Responsibility
Everyday users cannot place all bias on the model. Before generating, sharing, or reusing, check whether a prompt reduces a people to costume, whether permission is needed, and whether generation is disclosed. When communities respond, be willing to stop and correct. Each publish button can give a stereotype another route into circulation and training.
Make Refusal and Reciprocity Auditable
Governance can become auditable terms: data sets record provenance and permissions, model cards state exclusions, systems log query approval, outputs carry context and a challenge route, and contracts define reciprocity and withdrawal. These details do not solve power by themselves, but they show communities where a problem occurred and who must act.
Data Minimization Is Also Respect
Not everything that can be collected is worth collecting. If a model needs language examples but retains identity, location, family, and images as well, unnecessary risk grows. Data minimization should be decided with communities and implemented through fields, permissions, retention, and deletion tests, not only an ethics statement.
Representation Is Not Filling a Quota
A face, a few language samples, or a community name in a data set does not mean the community is represented correctly. Ask who was not invited, whose dialect was removed, whose difference was averaged away, and who can correct errors. AI fairness is not only data volume and accuracy; it includes community authority over categories, refusal, and reciprocity.
Let Governance Records Travel with the Model
Data sets, embeddings, prompt templates, and outputs can leave the original collaboration. If governance records live only in a proposal appendix, the next developer will not know the limits. Put provenance, permission, sensitivity, purpose, expiry, and a community contact into machine-readable fields and human-readable notes, checking them again at deployment, update, and transfer.
Continue asking by role
- Community data steward: Ask about evidence, limits, and feasible action.
- AI and RAG developer: Ask about evidence, limits, and feasible action.
- Museum and archive worker: Ask about evidence, limits, and feasible action.
- Everyday AI image user: Ask about evidence, limits, and feasible action.
Sources and further reading
Role-Driven Inquiry
Yuan Media AI | Guided Questions by Role
Select a role first, then review the complete questions from that perspective; clicking a question will display Yuan Media AI’s response in real time.
Select a role
Community data steward
Examine the article's evidence, governance boundaries, data rights, and actionable steps from the perspective of a Community data steward.
Choose a question
AI use and content-safety disclosure
This article draws on official and research sources. Evidence, limitations, and analysis are distinguished.