When AI Starts Translating Ancestors' Dreams: Can Large Language Models Preserve Endangered Languages?
Original Chinese title: 當AI開始翻譯祖先的夢:大型語言模型能保存瀕危語言嗎?
Large language models entering endangered language preservation may appear as technological benevolence, but they actually implicate corpus licensing, cultural context, and knowledge sovereignty. Language is not a dictionary or clean data that can be fed into a model; it connects to land, dreams, kinship terms, ancestral memory, and community responsibility. This article asks from the Two-Eyed Seeing perspective: Is AI assisting language revitalization, or turning living languages into beautiful digital specimens?
Ma Le Ve

# When AI Starts Translating Ancestors' Dreams: Can Large Language Models Preserve Endangered Languages?
From language revitalization to digital colonialism, a tug-of-war over memory, power, and algorithms
Author: Ma Le Ve / Indigenous University Advocate / Postdoctoral Researcher / Rukai
Editorial Preface
Large language models entering endangered language preservation may appear as technological benevolence, but they actually implicate corpus licensing, cultural context, and knowledge sovereignty. Language is not a dictionary or clean data that can be fed into a model; it connects to land, dreams, kinship terms, ancestral memory, and community responsibility. This article asks from the Two-Eyed Seeing perspective: Is AI assisting language revitalization, or turning living languages into beautiful digital specimens?
From "preservation" a step back: What exactly do we want to preserve?
When tech companies say AI can preserve endangered languages, that phrase sounds lovely. Like an algorithmic archaeological team carrying GPUs into the mountains, leaving voices for ancestors and retrieving mother tongues for children.
But here's the problem: Can language really be "preserved"?
If preservation means stuffing words, sentences, recordings, and translations into a database, then language becomes a specimen. Specimens are clean and quiet—so quiet that no child mispronounces at dinner, no elder corrects with a smile, and no word grows different meanings across families, mountain paths, or ritual occasions.
Language is not a database. Language is a living worldview.
The allure of large models: It speaks, but does it understand?
The most captivating thing about large language models is that they seem to know everything. Feed them an Indigenous language passage and they may translate; give them Romanized text and they might guess Chinese; ask them to design curriculum materials and they can even generate lesson activities.
The problem is that the hardest part of language revitalization isn't "producing more sentences"—it's whether those sentences stand in the right cultural relationships.
The same word may mean entirely different things in daily life, ceremonial contexts, dreams, kinship terms, and community histories. AI can calculate distances between words but may not understand why an elder pauses; it can mimic tone but doesn't know which words must never be spoken casually. Worse, sometimes it errs with such fluency that mistakes look like truth once typeset—tech's version of black humor: errors formatted suddenly become authoritative.
Two-Eyed Seeing: one eye on tools, the other on relationships
Two-Eyed Seeing doesn't call for rejecting AI nor sealing traditional knowledge in glass cabinets. It reminds us: one eye sees what technology can do; the other attends to the relationships that make language meaningful.
AI can assist with speech recognition, pronunciation practice, draft materials, vocabulary retrieval, and corpus organization—important work, especially where teacher staffing is thin, materials are scattered, and children have limited learning time. Tools truly lighten burdens.
But tools cannot replace relationships. Language revitalization requires families, schools, communities, churches, youth groups, elders, and institutions acting together. If AI merely takes the corpus away, leaves the model behind, writes reports, and closes funding cases, that isn't revitalization—it's digital collection.
Corpus sovereignty: "public" does not equal "usable"
Many Indigenous language materials appear public online: textbooks, videos, songs, dictionaries, teaching audio, research papers. Yet public doesn't mean ownerless; educational use doesn't automatically allow commercial models to train on them.
Indigenous data governance emphasizes collective interests, community control, responsibility, and ethics. Applied to AI language models, these principles become concrete questions: Who consents to corpus collection? Who can see training results? Who corrects errors? How is commercial value returned to communities? Which corpora must never enter the model?
Without answering these first, AI preservation of languages easily becomes another form of digital colonialism. Past colonialism drew maps over land; today platforms may slice language into tokens. Different names, similar logic.
Don't turn living language into beautiful specimens
The future worth anticipating isn't "one day AI speaks Indigenous languages better than community members." That would be the most absurd victory. Success in language revitalization means children willingly speaking with grandparents, young people naturally using mother tongues on short videos, Indigenous language teachers no longer fighting alone, and communities deciding how their data is preserved and used.
AI can be a tool but not an owner. It may help organize memory, cannot replace the owners of that memory; it may assist pronunciation practice, cannot decide what counts as correct culture.
The key to preserving endangered languages isn't teaching models ancestral speech—it's restoring living people's right to speak.
AI use and content-safety disclosure
This article was compiled and edited through Yuan Media AI's editorial workflow.