跳转至

文章背景与核心概要

尽管以人工智能驱动的教育工具经常被宣传为弥合全球教育获取差距的可扩展解决方案,但本文指出,这些系统的底层基础设施(从训练语料库到部署架构)系统性地边缘化了代表性不足语言的使用者。

以孟加拉语(Bengali)为案例研究,作者展示了结构性障碍如何在模型训练之前就已长期存在。他们指出了四个相互交织的关键失效点,这些失效点为这些语言社群制造了“结构性沉默”:1. 网络存在感差距;2. 训练词元(Token)赤字;3. 词元化惩罚;4. 连接性排斥。作者得出结论,数据集的稀缺不仅仅是一个技术障碍,更是机构优先事项和设计默认值的结果。他们倡导采用“离线优先”的设计策略来实现公平,并呼吁转变语言学和人工智能研究方向,以解决这些系统性不平等问题。


Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Authors: Avijit Roy, Proma Roy
Date: August 12, 2026
Subject: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
DOI: 10.48550/arXiv.2608.12278

Authors: Avijit Roy, Proma Roy
Date: August 12, 2026
Subject: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
DOI: 10.48550/arXiv.2608.12278


Summary

Summary

虽然以人工智能驱动的教育工具经常被宣传为弥合全球获取差距的可扩展解决方案,但本文指出,这些系统的底层基础设施(从训练语料库到部署架构)系统性地边缘化了代表性不足语言的使用者。

While AI-driven educational tools are often marketed as scalable solutions to bridge global access gaps, this paper argues that the underlying infrastructure of these systems—ranging from training corpora to deployment architectures—systematically marginalizes speakers of underrepresented languages.

孟加拉语(Bengali)作为案例研究,作者展示了结构性障碍如何在模型训练之前就已长期存在。他们指出了四个相互交织的关键失效点,这些失效点为这些语言社群制造了“结构性沉默”:

Using Bengali as a case study, the authors demonstrate how structural barriers exist long before a model is even trained. They identify four critical, interlocking failures that create a "structural silence" for these linguistic communities:

  1. 网络存在感差距: 尽管孟加拉语使用者占全球总人口近 4%,但孟加拉语在全球网络内容中所占比例却不到 0.5%。
  2. 训练词元赤字: 在主流多语言语料库中,英语词元与孟加拉语词元的比例高达惊人的 67:1。
  3. 词元化惩罚: 孟加拉语独特的元音附标文字(alphasyllabary)脚本遭遇了更高的词元丰度(token fertility),这加剧了本就存在的数据稀缺问题。
  4. 连接性排斥: 严重的数字鸿沟依然存在,农村地区的互联网普及率(36.5%)远远落后于城市地区(71.4%)。
  1. Web Presence Gap: Despite representing nearly 4% of the global population, Bengali accounts for less than 0.5% of global web content.
  2. Training-Token Deficit: There is a staggering 67:1 ratio between English and Bengali tokens in major multilingual corpora.
  3. Tokenization Penalty: The unique alphasyllabary script of Bengali suffers from higher token fertility, which exacerbates the existing data scarcity.
  4. Connectivity Exclusion: A significant digital divide persists, with rural internet penetration (36.5%) trailing far behind urban areas (71.4%).

作者得出结论,数据集的稀缺不仅仅是一个技术障碍,更是机构优先事项和设计默认值的结果。他们倡导采用“离线优先”的设计策略来实现公平,并呼吁转变语言学和人工智能研究方向,以解决这些系统性不平等问题。

The authors conclude that dataset scarcity is not merely a technical hurdle but a consequence of institutional priorities and design defaults. They advocate for an "offline-first" design strategy as a means of achieving equity and call for a shift in linguistics and AI research to address these systemic inequalities.


Access the Paper

Access the Paper


License & Metadata

License & Metadata

license icon 查看许可协议 (View License (CC BY 4.0))

license icon View License (CC BY 4.0)

  • 评论: 相关海报于 2026 年 4 月 30 日至 5 月 2 日在纽约州纽约市举行的国际语言学协会第 69 届年会(ILA 2026)上展示。
  • 提交历史: [v1] 2026年8月12日 星期三 17:17:25 UTC。
  • Comments: Associated poster presented at the 69th Annual Conference of the International Linguistic Association (ILA 2026), New York, NY, April 30-May 2, 2026.
  • Submission History: [v1] Wed, 12 Aug 2026 17:17:25 UTC.