跳转至

文章背景与核心概要

尽管像 GPT-4 和 Claude 这样的大语言模型(LLM)能够展现出令人惊叹的流畅表达能力,但它们是通过消耗数万亿词汇的语料来实现的——这一数据规模比人类儿童接触到的多出数千倍。这种“数据效率鸿沟”凸显了一个巨大的谜团:儿童如何在接触极少语言材料的情况下,就掌握了复杂的递归语言。目前,研究人员正在利用“婴儿规模(Baby-scale)”的 AI 模型来对这一过程进行逆向工程,希望借此开发出更高效率的 AI、保护濒危少数民族语言,并最终解答语言习得究竟是天生的生物学本能,还是统计学习的奇迹。


Kids Outlearn AI—and We Still Don’t Know Why

Summary

While modern Large Language Models (LLMs) like GPT-4 and Claude can mimic human fluency, they do so by consuming trillions of words—a scale of data thousands of times greater than what a human child encounters. This "data efficiency gap" highlights a profound mystery: how children master complex, recursive language with minimal exposure. Researchers are now using "baby-scale" AI models to reverse-engineer this process, hoping to create more efficient AI, preserve minority languages, and finally answer whether language acquisition is an innate biological instinct or a feat of statistical learning.


数据效率鸿沟

在长达十万年的时间里,人类是唯一能够完美流利使用语言的物种。如今,人工智能也加入了这一行列,但代价高昂。一个儿童在 12 岁之前大约听到 1 亿个词,而像 Llama 3.1 这样的现代大语言模型则是在 15 万亿个 Token(词元)上进行训练的。如果将用于训练大语言模型的文本打印出来,其长度将超过国际空间站的高度,而儿童的语言“输入”堆叠起来只有 20 米高。

认知科学家迈克尔·C·弗兰克(Michael C. Frank)指出了这种讽刺现实:“我们依然不得不砍伐森林、搜罗全人类知识的总和,才能重新复现这个在我们的客厅里用一年时间就能完成的奇迹。”

The Data Efficiency Gap

For 100,000 years, humans were the only entities capable of perfect language fluency. Today, AI has joined that club, but at a staggering cost. While a child hears roughly 100 million words by age 12, a modern LLM like Llama 3.1 is trained on 15 trillion tokens. If printed, the data used to train an LLM would reach past the International Space Station, whereas a child’s linguistic "input" would stack only 20 meters high.

Cognitive scientist Michael C. Frank notes the irony: "We still have to burn down a forest and scrape the entire sum of all human knowledge to re-create this milestone that happens in our living rooms over the course of a year."

语言学理论的演变

几十年来,关于儿童如何学习语言的争论几经转变: * 乔姆斯基学派观点: 在 20 世纪 50 年代,诺姆·乔姆斯基(Noam Chomsky)提出,语言过于复杂,仅靠统计学是无法学会的。他认为人类天生带有“语言本能”或内置的语法规则。 * 统计转向: 2010 年代神经网络和 Transformer 架构的崛起对上述观点发起了挑战。这些模型证明了统计学习确实能够产生句法,它们能够在没有显式规则编码的情况下,从海量数据中“学会”语法。

The Evolution of Linguistic Theory

The debate over how children learn language has shifted over decades: * The Chomskyan View: In the 1950s, Noam Chomsky argued that language is too complex to be learned through statistics alone. He proposed that children are born with an innate "language instinct" or hardwired grammatical rules. * The Statistical Turn: The rise of neural networks and transformer architectures in the 2010s challenged this. These models proved that statistical learning could produce syntax, effectively "learning" grammar from massive datasets without explicit rule-coding.

BabyLM:检验假设

为了缩小这一差距,研究人员推出了 BabyLM 竞赛,要求开发者在“符合发展规律”的数据(大约 1000 万到 1 亿个词)上训练模型。

研究结果令人感到意外: * 课程学习(即先从简单数据学起再过渡到复杂数据)的效果并没有预期的那么好。 * 表现最好的模型往往采用了完全不模仿人类生物特征的架构,这表明我们仍未掌握人类学习的“核心秘诀”。

BabyLM: Testing Hypotheses

To bridge the gap, researchers launched BabyLM, a competition that challenges developers to train models on "developmentally plausible" data—roughly 10 million to 100 million words.

The findings have been surprising: * Curriculum learning (starting with simple data and moving to complex) was less effective than expected. * The best-performing models often use architectures that don't mimic human biology at all, suggesting that we still haven't captured the "secret sauce" of human learning.

缺失的要素:具身性与主体性

当前的 AI 模型是脱离身体、被动观察文本的旁观者。与此形成鲜明对比的是,儿童是其环境中的积极参与者。他们运用感官、探索物理世界,并且——至关重要的是——参与社交互动。

诸如布伦登·莱克(Brenden Lake)和尤里·哈森(Uri Hasson)等研究人员目前正利用头戴式摄像机拍摄的素材(例如 SAYCam 和 BabyView 项目),为 AI 提供与幼儿相同的视觉和听觉输入。他们的假设是:通过模拟儿童如何通过探索和社交反馈来“选择”自己的数据,我们最终或许能够弥合数据效率的鸿沟。

“这简直是个奇迹……如果你用 3000 万个词去训练 GPT-2,你得到的只是一个乱码生成器,而不是一个孩子。” — 迈克尔·C·弗兰克,斯坦福大学

The Missing Ingredients: Embodiment and Agency

Current AI models are disembodied, passive observers of text. Children, by contrast, are active participants in their environment. They use their senses, explore their physical world, and—crucially—engage in social interactions.

Researchers like Brenden Lake and Uri Hasson are now using headcam footage (such as the SAYCam and BabyView projects) to provide AI with the same visual and auditory input as a toddler. The hypothesis is that by simulating how children "choose" their own data through exploration and social feedback, we might finally close the data efficiency gap.

“It’s just totally miraculous … If you train GPT-2 on 30 million words, you get a nonsense generator; you don’t get a kid.” — Michael C. Frank, Stanford University

为什么这很重要

弥合数据鸿沟不仅关乎打造更好的聊天机器人,它还带来了以下几点优势: 1. 民主化: 更小、更高效的模型使得没有庞大计算预算的高校和研究人员也能够参与到 AI 研究中来。 2. 语言保护: 高效的模型有助于重振那些缺乏当前 LLM 所需海量数据集的少数民族语言。 3. 自我发现: 通过将 AI 视为一种“模式生物”,科学家们可以验证那些在真实儿童身上无法研究的人类认知理论,从而有效地借助机器来理解人类心智。

Why It Matters

Closing the data gap is about more than just building better chatbots. It offers: 1. Democratization: Smaller, efficient models allow universities and researchers without massive compute budgets to contribute to AI. 2. Linguistic Preservation: Efficient models could help revitalize minority languages that lack the massive datasets required by current LLMs. 3. Self-Discovery: By treating AI as a "model organism," scientists can test theories about human cognition that are impossible to study in real children, effectively using the machine to understand the human mind.


a cradle with an LLM model hanging like a mobile over itSELMAN DESIGN
a retro computer with the word hello in script on the screen sits in a child's high chairSELMAN DESIGN