跳转至

AI智能体能否进行开放式AI研究?两项案例研究的初步证据

文章背景与核心概要

随着人工智能研发自动化预期的不断升温,评估AI智能体是否具备真正进行开放式研究的能力已成为一项关键挑战。现有的评估方法往往局限于狭窄且可验证的任务,或者依赖于负担过重的盲审机制。为了填补这一空白,本文作者引入了“影子评估”(shadow evaluations)框架:让AI智能体尝试解决高质量未发表论文中的核心开放性研究问题,并由原作者对结果进行评分。

通过对两篇NeurIPS 2026未发表论文进行为期六天、投入大量计算资源的测试,研究发现,尽管智能体能够在无人干预的情况下完成所有工程任务,但在更广泛的研究生命周期中表现乏力。最终,两篇论文均被原作者明确拒绝。研究总结了五种反复出现的失败模式,包括对发表标准的判断力不足、缺乏创造性的调整以及指令漂移等,为当前AI驱动的研究自动化局限性提供了重要的实证洞察。


📌 总结

随着对AI研发自动化的期望不断提高,评估AI智能体是否能够真正进行开放式研究仍然是一个严峻的挑战。现有的评估技术要么将智能体限制在狭窄、可验证的任务中,要么依赖于负担过重的盲审。为了弥补这一差距,作者引入了“影子评估”:这是一种评估框架,即AI智能体尝试解决高质量未发表论文的核心开放式研究问题,并由原作者对结果进行评分。

在两篇未发表的NeurIPS 2026论文中,我们对前沿智能体进行了为期六天的测试,并投入了大量的计算预算。结果显示,虽然智能体在没有人类干预的情况下成功执行了所有必要的工程任务,但它们在更广泛的研究生命周期中却举步维艰。因此,两篇论文都被原作者明确拒绝。该研究确定了五种反复出现的失败模式——包括对发表门槛的判断力差、对研究设计缺陷的反应缺乏创造力、以及指令漂移等——为当前AI驱动的研究自动化的局限性提供了重要的实证见解。

As expectations around AI R&D automation grow, evaluating whether AI agents can genuinely conduct open-ended research remains a critical challenge. Existing evaluation techniques either limit agents to narrow, verifiable tasks or rely on overstretched blind peer reviews. To bridge this gap, the authors introduce shadow evaluations: an evaluation framework where an AI agent attempts to solve the central, open-ended research question of a high-quality unpublished paper, and the original authors grade the results.

Testing frontier agents over six days with substantial compute budgets on two unpublished NeurIPS 2026 submissions revealed that while agents successfully executed all necessary engineering without human intervention, they struggled significantly with the broader research lifecycle. Consequently, both papers were rejected by their original authors. The study identifies five recurring failure modes—including poor publication-bar judgment, uncreative adaptations, and instruction drift—offering vital empirical insight into the current limitations of AI-driven research automation.


📝 摘要

关于AI爆炸式发展的预测取决于AI智能体能否实现AI研究的自动化。但关于智能体是否能够开展开放式AI研究的证据尚不充分。目前的评估要么在狭窄、可验证的任务上测试智能体(这排除了开放式研究),要么将AI生成的论文提交给盲审(这不仅负担过重、具有随机性,而且审稿质量较差)。

我们引入了第三种衡量AI研发自动化进展的方法。智能体承担高质量未发表论文的核心开放式研究问题,并由论文原作者对输出结果进行评分。我们称之为“影子评估”。我们对两篇未发表的NeurIPS 2026论文进行了影子评估,给予前沿智能体六天的时间和数千美元的计算资源。智能体在没有人类帮助的情况下完成了所有工程工作,但在回答研究问题方面却未能取得实质性进展。因此,两篇论文都被作者明确拒绝。

我们确定了五种反复出现的失败模式: 1. 对可发表研究标准的判断力差 2. 对研究设计缺陷的反应缺乏创造力 3. 从死胡同中回溯无效 4. 资源意识薄弱 5. 指令漂移

通过第二个模型和脚手架进行的稳健性检查重现了这些失败。我们发布了专家评审、调查回复、智能体存储库和日志。我们的结果提供了初步证据,表明今天的智能体可以完成AI研究的工程部分,但在研究生命周期的关键环节上仍面临困难。

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality.

We introduce a third way to measure progress towards AI R&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors.

We identify five recurring failure modes: 1. Poor judgment about the bar for publishable research 2. Uncreative responses to shortcomings in the research design 3. Ineffective backtracking from dead ends 4. Poor resource awareness 5. Instruction drift

A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.


🔗 链接与资源