K-Bench:衡量模型在真实科学智能体请求上的性能
文章背景与核心概要
当前的科学人工智能(AI)基准测试通常是为了便于自动化打分而过度设计的——它们严重依赖选择题、带有预定义参考答案的精心编排的智能体任务,或者是具有已知生成结构的模拟器。相比之下,K-Bench 01 引入了一种更具现实意义的评估框架,该框架直接构建自 K-Dense Web 真实用户流量中的第一轮请求样本。
通过在相同的沙盒环境中对九个前沿模型进行端到端测试(共完成 1,602 次运行)并采用严格的多裁判评估体系,这项研究揭示了当前前沿 AI 能力的关键洞察:首先是“阈值挑战”,在严格的标准下(即 8 分锚定标准代表领域科学家只需进行微小修改即可接受的工作),没有任何单一模型能够在所有三个盲测大模型裁判的标准下同时跨越这一及格线。其次是模型排名,虽然 gpt-5.6-sol 取得了最高的合并平均分(8.04),但其置信区间跨越了接受阈值,且三分之二的裁判将 claude-opus-5 排在第一位。因此,模型排序成了可靠的度量标准,而绝对分数仍取决于评估工具。最后是系统性弱点,在近 40,000 次独立评估中,47.6% 的分数低于 8 分阈值。其中科学准确度平均仅为 6.22,远低于沟通能力的 7.33;此外,过度夸大(overclaiming)成为最突出的失败标签,存在于 31.4% 的评估中。
最终,作者认为评估科学智能体需要超越单纯的排行榜排名,转而分析交付内容、声明内容和产生工件之间的联合分布。
Paper Metadata
- arXiv Identifier: arXiv:2608.21601 [cs.AI]
- DOI: 10.48550/arXiv.2608.21601
- Primary Subject: Artificial Intelligence (
cs.AI) - Secondary Subject: Computation and Language (
cs.CL) - Authors: Aubrey M. Brueckner, Darshil Patel, Yuhuan He, Timothy Kassis
- Submission History:
- [v1] Fri, 21 Aug 2026
- [v2] Wed, 2 Sep 2026 (Current version: textual corrections and clarifications)
- License: Creative Commons Attribution 4.0

Abstract
Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and they lack ground truth.
We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor is defined as work a domain scientist would accept with minor edits, no model clears the line under all three judges.
gpt-5.6-solhas the highest pooled mean, 8.04, but its 95% interval[7.80, 8.23]spans the threshold, and two of the three judges rankclaude-opus-5first instead.We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved. Across all 39,934 scored judgments — the eight dimension scores plus a holistic overall for each assessment, excluding not-applicable cells — 47.6% fall below the 8-point threshold. Difficulty is not uniform across the rubric: scientific accuracy averages 6.22 against 7.33 for communication, on identical denominators and in the same direction within every one of the nine models. The single leading failure tag is overclaiming, on 31.4% of assessments. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced.
科学人工智能的基准测试大多是为了打分而编写的:选择题、带有参考答案的精心编排的智能体任务,或者是具有已知生成结构的模拟器。而真正的科学请求则截然不同。它们往往描述不足、带有附件且缺乏真实标签(ground truth)。
我们报告了 K-Bench 01,这是一个评估框架,它构建自 K-Dense Web 真实用户流量中的第一轮请求样本,由九个前沿模型在相同的沙盒中进行端到端运行,共产生了 1,602 次完成的智能体运行。三个盲测的大模型裁判根据八个维度的评分标准对每次运行进行了打分。在一个将其 8 分锚定标准定义为“领域科学家只需进行微小修改即可接受的工作”的评分体系下,没有任何模型能在所有三个裁判的标准下全部及格。gpt-5.6-sol 拥有最高的合并平均分 8.04,但其 95% 的置信区间 [7.80, 8.23] 跨越了阈值,并且三位裁判中有两位将 claude-opus-5 排在第一位。
因此,我们报告系统的排序作为可复现的量,将绝对水平视为评估工具的一个属性,并将排行榜的顶端视为未决状态。在所有 39,934 项评分判断中(包括八个维度得分加上每次评估的整体得分,排除了不适用单元格),47.6% 的分数低于 8 分阈值。任务难度在各个评估维度上并不均匀:科学准确度的平均分为 6.22,而沟通能力平均分为 7.33(在相同的分母下,且在全部九个模型中均呈现相同的方向性趋势)。最主要的失败标签是过度夸大(overclaiming),占所有评估的 31.4%。我们认为,对于科学智能体而言,具有信息量的指标不是排行榜上的位置,而是交付内容、声明内容和产生工件之间的联合分布。
Key Takeaways & Access
- Full-Text Links: View PDF | Experimental HTML | TeX Source
- External Resources:
- NASA ADS
- Google Scholar
- Semantic Scholar
- 全文链接: 查看 PDF | 实验性 HTML | TeX 源码
- 外部资源:
- NASA ADS
- Google Scholar
- Semantic Scholar