文章背景与核心概要
当前的各类人工智能系统主要针对“回答现有问题”进行了深度优化,然而科学事业的真正瓶颈其实出现在更早的阶段:即发现真正值得研究的科学问题。本文介绍了一种新颖的框架,能够将可重复的研究语料库转化为带有多重排名的、可证伪的研究问题。该系统将证据表示为携带溯源信息的断言(claims),从而检测并分类跨论文之间的矛盾冲突(tensions),随后通过人工裁决来过滤信号。存活下来的信号会通过一个两阶段协议进行精炼与优先级排序,该协议明确区分了科学优先级与执行优先级。
研究人员在系外行星大气这一融合了学术文献、结构化星表以及太空望远镜档案的领域对该框架进行了历史回测。测试结果展现出卓越的预测有效性:所有基于2021年之前可用证据生成的问题,均在系统从未见过的2021年至2026年后续文献中得到了实质性的探讨——其中两个问题已被解答(包括一个其前提后来被明确驳回的问题),而排名最高的问题目前仍由独立学者提出且处于未解状态。这些发现表明,基于证据冲突的系统性问题 discovery(问题发现),能够精准浮现出在一线科学家后续研究中真正投入精力的关键问题。
Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting
Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting
Summary
Summary
While current artificial intelligence systems are heavily optimized for answering existing questions, the scientific enterprise is fundamentally bottlenecked at an earlier stage: discovering the right questions worth investigating.
While current artificial intelligence systems are heavily optimized for answering existing questions, the scientific enterprise is fundamentally bottlenecked at an earlier stage: discovering the right questions worth investigating.
This paper introduces a novel framework that transforms a reproducible research corpus into ranked, falsifiable research questions. By representing evidence as provenance-carrying claims, the system detects and types cross-paper tensions, which are then human-adjudicated. Surviving signals are refined and prioritized using a two-stage protocol separating scientific priority from execution priority.
This paper introduces a novel framework that transforms a reproducible research corpus into ranked, falsifiable research questions. By representing evidence as provenance-carrying claims, the system detects and types cross-paper tensions, which are then human-adjudicated. Surviving signals are refined and prioritized using a two-stage protocol separating scientific priority from execution priority.
Tested historically on the domain of exoplanet atmospheres—combining literature, structured catalogs, and space telescope archives—the framework demonstrated remarkable predictive validity. All questions generated from evidence available prior to 2021 were substantively engaged by subsequent literature (2021–2026) that the system had never seen: two were answered (including one whose premise was later explicitly refuted), and the top-ranked question remains independently posed and open. These findings suggest that systematic question discovery from evidence tensions effectively surfaces the exact questions working scientists subsequently choose to investigate.
Tested historically on the domain of exoplanet atmospheres—combining literature, structured catalogs, and space telescope archives—the framework demonstrated remarkable predictive validity. All questions generated from evidence available prior to 2021 were substantively engaged by subsequent literature (2021–2026) that the system had never seen: two were answered (including one whose premise was later explicitly refuted), and the top-ranked question remains independently posed and open. These findings suggest that systematic question discovery from evidence tensions effectively surfaces the exact questions working scientists subsequently choose to investigate.
Document Details
Document Details
- arXiv Identifier: arXiv:2608.09968 [cs.DL]
- Authors: Hui Mao
- Submitted: July 29, 2026
- Primary Subject: Digital Libraries (
cs.DL) - Secondary Subject: Artificial Intelligence (
cs.AI) - Length: 13 pages, 4 figures
- DOI: 10.48550/arXiv.2608.09968
- arXiv Identifier: arXiv:2608.09968 [cs.DL]
- Authors: Hui Mao
- Submitted: July 29, 2026
- Primary Subject: Digital Libraries (
cs.DL)- Secondary Subject: Artificial Intelligence (
cs.AI)- Length: 13 pages, 4 figures
- DOI: 10.48550/arXiv.2608.09968
Abstract
Abstract
Current AI systems are optimized for answering questions; the scientific enterprise is bottlenecked earlier, at discovering the questions worth investigating. We present a framework that turns a traceable, reproducible, scope controlled research corpus into ranked, falsifiable research questions: evidence is represented as provenance carrying claims; cross paper tensions are detected, typed, and human adjudicated; surviving signals are refined into questions and ranked by a two stage protocol separating scientific priority from execution priority. We instantiate the framework on exoplanet atmospheres, a domain that uniquely combines literature, structured catalogs, and space telescope archives. In a historical backtest, all questions generated from evidence available before 2021 were substantively engaged by the 2021 to 2026 literature the system never saw: two were answered, including one whose premise the community later explicitly refuted and the top ranked question is independently posed and still open. These results suggest that systematic question discovery from evidence tensions surfaces the questions working scientists subsequently invest in.
Current AI systems are optimized for answering questions; the scientific enterprise is bottlenecked earlier, at discovering the questions worth investigating. We present a framework that turns a traceable, reproducible, scope controlled research corpus into ranked, falsifiable research questions: evidence is represented as provenance carrying claims; cross paper tensions are detected, typed, and human adjudicated; surviving signals are refined into questions and ranked by a two stage protocol separating scientific priority from execution priority. We instantiate the framework on exoplanet atmospheres, a domain that uniquely combines literature, structured catalogs, and space telescope archives. In a historical backtest, all questions generated from evidence available before 2021 were substantively engaged by the 2021 to 2026 literature the system never saw: two were answered, including one whose premise the community later explicitly refuted and the top ranked question is independently posed and still open. These results suggest that systematic question discovery from evidence tensions surfaces the questions working scientists subsequently invest in.
Access & Resources
Access & Resources