文章背景与核心概要
粉末X射线衍射(XRD)是材料表征的基础方法,但实现可靠的端到端自动化一直是一个重大难题。一个自主的XRD智能体必须能够成功解释衍射数据、操作专门的精修软件、按逻辑顺序管理耦合参数,并严格区分单纯的数值改善与真正的物理有效性。
本文引入了 AutoXRD,这是一个专为粉末XRD分析设计的自主大语言模型(LLM)智能体框架,通过逐步精修、基于证据的操作和确定性的晶体学检查来构建分析流程。为了评估其有效性,作者提出了 XRDBench,这是一个包含两个不同赛道的综合评测基准:1. XRDBench-QA:100个侧重于科学推理和决策的有界诊断任务;2. XRDBench-E2E:34个可执行工作流,测试文件检查、软件执行、迭代精修和自动报告生成等端到端能力。
在评估了十个最新大模型的1,340次模型-任务运行中,模型的平均总分为 57.8/100(在QA任务上的得分为61.9,而在E2E工作流上则降至53.7)。虽然模型在精修历史评估和结果接受方面表现相对较好,但在指标化、相定量、Rietveld精修和精修动作选择方面遇到了巨大困难。
AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis
AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis
Authors: Yuetong Wu, Maojun Sun
Subjects: Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI)
arXiv: 2609.00070 [cond-mat.mtrl-sci]
Submitted: August 30, 2026
📌 Executive Summary
📌 Executive Summary
Powder X-ray diffraction (XRD) is a foundational method for materials characterization, but achieving reliable, end-to-end automation remains a significant hurdle. An autonomous XRD agent must successfully interpret diffraction data, operate specialized refinement software, manage coupled parameters in a logical sequence, and rigorously distinguish between mere numerical improvements and true physical validity.
This paper introduces AutoXRD, an autonomous Large Language Model (LLM) agent framework designed to structure powder-XRD analysis through stepwise refinement, evidence-grounded actions, and deterministic crystallographic checks. To evaluate its effectiveness, the authors present XRDBench, a comprehensive benchmark comprising two distinct tracks: 1. XRDBench-QA: 100 bounded diagnostic tasks targeting scientific reasoning and decision-making. 2. XRDBench-E2E: 34 executable workflows testing end-to-end capabilities like file inspection, software execution, iterative refinement, and automated reporting.
This paper introduces AutoXRD, an autonomous Large Language Model (LLM) agent framework designed to structure powder-XRD analysis through stepwise refinement, evidence-grounded actions, and deterministic crystallographic checks. To evaluate its effectiveness, the authors present XRDBench, a comprehensive benchmark comprising two distinct tracks: 1. XRDBench-QA: 100 bounded diagnostic tasks targeting scientific reasoning and decision-making. 2. XRDBench-E2E: 34 executable workflows testing end-to-end capabilities like file inspection, software execution, iterative refinement, and automated reporting.
Across 1,340 model-task runs evaluating ten recent LLMs, the models averaged an overall score of 57.8/100 (falling from 61.9 on QA tasks to 53.7 on E2E workflows). While models performed relatively well in refinement-history assessment and result acceptance, they struggled significantly with indexing, phase quantification, Rietveld refinement, and refinement-action selection.
Across 1,340 model-task runs evaluating ten recent LLMs, the models averaged an overall score of 57.8/100 (falling from 61.9 on QA tasks to 53.7 on E2E workflows). While models performed relatively well in refinement-history assessment and result acceptance, they struggled significantly with indexing, phase quantification, Rietveld refinement, and refinement-action selection.
Key Findings & Model Performance
- Top Overall Score: GPT-5.6 Sol achieved the highest overall score of 81.1.
- Best End-to-End Performance: GPT-5.6 Terra secured the highest XRDBench-E2E point estimate of 81.0.
- Best Value: GPT-5.6 Luna offered the optimal score-to-cost trade-off.
- Ablation Studies: All six individual components of the AutoXRD framework were proven to consistently enhance overall system performance.
- Common Failure Modes: Execution-trace analysis highlighted persistent challenges in coupled-parameter control, quantitative reasoning, evidence preservation, and workflow termination—pointing to a need for stronger scientific constraints and uncertainty-aware decision-making.
Key Findings & Model Performance
- Top Overall Score: GPT-5.6 Sol achieved the highest overall score of 81.1.
- Best End-to-End Performance: GPT-5.6 Terra secured the highest XRDBench-E2E point estimate of 81.0.
- Best Value: GPT-5.6 Luna offered the optimal score-to-cost trade-off.
- Ablation Studies: All six individual components of the AutoXRD framework were proven to consistently enhance overall system performance.
- Common Failure Modes: Execution-trace analysis highlighted persistent challenges in coupled-parameter control, quantitative reasoning, evidence preservation, and workflow termination—pointing to a need for stronger scientific constraints and uncertainty-aware decision-making.
🔗 Links & Resources
🔗 Links & Resources
- View PDF: arXiv:2609.00070 PDF
- HTML Version: arXiv HTML Preview
- DOI: 10.48550/arXiv.2609.00070