具备自主决策能力的自动显微镜基准测试:支持系统资质认证,但未必能泛化至未见任务
文章背景与核心概要
随着大语言模型(LLM)智能体被越来越广泛地部署于控制显微镜和同步辐射光束线等科学基础设施,建立可靠的设计范式变得至关重要。本文在53项显微镜基准测试中,评估了各种设计选择(如一、二、三智能体图拓扑结构,5种不同的LLM,检索增强生成(RAG)参数以及操作约束)。该评估涵盖了105种智能体配置、1,949次测试运行和49,109次RAG检索,展示了延迟、Token使用量、成本以及失效模式之间的明显权衡。然而,基于这些架构和测试结果训练出的代理模型,无法可靠地预测智能体在新的、未见任务上的性能。
总体而言,当前的基准测试对于资质认证、回归测试和直接比较非常有价值,但仅靠异构测试套件无法支撑一个与任务无关的全局配置模型。
摘要 (Abstract)
Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent and there are few well-established paradigms for how to engineer an agentic system. There are many choices to make when designing a microscopy agent, including the choice of LLM, the number of agents to use, agent responsibilities and delegation rules, retrieval-augmented generation parameters, and more.
大语言模型智能体正越来越多地被开发用于控制各种科学表征工具,包括显微镜和同步辐射光束线。针对物理基础设施的智能体自主控制研究尚处于萌芽阶段,目前几乎没有成熟的范式来指导如何设计一个智能体系统。在设计显微镜智能体时,研究人员面临诸多选择,包括LLM的选择、智能体数量、智能体的职责与分工规则、检索增强生成(RAG)参数等。
When designing and optimizing an agentic microscope controller, researchers not only want to ensure that the agent can correctly perform known tasks but also that the agent can generalize to new tasks that it has not encountered before. In this study, we develop a benchmark and trace-logging framework that reveals: 1. How different choices of agent architecture impact performance at microscopy tasks. 2. The limitations of benchmarks for predicting if a particular agent will perform well on unseen microscopy tasks.
在设计和优化自主显微镜控制器时,研究人员不仅希望确保智能体能够正确执行已知任务,还希望智能体能够泛化到未曾遇到过的新任务中。在本研究中,我们开发了一个基准测试与追踪日志框架,揭示了以下内容: 1. 不同的智能体架构选择如何影响显微镜任务的性能。 2. 基准测试在预测特定智能体是否能在未见显微镜任务中表现良好时的局限性。
The framework was used to evaluate one-, two-, and three-agent graph topologies, five LLMs, RAG and context parameters, and operational constraints across 53 microscopy benchmark tests. In total, 105 agent configurations, 1,949 individual test runs, and 49,109 RAG retrievals were recorded.
该框架被用于在53项显微镜基准测试中评估一、二、和三智能体图拓扑结构、5种LLM、RAG与上下文参数以及操作约束。总共记录了105种智能体配置、1,949次独立测试运行以及49,109次RAG检索。
Direct comparisons showed clear differences in latency, token use, cost, and failure mode between configurations. However, surrogate models trained on agent architecture and test results did not reliably predict an agent's performance on new, unseen tasks. These results show that these benchmarks are useful for qualification, regression testing, diagnosis, and direct comparison, but the current heterogeneous test suite does not support a task-independent global configuration model.
直接对比表明,不同配置在延迟、Token使用量、成本和失效模式上存在明显差异。然而,基于智能体架构和测试结果训练的代理模型无法可靠地预测智能体在新的、未见任务上的性能。这些结果表明,这些基准测试对于资质认证、回归测试、诊断和直接对比非常有用,但当前的异构测试套件尚不足以支持一个与任务无关的全局配置模型。
元数据与出版详情 (Metadata & Publication Details)
- arXiv ID: arXiv:2608.05266 [cs.AI]
- Authors: Nathan S. Johnson, Ian Abshire
- Submitted: August 5, 2026
- Primary Subject: Artificial Intelligence (
cs.AI)- Secondary Subjects: Materials Science (
cond-mat.mtrl-sci), Machine Learning (cs.LG)- Comments: 20 pages, 8 figures
- DOI: 10.48550/arXiv.2608.05266
- arXiv ID: arXiv:2608.05266 [cs.AI]
- 作者: Nathan S. Johnson, Ian Abshire
- 提交时间: 2026年8月5日
- 主学科: 人工智能 (
cs.AI) - 辅学科: 材料科学 (
cond-mat.mtrl-sci)、机器学习 (cs.LG) - 评论: 20页,8张图表
- DOI: 10.48550/arXiv.2608.05266