文章背景与核心概要
在航空等安全关键且受物理规律约束的领域中评估大语言模型(LLM),仅依赖传统的准确率指标已经远远不够。这是因为,表面上在数值上无限接近真实值的预测,仍可能违反运行约束、在物理场之间产生不一致的组合,或者输出无法直接使用的结构化结果。传统的评估方法无法可靠地识别出这些潜在的失效模式。
为了解决这一痛点,本文提出了 FLY-EVAL++——一种证据驱动的评估协议。该协议将确定性验证(涵盖协议合规性、物理可行性以及安全约束)与固定的评分准则引导聚合相结合,从而产出具备强可解释性的多维评分。通过对 66 个大语言模型在飞行轨迹与姿态预测(FTAP)任务上的评估,该协议揭示了“安全合规性”是区分模型行为最具决定性的维度。研究发现,在物理预测看似合理的模型之间,其安全得分的差距竟可高达 28 分以上,且普遍存在表面合理但实际违背安全约束、多步推理展开不稳定等频发故障。这些结果表明,在安全关键领域中,评估工作必须显式地衡量约束满足度和结构有效性,而不能仅依赖以准确率为中心的传统指标。
FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models
Summary
Evaluating Large Language Models (LLMs) in safety-critical, physics-governed domains like aviation requires moving beyond simple accuracy metrics. Traditional metrics fail because predictions that appear numerically close to ground truth can still violate operational constraints, mix physical fields inconsistently, or yield unusable outputs.
To address this, FLY-EVAL++ introduces an evidence-driven evaluation protocol that combines deterministic verification—checking protocol compliance, physical feasibility, and safety constraints—with fixed rubric-guided aggregation to produce interpretable, multi-dimensional scores. Evaluated across 66 LLMs on Flight Trajectory and Attitude Prediction (FTAP) tasks, the protocol reveals that safety compliance is the most critical differentiator in model behavior, exposing widespread issues like safety violations under seemingly plausible predictions and instability in multi-step rollouts.
Summary
Evaluating Large Language Models (LLMs) in safety-critical, physics-governed domains like aviation requires moving beyond simple accuracy metrics. Traditional metrics fail because predictions that appear numerically close to ground truth can still violate operational constraints, mix physical fields inconsistently, or yield unusable outputs.
To address this, FLY-EVAL++ introduces an evidence-driven evaluation protocol that combines deterministic verification—checking protocol compliance, physical feasibility, and safety constraints—with fixed rubric-guided aggregation to produce interpretable, multi-dimensional scores. Evaluated across 66 LLMs on Flight Trajectory and Attitude Prediction (FTAP) tasks, the protocol reveals that safety compliance is the most critical differentiator in model behavior, exposing widespread issues like safety violations under seemingly plausible predictions and instability in multi-step rollouts.
Metadata & Publication Details
- arXiv Identifier: arXiv:2609.04021 [cs.AI]
- Authors: Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang
- Submitted on: September 3, 2026
- Conference Publication: Proceedings of the Conference on Language Modeling (COLM), 2026
- Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- DOI: 10.48550/arXiv.2609.04021
Metadata & Publication Details
- arXiv Identifier: arXiv:2609.04021 [cs.AI]
- Authors: Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang
- Submitted on: September 3, 2026
- Conference Publication: Proceedings of the Conference on Language Modeling (COLM), 2026
- Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
- DOI: 10.48550/arXiv.2609.04021
Abstract
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing evaluation protocols do not measure these failure modes reliably. We propose FLY-EVAL++, an evidence-driven evaluation protocol that combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with fixed rubric-guided aggregation into interpretable multi-dimensional scores. We instantiate FLY-EVAL++ for Flight Trajectory and Attitude Prediction (FTAP) by extending the PilotBench setting with history-conditioned and multi-step prediction tasks. Across 66 LLMs, safety compliance is the most discriminative dimension of model behavior: models with comparable predictive performance differ by more than 28 points in safety score, and we observe recurrent failures including safety violations under physically plausible predictions and instability in multi-step rollouts. These results show that evaluation in safety-critical domains should measure constraint satisfaction and structured validity explicitly rather than rely on accuracy-centric reporting alone.
Abstract
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing evaluation protocols do not measure these failure modes reliably. We propose FLY-EVAL++, an evidence-driven evaluation protocol that combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with fixed rubric-guided aggregation into interpretable multi-dimensional scores. We instantiate FLY-EVAL++ for Flight Trajectory and Attitude Prediction (FTAP) by extending the PilotBench setting with history-conditioned and multi-step prediction tasks. Across 66 LLMs, safety compliance is the most discriminative dimension of model behavior: models with comparable predictive performance differ by more than 28 points in safety score, and we observe recurrent failures including safety violations under physically plausible predictions and instability in multi-step rollouts. These results show that evaluation in safety-critical domains should measure constraint satisfaction and structured validity explicitly rather than rely on accuracy-centric reporting alone.
Access & Resources
- Full-Text Links: View PDF | HTML Version | TeX Source
- License: Creative Commons Attribution 4.0 International

- External References: NASA ADS | Google Scholar | Semantic Scholar
Access & Resources
- Full-Text Links: View PDF | HTML Version | TeX Source
- License: Creative Commons Attribution 4.0 International
- External References: NASA ADS | Google Scholar | Semantic Scholar