AI天气模型是否会忽略极端天气?
文章背景与核心概要
长期以来,第一代AI天气模型在预测极端天气事件时表现不佳的批评屡见不鲜,这些结论主要基于对确定性回归系统进行再分析评估(reanalysis evaluations)得出。为了验证这一说法,研究人员在为期十个月的时间里,将11个物理预报系统和AI预报系统与欧洲的测风站、太阳能观测站以及雨量计站点的数据进行了对比评估。
该研究分析了 10米风速、2米气温、逐小时短波辐射累积量 以及 逐小时降水量 的相关指标,并在ERA5(1991–2020年)气候态系统下,以欧洲中期天气预报中心(ECMWF)的IFS系统为基准计算了平均绝对误差(MAE)。
研究结果表明,AI模型在分布尾部(即极端天气)并没有表现出统一的相对技能缺陷。相反,性能差异取决于具体的模型,而非AI天气模型作为一类技术的系统性缺陷。值得注意的是,无论是AI模型还是数值天气预报(NWP)系统,都表现出一种向观测分布中心靠拢的共同条件偏差。
文章元数据 (Article Metadata)
| 字段 | 详情 |
|---|---|
| arXiv 标识符 | arXiv:2608.09972 |
| 主要学科 | 大气与海洋物理学 (physics.ao-ph) |
| 次要学科 | 人工智能 (cs.AI)、机器学习 (cs.LG) |
| 提交日期 | 2026年7月31日 |
| DOI | 10.48550/arXiv.2608.09972 |
| 许可证 | CC BY-NC-SA 4.0 ![]() |
作者 (Authors)
- Marvin Vincent Gabler
- Roberto Molinaro
- Niall Siegenheim
- Henry Martin
- Mark Frey
- Niels Poulsen
- Philipp Seitz
- Olivier Lam
摘要 (Abstract)
据报道,第一代AI天气模型在预测极端天气时往往表现不佳,这主要是在基于再分析数据的确定性回归系统评估中得出的。我们在为期十个月的时间里,针对10米风速、2米气温、逐小时短波辐射累积量以及逐小时降水量,使用欧洲天气、太阳能和雨量计站点数据对11个物理预报和AI预报系统进行了验证,并在ERA5 1991-2020年气候态下以ECMWF IFS为基准对平均绝对误差(MAE)进行了评分。
First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European synoptic, solar, and rain-gauge stations over ten months for 10 m wind, 2 m temperature, hourly shortwave accumulation, and hourly precipitation, scoring mean absolute error (MAE) against ECMWF IFS in ERA5 1991-2020 climatological regimes.
在这些系统中,AI模型在分布尾部并没有显示出统一的相对技能缺陷: * 风速: Jua EPT-2.1 Europa 在全天候风速预测中领先(+8.4%)。EPT-2.1 Europa 和 DWD ICON Global 在烈风(gale-force wind)条件下均处于领先地位。 * 气温: Jua EPT-2 HRRR 在整体气温(+12.1%)以及炎热天气条件(+19.6 ± 2.2%)下均处于领先地位。相反,ECMWF AIFS 在高温尾部损失了 4.9 ± 2.0%,而 NOAA GFS 在同一条件下的损失达到 22.8 ± 2.0%。 * 太阳能: Jua EPT-2.1 Helios 在整体太阳能辐射(+10.2 ± 1.7%)、阴天条件(+16.4 ± 3.4%)以及晴空尾部(+24.8 ± 5.4%)均处于领先。 * 降水: 三个 Jua 模型在中等降水强度下性能提升了 14–15%,在 P75–P95 分位点提升了 9–11%;EPT-2 Reasoning 在 P95 分位点以上保持领先地位(+1.7 ± 0.5%)。
Among these systems, AI models do not show a uniform relative-skill deficit in the tails: * Wind: Jua EPT-2.1 Europa leads all-conditions wind (+8.4%). Both EPT-2.1 Europa and DWD ICON Global lead at gale-force wind. * Temperature: Jua EPT-2 HRRR leads temperature overall (+12.1%) and in the heat regime (+19.6 ± 2.2%). Conversely, ECMWF AIFS loses 4.9 ± 2.0% in the heat tail, and NOAA GFS loses 22.8 ± 2.0% in the same regime. * Solar: Jua EPT-2.1 Helios leads solar overall (+10.2 ± 1.7%), in overcast conditions (+16.4 ± 3.4%), and in the clear-sky tail (+24.8 ± 5.4%). * Precipitation: Three Jua models gain 14–15% at moderate intensity and 9–11% at P75–P95; EPT-2 Reasoning remains ahead above P95 (+1.7 ± 0.5%).
每一个模型(包括数值天气预报系统)都表现出一种共同的条件偏差,即倾向于向观测分布的中心靠拢,模型间的差异(spread)比这种共同信号要小好几个数量级。因此,极端天气下相对技能的缺失并不是AI天气模型这一类技术的固有属性,而是某些特定AI模型和物理模型的问题。
Every model—including numerical weather prediction systems—shows a shared conditional bias toward the center of the observed distribution, with an inter-model spread several times smaller than the shared signal. Missing relative skill at extremes is therefore not a property of AI weather models as a class, but of particular AI and physical models.
