文章背景与核心概要
长期以来,高分辨率(\(0.1^\circ\))机器学习全球天气 forecasting 模型的开发一直受制于大规模高分辨率训练数据的缺乏——数十年的再分析数据主要仅提供在较粗的 \(0.25^\circ\) 分辨率下。传统方法试图通过在有限的 \(0.1^\circ\) 样本上对粗分辨率模型进行微调来弥补这一差距,但这种迁移从根本上受到了粗分辨率预报中固有的不可逆信息丢失的阻碍。
为了克服这一难题,作者推出了 BaguanHR,这是一个将范式从“模型迁移”转变为“数据迁移”的新颖框架。通过证明超分辨率(SR)比预报具有更低的条件熵和输入放大效应,BaguanHR 确立了 SR 作为分辨率迁移更稳健载体的地位。该框架利用逐变量超分辨率(variable-wise super-resolution),从 ERA5 再分析数据中合成了大量 \(0.1^\circ\) 的训练数据。
该研究的主要成果包括: * 卓越的性能: BaguanHR 的表现超越了传统机器学习方法和 IFS-HRES,在 72 小时窗口内的超 85% 预报时效中实现了更高的准确率。 * 幂律缩放效应: 研究揭示了清晰的幂律缩放关系,即训练数据增加一倍,72 小时预报的均方根误差(RMSE)降低了 4.6%,120 小时预报的均方根误差降低了 4.9%。
这些结果证实,高分辨率机器学习天气预报的规模化本质上是一个数据限制问题,而逐变量超分辨率提供了一种简单而有效的方法,能够解锁历史粗分辨率再分析数据,以用于先进的高分辨率训练。
通过数据规模化突破高分辨率天气预报的极限
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
作者: 杨赵、牛培松、周甜、马子晴、马冠龙、金蓉、袁慧玲、孙亮
提交日期: 2026年7月31日
主要主题: 机器学习(cs.LG),并交叉收录于人工智能(cs.AI)和计算机视觉(cs.CV)
录用会议: ECCV 2026
标识符: arXiv:2608.14652 | DOI: 10.48550/arXiv.2608.14652
📌 摘要
📌 Summary
长期以来,高分辨率(\(0.1^\circ\))机器学习全球天气预报模型的开发一直受制于缺乏广泛的高分辨率训练数据——数十年的再分析数据主要仅以较粗的 \(0.25^\circ\) 分辨率提供。传统方法试图通过在有限的 \(0.1^\circ\) 样本上微调粗分辨率模型来弥合这一差距,但这种迁移从根本上受到粗分辨率预报中固有的不可逆信息丢失的阻碍。
The development of high-resolution (\(0.1^\circ\)) machine learning-based global weather forecasting models has historically been bottlenecked by a lack of extensive high-resolution training data—decades of reanalysis data are predominantly available only at a coarser \(0.25^\circ\) resolution. Traditional approaches attempt to bridge this gap by fine-tuning coarse-resolution models on limited \(0.1^\circ\) samples, but this transfer is fundamentally hindered by the irreversible information loss inherent in coarse-resolution forecasting.
为了克服这一困难,作者引入了 BaguanHR,这是一个将范式从模型迁移转向数据迁移的新颖框架。通过证明超分辨率(SR)比预报具有更低的条件熵和输入放大效应,BaguanHR 确立了 SR 作为分辨率迁移更稳健载体的地位。利用逐变量超分辨率,该框架从 ERA5 再分析中合成了大量 \(0.1^\circ\) 的训练数据。
To overcome this, the authors introduce BaguanHR, a novel framework that shifts the paradigm from model transfer to data transfer. By demonstrating that super-resolution (SR) features lower conditional entropy and input amplification than forecasting, BaguanHR establishes SR as a more robust vehicle for resolution transfer. Using variable-wise super-resolution, the framework synthesizes extensive \(0.1^\circ\) training data from ERA5 reanalysis.
该研究的主要成果包括: * 卓越性能: BaguanHR 优于传统机器学习方法和 IFS-HRES,在 72 小时窗口内超过 85% 的预报时效中实现了更高的准确率。 * 幂律缩放效应: 研究揭示了清晰的幂律缩放关系,其中训练数据增加一倍,使 72 小时预报的均方根误差(RMSE)降低了 4.6%,120 小时预报的均方根误差降低了 4.9%。
Key outcomes of the study include: * Superior Performance: BaguanHR outperforms both traditional ML methods and IFS-HRES, achieving superior accuracy across more than 85% of lead times within a 72-hour window. * Power-Law Scaling Effect: The research reveals a clear power-law scaling relationship, where a twofold increase in training data reduces Root Mean Squared Error (RMSE) by 4.6% for 72-hour forecasts and 4.9% for 120-hour forecasts.
这些结果证实,高分辨率机器学习天气预报的规模化主要是一个数据限制问题,并且逐变量超分辨率提供了一种简单而有效的解决方案,可以解锁历史粗分辨率再分析以进行高级的高分辨率训练。
These results confirm that scaling high-resolution ML weather forecasting is primarily a data limitation problem, and that variable-wise super-resolution provides a simple yet effective solution to unlock historical coarse-resolution reanalyses for advanced high-resolution training.
🔗 链接与资源
🔗 Links & Resources
- 全文访问:
- 查看 PDF
- HTML 版本(实验性)
- TeX 源码
- 许可证: 知识共享署名-非商业性使用-禁止演绎 4.0 国际许可协议

- 外部引用与工具:
- NASA ADS
- Google Scholar
- Semantic Scholar
- Full Text Access:
- View PDF
- HTML Version (Experimental)
- TeX Source
- License: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0
- External Citations & Tools:
- NASA ADS
- Google Scholar
- Semantic Scholar