文章背景与核心概要
在基于相机的卫星视觉感知领域,训练鲁棒的模型需要进行“仿真到真实(Sim2Real)”的数据构建,以此弥合合成渲染图(能提供精确的几何标注)与真实世界传感器图像之间的外观差异。然而,获取带有可靠姿态标签和组件级掩码的大规模真实卫星传感器图像非常困难,纯仿真渲染虽有完美标注却存在显着的领域鸿沟。
本文提出了一种全新的、面向卫星视觉Sim2Real数据构建的组件感知保结构风格迁移框架。该方法通过结合经过标定的真实图像采集、基于ArUco的相机位姿测量、CAD渲染以及组件掩码,构建了弱配对的真实-合成样本。随后,它从无标签真实图像中提取部件级的真实域风格码,并通过掩码对齐的调制将其注入到对应的合成卫星区域中。为了确保生成的图像能够用于下游传感器数据的监督,该方法将对抗训练与局部对比一致性、自正则化以及边缘保持约束相结合。实验证明,该方法在保持几何标注的同时显着提升了卫星视觉数据的Sim2Real生成质量,有效增强了下游姿态估计器的性能。
Component-Aware Structure-Preserving Style Transfer for Satellite Visual Sim2Real Data Construction
Summary
For camera-based satellite visual sensing, training robust models requires Sim2Real data construction—bridging the appearance gap between synthetic renders (which provide precise geometric annotations) and real-world sensor images.
This paper introduces a novel component-aware structure-preserving style transfer framework that builds weakly paired real-synthetic samples using calibrated real acquisition, ArUco-based camera-pose measurement, CAD rendering, and component masks. By extracting part-wise real-domain style codes and injecting them into corresponding synthetic satellite regions via mask-aligned modulation, the method successfully bridges the domain gap. Combined with adversarial training, local contrastive consistency, self-regularization, and edge-preserving constraints, the generated data effectively improves downstream sensor-data supervision (e.g., training the GDRNet pose estimator).
Metadata & Article Details
- arXiv ID: arXiv:2605.19624 [cs.CV]
- Subjects: Computer Vision and Pattern Recognition (
cs.CV); Artificial Intelligence (cs.AI)- Authors:
- Zongwu Xie
- Yonglong Zhang
- Yifan Yang
- Yang Liu
- Baoshi Cao
- Guanghu Xie
- Submission History:
- [v1] Tue, 19 May 2026
- [v2] Wed, 20 May 2026
- [v3] Fri, 21 Aug 2026 (Latest Version)
Abstract
For camera-based satellite visual sensing, Sim2Real data construction requires images that approach real-domain sensor appearance while retaining the annotations inherited from simulation. Real sensor images of satellite targets with reliable pose labels and component-level masks are difficult to acquire at scale, whereas synthetic rendering provides exact geometric annotations but suffers from a visible appearance gap.
This paper presents a component-aware structure-preserving style transfer framework for satellite visual synthetic-to-real data construction. The method builds weakly paired real--synthetic samples from calibrated real acquisition, ArUco-based camera-pose measurement, CAD rendering, and component masks. It then extracts part-wise real-domain style codes from unlabeled real images and injects them into corresponding synthetic satellite regions through mask-aligned modulation. To keep the generated images usable for downstream sensor-data supervision, adversarial training is combined with local contrastive consistency, self-regularization, and edge-preserving constraints.
Experiments are conducted on 5,000 rendered satellite images and 100 real images captured in a calibrated setup. The real images provide target-domain appearance references and final evaluation images, while the downstream GDRNet pose estimator is trained only on synthetic or translated synthetic images. Compared with representative image-translation baselines, the proposed method achieves the lowest image distribution discrepancy, with an FID of 54.32 and a KID of 0.048. When the translated data are used to train GDRNet in this target-domain adaptation setting, the ADD pass rate improves to 0.260 and the AUC improves to 0.611. These results indicate that component-level appearance transfer can improve annotation-preserving satellite visual Sim2Real data generation in the considered calibrated setup.
Access & Resources
- Full-Text Links:
- View PDF
- HTML Version (Experimental)
- TeX Source
- Digital Object Identifier (DOI): 10.48550/arXiv.2605.19624
- External Citations & Tools:
- NASA ADS
- Google Scholar
- Semantic Scholar