ARTICLE
14 March 2026

面向开放场景理解的遥感点云多模态融合与开放词汇技术综述

琼洁 王1
Show Less
1 中国电子信息产业发展研究院, 中国
TACS 2026 , 3(5), 110–112; https://doi.org/10.61369/TACS.2026050048
© 2026 by the Author(s). Licensee Art and Technology, USA. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution -Noncommercial 4.0 International License (CC BY-NC 4.0) ( https://creativecommons.org/licenses/by-nc/4.0/ )
Abstract

随着机载激光雷达、车载移动测量与无人机倾斜摄影等技术的快速发展,遥感场景中产生了海量高精度三维点云数据。传统遥感点云解译方法多建立在封闭类别监督范式之上,依赖大量点级标注,且在未知类别识别、跨场景迁移和复杂语义查询等方面存在明显局限。近年来,随着多模态预训练、开放词汇视觉理解以及大语言模型的发展,这些方式为遥感点云从“封闭集语义分割”走向“开放场景三维理解”提供了新的技术路径。基于此,本文围绕这一演进脉络,对相关研究进行系统梳理。首先,从遥感点云的大规模、稀疏性与弱纹理特征出发,概述了适用于大场景处理的点云表征骨干网络与自监督预训练范式。其次,本文重点总结了基于二维视觉- 语言模型的2D-to-3D 跨模态蒸馏与特征反投影方法与面向点云直接输入的三维语言模型构建方法这两类关键路线。最后,结合测绘遥感任务对几何精度、跨尺度一致性与工程部署的特殊要求,分析了现有方法在空间精确对齐、纯几何语义表达、长尾类别识别、数据集构建与可信落地等方面的主要瓶颈。

Keywords
遥感点云
开放词汇分割
多模态融合
三维基础模型
3D-LLM
跨模态蒸馏
References

[1] 杨必胜, 董震. 点云智能研究进展与趋势[J]. 测绘学报,2019,48(12):1575-1585.
[2] 胡伏原, 李晨露, 周涛, 等. 面向深度学习的三维点云补全算法综述[J]. 中国图象图形学报,2025,30(2):309-333. DOI:10.11834/jig.240124.
[3] 潘洁晨, 邢帅, 曹家印, 等. 基于深度学习的航空点云语义分割研究进展[J]. 地球信息科学学报, 2025,27(9):1999-2020. DOI:10.12082/dqxxkx.2025.250151.
[4] 张帅豪, 潘志刚. 遥感大模型: 综述与未来设想[J]. 遥感技术与应用,2025,40(1):1-13. DOI:10.11873/j.issn.1004-0323.2025.1.0001.
[5]Hu Q, Yang B, Xie L, et al. Randla-net: Efficient semantic segmentation of large-scale point clouds[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020: 11108-11117.
[6]Choy C, Gwak J Y, Savarese S. 4d spatio-temporal convnets: Minkowski convolutional neural networks[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 3075-3084.
[7] 郑智鸿, 宋海川. 基于组对比学习的弱监督三维点云语义分割方法[J]. 华东师范大学学报(自然科学版),2024(2):108-118. DOI:10.3969/j.issn.1000-5641.2024.02.012.
[8]Yu X, Tang L, Rao Y, et al. Point-bert: Pre-training 3d point cloud transformers with masked point modeling[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022: 19313-19322.
[9]Pang Y, Tay E H F, Yuan L, et al. Masked autoencoders for 3d point cloud selfsupervised
learning[J]. World Scientific Annual Review of Artificial Intelligence, 2023, 1: 2440001.
[10]Peng S, Genova K, Jiang C, et al. Openscene: 3d scene understanding with open vocabularies[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023: 815-824.
[11]Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//International conference on machine learning. PmLR, 2021: 8748-8763.
[12]Takmaz A, Fedele E, Sumner R W, et al. Openmask3d: Open-vocabulary 3d instance segmentation[J]. arXiv preprint arXiv:2306.13631, 2023.
[13]guyen P, Ngo T D, Kalogerakis E, et al. Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 4018-4028.
[14] 付琨, 卢宛萱, 刘小煜, 等. 遥感基础模型发展综述与未来设想[J]. 遥感学报, 2024, 28(7).DOI:10.11834/jrs.20233313.
[15]Hong Y, Zhen H, Chen P, et al. 3d-llm: Injecting the 3d world into large language models[J]. Advances in Neural Information Processing Systems, 2023, 36:20482-20494.
[16]Xu R, Wang X, Wang T, et al. Pointllm: Empowering large language models to understand point clouds[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024: 131-147.

Share
Back to top