Abstract:Task-oriented route planning from remote sensing imagery requires understanding of open-ended semantic constraints and fine-grained perception of complex land-cover environments. Existing methods relying on high-definition maps, road networks, vehicle-mounted sensors, or manually specified rules struggle to adapt to such scenarios. A neuro-symbolic collaborative route planner(NSCRP) for route planning in remote sensing is proposed. First, a multimodal large language model is utilized to translate natural-language tasks into structured traversability rules and preference representations. A stable semantic land-cover map is constructed by integrating multiple semantic segmentation models. Then, semantic constraints and environmental perception results are uniformly mapped into a pixel-level cost map. An interpretable path with the minimum accumulated cost is generated via heuristic search on the predicted cost map. Experiments based on NeSy-Route show that NSCRP outperforms existing approaches in constraint satisfaction, path quality, and overall planning ability.
[1] Dell'Acqua F, Gamba P. Remote sensing and earthquake damage assessment: experiences, limits, and perspectives[J]. Proceedings of the IEEE, 2012, 100(10): 2876-2890. [2] Zhang X, Sun Y L, Shang K, et al. Crop classification based on feature band set construction and object-oriented approach using hyperspectral images[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2016, 9(9): 4117-4128. [3] Li T, Wang C Q, Meng M Q H, et al. Coverage sampling planner for UAV-enabled environmental exploration and field mapping[C]//Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. Washington, USA: IEEE, 2019: 2509-2516. [4] Marsocci V, Jia Y R, Le Bellier G, et al. PANGAEA: assessing geospatial foundation models capabilities through a global and inclusive benchmark[J]. IEEE Geoscience and Remote Sensing Magazine, 2026, 14(1): 245-285. [5] Zhu Z, Zhou Y Y, Seto K C, et al. Understanding an urbanizing planet: strategic directions for remote sensing[J]. Remote Sensing of Environment, 2019, 228: 164-182. [6] Kaku K.Satellite remote sensing for disaster management support: a holistic and staged approach based on case studies in Sentinel Asia[J]. International Journal of Disaster Risk Reduction, 2019, 33: 417-432. [7] Boccardo P, Giulio Tonolo F. Remote sensing role in emergency mapping for disaster response[M]//Lollino G, Manconi A, Gu-zzetti F, et al., eds. Engineering Geology for Society and Territory. Berlin, Germany: Springer, 2015, V: 17-24. [8] Lei T J, Wang J B, Li X Y, et al. Flood disaster monitoring and emergency assessment based on multi-source remote sensing observations[J/OL]. Water, 2022, 14(14). https://doi.org/10.3390/w14142207. [9] Xia G S, Hu J W, Hu F, et al. AID: a benchmark data set for performance evaluation of aerial scene classification[J]. IEEE Transactions on Geoscience and Remote Sensing, 2017, 55(7): 3965-3981. [10] Li K, Wan G, Cheng G, et al. Object detection in optical remote sensing images: a survey and a new benchmark[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 159: 296-307. [11] Lee D H, Hong J H, Seo H W, et al. KFGOD: a fine-grained object detection dataset in KOMPSAT satellite imagery[J/OL]. Remote Sensing, 2025, 17(22). https://doi.org/10.3390/rs17223774. [12] Mei S H, Lian J W, Wang X F, et al. A comprehensive study on the robustness of deep learning-based image classification and object detection in remote sensing: surveying and benchmarking[J/OL]. Journal of Remote Sensing, 2024, 4. https://spj.science.org/doi/epdf/10.34133/remotesensing.0219. [13] Wang J J, Zheng Z, Ma A L, et al. LoveDA: a remote sensing land-cover dataset for domain adaptive semantic segmentation[EB/OL]. [2026-04-23]. https://arxiv.org/pdf/2110.08733. [14] Boguszewski A, Batorski D, Ziemba-Jankowska N, et al. LandCover.ai: dataset for automatic mapping of buildings, woodlands, water and roads from aerial imagery[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Washington, USA: IEEE, 2021: 1102-1110. [15] Xia J S, Yokoya N, Adriano B, et al. OpenEarthMap: a benchmark dataset for global high-resolution land cover mapping[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Washington, USA: IEEE, 2023: 6254-6264. [16] Li X, Ding J, Elhoseiny M. VRSBench: a versatile vision-language benchmark dataset for remote sensing image understanding[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2024: 3229-3242. [17] Luo J W, Pang Z, Zhang Y J, et al. SkySenseGPT: a fine-grained instruction tuning dataset and model for remote sensing vision-language understanding[EB/OL].[2026-04-23]. https://arxiv.org/pdf/2406.10100. [18] Shao H, Qian S J, Xiao H, et al. Visual CoT: advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2024: 8612-8642. [19] Muhtar D, Li Z S, Gu F, et al. LHRS-Bot: empowering remote sensing with VGI-enhanced large multimodal language model[C]//Proceedings of the 18th European Conference on Computer Vision. Berlin, Germany: Springer, 2024: 440-457. [20] Li K, Dong F Y, Wang D, et al. Show me what and where has changed question answering and grounding for remote sensing change detection[EB/OL].[2026-04-23]. https://arxiv.org/pdf/2410.23828. [21] Wang F X, Wang H Z, Guo Z H, et al. XLRS-Bench: could your multimodal LLMs understand extremely large ultra-high-resolution remote sensing imagery[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2025: 14325-14336. [22] Yang M, Zhou Z, Tian S Y, et al. NeSy-Route: a neuro-symbolic benchmark for constrained route planning in remote sensing[EB/OL]. [2026-04-23]. https://arxiv.org/pdf/2603.16307. [23] Gemini Team. Gemini: a family of highly capable multimodal mo-dels[EB/OL]. [2026-04-23]. https://arxiv.org/pdf/2312.11805. [24] Open AI. OpenAI GPT-5 system card[EB/OL]. [2026-04-23]. https://arxiv.org/pdf/2601.03267. [25] Qwen Team.Qwen3-VL technical report[EB/OL]. [2026-04-23]. https://arxiv.org/pdf/2511.21631. [26] An X, Xie Y, Yang K C, et al. LLaVA-OneVision-1.5: fully open framework for democratized multimodal training[EB/OL]. [2026-04-23]. https://arxiv.org/pdf/2509.23661. [27] Wang W Y, Gao Z W, Gu L X, et al. InternVL3.5: advancing open-source multimodal models in versatility, reasoning, and efficiency[EB/OL]. [2026-04-23]. https://arxiv.org/pdf/2508.18265. [28] Qwen Team.Qwen3.5: accelerating productivity with native multimodal agents[EB/OL]. [2026-04-23]. https://qwen.ai/blog?id=qwen3.5. [29] Kuckreja K, Danish M S, Naseer M, et al. GeoChat: grounded large vision-language model for remote sensing[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2024: 27831-27840. [30] Cheng B W, Misra I, Schwing A G, et al. Masked-attention mask transformer for universal image segmentation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2022: 1280-1289. [31] Li K Y, Liu R X, Cao X Y, et al. SegEarth-OV: towards training-free open-vocabulary segmentation for remote sensing images[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2025: 10545-10556. [32] Xie E Z, Wang W H, Yu Z D, et al. SegFormer: simple and efficient design for semantic segmentation with transformers[C]//Proceedings of the 35th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2021: 12077-12090. [33] Hu R Z, Bai S Y, Wen W S, et al. Towards high-definition vector map construction based on multi-sensor integration for intelligent vehicles: systems and error quantification[J]. IET Intelligent Transport Systems, 2024, 18(8): 1477-1493. [34] Alqobali R, Alshmrani M, Alnasser R, et al. A survey on robot semantic navigation systems for indoor environments[J/OL]. Applied Sciences, 2024, 14(1). https://doi.org/10.3390/app14010089. [35] Hart P E, Nilsson N J, Raphael B.A formal basis for the heuristic determination of minimum cost paths[J]. IEEE Transactions on Systems Science and Cybernetics, 1968, 4(2): 100-107. [36] Daniel K, Nash A, Koenig S, et al. Theta*: any-angle path pla-nning on grids[J]. Journal of Artificial Intelligence Research, 2010, 39(1): 533-579.