Mobile Robot SLAM Method Inspired by Cognitive Mechanisms of Visual Cortex and Entorhinal-Hippocampal System
LIAO Yishen1, YU Naigong2, YU Hejie3, WANG Chenghua1
1. School of Virtual Reality and Modern Industry, Jiangxi University of Finance and Economics, Nanchang 330032; 2. School of Information Science and Technology, Beijing University of Technology, Beijing 100124; 3. Department of Precision Instrument, Tsinghua University, Beijing 100084
摘要 针对传统SLAM方法在大规模、动态及视觉退化场景中存在的漂移累积、闭环鲁棒性不足等问题,文中提出受视皮层与内嗅-海马认知机理启发的移动机器人SLAM方法(Mobile Robot SLAM Method Inspired by Cognitive Mechanisms of Visual Cortex and Entorhinal-Hippocampal System, VEH-SLAM).首先,以内嗅-海马路径积分模型的位姿估计为基础,构建受视图细胞机制启发的情景记忆网络,利用最大后验估计与后验熵驱动的自适应阈值机制,显著提升闭环检测的准确性与鲁棒性.然后,在基础地图学习模块的基础上,提出融合空间密度修剪与神经尺度伸缩的高效认知地图学习算法,旨在有效控制地图规模的同时自适应调节编码分辨率,从而兼顾建图的快速性与精确性.在7个公开数据集上的实验表明,VEH-SLAM在其中5个数据集上100%精确度下最大召回率取得最优值.在建图方面,方法大幅降低总耗时,并在部分序列上获得最低地图误差.室内外物理实验进一步验证其实时性与工程可行性.
Abstract:To address the problem of drift accumulation and insufficient loop-closure robustness of traditional SLAM methods in large-scale, dynamic and visually degraded scenes, a mobile robot SLAM method inspired by cognitive mechanisms of visual cortex and entorhinal-hippocampal system(VEH-SLAM) is proposed. Based on the pose estimation of entorhinal-hippocampal path integration model, VEH-SLAM is designed to construct a view-cell-inspired episodic memory network. Maximum a posteriori estimation and a posterior entropy-driven adaptive threshold mechanism are utilized to significantly improve the accuracy and robustness of loop-closure detection. In addition, an efficient cognitive map learning algorithm is developed based on the basic map learning module. Spatial-density pruning and neural-scale stretching are integrated into the algorithm. The map size is effectively controlled, and the encoding resolution is adaptively adjusted. Therefore, mapping speed and accuracy are balanced. VEH-SLAM is shown to achieve the highest recall at 100% precision on five of the seven public datasets. In terms of mapping, the total time is significantly reduced, and the lowest map errors are obtained on some sequences. Real-time performance and engineering feasibility are further verified through indoor and outdoor physical experiments.
[1] Xu K, Hao Y F, Yuan S H, et al. AirSLAM: an efficient and illumination-robust point-line visual SLAM system[J]. IEEE Transactions on Robotics, 2025, 41: 1673-1692. [2] Strasdat H, Montiel J M M, Davison A J. Visual SLAM: why filter[J]. Image and Vision Computing, 2012, 30(2): 65-77. [3] Cadena C, Carlone L, Carrillo H, et al. Past, present, and future of simultaneous localization and mapping: toward the robust-perception age[J]. IEEE Transactions on Robotics, 2016, 32(6): 1309-1332. [4] Fenton A A. Remapping revisited: how the hippocampus represents different spaces[J]. Nature Reviews Neuroscience, 2024, 25(6): 428-448. [5] Shao Q M, Chen L G, Li X W, et al. A non-canonical visual cortical-entorhinal pathway contributes to spatial navigation[J/OL]. Nature Communications, 2024, 15(1). https://www.nature.com/articles/s41467-024-48483-y.pdf. [6] Coogan T A, Burkhalter A. Hierarchical organization of areas in rat visual cortex[J]. Journal of Neuroscience, 1993, 13(9): 3749-3772. [7] Rolls E T. Hippocampal spatial view cells, place cells, and concept cells: view representations[J]. Hippocampus, 2023, 33(5): 667-687. [8] Cooper R A, Ritchey M. Progression from feature-specific brain activity to hippocampal binding during episodic encoding[J]. Journal of Neuroscience, 2020, 40(8): 1701-1709. [9] Peng J J, Throm B, Jazi M N, et al. Grid cells accurately track movement during path integration-based navigation despite switching reference frames[J]. Nature Neuroscience, 2025, 28(10): 2092-2105. [10] Etienne A S, Jeffery K J. Path integration in mammals[J]. Hippocampus, 2004, 14(2): 180-192. [11] Chen G F, King J A, Burgess N, et al. How vision and movement combine in the hippocampal place code[J]. PNAS, 2013, 110(1): 378-383. [12] Priestley J B, Bowler J C, Rolotti S V, et al. Signatures of rapid plasticity in hippocampal CA1 representations during novel experien-ces[J]. Neuron, 2022, 110(12): 1978-1992. [13] Harland B, Contreras M, Souder M, et al. Dorsal CA1 hippocampal place cells form a multi-scale representation of megaspace[J]. Current Biology, 2021, 31(10): 2178-2190. [14] Grisetti G, Stachniss C, Burgard W. Improved techniques for grid mapping with Rao-Blackwellized particle filters[J]. IEEE Transactions on Robotics, 2007, 23(1): 34-46. [15] Zhang J, Singh S. LOAM: lidar odometry and mapping in real-time[EB/OL]. [2026-05-27].https://www.roboticsproceedings.org/rss10/p07.pdf. [16] Bailey T, Nieto J, Guivant J, et al. Consistency of the EKF-SLAM algorithm[C]//Proceedings of the IEEE/RSJ International Confe-rence on Intelligent Robots and Systems. Washington, USA: IEEE, 2006: 3562-3568. [17] Mur-Artal R, Montiel J M M, Tardos J D. ORB-SLAM: a versatile and accurate monocular SLAM system[J]. IEEE Transactions on robotics, 2015, 31(5): 1147-1163. [18] Campos C, Elvira R, Gómez Rodríguez J J, et al. ORB-SLAM3: an accurate open-source library for visual, visual-inertial, and mul-timap SLAM[J]. IEEE Transactions on Robotics, 2021, 37(6): 1874-1890. [19] Cummins M, Newman P. Appearance-only SLAM at large scale with FAB-MAP 2.0[J]. International Journal of Robotics Research, 2011, 30(9): 1100-1123. [20] Milford M J, Wyeth G F. SeqSLAM: visual route-based navigation for sunny summer days and stormy winter nights[C]//Proceedings of the IEEE International Conference on Robotics and Automation. Washington, USA: IEEE, 2012: 1643-1649. [21] Garcia-Fidalgo E, Ortiz A. iBoW-LCD: an appearance-based loop-closure detection approach using incremental bags of binary words[J]. IEEE Robotics and Automation Letters, 2018, 3(4): 3051-3057. [22] Tsintotas K A, Bampis L, Gasteratos A. Tracking-DOSeqSLAM: a dynamic sequence-based visual place recognition paradigm[J]. IET Computer Vision, 2021, 15(4): 258-273. [23] Zhang B S, Xian Y L, Ma X G. LSGDDN-LCD: an appearance-based loop closure detection using local superpixel grid descriptors and incremental dynamic nodes[J/OL]. Computers and Electrical Engineering, 2024, 119. https://doi.org/10.1016/j.compeleceng.2024.109477. [24] Zhou Y H, Sun M L. A visual SLAM loop closure detection method based on lightweight Siamese capsule network[J/OL]. Scientific Reports, 2025, 15(1). https://www.nature.com/articles/s41598-025-90511-4.pdf. [25] Kümmerle R, Grisetti G, Strasdat H, et al. g2o: a general framework for graph optimization[C]//Proceedings of the IEEE International Conference on Robotics and Automation. Washington, USA: IEEE, 2011: 3607-3613. [26] Kaess M, Johannsson H, Roberts R, et al. iSAM2: incremental smoothing and mapping using the Bayes tree[J]. International Journal of Robotics Research, 2012, 31(2): 216-235. [27] Milford M J, Wyeth G F. Mapping a suburb with a single camera using a biologically inspired SLAM system[J]. IEEE Transactions on Robotics, 2008, 24(5): 1038-1053. [28] Yuan M L, Tian B, Shim V A, et al. An entorhinal-hippocampal model for simultaneous cognitive map building[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2015, 29(1): 586-592. [29] Yu N G, Zhai Y J, Yuan Y H, et al. A bionic robot navigation algorithm based on cognitive mechanism of hippocampus[J]. IEEE Transactions on Automation Science and Engineering, 2019, 16(4): 1640-1652. [30] Liang C, Zhao Y M, Yu R X, et al. ViT-RatSLAM: practical integration of vision Transformer-based visual place recognition into bio-inspired SLAM for robust topological mapping[C]//Procee-dings of the 12th International Conference on Electrical Enginee-ring, Control and Robotics. Washington, USA: IEEE, 2026: 249-255. [31] Pirozzo B M, Fernandez-Leon J A, de Paula M, et al. Rat-SLAM with neuromorphic computing[J/OL]. Neurocomputing, 2026, 684. https://doi.org/10.1016/j.neucom.2026.133572. [32] Zeng T P, Si B L. A brain-inspired compact cognitive mapping system[J]. Cognitive Neurodynamics, 2021, 15(1): 91-101. [33] Pizzino C A P, Costa R R, Mitchell D, et al. NeoSLAM: long-term SLAM using computational models of the brain[J/OL]. Sensors, 2024, 24(4). https://doi.org/10.3390/s24041143. [34] 廖诣深,于贺捷,于乃功,等.鼠脑内嗅-海马结构启发的移动机器人仿生路径积分模型[J].模式识别与人工智能, 2025, 38(10): 876-892. (Liao Y S, Yu H J, Yu N G, et al. Bionic path integration model for mobile robots inspired by entorhinal-hippocampal structure of rat brain[J]. Pattern Recognition and Artificial Intelligence, 2025, 38(10): 876-892.) [35] Oliva A, Torralba A. Modeling the shape of the scene: a holistic representation of the spatial envelope[J]. International Journal of Computer Vision, 2001, 42: 145-175. [36] Ma W J, Beck J M, Latham P E, et al. Bayesian inference with pro-babilistic population codes[J]. Nature Neuroscience, 2006, 9(11): 1432-1438. [37] Ma W J, Jazayeri M. Neural coding of uncertainty and probability[J]. Annual Review of Neuroscience, 2014, 37(1): 205-220. [38] Arroyo R, Alcantarilla P F, Bergasa L M, et al. Fast and effective visual place recognition using binary codes and disparity information[C]//Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems. Washington, USA: IEEE, 2014: 3089-3094.