In complex acoustic environments, hearing-impaired listeners often struggle to reliably track a target speech stream from multi-talker mixtures. Conventional hearing aids can amplify speech, but fail to capture the listener's subjective attentional focus. Personalized hearing assistance requirements in realistic settings are therefore difficult to be met. Auditory attention decoding based on electroencephalogram(EEG) signals is a promising paradigm for neuro-steered hearing assistance. In this paper, a target speaker identification method based on EEG auditory attention cue disentanglement and cross-modal contrastive learning(DCLNet) is proposed for the auditory attention decoding problem. Motivated by the hierarchical neural processing mechanism of auditory attention, a static-dynamic disentangled representation learning framework is adopted. The EEG signal is factorized into a static latent variable for encoding speaker identity and a dynamic latent sequence for capturing time-varying auditory tracking cues. Then, a pre-trained speaker encoder is further incorporated as an external identity prior, and a speaker-aware bidirectional cross-modal contrastive learning is employed to align the identity representations between EEG and speech modalities. Experiments on the DTU and KUL datasets demonstrate that DCLNet consistently improves target speaker identification performance and yields more discriminative speaker-identity representations in EEG.
To address the problem of drift accumulation and insufficient loop-closure robustness of traditional SLAM methods in large-scale, dynamic and visually degraded scenes, a mobile robot SLAM method inspired by cognitive mechanisms of visual cortex and entorhinal-hippocampal system(VEH-SLAM) is proposed. Based on the pose estimation of entorhinal-hippocampal path integration model, VEH-SLAM is designed to construct a view-cell-inspired episodic memory network. Maximum a posteriori estimation and a posterior entropy-driven adaptive threshold mechanism are utilized to significantly improve the accuracy and robustness of loop-closure detection. In addition, an efficient cognitive map learning algorithm is developed based on the basic map learning module. Spatial-density pruning and neural-scale stretching are integrated into the algorithm. The map size is effectively controlled, and the encoding resolution is adaptively adjusted. Therefore, mapping speed and accuracy are balanced. VEH-SLAM is shown to achieve the highest recall at 100% precision on five of the seven public datasets. In terms of mapping, the total time is significantly reduced, and the lowest map errors are obtained on some sequences. Real-time performance and engineering feasibility are further verified through indoor and outdoor physical experiments.
Triadic concept analysis, as a three-dimensional extension of formal concept analysis, takes triadic concepts as its basic units to enable knowledge discovery in three-dimensional data. When the data change dynamically, the update of triadic concepts is inevitable. However, the recomputation of all triadic concepts is time-consuming. Since dynamic changes in a triadic context essentially involve the alterations in triadic relations, a method for triadic concept updating when an object-attribute-condition triple is added to the triadic context is proposed. First, eight distinct scenarios of the newly added triple are classified into three updating types. Then, a formal context induced by the triadic context and its formal concepts are introduced and combined with the candidate set. The updating methods for triadic concepts in various cases are presented. Finally, an algorithm for updating triadic concepts is developed, and experiments are conducted to verify the effectiveness of the proposed algorithm.
The high macroscopic visual similarity between authentic and fake patina on ancient coins limits the discriminative capability of traditional spatial-domain models. Existing methods struggle to solve the continuous scoring problem of the physical evolution degree of patina. To address the above issues, a Chinese ancient coin patina authentication method based on frequency-domain features and multi-task learning(FDF-MTL) is proposed. First, a frequency-domain feature extraction and enhancement module is designed. Patch-based discrete cosine transform is adopted to capture the essential differences in frequency energy distribution between authentic patina and fake patina. A dual-path pooling strategy is utilized to perceive abnormal frequencies. Second, a multi-scale spatial-frequency modulation fusion module is constructed to transform frequency-domain features into attention weights for modulating spatial-domain features, thereby avoiding semantic conflicts. Finally, a fused feature screening module is designed to suppress the interference of coin inscription contours, and a multi-task learning framework is constructed to solve the scoring problem. The joint optimization of authenticity discrimination and continuous patina score regression is achieved. Experiments on a self-built dataset demonstrate the superior performance of the proposed method in both authenticity discrimination and patina scoring tasks.
Existing temporal knowledge graph completion models struggle to effectively model the deep interactions between semantic information and multi-granularity temporal information and lack explicit the modeling of temporal sensitivity of historical information. To solve these problems, an approach for quaternion-based representation and explicit historical enhancement for temporal knowledge graph completion(QR-EH) is proposed. First, the semantic information of entities and relations, and the multi-granularity temporal information are mapped into the quaternion space. Deep interactions between semantics and time are achieved through Hamiltonian products. Then, an explicit historical retrieval module is designed. A temporal modulation mechanism is defined by this module. Time-sensitive historical patterns are accurately captured by weighting historical recurring events. Finally, an adaptive score fusion module is constructed. Global spatio-temporal information and historical recurring information are weighted and fused by this module. The final prediction results are generated. Experiments on three publicly available datasets demonstrate that QR-EH deeply mines temporal evolution and historical recurrence patterns and exhibits good generalization and interpretability.
Existing RGB-D salient object detection methods predominantly employ homogenized fusion strategies and neglect the inherent discrepancies across spatial and semantic dimensions. To address these issues, a global fusion and progressive decoding network for RGB-D salient object detection is proposed in this paper. An asymmetric dual-stream encoder architecture is utilized, and differentiated backbone networks tailored to the distinct information densities of RGB and depth modalities are designed. First, a detail-aware module is developed to replace traditional symmetric fusion. Global RGB texture features are employed as spatial guidance for the asymmetric filtering of low-quality depth features with fine boundaries preserved and sensor noise suppressed. Then, a semantic aggregation module is developed to construct parallel multi-scale receptive fields. A channel attention mechanism is introduced for cross-modal semantic purification, optimizating precise localization and semantic alignment for multi-scale targets. Finally, a refined decoding module is constructed to cooperate with large receptive fields to extract channel attention through a dual-branch residual structure, effectively compensating for information loss and achieving a progressive reconstruction of salient objects from high-level semantics to low-level details. Experiments on six public datasets demonstrate that the proposed network achieves superior detection performance.