Target Speaker Identification Based on EEG Auditory Attention Cue Disentanglement and Cross-Modal Contrastive Learning
LING Xilei1, WANG Haixian1
1. Key Laboratory of Child Development and Learning Science of Ministry of Education, School of Biological Science and Medical Engineering, Southeast University, Nanjing 211189
Abstract:In complex acoustic environments, hearing-impaired listeners often struggle to reliably track a target speech stream from multi-talker mixtures. Conventional hearing aids can amplify speech, but fail to capture the listener's subjective attentional focus. Personalized hearing assistance requirements in realistic settings are therefore difficult to be met. Auditory attention decoding based on electroencephalogram(EEG) signals is a promising paradigm for neuro-steered hearing assistance. In this paper, a target speaker identification method based on EEG auditory attention cue disentanglement and cross-modal contrastive learning(DCLNet) is proposed for the auditory attention decoding problem. Motivated by the hierarchical neural processing mechanism of auditory attention, a static-dynamic disentangled representation learning framework is adopted. The EEG signal is factorized into a static latent variable for encoding speaker identity and a dynamic latent sequence for capturing time-varying auditory tracking cues. Then, a pre-trained speaker encoder is further incorporated as an external identity prior, and a speaker-aware bidirectional cross-modal contrastive learning is employed to align the identity representations between EEG and speech modalities. Experiments on the DTU and KUL datasets demonstrate that DCLNet consistently improves target speaker identification performance and yields more discriminative speaker-identity representations in EEG.
[1] Cherry E C. Some experiments on the recognition of speech, with one and with two ears[J]. Journal of the Acoustical Society of America, 1953, 25(5): 975-979. [2] Wu Y H, Stangl E, Chipara O, et al. Efficacy and effectiveness of advanced hearing aid directional and noise reduction technologies for older adults with mild to moderate hearing loss[J]. Ear and Hea-ring, 2019, 40(4): 805-822. [3] Favre-Félix A, Graversen C, Bhuiyan T A, et al. Absolute eye gaze estimation with biosensors in hearing aids[J/OL]. Frontiers in Neu-roscience, 2019, 13. https://doi.org/10.3389/fnins.2019.01294. [4] Gogate M, Dashtipour K, Adeel A, et al. CochleaNet: a robust language-independent audio-visual model for real-time speech enhancement[J]. Information Fusion, 2020, 63: 273-285. [5] O'Sullivan J, Chen Z, Herrero J, et al. Neural decoding of attentional selection in multi-speaker environments without access to clean sources[J/OL]. Journal of Neural Engineering, 2017, 14(5). https://iopscience.iop.org/article/10.1088/1741-2552/aa7ab4/pdf. [6] Fiedler L, Wöstmann M, Graversen C, et al. Single-channel in-ear-EEG detects the focus of auditory attention to concurrent tone streams Single-channel in-ear-EEG detects the focus of auditory attention to concurrent tone streams and mixed speech[J/OL]. Journal of Neural Engineering, 2017, 14(3). https://iopscience.iop.org/article/10.1088/1741-2552/aa66dd/pdf. [7] Gevins A, Smith M E, McEvoy L K, et al. Electroencephalographic imaging of higher brain function[J]. Philosophical Transactions of the Royal Society of London: Biological Sciences, 1999, 354(1387): 1125-1134. [8] Pu W Q, Xiao J J, Zhang T, et al. A joint auditory attention decoding and adaptive binaural beamforming algorithm for hearing devices[C]//Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. Washington, USA: IEEE, 2019: 311-315. [9] Aroudi A, Doclo S. Cognitive-driven binaural beamforming using EEG-based auditory attention decoding[J]. IEEE/ACM Transac-tions on Audio, Speech, and Language Processing, 2020, 28: 862-875. [10] Geirnaert S, Francart T, Bertrand A. Fast EEG-based decoding of the directional focus of auditory attention using common spatial pa-tterns[J]. IEEE Transactions on Biomedical Engineering, 2021, 68(5): 1557-1568. [11] Han C, O'Sullivan J, Luo Y, et al. Speaker-independent auditory attention decoding without access to clean speech sources[J/OL]. Science Advances, 2019, 5(5). https://www.science.org/doi/epdf/10.1126/sciadv.aav6134. [12] Borgström B J, Brandstein M S, Ciccarelli G A, et al. Speaker separation in realistic noise environments with applications to a cognitively-controlled hearing aid[J]. Neural Networks, 2021, 140: 136-147. [13] Han C, Choudhari V, Li Y H, et al. Improved decoding of attentional selection in multi-talker environments with self-supervised learned speech representation[C/OL]//Proceedings of the 45th Annual International Conference of the IEEE Engineering in Medi-cine & Biology Society.2023. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10340191. [14] O'Sullivan J A, Power A J, Mesgarani N, et al. Attentional selec-tion in a cocktail party environment can be decoded from single-trial EEG[J]. Cerebral Cortex, 2015, 25(7): 1697-1706. [15] Biesmans W, Das N, Francart T, et al. Auditory-inspired speech envelope extraction methods for improved EEG-based auditory attention detection in a cocktail party scenario[J]. IEEE Transac-tions on Neural Systems and Rehabilitation Engineering, 2017, 25(5): 402-412. [16] de Cheveigné A, Wong D D E, Di Liberto G M, et al. Decoding the auditory brain with canonical component analysis[J]. Neuroimage, 2018, 172: 206-216. [17] de Taillez T, Kollmeier B, Meyer B T. Machine learning for decoding listeners' attention from electroencephalography evoked by continuous speech[J]. European Journal of Neuroscience, 2020, 51(5): 1234-1241. [18] Kuruvila I, Muncke J, Fischer E, et al. Extracting the auditory attention in a dual-speaker scenario from EEG using a joint CNN-LSTM model[J/OL]. Frontiers in Physiology, 2021, 12. https://doi.org/10.3389/fphys.2021.700655. [19] Cai S Q, Li P W, Su E Z, et al. Auditory attention detection via cross-modal attention[J/OL]. Frontiers in Neuroscience, 2021, 15. https://doi.org/10.3389/fnins.2021.652058. [20] Cai S Q, Su E Z, Xie L H, et al. EEG-based auditory attention detection via frequency and channel neural attention[J]. IEEE Transactions on Human-Machine Systems, 2022, 52(2): 256-266. [21] Chen X Y, Du C D, Zhou Q Y, et al. Auditory attention decoding with task-related multi-view contrastive learning[C]//Proceedings of the 31st ACM International Conference on Multimedia. New York, USA: ACM, 2023: 6025-6033. [22] Ding Y T, Wang L, Lu J, et al. EEG-based auditory attention decoding for speaker identification under mixed-speech hearing-assistive conditions[J/OL]. IEEE Transactions on Biomedical Engineering, 2025. https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11311555. [23] Chen Y B, Chen N, Niu Y X, et al. Neuroscience prior knowledge guided EEG representation disentanglement for auditory attention decoding[J/OL]. Hearing Research, 2026, 475. https://doi.org/10.1016/j.heares.2026.109625. [24] Desplanques B, Thienpondt J, Demuynck K. ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification[EB/OL]. [2026-04-20]. https://arxiv.org/pdf/2005.07143. [25] Bai J W, Wang W R, Gomes C. Contrastively disentangled sequential variational autoencoder[C]//Proceedings of the 35th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2021: 10105-10118. [26] Lawhern V J, Solon A J, Waytowich N R, et al. EEGNet: a compact convolutional network for EEG-based brain-computer interfaces[J/OL]. Journal of Neural Engineering, 2018, 15(5). https://iopscience.iop.org/article/10.1088/1741-2552/aace8c/pdf. [27] Chou J C, YEH C C, LEE H Y. One-shot voice conversion by separating speaker and content representations with instance normalization[EB/OL]. [2026-04-20]. https://arxiv.org/pdf/1904.05742. [28] Kim D, Chang J H. Improving target sound extraction via disentangled codec representations with privileged knowledge distillation[EB/OL]. [2026-04-20]. https://openreview.net/pdf?id=rew03VaNUJ. [29] Cheng P Y, Hao W T, Dai S Y, et al. CLUB: A contrastive log-ratio upper bound of mutual information[C]//Proceedings of the 37th International Conference on Machine Learning. San Diego, USA: JMLR, 2020: 1779-1788. [30] Li Y G, Liang F, Zhao L C, et al. Supervision exists everywhere: a data efficient contrastive language-image pre-training paradigm[EB/OL].[2026-04-20]. https://arxiv.org/pdf/2110.05208. [31] Fuglsang S A, Dau T, Hjortkjær J. Noise-robust cortical tracking of attended speech in real-world acoustic scenes[J]. NeuroImage, 2017, 156: 435-444. [32] Liao Y, Zhang Y H, Han Q S, et al. Frequency-based alignment of EEG and audio signals using contrastive learning and SincNet for auditory attention detection[EB/OL].[2026-04-20]. https://arxiv.org/pdf/2503.04156. [33] Weston P S J, Hunter M D, Sokhi D S, et al. Discrimination of voice gender in the human auditory cortex[J]. NeuroImage, 2015, 105: 208-214. [34] Zakaria M N, Wahab N A A, Maamor N, et al. Auditory brainstem response(ABR) findings in males and females with comparable head sizes at supra-threshold and threshold levels[J]. Neurology, Psychiatry and Brain Research, 2019, 32: 4-7. [35] van der Ryck I, Heintz N, Rotaru I, et al. Effects of speaker and listener sex on auditory attention decoding performance[J/OL]. Hearing Research, 2025, 468. https://doi.org/10.1016/j.heares.2025.109444. [36] Rossi A F, Pessoa L, Desimone R, et al. The prefrontal cortex and the executive control of attention[J]. Experimental Brain Research, 2009, 192(3): 489-497. [37] Hill K T, Miller L M. Auditory attentional control and selection during cocktail party listening[J]. Cerebral Cortex, 2010, 20(3): 583-590. [38] Hötting K, Rösler F, Röder B. Crossmodal and intermodal atten-tion modulate event-related brain potentials to tactile and auditory stimuli[J]. Experimental Brain Research, 2003, 148(1): 26-37.