Textual Causal Feature Selection via Contrastive Learning and Large Language Model
YANG Jiahao1,2, YANG Juntao1,2, CAO Dayuan1,2, XIONG Shoujiu1,2, YU Kui1,2
1. School of Computer Science and Information Engineering, Hefei University of Technology, Hefei 230601; 2. Key Laboratory of Knowledge Engineering with Big Data of Ministry of Education, Hefei University of Technology, Hefei 230601
摘要 从文本数据中学习与目标变量相关的因果特征,是提升模型可解释性与决策可靠性的重要途径之一.然而,文本因果特征选择仍面临两方面挑战:1)潜在特征通常隐含于非结构化文本中,且句子表示常混合潜在特征语义与情感信息,导致潜在特征提取结果不稳定;2)潜在特征在提取、赋值与补充过程中动态变化,难以实现稳定高效的因果特征选择.针对上述问题,文中提出基于对比学习与大语言模型的文本因果特征选择方法(Textual Causal Feature Selection via Contrastive Learning and Large Language Model, TCFS).首先,针对潜在特征提取不稳定的问题,设计解耦潜在特征提取方法,并结合迭代聚类优化,构建语义一致、结构稳定的潜在特征簇.然后,针对因果特征动态变化带来的选择困难,提出动态因果特征选择方法,动态维护与目标变量相关的马尔可夫毯,实现关键因果特征的稳定识别与冗余剔除.最后,通过潜在特征赋值将非结构化文本转化为结构化数值数据,并引入缺失特征发现机制,对难以解释的样本进行分析与迭代补充,持续优化潜在特征集合与特征选择结果.实验表明,TCFS能有效提升文本场景中潜在特征提取的稳定性和因果特征选择的准确性,为可解释预测与文本因果分析提供有效支持.
Abstract:Learning causal features related to a target variable from textual data is considered as an important approach to improve model interpretability and decision reliability. However, textual causal feature selection still faces two challenges. Latent features are usually embedded in unstructured text, and latent feature semantics are often mixed with sentiment information in sentence representations. Consequently, latent feature extraction results are unstable. Stable and efficient causal feature selection becomes difficult due to the dynamic changes of latent features during extraction, assignment and supplementation. To address these issues, a method for textual causal feature selection via contrastive learning and large language model(TCFS) is proposed in this paper. First, a decoupled latent feature extraction method is designed to address the instability of latent feature extraction, and iterative clustering optimization is introduced to construct semantically consistent and structurally stable latent feature clusters. Second, a dynamic causal feature selection method is proposed, and the stable identification of key causal features and the removal of redundant features are achieved by dynamically maintaining the Markov blanket associated with the target variable. Finally, unstructured text is transformed into structured numerical data through latent feature assignment, and a missing-feature discovery mechanism is introduced to analyze hard-to-explain samples and iteratively supplement the missing features, thereby continuously optimizing the latent feature set and the feature selection results. Experiments demonstrate that TCFS effectively improves the stability of latent feature extraction and the accuracy of causal feature selection in text scenarios and provides effective support for interpretable prediction and textual causal analysis.
杨嘉豪, 杨俊涛, 曹大元, 熊寿久, 俞奎. 基于对比学习与大语言模型的文本因果特征选择方法[J]. 模式识别与人工智能, 2026, 39(8): 665-682.
YANG Jiahao, YANG Juntao, CAO Dayuan, XIONG Shoujiu, YU Kui. Textual Causal Feature Selection via Contrastive Learning and Large Language Model. Pattern Recognition and Artificial Intelligence, 2026, 39(8): 665-682.
[1] Ling Z L, Li Y, Zhang Y W, et al. A light causal feature selection approach to high-dimensional data[J]. IEEE Transactions on Knowledge and Data Engineering, 2023, 35(8): 7639-7650. [2] Yu K, Liu L, Li J Y.A unified view of causal and non-causal feature selection[J]. ACM Transactions on Knowledge Discovery from Data, 2021, 15(4): 1-46. [3] Pekar V, Candi M, Beltagui A, et al. Explainable text-based features in predictive models of crowdfunding campaigns[J]. Annals of Operations Research, 2025, 354(1): 367-397. [4] Han Y, Bruggeman R, Peper J, et al.Extracting latent needs from online reviews through deep learning based language model[C]//Proceedings of the International Conference on Engineering Design. Cambridge, UK: Cambridge University Press, 2023: 1855-1864. [5] Kamthawee K, Udomcharoenchaikit C, Nutanong S.MIST: mutual information maximization for short text clustering[C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics(Long Papers). Stroudsburg, USA: ACL, 2024: 11309-11324. [6] Liu C X, Chen Y Q, Liu T L, et al.Discovery of the hidden world with large language models[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems. Cambridge, USA: MIT Press, 2024: 102307-102365. [7] Zhu L X, Zhao R C, Pergola G, et al. Disentangling aspect and stance via a Siamese autoencoder for aspect clustering of vaccination opinions[C]//Findings of the Association for Computational Linguistics. Stroudsburg, USA: ACL, 2023: 1827-1842. [8] Wang Y X, Cao F Y, Yu K, et al. Federated causal structure lear-ning with non-identical variable sets[C]//Proceedings of the 42nd International Conference on Machine Learning. San Diego, USA: JMLR, 2025: 62445-62466. [9] Marutho D, Muljono, Rustad S,et al. Optimizing aspect-based sentiment analysis using sentence embedding transformer, Bayesian search clustering, and sparse attention mechanism[J/OL]. Journal of Open Innovation: Technology, Market, and Complexity, 2024, 10(1). https://doi.org/10.1016/j.joitmc.2024.100211. [10] Xu E J, Tang B, Liu X, et al. Automatic Aspect-Based Sentiment Analysis(AABSA) from Customer Reviews[EB/OL].[2026-03-03]. https://ceur-ws.org/Vol-2614/AffCon20_session1_automatically.pdf. [11] 闫小强,卢耀恩,娄铮铮,等.基于并行信息瓶颈的多语种文本聚类算法[J].模式识别与人工智能, 2017, 30(6): 559-568. (Yan X Q, Lu Y E, Lou Z Z, et al. Multilingual documents clustering algorithm based on parallel information bottleneck[J]. Pa-ttern Recognition and Artificial Intelligence, 2017, 30(6): 559-568.) [12] 路荣,项亮,刘明荣,等.基于隐主题分析和文本聚类的微博客中新闻话题的发现[J].模式识别与人工智能, 2012, 25(3): 382-387. (Lu R, Xiang L, Liu M R, et al. Discovering News Topics from Microblogs Based on Hidden Topics Analysis and Text Clustering[J]. Pattern Recognition and Artificial Intelligence, 2012, 25(3): 382-387.) [13] Li B H, Zhou H, He J X, et al. On the sentence embeddings from pre-trained language models[C]//Proceedings of the Conference on Empirical Methods in Natural Language Processing. Stroudsburg, USA: ACL, 2020: 9119-9130. [14] Pang-Naylor K, Manivasagan S, Zhong A T, et al. Controllable clustering with LLM-driven embeddings[C]//Proceedings of the Conference on Empirical Methods in Natural Language Processing(Industry Track). Stroudsburg, USA: ACL, 2025: 686-702. [15] Jiang T, Huang S H, Luan Z Z, et al. Scaling sentence embe-ddings with large language models[C]//Proceedings of the Confe-rence on Empirical Methods in Natural Language Processing. Stroudsburg, USA: ACL, 2024: 3182-3196. [16] Zhao R C, Gui L, He Y L.Cone: unsupervised contrastive opi-nion extraction[C]//Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, USA: ACM, 2023: 1066-1075. [17] Colombo D, Maathuis M H.Order-independent constraint-based cau-sal structure learning[J]. The Journal of Machine Learning Research, 2014, 15(1): 3741-3782. [18] Spirtes P, Glymour C N, Scheines R.Causation, prediction, and search[M]. Cambridge, USA: MIT Press, 2000. [19] Guo X J, Yu K, Cao F Y, et al. Error-aware Markov blanket learning for causal feature selection[J]. Information Sciences, 2022, 589: 849-877. [20] 吴兴宇,江兵兵,吕胜飞,等.基于马尔科夫边界发现的因果特征选择算法综述[J].模式识别与人工智能, 2022, 35(5): 422-438. (Wu X Y, Jiang B B, Lü S F, et al. A survey on causal feature selection based on Markov boundary discovery[J]. Pattern Recognition and Artificial Intelligence, 2022, 35(5): 422-438.) [21] Ling Z L, Guo M X, Wu X Y, et al. Gradient-based causal feature selection[C]//Proceedings of the 34th International Joint Conference on Artificial Intelligence. San Francisco, USA: IJCAI,2025: 5716-5722. [22] Suzuki H, Kanamori K, Takagi T, et al. I-CAM-UV: integrating causal graphs over non-identical variable sets using causal additive models with unobserved variables[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2026, 40(30): 25745-25752. [23] Wang F, Wang X H, Wei W,et al. An incremental feature selection approach for dynamic feature variation[J/OL]. Neurocompu-ting, 2024, 570. https://doi.org/10.1016/j.neucom.2023.127138. [24] Zhuo S D, Qiu J J, Wang C D, et al. Online feature selection with varying feature spaces[J]. IEEE Transactions on Knowledge and Data Engineering, 2024, 36(9): 4806-4819. [25] Wu X D, Yu K, Wang H, et al. Online streaming feature selection[EB/OL].[2026-03-03]. https://icml.cc/Conferences/2010/papers/238.pdf. [26] Wu X D, Yu K, Ding W, et al. Online feature selection with streaming features[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013, 35(5): 1178-1192. [27] Zhou J, Foster D, Stine R, et al. Streaming feature selection using alpha-investing[C]//Proceedings of the 11th ACM SIGKDD International Conference on Knowledge Discovery in Data Mining. New York, USA: ACM, 2005: 384-393.