[1]汤龙,雷震春.基于通道注意力和特征融合的伪造语音检测研究[J].计算机技术与发展,2025,(10):131-138.[doi:10.20165/j.cnki.ISSN1673-629X.2025.0136]
 TANG Long,LEI Zhen-chun.Research on Deepfake Speech Detection Based on Channel Attention and Feature Fusion[J].,2025,(10):131-138.[doi:10.20165/j.cnki.ISSN1673-629X.2025.0136]
点击复制

基于通道注意力和特征融合的伪造语音检测研究

《计算机技术与发展》[ISSN:1006-6977/CN:61-1281/TN]

卷:
期数:
2025年10期
页码:
131-138
栏目:
人工智能
出版日期:
2025-10-10

文章信息/Info

Title:
Research on Deepfake Speech Detection Based on Channel Attention and Feature Fusion
文章编号:
1673-629X(2025)10-0131-08
作者:
汤龙雷震春
江西师范大学计算机信息工程学院,江西南昌 360111
Author(s):
TANG LongLEI Zhen-chun
School of Computer and Information Engineering,Jiangxi Normal University,Nanchang 360111,China
关键词:
伪造语音检测对数高斯概率特征通道注意力深度卷积多尺度上下文信息特征融合
Keywords:
deepfake speech detectionlog-Gaussian probability featureschannel attentiondepth-wise convolutionmulti-scale context informationfeature fusion
分类号:
TP391
DOI:
10.20165/j.cnki.ISSN1673-629X.2025.0136
摘要:
随着深度学习技术的迅猛发展,语音伪造技术对自动说话人验证系统的安全性构成严峻挑战,语音伪造检测系统依旧面临准确率不足、场景单一等问题。该文提出了一种结合通道注意力和特征融合的伪造语音检测方法,以解决语音伪造检测系统面临的一系列问题。为了聚集丰富的上下文信息和融合尺度不一致的特征,该文提出了双分支通道注意力模块,利用深度卷积沿通道维度聚合多尺度上下文信息,同时在两个分支上捕捉全局和局部特征信息;然后提出了注意力特征融合模块,将LFCC特征经过真实语音GMM和欺骗语音GMM得到对数高斯概率特征,随后基于注意力进行特征融合以学习具有通道上下文信息和全局局部特征信息的交互特征,解决了特征融合机制场景单一的问题。与基线系统相比,文中最佳系统AFF-ResNet在ASVSpoof2021LA数据集上的EER和min t-DCF分别降低37.5%和15.3%。实验结果表明,该方法显著提升了语音欺骗检测的准确率。
Abstract:
With the rapid development of deep learning technologies,speech spoofing techniques pose serious threats to the security of au-tomatic speaker verification systems. Speech spoofing detection systems still face challenges such as insufficient accuracy and limited ap-plication scenarios. We propose a speech spoofing detection method combining channel attention and feature fusion to address the challenges faced by deepfake speech detection systems. To gather rich contextual information and fuse features with inconsistent scales,a dual-branch channel attention module is proposed,which uses depth-wise convolution along the channel dimension to aggregate multi-scale contextual information,while capturing both global and local feature information in two separate branches. Then,an attention feature fusion module is proposed,where the LFCC features are processed through the bonafide speech GMM and spoofed speech GMM to obtain Log-Gaussian probability features. Attention-based feature fusion is then applied to learn interactive features that incorporate both channel context information and global- local feature information, addressing the issue of a single scenario in the feature fusion mechanism. Compared to the baseline system,the proposed best system,AFF-ResNet,reduces the EER and min t-DCF by 37.5% and 15. 3%,respectively,on the ASVSpoof2021LA dataset. Experimental results show that the proposed method significantly improves the accuracy of speech spoofing detection.
更新日期/Last Update: 2025-10-10