融合分组建模与多维注意力机制的轻量级说话人确认方法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TN912.34

基金项目:

国家自然科学基金(62064003);认知无线电与信息处理教育部重点实验室项目(CRKL230103)


[ ]A Lightweight Speech Enhancement Method Based on Grouped Dual-path LSTM
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    近年来,深度学习在说话人确认任务中取得显著进展,但现有模型普遍存在参数规模大、时序建模能力有限、难以在边缘设备部署等问题。为此,提出一种融合分组建模与多维注意力机制的轻量级说话人确认模型。该模型基于改进的时间延迟神经网络结构,设计通道分组时序编码模块(Channel-Split Temporal Encoding Block, CSTE-Block)模块与多尺度分组残差编码(Grouped Scaled Res2Net, GSRes2Block)模块,用以增强帧级特征的时序建模与多尺度表达能力;引入说话人聚合优化模块(Speaker Embedding Refinement Module Block, SERMBlock)与多维注意力机制,提升全局上下文的建模深度与特征聚焦能力。全局嵌入通过注意力统计池化生成,并采用加性角度边界Softmax分类器与监督对比损失(Supervised Contrastive Loss)联合优化嵌入空间判别性。实验在CN-Celeb数据集上进行,结果表明,所提模型在EER和minDCF方面分别达到14.82%和0.6948,参数量较ECAPA-TDNN减少约69.4%;同时在VoxCeleb1跨语种验证集上展现良好鲁棒性。复杂度分析显示,该模型兼具高性能与高效率,具备优良的部署适应性与实际应用潜力。

    Abstract:

    In recent years, deep learning has made significant progress in speaker verification tasks. However, existing models often suffer from large parameter sizes, limited temporal modeling capability, and poor deployability on edge devices. To address these issues, a lightweight speaker verification model incorporating grouped modeling and multi-dimensional attention mechanisms is proposed. The model is built upon an improved Time-Delay Neural Network (TDNN) architecture and introduces a Channel-Split Temporal Encoding Block (CSTE-Block) and a Grouped Scaled Res2Net block (GSRes2Block) to enhance frame-level temporal modeling and multi-scale feature representation. In addition, a Speaker Embedding Refinement Module block (SERMBlock) and multi-dimensional attention are employed to improve global contextual modeling and feature focusing. A global speaker embedding is generated via attentive statistical pooling and jointly optimized using an Additive Angular Margin Softmax classifier and Supervised Contrastive Loss to enhance the discriminative power of the embedding space. Experiments on the CN-Celeb dataset demonstrate that the proposed model achieves an Equal Error Rate (EER) of 14.82% and a minimum Detection Cost Function (minDCF) of 0.6948, with a parameter reduction of approximately 69.4% compared to the ECAPA-TDNN baseline. The model also exhibits strong cross-lingual robustness on the VoxCeleb1 evaluation set. Complexity analysis further confirms that the model offers a favorable trade-off between performance and computational efficiency, showing excellent potential for real-world deployment.

    参考文献
    相似文献
    引证文献
引用本文

郑展恒,李嘉麒,王健,等. 融合分组建模与多维注意力机制的轻量级说话人确认方法[J]. 科学技术与工程, 2026, 26(27): 11804-11811.
Zheng Zhanheng, Li Jiaqi, Wang Jian, et al.[ ]A Lightweight Speech Enhancement Method Based on Grouped Dual-path LSTM[J]. Science Technology and Engineering,2026,26(27):11804-11811.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-26
  • 最后修改日期:2026-07-16
  • 录用日期:2026-03-25
  • 在线发布日期: 2026-09-30
  • 出版日期:
×
喜报|《科学技术与工程》3篇论文成功入选 “第二十八届中国科协年会影响力提名论文”
2026年会通知 | “技术经济学驱动智能经济生态构建与治理变革”——中国技术经济学会第三十三届学术年会(2026)会议通知暨征文启事(第一轮)
亟待确认版面费归属稿件,敬请作者关注