基于分组双路径LSTM的轻量级语音增强方法
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TN912.35

基金项目:

国家自然科学基金(62064003);认知无线电与信息处理教育部重点实验室项目(CRKL230103)


A lightweight speech enhancement method based on grouped dual-path LSTMZHENG Zhan-heng1,2,3, XU Jia-yu1,2*, WANG Jian 1,2, LI Qi1
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    近年来,深度学习方法在语音增强领域取得显著进展,但现有模型普遍存在参数冗余、推理开销大的问题,难以满足边缘计算设备对低资源占用与实时性的要求。针对这一问题,提出一种高效轻量化的语音增强框架。该模型前端采用频带合并(band merging, BM)模块,压缩高频冗余信息;引入分组融合卷积(group fusion convolution, GF-Conv),结合扩张深度卷积与时频混合扩张残差注意力模块(time-frequency hybrid dilated residual attention block, TF-HDRA)提升多尺度建模能力;同时,通过通道聚合(channel aggregation, CA)模块增强跨通道特征交互。增强层采用分组双路径长短期记忆(grouped dual-path LSTM, GDPLSTM)模块,结合通道分组与双路径机制,在保持出色时频建模能力的同时,实现模型结构的轻量化,有效降低参数量与推理开销。实验在VoiceBank+DEMAND数据集上进行,结果表明,所提模型语音感知质量为2.99,参数量为51.85 K,计算量为145.23 MMac,实现了性能与计算效率的良好平衡。

    Abstract:

    In recent years, deep learning methods have made significant progress in the field of speech enhancement, but the existing models generally have the problems of redundant parameters and high inference overhead, which are difficult to meet the requirements of edge computing devices for low resource occupation and real-time performance. In order to solve this problem, an efficient and lightweight speech enhancement framework was proposed. The front-end of the model uses the band merging to compress high-frequency redundant information. Group fusion convolution was introduced, and the dilated deep convolution and time-frequency hybrid dilated residual attention module were combined to improve the multi-scale modeling ability. At the same time, the cross-channel feature interaction is enhanced by the channel aggregation module. The group dual-path long short-term memory module is used in the enhancement layer, which combines channel grouping and dual-path mechanism, which not only maintains excellent time-frequency modeling capabilities, but also realizes the lightweight of the model structure, and effectively reduces the number of parameters and inference overhead. Experiments are carried out on the VoiceBank DEMAND dataset, and the results show that the voice perception quality of the proposed model is 2.99, the parameter quantity is 51.85 K, and the computational amount is 145.23 MMac, which achieves a good balance between performance and computing efficiency.

    参考文献
    相似文献
    引证文献
引用本文

郑展恒,徐佳瑜,王健,等. 基于分组双路径LSTM的轻量级语音增强方法[J]. 科学技术与工程, 2026, 26(20): 8712-8719.
Zheng Zhanheng, Xu Jiayu, Wang Jian, et al. A lightweight speech enhancement method based on grouped dual-path LSTMZHENG Zhan-heng1,2,3, XU Jia-yu1,2*, WANG Jian 1,2, LI Qi1[J]. Science Technology and Engineering,2026,26(20):8712-8719.

复制
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-08-05
  • 最后修改日期:2026-04-16
  • 录用日期:2026-01-31
  • 在线发布日期: 2026-07-27
  • 出版日期:
×
2026年会通知 | “技术经济学驱动智能经济生态构建与治理变革”——中国技术经济学会第三十三届学术年会(2026)会议通知暨征文启事(第一轮)
亟待确认版面费归属稿件,敬请作者关注