Abstract:In recent years, deep learning has made significant progress in speaker verification tasks. However, existing models often suffer from large parameter sizes, limited temporal modeling capability, and poor deployability on edge devices. To address these issues, a lightweight speaker verification model incorporating grouped modeling and multi-dimensional attention mechanisms is proposed. The model is built upon an improved Time-Delay Neural Network (TDNN) architecture and introduces a Channel-Split Temporal Encoding Block (CSTE-Block) and a Grouped Scaled Res2Net block (GSRes2Block) to enhance frame-level temporal modeling and multi-scale feature representation. In addition, a Speaker Embedding Refinement Module block (SERMBlock) and multi-dimensional attention are employed to improve global contextual modeling and feature focusing. A global speaker embedding is generated via attentive statistical pooling and jointly optimized using an Additive Angular Margin Softmax classifier and Supervised Contrastive Loss to enhance the discriminative power of the embedding space. Experiments on the CN-Celeb dataset demonstrate that the proposed model achieves an Equal Error Rate (EER) of 14.82% and a minimum Detection Cost Function (minDCF) of 0.6948, with a parameter reduction of approximately 69.4% compared to the ECAPA-TDNN baseline. The model also exhibits strong cross-lingual robustness on the VoxCeleb1 evaluation set. Complexity analysis further confirms that the model offers a favorable trade-off between performance and computational efficiency, showing excellent potential for real-world deployment.