《電子技術(shù)應(yīng)用》
您所在的位置:首頁 > 通信與網(wǎng)絡(luò) > 設(shè)計應(yīng)用 > 基于生成對抗網(wǎng)絡(luò)合成噪聲的語音增強方法研究
基于生成對抗網(wǎng)絡(luò)合成噪聲的語音增強方法研究
2020年電子技術(shù)應(yīng)用第11期
夏 鼎,徐文濤
南京航空航天大學(xué) 理學(xué)院,,江蘇 南京211106
摘要: 在語音增強領(lǐng)域,,深度神經(jīng)網(wǎng)絡(luò)通過對大量含有不同噪聲的語音以監(jiān)督學(xué)習(xí)方式進行訓(xùn)練建模,從而提升網(wǎng)絡(luò)的語音增強能力,。然而不同類型噪聲的獲取成本較大,,噪聲類型難以全面采集,影響了模型的泛化能力,。針對這個問題,,提出一種基于生成對抗網(wǎng)絡(luò)(Generative Adversarial Networks,GAN)的噪聲數(shù)據(jù)樣本增強方法,,該方法對真實噪聲數(shù)據(jù)進行學(xué)習(xí),,根據(jù)數(shù)據(jù)特征合成虛擬噪聲,,以此擴充訓(xùn)練集中噪聲數(shù)據(jù)的數(shù)量和類型。通過實驗驗證,,所采用的噪聲合成方法能夠有效擴展訓(xùn)練集中噪聲來源,,增強模型的泛化能力,有效提高語音信號去噪處理后的信噪比和可理解性,。
中圖分類號: TN912.3
文獻標識碼: A
DOI:10.16157/j.issn.0258-7998.200327
中文引用格式: 夏鼎,,徐文濤. 基于生成對抗網(wǎng)絡(luò)合成噪聲的語音增強方法研究[J].電子技術(shù)應(yīng)用,2020,,46(11):56-59,,64.
英文引用格式: Xia Ding,Xu Wentao. Research on speech enhancement method based on generating noise using GAN[J]. Application of Electronic Technique,,2020,,46(11):56-59,64.
Research on speech enhancement method based on generating noise using GAN
Xia Ding,,Xu Wentao
School of Science,,Nanjing University of Aeronautics and Astronautics,Nanjing 211106,,China
Abstract: In the field of speech enhancement, deep neural network can improve the enhancement ability of the model by training and modeling a large number of data with different noises in the supervised learning way. However, the acquisition cost of different types of noise is large and the noise types are difficult to be comprehensive, which affects the generalization ability of the model. Aiming at this problem, this paper proposes a noise data augmentation method based on generative adversarial network(GAN), which learns from the real noise data and synthesizes virtual noises according to the data features, so as to expand the number and type of the noise data in the training set. Experimental results show that the method of noise synthesis adopted in this article can effectively expand the source of noise in the training set, enhance the generalization ability of the model, and effectively improve the signal-to-noise ratio and intelligibility of speech signal after denoising.
Key words : speech enhancement,;generative adversarial network;data augmentation

0 引言

    在語音信號處理的過程中,,背景噪聲和環(huán)境干擾嚴重影響了信號處理的可靠性,,需要通過語音增強處理方法去除信號中的噪聲干擾,改善含噪語音的質(zhì)量,。因此,,語音增強技術(shù)在語音識別、聽力輔助和語音通信等領(lǐng)域中具有非常重要的作用,。

    傳統(tǒng)的語音增強方法有譜減法[1],、維納濾波[2-3]以及之后出現(xiàn)的基于統(tǒng)計模型的處理方法[4]等,這些方法都是基于已知噪聲的統(tǒng)計特性來進行建模,,得到噪聲的功率譜信息,,對含噪語音信號進行降噪處理,以估計純凈語音信號,。這些傳統(tǒng)方法的準確性嚴重依賴數(shù)據(jù)特征工程處理方法和數(shù)據(jù)類型,,對于未知的噪聲干擾,其適應(yīng)能力較差[5],。隨著人工智能的發(fā)展,,深度神經(jīng)網(wǎng)絡(luò)被應(yīng)用于語音增強領(lǐng)域[6]。利用深層神經(jīng)網(wǎng)絡(luò)的特征學(xué)習(xí),,可以將含噪語音映射為純凈語音,,達到去除噪聲的目的,。為了提高深度神經(jīng)網(wǎng)絡(luò)進行語音增強方法的泛化能力,最直接的手段是進行數(shù)據(jù)增強,,包括增加數(shù)據(jù)的多樣性、擴大數(shù)據(jù)集等,。實驗表明,,在深度神經(jīng)網(wǎng)絡(luò)訓(xùn)練的過程中采用更多種類的噪聲數(shù)據(jù),語音信噪比質(zhì)量可以顯著提高[7-8],。但是,,真實的噪聲數(shù)據(jù)獲取難度較大,成本較高,,這限制了網(wǎng)絡(luò)去噪能力的適用性,。針對這一問題,本文基于生成對抗網(wǎng)絡(luò)GAN設(shè)計了一種訓(xùn)練數(shù)據(jù)集增強方法,,通過生成虛擬噪聲,,擴充訓(xùn)練集中噪聲數(shù)據(jù)的類型和數(shù)量,提高模型的泛化能力,。




本文詳細內(nèi)容請下載:http://forexkbc.com/resource/share/2000003050




作者信息:

夏  鼎,,徐文濤

(南京航空航天大學(xué) 理學(xué)院,江蘇 南京211106)

此內(nèi)容為AET網(wǎng)站原創(chuàng),,未經(jīng)授權(quán)禁止轉(zhuǎn)載,。