English  |  正體中文  |  简体中文  |  Items with full text/Total items : 65317/65317 (100%)
Visitors : 21336340      Online Users : 698
RC Version 7.0 © Powered By DSPACE, MIT. Enhanced by NTU Library IR team.
Scope Tips:
  • please add "double quotation mark" for query phrases to get precise results
  • please goto advance search for comprehansive author search
  • Adv. Search
    HomeLoginUploadHelpAboutAdminister Goto mobile version


    Please use this identifier to cite or link to this item: http://ir.lib.ncu.edu.tw/handle/987654321/8920


    Title: 國語語音強健辨認之研究;Robust speech recognition in noisy environments
    Authors: 黃國彰;Kuo-Chang Huang
    Contributors: 電機工程研究所
    Keywords: 強健特徵參數;模型補償;robust features;model compensation
    Date: 2003-05-30
    Issue Date: 2009-09-22 11:37:41 (UTC+8)
    Publisher: 國立中央大學圖書館
    Abstract: Despite sophisticated present day automatic speech recognition (ASR) techniques, a single recognizer is usually incapable of accounting for the varying conditions in a typical natural environment. Higher robustness to a range of noise cases can potentially be achieved by combining the results of several recognizers operating in parallel. To overcome this problem and improve the performance of speech recognition systems in additive conditions, special attention should be paid to the problem of robust feature and compensation of models. This thesis is concerned with the problem of noise-resistance applied to automatic speaker-independent speech recognition. The two problems of the model compensation and robust feature are treated in this work. In model compensation stage, first, we investigate a projection-based group delay scheme (PGDS) likelihood measure that significantly reduces noise contamination in speech recognition. Because the norm of the cepstral/GDS vector will be shrinked when the speech signals are corrupted by additive noise, the HMM parameters, namely, the mean vector and the covariance matrix, need to be furthermore modified. The proposed approach compensates the mean vector using a projection-based scale factor and the mean compensation bias, and fits the covariance matrix using a variance adaptive function. The bias and variance adaptive functions estimated from the training and/or testing data were used to balance the mismatch between different environments. Lastly, a state duration method was utilized to deal with the problem that the additive noise segments the error path in Viterbi decoding. Secondly, we proposed a model compensation method that is similar to parallel model combination. The basis of the method is the fact that the autocorrelation function of the signal resulting from the addition of two statistically independent signals is equal to the sum of their individual autocorrelation functions. Therefore, in adjusting a clean model, its state spectral representation is transformed from the autoregressive, or cepstral, domain to the autocorrelation domain. Then, the autocorrelation of the clean model is added to a sample of the autocorrelation of the additive noise, resulting in the autocorrelation of the noisy signal, which is transformed back to the original spectral representation. At the end of this process, an adjusted model results with better capabilities of handling the noisy signal. Most speech recognition systems are based on cepstral coefficients and their first- and second order derivatives. The derivatives are normally approximated by fitting a linear regression line to a fixed-length segment of consecutive frames. The time resolution and smoothness of the estimated derivative depends on the length of the segment. Herein, we present an approach to improve the representation of speech dynamics, which is based on the combination of multiple time resolutions. To illustrate the procedure, we take two different sets of feature combinations. In the first system, we combine separated input used different features, i.e. the cepstral and group delay spectrum coefficients leading to higher performance in all noise condition. In the second system, we extract feature over variable sized windows of three or five times the original windows size. Capturing different information in different feature combination or in multi-scale features being more robust to noise, the robust integration system gained a significant performance improvement in both clean speech and in real environmental noise. Despite sophisticated present day automatic speech recognition (ASR) techniques, a single recognizer is usually incapable of accounting for the varying conditions in a typical natural environment. Higher robustness to a range of noise cases can potentially be achieved by combining the results of several recognizers operating in parallel. To overcome this problem and improve the performance of speech recognition systems in additive conditions, special attention should be paid to the problem of robust feature and compensation of models. This thesis is concerned with the problem of noise-resistance applied to automatic speaker-independent speech recognition. The two problems of the model compensation and robust feature are treated in this work. In model compensation stage, first, we investigate a projection-based group delay scheme (PGDS) likelihood measure that significantly reduces noise contamination in speech recognition. Because the norm of the cepstral/GDS vector will be shrinked when the speech signals are corrupted by additive noise, the HMM parameters, namely, the mean vector and the covariance matrix, need to be furthermore modified. The proposed approach compensates the mean vector using a projection-based scale factor and the mean compensation bias, and fits the covariance matrix using a variance adaptive function. The bias and variance adaptive functions estimated from the training and/or testing data were used to balance the mismatch between different environments. Lastly, a state duration method was utilized to deal with the problem that the additive noise segments the error path in Viterbi decoding. Secondly, we proposed a model compensation method that is similar to parallel model combination. The basis of the method is the fact that the autocorrelation function of the signal resulting from the addition of two statistically independent signals is equal to the sum of their individual autocorrelation functions. Therefore, in adjusting a clean model, its state spectral representation is transformed from the autoregressive, or cepstral, domain to the autocorrelation domain. Then, the autocorrelation of the clean model is added to a sample of the autocorrelation of the additive noise, resulting in the autocorrelation of the noisy signal, which is transformed back to the original spectral representation. At the end of this process, an adjusted model results with better capabilities of handling the noisy signal. Most speech recognition systems are based on cepstral coefficients and their first- and second order derivatives. The derivatives are normally approximated by fitting a linear regression line to a fixed-length segment of consecutive frames. The time resolution and smoothness of the estimated derivative depends on the length of the segment. Herein, we present an approach to improve the representation of speech dynamics, which is based on the combination of multiple time resolutions. To illustrate the procedure, we take two different sets of feature combinations. In the first system, we combine separated input used different features, i.e. the cepstral and group delay spectrum coefficients leading to higher performance in all noise condition. In the second system, we extract feature over variable sized windows of three or five times the original windows size. Capturing different information in different feature combination or in multi-scale features being more robust to noise, the robust integration system gained a significant performance improvement in both clean speech and in real environmental noise.
    Appears in Collections:[電機工程研究所] 博碩士論文

    Files in This Item:

    File SizeFormat
    0KbUnknown512View/Open


    All items in NCUIR are protected by copyright, with all rights reserved.

    社群 sharing

    ::: Copyright National Central University. | 國立中央大學圖書館版權所有 | 收藏本站 | 設為首頁 | 最佳瀏覽畫面: 1024*768 | 建站日期:8-24-2009 :::
    DSpace Software Copyright © 2002-2004  MIT &  Hewlett-Packard  /   Enhanced by   NTU Library IR team Copyright ©   - Feedback  - 隱私權政策聲明