The intelligibility of noise-vocoded speech:spectral information available from across-channel comparison of amplitude envelopes

Abstract

Noise-vocoded (NV) speech is often regarded as conveying phonetic information primarily through temporal-envelope cues rather than spectral cues. However, listeners may infer the formant frequencies in the vocal-tract output—a key source of phonetic detail—from across-band differences in amplitude when speech is processed through a small number of channels. The potential utility of this spectral information was assessed for NV speech created by filtering sentences into six frequency bands, and using the amplitude envelope of each band (=30 Hz) to modulate a matched noise-band carrier (N). Bands were paired, corresponding to F1 (˜N1 + N2), F2 (˜N3 + N4) and the higher formants (F3' ˜ N5 + N6), such that the frequency contour of each formant was implied by variations in relative amplitude between bands within the corresponding pair. Three-formant analogues (F0 = 150 Hz) of the NV stimuli were synthesized using frame-by-frame reconstruction of the frequency and amplitude of each formant. These analogues were less intelligible than the NV stimuli or analogues created using contours extracted from spectrograms of the original sentences, but more intelligible than when the frequency contours were replaced with constant (mean) values. Across-band comparisons of amplitude envelopes in NV speech can provide phonetically important information about the frequency contours of the underlying formants.

Publication DOI: https://doi.org/10.1098/rspb.2010.1554
Divisions: College of Health & Life Sciences > School of Psychology
College of Health & Life Sciences > Clinical and Systems Neuroscience
College of Health & Life Sciences > School of Optometry > Optometry
College of Health & Life Sciences > School of Optometry > Centre for Vision and Hearing Research
Aston University (General)
Additional Information: © 2010 The Royal Society. The intelligibility of noise-vocoded speech: spectral information available from across-channel comparison of amplitude envelopes. Brian Roberts, Robert J. Summers, Peter J. Bailey. Published 10 November 2010.DOI: 10.1098/rspb.2010.1554
Uncontrolled Keywords: noise-vocoded speech,spectral cues,formant frequencies,intelligibility,General Agricultural and Biological Sciences,General Biochemistry,Genetics and Molecular Biology,General Environmental Science,General Immunology and Microbiology,General Medicine
Publication ISSN: 1471-2954
Last Modified: 16 Dec 2024 08:07
Date Deposited: 19 Apr 2012 09:09
Full Text Link:
Related URLs: http://www.scop ... tnerID=8YFLogxK (Scopus URL)
http://rspb.roy ... t/278/1711/1595 (Publisher URL)
PURE Output Type: Article
Published Date: 2011-05-22
Published Online Date: 2010-11-10
Authors: Roberts, Brian (ORCID Profile 0000-0002-4232-9459)
Summers, Robert J. (ORCID Profile 0000-0003-4857-7354)
Bailey, Peter J.

Download

[img]

Version: Accepted Version

| Preview

Export / Share Citation


Statistics

Additional statistics for this record