PENS: a confidence parameter estimating the number of speakers
DOI:
https://doi.org/10.36505/ExLing-2008/02/0045/000104Abstract
Is it possible to know how many speakers are speaking simultaneously in a case of speech overlap? While the human brain—a creation not yet mastered—manages to do this and even to understand the meaning of the mixed speech, it is not yet the case for existing automatic systems. For this task, we propose a new method able to estimate the number of speakers in a mixture of speech signals. The algorithm developed here is based on the computation of the statistical characteristics of the seventh Mel coefficient extracted by spectral analysis from the speech signal. This algorithm, which uses a confidence parameter that we called PENS, is tested on seven different sets of the ORATOR database, each containing seven multi-speaker files. Results show that the PENS parameter permits a clear discrimination, without any ambiguity, between a mono-speaker signal (only one speaker is speaking) and a mixed-speaker signal (several speakers are speaking simultaneously). Moreover, in the case of mixed speech signals, it permits an estimation of the number of speakers with good precision, especially when the number of speakers is less than four.
References
Arai, T. 2003. Estimating Number of Speakers by the Modulation Characteristics of Speech. ICASSP, 197-200.
Asano, F., Yamamoto, K., Ogata, J., Yamada, M. and Nakamura, M. 2007. Detection and separation of speech events in meeting recordings using a microphone array. EURASIP Journal on Audio, Speech, and Music, Volume 2007, ID 27616.
Overlapping speech. [http://changingminds.org/techniques/conversation/interrupting/overlap_speech.htm](http://changingminds.org/techniques/conversation/interrupting/overlap_speech.htm)
Lee, H.S. and Tsoi, A.C. 1995. Application of multi-layer perceptron in estimating speech / noise characteristics for speech recognition in noisy environment. Speech Communication, 17, 59-76.
Quast, H. 2002. Automatic recognition of nonverbal speech. An Approach to model the perception of para- and extraling. Vocal Commun. with Neural Net. Mach. Per. Inst. for Neural Comput. UC San Diego, June 28 2002.
Downloads
Published
Issue
Section
License
Copyright (c) 2008 Siham Ouamour (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Articles are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are properly credited.