PENS: a confidence parameter estimating the number of speakers

Authors

  • Siham Ouamour University of Sciences and Technology Houari Boumediene, Algeria Author
  • Mhania Guerti National Polytechnic School of Algiers, Algeria Author
  • Halim Sayoud University of Sciences and Technology Houari Boumediene, Algeria Author

DOI:

https://doi.org/10.36505/ExLing-2008/02/0045/000104

Abstract

Is it possible to know how many speakers are speaking simultaneously in a case of speech overlap? While the human brain—a creation not yet mastered—manages to do this and even to understand the meaning of the mixed speech, it is not yet the case for existing automatic systems. For this task, we propose a new method able to estimate the number of speakers in a mixture of speech signals. The algorithm developed here is based on the computation of the statistical characteristics of the seventh Mel coefficient extracted by spectral analysis from the speech signal. This algorithm, which uses a confidence parameter that we called PENS, is tested on seven different sets of the ORATOR database, each containing seven multi-speaker files. Results show that the PENS parameter permits a clear discrimination, without any ambiguity, between a mono-speaker signal (only one speaker is speaking) and a mixed-speaker signal (several speakers are speaking simultaneously). Moreover, in the case of mixed speech signals, it permits an estimation of the number of speakers with good precision, especially when the number of speakers is less than four.

References

Arai, T. 2003. Estimating Number of Speakers by the Modulation Characteristics of Speech. ICASSP, 197-200.

Asano, F., Yamamoto, K., Ogata, J., Yamada, M. and Nakamura, M. 2007. Detection and separation of speech events in meeting recordings using a microphone array. EURASIP Journal on Audio, Speech, and Music, Volume 2007, ID 27616.

Overlapping speech. [http://changingminds.org/techniques/conversation/interrupting/overlap_speech.htm](http://changingminds.org/techniques/conversation/interrupting/overlap_speech.htm)

Lee, H.S. and Tsoi, A.C. 1995. Application of multi-layer perceptron in estimating speech / noise characteristics for speech recognition in noisy environment. Speech Communication, 17, 59-76.

Quast, H. 2002. Automatic recognition of nonverbal speech. An Approach to model the perception of para- and extraling. Vocal Commun. with Neural Net. Mach. Per. Inst. for Neural Comput. UC San Diego, June 28 2002.

Downloads

Published

01-01-2008

How to Cite

PENS: a confidence parameter estimating the number of speakers. (2008). Linguistic Proceedings Series, 2(1), 177-180. https://doi.org/10.36505/ExLing-2008/02/0045/000104