Voice Activity Detector (VAD) based on long-term phonetic features

Authors

  • Andrey Barabanov Saint Petersburg State University, Russia Author
  • Daniil Kocharov Saint Petersburg State University, Russia Author
  • Sergey Salishev Saint Petersburg State University, Russia Author
  • Pavel Skrelin Saint Petersburg State University, Russia Author
  • Mikhail Moiseev Saint Petersburg State University, Russia Author

DOI:

https://doi.org/10.36505/ExLing-2016/07/0005/000264

Keywords:

Voice Activity Detector, classification, decision tree ensemble, auditory masking, phonetic features

Abstract

We propose a VAD using long-term phonetically motivated features with auditory masking, and pre-trained decision tree based classifier, which allows capturing syllable level structure of speech and discriminating it from common noise types. The algorithm demonstrates on test dataset almost 100% acceptance of clear voice for English, Chinese, Russian, and Polish speech and 100% rejection of stationary noises independently of loudness with low computational cost.

 

References

Fant, G. 1960. Acoustic theory of speech production: With calculations based on x-ray studies of Russian articulations.

Fastl, H., Zwicker, E. 2006. Psychoacoustics: facts and models, vol. 22. Springer Science & Business Media.

Zhou, Z.H. 2012. Ensemble methods: foundations and algorithms. CRC Press.

Downloads

Published

01-01-2016

How to Cite

Voice Activity Detector (VAD) based on long-term phonetic features. (2016). Linguistic Proceedings Series, 7(1), 33-36. https://doi.org/10.36505/ExLing-2016/07/0005/000264

Similar Articles

21-30 of 199

You may also start an advanced similarity search for this article.