MORPHEMIA: a semi-supervised algorithm for the segmentation of Modern Greek words into morphemes
DOI:
https://doi.org/10.36505/ExLing-2008/02/0030/000089Keywords:
machine-learning, morphologically rich languages, statistical language modeling, morphological lexiconsAbstract
The present paper reports on MORPHEMIA, a semi-supervised machine-learning algorithm designed to segment Modern Greek (MG) words into morphemes. The algorithm segments its input iteratively. During its first iteration, the algorithm uses its a priori linguistic knowledge. At the end of each successful iteration, the algorithm extracts new morphological knowledge which is utilised during its next iteration. Thus, with each successful iteration, the algorithm segments an increasing amount of its input data. The algorithm uses a metric to decide whether a given extracted piece of morphological knowledge will improve its performance and only accepts it if it will. Thus, its output gradually improves in quality. MORPHEMIA terminates its operation when new knowledge can no longer be extracted from its input data.
References
Kirchhoff, K. and Sarikaya, R. 2007. Processing morphologically rich languages. Research Tutorial at the 8th Annual Conference of the International Speech Communication Association, Interspeech 2007. Antwerp, Belgium.
Siivola, V., Creutz, M. and Kurimo, M. 2007. Morfessor and VariKN machine learning tools for speech and language Technology. Proceedings of the 8th Annual Conference of the International Speech Communication Association, Interspeech 2007. Antwerp, Belgium.
Downloads
Published
Issue
Section
License
Copyright (c) 2008 Constandinos Kalimeris (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Articles are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are properly credited.