Amharic Text Chuncker Using Machine Learning
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Amharic is one of the most widely spoken languages in Ethiopia and is commonly used in education, government offices, media and everyday communication. However, the development of Natural Language Processing (NLP) tools for Amharic is still limited mainly because of the lack of large, annotated datasets and effective language processing systems. One key NLP task is text chunking which means grouping words into meaningful parts such as noun phrases and verb phrases. This process is very important for many applications like machine translation, information retrieval,
Question answering and text summarization. In this study an Amharic text chunking system is developed using machine learning techniques to overcome the limitations of previous rule-based and statistical approaches. The main goal is to design and evaluate models that can automatically identify phrase structures in Amharic text focusing on Conditional Random Fields (CRF) and Memory-Based Learning (MBL) methods using different features and context window sizes.
