Amharic Text Chuncker Using Machine Learning

Thumbnail Image

Date

2026

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Amharic is one of the most widely spoken languages in Ethiopia and is commonly used in education, government offices, media and everyday communication. However, the development of Natural Language Processing (NLP) tools for Amharic is still limited mainly because of the lack of large, annotated datasets and effective language processing systems. One key NLP task is text chunking which means grouping words into meaningful parts such as noun phrases and verb phrases. This process is very important for many applications like machine translation, information retrieval, Question answering and text summarization. In this study an Amharic text chunking system is developed using machine learning techniques to overcome the limitations of previous rule-based and statistical approaches. The main goal is to design and evaluate models that can automatically identify phrase structures in Amharic text focusing on Conditional Random Fields (CRF) and Memory-Based Learning (MBL) methods using different features and context window sizes.

Description

Keywords

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By