EMAD: A Bridge Tagset for Unifying Arabic POS Annotations

Omar Kallas, Go Inoue, Nizar Habash


Abstract
There have been many attempts to model the morphological richness and complexity of Arabic, leading to numerous Part-of-Speech (POS) tagsets that differ in terms of (a) which morphological features they represent, (b) how they represent them, and (c) the degree of specification of said features. Tagset granularity plays an important role in determining how annotated data can be used and for what applications. Due to the diversity among existing tagsets, many annotated corpora for Arabic cannot be easily combined, which exacerbates the Arabic resource poverty situation. In this work, we propose an intermediate tagset designed to facilitate the conversion and unification of different tagsets used to annotate Arabic corpora. This new tagset acts as a bridge between different annotation schemes, simplifying the integration of annotated corpora and promoting collaboration across the projects using them.
Anthology ID:
2024.lrec-main.500
Volume:
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Month:
May
Year:
2024
Address:
Torino, Italia
Editors:
Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, Nianwen Xue
Venues:
LREC | COLING
SIG:
Publisher:
ELRA and ICCL
Note:
Pages:
5637–5643
Language:
URL:
https://aclanthology.org/2024.lrec-main.500
DOI:
Bibkey:
Cite (ACL):
Omar Kallas, Go Inoue, and Nizar Habash. 2024. EMAD: A Bridge Tagset for Unifying Arabic POS Annotations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 5637–5643, Torino, Italia. ELRA and ICCL.
Cite (Informal):
EMAD: A Bridge Tagset for Unifying Arabic POS Annotations (Kallas et al., LREC-COLING 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.lrec-main.500.pdf