The Utilization of CAMeL Tools in the Part-of-Speech Analysis of Arabic in the Short Story Abdullah wa al-Ushfur

Authors

  • Khairina Nasution Arabic Literature Study Program, Faculty of Cultural Sciences, Universitas Sumatera Utara, Indonesia
  • Rima Mella Arabic Literature Study Program, Faculty of Cultural Sciences, Universitas Sumatera Utara, Indonesia
  • Rozanna Mulyani Arabic Literature Study Program, Faculty of Cultural Sciences, Universitas Sumatera Utara, Indonesia

DOI:

https://doi.org/10.61320/jolcc.v4i1.278-296

Abstract

This study aims to analyze Arabic word categories, namely ism, fi'il, and harf, in the short story Abdullah wa al-Ushfur using CAMeL Tools. The study employed a qualitative descriptive approach assisted by Natural Language Processing (NLP) technology. The data consisted of the text of Abdullah wa al-Ushfur, obtained from the muthala'ah materials of the Kulliyatul Mu'allimin al-Islamiyah (KMI) second grade. The analysis was conducted using the Part-of-Speech (POS) Tagging feature in CAMeL Tools to automatically identify word classes based on their grammatical functions. The findings revealed that CAMeL Tools successfully identified various word classes, which were subsequently classified into the three primary categories of Arabic grammar: ism, fi'il, and harf. The category of ism was frequently used to refer to characters, objects, places, and conditions within the story, while fi'il contributed to the development of the narrative through the characters' actions. Meanwhile, harf functioned as a linking element that maintained the coherence of sentence structures. The findings suggest that CAMeL Tools can support Arabic linguistic analysis in a more systematic and efficient manner. This study is expected to contribute to technology-based Arabic linguistic studies and serve as a reference for future research integrating NLP into the analysis of Arabic literary texts.

References

Abdelali, A., Darwish, K., Durrani, N., & Mubarak, H. (2016). Farasa: A fast and furious segmenter for Arabic. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations (pp. 11–16).

Alfaidi, A., Alwadei, H., Alshutayri, A., & Alahdal, S. (2023). Exploring the performance of Farasa and CAMeL taggers for Arabic dialect tweets. International Arab Journal of Information Technology, 20(3), 349–356.

Alkuhlani, S., Habash, N., & Roth, R. (2023). Advances in Arabic morphological analysis and disambiguation: A review. ACM Computing Surveys, 56(4), 1–35.

Alluhaibi, R. (2021). A comparative study of Arabic part of speech taggers using literary text samples from Saudi novels. Information, 12(12), 523. https://doi.org/10.3390/info12120523

Alrayani, B., Kalkatawi, M., Abulkhair, M., & Abukhodair, F. (2024). From customer's voice to decision-maker insights: Textual analysis framework for Arabic reviews of Saudi Arabia's super app. Applied Sciences, 14(16), 6952. https://doi.org/10.3390/app14166952

Darwish, K., Habash, N., Abbas, M., Al-Khalifa, H. S., Al-Natsheh, H. T., El-Beltagy, S. R., Bouamor, H., Bouzoubaa, K., Cavalli-Sforza, V., El-Hajj, W., Jarrar, M., & Mubarak, H. (2021). A panoramic survey of natural language processing in the Arab world. Communications of the ACM, 64(12), 72–81.

Elshabrawy, A., AbuOdeh, M., Inoue, G., & Habash, N. (2023). CamelParser 2.0: A state-of-the-art dependency parser for Arabic. In Proceedings of the First Arabic Natural Language Processing Conference (pp. 170–180).

Eryani, F., Khalifa, S., & Habash, N. (2024). Recent trends in Arabic natural language processing resources and tools. In Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 1–10).

Habash, N. (2010). Introduction to Arabic Natural Language Processing. Morgan & Claypool Publishers.

Habash, N., Diab, M., & Rambow, O. (2012). Conventional orthography for dialectal Arabic. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012) (pp. 711–718).

Habash, N., Eskander, R., & Hawwari, A. (2019). A morphological analyzer for Arabic dialects. Natural Language Engineering, 25(2), 195–225.

Habash, N., Rambow, O., & Roth, R. (2009). MADA+TOKAN: A toolkit for Arabic tokenization, diacritization, morphological disambiguation, POS tagging, stemming and lemmatization. In Proceedings of the 2nd International Conference on Arabic Language Resources and Tools (pp. 102–109).

Hamed, I., Eryani, F., Palfreyman, D., & Habash, N. (2024). ZAEBUC-Spoken: A multilingual multidialectal Arabic-English speech corpus. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 17770–17782).

Jurafsky, D., & Martin, J. H. (2023). Speech and Language Processing (3rd ed., draft).

Kallas, O., Inoue, G., & Habash, N. (2024). EMAD: A bridge tagset for unifying Arabic POS annotations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 5637–5643).

Mubarak, H., Darwish, K., & Magdy, W. (2021). Arabic language processing: Challenges and opportunities. ACM Transactions on Asian and Low-Resource Language Information Processing, 20(5), 1–23.

Obeid, O., Zalmout, N., Khalifa, S., Taji, D., Oudah, M., Alhafni, B., Inoue, G., Eryani, F., Erdmann, A., & Habash, N. (2020). CAMeL Tools: An open source Python toolkit for Arabic natural language processing. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 7022–7032).

Pasha, A., Al-Badrashiny, M., Diab, M. T., El Kholy, A., Eskander, R., Habash, N., Pooleery, M., Rambow, O., & Roth, R. (2014). MADAMIRA: A fast, comprehensive tool for morphological analysis and disambiguation of Arabic. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC 2014) (pp. 1094–1101).

Sleiman, N. M., Hussein, A. A., Kuflik, T., & Minkov, E. (2024). Automatic era identification in classical Arabic poetry. Applied Sciences, 14(18), 8240. https://doi.org/10.3390/app14188240

Zalmout, N., Khalifa, S., Taji, D., Oudah, M., Alhafni, B., Inoue, G., Eryani, F., Erdmann, A., & Habash, N. (2021). Adapting language models for Arabic NLP. In Proceedings of the Arabic Natural Language Processing Workshop (pp. 1–12).

Zeroual, I., Lakhouaja, A., & Belahbib, R. (2017). Towards a standard part-of-speech tagset for the Arabic language. Journal of King Saud University – Computer and Information Sciences, 29(1), 43–50.

Downloads

Published

2026-07-23

How to Cite

Nasution, K., Mella, R., & Mulyani , R. (2026). The Utilization of CAMeL Tools in the Part-of-Speech Analysis of Arabic in the Short Story Abdullah wa al-Ushfur. Journal of Linguistics, Culture and Communication, 4(1), 278–296. https://doi.org/10.61320/jolcc.v4i1.278-296