The Utilization of CAMeL Tools in the Part-of-Speech Analysis of Arabic in the Short Story Abdullah wa al-Ushfur
DOI:
https://doi.org/10.61320/jolcc.v4i1.278-296Abstract
This study aims to analyze Arabic word categories, namely ism, fi'il, and harf, in the short story Abdullah wa al-Ushfur using CAMeL Tools. The study employed a qualitative descriptive approach assisted by Natural Language Processing (NLP) technology. The data consisted of the text of Abdullah wa al-Ushfur, obtained from the muthala'ah materials of the Kulliyatul Mu'allimin al-Islamiyah (KMI) second grade. The analysis was conducted using the Part-of-Speech (POS) Tagging feature in CAMeL Tools to automatically identify word classes based on their grammatical functions. The findings revealed that CAMeL Tools successfully identified various word classes, which were subsequently classified into the three primary categories of Arabic grammar: ism, fi'il, and harf. The category of ism was frequently used to refer to characters, objects, places, and conditions within the story, while fi'il contributed to the development of the narrative through the characters' actions. Meanwhile, harf functioned as a linking element that maintained the coherence of sentence structures. The findings suggest that CAMeL Tools can support Arabic linguistic analysis in a more systematic and efficient manner. This study is expected to contribute to technology-based Arabic linguistic studies and serve as a reference for future research integrating NLP into the analysis of Arabic literary texts.
References
Abdelali, A., Darwish, K., Durrani, N., & Mubarak, H. (2016). Farasa: A fast and furious segmenter for Arabic. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations (pp. 11–16).
Alfaidi, A., Alwadei, H., Alshutayri, A., & Alahdal, S. (2023). Exploring the performance of Farasa and CAMeL taggers for Arabic dialect tweets. International Arab Journal of Information Technology, 20(3), 349–356.
Alkuhlani, S., Habash, N., & Roth, R. (2023). Advances in Arabic morphological analysis and disambiguation: A review. ACM Computing Surveys, 56(4), 1–35.
Alluhaibi, R. (2021). A comparative study of Arabic part of speech taggers using literary text samples from Saudi novels. Information, 12(12), 523. https://doi.org/10.3390/info12120523
Alrayani, B., Kalkatawi, M., Abulkhair, M., & Abukhodair, F. (2024). From customer's voice to decision-maker insights: Textual analysis framework for Arabic reviews of Saudi Arabia's super app. Applied Sciences, 14(16), 6952. https://doi.org/10.3390/app14166952
Darwish, K., Habash, N., Abbas, M., Al-Khalifa, H. S., Al-Natsheh, H. T., El-Beltagy, S. R., Bouamor, H., Bouzoubaa, K., Cavalli-Sforza, V., El-Hajj, W., Jarrar, M., & Mubarak, H. (2021). A panoramic survey of natural language processing in the Arab world. Communications of the ACM, 64(12), 72–81.
Elshabrawy, A., AbuOdeh, M., Inoue, G., & Habash, N. (2023). CamelParser 2.0: A state-of-the-art dependency parser for Arabic. In Proceedings of the First Arabic Natural Language Processing Conference (pp. 170–180).
Eryani, F., Khalifa, S., & Habash, N. (2024). Recent trends in Arabic natural language processing resources and tools. In Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 1–10).
Habash, N. (2010). Introduction to Arabic Natural Language Processing. Morgan & Claypool Publishers.
Habash, N., Diab, M., & Rambow, O. (2012). Conventional orthography for dialectal Arabic. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC 2012) (pp. 711–718).
Habash, N., Eskander, R., & Hawwari, A. (2019). A morphological analyzer for Arabic dialects. Natural Language Engineering, 25(2), 195–225.
Habash, N., Rambow, O., & Roth, R. (2009). MADA+TOKAN: A toolkit for Arabic tokenization, diacritization, morphological disambiguation, POS tagging, stemming and lemmatization. In Proceedings of the 2nd International Conference on Arabic Language Resources and Tools (pp. 102–109).
Hamed, I., Eryani, F., Palfreyman, D., & Habash, N. (2024). ZAEBUC-Spoken: A multilingual multidialectal Arabic-English speech corpus. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 17770–17782).
Jurafsky, D., & Martin, J. H. (2023). Speech and Language Processing (3rd ed., draft).
Kallas, O., Inoue, G., & Habash, N. (2024). EMAD: A bridge tagset for unifying Arabic POS annotations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (pp. 5637–5643).
Mubarak, H., Darwish, K., & Magdy, W. (2021). Arabic language processing: Challenges and opportunities. ACM Transactions on Asian and Low-Resource Language Information Processing, 20(5), 1–23.
Obeid, O., Zalmout, N., Khalifa, S., Taji, D., Oudah, M., Alhafni, B., Inoue, G., Eryani, F., Erdmann, A., & Habash, N. (2020). CAMeL Tools: An open source Python toolkit for Arabic natural language processing. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 7022–7032).
Pasha, A., Al-Badrashiny, M., Diab, M. T., El Kholy, A., Eskander, R., Habash, N., Pooleery, M., Rambow, O., & Roth, R. (2014). MADAMIRA: A fast, comprehensive tool for morphological analysis and disambiguation of Arabic. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC 2014) (pp. 1094–1101).
Sleiman, N. M., Hussein, A. A., Kuflik, T., & Minkov, E. (2024). Automatic era identification in classical Arabic poetry. Applied Sciences, 14(18), 8240. https://doi.org/10.3390/app14188240
Zalmout, N., Khalifa, S., Taji, D., Oudah, M., Alhafni, B., Inoue, G., Eryani, F., Erdmann, A., & Habash, N. (2021). Adapting language models for Arabic NLP. In Proceedings of the Arabic Natural Language Processing Workshop (pp. 1–12).
Zeroual, I., Lakhouaja, A., & Belahbib, R. (2017). Towards a standard part-of-speech tagset for the Arabic language. Journal of King Saud University – Computer and Information Sciences, 29(1), 43–50.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Khairina, Rima Mella

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
License and Copyright Agreement
In submitting the manuscript to the journal, the authors certify that:
- Their co-authors authorize them to enter into these arrangements.
- The work described has not been formally published before, except in the form of an abstract or as part of a published lecture, review, thesis, or overlay journal.
- That it is not under consideration for publication elsewhere,
- That its publication has been approved by all the author(s) and by the responsible authorities – tacitly or explicitly – of the institutes where the work has been carried out.
- They secure the right to reproduce any material that has already been published or copyrighted elsewhere.
- They agree to the following license and copyright agreement.
Copyright
Authors who publish in the Journal of Linguistics, Culture, and Communication agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) before and during the submission process, as it can lead to productive exchanges and earlier and greater citation of published work.
Licensing for Data Publication
Journal of Linguistics, Culture, and Communication use a variety of waivers and licenses that are specifically designed for and appropriate for the treatment of data:
- Open Data Commons Attribution License, http://www.opendatacommons.org/licenses/by/1.0/ (default)
- Creative Commons CC-Zero Waiver, http://creativecommons.org/publicdomain/zero/1.0/
- Open Data Commons Public Domain Dedication and Licence, http://www.opendatacommons.org/licenses/pddl/1-0/
Other data publishing licenses may be allowed as exceptions (subject to approval by the editor on a case-by-case basis) and should be justified with a written statement from the author, which will be published with the article.







.jpg)

