All systems operational
Q2 2024

Ar-CM-ViMETA: Arabic Image Captioning based on Concept Model and Vision-based Multi-Encoder Transformer Architecture

Asmaa Osman · Mohamed Shalaby · Mona Soliman · Khaled Elsayed
10.34028/iajit/21/3/9 379 Views 6 Citations
6
Citations
379
Views
Abstract

Image captioning is a major artificial intelligence research field that involves visual interpretation and linguistic description of a corresponding image. Successful image captioning relies on acquiring as much information as feasible from the original image. One of these essential bits of knowledge is the topic or the concept that the image is associated with. Recently, concept modeling technique has been utilized in English image captioning for completely capturing the image contexts and make use of these contexts to produce more accurate image descriptions. In this paper, a concept-based model is proposed for Arabic Image Captioning (AIC). A novel Vision-based Multi-Encoder Transformer Architecture (ViMETA) is proposed for handling the multi-outputs result from the concept modeling technique while producing the image caption. BiLingual Evaluation Understudy (BLEU) and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) standard metrics have been used to evaluate the proposed model using the Flickr8K dataset with Arabic captions. Furthermore, qualitative analysis has been conducted to compare the produced captions of the proposed model with the ground truth descriptions. Based on the experimental results, the proposed model outperformed the related works both quantitatively, using BLEU and ROUGE metrics, and qualitatively

Cite this Article (APA)
Asmaa, O., Mohamed, S., Mona, S., Khaled, E. (2024). Ar-CM-ViMETA: Arabic Image Captioning based on Concept Model and Vision-based Multi-Encoder Transformer Architecture. The International Arab Journal of Information Technology. https://doi.org/10.34028/iajit/21/3/9
Related Papers
Perception of Natural Scenes: Objects Detection and Segmentations using Saliency Map with AlexNet
Muhammad Waqas Ahmed; Abdulwahab Alazeb; Naif Al Mudawi; Touseef Sadiq; Bayan Al · 2025
21
cites
390
Agile Proactive Cybercrime Evidence Analysis Model for Digital Forensics
Mohammad Al-Mousa; Waleed Amer; Mosleh Abualhaj; Sultan Albilasi; Ola Nasir; Gha · 2025
16
cites
387
14
cites
397
Heart Disease Diagnosis Using Decision Trees with Feature Selection Method
Alaa Sheta; Walaa El-Ashmawi; Abdelkarim Baareh · 2024
13
cites
383
Access
View Full Text via DOI
Published in
ISSN 1683-3198
Quartile Q2
AMS Score 100
Field Computer Science & AI
Publisher Zarqa University / Colleges of Comp
Country 🇯🇴 Jordan
View Journal Profile →
Authors
Publication Details
Year 2024
Language English
Added 30 Jul 2026