Classify and Label the Content from the Cross-Cultural Misogynistic Meme Detection Using Zero-shot Prompts
DOI:
https://doi.org/10.65563/jeaai.v2i3.109Keywords:
Multimodal Learning, Zero-shot Prompts, Large Language Models (LLMs), Cross-Cultural Detection, Misogynistic Classification, Meme DetectionAbstract
Online misogyny increasingly appears in multimodal formats such as memes. Memes combine text and images to convey humor, sarcasm, and ideology. Misogynistic meaning is often implicit and culturally grounded. A meme interpreted as harmful in one cultural setting may be perceived differently in another. Based on the complex scenario of cross-cultural misogynistic meme detection, we proposed to conduct predictions using large language models (LLMs) through zero-shot prompts in the absence
of training datasets. This enables our method to effectively address the key pain points such as the difficulty in annotating the training dataset, the high cost and long time consumption of manually labeling the training dataset labels. We applied this zero-shot prompts method to the Tamil (India) MDMD dataset. The best model results achieved Macro F1-score and Accuracy of 0.628995 and 0.696629 respectively for Task A: Single-culture prediction. Overall contexts including Indian context, Chinese
context, and Western context, the best model results achieved Macro F1-score and Accuracy of 0.601143 and 0.695693 respectively for Task B: Cross-cultural prediction. This approach of deliberately not using the training dataset for fine-tuning and training large language models (LLMs) or pre-trained models enables our method to have greater flexibility in real-world scenarios. For further
information, the project code and user’s guide is available from https://github.com/WangKongQiang/CC-MMD_Grand_Challenge_ICMI_2026.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Kongqiang Wang, Qingli Tan, Peng Zhang

This work is licensed under a Creative Commons Attribution 4.0 International License.