Impact of emoji exclusion on the performance of Arabic sarcasm detection models

Aleryani, Ghalyah H.; Deabes, Wael; Albishre, Khaled; Abdel-Hakim, Alaa E.

Full-text links:

Download:

PDF only

Current browse context:

cs.CL

< prev | next >

new | recent | 2405

Computer Science > Computation and Language

Title: Impact of emoji exclusion on the performance of Arabic sarcasm detection models

Authors: Ghalyah H. Aleryani, Wael Deabes, Khaled Albishre, Alaa E. Abdel-Hakim

(Submitted on 3 May 2024)

Abstract: The complex challenge of detecting sarcasm in Arabic speech on social media is increased by the language diversity and the nature of sarcastic expressions. There is a significant gap in the capability of existing models to effectively interpret sarcasm in Arabic, which mandates the necessity for more sophisticated and precise detection methods. In this paper, we investigate the impact of a fundamental preprocessing component on sarcasm speech detection. While emojis play a crucial role in mitigating the absence effect of body language and facial expressions in modern communication, their impact on automated text analysis, particularly in sarcasm detection, remains underexplored. We investigate the impact of emoji exclusion from datasets on the performance of sarcasm detection models in social media content for Arabic as a vocabulary-super rich language. This investigation includes the adaptation and enhancement of AraBERT pre-training models, specifically by excluding emojis, to improve sarcasm detection capabilities. We use AraBERT pre-training to refine the specified models, demonstrating that the removal of emojis can significantly boost the accuracy of sarcasm detection. This approach facilitates a more refined interpretation of language, eliminating the potential confusion introduced by non-textual elements. The evaluated AraBERT models, through the focused strategy of emoji removal, adeptly navigate the complexities of Arabic sarcasm. This study establishes new benchmarks in Arabic natural language processing and presents valuable insights for social media platforms.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2405.02195 [cs.CL]
	(or arXiv:2405.02195v1 [cs.CL] for this version)

Submission history

From: Ghalyah Aleryani Ghalyah [view email]
[v1] Fri, 3 May 2024 15:51:02 GMT (628kb)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2405.02195

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Computer Science > Computation and Language

Title: Impact of emoji exclusion on the performance of Arabic sarcasm detection models

Submission history