Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Thesis . 2026
License: CC BY
Data sources: Datacite
ZENODO
Thesis . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

Emojis as Affective Signals for Sarcasm Detection: An Empirical Analysis Informing EmoCentricSarcBERT and Its Application in the RADMAD Framework for Toxicity Mitigation in Online Discourse

Authors: Grover, Vandita;

Emojis as Affective Signals for Sarcasm Detection: An Empirical Analysis Informing EmoCentricSarcBERT and Its Application in the RADMAD Framework for Toxicity Mitigation in Online Discourse

Abstract

Abstract The absence of paralinguistic cues in online discourse creates a contextual void, complicating classification of communicative intent such as sarcasm. Given the prevalent use of sarcasm to veil online toxicity, robust detection mechanisms are critical for effective content moderation. This thesis presents an integrated pipeline for an emoji-centric approach to sarcasm detection, transitioning from foundational statistical validation to the deployment of a resource-constrained moderation framework. Initial hypothesis testing quantitatively validates that over 20% of online discourse utilizes emojis, establishing their viability as computational signals in text analytics. To address the inherent scarcity of emoji-rich sarcastic data, a novel dataset, SarcOji, is curated alongside two test sets: SarcOjiTest1 from existing benchmarks and SarcOjiTest2 from X (Twitter). Leveraging this data, novel Sarcasm-aware emoji embeddings are engineered to capture context from surrounding text based on word-emoji co-occurrences. Empirical evaluations isolate three primary drivers required for emoji-centric sarcasm classification: explicit modeling of user intent via their most frequent emoji (MaxEmoji), contextual grounding through the proposed embeddings, and the integration of attention mechanisms. These drivers are subsequently unified within the EmoCentricSarcBERT architecture, which successfully captures deep bidirectional relationships between text and emojis to achieve robust F1, MCC, and ROC-AUC scores of 65.54%, 0.33, and 71.3% for SarcOjiTest1, and 53.33%, 0.35, and 74.18% for SarcOjiTest2, respectively. This empirical success motivates the real-world deployment of EmoCentricSarcBERT, where knowledge distillation and quantization systematically compress this model into highly efficient low-resource architectures. This optimization culminates in the Resource-Aware Decentralized Moderation and Deployment (RADMAD) framework for sarcastic toxicity. Designed for low-resource edge devices, RADMAD functions as a proactive, nudge-based moderation system, dynamically calculating sarcasm scores to mitigate this toxicity at the source. This deployable pipeline for real-time content moderation on resource-constrained devices yields ten core contributions: three datasets, Sarcasm-aware embeddings, EmoCentricSarcBERT, four compressed models, and the RADMAD framework.

How to Cite This Thesis To reference the extended methodology, or overall compilation, please cite this document: Grover, V. (2026). Emojis as Affective Signals for Sarcasm Detection: An Empirical Analysis Informing EmoCentricSarcBERT and Its Application in the RADMAD Framework for Toxicity Mitigation in Online Discourse [Thesis]. Zenodo. https://doi.org/10.5281/zenodo.21297666 Peer-Reviewed Foundations This independent research compilation was developed utilizing methodologies and findings published in peer-reviewed journals. If you utilize the SarcOji dataset or the EmoCentricSarcBERT architecture in your own research, please additionally cite the underlying peer-reviewed papers: Grover, V. & Banati, H., 2022. DUCS at SemEval-2022 Task 6: Exploring Emojis and Sentiments for Sarcasm Detection. Seattle, Association for Computational Linguistics, pp. 1005-1011. Grover, V. & Banati, H., 2022. Understanding the sarcastic nature of emojis with SarcOji. Seattle, Association for Computational Linguistics, pp. 29--39. Grover, V. & Banati, H., 2023. EmoRile: a personalised emoji prediction scheme based on user profiling. International Journal of Business Intelligence and Data Mining, 22(4), pp. 470-485. Grover, V. & Banati, H., 2024. An attention approach to emoji focused sarcasm detection. Heliyon, 10(17), p. e36398. Grover, V. & Banati, H., 2026. An emoji centric approach to sarcasm detection in online discourse. Scientific Reports, 16(1), p. 3891..

Open Source Assets & Repositories Training Dataset (SarcOji): https://github.com/VanditaGroverKapila/SarcOji Test Datasets: https://github.com/VanditaGroverKapila/SarcOjiTestSets Embeddings: https://github.com/VanditaGroverKapila/SarcasmAwareEmojiEmbeddings Base Transformer Model: EmoCentricSarcBERT: https://huggingface.co/Vandita/Bert-finetuned-Sarc Compressed Models: DistilledEmoCentricSarcBERT: https://huggingface.co/Vandita/DistilledEmoCentricSarcBERT TinyEmoCentricSarcBERT: https://huggingface.co/Vandita/TinyEmoCentricSarcBERT MobileEmoCentricSarcBERT: https://huggingface.co/Vandita/MobileEmoCentricSarcBERT QuantizedEmoCentricSarcBERT: https://huggingface.co/Vandita/QuantizedEmoCentricSarcBERT The RADMAD Framework: Detailed in Chapter 6 of the thesis.

Keywords

emojis, hate-speech detection, text analytics, Sentiment Analysis, emotion detection, sarcasm classification, sarcasm detection, emojis in sarcasm detection, Sentiment Analysis/classification

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!