Identification of Negative Intent in Hindi Social Media Posts Using NLP Techniques
ID:50
View Protection:ATTENDEE
Updated Time:2026-07-27 15:35:23
Hits:13
Online
Abstract
Nowadays, with the rapid growth of social media, the amount of content created by Hindi-speaking users is also increasing continuously. This content also includes negative material such as hate speech, online threats, exclusionary posts, and abusive language, which are becoming more common and difficult to detect because this content is written in the Devanagari script. In this paper, we developed an end-to-end Natural Language Processing (NLP) pipeline to classify Hindi social media posts into negative and non-negative categories. This pipeline includes Unicode normalization, Devanagarispecific noise removal, word-level and character n-gram TF-IDF feature extraction, and lexicon-based scoring techniques. with a deep learning approach employing fine-tuned MuRIL (Multilingual Representations for Indian Languages). Experimental evaluation demonstrates that the proposed MuRIL based model achieves F1-score of 0.922 and AUC-score of 0.967, outperforming classical ML baselines (SVM: F1 = 0.872, LR: F1 = 0.810) by a significant margin. The hybrid architecture combining rule-based lexicon features with transformer embeddings yields both high performance and interoperability. We further provide detailed analysis of feature importance, confusion matrices, ROC and Precision- Recall curves, and a comparative study of all models.
Submission Author
Priti Tiwari
Amity University
Submit Comment