📣

Follow ACM MAJU on Instagram, Facebook, LinkedIn, YouTube, and TikTok.

Ad space ad

Sponsored

Final Year Project

ASR - Enabled Roman Urdu Hate Speech Detection (ARUHSD)

Team

MA Maria Asad AS Asmar Sohail AL Almas Suleman
Natural Language Processing Deep Learning Automated Speech Recognition Roman Urdu Hate Speech Detection Hate Speech Classification Offensive Speech Detection Speech-to-Text Text Classification Multilingual NLP Custom Dataset Low-Resource Language Content Moderation

Abstract

Natural Language Processing (NLP) and Speech Recognition technologies are rapidly advancing, enabling intelligent systems capable of understanding and analyzing spoken language. However, detecting hate and offensive speech in multilingual regions such as Pakistan remains a significant challenge, particularly in informal and widely used variants of Roman Urdu. During our research, we observed that existing Roman Urdu textual datasets within the Pakistani context are largely inadequate, unstructured, and inconsistent. Moreover,no authentic or no publicly available Roman Urdu speech datasets could be identified, highlighting a critical lack of multimodal resources for this low-resource language. As a result, efforts to develop robust ASR-enabled hate speech detection systems for Roman Urdu have been severely constrained. To address these limitations, firstly developed a customized multimodal Roman Urdu dataset comprising over 10,000 text and corresponding audio samples. The dataset was manually curated using standard sentences that were purposely constructed, along with audio recordings collected from speakers of different ages, genders, and accents to ensure diversity and improve model robustness. This process was developed to create an ASR-Enabled Roman Urdu Hate Speech Detection (ARUHSD) system. The proposed system performs automatic speech recognition to transcribe Urdu speech into Roman Urdu text. Subsequently, it applies a deep learning–based NLP model to classify the transcribed content into these four categories: Threatening, Anti-National. Neutral and Abusive Experimental results demonstrate that the proposed system achieves efficient overall classification accuracy and strong prior approaches in consistently detecting offensive and hateful speech. These findings of the experiment indicate that ARUHSD provides a scalable and efficient solution for multimodal hate speech detection, contributing toward safer, more responsible, and inclusive communication within the Urdu-speaking community.
Ad space ad

Sponsored

Ad space ad

Sponsored

Ad space ad

Sponsored